Strategic Chain-of-Thought
Yu Wang, Shiwan Zhao, Zhihu Wang, et al.
SCoT improves LLM reasoning by first eliciting a problem-solving strategy before generating Chain-of-Thought steps, achieving significant gains on reasoning benchmarks.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Yu Wang, Shiwan Zhao, Zhihu Wang, et al.
SCoT improves LLM reasoning by first eliciting a problem-solving strategy before generating Chain-of-Thought steps, achieving significant gains on reasoning benchmarks.
Paul Bergmann, Kilian Batzner, Michael Fauser, et al.
Introduces MVTec AD, a comprehensive dataset with 5354 high-resolution images across 15 categories for unsupervised anomaly detection, and benchmarks state-of-the-art methods.
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, et al.
DeepLab advances semantic segmentation via atrous convolution, ASPP, and fully connected CRFs, achieving state-of-the-art on multiple benchmarks.
Farieda Gaber, Maqsood Shaik, Fabio Allega, et al.
Benchmarks LLMs and a RAG workflow on 2000 MIMIC-derived medical cases for triage, referral, and diagnosis support.
Nourhan Ibrahim, Samar AboulEla, Ahmed Ibrahim, et al.
This survey classifies LLM-KG integration into three paradigms—KG-augmented LLMs, LLM-augmented KGs, and synergized frameworks—and evaluates their methodologies, metrics, benchmarks, and challenges.
Zheyang Xiong, Vasileios Papageorgiou, Kangwook Lee, et al.
Proposes finetuning LLMs on synthetic key-value retrieval data to improve long-context retrieval and reasoning without harming general benchmarks.
Yibo Yan, Jiamin Su, Jianxiang He, et al.
First comprehensive survey of mathematical reasoning in multimodal LLMs, reviewing over 200 studies since 2021 across benchmarks, methodologies, and challenges.
Shen Nie, Fengqi Zhu, Zebin You, et al.
LLaDA is a diffusion model trained from scratch for language modeling that matches autoregressive LLMs like LLaMA3 8B across benchmarks and solves the reversal curse.
Yifei Zhou, Sergey Levine, J. Weston, et al.
Self-Challenging framework enables LLM agents to self-generate high-quality training tasks via Code-as-Task, achieving over two-fold improvement on tool-use benchmarks without human annotation.
Thang Luong, Dawsen Hwang, Hoang Nguyen, et al.
Introduces IMO-Bench, a suite of Olympiad-level benchmarks for robust mathematical reasoning, enabling gold-level IMO performance via Gemini Deep Think.
Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, et al.
First comprehensive overview of metacognition in LLMs, taxonomizing methods, benchmarks, and techniques to measure and improve metacognitive abilities.
Unknown
CHASE is a framework for synthetically generating challenging LLM evaluation benchmarks by building complex problems from simpler components and hiding solution elements within context.