Strategic Chain-of-Thought
Yu Wang, Shiwan Zhao, Zhihu Wang, et al.
SCoT improves LLM reasoning by first eliciting a problem-solving strategy before generating Chain-of-Thought steps, achieving significant gains on reasoning benchmarks.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Yu Wang, Shiwan Zhao, Zhihu Wang, et al.
SCoT improves LLM reasoning by first eliciting a problem-solving strategy before generating Chain-of-Thought steps, achieving significant gains on reasoning benchmarks.
Annie Wong, Thomas H. W. Back, A. Plaat, et al.
Evaluates prompting strategies in dynamic environments, finding strategic prompting can close performance gaps but reveals persistent reasoning limitations in LLMs.
Chengshuai Zhao, Zhen Tan, Pingchuan Ma, et al.
Proposes a data distribution lens to understand when and why Chain-of-Thought reasoning succeeds or fails, revealing it as a brittle mirage beyond training distributions.
Kaiwen Wei, Rui Shan, Dongsheng Zou, et al.
MIRAGE enhances test-time scaling for medical QA by combining multi-path parallel inference with structured knowledge graph retrieval to reduce error accumulation and improve traceability.
Jundong Xu, Hao Fei, Liangming Pan, et al.
Proposes SymbCoT, a fully LLM-based framework integrating symbolic expressions and logic rules with Chain-of-Thought prompting to enhance logical reasoning.
J. Kirchner, Yining Chen, Harri Edwards, et al.
Proposes legibility training via a Prover-Verifier Game to make LLM chain-of-thought reasoning easier for humans to verify, improving trust in model outputs.
Amirmohammad Izadi, Mohammadali Banayeeanzade, Fatemeh Askari, et al.
VISER augments visual inputs with low-level spatial structures to improve LVLM visual reasoning, achieving gains of 25% in visual search, 26.8% in counting, and 9.5% in spatial relationships.
Adrian de Wynter
This paper argues that in-context learning fits the mathematical definition of learning but empirically shows limited generalization to unseen tasks, with accuracy insensitive to exemplar distribution and prompt style.
Unknown
PIR optimizes chain-of-thought data by pruning low-importance reasoning steps, improving accuracy and reducing token usage.
Unknown
ThinkPO uses DPO on existing short and long CoT data to improve LLM reasoning without costly new long CoT data.
Unknown
Proposes Direct Judgement Preference Optimization to enhance LLM judges via Chain-of-Thought Critique, Standard Judgement, and Response Deduction for rating, comparison, and classification.
Unknown
NuminaMath introduces a massive 860k problem-solution dataset with chain-of-thought traces to boost LLM mathematical reasoning, winning the 1st AIMO Progress Prize.