A Primer on Post-Training Reasoning Data
Yaoming Li, Guangxiang Zhao, Qilong Shi, et al.
First primer synthesizing over 150 studies on post-training reasoning data, organizing the field around four key questions.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Yaoming Li, Guangxiang Zhao, Qilong Shi, et al.
First primer synthesizing over 150 studies on post-training reasoning data, organizing the field around four key questions.
Tanishq Kumar, Zachary Ankner, B. Spector, et al.
Proposes precision-aware scaling laws for language models, showing low precision reduces effective parameter count and post-training quantization can harm models trained on more data.
A. Sheshadri, John Hughes, Julian Michael, et al.
This paper analyzes 25 language models to understand why only 5 exhibit alignment faking, finding that post-training variations in refusal behavior largely explain differences.
Aayush Karan, Yilun Du
Proposes a simple iterative sampling algorithm that elicits reasoning from base LLMs at inference time, matching or outperforming RL post-training on single-shot tasks without additional training or verifiers.
Unknown
Eagle 2 introduces a data-centric post-training strategy and a tiled mixture of vision encoders to build high-performing vision-language models.
Unknown
Gemma 3 is a multimodal language model with vision understanding, 128k context, and improved performance via distillation and a novel post-training recipe.
Unknown
Phi-4 is a 14B language model that achieves strong reasoning performance, especially in STEM, by prioritizing data quality through synthetic data, curated organic seeds, and innovative post-training techniques.
Unknown
Tulu V3 presents a fully open post-training recipe for Llama 3.1 models, achieving state-of-the-art performance via SFT, DPO, and RLVR.
Shuo Yang, Ying Sheng, Joseph E. Gonzalez, et al.
Double Sparsity reduces KV cache access in LLMs via post-training sparse attention combining token and channel sparsity.
Jinuk Kim, Marwa El Halabi, Wonpyo Park, et al.
GuidedQuant improves post-training quantization of large language models by using end-loss guidance to better preserve model accuracy.
Unknown
Proposes post-training world models via reinforcement learning to improve their generality and task performance.
Dongfang Li, Xiaodong Luo, Ruoyu Sun, et al.
Full-stack optimization for post-training trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and domain-specialized OR models outperforming GPT-5.4-Mini.