Journal Article
Large Language Models
Featured

DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning

Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Zhenhua Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bowen Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai(Individual Differences), Fuli Luo(Individual Differences), Guangbo Hao, Guan-Ting Chen, Guowei Li, Hongjun Zhang, Hanwei Xu, Honghui Ding, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jingchang Chen, Jingyang Yuan(Shanghai Jinyuan Senior High School), Jiagang Tu(Shanghai Jinyuan Senior High School), Junjie Qiu, Junlong Li, Jiali Cai, Jiaqi Ni, Jian Liang, Jing Chen, Kai Dong(University of Science and Technology of China), Kai Hu(University of Science and Technology of China), Kaichao You, Kaige Gao, Kang Guan(Peking University), Kexin Huang(Peking University), Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, L. Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, Rong Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Sansan Li, Shuang Zhou, Shaoqing Wu, Tao Yun
September 17, 2025Nature5,469 citations

5.5k

Citations

932

Influential Citations

Nature

Venue

2025

Year

Abstract

Abstract General reasoning represents a long-standing and formidable challenge in artificial intelligence (AI). Recent breakthroughs, exemplified by large language models (LLMs) 1,2 and chain-of-thought (CoT) prompting 3 , have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent on extensive human-annotated demonstrations and the capabilities of models are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labelled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions and STEM fields, surpassing its counterparts trained through conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically used to guide and enhance the reasoning capabilities of smaller models.

Analysis

Why This Paper Matters

This paper, published in Nature with over 5,400 citations, represents a paradigm shift in how we train large language models for reasoning. Traditionally, chain-of-thought reasoning has required extensive human-annotated demonstrations, which are costly and difficult to scale. DeepSeek-R1 demonstrates that pure reinforcement learning—using only reward signals from verifiable tasks—can incentivize LLMs to develop sophisticated reasoning strategies on their own.

The significance is twofold. First, it dramatically reduces the need for human labor in creating training data for reasoning tasks. Second, and more importantly, the RL-trained models develop emergent reasoning patterns—such as self-reflection, verification, and dynamic strategy adaptation—that are not explicitly programmed. These patterns mirror human-like metacognitive processes and suggest that RL can unlock capabilities that supervised learning cannot easily instill.

Technical Contributions

  • Pure RL training pipeline: The framework uses only reward signals from verifiable tasks (e.g., math problems with known answers, coding problems with test cases) to train the model, eliminating the need for human-annotated reasoning chains.
  • Emergent reasoning patterns: The model spontaneously develops self-reflection (checking its own reasoning steps), verification (validating intermediate results), and dynamic strategy adaptation (changing approach when stuck).
  • Scalable reward design: The authors design reward functions that are automatically computable, enabling training on large-scale datasets without human intervention.
  • Knowledge distillation: The emergent reasoning patterns from the large RL-trained model are systematically used to guide and enhance smaller models, enabling broader deployment.

Results

The paper reports that the RL-trained model achieves superior performance on verifiable tasks including mathematics, coding competitions, and STEM fields. It surpasses counterparts trained through conventional supervised learning on human demonstrations. Specific metrics are not detailed in the abstract, but the high citation count (5,469) and publication in Nature indicate strong empirical validation. The emergent reasoning patterns are shown to be transferable to smaller models, improving their reasoning capabilities.

Significance

DeepSeek-R1 has broad implications for the AI field. It challenges the prevailing assumption that high-quality reasoning requires human-curated demonstrations, suggesting that RL with verifiable rewards can be a more scalable and potentially more powerful alternative. This could accelerate progress in AI for science, mathematics, and engineering by enabling models to self-improve through practice on verifiable problems. The ability to distill reasoning patterns into smaller models also makes advanced reasoning more accessible for practical applications. However, the approach is currently limited to domains where automatic verification is possible, leaving open questions about how to extend it to open-ended or subjective reasoning tasks.