QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
FreeLong-context large reasoning models through RL
FreeFree tier
Inputs: textOutputs: text
About QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
QwenLong-L1 is a framework that adapts short-context large reasoning models (LRMs) to long-context scenarios via progressive context scaling, combining warm-up supervised fine-tuning (SFT), curriculum-guided phased reinforcement learning (RL), and difficulty-aware retrospective sampling. It achieves state-of-the-art performance on seven long-context document question-answering benchmarks, outperforming OpenAI o3-mini and Qwen3-235B-A22B, and performing on par with Claude-3.7-Sonnet-Thinking.
Key Features
Progressive context scaling
Warm-up supervised fine-tuning (SFT) stage
Curriculum-guided phased reinforcement learning
Difficulty-aware retrospective sampling
Adapts short-context LRMs to long-context scenarios
Pros & Cons
Pros
- Outperforms OpenAI o3-mini and Qwen3-235B-A22B on long-context QA benchmarks
- Achieves performance on par with Claude-3.7-Sonnet-Thinking
- Introduces novel RL techniques for stable training in long-context settings
Cons
- Evaluation limited to document question-answering benchmarks; performance on other long-context tasks not yet verified
Best For
Long-context document question-answeringInformation-intensive environments requiring robust reasoning
FAQ
What is QwenLong-L1?
QwenLong-L1 is a framework that extends short-context large reasoning models to handle long-context inputs using reinforcement learning and progressive context scaling.
How does QwenLong-L1 improve long-context reasoning?
It uses a warm-up SFT stage, curriculum-guided phased RL, and difficulty-aware retrospective sampling to stabilize training and incentivize exploration.