Reasoning over Longer Horizons via RL
S. Motwani, Alesia Ivanova, Ziyang Cai, et al.
Introduces a scalable method using curriculum RL on synthetically composed short-horizon data to boost long-horizon reasoning, achieving up to 2.06x accuracy gains on competition-level benchmarks.