Does RL Incentivize Reasoning Capacity in LLMs Beyond the Base Model
Unknown
RLVR improves low-k pass@k efficiency but restricts high-k reasoning capacity by reducing exploration of successful paths already present in base models.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
RLVR improves low-k pass@k efficiency but restricts high-k reasoning capacity by reducing exploration of successful paths already present in base models.
Unknown
Trains an LLM with RL using a verifiable reward on just one example, matching performance of larger datasets in math reasoning.
Unknown
A large language model built on Mistral Small, enhanced with SFT, RLVR, and inference optimization for Indian languages and reasoning tasks.
Unknown
Tulu V3 presents a fully open post-training recipe for Llama 3.1 models, achieving state-of-the-art performance via SFT, DPO, and RLVR.
Unknown
Proposes post-training world models via reinforcement learning to improve their generality and task performance.
Md Tanvirul Alam
Trace is a taxonomy-guided environment for multidomain visual reasoning that generates verifiable training data, improving VLMs by 3-4% on external benchmarks via RLVR.