Gemini 2.5 Pro Capable of Winning Gold at IMO 2025
FreeGold medal performance on IMO 2025 via model-agnostic verification and refinement
About Gemini 2.5 Pro Capable of Winning Gold at IMO 2025
This research paper introduces a model-agnostic verification-and-refinement pipeline that achieves a gold medal level on the International Mathematical Olympiad (IMO) 2025. The pipeline uses carefully designed prompts to iteratively verify and refine solutions generated by leading large language models such as Gemini 2.5 Pro, Grok-4, and GPT-5. It correctly solved 5 out of 6 problems (approximately 85.7% accuracy), significantly surpassing the baseline accuracies of the models when used alone (which ranged from 21.4% to 38.1%). The work emphasizes that advanced AI reasoning requires both powerful base models and effective methodologies to harness their full potential.
Key Features
Pros & Cons
- Achieves gold-medal level on the world’s hardest high-school math competition
- Model-agnostic, allowing use with any advanced LLM
- Substantially improves base model accuracy (e.g., 31.6% → 85.7% for Gemini 2.5 Pro)
- Relies on careful prompt engineering, which may require domain expertise
- Only tested on a single year’s problems (IMO 2025)
- Effectiveness still depends on the capabilities of the underlying LLM