Gemini 2.5 Pro Capable of Winning Gold at IMO 2025 logo

Gemini 2.5 Pro Capable of Winning Gold at IMO 2025

Free

Gold medal performance on IMO 2025 via model-agnostic verification and refinement

FreeFree tier
Type
Open Source

About Gemini 2.5 Pro Capable of Winning Gold at IMO 2025

This research paper introduces a model-agnostic verification-and-refinement pipeline that achieves a gold medal level on the International Mathematical Olympiad (IMO) 2025. The pipeline uses carefully designed prompts to iteratively verify and refine solutions generated by leading large language models such as Gemini 2.5 Pro, Grok-4, and GPT-5. It correctly solved 5 out of 6 problems (approximately 85.7% accuracy), significantly surpassing the baseline accuracies of the models when used alone (which ranged from 21.4% to 38.1%). The work emphasizes that advanced AI reasoning requires both powerful base models and effective methodologies to harness their full potential.

Key Features

Model-agnostic design works with Gemini 2.5 Pro, Grok-4, and GPT-5
Verification-and-refinement pipeline using carefully crafted prompts
Achieved 85.7% accuracy (5/6 problems) on IMO 2025
Significant improvement over baseline model accuracies (21.4%–38.1%)
Avoids data contamination by using problems from the recent 2025 competition

Pros & Cons

Pros
  • Achieves gold-medal level on the world’s hardest high-school math competition
  • Model-agnostic, allowing use with any advanced LLM
  • Substantially improves base model accuracy (e.g., 31.6% → 85.7% for Gemini 2.5 Pro)
Cons
  • Relies on careful prompt engineering, which may require domain expertise
  • Only tested on a single year’s problems (IMO 2025)
  • Effectiveness still depends on the capabilities of the underlying LLM

Best For

Solving Olympiad-level mathematics problemsAdvancing AI reasoning capabilities for complex, creative tasksBenchmarking and improving LLM performance on rigorous mathematical challenges

FAQ

What is the verification-and-refinement pipeline?
The pipeline is a model-agnostic methodology that uses carefully designed prompts to generate, verify, and refine candidate solutions iteratively, significantly improving accuracy on challenging math problems.
Which models were tested with this pipeline?
The pipeline was tested with Gemini 2.5 Pro, Grok-4, and GPT-5. All three showed substantial accuracy gains over their baseline performances.
How many problems did the pipeline solve correctly on IMO 2025?
It correctly solved 5 out of 6 problems, corresponding to approximately 85.7% accuracy.