Preprint
Machine Learning

OmniMath

October 1, 2024

0

Citations

0

Influential Citations

Venue

2024

Year

Abstract

A comprehensive benchmark dataset designed to evaluate the mathematical reasoning abilities of LLMs at the Olympiad level, comprising 4428 competition-level problems across 33 sub-domains and 10 difficulty levels, with a rigorous evaluation process utilizing GPT-4o and an open-source verifier, OmniJudge, to address the limitations of existing benchmarks that are now easily solved by advanced LLMs.

Analysis

Why This Paper Matters

As large language models (LLMs) continue to advance, existing mathematical reasoning benchmarks have become saturated, with many models achieving near-perfect scores. This saturation limits the ability to differentiate between models and identify areas for improvement. OmniMath addresses this gap by introducing a benchmark focused on Olympiad-level problems, which require deep reasoning, multi-step logic, and domain-specific knowledge. The inclusion of 33 sub-domains and 10 difficulty levels ensures comprehensive coverage and granular evaluation.

The paper also introduces OmniJudge, an open-source verifier that provides a transparent and reproducible evaluation process. This is crucial for the research community, as proprietary evaluation methods can hinder progress and comparability. By making both the benchmark and verifier open-source, OmniMath promotes collaboration and standardization in LLM evaluation.

Technical Contributions

  • Large-scale benchmark: 4428 competition-level problems across 33 sub-domains (e.g., algebra, geometry, number theory) and 10 difficulty levels, providing fine-grained evaluation.
  • OmniJudge verifier: An open-source tool that uses GPT-4o to verify model outputs, ensuring rigorous and consistent scoring.
  • Addressing benchmark saturation: Focuses on problems that are challenging for current advanced LLMs, pushing the frontier of mathematical reasoning evaluation.
  • Comprehensive domain coverage: Includes a wide range of mathematical sub-domains, enabling targeted analysis of model strengths and weaknesses.

Results

The abstract does not provide specific performance metrics or comparisons with other models. The primary contribution is the benchmark itself and the evaluation methodology. Future work will likely report results on OmniMath using various LLMs, providing insights into their mathematical reasoning capabilities at the Olympiad level.

Significance

OmniMath has the potential to become a standard benchmark for evaluating high-level mathematical reasoning in LLMs. By focusing on Olympiad-level problems, it challenges models to go beyond pattern matching and demonstrate genuine reasoning. The open-source nature of both the benchmark and OmniJudge fosters reproducibility and community-driven improvements. This work could accelerate progress in AI for mathematics, with applications in education, automated theorem proving, and scientific discovery.