ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
A comprehensive benchmark dataset designed to evaluate the mathematical reasoning abilities of LLMs at the Olympiad level, comprising 4428 competition-level problems across 33 sub-domains and 10 difficulty levels, with a rigorous evaluation process utilizing GPT-4o and an open-source verifier, OmniJudge, to address the limitations of existing benchmarks that are now easily solved by advanced LLMs.
As large language models (LLMs) continue to advance, existing mathematical reasoning benchmarks have become saturated, with many models achieving near-perfect scores. This saturation limits the ability to differentiate between models and identify areas for improvement. OmniMath addresses this gap by introducing a benchmark focused on Olympiad-level problems, which require deep reasoning, multi-step logic, and domain-specific knowledge. The inclusion of 33 sub-domains and 10 difficulty levels ensures comprehensive coverage and granular evaluation.
The paper also introduces OmniJudge, an open-source verifier that provides a transparent and reproducible evaluation process. This is crucial for the research community, as proprietary evaluation methods can hinder progress and comparability. By making both the benchmark and verifier open-source, OmniMath promotes collaboration and standardization in LLM evaluation.
The abstract does not provide specific performance metrics or comparisons with other models. The primary contribution is the benchmark itself and the evaluation methodology. Future work will likely report results on OmniMath using various LLMs, providing insights into their mathematical reasoning capabilities at the Olympiad level.
OmniMath has the potential to become a standard benchmark for evaluating high-level mathematical reasoning in LLMs. By focusing on Olympiad-level problems, it challenges models to go beyond pattern matching and demonstrate genuine reasoning. The open-source nature of both the benchmark and OmniJudge fosters reproducibility and community-driven improvements. This work could accelerate progress in AI for mathematics, with applications in education, automated theorem proving, and scientific discovery.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba