MathEval
Freea comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.
About MathEval
MathEval is a comprehensive benchmarking platform designed to evaluate the mathematical abilities of large language models (LLMs). It covers 20 specialized math evaluation sets with nearly 30,000 problems, spanning arithmetic, primary/secondary school competition math, and portions of higher mathematics. The platform supports zero-shot and few-shot evaluation, and is compatible with Hugging Face models, API models, and custom open-source models. It provides a one-stop reference for cross-model comparison and aims to guide further improvement of LLM mathematical reasoning. The tool is free and open-source, regularly updated with new datasets and models.
Key Features
Pros & Cons
- Comprehensive coverage of 20 math datasets and nearly 30,000 problems
- Free and open-source, with no usage restrictions
- Supports a wide variety of model types (HF, API, custom)
- Regular updates with new datasets and model evaluations
- Multi-dimensional analysis for detailed insights
- Currently limited to evaluating mathematical abilities only
- Requires technical knowledge to set up and run evaluations
- User interface is primarily in Chinese, which may be a barrier for non-Chinese speakers