MathEval logo

MathEval

Free

a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.

FreeFree tier
Inputs: textOutputs: text
Type
Open Source

About MathEval

MathEval is a comprehensive benchmarking platform designed to evaluate the mathematical abilities of large language models (LLMs). It covers 20 specialized math evaluation sets with nearly 30,000 problems, spanning arithmetic, primary/secondary school competition math, and portions of higher mathematics. The platform supports zero-shot and few-shot evaluation, and is compatible with Hugging Face models, API models, and custom open-source models. It provides a one-stop reference for cross-model comparison and aims to guide further improvement of LLM mathematical reasoning. The tool is free and open-source, regularly updated with new datasets and models.

Key Features

20 specialized math evaluation datasets covering arithmetic, competition math, and higher mathematics
Supports zero-shot and few-shot evaluation modes
Compatible with Hugging Face models, API models, and custom open-source models
Multi-dimensional evaluation across difficulty levels and mathematical topics
Flexible extension to easily add new evaluation sets
Regularly updated with new datasets and model results

Pros & Cons

Pros
  • Comprehensive coverage of 20 math datasets and nearly 30,000 problems
  • Free and open-source, with no usage restrictions
  • Supports a wide variety of model types (HF, API, custom)
  • Regular updates with new datasets and model evaluations
  • Multi-dimensional analysis for detailed insights
Cons
  • Currently limited to evaluating mathematical abilities only
  • Requires technical knowledge to set up and run evaluations
  • User interface is primarily in Chinese, which may be a barrier for non-Chinese speakers

Best For

Comparing mathematical abilities of large language models across diverse problem typesEvaluating LLM performance on arithmetic, competition math, and higher mathResearch on improving mathematical reasoning capabilities of LLMsBenchmarking academic and commercial models for math-related tasks

FAQ

What types of models can be evaluated with MathEval?
MathEval supports Hugging Face models, API models (e.g., GPT-4, Qwen2), and custom open-source models.
How can I add a new model or dataset to the benchmark?
You can request to join the evaluation or collaborate by sending your requirements to the contact email provided on the website.
Is MathEval free to use?
Yes, MathEval is free and open-source. The computing support is provided by a national open innovation platform.