Qwen2-Math logo

Qwen2-Math

Paid

State-of-the-art math language model series outperforming GPT-4o

4.7
Inputs: textOutputs: text
Type
Saas

About Qwen2-Math

Qwen2-Math is a series of specialized large language models for mathematics, built upon the Qwen2 LLM foundation. Developed by the Qwen Team, these models are pre-trained on a meticulously designed mathematics-specific corpus comprising high-quality web texts, books, codes, exam questions, and synthetic data. The instruction-tuned variant, Qwen2-Math-Instruct, incorporates a math-specific reward model and reinforcement learning via Group Relative Policy Optimization (GRPO) to enhance reasoning. Available in 1.5B, 7B, and 72B parameter sizes, Qwen2-Math significantly outperforms open-source models and rivals closed-source models like GPT-4o, Claude-3.5-Sonnet, and Gemini-1.5-Pro on English and Chinese math benchmarks.

Key Features

Built upon Qwen2 large language models
Pre-trained on a large-scale, high-quality mathematics-specific corpus (web texts, books, codes, exam questions, synthetic data)
Instruction-tuned with math-specific reward model and GRPO reinforcement learning
Available in 1.5B, 7B, and 72B parameter sizes
Supports English and (coming soon) bilingual English + Chinese
Evaluated on benchmarks including GSM8K, Math, MMLU-STEM, OlympiadBench, CollegeMath, GaoKao, AIME2024, AMC2023, and Chinese math exams
Outperforms GPT-4o, Claude-3.5-Sonnet, Gemini-1.5-Pro, and Llama-3.1-405B on math tasks

Pros & Cons

Pros
  • Outperforms leading open-source and closed-source models on math benchmarks
  • Specialized math corpus enhances mathematical reasoning capabilities
  • Instruction-tuned with reward model for improved accuracy
  • Varied model sizes to suit different computational resources
  • Open-source with weights available on Hugging Face, ModelScope, and GitHub
Cons
  • Currently primarily supports English; bilingual support is planned but not yet released
  • Large model (72B) requires significant computational resources for inference
  • May generate incorrect solutions for complex problems; solutions are not guaranteed
  • Limited to mathematical tasks; not a general-purpose model

Best For

Solving arithmetic and mathematical problemsAnswering math competition questions (e.g., AIME, AMC, IMO-style)Assisting in math education and tutoringResearch in mathematical reasoning for AI

Alternatives to Qwen2-Math

FAQ

What sizes are available for Qwen2-Math?
Qwen2-Math is available in 1.5B, 7B, and 72B parameter sizes.
Does Qwen2-Math support languages other than English?
Currently, Qwen2-Math mainly supports English. The team plans to release bilingual (English and Chinese) math models soon.
How does Qwen2-Math compare to GPT-4o?
On a series of math benchmarks, Qwen2-Math-72B-Instruct outperforms GPT-4o, as well as Claude-3.5-Sonnet, Gemini-1.5-Pro, and Llama-3.1-405B.
Is Qwen2-Math open-source?
Yes, the models are open-source and available on GitHub, Hugging Face, and ModelScope.