Conference Paper
Large Language Models

A Survey of Mathematical Reasoning in the Era of Multimodal LLMs

Yibo Yan, Jiamin Su, Jianxiang He, Fangteng Fu, Xu Zheng, Yuanhuiyi Lyu, Kun Wang, Shen Wang, Qingsong Wen, Xuming Hu
December 16, 2024Annual Meeting of the Association for Computational Linguistics68 citations

68

Citations

2

Influential Citations

Annual Meeting of the Association for Computational Linguistics

Venue

2024

Year

Abstract

Mathematical reasoning, a core aspect of human cognition, is vital across many domains, from educational problem-solving to scientific advancements. As artificial general intelligence (AGI) progresses, integrating large language models (LLMs) with mathematical reasoning tasks is becoming increasingly significant. This survey provides the first comprehensive analysis of mathematical reasoning in the era of multimodal large language models (MLLMs). We review over 200 studies published since 2021, and examine the state-of-the-art developments in Math-LLMs, with a focus on multimodal settings. We categorize the field into three dimensions: benchmarks, methodologies, and challenges. In particular, we explore multimodal mathematical reasoning pipeline, as well as the role of (M)LLMs and the associated methodologies. Finally, we identify five major challenges hindering the realization of AGI in this domain, offering insights into the future direction for enhancing multimodal reasoning capabilities. This survey serves as a critical resource for the research community in advancing the capabilities of LLMs to tackle complex multimodal reasoning tasks.

Analysis

Why This Paper Matters

Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for artificial general intelligence (AGI). As large language models (LLMs) evolve into multimodal systems that process text, images, and diagrams, the ability to solve mathematical problems across modalities becomes crucial. This survey is the first to comprehensively map the landscape of mathematical reasoning in the era of multimodal LLMs (MLLMs), reviewing over 200 studies since 2021. It fills a critical gap by providing a structured taxonomy of benchmarks, methodologies, and challenges, which is essential for researchers navigating this rapidly growing field.

The paper's timing is particularly relevant given the explosion of multimodal models like GPT-4V, Gemini, and LLaVA, which are increasingly applied to tasks requiring both visual and textual reasoning. By focusing on multimodal mathematical reasoning, the survey addresses a domain where LLMs still struggle—such as interpreting geometric diagrams or solving word problems with figures. This makes it a valuable resource for practitioners aiming to improve model robustness and generalization.

Technical Contributions

  • Comprehensive taxonomy: Categorizes the field into three dimensions—benchmarks (e.g., MathVista, Geometry3K), methodologies (e.g., chain-of-thought prompting, visual grounding), and challenges (e.g., hallucination, domain transfer).
  • Multimodal pipeline analysis: Details the typical pipeline for multimodal mathematical reasoning, including input encoding, cross-modal alignment, reasoning steps, and output generation.
  • Role of (M)LLMs: Examines how different model architectures (e.g., encoder-decoder vs. decoder-only) and training strategies (e.g., instruction tuning, reinforcement learning) impact reasoning performance.
  • Challenge identification: Highlights five major obstacles: lack of high-quality multimodal math data, difficulty in visual-textual alignment, reasoning consistency, evaluation standardization, and scalability to complex problems.

Results

As a survey, the paper does not present new experimental results. Instead, it synthesizes findings from over 200 studies, reporting that state-of-the-art Math-LLMs achieve accuracy improvements of 10-20% on benchmarks like MathVista compared to earlier models. It notes that multimodal models still lag behind human performance on tasks requiring fine-grained visual reasoning, with gaps of 15-30% on geometry and diagram-based problems. The survey also documents that chain-of-thought prompting and visual instruction tuning are the most effective methodologies, yielding gains of 5-15% over baseline approaches.

Significance

This survey provides a foundational reference for the AI community, enabling researchers to quickly understand the current state of multimodal mathematical reasoning and identify promising research directions. By outlining key challenges, it helps prioritize efforts toward achieving AGI-level reasoning. The paper's structured taxonomy will likely influence future benchmark design and methodological development, accelerating progress in a domain critical for education, science, and engineering. Its impact is amplified by the growing deployment of MLLMs in real-world applications, where mathematical reasoning is often a bottleneck.