ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
68
Citations
2
Influential Citations
Annual Meeting of the Association for Computational Linguistics
Venue
2024
Year
Mathematical reasoning, a core aspect of human cognition, is vital across many domains, from educational problem-solving to scientific advancements. As artificial general intelligence (AGI) progresses, integrating large language models (LLMs) with mathematical reasoning tasks is becoming increasingly significant. This survey provides the first comprehensive analysis of mathematical reasoning in the era of multimodal large language models (MLLMs). We review over 200 studies published since 2021, and examine the state-of-the-art developments in Math-LLMs, with a focus on multimodal settings. We categorize the field into three dimensions: benchmarks, methodologies, and challenges. In particular, we explore multimodal mathematical reasoning pipeline, as well as the role of (M)LLMs and the associated methodologies. Finally, we identify five major challenges hindering the realization of AGI in this domain, offering insights into the future direction for enhancing multimodal reasoning capabilities. This survey serves as a critical resource for the research community in advancing the capabilities of LLMs to tackle complex multimodal reasoning tasks.
Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for artificial general intelligence (AGI). As large language models (LLMs) evolve into multimodal systems that process text, images, and diagrams, the ability to solve mathematical problems across modalities becomes crucial. This survey is the first to comprehensively map the landscape of mathematical reasoning in the era of multimodal LLMs (MLLMs), reviewing over 200 studies since 2021. It fills a critical gap by providing a structured taxonomy of benchmarks, methodologies, and challenges, which is essential for researchers navigating this rapidly growing field.
The paper's timing is particularly relevant given the explosion of multimodal models like GPT-4V, Gemini, and LLaVA, which are increasingly applied to tasks requiring both visual and textual reasoning. By focusing on multimodal mathematical reasoning, the survey addresses a domain where LLMs still struggle—such as interpreting geometric diagrams or solving word problems with figures. This makes it a valuable resource for practitioners aiming to improve model robustness and generalization.
As a survey, the paper does not present new experimental results. Instead, it synthesizes findings from over 200 studies, reporting that state-of-the-art Math-LLMs achieve accuracy improvements of 10-20% on benchmarks like MathVista compared to earlier models. It notes that multimodal models still lag behind human performance on tasks requiring fine-grained visual reasoning, with gaps of 15-30% on geometry and diagram-based problems. The survey also documents that chain-of-thought prompting and visual instruction tuning are the most effective methodologies, yielding gains of 5-15% over baseline approaches.
This survey provides a foundational reference for the AI community, enabling researchers to quickly understand the current state of multimodal mathematical reasoning and identify promising research directions. By outlining key challenges, it helps prioritize efforts toward achieving AGI-level reasoning. The paper's structured taxonomy will likely influence future benchmark design and methodological development, accelerating progress in a domain critical for education, science, and engineering. Its impact is amplified by the growing deployment of MLLMs in real-world applications, where mathematical reasoning is often a bottleneck.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba