A Survey of Deep Learning for Mathematical Reasoning
Pan LuLiang QiuWenhao YuSean WelleckKai-Wei Chang
Provides a systematic taxonomy and evaluation of over 180 studies covering neural architectures, large language models, and benchmarks for automated mathematical problem solving and theorem proving.
Mathematical reasoning is a fundamental pillar of human intelligence and decision-making across engineering, science, and finance. While artificial intelligence systems have achieved notable success in general language tasks, enabling machines to reliably solve math word problems, prove formal theorems, and interpret multi-modal geometric figures remains a critical bottleneck. The article aims to establish a clear taxonomy of mathematical reasoning tasks, evaluate the performance and limitations of deep learning methods over the past decade, and provide actionable future directions for the field.
The authors conducted a comprehensive review of over 180 studies published between 2013 and 2022 across the machine learning and natural language processing communities. The analysis categorizes benchmarks into key areas such as math word problems, formal and informal theorem proving, geometry problem solving, and general quantitative question answering. It systematically examines three generations of methods: specialized neural networks (sequence-to-sequence, graph-based, and attention architectures), pre-trained language models, and recent large language models utilizing in-context learning and chain-of-thought prompting.
The review identifies critical breakthroughs alongside significant technical vulnerabilities. First, while scaling models enhances performance—exemplified by Google's Minerva achieving 75.0% on advanced STEM benchmarks and GPT-3 reaching 93.0% on MultiArith—state-of-the-art systems remain brittle. When tested on SVAMP, an elementary benchmark with slight phrasing variations, top systems degrade sharply, with Graph2Tree achieving only 43.8% and GPT-3 reaching 63.7%. Second, current tokenization methods inherently fail at numerical representation; models break multi-digit numbers into arbitrary sub-word fragments, leading to systematic calculation errors and poor generalization on large numbers. Third, outcome-based approaches (such as self-consistency sampling) and process-based frameworks (like decomposing problems into sub-tasks or delegating math execution to external computer programs) significantly outperform standard single-pass prompting. Finally, the research landscape remains heavily biased toward text-only English datasets, leaving multi-modal contexts (diagrams and tables) and low-resource domains severely underexplored.
These findings indicate that high benchmark scores can mask a lack of true reasoning capability. The presence of ungrounded outputs and hallucinated logical steps creates major operational and safety risks for organizations attempting to deploy AI for automated finance, engineering, or scientific analytics without oversight. Because systems are highly sensitive to superficial prompt changes and struggle with basic number representations, standard language models cannot be trusted as standalone mathematical solvers.
To address these limitations, stakeholders and practitioners should avoid relying on single-pass model outputs and instead implement hybrid workflows, such as program-aided execution that routes deterministic calculations to verified computing engines. Decision-makers should support research into better numerical encoding formats (such as scientific notation), expand multi-modal benchmark datasets, and integrate reinforcement learning from human or formal theorem feedback. Because this survey focuses on published literature through 2022, continued caution and empirical pilots are necessary as model capabilities rapidly evolve.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). Introduces the seminal MATH competition benchmark that serves as the central evaluation standard throughout the survey of deep learning for mathematical reasoning.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Establishes chain-of-thought prompting, the foundational multi-step inference technique surveyed as a prerequisite for neural mathematical problem solving.
- Paper: Solving Quantitative Reasoning Problems with Language Models, Aitor Lewkowycz et al. (2022). Presents Minerva, demonstrating large-scale pretraining on mathematical text and serving as a critical baseline discussed in the survey.
- Paper: LILA: A Unified Benchmark for Mathematical Reasoning, Swaroop Mishra et al. (2022). Provides a comprehensive unified benchmark (LILA) covering diverse mathematical abilities and program synthesis baselines evaluated in the survey.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Introduces self-consistency decoding across sampled reasoning paths, a core inference-time method reviewed in the survey.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). Proposes Program of Thoughts prompting to decouple logical reasoning from deterministic computation, a key methodology highlighted in mathematical reasoning literature.
- Paper: Are NLP Models really able to Solve Simple Math Word Problems?, Arkil Patel et al. (2021). Critically assesses standard arithmetic word problem benchmarks and introduces SVAMP to evaluate robustness against superficial cues.
- Paper: NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks, Swaroop Mishra et al. (2022). Supplies the NumGLUE benchmark for evaluating multi-task arithmetic reasoning across diverse linguistic formats.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot chain-of-thought prompting on arithmetic benchmarks, establishing an essential baseline for eliciting step-by-step reasoning without exemplars.
- Paper: DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs, Dheeru Dua et al. (2019). Introduces the DROP benchmark, pioneering the integration of discrete mathematical reasoning into paragraph reading comprehension.
- Paper: DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, Zhihong Shao et al. (2024). Advances mathematical reasoning in open models through large-scale math web pre-training and Group Relative Policy Optimization.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). Extends post-training methodologies by showing that large-scale reinforcement learning without supervised warm-starts incentivizes emergent reasoning behaviors.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Builds upon process supervision by establishing an automated framework to verify and reinforce intermediate reasoning steps without human annotation.
- Paper: ReFT: Reasoning with Reinforced Fine-Tuning, Luong Quoc Trung et al. (2024). Combines supervised fine-tuning with online reinforcement learning to explore and generalize across diverse mathematical solution paths.
- Paper: Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry2, Yuri Chervonyi et al. (2025). Applies advanced neuro-symbolic search and theorem proving to achieve gold-medalist performance on Olympiad-level Euclidean geometry.
- Paper: OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems, Chaoqun He et al. (2024). Expands evaluation beyond traditional text-only math datasets to multimodal, bilingual Olympiad competition benchmarks.
- Paper: MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations, Kaixuan Huang et al. (2025). Probes the limits and memorization patterns of advanced reasoning models by introducing structural perturbations to the standard MATH benchmark.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). Examines whether LLM benchmark accuracy reflects robust logical problem-solving or dataset contamination using symbolic template variations.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). Refines process-level verification by reformulating step-by-step scoring as a sequential decision-making process using Q-value rankings.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). Scales inference-time compute by coordinating multiple reasoning trees with dynamic pruning and self-correction to enhance mathematical problem solving.
