Are NLP Models really able to Solve Simple Math Word Problems?
Arkil PatelSatwik BhattamishraNavin Goyal
Demonstrates that current math word problem solvers rely on shallow heuristics rather than mathematical reasoning, and introduces the SVAMP challenge benchmark to accurately evaluate arithmetic comprehension in language models.
Recent advances in natural language processing have led to high test scores on standard arithmetic word problem benchmarks, creating a common belief that elementary-level math problems are effectively solved. Consequently, research and development resources have shifted toward more complex challenges. However, high accuracy on existing benchmarks may reflect reliance on superficial cues rather than genuine mathematical reasoning, creating performance risks for real-world automated educational tools.
The article evaluates whether current state-of-the-art models truly solve simple, elementary-school math word problems or merely exploit statistical shortcuts in standard benchmarks. To test this, the authors introduce a new challenge dataset to reliably measure model robustness against basic reasoning and structural variations.
The authors conducted an empirical investigation using leading model architectures (Sequence-to-Sequence, Goal-Driven Tree-Structured models, and Graph-to-Tree models) across two standard benchmarks: MAWPS (2,373 problems) and ASDiv-A (1,218 problems). They first stripped the question sentences from the problems to test if models could solve them using only the descriptive narrative. Next, they evaluated a constrained baseline stripped of word-order awareness to see if bag-of-words keyword associations sufficed. Finally, the authors constructed SVAMP, a challenge dataset of 1,000 elementary-level problems created by applying controlled variations—altering question focus, modifying reasoning steps, or inserting irrelevant details—to seed problems.
The investigation produced four critical findings. First, existing benchmarks are heavily compromised by superficial artifacts: top models solved 64.4% of ASDiv-A problems and 77.7% of MAWPS problems with the question entirely removed. Second, a constrained model lacking any word-order understanding achieved 77.9% accuracy on MAWPS and 51.2% on ASDiv-A by latching onto single trigger words. Third, when evaluated on SVAMP, top-performing models experienced severe performance drops, with the best model achieving only 43.8% accuracy despite the dataset containing no higher-level operations. Fourth, model accuracy degraded substantially as problem complexity grew slightly, dropping from 78.3% on two-number problems in SVAMP to just 25.4% on problems containing three or four numbers.
These findings demonstrate that high benchmark accuracies are misleading; current models rely heavily on shallow heuristics rather than actual language understanding or numerical reasoning. Deploying these systems into real-world settings, such as automated tutoring, carries substantial performance and user-trust risks because slight changes in phrasing cause models to fail unpredictably. The assumption that basic arithmetic word problems are solved is incorrect.
Stakeholders and developers should halt the premature shift away from elementary reasoning tasks and incorporate challenge sets like SVAMP into evaluation pipelines to ensure robust testing. Training regimes must combine standard datasets with structural variations to mitigate shallow heuristic exploitation. Moving forward, research should focus on developing model architectures capable of genuine contextual reasoning and number binding before deploying automated solvers in high-stakes educational applications.
The scope of the article is limited to English-language, one-unknown arithmetic problems up to a fourth-grade level. While the results clearly demonstrate the brittleness of current language models on these benchmarks, further research is required to determine whether these vulnerabilities generalize across other languages and more complex multi-variable domains.
- Paper: Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference, R. Thomas McCoy et al. (2019). This foundational work establishes how NLP models rely on shallow, non-generalizable heuristics to succeed on standard benchmarks, providing the direct conceptual methodology for diagnosing heuristic shortcuts in math word problem solvers.
- Paper: DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs, Dheeru Dua et al. (2019). This paper introduces benchmark design principles for discrete arithmetic reasoning and adversarial evaluation in NLP, serving as essential background for testing numerical reasoning beyond pattern matching.
- Paper: Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, Peter Clark et al. (2018). This study pioneeringly isolates challenge sets that foil shallow retrieval heuristics to test genuine question-answering reasoning, laying the groundwork for creating diagnostic challenge sets like SVAMP.
- Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). This paper introduces the GSM8K benchmark and solution verification to specifically tackle multi-step reasoning failures in elementary math word problems revealed by SVAMP.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This work demonstrates how eliciting intermediate chain-of-thought reasoning steps enables language models to overcome the shallow pattern-matching limitations documented in SVAMP.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). This paper introduces subproblem decomposition strategies to directly address the vulnerability of language models when solving multi-step arithmetic word problems.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). This benchmark study extends the challenge-dataset paradigm of SVAMP by generating controlled symbolic variations to evaluate the true reasoning robustness and perturbation sensitivity of modern LLMs.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). This research overcomes shallow linguistic heuristics in word problem solving by disentangling semantic reasoning from arithmetic computation via generated code execution.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This work investigates process-level supervision to ensure mathematical solvers execute genuinely valid step-by-step logic rather than relying on superficial outcome heuristics.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). Following the demonstration that elementary math benchmarks are vulnerable to shallow heuristics, this benchmark scales mathematical reasoning evaluation to complex competition-level problems requiring rigorous multi-step deduction.
