MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
Kaixuan HuangJiacheng GuoZihao LiXiang JiJiawei GeWenzhe LiYingqing GuoTianle CaiHui YuanRunzhe Wang
Introduces MATH-Perturb to benchmark large language models against fundamental problem modifications that break original solution paths, revealing that leading reasoning models suffer steep accuracy drops by blindly applying memorized problem-solving heuristics.
Recent advances have enabled large language models to achieve high scores on standard mathematical reasoning benchmarks. However, these strong results raise a critical question for decision-makers: do these systems possess genuine analytical reasoning capabilities, or are they relying on memorized patterns from their training data? Prior evaluations primarily tested model resilience using simple modifications, such as changing numerical values, which leave the fundamental problem-solving method unchanged. The article addresses this evaluation gap by investigating how models perform when confronted with structural changes that invalidate previously learned solution steps.
The main objective of the article is to benchmark and analyze the mathematical reasoning abilities of large language models when subjected to hard perturbations—modifications that appear textually similar to known problems but require fundamentally different solution strategies.
To evaluate this, a team of 12 mathematics experts curated two new evaluation sets from 279 challenging high school competition problems in the MATH benchmark: MATH-P-Simple (containing non-essential numerical or surface-level edits) and MATH-P-Hard (containing minimal edits that alter core conditions and demand deeper techniques). The researchers evaluated 18 leading language models, including advanced reasoning systems, proprietary commercial models, and open-source models, under zero-shot and in-context learning settings without computational tool assistance.
The evaluation yielded several key findings. First, all 18 models experienced significant performance drops of roughly 10% to 25% on MATH-P-Hard compared to the original problems; for example, o1-mini dropped by 16.49% (from 94.27% to 78.49%) and gemini-2.0-flash-thinking dropped by 12.90% (from 92.47% to 78.14%). Second, qualitative error analysis revealed a subtle form of memorization: models frequently recognized familiar problem patterns and blindly applied learned solution techniques without verifying whether altered constraints made those steps invalid. For top-tier models, memorization accounted for an estimated 25% to 40% of their errors on hard perturbations. Third, providing the original problem and solution as an in-context demonstration yielded marginal overall gains (often below 5%) on hard perturbations because the demonstration frequently misled the models into repeating the original, now incorrect, solution method.
These findings indicate that high benchmark scores may mask critical operational vulnerabilities. In high-stakes applications—such as scientific research, automated engineering, and educational tutoring—models risk generating convincing but flawed solutions when faced with novel, out-of-distribution variations of familiar problems. Relying on superficial prompt demonstrations or standard fine-tuning on limited problem sets does not resolve this issue and may even increase the risk of misleading failures.
The article recommends that AI developers and deploying organizations shift evaluation frameworks toward hard perturbation benchmarks to accurately assess reasoning robustness. Organizations should exercise caution when using retrieval-augmented or demonstration-based prompting with closely matching problems, as these examples can cause models to misapply past methods. Further research must prioritize developing training and verification techniques that explicitly evaluate whether problem constraints match specific solution strategies before executing them.
These conclusions are bounded by the study's scope of 279 competition-level mathematical problems and the exclusion of external tools such as code interpreters. Nevertheless, the rigorous multi-expert annotation process and consistent performance degradation observed across 18 diverse models provide high confidence that current systems struggle with structural problem variations.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). This paper introduces the foundational MATH benchmark of competition-level problems from which MATH-Perturb curates its evaluation subsets and analyzes memorization versus reasoning.
- Paper: Are NLP Models really able to Solve Simple Math Word Problems?, Arkil Patel et al. (2021). This work pioneered the approach of testing whether math problem solvers rely on superficial patterns versus genuine understanding by introducing controlled variations and question perturbations in SVAMP.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). This study demonstrates how large language models suffer performance drops under symbolic and numerical perturbations on grade-school math, establishing the groundwork for evaluating harder structural perturbations on competition math.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). This paper establishes the vulnerability of reasoning models to irrelevant context and superficial distractions, motivating the need to evaluate deeper, constraint-altering perturbations.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This foundational paper introduced chain-of-thought prompting for eliciting step-by-step mathematical reasoning, which serves as the baseline reasoning paradigm analyzed in MATH-Perturb.
- Paper: DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, Zhihong Shao et al. (2024). This work establishes modern reinforcement learning and pre-training techniques for mathematical problem-solving on the MATH benchmark, representing the base architecture styles evaluated in MATH-Perturb.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This study introduces step-by-step process supervision on the MATH dataset, providing background on intermediate reasoning verification methods discussed in MATH-Perturb.
- Paper: Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation, Yanqi Dai et al. (2026). This work directly addresses the reasoning brittleness identified on hard competition problems by reforming questions and tuning reinforcement learning policies to prioritize difficult mathematical variations.
- Paper: Reasoning Models Generate Societies of Thought, Junsol Kim et al. (2026). This paper investigates internal multi-perspective deliberations in reasoning models like DeepSeek-R1, offering an interpretability perspective on why models struggle or succeed when self-correcting on complex problems.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). This paper extends the critique of superficial benchmark accuracy by developing methods to evaluate step validity and redundancy in intermediate mathematical reasoning.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). This research builds on verification deficiencies by proposing Q-value ranking process reward models that better evaluate whether sequential steps actually lead to valid solutions.
- Paper: Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?, Yang Yue et al. (2025). This study investigates whether reinforcement learning with verifiable rewards truly expands foundational reasoning capacity or merely explores existing patterns within base models.
