MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Kaixuan HuangJiacheng GuoZihao LiXiang JiJiawei GeWenzhe LiYingqing GuoTianle CaiHui YuanRunzhe Wang

article2025ICML72 citations

Introduces MATH-Perturb to benchmark large language models against fundamental problem modifications that break original solution paths, revealing that leading reasoning models suffer steep accuracy drops by blindly applying memorized problem-solving heuristics.

Listen

Recent advances have enabled large language models to achieve high scores on standard mathematical reasoning benchmarks. However, these strong results raise a critical question for decision-makers: do these systems possess genuine analytical reasoning capabilities, or are they relying on memorized patterns from their training data? Prior evaluations primarily tested model resilience using simple modifications, such as changing numerical values, which leave the fundamental problem-solving method unchanged. The article addresses this evaluation gap by investigating how models perform when confronted with structural changes that invalidate previously learned solution steps.

The main objective of the article is to benchmark and analyze the mathematical reasoning abilities of large language models when subjected to hard perturbations—modifications that appear textually similar to known problems but require fundamentally different solution strategies.

To evaluate this, a team of 12 mathematics experts curated two new evaluation sets from 279 challenging high school competition problems in the MATH benchmark: MATH-P-Simple (containing non-essential numerical or surface-level edits) and MATH-P-Hard (containing minimal edits that alter core conditions and demand deeper techniques). The researchers evaluated 18 leading language models, including advanced reasoning systems, proprietary commercial models, and open-source models, under zero-shot and in-context learning settings without computational tool assistance.

The evaluation yielded several key findings. First, all 18 models experienced significant performance drops of roughly 10% to 25% on MATH-P-Hard compared to the original problems; for example, o1-mini dropped by 16.49% (from 94.27% to 78.49%) and gemini-2.0-flash-thinking dropped by 12.90% (from 92.47% to 78.14%). Second, qualitative error analysis revealed a subtle form of memorization: models frequently recognized familiar problem patterns and blindly applied learned solution techniques without verifying whether altered constraints made those steps invalid. For top-tier models, memorization accounted for an estimated 25% to 40% of their errors on hard perturbations. Third, providing the original problem and solution as an in-context demonstration yielded marginal overall gains (often below 5%) on hard perturbations because the demonstration frequently misled the models into repeating the original, now incorrect, solution method.

These findings indicate that high benchmark scores may mask critical operational vulnerabilities. In high-stakes applications—such as scientific research, automated engineering, and educational tutoring—models risk generating convincing but flawed solutions when faced with novel, out-of-distribution variations of familiar problems. Relying on superficial prompt demonstrations or standard fine-tuning on limited problem sets does not resolve this issue and may even increase the risk of misleading failures.

The article recommends that AI developers and deploying organizations shift evaluation frameworks toward hard perturbation benchmarks to accurately assess reasoning robustness. Organizations should exercise caution when using retrieval-augmented or demonstration-based prompting with closely matching problems, as these examples can cause models to misapply past methods. Further research must prioritize developing training and verification techniques that explicitly evaluate whether problem constraints match specific solution strategies before executing them.

These conclusions are bounded by the study's scope of 279 competition-level mathematical problems and the exclusion of external tools such as code interpreters. Nevertheless, the rigorous multi-expert annotation process and consistent performance degradation observed across 18 diverse models provide high confidence that current systems struggle with structural problem variations.

arXiv: 2502.06453
Cover for MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Abstract

Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations – modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycks et al., 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (−16.49%) and gemini-2.0-flash-thinking (−12.9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models. The project is available here.

Table of Contents

  • 1. Introduction
  • 2. Dataset Curation
  • Benchmark Overview and Statistics.
  • Common Strategies for Perturbations.
  • 3. Experimental Results
  • 3.1. Benchmarking the Performance of LLMs
  • 3.2. Failure Mode Analysis
  • 3.3. Is Mode Collapse a Problem?
  • 3.4. Does In-context Learning Help or Hurt?
  • 4. Related Work
  • Perturbations to Existing Mathematical Benchmarks.
  • 5. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Version Information of the Models
  • B. Benchmark Statistics
  • C. Additional Experimental Results
  • C.1. Categorizing Model Responses Across Problem Variations
  • C.2. Is Mode Collapse a Problem?
  • C.3. The Effect of In-Context Learning
  • C.4. Ablation Study: In-Context Learning with the Original Example v.s. In-Context Learning with a Random Example
  • C.5. Inference-time Scaling Behaviors

Knowls

  1. Knowl 1 — MATH-Perturb Benchmark: Simple and Hard Mathematical Perturbations

    definition

    The MATH-Perturb benchmark is designed to assess whether large language model (LLM) mathematical reasoning reflects genuine problem-solving or memorization of solution patterns. It is constructed from 279279 level-5 (hardest difficulty) seed problems from the MATH dataset, split across 164164 training problems and 115115 test problems across seven subjects: Algebra (7979), Intermediate Algebra (4848), Counting & Probability (3838), Number Theory (3636), Prealgebra (3535), Precalculus (2222), and Geometry (2121).

    For each seed problem, two perturbed variants are created:

    1. MATH-P-Simple (Simple Perturbations): Non-essential modifications (such as changing non-critical numerical values, requesting a different but related quantity, or altering non-essential constraints) such that the perturbed problem is solvable using the exact same mathematical reasoning path and problem-solving method as the original problem.

    2. MATH-P-Hard (Hard Perturbations): Lexically minimal but structurally fundamental modifications (such as raising polynomial degrees, altering parameters to eliminate brute-force search, or removing simplifying conditions like symmetry, linearity, or reducibility) such that the original solution steps become invalid and deeper mathematical principles or harder problem-solving skills are required.

    Both splits enforce two structural requirements:

    • Minimal Edits: Problems maintain low lexical edit distances and high embedding cosine similarities to their original counterparts (mean reciprocal ranks of 0.9950.995 for MATH-P-Simple and 0.9860.986 for MATH-P-Hard when retrieving the original problem via text-embedding-3-large).
    • Changed Answers: The ground-truth answer for every perturbed problem is strictly different from the original problem's answer to prevent models from succeeding via exact answer memorization.
  2. Knowl 2 — Zero-Shot Chain-of-Thought Reasoning Degradation Under Hard Perturbations

    data/table

    Zero-shot chain-of-thought (CoT) evaluation on 18 large language models demonstrates that while performance on MATH-P-Simple is relatively close to original MATH level-5 problems, performance sharply deteriorates on MATH-P-Hard across all model families (long-CoT reasoning models, closed-source foundation models, open-weight general models, and math-specialized models).

    Model Original (%) MATH-P-Simple (%) MATH-P-Hard (%)
    All Train Test All Train Test All Train Test
    Gemini-2.0-flash-thinking-exp 92.47 92.68 92.17 91.04 87.80 95.65 78.14 77.44 79.13
    o1-preview 87.81 88.41 86.96 87.81 87.80 87.83 72.40 73.78 70.43
    o1-mini 94.27 93.90 94.78 94.98 93.29 97.39 78.49 79.27 77.39
    Gemini-2.0-flash-exp 88.17 87.20 89.57 82.80 81.71 84.35 67.03 68.29 65.22
    Gemini-1.5-pro 77.78 77.44 78.26 77.42 76.83 78.26 56.63 56.10 57.39
    GPT-4o 67.03 68.90 64.35 62.01 60.98 63.48 39.43 37.80 41.74
    GPT-4-turbo 56.99 55.49 59.13 55.20 56.71 53.04 34.41 36.59 31.30
    Claude-3.5-Sonnet 64.52 62.80 66.96 58.42 57.32 60.00 38.71 38.41 39.13
    Claude-3-Opus 41.94 39.02 46.09 41.94 39.63 45.22 26.52 25.00 28.70
    Llama-3.1-8B-Instruct 36.56 45.12 24.35 31.54 35.37 26.09 10.04 10.98 8.70
    Gemma-2-9b-it 27.60 28.05 26.96 27.60 30.49 23.48 11.83 12.80 10.43
    Phi-3.5-mini-instruct 26.16 27.44 24.35 28.67 26.83 31.30 14.34 15.24 13.04
    Deepseek-math-7b-rl 37.28 42.68 29.57 33.33 35.37 30.43 13.62 15.85 10.43
    Qwen2.5-Math-7B-Instruct 58.78 59.15 58.26 51.61 50.00 53.91 27.24 29.88 23.48
    Mathstral-7b-v0.1 36.56 43.29 26.96 36.20 42.07 27.83 14.70 16.46 12.17
    NuminaMath-7B-CoT 43.73 51.22 33.04 40.14 44.51 33.91 17.20 18.90 14.78
    MetaMath-13B-V1.0 21.15 32.32 5.22 7.53 7.32 7.83 5.73 4.88 6.96
    MAmmoTH2-8B 12.90 11.59 14.78 17.92 17.07 19.13 7.53 10.37 3.48

    All evaluated models experience substantial absolute accuracy drops ranging from 10%10\% to over 25%25\% on MATH-P-Hard compared to the original problems (e.g., o1-mini falls from 94.27%94.27\% to 78.49%78.49\%, and Gemini-2.0-flash-thinking falls from 92.47%92.47\% to 78.14%78.14\%), demonstrating sensitivity to shifts in underlying mathematical reasoning patterns.

  3. Knowl 3 — Conceptual Shortcut Overfitting as a Distinct Failure Mode in LLM Reasoning

    empirical result

    Detailed failure analysis of language models on hard perturbations reveals a distinct form of memorization characterized by shortcut overfitting rather than verbatim recitation. When evaluated on problems where the model successfully solves the original or simple variant but fails on the hard variant (accounting for 20%–47%20\%\text{--}47\% of total problems across models), errors frequently manifest as:

    1. Presuming Original Assumptions: The model recognizes the surface question but proceeds under the seed problem's assumptions rather than the modified constraints.
    2. Blind Technique Application: The model applies factorization, simplification, or bounding techniques that were valid for the original algebraic structure but do not apply to the modified equation.
    3. Outputting Target Quantities of Seed Problems: The model produces the desired output format of the unprovided original problem (for example, listing all valid integers instead of identifying only the smallest integer as requested in the modified prompt).

    Manual inspection reveals that this conceptual shortcut memorization accounts for approximately 40%40\% of all errors made by o1-mini and 25%25\% of all errors made by Claude-3.5-Sonnet on hard-perturbed tasks.

  4. Knowl 4 — In-Context Learning Decomposition: ICL Gain vs. Misleading Effect on Hard Perturbations

    empirical result

    Providing the original, unmodified problem and its solution as a one-shot in-context learning (ICL) demonstration yields divergent effects on simple versus hard problem perturbations:

    • On MATH-P-Simple, one-shot demonstration with the original problem-solution pair consistently increases model accuracy across nearly all architectures.
    • On MATH-P-Hard, the demonstration introduces two competing forces:
      1. ICL Effect (nwrong→correctn_{\text{wrong} \to \text{correct}}): Helpful mathematical background provided in the demonstration converts previously failed problems into correct solutions (24%–40%24\%\text{--}40\% of errors fixed in large models; 2%–15%2\%\text{--}15\% in small models).
      2. Misleading Effect (ncorrect→wrongn_{\text{correct} \to \text{wrong}}): The subtle structural differences between the demonstration and the query mislead the model into blindly copying inappropriate solution paths, causing previously correct problems to fail (18%–40%18\%\text{--}40\% of correct predictions corrupted in large models; 4%–15%4\%\text{--}15\% in small models).

    Because the misleading effect largely offsets the ICL gains, one-shot prompting with original problems produces only marginal net accuracy improvements (<5%<5\%) or net degradations on MATH-P-Hard.

  5. Knowl 5 — Low Prevalence of Verbatim Mode Collapse in Perturbed Reasoning

    empirical result

    Mode collapse—defined as a failure to detect prompt perturbations causing the model's generated output to collapse to the exact text or numerical answer of the training problem—is rare in modern LLMs:

    • The fraction of errors where a model's final numerical answer matches the ground-truth answer of the original seed problem (nsame/ntotaln_{\text{same}} / n_{\text{total}}) is below 10%10\% for most models on MATH-P-Hard, reaching at most 16.39%16.39\% for Gemini-2.0-flash-thinking-exp, 15.00%15.00\% for o1-mini, and 14.13%14.13\% for Gemini-2.0-flash-exp.
    • Across all models and problem pairs evaluated, only a single instance (in Gemma-2-9b-it) produced an identical full text response between the original and perturbed problem.
    • Normalized string edit distances between responses generated for perturbed problems and original problems remain high (averaging between 0.550.55 and 1.51.5 across most architectures).

    This indicates that failure on perturbed mathematical reasoning is driven by subtle conceptual transfer errors and incorrect reasoning adaptations rather than verbatim training text regurgitation.

  6. Knowl 6 — Taxonomy of Model Response Behaviors Across Problem Perturbation Variants

    model/method

    To analyze how models generalize across structural perturbations, model responses across the original problem, the MATH-P-Simple variant, and the MATH-P-Hard variant are categorized into four mutually exclusive cases:

    • Case I (Robust Success): The model correctly solves at least one of the easier variants (Original or MATH-P-Simple) and also correctly solves the MATH-P-Hard modification.
    • Case II (Base Incompetence): The model fails both the Original and MATH-P-Simple variants, and fails the MATH-P-Hard modification.
    • Case III (Anomalous Generalization): The model fails both easier variants (Original and MATH-P-Simple) but correctly solves the MATH-P-Hard modification (consistently <10%<10\% across models, attributable to misalignment between model capabilities and human difficulty perceptions).
    • Case IV (Hard Perturbation Vulnerability): The model correctly solves at least one easier variant (Original or MATH-P-Simple) but fails on the MATH-P-Hard modification.

    Strong reasoning models exhibit higher proportions of Case I (e.g., o1-mini: 78.14%78.14\%; Gemini-2.0-flash-thinking-exp: 75.99%75.99\%) and lower proportions of Case II (1.43%1.43\% and 1.79%1.79\% respectively), while Case IV comprises 20.07%–47.67%20.07\%\text{--}47.67\% across evaluated models.

  7. Knowl 7 — Ablation: Original Problem Demonstrations vs. Random In-Category Demonstrations

    empirical result

    In one-shot in-context learning evaluations, prompting with the specific seed problem and solution outperforms prompting with a randomly selected problem-solution pair from the same mathematical subject category, on both MATH-P-Simple and MATH-P-Hard splits.

    On MATH-P-Simple, original demonstrations achieve substantially higher accuracy than random category demonstrations across all models (e.g., Claude-3.5-Sonnet: 83.15%83.15\% vs. 62.37%62.37\%; GPT-4o: 77.06%77.06\% vs. 63.08%63.08\%; Gemini-1.5-pro: 88.17%88.17\% vs. 75.99%75.99\%; o1-mini: 94.98%94.98\% vs. 92.83%92.83\%).

    On MATH-P-Hard, despite the misleading effect induced by close lexical similarity, original demonstrations still outperform random demonstrations for almost all models (e.g., Claude-3.5-Sonnet: 49.46%49.46\% vs. 40.86%40.86\%; GPT-4o: 43.01%43.01\% vs. 37.28%37.28\%; Gemini-1.5-pro: 60.57%60.57\% vs. 51.97%51.97\%; o1-mini: 78.49%78.49\% vs. 75.99%75.99\%), indicating that relevant mathematical context and domain theorems transfer more utility than general format priming.

  8. Knowl 8 — Inference-Time Scaling and Self-Consistency on Perturbed Math Benchmarks

    empirical result

    Inference-time compute scaling via repeated independent sampling (N=64N=64 for Llama-3.1-8B-Instruct and Qwen2.5-Math-7B-Instruct; N=8N=8 for o1-mini) improves performance across Original, MATH-P-Simple, and MATH-P-Hard benchmarks under both pass@k\text{pass}@k and self-consistency (majority voting) metrics.

    Given NN independent candidate solutions where cc answers are correct, the unbiased estimator for pass@k\text{pass}@k across problems is defined as: pass@k=Eproblem[1−(N−ck)(Nk)]\text{pass}@k = \mathbb{E}_{\text{problem}} \left[ 1 - \frac{\binom{N-c}{k}}{\binom{N}{k}} \right]

    While increasing kk consistently monotonically improves accuracy for all three splits, the absolute performance deficit between MATH-P-Hard and the easier variants (Original and MATH-P-Simple) remains persistent as compute scales, indicating that inference-time search alone does not fully bridge the out-of-distribution reasoning gap introduced by hard perturbations.

Coverage note — Omitted individual hyperparameter configs for standard model APIs, general dataset licenses, and qualitative prompt formatting strings from the appendices as they represent routine evaluation setup rather than primary scientific findings.

References

  1. 1.Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A. A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., Behl, H., Benhaim, A., Bilenko, M., Bjorck, J., Bubeck, S., Cai, M., Cai, Q., Chaudhary, V., Chen, D., Chen, D., Chen, W., Chen, Y.-C., Chen, Y.-L., Cheng, H., Chopra, P., Dai, X., Dixon, M., Eldan, R., Fragoso, V., Gao, J., Gao, M., Gao, M., Garg, A., Giorno, A. D., Goswami, A., Gunasekar, S., Haider, E., Hao, J., Hewett, R. J., Hu, W., Huynh, J., Iter, D., Jacobs, S. A., Javaheripi, M., Jin, X., Karampatziakis, N., Kauffmann, P., Khademi, M., Kim, D., Kim, Y. J., Kurilenko, L., Lee, J. R., Lee, Y. T., Li, Y., Li, Y., Liang, C., Liden, L., Lin, X., Lin, Z., Liu, C., Liu, L., Liu, M., Liu, W., Liu, X., Luo, C., Madan, P., Mahmoudzadeh, A., Majercak, D., Mazzola, M., Mendes, C. C. T., Mitra, A., Modi, H., Nguyen, A., Norick, B., Patra, B., Perez-Becker, D., Portet, T., Pryzant, R., Qin, H., Radmilac, M., Ren, L., de Rosa, G., Rosset, C., Roy, S., Ruwase, O., Saarikivi, O., Saied, A., Salim, A., Santacroce, M., Shah, S., Shang, N., Sharma, H., Shen, Y., Shukla, S., Song, X., Tanaka, M., Tupini, A., Vaddamanu, P., Wang, C., Wang, G., Wang, L., Wang, S., Wang, X., Wang, Y., Ward, R., Wen, W., Witte, P., Wu, H., Wu, X., Wyatt, M., Xiao, B., Xu, C., Xu, J., Xu, W., Xue, J., Yadav, S., Yang, F., Yang, J., Yang, Y., Yang, Z., Yu, D., Yuan, L., Zhang, C., Zhang, C., Zhang, J., Zhang, L. L., Zhang, Y., Zhang, Y., Zhang, Y., and Zhou, X. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv.org/abs/2404.14219.
  2. 2.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  3. 3.Anthropic. Claude-3-5-sonnet. 2024. URL https://www.anthropic.com/news/claude-3-5-sonnet.
  4. 4.Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pp. 2397–2430. PMLR, 2023.
  5. 5.Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787, 2024.
  6. 6.Brown, H., Lee, K., Mireshghallah, F., Shokri, R., and Tramèr, F. What does it mean for a language model to preserve privacy? In Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, pp. 2280–2292, 2022.
  7. 7.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  8. 8.Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pp. 2633–2650, 2021.
  9. 9.Carlini, N., Jagielski, M., Zhang, C., Papernot, N., Terzis, A., and Tramer, F. The privacy onion effect: Memorization is relative. Advances in Neural Information Processing Systems, 35:13263–13276, 2022.
  10. 10.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code. 2021.
  11. 11.Chen, T., Asai, A., Mireshghallah, N., Min, S., Grimmelmann, J., Choi, Y., Hajishirzi, H., Zettlemoyer, L., and Koh, P. W. Copybench: Measuring literal and non-literal reproduction of copyright-protected text in language model generation. arXiv preprint arXiv:2407.07087, 2024.
  12. 12.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021a.
  13. 13.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021b.
  14. 14.DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Bao, H., Xu, H., Wang, H., Ding, H., Xin, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Wang, J., Chen, J., Yuan, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Ye, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Zhao, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z. DeepSeek-R1: Incentivizing reasoning capability in llms via reinforcement learning, 2025. URL https://arxiv.org/abs/2501.12948.
  15. 15.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  16. 16.Feldman, V. Does learning require memorization? a short tale about a long tail. In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, pp. 954–959, 2020.
  17. 17.Feldman, V. and Zhang, C. What neural networks memorize and why: Discovering the long tail via influence estimation. Advances in Neural Information Processing Systems, 33:2881–2891, 2020.
  18. 18.Gulati, A., Miranda, B., Chen, E., Xia, E., Fronsdal, K., de Moraes Dumont, B., and Koyejo, S. Putnam-AXIOM: A functional and static benchmark for measuring higher level mathematical reasoning. In The 4th Workshop on Mathematical Reasoning and AI at NeurIPS’24, 2024. URL https://openreview.net/forum?id=YXnwlZe0yf.
  19. 19.He, C., Luo, R., Bai, Y., Hu, S., Thai, Z. L., Shen, J., Hu, J., Han, X., Huang, Y., Zhang, Y., Liu, J., Qi, L., Liu, Z., and Sun, M. OlympiadBench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.
  20. 20.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the math dataset. NeurIPS, 2021.
  21. 21.Hernandez, D., Brown, T., Conerly, T., DasSarma, N., Drain, D., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Henighan, T., Hume, T., et al. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022.
  22. 22.Huang, Y., Gupta, S., Zhong, Z., Li, K., and Chen, D. Privacy implications of retrieval-based language models. arXiv preprint arXiv:2305.14888, 2023.
  23. 23.jylin04, JackS, Karvonen, A., and Can. Othellogpt learned a bag of heuristics. https://www.lesswrong.com/posts/gcpNuEZnxAPayaKBY/othellogpt-learned-a-bag-of-heuristics-1. Accessed on Date (2025-01-28).
  24. 24.Karamolegkou, A., Li, J., Zhou, L., and Søgaard, A. Copyright violations and large language models. arXiv preprint arXiv:2310.13771, 2023.
  25. 25.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  26. 26.Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021.
  27. 27.Li, J., Beeching, E., Tunstall, L., Lipkin, B., Soletskyi, R., Huang, S. C., Rasul, K., Yu, L., Jiang, A., Shen, Z., Qin, Z., Dong, B., Zhou, L., Fleureau, Y., Lample, G., and Polu, S. Numinamath. https://github.com/project-numina/aimo-progress-prize, 2024a.
  28. 28.Li, Q., Cui, L., Zhao, X., Kong, L., and Bi, W. Gsm-plus: A comprehensive benchmark for evaluating the robustness of llms as mathematical problem solvers. arXiv preprint arXiv:2402.19255, 2024b.
  29. 29.Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  30. 30.Liu, J., Xia, C. S., Wang, Y., and Zhang, L. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems, 36, 2024.
  31. 31.Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M. Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229, 2024.
  32. 32.Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K. Scalable extraction of training data from (production) language models. arXiv preprint arXiv:2311.17035, 2023.
  33. 33.Nikankin, Y., Reusch, A., Mueller, A., and Belinkov, Y. Arithmetic without algorithms: Language models solve math with a bag of heuristics. arXiv preprint arXiv:2410.21272, 2024.
  34. 34.OpenAI. OpenAI o1. 2024. URL https://openai.com/index/openai-o1-system-card/.
  35. 35.Patel, A., Bhattamishra, S., and Goyal, N. Are NLP models really able to solve simple math word problems? In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2080–2094, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.168. URL https://aclanthology.org/2021.naacl-main.168/.
  36. 36.Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022, 2023.
  37. 37.Shah, V., Yu, D., Lyu, K., Park, S., Yu, J., He, Y., Ke, N. R., Mozer, M., Bengio, Y., Arora, S., et al. Ai-assisted generation of difficult math questions. arXiv preprint arXiv:2407.21009, 2024.
  38. 38.Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  39. 39.Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E. H., Schärli, N., and Zhou, D. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pp. 31210–31227. PMLR, 2023a.
  40. 40.Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L. Detecting pretraining data from large language models. arXiv preprint arXiv:2310.16789, 2023b.
  41. 41.Srivastava, S., PV, A., Menon, S., Sukumar, A., Philipose, A., Prince, S., Thomas, S., et al. Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap. arXiv preprint arXiv:2402.19450, 2024.
  42. 42.Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024a.
  43. 43.Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024b.
  44. 44.Team, Q. QwQ: Reflect deeply on the boundaries of the unknown, November 2024. URL https://qwenlm.github.io/blog/qwq-32b-preview/.
  45. 45.Tirumala, K., Markosyan, A., Zettlemoyer, L., and Aghajanyan, A. Memorization without overfitting: Analyzing the training dynamics of large language models. Advances in Neural Information Processing Systems, 35:38274–38290, 2022.
  46. 46.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  47. 47.Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. arXiv preprint arXiv:2406.01574, 2024.
  48. 48.Wei, B., Shi, W., Huang, Y., Smith, N. A., Zhang, C., Zettlemoyer, L., Li, K., and Henderson, P. Evaluating copyright takedown methods for language models. arXiv preprint arXiv:2406.18664, 2024.
  49. 49.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  50. 50.Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y. An empirical analysis of compute-optimal inference for problem-solving with language models. 2024.
  51. 51.Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y. Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks. arXiv preprint arXiv:2307.02477, 2023.
  52. 52.Xie, C., Huang, Y., Zhang, C., Yu, D., Chen, X., Lin, B. Y., Li, B., Ghazi, B., and Kumar, R. On memorization of large language models in logical reasoning. arXiv preprint arXiv:2410.23123, 2024.
  53. 53.Yan, F., Mao, H., Ji, C. C.-J., Zhang, T., Patil, S. G., Stoica, I., and Gonzalez, J. E. Berkeley function calling leaderboard. https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html, 2024.
  54. 54.Yang, A., Zhang, B., Hui, B., Gao, B., Yu, B., Li, C., Liu, D., Tu, J., Zhou, J., Lin, J., et al. Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement. arXiv preprint arXiv:2409.12122, 2024.
  55. 55.Yu, L., Jiang, W., Shi, H., YU, J., Liu, Z., Zhang, Y., Kwok, J., Li, Z., Weller, A., and Liu, W. Metamath: Bootstrap your own mathematical questions for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=N8N0hgNDRt.
  56. 56.Yue, X., Zheng, T., Zhang, G., and Chen, W. Mammoth2: Scaling instructions from the web. Advances in Neural Information Processing Systems, 2024.
  57. 57.Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. Understanding deep learning (still) requires rethinking generalization. Communications of the ACM, 64(3):107–115, 2021.
  58. 58.Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., and Carlini, N. Counterfactual memorization in neural language models. Advances in Neural Information Processing Systems, 36:39321–39362, 2023.
  59. 59.Zhang, H., Da, J., Lee, D., Robinson, V., Wu, C., Song, W., Zhao, T., Raja, P., Slack, D., Lyu, Q., et al. A careful examination of large language model performance on grade school arithmetic. arXiv preprint arXiv:2405.00332, 2024.
  60. 60.Zheng, C., Zhou, H., Meng, F., Zhou, J., and Huang, M. On large language models’ selection bias in multi-choice questions. arXiv preprint arXiv:2309.03882, 2023.
  61. 61.Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., and Hou, L. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911, 2023.
  62. 62.Zhou, Z., Liu, S., Ning, M., Liu, W., Wang, J., Wong, D. F., Huang, X., Wang, Q., and Huang, K. Is your model really a good math reasoner? evaluating mathematical reasoning with checklist. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=nDvgHIBRxQ.
  63. 63.Zou, C., Guo, X., Yang, R., Zhang, J., Hu, B., and Zhang, H. Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models. arXiv preprint arXiv:2411.00836, 2024.

Citation

MLA
Huang, K., et al. “MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities Against Hard Perturbations”. arXiv, 2025, http://arxiv.org/abs/2502.06453v2.
APA
Huang, K., Guo, J., Li, Z., Ji, X., Ge, J., Li, W., Guo, Y., Cai, T., Yuan, H., Wang, R., Wu, Y., Yin, M., Tang, S., Huang, Y., Jin, C., Chen, X., Zhang, C., & Wang, M. (2025). MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations. arXiv. http://arxiv.org/abs/2502.06453v2
Chicago
Huang, K., J. Guo, Z. Li, et al. 2025. “MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities Against Hard Perturbations”. arXiv. http://arxiv.org/abs/2502.06453v2.
Harvard
Huang, K. et al. (2025) “MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.06453v2.
Vancouver
1. Huang K, Guo J, Li Z, et al (2025) MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations. arXiv

BibTeX

@article{huang2025math,
  title = {MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations},
  author = {Huang, Kaixuan and Guo, Jiacheng and Li, Zihao and Ji, Xiang and Ge, Jiawei and Li, Wenzhe and Guo, Yingqing and Cai, Tianle and Yuan, Hui and Wang, Runzhe and Wu, Yue and Yin, Ming and Tang, Shange and Huang, Yangsibo and Jin, Chi and Chen, Xinyun and Zhang, Chiyuan and Wang, Mengdi},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.06453v2},
  eprint = {2502.06453}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/