MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning

Chengpeng LiZheng YuanHongyi YuanGuanting DongKeming LuJiancan WuChuanqi TanXiang WangChang Zhou

article2024ACL55 citations

Presents a systematic investigation into query and response augmentation for mathematical reasoning, establishing state-of-the-art open-source models while identifying quantitative scaling laws and critical out-of-domain generalization limits between GSM8K and MATH.

Listen

Large language models often struggle with complex multi-step mathematical reasoning. While leading proprietary models exhibit strong capabilities, open-source alternatives lag behind. Fine-tuning models on synthetic datasets created through automated data augmentation offers a promising path to bridge this gap, yet organizations lack clear insight into which augmentation strategies work best, how model performance scales with data volume, and whether improvements generalize beyond specific training domains.

The article evaluates data augmentation techniques across query rephrasing and multi-path reasoning to determine optimal training recipes, quantify performance scaling laws, and measure transferability across distinct mathematical benchmarks.

To conduct this evaluation, the authors generated synthetic datasets, AugGSM8K and AugMATH, using proprietary language models to expand standard elementary and high-school math benchmarks. They applied five query modification techniques—such as increasing complexity, introducing fractions or percentages, and combining concepts—alongside multi-path chain-of-thought response generation. They then fine-tuned open-source LLaMA family models ranging from 7 billion to 70 billion parameters, producing a specialized series dubbed MuggleMath.

The investigation yielded several critical findings. First, MuggleMath achieved new state-of-the-art results among open-source models, improving 70-billion-parameter performance on GSM8K to 82.7 percent and on MATH to 36.3 percent, up from baseline scores of 63.2 percent and 14.4 percent, respectively. Second, performance followed a predictable log-linear scaling curve with data volume on grade-school math and a segmented log-linear relationship on advanced math, matching the sample efficiency of human-written data. Third, increasing problem complexity proved to be the single most effective individual query strategy, while mixing diverse augmentation strategies achieved the highest overall performance as data scaled. Fourth, augmenting harder and previously failed problems yielded substantially larger accuracy improvements than augmenting easier questions. Finally, performance gains failed to generalize across domains: training extensively on elementary word problems provided almost no benefit on advanced multi-topic mathematics, and vice versa.

These findings indicate that targeted synthetic data generation is highly cost-effective for matching the performance of expensive human curation within a given problem type. However, leaders should recognize that high performance on narrow mathematical benchmarks reflects domain-specific task mastery rather than generalized reasoning capability. Embedding space analyses confirm that synthetic variations remain clustered tightly around their seed distributions without expanding model breadth.

Organizations developing reasoning models should prioritize creating diverse seed datasets spanning all target domains rather than over-scaling synthetic variations of a single domain. Practical implementation pipelines should focus synthetic generation on complex edge cases and failed problems to maximize training efficiency. Future initiatives should establish robust multi-domain data generation frameworks and evaluate broader out-of-distribution transfer before deploying fine-tuned models into varied production environments.

Confidence in these findings is high regarding in-domain benchmark performance and predictable scaling trends. However, practitioners should note limitations: data generation relied on proprietary model prompts that may vary across revisions, and the evaluation was confined strictly to mathematical reasoning benchmarks rather than open-ended enterprise reasoning tasks.

Cover for MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning

Abstract

In math reasoning with large language models (LLMs), fine-tuning data augmentation by query evolution and diverse reasoning paths is empirically verified effective, profoundly narrowing the gap between open-sourced LLMs and cutting-edge proprietary LLMs. In this paper, we conduct an investigation for such data augmentation in math reasoning and are intended to answer: (1) What strategies of data augmentation are more effective; (2) What is the scaling relationship between the amount of augmented data and model performance; and (3) Can data augmentation incentivize generalization to out-of-domain mathematical reasoning tasks? To this end, we create two new dataset AugGSM8K and AugMATH, by complicating and diversifying the queries and sampling multiple reasoning paths from GSM8K and MATH. We obtained a series of LLMs called MuggleMath by fine-tuning LLaMA models on AugGSM8K and AugMATH. MuggleMath substantially achieves new state-of-the-art on GSM8K and MATH. A log-linear relationship and a segmented log-linear are presented between MuggleMath’s performance and the amount of augmented data on GSM8K and MATH, respectively. We also find that it is weak in out-of-domain math reasoning generalization from AugGSM8K to MATH and from AugMATH to GSM8K, which suggests that augmenting queries that cover a broader range of subjects is more beneficial for generalization.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Dataset Augmentation
  • 3.3 RQ1. What Strategies of Data Augmentation are More Effective
  • 3.4 RQ2. What is the Scaling Relationship between the Amount of Augmented Data and Model Performance
  • 3.5 RQ3. Can Data Augmentation Incentivize Generalization to Out-of-domain Mathematical Reasoning Tasks?
  • 4 Discussion
  • 4.1 Training set vs. Test set accuracy
  • 4.2 Make more augmentation on harder problems
  • 4.3 Result on the Perturbed Test Set
  • 5 Conclusion
  • Limitations
  • 6 Acknowledgement
  • References
  • A Augmentation on MATH: AugMATH
  • A.1 Query and response augmentation on MATH
  • A.2 Relationship of performance and augmented data size
  • A.3 Generalizability of AugMATH to GSM8K
  • A.4 Generalizability on other OOD datasets
  • B Instruction prompt for training and inference
  • C Query augmentation prompt for GSM8K
  • D Response augmentation prompt for GSM8K
  • E Query and response augmentation prompt for MATH
  • F Response filter
  • G Case Study of GSM8K
  • H Difficulty level definition on GSM8K
  • I Discussion of augmentation methods for GSM8K
  • I.1 How to improve existing data augmentation methods?
  • I.2 Will mixed-augmentation has advantages over increase the complexity if we enlarge the size of data?
  • J The effect of augmentation response quality of GSM8K
  • J.1 Under what circumstances can wrong answers have a positive effect?
  • J.2 The relationship between response quality and data volume.
  • J.3 The Accuracy and difficulty of different query augmentation categories
  • K A more detailed comparison of GSM8K, MATH, AugGSM8K, and AugMATH
  • L Detailed Experimental results

Knowls

  1. Knowl 1 — MuggleMath raises in-domain accuracy on GSM8K and MATH

    empirical result

    The authors fine-tuned LLaMA-1 7B and LLaMA-2 7B, 13B, and 70B models on augmented mathematical reasoning data, calling the resulting models MuggleMath. Fine-tuning used AdamW for three epochs, learning rate 2×10−52\times10^{-5}, warmup ratio 0.03, cosine learning-rate scheduling, and full-model updates; evaluation used greedy decoding. In model order LLaMA-1 7B, LLaMA-2 7B, LLaMA-2 13B, and LLaMA-2 70B, GSM8K accuracy (%) was 35.9, 41.6, 50.0, and 63.2 for the original-data baselines, versus 66.0, 69.8, 74.3, and 82.7 for MuggleMath. On MATH, the corresponding baseline scores were 4.8, 5.8, 6.0, and 14.4, while MuggleMath scored 18.4, 25.8, 30.7, and 36.3. For the LLaMA-2 sizes, MuggleMath also exceeded the reported WizardMath scores: on GSM8K, 69.8/74.3/82.7 versus 54.9/63.9/81.6; on MATH, 25.8/30.7/36.3 versus 10.7/14.0/22.7.

  2. Knowl 2 — Construction of AugGSM8K and AugMATH

    model/method

    AugGSM8K was built from the 7,473 GSM8K training problems by generating five new queries per original problem, or 37,365 queries per augmentation run. Three query-generation runs used GPT-3.5 for two runs and GPT-4 for one; one response was generated for each query, and manual filtering left approximately 30,000 query-response pairs per run. The GSM8K query variations were designed to change specific numbers, introduce fractions or percentages, combine concepts, add conditional statements, or increase problem complexity. Responses were generated with GPT-3.5 or GPT-4 using structured prompting; filtering removed responses without a final answer, responses longer than 1,500 characters, and extraneous text. AugMATH was built from 7,500 MATH training problems: each problem was varied in five ways and each variation was sampled eight times with GPT-4 at temperature 0.7, yielding 300,000 synthetic examples. Its variations were allowed to extend beyond the five GSM8K templates to reflect MATH's wider range of subjects. Exact generated examples may not be reproducible from the same prompts, although the authors report similar performance improvements when generation temperature is varied.

  3. Knowl 3 — Query complexity and diversity affect augmentation performance

    empirical result

    On GSM8K, each of five query-augmentation strategies improved over training on the original set alone. In the order LLaMA-1 7B, LLaMA-2 7B, and LLaMA-2 13B, baseline accuracy (%) was 35.9/41.6/50.0; changing numbers gave 41.5/48.5/54.1; adding fractions or percentages gave 41.2/46.2/54.4; combining concepts gave 41.1/47.5/56.1; adding a conditional statement gave 41.7/45.8/56.4; and increasing complexity gave 42.4/48.6/57.6. A mixed set containing examples from different strategies scored 42.2/48.2/57.2 at the smaller tested size. Increasing complexity was generally strongest among individual strategies, while at 19,000 augmented examples the mixed strategy scored 55.3% for LLaMA-2 7B, above 54.0% for complexity-only; this supports combining query types when scaling the data. The source of the query generator (GPT-3.5 versus GPT-4) and GPT-4 response temperature (0 versus 1) had no reported significant effect. In contrast, one-shot structured response prompting outperformed zero-shot prompting by 3.6, 3.7, and 3.3 percentage points for LLaMA-1 7B, LLaMA-2 7B, and LLaMA-2 13B, respectively; GPT-4 responses also outperformed GPT-3.5 responses in a roughly size-matched comparison.

  4. Knowl 4 — GSM8K accuracy scales log-linearly with augmented query volume

    empirical result

    For GSM8K, the authors fitted accuracy as a function of augmented-query volume over the tested range of approximately 13,000–97,000 examples. Let xx be query volume in thousands and yy be GSM8K test accuracy in percentage points. The fitted relationships were y=10.7log⁡(x)+13.2y=10.7\log(x)+13.2 for LLaMA-1 7B, y=9.8log⁡(x)+21.3y=9.8\log(x)+21.3 for LLaMA-2 7B, and y=7.6log⁡(x)+36.3y=7.6\log(x)+36.3 for LLaMA-2 13B. The fits predicted intermediate and extrapolated results closely: at x=17x=17, predictions/observations were 43.4/43.4, 49.2/50.0, and 57.7/56.0; at x=104x=104, they were 62.7/62.1, 67.0/66.7, and 71.4/70.8, respectively. The fitted gains from doubling query volume were 7.4, 6.8, and 5.3 percentage points for those models, similar to the paper's cited estimates for doubling human-written data (6.5, 6.6, and 5.5 points). The authors caution that this log-linear fit cannot hold across all data sizes because accuracy is bounded.

  5. Knowl 5 — MATH shows a segmented log-linear scaling pattern

    empirical result

    For LLaMA-2 7B fine-tuned on the original MATH training set plus varying amounts of AugMATH, performance on the MATH test set followed two separately fitted log-linear regimes. Let xx be the augmented-data volume in thousands and yy be test accuracy in percentage points. In the 7.5–37.5 thousand range, the fit was y=1.47+2.45log⁡(x)y=1.47+2.45\log(x), with coefficient of determination R2=0.992R^2=0.992; in the 82.5–307.5 thousand range, it was y=−13.56+6.33log⁡(x)y=-13.56+6.33\log(x), with R2=0.985R^2=0.985. The reported accuracies at augmented-data volumes 7.5k, 15k, 22.5k, 30k, 37.5k, 82.5k, 157.5k, 232.5k, and 307.5k were 6.5%, 8.0%, 9.0%, 9.8%, 10.6%, 14.4%, 18.7%, 20.3%, and 23.1%, respectively; training on MATH without added AugMATH scored 2.5%. The separated fits indicate that gains per data doubling were larger in the high-volume regime than in the low-volume regime.

  6. Knowl 6 — More reasoning responses help until model-dependent saturation

    empirical result

    With GSM8K augmented queries held fixed, the authors increased the number of generated responses per query from one to five. In the order LLaMA-1 7B, LLaMA-2 7B, and LLaMA-2 13B, accuracy (%) was 53.0/57.0/65.5 with one response, 55.9/61.4/67.0 with two, 61.3/64.4/68.4 with three, 60.1/63.8/69.1 with four, and 60.7/63.6/71.6 with five. Thus, the 7B models stopped improving consistently beyond three responses per query, while the 13B model continued to improve over the tested range of 37,000–157,000 examples. Majority-vote filtering, which discarded queries whose sampled responses all gave different answers, generally reduced performance at comparable data scales: with three responses, filtered scores were 56.4/60.7/65.7; with five, they were 58.9/62.5/68.3. The authors suggest that filtering may remove useful reasoning traces, reduce the number of retained queries, or both.

  7. Knowl 7 — Augmentation transfers weakly between GSM8K and MATH

    empirical result

    Augmentation primarily improved performance on the source distribution rather than transferring to the other benchmark. In a transfer experiment evaluated on a 500-question MATH test sample, LLaMA-2 7B and 13B trained on an augmented GSM8K subset and then fine-tuned on MATH scored 5.6% and 7.8%; a larger GSM8K-augmentation mixture followed by MATH fine-tuning scored 8.4% and 9.4%. Results varied by training setting and model, and the authors characterize the overall benefit from GSM8K augmentation to MATH as marginal. In the reverse direction, LLaMA-2 7B's GSM8K score rose from 40.3% with MATH training alone to 45.4% with the full 300,000-example AugMATH set, whereas 30,000 AugGSM8K examples raised it to 58%. A t-SNE visualization of LLaMA-2 7B representations from the 15th layer's final problem token showed AugGSM8K problems close to GSM8K problems, while MATH problems were largely separated. The authors propose this distributional difference, alongside differences in difficulty, response style, and mathematical content, as an explanation for weak transfer.

  8. Knowl 8 — Targeting difficult or previously failed problems improves augmentation yield

    empirical result

    For GSM8K, the authors categorized source problems by the number of formulas in their reasoning: fewer than three was easy, exactly three was medium, and more than three was hard. Fine-tuning on queries augmented from hard problems yielded test accuracy of 43.0%, 51.3%, and 58.8% for LLaMA-1 7B, LLaMA-2 7B, and LLaMA-2 13B, respectively; random-problem augmentation yielded 43.4%, 50.0%, and 56.0%. Augmenting problems that the SFT model had answered incorrectly yielded 46.2%, 49.5%, and 55.4%, compared with 43.6%, 49.2%, and 54.2% for random augmentation. The hard-problem advantage was clearest for the LLaMA-2 models, and wrong-problem augmentation exceeded random augmentation for all three. On test problems grouped by difficulty, the LLaMA-2 7B model scored 55%, 42%, and 21% on easy, medium, and hard questions; MuggleMath scored 73%, 70%, and 64% on those groups.

  9. Knowl 9 — Incorrect synthetic reasoning can help weaker models but harm stronger ones

    empirical result

    The authors formed three response datasets containing the same 20,500 queries: responses estimated to be correct by majority voting, responses estimated to be incorrect, and a 50/50 mixture. On LLaMA-2 7B-SFT, whose GSM8K test accuracy before this experiment was 41.6%, fine-tuning on the all-correct, all-incorrect, and mixed datasets yielded 58.7%, 46.1%, and 52.4%, respectively. Thus, even the dataset estimated to contain only incorrect final answers improved this weaker model, which the authors suggest may reflect useful correct intermediate reasoning. On the stronger MuggleMath-7B model, the same datasets yielded 67.0%, 54.7%, and 57.0%, compared with its reported 68.4% starting accuracy; lower estimated response quality was associated with greater degradation. Correctness here is an estimate based on agreement among sampled answers, not an independently verified label.

  10. Knowl 10 — Augmented models perform better on perturbed and related math benchmarks

    empirical result

    On two perturbed GSM8K test sets, MuggleMath trained on AugGSM8K substantially outperformed models fine-tuned on the original GSM8K training set. On Change-Test, created by changing numbers and corresponding answers in 1,211 questions, the original-data SFT scores for LLaMA-1 7B, LLaMA-2 7B, and LLaMA-2 13B were 26.2%, 30.1%, and 38.6%; MuggleMath scored 60.1%, 62.8%, and 67.1%. On Aug-Test, containing 1,378 test questions varied using the training augmentation approach, the respective SFT scores were 14.2%, 17.2%, and 22.4%, versus 40.1%, 44.3%, and 49.3% for MuggleMath. The authors also evaluated LLaMA-2 7B MuggleMath variants on five additional datasets. Scores for AugGSM8K-trained versus AugMATH-trained variants were GSM-Hard 36.5/17.1, SVAMP 70.9/53.0, TabMWP 35.6/41.9, ASDiv 69.5/59.7, and MAWPS 89.8/68.5. They caution that these datasets may be relatively close to GSM8K rather than representing strongly out-of-distribution mathematical reasoning.

Coverage note — Detailed per-subject transfer tables, training-versus-test accuracy correlations, and individual reasoning case studies are omitted because they provide supplementary breakdowns rather than additional main findings.

References

  1. 1.Alibaba. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  2. 2.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  3. 3.BaichuanInc. 2023. Baichuan 2. technical report. arXiv preprint arXiv:2309.16609.
  4. 4.Tiffany Tianhui Cai, Hongseok Namkoong, and Steve Yadlowsky. 2023. Diagnosing model performance under distribution shift.
  5. 5.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  7. 7.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  8. 8.Elliot Creager, Jörn-Henrik Jacobsen, and Richard S. Zemel. 2021. Environment inference for invariant learning. pages 2189–2200. PMLR.
  9. 9.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233.
  10. 10.Guanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024. How abilities in large language models are affected by supervised fine-tuning data composition.
  11. 11.Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021. A survey of data augmentation approaches for NLP. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 968–988, Online. Association for Computational Linguistics.
  12. 12.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. PAL: program-aided language models. In ICML 2023.
  13. 13.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021a. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874.
  14. 14.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021b. Measuring mathematical problem solving with the MATH dataset.
  15. 15.Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022. Large language models can self-improve.
  16. 16.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018. Progressive growing of gans for improved quality, stability, and variation. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings.
  17. 17.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1152–1157, San Diego, California. Association for Computational Linguistics.
  18. 18.Kun Kuang, Ruoxuan Xiong, Peng Cui, Susan Athey, and Bo Li. 2020. Stable prediction with model misspecification and agnostic distribution shift. pages 4485–4492. AAAI Press.
  19. 19.Ariel N. Lee, Cole J. Hunter, and Nataniel Ruiz. 2023. Platypus: Quick, cheap, and powerful refinement of llms.
  20. 20.Shanglin Lei, Guanting Dong, Xiaoping Wang, Keheng Wang, and Sirui Wang. 2024. Instructerc: Reforming emotion recognition in conversation with a retrieval multi-task llms framework.
  21. 21.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022a. Solving quantitative reasoning problems with language models.
  22. 22.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022b. Solving quantitative reasoning problems with language models. In NeurIPS.
  23. 23.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023a. Let’s verify step by step. arXiv preprint arXiv:2305.20050.
  24. 24.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023b. Let’s verify step by step.
  25. 25.Zachary C. Lipton, Yu-Xiang Wang, and Alexander J. Smola. 2018. Detecting and correcting for label shift with black box predictors. pages 3128–3136. PMLR.
  26. 26.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. 2023. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688.
  27. 27.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. 2023. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. In ICLR.
  28. 28.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023a. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct.
  29. 29.Haoran Luo, Haihong E, Zichen Tang, Shiyao Peng, Yikai Guo, Wentai Zhang, Chenghao Ma, Guanting Dong, Meina Song, Wei Lin, Yifan Zhu, and Luu Anh Tuan. 2024. Chatkbqa: A generate-then-retrieve framework for knowledge base question answering with fine-tuned large language models.
  30. 30.Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023b. Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568.
  31. 31.Yuzhou Mao, Liu Yu, Yi Yang, Fan Zhou, and Ting Zhong. 2023. Debiasing intrinsic bias and application bias jointly via invariant risk minimization (student abstract). AAAI Press.
  32. 32.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020. A diverse corpus for evaluating and developing english math word problem solvers. In ACL.
  33. 33.Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. 2023. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707.
  34. 34.Ansong Ni, Jeevana Priya Inala, Chenglong Wang, Alex Polozov, Christopher Meek, Dragomir Radev, and Jianfeng Gao. 2023. Learning math reasoning from self-sampled correct and partially-correct solutions. In The Eleventh International Conference on Learning Representations.
  35. 35.OpenAI. 2023. Gpt-4 technical report.
  36. 36.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  37. 37.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In NAACL-HLT.
  38. 38.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon LLM: outperforming curated corpora with web data, and web data only.
  39. 39.Ru Peng, Qiuyang Duan, Haobo Wang, Jiachen Ma, Yanbo Jiang, Yongjun Tu, Xiu Jiang, and Junbo Zhao. 2023. Came: Contrastive automated model evaluation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20121–20132.
  40. 40.Ru Peng, Heming Zou, Haobo Wang, Yawen Zeng, Zenan Huang, and Junbo Zhao. 2024. Energy-based automated model evaluation. arXiv preprint arXiv:2401.12689.
  41. 41.Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen. 2015. Causal inference using invariant prediction: identification and confidence intervals.
  42. 42.Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023. Code llama: Open foundation models for code.
  43. 43.Bernhard Schölkopf, Dominik Janzing, Jonas Peters, Eleni Sgouritsa, Kun Zhang, and Joris M. Mooij. 2012. On causal and anticausal learning.
  44. 44.Zheyan Shen, Peng Cui, Tong Zhang, and Kun Kuang. 2020. Stable learning via sample reweighting. pages 5692–5699. AAAI Press.
  45. 45.Xiaoshuai Song, Keqing He, Pei Wang, Guanting Dong, Yutao Mou, Jingang Wang, Yunsen Xian, Xunliang Cai, and Weiran Xu. 2023. Large language models meet open-world intent discovery and recognition: An evaluation of chatgpt.
  46. 46.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  47. 47.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science.
  48. 48.InternLM Team. 2023a. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM.
  49. 49.MosaicML NLP Team. 2023b. Introducing mpt-7b: A new standard for open-source, commercially usable llms. Accessed: 2023-05-05.
  50. 50.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models.
  51. 51.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  52. 52.Dustin Tran, Jeremiah Z. Liu, Michael W. Dusenberry, Du Phan, Mark Collier, Jie Ren, Kehang Han, Zi Wang, Zelda Mariet, Huiyi Hu, Neil Band, Tim G. J. Rudner, Karan Singhal, Zachary Nado, Joost van Amersfoort, Andreas Kirsch, Rodolphe Jenatton, Nithum Thain, Honglin Yuan, Kelly Buchanan, Kevin Murphy, D. Sculley, Yarin Gal, Zoubin Ghahramani, Jasper Snoek, and Balaji Lakshminarayanan. 2022. Plex: Towards reliability using pretrained large model extensions.
  53. 53.Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek, Yuhuai Wu, Henryk Michalewski, and Piotr Miłos. 2023. Focused transformer: Contrastive training for context scaling. arXiv preprint arXiv:2307.03170.
  54. 54.Z. Lin Y. Sheng Z. Wu H. Zhang L. Zheng S. Zhuang Y. Zhuang J. Gonzalez I. Stoica W. Chiang, Z. Li and E. Xing. 2023. icuna: An open-source chatbot impressing gpt-4 with 90Technical Report.
  55. 55.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  56. 56.Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, and Tao Qin. 2021. Generalizing to unseen domains: A survey on domain generalization. pages 4627–4635.
  57. 57.Pei Wang, Yejie Wang, Muxi Diao, Keqing He, Guanting Dong, and Weiran Xu. 2024a. Multi-perspective consistency enhances confidence estimation in large language models. arXiv preprint arXiv:2402.11279.
  58. 58.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023a. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  59. 59.Yejie Wang, Keqing He, Guanting Dong, Pei Wang, Weihao Zeng, Muxi Diao, Yutao Mou, Mengdi Zhang, Jingang Wang, Xunliang Cai, and Weiran Xu. 2024b. Dolphcoder: Echo-locating code large language models with diverse and multi-objective instruction tuning.
  60. 60.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023b. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13484–13508, Toronto, Canada. Association for Computational Linguistics.
  61. 61.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  62. 62.Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2023. Magicoder: Source code is all you need.
  63. 63.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions.
  64. 64.Mingfeng Xue, Dayiheng Liu, Kexin Yang, Guanting Dong, Wenqiang Lei, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2023. Occuquest: Mitigating occupational bias for inclusive large language models.
  65. 65.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models.
  66. 66.Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Keming Lu, Chuanqi Tan, Chang Zhou, and Jingren Zhou. 2023a. Scaling relationship on learning mathematical reasoning with large language models.
  67. 67.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, and Songfang Huang. 2023b. How well do large language models perform in arithmetic tasks? arXiv preprint arXiv:2304.02015.
  68. 68.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023a. Mammoth: Building math generalist models through hybrid instruction.
  69. 69.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023b. MAmmoTH: Building math generalist models through hybrid instruction tuning. arXiv preprint arXiv:2309.05653.
  70. 70.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022. STar: Bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems.
  71. 71.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  72. 72.Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023a. Instruction-following evaluation for large language models.
  73. 73.Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy. 2023b. Domain generalization: A survey. IEEE Trans. Pattern Anal. Mach. Intell., pages 4396–4415.
  74. 74.Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang. 2023a. Solving math word problems via cooperative reasoning induced language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4471–4485, Toronto, Canada. Association for Computational Linguistics.
  75. 75.Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang. 2023b. Solving math word problems via cooperative reasoning induced language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4471–4485, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Li, C., et al. “MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10230–58, https://doi.org/10.18653/v1/2024.acl-long.551.
APA
Li, C., Yuan, Z., Yuan, H., Dong, G., Lu, K., Wu, J., Tan, C., Wang, X., & Zhou, C. (2024). MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10230–10258. https://doi.org/10.18653/v1/2024.acl-long.551
Chicago
Li, C., Z. Yuan, H. Yuan, et al. 2024. “MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10230–58. https://doi.org/10.18653/v1/2024.acl-long.551.
Harvard
Li, C. et al. (2024) “MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10230–10258. Available at: https://doi.org/10.18653/v1/2024.acl-long.551.
Vancouver
1. Li C, Yuan Z, Yuan H, Dong G, Lu K, Wu J, Tan C, Wang X, Zhou C (2024) MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10230–10258

BibTeX

@inproceedings{li-etal-2024-mugglemath,
    title = "{M}uggle{M}ath: Assessing the Impact of Query and Response Augmentation on Math Reasoning",
    author = "Li, Chengpeng  and
      Yuan, Zheng  and
      Yuan, Hongyi  and
      Dong, Guanting  and
      Lu, Keming  and
      Wu, Jiancan  and
      Tan, Chuanqi  and
      Wang, Xiang  and
      Zhou, Chang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.551/",
    doi = "10.18653/v1/2024.acl-long.551",
    pages = "10230--10258"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/