MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
Chengpeng LiZheng YuanHongyi YuanGuanting DongKeming LuJiancan WuChuanqi TanXiang WangChang Zhou
Presents a systematic investigation into query and response augmentation for mathematical reasoning, establishing state-of-the-art open-source models while identifying quantitative scaling laws and critical out-of-domain generalization limits between GSM8K and MATH.
Large language models often struggle with complex multi-step mathematical reasoning. While leading proprietary models exhibit strong capabilities, open-source alternatives lag behind. Fine-tuning models on synthetic datasets created through automated data augmentation offers a promising path to bridge this gap, yet organizations lack clear insight into which augmentation strategies work best, how model performance scales with data volume, and whether improvements generalize beyond specific training domains.
The article evaluates data augmentation techniques across query rephrasing and multi-path reasoning to determine optimal training recipes, quantify performance scaling laws, and measure transferability across distinct mathematical benchmarks.
To conduct this evaluation, the authors generated synthetic datasets, AugGSM8K and AugMATH, using proprietary language models to expand standard elementary and high-school math benchmarks. They applied five query modification techniques—such as increasing complexity, introducing fractions or percentages, and combining concepts—alongside multi-path chain-of-thought response generation. They then fine-tuned open-source LLaMA family models ranging from 7 billion to 70 billion parameters, producing a specialized series dubbed MuggleMath.
The investigation yielded several critical findings. First, MuggleMath achieved new state-of-the-art results among open-source models, improving 70-billion-parameter performance on GSM8K to 82.7 percent and on MATH to 36.3 percent, up from baseline scores of 63.2 percent and 14.4 percent, respectively. Second, performance followed a predictable log-linear scaling curve with data volume on grade-school math and a segmented log-linear relationship on advanced math, matching the sample efficiency of human-written data. Third, increasing problem complexity proved to be the single most effective individual query strategy, while mixing diverse augmentation strategies achieved the highest overall performance as data scaled. Fourth, augmenting harder and previously failed problems yielded substantially larger accuracy improvements than augmenting easier questions. Finally, performance gains failed to generalize across domains: training extensively on elementary word problems provided almost no benefit on advanced multi-topic mathematics, and vice versa.
These findings indicate that targeted synthetic data generation is highly cost-effective for matching the performance of expensive human curation within a given problem type. However, leaders should recognize that high performance on narrow mathematical benchmarks reflects domain-specific task mastery rather than generalized reasoning capability. Embedding space analyses confirm that synthetic variations remain clustered tightly around their seed distributions without expanding model breadth.
Organizations developing reasoning models should prioritize creating diverse seed datasets spanning all target domains rather than over-scaling synthetic variations of a single domain. Practical implementation pipelines should focus synthetic generation on complex edge cases and failed problems to maximize training efficiency. Future initiatives should establish robust multi-domain data generation frameworks and evaluate broader out-of-distribution transfer before deploying fine-tuned models into varied production environments.
Confidence in these findings is high regarding in-domain benchmark performance and predictable scaling trends. However, practitioners should note limitations: data generation relied on proprietary model prompts that may vary across revisions, and the evaluation was confined strictly to mathematical reasoning benchmarks rather than open-ended enterprise reasoning tasks.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). Read the paper that introduced the MATH benchmark first: MuggleMath uses MATH to generate training data and measure its advanced-math results.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This paper establishes chain-of-thought prompting, the step-by-step response format that MuggleMath augments with multiple reasoning paths.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Its self-consistency method provides the multi-path reasoning precedent that helps explain MuggleMath’s augmentation of solution responses.
- Paper: Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation, Yanqi Dai et al. (2026). MathForge carries MuggleMath’s emphasis on harder reformulated problems forward into difficulty-aware training and reinforcement learning.
- Paper: WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct, Haipeng Luo et al. (2025). WizardMath extends synthetic math-question evolution into a larger training pipeline that combines supervised fine-tuning with reinforcement learning.
