Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Peiyi WangLei LiZhihong ShaoRunxin XuDamai DaiYifei LiDeli ChenYu WuZhifang Sui
Presents Math-Shepherd, an automated framework that generates step-by-step process supervision data without human annotators to train reward models that substantially boost mathematical reasoning in language models through verification and reinforcement learning.
Large language models frequently struggle with complex, multi-step mathematical reasoning. While evaluating model outputs step by step through process reward models provides far more reliable guidance than simply evaluating the final answer, training these step-level verifiers has traditionally required massive, expensive human annotation.
The article demonstrates an automated framework called MATH-SHEPHERD that evaluates and reinforces language models step by step without requiring human annotations. The study evaluates how effectively this automated process supervision performs both as an output verifier to rank candidate answers and as a reward model for step-by-step reinforcement learning.
The researchers developed an automated annotation pipeline that determines the quality of any intermediate reasoning step by using a language model to generate multiple possible continuations to the final solution. If a step frequently leads to the correct ground-truth answer, it receives a high reward score. Using this approach, the authors generated hundreds of thousands of step-level training examples and evaluated models ranging from 7 billion to 70 billion parameters across standard benchmarks (GSM8K and MATH) and an out-of-distribution high school exam dataset.
The findings establish that automated step-by-step supervision significantly boosts mathematical problem-solving performance. First, process reinforcement learning enhanced base model accuracy; for example, a 7-billion parameter Mistral model improved from 77.9% to 84.1% on GSM8K and from 28.6% to 33.0% on the harder MATH benchmark. Second, using MATH-SHEPHERD as a verifier to select among candidate solutions further elevated accuracy, bringing the same model to 89.1% on GSM8K and 43.5% on MATH. Third, the system scaled effectively to larger models, enabling a 67-billion parameter model to achieve 93.3% on GSM8K and 48.1% on MATH without external tools. Fourth, automated process rewards proved more data-efficient and robust than whole-solution outcome rewards, outperforming human-annotated datasets and improving performance on an out-of-distribution Hungarian national mathematics exam by 9 points over outcome-based verification.
These results imply that automated step-level verification can substantially reduce the cost and turnaround time of training reliable reasoning systems. By pinpointing exactly where logical errors occur without human labelers, organizations can deploy higher-accuracy models while lowering manual data curation costs and hallucination risks.
Organizations developing reasoning models should transition from whole-output evaluation to automated step-level verification and integrate step-by-step reinforcement learning. Teams should pair smaller generators with stronger, larger verifier models for optimal selection quality, as weaker verifiers can degrade the accuracy of larger base models. Future technical work should explore iterative training cycles where the verifier and generator improve in alternating stages.
Key limitations include the high initial computational cost required to generate multiple solution completions during the annotation phase, as well as inherent label noise from automated estimation. Nonetheless, the resulting models demonstrate high empirical accuracy, and adopting efficient inference architectures can further reduce computational overhead.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This seminal work establishes the foundational methodology of process supervision and step-level reward modeling on mathematical reasoning, which Math-Shepherd directly automates to eliminate human annotation costs.
- Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). This foundational paper introduces verifier-based outcome reward modeling for mathematical word problems, establishing the paradigm of using verifiers to select top candidate reasoning chains.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). This paper presents the MATH dataset and benchmark suite that serves as a primary evaluation testbed throughout Math-Shepherd's experiments.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). This work introduces sampling multiple reasoning paths to evaluate consistency and accuracy, inspiring the multi-continuation rollout technique used by Math-Shepherd to compute step-level rewards.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This paper establishes chain-of-thought prompting, the core step-by-step reasoning structure that Math-Shepherd assesses and reinforces at intermediate stages.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). This paper builds directly upon step-level process reward modeling by framing intermediate verification as sequential Q-value ranking rather than independent step classification.
- Paper: Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision, Zhiqing Sun et al. (2024). This work extends process reward modeling to scalable oversight by investigating how process-level evaluators generalize from easy to hard mathematical problems without manual labels on complex tasks.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). This work extends intermediate step verification by formulating reward modeling as generative next-token prediction with chain-of-thought rationales rather than pure discriminative scoring.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). This paper expands automated step-level verification by developing fine-grained metrics for evaluating step validity and redundancy in mathematical chains of thought.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). This study applies process-based reward models at test time to guide tree search and iterative solution refinement dynamically across problem difficulty tiers.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). This survey provides an extensive synthesis of reinforced reasoning methods in large language models, framing automated step-supervision techniques like Math-Shepherd within the broader frontier of large reasoning models.
