Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Peiyi WangLei LiZhihong ShaoRunxin XuDamai DaiYifei LiDeli ChenYu WuZhifang Sui

article2024ACL802 citations

Presents Math-Shepherd, an automated framework that generates step-by-step process supervision data without human annotators to train reward models that substantially boost mathematical reasoning in language models through verification and reinforcement learning.

Listen

Large language models frequently struggle with complex, multi-step mathematical reasoning. While evaluating model outputs step by step through process reward models provides far more reliable guidance than simply evaluating the final answer, training these step-level verifiers has traditionally required massive, expensive human annotation.

The article demonstrates an automated framework called MATH-SHEPHERD that evaluates and reinforces language models step by step without requiring human annotations. The study evaluates how effectively this automated process supervision performs both as an output verifier to rank candidate answers and as a reward model for step-by-step reinforcement learning.

The researchers developed an automated annotation pipeline that determines the quality of any intermediate reasoning step by using a language model to generate multiple possible continuations to the final solution. If a step frequently leads to the correct ground-truth answer, it receives a high reward score. Using this approach, the authors generated hundreds of thousands of step-level training examples and evaluated models ranging from 7 billion to 70 billion parameters across standard benchmarks (GSM8K and MATH) and an out-of-distribution high school exam dataset.

The findings establish that automated step-by-step supervision significantly boosts mathematical problem-solving performance. First, process reinforcement learning enhanced base model accuracy; for example, a 7-billion parameter Mistral model improved from 77.9% to 84.1% on GSM8K and from 28.6% to 33.0% on the harder MATH benchmark. Second, using MATH-SHEPHERD as a verifier to select among candidate solutions further elevated accuracy, bringing the same model to 89.1% on GSM8K and 43.5% on MATH. Third, the system scaled effectively to larger models, enabling a 67-billion parameter model to achieve 93.3% on GSM8K and 48.1% on MATH without external tools. Fourth, automated process rewards proved more data-efficient and robust than whole-solution outcome rewards, outperforming human-annotated datasets and improving performance on an out-of-distribution Hungarian national mathematics exam by 9 points over outcome-based verification.

These results imply that automated step-level verification can substantially reduce the cost and turnaround time of training reliable reasoning systems. By pinpointing exactly where logical errors occur without human labelers, organizations can deploy higher-accuracy models while lowering manual data curation costs and hallucination risks.

Organizations developing reasoning models should transition from whole-output evaluation to automated step-level verification and integrate step-by-step reinforcement learning. Teams should pair smaller generators with stronger, larger verifier models for optimal selection quality, as weaker verifiers can degrade the accuracy of larger base models. Future technical work should explore iterative training cycles where the verifier and generator improve in alternating stages.

Key limitations include the high initial computational cost required to generate multiple solution completions during the annotation phase, as well as inherent label noise from automated estimation. Nonetheless, the resulting models demonstrate high empirical accuracy, and adopting efficient inference architectures can further reduce computational overhead.

arXiv: 2312.08935
  • Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). This seminal work establishes the foundational methodology of process supervision and step-level reward modeling on mathematical reasoning, which Math-Shepherd directly automates to eliminate human annotation costs.
  • Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). This foundational paper introduces verifier-based outcome reward modeling for mathematical word problems, establishing the paradigm of using verifiers to select top candidate reasoning chains.
  • Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). This paper presents the MATH dataset and benchmark suite that serves as a primary evaluation testbed throughout Math-Shepherd's experiments.
  • Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). This work introduces sampling multiple reasoning paths to evaluate consistency and accuracy, inspiring the multi-continuation rollout technique used by Math-Shepherd to compute step-level rewards.
  • Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This paper establishes chain-of-thought prompting, the core step-by-step reasoning structure that Math-Shepherd assesses and reinforces at intermediate stages.
Cover for Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Abstract

In this paper, we present an innovative process-oriented math process reward model called MATH-SHEPHERD, which assigns a reward score to each step of math problem solutions. The training of MATH-SHEPHERD is achieved using automatically constructed process-wise supervision data, breaking the bottleneck of heavy reliance on manual annotation in existing work. We explore the effectiveness of MATH-SHEPHERD in two scenarios: 1) Verification: MATH-SHEPHERD is utilized for reranking multiple outputs generated by Large Language Models (LLMs); 2) Reinforcement Learning (RL): MATH-SHEPHERD is employed to reinforce LLMs. With MATH-SHEPHERD, a series of open-source LLMs demonstrates exceptional performance. For instance, process RL with MATH-SHEPHERD significantly enhances Mistral-7B (77.9%→84.1% on GSM8K and 28.6%→33.0% on MATH). The accuracy can be further improved to 89.1% and 43.5% on two benchmarks with verification of MATH-SHEPHERD. We believe that automatic process supervision holds significant potential for the future evolution of LLMs.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Methodology
  • 3.1 Task Formulation
  • 3.2 Reward Models
  • 3.3 Automatic Process Annotation
  • 3.3.1 Definition
  • 3.3.2 Solution
  • 3.4 Ranking for Verification
  • 3.5 Reinforcement Learning
  • 4 Experiments
  • 4.1 Main Results
  • 5 Analysis
  • 5.1 Performance with Different Number of Candidate Solutions
  • 5.2 Quality of Automatic Process Annotations
  • 5.3 Influence of the Number of Data
  • 5.4 Out-of-distribution Performance
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Influence of the Model Size for Verification
  • B Case Study

Knowls

  1. Knowl 1 — Potential-based definition of process quality

    definition

    MATH-SHEPHERD defines the quality of an intermediate mathematical reasoning step by its potential to lead to the correct final answer. A step is considered good when subsequent reasoning can use it to deduce a well-founded answer, even if the step itself does not contain a complete solution. This criterion is intentionally outcome-grounded and can contain noise, because a correct final answer does not guarantee that every preceding step is logically valid.

  2. Knowl 2 — Automatic completion-based process annotation

    algorithm

    For a mathematical problem pp with gold answer a∗a^* and a solution divided into steps s1,…,sKs_1,\ldots,s_K, MATH-SHEPHERD labels each step without human process annotations. For every step sis_i, a completer generates NN independent continuations from that step. The jj-th continuation produces a finalized solution and final answer aja_j, giving the answer set A={aj}j=1NA=\{a_j\}_{j=1}^{N}.

    The hard-estimation label is binary:

    yiHE={1,if there exists aj∈A such that aj=a∗,0,otherwise.y_i^{\mathrm{HE}}=\begin{cases}1,&\text{if there exists }a_j\in A\text{ such that }a_j=a^*,\\0,&\text{otherwise.}\end{cases}

    The soft-estimation label is the fraction of completions reaching the gold answer:

    yiSE=1N∑j=1N1(aj=a∗),yiSE∈[0,1].y_i^{\mathrm{SE}}=\frac{1}{N}\sum_{j=1}^{N}\mathbf{1}(a_j=a^*),\qquad y_i^{\mathrm{SE}}\in[0,1].

    The resulting step labels train a binary process reward model. If ri∈(0,1)r_i\in(0,1) is the model's sigmoid score for step sis_i and KK is the number of steps, the training loss is

    LPRM=−∑i=1K[yilog⁡ri+(1−yi)log⁡(1−ri)].\mathcal{L}_{\mathrm{PRM}}=-\sum_{i=1}^{K}\left[y_i\log r_i+(1-y_i)\log(1-r_i)\right].

    The method therefore replaces expensive human step annotations with repeated model completion followed by exact comparison of each completion's final answer with a∗a^*.

  3. Knowl 3 — Step-wise verification and answer aggregation

    model/method

    MATH-SHEPHERD is a process reward model (PRM): it assigns a score to every reasoning step, whereas an outcome reward model (ORM) assigns one score to an entire solution. For verification, the PRM score of a solution S=(s1,…,sK)S=(s_1,\ldots,s_K) is the minimum score among its steps, RM(p,S)=min⁡1≤k≤KrkRM(p,S)=\min_{1\leq k\leq K}r_k, where pp is the problem and rkr_k is the PRM score of step sks_k. The minimum-score rule treats the weakest step as the solution's overall reliability.

    Given NN candidate solutions S1,…,SNS_1,\ldots,S_N for problem pp, with final answers a1,…,aNa_1,\ldots,a_N, reward-model reranking combined with self-consistency predicts

    a^=arg⁡max⁡a∑i=1N1(ai=a) RM(p,Si).\hat a=\arg\max_{a}\sum_{i=1}^{N}\mathbf{1}(a_i=a)\,RM(p,S_i).

    Here RMRM can be an ORM score or the minimum-step PRM score. Self-consistency alone corresponds to choosing the answer with the largest unweighted vote count; the combined method weights each vote by the candidate solution's reward.

  4. Knowl 4 — Step-by-step reinforcement learning rewards

    model/method

    The paper uses MATH-SHEPHERD to supervise Proximal Policy Optimization (PPO) at reasoning-step boundaries rather than only at the end of a response. For a response containing nn tokens, let tt denote a token position and let EOS\mathrm{EOS} be the set of positions marking the end of reasoning steps. Outcome reinforcement learning gives a reward only at the final token:

    rtoutcome={0,t≠n,rORM,t=n,r_t^{\mathrm{outcome}}=\begin{cases}0,&t\neq n,\\r_{\mathrm{ORM}},&t=n,\end{cases}

    where rORMr_{\mathrm{ORM}} is the ORM score for the complete response. Process reinforcement learning gives the PRM score at every step boundary and zero elsewhere:

    rtprocess={0,t∉EOS,rPRM,t,t∈EOS,r_t^{\mathrm{process}}=\begin{cases}0,&t\notin\mathrm{EOS},\\r_{\mathrm{PRM},t},&t\in\mathrm{EOS},\end{cases}

    where rPRM,tr_{\mathrm{PRM},t} is the MATH-SHEPHERD score for the reasoning step ending at position tt. This provides PPO with localized feedback about intermediate reasoning decisions.

  5. Knowl 5 — Training and evaluation configuration

    experimental setup

    Experiments use GSM8K and MATH. The full GSM8K test set is used for verification and reinforcement-learning evaluation; verification on MATH uses MATH500, a 500-problem subset, while reinforcement-learning evaluation uses the full MATH test set. Verification samples 256 candidate solutions per problem and reports the mean accuracy over three sampling groups. Greedy-decoding accuracy is used after reinforcement learning.

    The evaluated language models are LLaMA2-7B/13B/70B, Llemma-7B/34B, Mistral-7B, and DeepSeek-67B. Generators and completers are trained for three epochs on MetaMATH. To create reward-model data, 7B and 13B models are trained for one epoch on the GSM8K and MATH training sets, 15 solutions are sampled per problem, duplicates are removed, and each step is automatically annotated using LLemma-7B as completer with N=8N=8. This produces approximately 170,000 GSM8K solutions and 270,000 MATH solutions.

    For verification, reward models are based on LLaMA2-70B for GSM8K and Llemma-34B for MATH. For reinforcement learning, Mistral-7B supplies the reward model used to supervise LLaMA2-7B and Mistral-7B generators. The experiments use hard estimation, a one-epoch reward-model learning rate of 10−610^{-6}, PPO learning rates of 4×10−74\times10^{-7} for LLaMA2-7B and 10−710^{-7} for Mistral-7B, KL coefficient 0.040.04, cosine scheduling with minimum learning rate 10−810^{-8}, and maximum sequence length 512. Rejective-sampling fine-tuning uses eight sampled responses per MetaMATH question.

  6. Knowl 6 — Verification performance across open-source models

    data/table

    With 256 candidate solutions per problem, MATH-SHEPHERD generally outperforms self-consistency and ORM verification. The results below are accuracy percentages; GSM8K uses the full test set and MATH uses MATH500. The reward models are trained from LLaMA2-70B for GSM8K and Llemma-34B for MATH.

    Generator Verifier GSM8K MATH500
    LLaMA2-70B: MetaMATH Self-Consistency 88.0 39.4
    ORM 91.8 40.4
    Self-Consistency + ORM 92.0 42.0
    MATH-SHEPHERD 93.2 44.5
    Self-Consistency + MATH-SHEPHERD 92.4 45.2
    Llemma-34B: MetaMATH Self-Consistency 82.6 44.2
    ORM 90.0 43.7
    Self-Consistency + ORM 89.6 45.4
    MATH-SHEPHERD 90.9 46.0
    Self-Consistency + MATH-SHEPHERD 89.7 47.3
    DeepSeek-67B: MetaMATH Self-Consistency 88.2 45.4
    ORM 92.6 45.3
    Self-Consistency + ORM 92.4 47.0
    MATH-SHEPHERD 93.3 47.0
    Self-Consistency + MATH-SHEPHERD 92.5 48.1

    MATH-SHEPHERD is strongest as a standalone verifier for every listed generator on GSM8K and improves over ORM particularly clearly on the harder MATH benchmark. Combining it with self-consistency helps on MATH but can reduce performance on GSM8K, indicating that vote aggregation is not universally beneficial when the reward model is already strong.

  7. Knowl 7 — Process reinforcement learning improves mathematical reasoning

    data/table

    Process-supervised PPO improves greedy-decoding accuracy more than rejective-sampling fine-tuning or outcome-supervised PPO for both evaluated generators. Accuracy is reported on GSM8K and MATH.

    Model and training GSM8K MATH
    LLaMA2-7B: MetaMATH 66.6 19.2
    LLaMA2-7B + RFT 68.5 19.9
    LLaMA2-7B + Outcome RL 70.8 20.8
    LLaMA2-7B + Process RL 73.2 21.6
    Mistral-7B: MetaMATH 77.9 28.6
    Mistral-7B + RFT 79.0 29.9
    Mistral-7B + Outcome RL 81.8 31.3
    Mistral-7B + Process RL 84.1 33.0

    For Mistral-7B, process RL raises accuracy from 77.9% to 84.1% on GSM8K and from 28.6% to 33.0% on MATH. Verification and process RL are complementary: for Mistral-7B, self-consistency plus MATH-SHEPHERD reaches 89.1% on GSM8K and 43.5% on MATH after process RL, compared with 86.3% and 38.3% before process RL. The paper also reports that reward-model-only verification after RL can underperform self-consistency, suggesting that the initial reward model may not fully match the distribution produced by the RL-enhanced generator.

  8. Knowl 8 — Quality of automatically generated process labels

    empirical result

    Manual evaluation of 160 GSM8K reasoning steps shows that automatic process annotation can achieve high quality. Using LLaMA2-70B trained on MetaMATH as the completer, hard estimation reaches 86% label accuracy when N=4N=4. Increasing NN beyond this point reduces hard-label accuracy because more completions create additional false positives. Soft estimation becomes progressively closer to the human-annotated label distribution as NN increases, although the verifier's downstream performance shows no substantial difference between models trained with soft and hard labels.

    Automatic completion-based annotation also outperforms the compared NLI- and rule-based methods:

    Method Model Accuracy
    DIVERSE-NLI DeBERTa 61.3
    DIVERSE-NLI LLaMA2-13B 75.6
    DIVERSE-Rule – 75.0
    MATH-SHEPHERD LLaMA2-13B (N=4N=4) 85.0

    Completer capability and training data quality materially affect annotation quality: larger completers produce better labels, and completers trained on MetaMATH or the original GSM8K data have lower evaluation loss than a model trained on a weakened dataset that excludes the evaluation questions. The authors therefore infer, with appropriate qualification, that stronger base models and better training data could further improve automatic annotation.

  9. Knowl 9 — Data efficiency, model scaling, and out-of-distribution generalization

    empirical result

    MATH-SHEPHERD shows higher verification data efficiency than ORM: with approximately 10,000 training solutions, PRM accuracy is about 4 percentage points higher than ORM accuracy, and PRM appears to have a higher performance ceiling. On MATH, the automatically generated reward-model dataset also outperforms the human-annotated PRM800K baseline in the reported comparisons; the automatic dataset is approximately four times larger and is generated from open-source-model outputs, reducing a distribution mismatch with the evaluated generators.

    Verifier size matters. PRM remains better than self-consistency and ORM across 7B, 13B, and 70B model settings. Larger reward models are more robust as the number of candidate solutions increases. A larger reward model substantially improves verification of a smaller generator, whereas using a smaller reward model to verify a larger generator can perform worse than self-consistency.

    The method generalizes beyond GSM8K and MATH on a Hungarian national final examination containing 33 questions worth 100 total points. Llemma-34B generates 256 candidates per question; greedy decoding scores 46.0, ORM selection scores 54.0, and MATH-SHEPHERD selection scores 63.0. Thus PRM exceeds ORM by 9 points in this out-of-distribution evaluation. In a representative problem, MATH-SHEPHERD selected the correct solution from 256 candidates when ORM did not and assigned lower scores to solutions containing arithmetic errors in intermediate steps.

  10. Knowl 10 — Limitations of completion-based supervision

    limitation

    Automatic process annotation requires a completer to generate NN continuations for every reasoning step, so its computational cost grows with the number of completions and can be substantial despite remaining lower than human annotation cost. The paper suggests that more efficient inference techniques could mitigate this expense.

    The generated labels are also noisy: a continuation may reach the correct final answer despite an invalid intermediate step, and larger NN can increase false positives. Although the resulting PRM performs well and exceeds the reported PRM800K comparison, the separate effects of annotation noise and data-distribution differences are not fully determined. The authors identify systematic comparison of human and automatic process labels, and combining both sources of supervision, as future work.

Coverage note — The detailed per-step Hungarian case-study transcript and exact plotted values for every generator/verifier-size curve were omitted because the summarized results capture their substantive contribution without reproducing secondary examples.

References

  1. 1.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  2. 2.Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics. arXiv preprint arXiv:2310.10631, 2023.
  3. 3.Zhen Bi, Ningyu Zhang, Yinuo Jiang, Shumin Deng, Guozhou Zheng, and Huajun Chen. When do program-of-thoughts work for reasoning? arXiv preprint arXiv:2308.15452, 2023.
  4. 4.Liang Chen, Yichi Zhang, Shuhuai Ren, Haozhe Zhao, Zefan Cai, Yuchi Wang, Peiyi Wang, Tianyu Liu, and Baobao Chang. Towards end-to-end embodied decision making via multi-modal large language model: Explorations with gpt4-vision and beyond. arXiv preprint arXiv:2310.02071, 2023.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  6. 6.Rémi Coulom. Efficient selectivity and backup operators in monte-carlo tree search. In International conference on computers and games, pp. 72–83. Springer, 2006.
  7. 7.DeepSeek. Deepseek llm: Let there be answers. https://github.com/deepseek-ai/DeepSeek-LLM, 2023.
  8. 8.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720, 2022.
  9. 9.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Minlie Huang, Nan Duan, Weizhu Chen, et al. Tora: A tool-integrated reasoning agent for mathematical problem solving. arXiv preprint arXiv:2309.17452, 2023.
  10. 10.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  11. 11.High-flyer. Hai-llm: Efficient and lightweight training tool for large models, 2023. URL https://www.high-flyer.cn/en/blog/hai-llm.
  12. 12.Jie Huang and Kevin Chen-Chuan Chang. Towards reasoning in large language models: A survey. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Findings of the Association for Computational Linguistics: ACL 2023, pp. 1049–1065, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.67. URL https://aclanthology.org/2023.findings-acl.67.
  13. 13.Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. Large language models cannot self-correct reasoning yet. arXiv preprint arXiv:2310.01798, 2023.
  14. 14.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  15. 15.Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. Challenges and applications of large language models. arXiv preprint arXiv:2307.10169, 2023.
  16. 16.Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, and Lu Wang. Grace: Discriminator-guided chain-of-thought reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 15299–15328, 2023.
  17. 17.Levente Kocsis and Csaba Szepesvári. Bandit based monte-carlo planning. In European conference on machine learning, pp. 282–293. Springer, 2006.
  18. 18.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, pp. 611–626, 2023.
  19. 19.Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pp. 19274–19286. PMLR, 2023.
  20. 20.Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu, Xu Sun, et al. M3it: A large-scale dataset towards multi-modal multilingual instruction tuning. arXiv preprint arXiv:2306.04387, 2023a.
  21. 21.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. Making language models better reasoners with step-aware verifier. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5315–5333, Toronto, Canada, July 2023b. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.291. URL https://aclanthology.org/2023.acl-long.291.
  22. 22.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  23. 23.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. arXiv preprint arXiv:2308.09583, 2023.
  24. 24.Qianli Ma, Haotian Zhou, Tingkai Liu, Jianbo Yuan, Pengfei Liu, Yang You, and Hongxia Yang. Let’s reward step by step: Step-level reward model as the navigators for reasoning. arXiv preprint arXiv:2310.10080, 2023.
  25. 25.OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. doi: 10.48550/arXiv.2303.08774. URL https://doi.org/10.48550/arXiv.2303.08774.
  26. 26.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  27. 27.Sarah Pan, Vladislav Lialin, Sherin Muckatira, and Anna Rumshisky. Let’s reinforce step by step. arXiv preprint arXiv:2311.05821, 2023.
  28. 28.Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, pp. 1–22, 2023.
  29. 29.David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484–489, 2016.
  30. 30.Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, Ye Tian, and Sujian Li. Restgpt: Connecting large language models with real-world applications via restful apis. corr, abs/2306.06624, 2023. doi: 10.48550. arXiv preprint arXiv.2306.06624.
  31. 31.Maciej Swiechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mandziuk. Monte carlo tree search: A review of recent modifications and applications. Artificial Intelligence Review, 56(3):2497–2562, 2023.
  32. 32.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  33. 33.Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275, 2022.
  34. 34.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023a.
  35. 35.Peiyi Wang, Lei Li, Liang Chen, Feifan Song, Binghuai Lin, Yunbo Cao, Tianyu Liu, and Zhifang Sui. Making large language models better reasoners with alignment. arXiv preprint arXiv:2309.02144, 2023b.
  36. 36.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023c.
  37. 37.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023d. URL https://openreview.net/pdf?id=1PL1NIMMrw.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS, 2022.
  39. 39.Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. Fine-grained human feedback gives better rewards for language model training. arXiv preprint arXiv:2306.01693, 2023.
  40. 40.Heming Xia, Tao Ge, Furu Wei, and Zhifang Sui. Lossless speedup of autoregressive translation with generalized aggressive decoding. arXiv preprint arXiv:2203.16487, 2022.
  41. 41.Fei Yu, Anningzhe Gao, and Benyou Wang. Outcome-supervised verifiers for planning in mathematical reasoning. arXiv preprint arXiv:2311.09724, 2023a.
  42. 42.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. arXiv preprint arXiv:2309.12284, 2023b.
  43. 43.Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou. Scaling relationship on learning mathematical reasoning with large language models. arXiv preprint arXiv:2308.01825, 2023.
  44. 44.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mammoth: Building math generalist models through hybrid instruction tuning. arXiv preprint arXiv:2309.05653, 2023.
  45. 45.Yifan Zhang, Jingqin Yang, Yang Yuan, and Andrew Chi-Chih Yao. Cumulative reasoning with large language models. arXiv preprint arXiv:2308.04371, 2023.
  46. 46.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  47. 47.Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang. Solving math word problems via cooperative reasoning induced language models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 4471–4485, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.245. URL https://aclanthology.org/2023.acl-long.245.

Citation

MLA
(王培懿), P. W., et al. “Math-Shepherd: Verify and Reinforce LLMs Step-by-step Without Human Annotations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 9426–39, https://doi.org/10.18653/v1/2024.acl-long.510.
APA
(王培懿), P. W., Li, L., Shao, Z., Xu, R., Dai, D., Li, Y., Chen, D., Wu, Y., & Sui, Z. (2024). Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9426–9439. https://doi.org/10.18653/v1/2024.acl-long.510
Chicago
(王培懿), P. W., L. Li, Z. Shao, et al. 2024. “Math-Shepherd: Verify and Reinforce LLMs Step-by-step Without Human Annotations”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9426–39. https://doi.org/10.18653/v1/2024.acl-long.510.
Harvard
(王培懿), P.W. et al. (2024) “Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9426–9439. Available at: https://doi.org/10.18653/v1/2024.acl-long.510.
Vancouver
1. (王培懿) PW, Li L, Shao Z, Xu R, Dai D, Li Y, Chen D, Wu Y, Sui Z (2024) Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9426–9439

BibTeX

@inproceedings{wang-etal-2024-math,
    title = "Math-Shepherd: Verify and Reinforce {LLM}s Step-by-step without Human Annotations",
    author = "Wang, Peiyi  and
      Li, Lei  and
      Shao, Zhihong  and
      Xu, Runxin  and
      Dai, Damai  and
      Li, Yifei  and
      Chen, Deli  and
      Wu, Yu  and
      Sui, Zhifang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.510/",
    doi = "10.18653/v1/2024.acl-long.510",
    pages = "9426--9439"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/