Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Fangkai JiaoChengwei QinZhengyuan LiuNancy F. ChenShafiq Joty

article2024EMNLP54 citationsOutstanding Paper Award

Develops a framework that synthesizes step-level process rewards from offline outcome simulations and trains 7B models via Direct Preference Optimization, enabling them to outperform GPT-3.5-Turbo on complex logical reasoning without expensive human supervision or high-latency search algorithms.

Listen

Large language models frequently struggle with complex problem-solving because they can generate flawed, hallucinated, or unfaithful reasoning steps even when they arrive at the correct final answer. Existing methods to improve reasoning quality typically rely on either costly human step-by-step supervision or online search algorithms like Monte Carlo Tree Search, which introduce significant latency and computational overhead during deployment.

The main objective of the article is to demonstrate an efficient framework that teaches language models planning-based reasoning without relying on online search or expensive human annotations, by synthesizing step-level process rewards from final outcome data and optimizing the model via preference learning.

The researchers developed an offline simulation approach that samples intermediate reasoning steps from model-generated solutions and attempts to complete each path multiple times. The frequency of reaching the correct final answer serves as an estimated return to train a process reward model. This reward model then scores complete reasoning trajectories, constructing preference pairs that contrast high-quality reasoning against lower-quality paths. The target models are fine-tuned using Direct Preference Optimization (DPO), integrating both outcome and process supervision across logical reasoning benchmarks (LogiQA-v2 and ReClor) and mathematical reasoning benchmarks (GSM8K and MATH).

The analysis yielded several key findings. First, process-supervised Direct Preference Optimization (pDPO) consistently outperformed standard outcome-only DPO across benchmarks; on LogiQA-v2, a fine-tuned 7-billion-parameter model achieved 55.5% accuracy compared to 53.1% for standard DPO, outperforming larger foundation models such as GPT-3.5-Turbo (45.4%) and Mixtral-8x7B (49.5%). Second, the framework showed high data efficiency, as pDPO using only 40% of outcome annotations matched the performance of standard DPO trained on 100% of the data. Third, automated evaluations via GPT-4 showed that pDPO produced substantially superior reasoning rationales, winning 67.8% of pairwise comparisons overall, 52.5% on reasonableness, and 59.4% on conciseness. Finally, DPO-based training demonstrated major computational advantages over reinforcement learning baselines like PPO, completing training in under 16 hours on four H100 GPUs compared to over 40 hours for PPO.

These results demonstrate that planning capabilities can be distilled directly into model weights during offline training, eliminating the high inference latencies associated with search algorithms while drastically cutting the annotation costs of process supervision. For decision-makers, this translates to faster deployment times, reduced cloud serving expenses, and more transparent, reliable model explanations in high-stakes reasoning domains.

Organizations developing reasoning-intensive AI applications should adopt offline trajectory exploration and process reward modeling to enhance small-to-medium language models rather than relying exclusively on larger proprietary APIs or slow online search techniques. Before wide deployment, teams must carefully calibrate reward confidence margins, as the source shows that overly aggressive margins can degrade performance by introducing false positives.

While the findings provide high confidence for standard logical and arithmetic problems, limitations remain regarding extreme context lengths and tasks requiring deep multi-step code generation. The offline simulation phase still demands significant initial computational resources, and performance gains depend partly on the baseline capabilities of the underlying base model.

Cover for Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Abstract

Large Language Models (LLMs) have demonstrated significant potential in handling complex reasoning tasks through step-by-step rationale generation. However, recent studies have raised concerns regarding the hallucination and flaws in their reasoning process. Substantial efforts are being made to improve the reliability and faithfulness of the generated rationales. Some approaches model reasoning as planning, while others focus on annotating for process supervision. Nevertheless, the planning-based search process often results in high latency due to the frequent assessment of intermediate reasoning states and the extensive exploration space. Additionally, supervising the reasoning process with human annotation is costly and challenging to scale for LLM training. To address these issues, in this paper, we propose a framework to learn planning-based reasoning through Direct Preference Optimization (DPO) on collected trajectories, which are ranked according to our synthesized process rewards. Our results on challenging logical reasoning benchmarks demonstrate the effectiveness of our learning framework, showing that our 7B model can surpass the strong counterparts like GPT-3.5-Turbo.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 LLMs for Reasoning
  • 2.2 Improving LLMs via Sparse Feedback
  • 3 Method
  • 3.1 Formal Definition of Natural Language Reasoning
  • 3.2 Estimate Process Rewards via Offline Simulation
  • 3.3 Synthesized Process Reward Model
  • 3.4 Reward Annotation and Preference Dataset Construction
  • 3.5 Direct Preference Optimization
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Evaluation and Implementation
  • 5 Results and Analysis
  • 5.1 Overall Results on Logical Reasoning
  • 5.2 Improvements by Iterative Training
  • 5.3 Results on Mathematical Reasoning
  • 5.4 Reliance on Annotations of Outcome Supervision
  • 5.5 Auto-evaluation of Rationale Quality by GPT-4
  • 5.6 Analysis of Predicted Rewards
  • 5.7 Case Study
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A. Baseline
  • B. Evaluation Details
  • C. Implementation Details
  • C.1 Data Preparation
  • C.1.1 Training Data Collection For PRM
  • C.2 Hyper-Parameters
  • C.3 Training
  • D. Compared with MATH-Shepherd
  • E. Effect of Different Reward Margins

Knowls

  1. Knowl 1 — Offline learning framework for planning-based reasoning

    model/method

    The framework treats an LLM’s reasoning as a trajectory of alternating actions and states, with both generated by the policy model. For an input consisting of a prompt, context, and question, the terminal outcome receives reward 1 if it matches the ground-truth answer and 0 otherwise. Rather than using online tree search to evaluate intermediate reasoning states at inference time, the method collects trajectories offline: it samples complete seed solutions, selects intermediate prefixes, and has the policy model complete each prefix repeatedly. The resulting outcome-based estimates train a process reward model (PRM); that PRM scores complete trajectories to form preferences for Direct Preference Optimization (DPO). The aim is to transfer planning-style evaluation into the policy so that the learned model produces better reasoning without requiring the same repeated search and evaluation at inference.

  2. Knowl 2 — Estimating intermediate-step value by repeated continuation

    equation

    For an input with ground-truth answer yy, let uu be an intermediate reasoning prefix, and sample KK continuations from the policy model beginning at uu. The paper’s raw value label is the number of completed continuations whose final answer matches yy; with fixed KK, this count is proportional to the empirical success rate.

    v^(u,y)=∑k=1K1[answer⁡(τk∣u)=y]\hat{v}(u,y)=\sum_{k=1}^{K}\mathbf{1}[\operatorname{answer}(\tau^{k}\mid u)=y]

    Here, KK is the positive integer number of sampled completions, τk∣u\tau^{k}\mid u is the kkth completed trajectory continuing prefix uu, and 1[⋅]\mathbf{1}[\cdot] is 1 when its condition holds and 0 otherwise. The ground-truth outcome supplies the supervision; no human step-level correctness labels are required.

  3. Knowl 3 — PRM training and multiplicative trajectory scoring

    model/method

    The PRM is trained to classify an incomplete reasoning prefix by its raw continuation value: the count of successful completions among the KK sampled continuations. For an input xx and prefix uu, it predicts a categorical distribution pi(x,u)p_i(x,u) over success-count labels i∈{0,…,K}i\in\{0,\ldots,K\} and is fitted using cross-entropy against the observed count. To score a step, the method sums probabilities for counts at or above a threshold CC; it scores a full trajectory by multiplying these step scores across its reasoning steps.

    qt=∑i=CKpi(x,ut),Rproc(τ)=∏t=1Tqtq_t=\sum_{i=C}^{K}p_i(x,u_t),\qquad R_{\mathrm{proc}}(\tau)=\prod_{t=1}^{T}q_t

    Here, utu_t is the prefix at step tt, TT is the number of scored steps, and RprocR_{\mathrm{proc}} is the process-based trajectory score. The cutoff CC requires a minimum number of successful simulated continuations before a step contributes probability mass, reducing reliance on weak evidence from prefixes with few successful completions. The experiments set C=2C=2 for logical reasoning and C=3C=3 for mathematical reasoning. The PRM can score an action or a state; the method describes scoring actions for simplicity.

  4. Knowl 4 — Process-supervised DPO preference construction

    model/method

    The method combines two kinds of trajectory preferences. An outcome-supervised pair consists of a trajectory that reaches the correct answer and one that does not. A process-supervised pair consists of two trajectories that both reach the correct answer, but whose PRM trajectory scores differ by more than a margin σ\sigma; the higher-scoring trajectory is preferred. DPO trains the policy on the union of these pairs, so final-answer correctness and the estimated quality of intermediate reasoning both affect learning.

    LDPO=−E(x,τw,τl)[log⁡sigmoid⁡(β(log⁡πθ(τw∣x)πref(τw∣x)−log⁡πθ(τl∣x)πref(τl∣x)))]\mathcal{L}_{\mathrm{DPO}}=-\mathbb{E}_{(x,\tau_w,\tau_l)}\left[\log\operatorname{sigmoid}\left(\beta\left(\log\frac{\pi_\theta(\tau_w\mid x)}{\pi_{\mathrm{ref}}(\tau_w\mid x)}-\log\frac{\pi_\theta(\tau_l\mid x)}{\pi_{\mathrm{ref}}(\tau_l\mid x)}\right)\right)\right]

    Here, xx is the task input; τw\tau_w and τl\tau_l are the preferred and rejected trajectories; πθ\pi_\theta is the trainable policy; πref\pi_{\mathrm{ref}} is the reference policy initialized from the pre-DPO model; β\beta controls the policy’s divergence from the reference; and sigmoid⁡\operatorname{sigmoid} is the logistic function. The margin σ\sigma applies when constructing process-supervised pairs, not in the DPO loss itself.

  5. Knowl 5 — Logical reasoning benchmark results

    empirical result

    On the ReClor and LogiQA-v2 test sets, process-supervised DPO (pDPO) improved over outcome-only DPO when training on LogiQA-v2, and substantially improved ReClor test performance when training on ReClor. The following accuracies are percentages, with columns ordered as LogiQA-v2 then ReClor:

    • Training on LogiQA-v2: SFT achieved 45.5 / 53.4; DPO achieved 53.1 / 60.4; pDPO achieved 55.5 / 61.7.
    • Training on ReClor: SFT achieved 44.5 / 48.8; DPO achieved 47.5 / 51.3; pDPO achieved 47.4 / 53.5.
    • Reference foundation-model results were 45.4 / 53.7 for GPT-3.5-Turbo and 49.5 / 56.7 for Mixtral-8×7B-Instruct.

    Thus, LogiQA-v2-trained pDPO gained 2.4 percentage points on LogiQA-v2 and 1.3 on ReClor over the corresponding DPO model, and exceeded the reported GPT-3.5-Turbo and Mixtral results on both benchmarks. ReClor-trained pDPO was slightly below ReClor-trained DPO on LogiQA-v2 but reached 2.2 points higher accuracy on ReClor.

  6. Knowl 6 — Iterative training performance and resource use

    empirical result

    In a second training round, the researchers used LogiQA-v2-trained Llama-2-7B-pDPO as the starting model and trained on newly sampled solutions. Test accuracies, again listed as LogiQA-v2 / ReClor percentages, were 56.7 / 61.0 for iterative DPO, 57.3 / 61.8 for iterative pDPO, 56.2 / 61.2 for process-reward PPO, and 57.3 / 61.7 for process-reward GRPO. All four methods improved in-domain performance relative to the initial LogiQA-v2-trained pDPO result of 55.5 / 61.7. Iterative pDPO slightly exceeded GRPO on ReClor and matched it on LogiQA-v2; it also exceeded PPO on both test sets. DPO-based training took under 16 hours on four NVIDIA H100 GPUs, while PPO and GRPO each required over 40 hours on the same hardware.

  7. Knowl 7 — Mathematical reasoning results across two model families

    empirical result

    The mathematical-reasoning evaluation used GSM8K and MATH test accuracy, reported as percentages in that order. For Gemma, Gemma-2B-SFT scored 45.8 / 14.1, outcome-only DPO scored 50.6 / 16.0, and pDPO scored 52.8 / 15.7; Gemma-7B-Instruct scored 46.4 / 24.3. Thus pDPO improved Gemma-2B GSM8K accuracy over DPO and the 7B instruction model, but its MATH accuracy was slightly below Gemma-2B-DPO.

    For DeepSeekMath-7B-Instruct, the base model scored 82.3 / 45.1, DPO scored 82.4 / 46.3, and pDPO scored 82.3 / 46.8. Process-supervised DPO therefore achieved the highest MATH result in this group, while GSM8K performance was essentially unchanged from the base model and marginally below DPO. Gemma training used a 25,000-question MetaMath subset; DeepSeekMath training used a 55,000-question subset augmented from MATH training data. Results were averaged over three runs except for the SFT result.

  8. Knowl 8 — Performance with fewer outcome annotations

    empirical result

    The researchers evaluated LogiQA-v2 training subsets containing 40%, 60%, and 80% of the original questions. Across these subset sizes, pDPO consistently outperformed outcome-only DPO. With 40% of the training questions—3,234 annotated examples—pDPO exceeded the base SFT model by a substantial margin and reached 53.5% test accuracy, close to the 53.9% reported for DPO using the full training set. The PRM itself was trained using outcome annotations from only 10% of the LogiQA-v2 training questions. These results indicate that the proposed method can make effective use of sparse outcome supervision, although it still depends on labeled final answers.

  9. Knowl 9 — GPT-4 comparison of rationale quality

    empirical result

    GPT-4 compared pDPO and DPO rationales on 261 LogiQA-v2 questions for which both models produced the correct answer. It judged each rationale for reasonableness, conciseness, and logical consistency, and also provided an overall judgment. The percentages below are ordered as pDPO win / tie / DPO win: reasonableness, 52.5 / 22.2 / 25.3; conciseness, 59.4 / 0.0 / 40.6; logical consistency, 24.1 / 54.8 / 21.1; and overall, 67.8 / 0.0 / 32.2. The evaluation therefore favored pDPO on overall quality, reasonableness, and conciseness, while logical-consistency judgments were predominantly ties.

  10. Knowl 10 — Resource-related limitations of offline simulation

    limitation

    The repeated-completion simulation used to estimate intermediate-step values remains resource-intensive. The authors report that this constrained their ability to evaluate the approach on competition-level code generation, which requires long-context generation, and on larger policy models. Consequently, the reported results do not establish how well the method scales to those settings.

Coverage note — The reward-margin sensitivity experiment, reward-versus-length plots, and individual qualitative case study are omitted as secondary analyses or illustrations; the framework, principal benchmark results, annotation-scale analysis, and stated limitation are included.

References

  1. 1.Mohammad Gheshlaghi Azar, Mark Rowland, Bilal Piot, Daniel Guo, Daniele Calandriello, Michal Valko, and Rémi Munos. 2023. A general theoretical paradigm to understand learning from human preferences. CoRR, abs/2310.12036.
  2. 2.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosiute, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemí Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022. Constitutional AI: harmlessness from AI feedback. CoRR, abs/2212.08073.
  4. 4.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, YinTat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, MarcoTulio Ribeiro, and Yi Zhang. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4.
  5. 5.Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In NeurIPS, pages 4299–4307.
  6. 6.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. CoRR, abs/2110.14168.
  7. 7.Rémi Coulom. 2006. Efficient selectivity and backup operators in monte-carlo tree search. In Computers and Games, 5th International Conference, volume 4630 of Lecture Notes in Computer Science, pages 72–83. Springer.
  8. 8.Google DeepMind Gemma Team. 2024. Gemma: Open models based on gemini research and technology. Preprint, arXiv:2403.08295.
  9. 9.Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu. 2023. Reasoning with language model is planning with world model. CoRR, abs/2305.14992.
  10. 10.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. NeurIPS.
  11. 11.Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the knowledge in a neural network. CoRR, abs/1503.02531.
  12. 12.Jie Huang and Kevin Chen-Chuan Chang. 2023. Towards reasoning in large language models: A survey. In Findings of ACL, pages 1049–1065. ACL.
  13. 13.Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. CoRR, abs/2311.05232.
  14. 14.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024. Mixtral of experts. Preprint, arXiv:2401.04088.
  15. 15.Fangkai Jiao, Zhiyang Teng, Shafiq R. Joty, Bosheng Ding, Aixin Sun, Zhengyuan Liu, and Nancy F. Chen. 2024. Exploring self-supervised logic-enhanced training for large language models. In NAACL. ACL.
  16. 16.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles.
  17. 17.Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamile Lukosiute, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez. 2023. Measuring faithfulness in chain-of-thought reasoning. CoRR, abs/2307.13702.
  18. 18.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023a. Let’s verify step by step. arXiv preprint arXiv:2305.20050.
  19. 19.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023b. Let’s verify step by step. CoRR, abs/2305.20050.
  20. 20.Hanmeng Liu, Jian Liu, Leyang Cui, Nan Duan, Ming Zhou, and Yue Zhang. 2022. Logiqa2.0 dataset - logical reasoning in mrc and nli tasks. TASLP.
  21. 21.Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 2023. Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct. CoRR, abs/2308.09583.
  22. 22.OpenAI. 2023. Gpt-4 technical report. Technical report.
  23. 23.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  24. 24.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. CoRR, abs/2305.18290.
  25. 25.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. CoRR, abs/1707.06347.
  26. 26.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. CoRR, abs/2402.03300.
  27. 27.Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. In NeurIPS.
  28. 28.Avi Singh, John D. Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J. Liu, James Harrison, Jaehoon Lee, Kelvin Xu, Aaron Parisi, Abhishek Kumar, Alex Alemi, Alex Rizkowsky, Azade Nova, Ben Adlam, Bernd Bohnet, Gamaleldin F. Elsayed, Hanie Sedghi, Igor Mordatch, Isabelle Simpson, Izzeddin Gur, Jasper Snoek, Jeffrey Pennington, Jiri Hron, Kathleen Kenealy, Kevin Swersky, Kshiteej Mahajan, Laura Culp, Lechao Xiao, Maxwell L. Bileschi, Noah Constant, Roman Novak, Rosanne Liu, Tris Warkentin, Yundi Qian, Yamini Bansal, Ethan Dyer, Behnam Neyshabur, Jascha Sohl-Dickstein, and Noah Fiedel. 2023. Beyond human data: Scaling self-training for problem-solving with language models. CoRR, abs/2312.06585.
  29. 29.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  30. 30.Jonathan Uesato, Nate Kushman, Ramana Kumar, H. Francis Song, Noah Y. Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. 2022. Solving math word problems with process- and outcome-based feedback. CoRR, abs/2211.14275.
  31. 31.Bin Wang, Zhengyuan Liu, Xin Huang, Fangkai Jiao, Yang Ding, Ai Ti Aw, and Nancy F. Chen. 2024. Seaeval for multilingual foundation models: From cross-lingual alignment to cultural reasoning. In NAACL. ACL.
  32. 32.Peiyi Wang, Lei Li, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen, Y. Wu, and Zhifang Sui. 2023. Math-shepherd: Verify and reinforce llms step-by-step without human annotations. CoRR, abs/2312.08935.
  33. 33.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  34. 34.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. CoRR, abs/2304.12244.
  35. 35.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023a. Tree of thoughts: Deliberate problem solving with large language models. CoRR, abs/2305.10601.
  36. 36.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023b. React: Synergizing reasoning and acting in language models. In ICLR. OpenReview.net.
  37. 37.Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2023. Metamath: Bootstrap your own mathematical questions for large language models. CoRR, abs/2309.12284.
  38. 38.Ping Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissy, Gargi Ghosh, Mona T. Diab, and Asli Celikyilmaz. 2022. ALERT: adapting language models to reasoning tasks. CoRR, abs/2212.08286.
  39. 39.Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020. Reclor: A reading comprehension dataset requiring logical reasoning. In ICLR. OpenReview.
  40. 40.Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou. 2023. Scaling relationship on learning mathematical reasoning with large language models. CoRR, abs/2308.01825.
  41. 41.Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023. Mammoth: Building math generalist models through hybrid instruction tuning. CoRR, abs/2309.05653.
  42. 42.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. CoRR, abs/2306.05685.
  43. 43.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023a. LIMA: less is more for alignment. CoRR, abs/2305.11206.
  44. 44.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. 2023b. Least-to-most prompting enables complex reasoning in large language models. In ICLR. OpenReview.net.

Citation

MLA
Jiao, F., et al. “Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 334–50, https://doi.org/10.18653/v1/2024.emnlp-main.20.
APA
Jiao, F., Qin, C., Liu, Z., Chen, N., & Joty, S. (2024). Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 334–350. https://doi.org/10.18653/v1/2024.emnlp-main.20
Chicago
Jiao, F., C. Qin, Z. Liu, N. Chen, and S. Joty. 2024. “Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 334–50. https://doi.org/10.18653/v1/2024.emnlp-main.20.
Harvard
Jiao, F. et al. (2024) “Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 334–350. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.20.
Vancouver
1. Jiao F, Qin C, Liu Z, Chen N, Joty S (2024) Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 334–350

BibTeX

@inproceedings{jiao-etal-2024-learning,
    title = "Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing",
    author = "Jiao, Fangkai  and
      Qin, Chengwei  and
      Liu, Zhengyuan  and
      Chen, Nancy F.  and
      Joty, Shafiq",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.20/",
    doi = "10.18653/v1/2024.emnlp-main.20",
    pages = "334--350"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/