Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
Fangkai JiaoChengwei QinZhengyuan LiuNancy F. ChenShafiq Joty
Develops a framework that synthesizes step-level process rewards from offline outcome simulations and trains 7B models via Direct Preference Optimization, enabling them to outperform GPT-3.5-Turbo on complex logical reasoning without expensive human supervision or high-latency search algorithms.
Large language models frequently struggle with complex problem-solving because they can generate flawed, hallucinated, or unfaithful reasoning steps even when they arrive at the correct final answer. Existing methods to improve reasoning quality typically rely on either costly human step-by-step supervision or online search algorithms like Monte Carlo Tree Search, which introduce significant latency and computational overhead during deployment.
The main objective of the article is to demonstrate an efficient framework that teaches language models planning-based reasoning without relying on online search or expensive human annotations, by synthesizing step-level process rewards from final outcome data and optimizing the model via preference learning.
The researchers developed an offline simulation approach that samples intermediate reasoning steps from model-generated solutions and attempts to complete each path multiple times. The frequency of reaching the correct final answer serves as an estimated return to train a process reward model. This reward model then scores complete reasoning trajectories, constructing preference pairs that contrast high-quality reasoning against lower-quality paths. The target models are fine-tuned using Direct Preference Optimization (DPO), integrating both outcome and process supervision across logical reasoning benchmarks (LogiQA-v2 and ReClor) and mathematical reasoning benchmarks (GSM8K and MATH).
The analysis yielded several key findings. First, process-supervised Direct Preference Optimization (pDPO) consistently outperformed standard outcome-only DPO across benchmarks; on LogiQA-v2, a fine-tuned 7-billion-parameter model achieved 55.5% accuracy compared to 53.1% for standard DPO, outperforming larger foundation models such as GPT-3.5-Turbo (45.4%) and Mixtral-8x7B (49.5%). Second, the framework showed high data efficiency, as pDPO using only 40% of outcome annotations matched the performance of standard DPO trained on 100% of the data. Third, automated evaluations via GPT-4 showed that pDPO produced substantially superior reasoning rationales, winning 67.8% of pairwise comparisons overall, 52.5% on reasonableness, and 59.4% on conciseness. Finally, DPO-based training demonstrated major computational advantages over reinforcement learning baselines like PPO, completing training in under 16 hours on four H100 GPUs compared to over 40 hours for PPO.
These results demonstrate that planning capabilities can be distilled directly into model weights during offline training, eliminating the high inference latencies associated with search algorithms while drastically cutting the annotation costs of process supervision. For decision-makers, this translates to faster deployment times, reduced cloud serving expenses, and more transparent, reliable model explanations in high-stakes reasoning domains.
Organizations developing reasoning-intensive AI applications should adopt offline trajectory exploration and process reward modeling to enhance small-to-medium language models rather than relying exclusively on larger proprietary APIs or slow online search techniques. Before wide deployment, teams must carefully calibrate reward confidence margins, as the source shows that overly aggressive margins can degrade performance by introducing false positives.
While the findings provide high confidence for standard logical and arithmetic problems, limitations remain regarding extreme context lengths and tasks requiring deep multi-step code generation. The offline simulation phase still demands significant initial computational resources, and performance gains depend partly on the baseline capabilities of the underlying base model.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). It establishes why evaluating intermediate reasoning steps can outperform outcome-only supervision, the central premise behind the source’s process rewards.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Its rollout-based synthetic step scoring directly prefigures the source’s use of simulated continuations to train process rewards without human step labels.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). It advances process reward modeling by ranking steps according to their expected downstream success rather than scoring each step independently.
- Paper: Reward-Guided Speculative Decoding for Efficient LLM Reasoning, Baohao Liao et al. (2025). It carries process rewards into inference-time efficiency, using step scores to decide when a larger model must intervene.
- Paper: Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models, Raj Jaiswal et al. (2026). It extends step-level reward training to small-model physics reasoning, adding targeted error feedback while avoiding inference-time verification overhead.
