Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement
Weimin XiongYifan SongXiutian ZhaoWenhao WuXun WangKe WangCheng LiWei PengSujian Li
Proposes a framework that uses Monte Carlo estimation to derive step-level process rewards and construct contrastive action pairs, significantly improving language model agent training over standard outcome-only reinforcement methods across interactive environments.
Large language model agents are increasingly deployed to solve complex, multi-step interactive tasks such as online purchasing, database querying, and household automation. However, prevailing training methods primarily optimize agents based on final outcome rewards, treating entire task trajectories as single units. This creates operational vulnerabilities, as models can succeed by accident or fail due to uncorrected intermediate missteps without receiving granular feedback on where their reasoning broke down. Providing human supervision for every single intermediate action is prohibitively expensive, leaving a critical gap in scalable, fine-grained process optimization.
The article introduces and evaluates the Iterative Step-Level Process Refinement framework to address this challenge. The primary objective is to demonstrate that integrating automated, step-level process supervision improves decision accuracy and task completion efficiency across diverse large language model agents without requiring manual intermediate annotation.
The evaluated approach uses a multi-stage training pipeline. First, a base agent acquires foundational task capabilities via supervised fine-tuning on expert trajectories. Next, it estimates intermediate step rewards automatically by using a Monte Carlo sampling method, where a fixed scorer model simulates potential futures from each state to calculate expected success. The agent then explores paths originating along expert demonstration trajectories, comparing its actions against expert choices to identify mistakes and build contrastive action pairs. Finally, the agent undergoes iterative offline optimization combining step-level direct preference optimization, outcome-level direct preference optimization, and supervised fine-tuning losses. The framework was evaluated across three distinct interactive benchmarks: WebShop for online shopping, InterCodeSQL for SQL database execution, and ALFWorld for embodied household tasks, testing base models including Llama-2-7B, Llama-2-13B, Llama-3-8B, and Mistral-7B.
The evaluation produced four key findings. First, the proposed framework consistently outperformed leading outcome-based optimization methods, exceeding the prior state-of-the-art technique by 5.8% on WebShop, 7.2% on InterCodeSQL, and up to 3.2% on ALFWorld. Second, training smaller open-source models with this framework enabled them to outperform leading proprietary baselines; for instance, a Llama-2-7B model improved its average reward score from 5.5 to 69.4, surpassing GPT-4's 45.7 average. Third, the method substantially increased the average reward per step, eliminating blind exploration loops and generating more direct, error-free execution paths. Fourth, empirical analysis verified that automated sampling provides accurate step assessments up to 82% of the time, and training a standalone step-reward model can transfer effectively across different base model architectures to reduce training overhead.
These findings indicate that process-level supervision provides vital stability and performance gains that outcome-only feedback cannot deliver. For organizations building autonomous language model systems, this approach reduces operational execution risk, improves decision quality, and cuts inference interaction costs by shortening action paths. Although generating step rewards via Monte Carlo sampling increases initial training time by roughly double compared to outcome-only baselines, this computation occurs entirely offline, resulting in no added latency during live deployment.
Organizations developing interactive agents should transition from outcome-only fine-tuning to mixed process-and-outcome optimization architectures. When adopting this method, teams should calibrate iteration counts carefully, as empirical results show performance peaks within three to four rounds before excessive exploration causes overfitting. Furthermore, teams requiring faster training cycles can deploy a learned step-reward model as a viable alternative to raw Monte Carlo sampling.
Readers should note specific boundaries regarding these findings. Iterative preference tuning on self-generated actions risks overfitting when baseline training data is scarce. Additionally, while the framework categorizes actions into win-or-lose pairs, it does not yet utilize the exact numerical reward magnitudes to prioritize fixing severe errors before minor ones. Nonetheless, because findings remained robust across multiple architectures and held up on unseen generalization tasks, there is high confidence in the framework's effectiveness for structured interactive agent environments.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). Introduces process-supervised reward models that verify reasoning step by step rather than solely evaluating final outcomes, providing foundational motivation for step-level process refinement in LLM agents.
- Paper: A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning, Stephane Ross et al. (2010). Establishes the DAgger framework for iterative imitation learning where an agent explores around expert demonstrations to gather corrective feedback, directly underlying the iterative trajectory exploration in IPR.
- Paper: Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping, Andrew Y. Ng et al. (1999). Provides the core theoretical foundation for shaping intermediate state-level rewards to improve sample efficiency without distorting the optimal policy.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). Demonstrates iterative self-refinement and corrective feedback loops for large language model outputs, offering key intuition for iterative trajectory improvement.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). Extends step-level process verification by framing intermediate reasoning step evaluation through sequential Q-value rankings rather than independent reward classification.
- Paper: TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents, Leitian Tao et al. (2026). Builds on step-level credit assignment for interactive LLM agents by estimating turn-level temporal-difference rewards without requiring explicit step annotations.
- Paper: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments, Hongjin Su et al. (2025). Advances agent trajectory generation and refinement by proposing an environment-adaptive data-centric framework that constructs training tasks from interaction histories.
- Paper: Agent Workflow Memory, Zora Zhiruo Wang et al. (2025). Applies step-level agent learning principles to extract and retain reusable procedural workflows in memory across multi-step digital environments.
- Paper: Sample-Efficient Learning from Agent Experience, Chenhui Gou et al. (2026). Explores sample-efficient distillation of multi-turn agent interaction traces into model weights, continuing the study of agent trajectory optimization.
