Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement

Weimin XiongYifan SongXiutian ZhaoWenhao WuXun WangKe WangCheng LiWei PengSujian Li

article2024EMNLP108 citations

Proposes a framework that uses Monte Carlo estimation to derive step-level process rewards and construct contrastive action pairs, significantly improving language model agent training over standard outcome-only reinforcement methods across interactive environments.

Listen

Large language model agents are increasingly deployed to solve complex, multi-step interactive tasks such as online purchasing, database querying, and household automation. However, prevailing training methods primarily optimize agents based on final outcome rewards, treating entire task trajectories as single units. This creates operational vulnerabilities, as models can succeed by accident or fail due to uncorrected intermediate missteps without receiving granular feedback on where their reasoning broke down. Providing human supervision for every single intermediate action is prohibitively expensive, leaving a critical gap in scalable, fine-grained process optimization.

The article introduces and evaluates the Iterative Step-Level Process Refinement framework to address this challenge. The primary objective is to demonstrate that integrating automated, step-level process supervision improves decision accuracy and task completion efficiency across diverse large language model agents without requiring manual intermediate annotation.

The evaluated approach uses a multi-stage training pipeline. First, a base agent acquires foundational task capabilities via supervised fine-tuning on expert trajectories. Next, it estimates intermediate step rewards automatically by using a Monte Carlo sampling method, where a fixed scorer model simulates potential futures from each state to calculate expected success. The agent then explores paths originating along expert demonstration trajectories, comparing its actions against expert choices to identify mistakes and build contrastive action pairs. Finally, the agent undergoes iterative offline optimization combining step-level direct preference optimization, outcome-level direct preference optimization, and supervised fine-tuning losses. The framework was evaluated across three distinct interactive benchmarks: WebShop for online shopping, InterCodeSQL for SQL database execution, and ALFWorld for embodied household tasks, testing base models including Llama-2-7B, Llama-2-13B, Llama-3-8B, and Mistral-7B.

The evaluation produced four key findings. First, the proposed framework consistently outperformed leading outcome-based optimization methods, exceeding the prior state-of-the-art technique by 5.8% on WebShop, 7.2% on InterCodeSQL, and up to 3.2% on ALFWorld. Second, training smaller open-source models with this framework enabled them to outperform leading proprietary baselines; for instance, a Llama-2-7B model improved its average reward score from 5.5 to 69.4, surpassing GPT-4's 45.7 average. Third, the method substantially increased the average reward per step, eliminating blind exploration loops and generating more direct, error-free execution paths. Fourth, empirical analysis verified that automated sampling provides accurate step assessments up to 82% of the time, and training a standalone step-reward model can transfer effectively across different base model architectures to reduce training overhead.

These findings indicate that process-level supervision provides vital stability and performance gains that outcome-only feedback cannot deliver. For organizations building autonomous language model systems, this approach reduces operational execution risk, improves decision quality, and cuts inference interaction costs by shortening action paths. Although generating step rewards via Monte Carlo sampling increases initial training time by roughly double compared to outcome-only baselines, this computation occurs entirely offline, resulting in no added latency during live deployment.

Organizations developing interactive agents should transition from outcome-only fine-tuning to mixed process-and-outcome optimization architectures. When adopting this method, teams should calibrate iteration counts carefully, as empirical results show performance peaks within three to four rounds before excessive exploration causes overfitting. Furthermore, teams requiring faster training cycles can deploy a learned step-reward model as a viable alternative to raw Monte Carlo sampling.

Readers should note specific boundaries regarding these findings. Iterative preference tuning on self-generated actions risks overfitting when baseline training data is scarce. Additionally, while the framework categorizes actions into win-or-lose pairs, it does not yet utilize the exact numerical reward magnitudes to prioritize fixing severe errors before minor ones. Nonetheless, because findings remained robust across multiple architectures and held up on unseen generalization tasks, there is high confidence in the framework's effectiveness for structured interactive agent environments.

arXiv: 2406.11176
Cover for Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement

Abstract

Large language model agents have exhibited exceptional performance across a range of complex interactive tasks. Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they primarily concentrate on outcome rewards, which may lead to errors or suboptimal actions due to the absence of process supervision signals. In this paper, we introduce the Iterative step-level Process Refinement (IPR) framework, which provides detailed step-by-step guidance to enhance agent training. Specifically, we adopt the Monte Carlo method to estimate step-level rewards. During each iteration, the agent explores along the expert trajectory and generates new actions. These actions are then evaluated against the corresponding step of expert trajectory using step-level rewards. Such comparison helps identify discrepancies, yielding contrastive action pairs that serve as training data for the agent. Our experiments on three complex agent tasks demonstrate that our framework outperforms a variety of strong baselines. Moreover, our analytical findings highlight the effectiveness of IPR in augmenting action efficiency and its applicability to diverse models†.

Table of Contents

  • 1 Introduction
  • 2 Task Formulation
  • 3 Method
  • 3.1 Supervised Fine-tuning
  • 3.2 Step-level Reward Acquisition
  • 3.3 Iterative Agent Optimization
  • 4 Experiments
  • 4.1 Experiment Settings
  • 4.2 Results
  • 5 Analysis
  • 5.1 Different Base Models
  • 5.2 Ablation Study
  • 5.3 Step Reward Estimation Quality
  • 5.4 Average Reward Per Step
  • 5.5 Exploration of Step Reward Modeling
  • 6 Related Work
  • 6.1 LLM as Agents
  • 6.2 Step-level Process Supervision
  • 6.3 Self-Improvement
  • 7 Conclusion
  • Limitations
  • Acknowledgement
  • References
  • A Dataset Details
  • B Details of the Scoring Function
  • C Training Efficiency Analysis
  • D Expert Trajectories Collection
  • E Case Study
  • Case Study of ALFWorld
  • ETO
  • Actions of blind exploration
  • IPR
  • F Prompt for Evaluation
  • Instruction Prompt for WebShop
  • Instruction Prompt for InterCodeSQL
  • Instruction Prompt for ALFWorld

Knowls

  1. Knowl 1 — Iterative step-level process refinement framework

    model/method

    The Iterative step-level Process Refinement (IPR) framework trains an LLM agent from expert trajectories while adding process supervision at individual decision steps. It first creates a base agent by supervised fine-tuning, then uses that agent as a fixed scorer to estimate step rewards through sampled future completions. During each refinement iteration, the current agent follows prefixes of expert trajectories, explores alternative actions, and compares those actions with the expert action at the same state. Inferior alternatives are paired with the corresponding expert continuations as contrastive training examples. The agent is then updated offline using outcome-level preference learning, step-level preference learning, and supervised learning on successful trajectories. The updated agent becomes the base agent for the next iteration, and the procedure stops at a predetermined iteration limit. This design avoids direct online reinforcement learning while supplying action-level supervision in environments that provide only final outcome rewards.

  2. Knowl 2 — Agent task formulation and expert initialization

    model/method

    IPR models an interactive language-agent task as a partially observable Markov decision process (U,S,A,O,T,R)(U,S,A,O,T,R), where UU is the natural-language instruction space, SS is the environment-state space, AA is the action space, OO is the observation space, T:S×A→ST:S\times A\rightarrow S is the transition function, and R:S×A→[0,1]R:S\times A\rightarrow[0,1] is the environment reward function. At time tt, an agent with policy πθ\pi_\theta receives the history

    et−1=(u,a1,o1,…,at−1,ot−1)e_{t-1}=(u,a_1,o_1,\ldots,a_{t-1},o_{t-1})

    and samples an action at∼πθ(⋅∣et−1)a_t\sim\pi_\theta(\cdot\mid e_{t-1}). The interaction produces an observation oto_t and continues until task completion or a maximum length nn. The complete trajectory is en=(u,a1,o1,…,an,on)e_n=(u,a_1,o_1,\ldots,a_n,o_n), and its final outcome reward is ro(u,en)∈[0,1]r_o(u,e_n)\in[0,1].

    IPR initializes the agent with supervised fine-tuning on a dataset DD of successful ReAct-style expert trajectories. For a trajectory of length nn, the initialization objective is

    LSFT(θ)=−Ee∼D[log⁡πθ(e∣u)]=−Ee∼D[∑t=1nlog⁡πθ(at∣et−1)].\mathcal{L}_{\mathrm{SFT}}(\theta)=-\mathbb{E}_{e\sim D}\left[\log\pi_\theta(e\mid u)\right]=-\mathbb{E}_{e\sim D}\left[\sum_{t=1}^{n}\log\pi_\theta(a_t\mid e_{t-1})\right].

    Here θ\theta denotes trainable policy parameters, uu is the task instruction, and et−1e_{t-1} is the trajectory history before action ata_t.

  3. Knowl 3 — Monte Carlo acquisition of step-level rewards

    equation

    IPR estimates the quality of an action from the expected final task reward obtained after taking that action and continuing the task. Let st∈Ss_t\in S be the environment state at step tt, at∈Aa_t\in A the current action, and et−1e_{t-1} the history before that action. A fixed scorer policy πs\pi_s generates a continuation et:me_{t:m} from that history, where m≥tm\ge t is the resulting terminal or evaluated trajectory length. The step reward is defined as

    rs(st,at)=Eet:m∼πs(⋅∣et−1)[ro(u,em)].r_s(s_t,a_t)=\mathbb{E}_{e_{t:m}\sim\pi_s(\cdot\mid e_{t-1})}\left[r_o(u,e_m)\right].

    Because this expectation is generally intractable, IPR samples NN continuations {e(i)}i=1N\{e^{(i)}\}_{i=1}^{N} from πs\pi_s and estimates the reward by

    \begin{cases} \frac{1}{N}\sum_{i=1}^{N}r_o(u,e^{(i)}), & t<n,\\ r_o(u,e_n), & t=n. \end{cases}$$ The scorer $\pi_s$ is an SFT-trained agent with fixed parameters, $r_o$ is the environment’s final outcome reward in $[0,1]$, and $N$ is the Monte Carlo sample count. Thus, actions are judged by their estimated ability to lead to eventual task success rather than by manually annotated step labels.
  4. Knowl 4 — Step-wise construction of contrastive trajectories

    algorithm

    For every expert trajectory en=(u,a1,o1,…,an,on)e_n=(u,a_1,o_1,\ldots,a_n,o_n), IPR uses each expert prefix et−1e_{t-1} as a controlled exploration context. The current base agent generates an alternative continuation

    e^t:m=(a^t,o^t,…,a^m,o^m).\hat e_{t:m}=(\hat a_t,\hat o_t,\ldots,\hat a_m,\hat o_m).

    The expert action ata_t and alternative action a^t\hat a_t are scored from the same state sts_t using the Monte Carlo step-reward estimator. A contrastive error example is retained only when

    rs(st,a^t)<rs(st,at)−τr_s(s_t,\hat a_t)<r_s(s_t,a_t)-\tau

    and

    ro(u,e^m)<ro(u,en),r_o(u,\hat e_m)<r_o(u,e_n),

    where τ≥0\tau\ge0 is a filtering margin. The expert continuation from step tt is labeled the winning continuation et:nwe^w_{t:n}, and the generated continuation is labeled the losing continuation et:mle^l_{t:m}. The resulting step-level dataset is

    Ds={(et−1,et:nw,et:ml)},D_s=\{(e_{t-1},e^w_{t:n},e^l_{t:m})\},

    which compares continuations conditioned on the same prefix. IPR also forms an outcome-level dataset

    Dt={(u,enw,eml)},D_t=\{(u,e^w_n,e^l_m)\},

    by pairing successful expert trajectories with lower-reward generated trajectories. Exploration is restricted to expert prefixes, so the agent obtains informative mistakes while retaining an immediately available correct action for contrastive learning.

  5. Knowl 5 — Mixture of outcome preference, step preference, and supervised losses

    equation

    IPR updates the policy offline with three losses. Let πθ\pi_\theta be the trainable policy, πref\pi_{\mathrm{ref}} the fixed reference policy used for preference optimization, σ(x)=1/(1+e−x)\sigma(x)=1/(1+e^{-x}) the logistic sigmoid, and β>0\beta>0 the DPO temperature parameter. For outcome-level pairs (u,enw,eml)∼Dt(u,e_n^w,e_m^l)\sim D_t, the outcome preference loss is

    Lo-DPO=−E(u,enw,eml)∼Dt[log⁡σ(βlog⁡πθ(enw∣u)πref(enw∣u)−βlog⁡πθ(eml∣u)πref(eml∣u))].\mathcal{L}_{\mathrm{o\text{-}DPO}}=-\mathbb{E}_{(u,e_n^w,e_m^l)\sim D_t}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(e_n^w\mid u)}{\pi_{\mathrm{ref}}(e_n^w\mid u)}-\beta\log\frac{\pi_\theta(e_m^l\mid u)}{\pi_{\mathrm{ref}}(e_m^l\mid u)}\right)\right].

    For step-level pairs (et−1,et:nw,et:ml)∼Ds(e_{t-1},e_{t:n}^w,e_{t:m}^l)\sim D_s, the process preference loss is

    Ls-DPO=−E(et−1,et:nw,et:ml)∼Ds[log⁡σ(βlog⁡πθ(et:nw∣et−1)πref(et:nw∣et−1)−βlog⁡πθ(et:ml∣et−1)πref(et:ml∣et−1))].\mathcal{L}_{\mathrm{s\text{-}DPO}}=-\mathbb{E}_{(e_{t-1},e_{t:n}^w,e_{t:m}^l)\sim D_s}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(e_{t:n}^w\mid e_{t-1})}{\pi_{\mathrm{ref}}(e_{t:n}^w\mid e_{t-1})}-\beta\log\frac{\pi_\theta(e_{t:m}^l\mid e_{t-1})}{\pi_{\mathrm{ref}}(e_{t:m}^l\mid e_{t-1})}\right)\right].

    To increase the absolute likelihood of successful behavior rather than only the relative preference margin, IPR adds

    LSFT=−E(u,enw,eml)∼Dt[log⁡πθ(enw∣u)].\mathcal{L}_{\mathrm{SFT}}=-\mathbb{E}_{(u,e_n^w,e_m^l)\sim D_t}\left[\log\pi_\theta(e_n^w\mid u)\right].

    The complete update objective is

    L=Lo-DPO+Ls-DPO+LSFT.\mathcal{L}=\mathcal{L}_{\mathrm{o\text{-}DPO}}+\mathcal{L}_{\mathrm{s\text{-}DPO}}+\mathcal{L}_{\mathrm{SFT}}.

    Outcome-DPO learns which full trajectories succeed, step-DPO identifies which continuation is better after a particular prefix, and SFT prevents preference optimization from neglecting the absolute probability of the successful trajectory.

  6. Knowl 6 — Benchmark and training configuration

    experimental setup

    IPR was evaluated using average reward on WebShop for web navigation, InterCodeSQL for interactive SQL querying, and ALFWorld for textual household tasks. WebShop and InterCodeSQL provide rewards on a 00–11 scale, whereas ALFWorld provides a binary completion reward. Expert trajectories were collected in the ReAct format using GPT-4 and filtered for successful outcomes; ALFWorld also supplied human trajectories, with GPT-4 used to add missing intermediate thoughts.

    Could not parse LaTeX table

    The ALFWorld test set contains 140 seen and 134 unseen cases. The primary base model was Llama-2-7B, trained for 3 epochs with batch size 48 using AdamW and a cosine learning-rate schedule. The step-reward scorer used temperature 1 and N=5N=5 Monte Carlo samples. Alternative-action generation used temperature 0, with filtering margins τ=0.5\tau=0.5 for ALFWorld, 0.010.01 for WebShop, and 0.10.1 for InterCodeSQL. Learning rates were searched from 1×10−51\times10^{-5} to 5×10−55\times10^{-5} and DPO values of β\beta from 0.10.1 to 0.50.5; the maximum number of IPR iterations was 4. Generation used vLLM on 8 NVIDIA A100 80GB GPUs. Comparisons included prompt-only GPT-4, GPT-3.5-Turbo, and Llama-2-7B-Chat; SFT, PPO, rejection-sampling fine-tuning, and ETO as outcome-refinement baselines; and Step-PPO as a process-refinement baseline.

  7. Knowl 7 — IPR improves performance across three interactive tasks

    data/table

    The following results compare average rewards for prompt-based, outcome-refinement, and process-refinement methods. The IPR model is Llama-2-7B with the complete proposed training procedure; the best performance across iterations is reported for ETO and IPR. IPR is the strongest method on every benchmark split, including the out-of-domain ALFWorld unseen split.

    Could not parse LaTeX table

    Relative to ETO, the paper reports improvements of 5.8% on WebShop, 7.2% on InterCodeSQL, 2.5% on ALFWorld seen, and 3.2% on ALFWorld unseen, with a 4.5% average improvement. The untrained Llama-2-7B reaches only 5.5 average reward, while IPR raises it to 69.4 and surpasses GPT-4’s 45.7 average reward. Step-PPO performs competitively on InterCodeSQL but is substantially less reliable on the other tasks, consistent with instability from direct reinforcement-learning optimization.

  8. Knowl 8 — IPR transfers across base language models

    data/table

    The authors tested supervised fine-tuning, ETO, and IPR with three different base language models on WebShop and InterCodeSQL. IPR achieves the highest score for every model and task, including the weaker Mistral-7B setting. Stronger initial SFT agents generally retain an advantage after IPR, suggesting that the quality of the starting agent affects the quality of subsequent exploration and refinement.

    Could not parse LaTeX table

    For Mistral-7B, IPR improves over ETO by 3.4 WebShop points and 4.6 InterCodeSQL points. For Llama-2-13B, it improves by 3.3 and 3.0 points, and for Llama-3-8B by 5.8 and 2.3 points, respectively. The best post-IPR WebShop result is 72.2 from Llama-2-13B, while the best post-IPR InterCodeSQL result is 68.1 from Llama-3-8B.

  9. Knowl 9 — Ablations and process-level analyses validate the design

    empirical result

    Removing any component of the IPR objective reduces performance on the ALFWorld unseen set, WebShop, or InterCodeSQL. Removing the SFT term causes the largest degradation, and removing step-DPO hurts more than removing outcome-DPO, supporting the importance of process supervision. Performance improves through early iterations but can decline after excessive exploration, consistent with overfitting to the limited training tasks.

    Could not parse LaTeX table

    On WebShop, the ordering of contrastive pairs produced by Monte Carlo step rewards reaches accuracy up to 82% against an independently expanded page-level scoring function. Llama-2-13B gives the most accurate annotations at the same sample count, and increasing the number of samples generally improves accuracy. Using N=5N=5 is reported to perform comparably to larger sample counts while limiting inference cost. IPR also raises the average estimated reward per action over both SFT and ETO on WebShop, InterCodeSQL, and ALFWorld, with especially pronounced gains on InterCodeSQL.

    The authors additionally trained a Llama-2-7B step-reward model with mean-squared-error loss on 70,000 actions labeled by Monte Carlo estimates. Replacing Monte Carlo scoring with this learned model gives the following results:

    Could not parse LaTeX table

    The learned reward model improves over ETO and transfers to Llama-3-8B even though Llama-3-8B actions were not included in its training data, but it remains less accurate than direct Monte Carlo scoring. On WebShop, three rounds of training take approximately 1 hour for SFT, 2.5 hours for ETO, and 5.3 hours for IPR; the authors report that IPR obtains nearly a 6% performance improvement over ETO at less than three times ETO’s training duration.

  10. Knowl 10 — Stated limitations of IPR

    limitation

    The paper identifies three limitations. First, iterative preference learning from self-generated samples can overfit when the training set contains few tasks; augmenting the task set with GPT-4-generated examples is suggested as a possible mitigation. Second, IPR mainly uses step rewards to identify incorrect actions and construct contrastive pairs, rather than exploiting the numerical severity of those rewards; curriculum strategies that correct more severe errors first are proposed as future work. Third, the learned step-reward model was trained on only one agent task, so its generalization across environments is not established. A broadly applicable cross-task step-reward model remains an open direction.

Coverage note — Detailed benchmark-specific scoring formulas, evaluation prompts, and qualitative trajectory transcripts were omitted because they provide reproduction and illustration details rather than additional load-bearing method, theory, or quantitative findings.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Konstantinos Bousmalis, Giulia Vezzani, Dushyant Rao, Coline Manon Devin, Alex X Lee, Maria Bauza Villalonga, Todor Davchev, Yuxiang Zhou, Agrim Gupta, Akhil Raju, et al. 2023. Robocat: A self-improving generalist agent for robotic manipulation. Transactions on Machine Learning Research.
  3. 3.Baian Chen, Chang Shu, Ehsan Shareghi, Nigel Collier, Karthik Narasimhan, and Shunyu Yao. 2023. Fireact: Toward language agent fine-tuning. arXiv preprint arXiv:2310.05915.
  4. 4.Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. 2024. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335.
  5. 5.Alex Havrilla, Sharath Raparthy, Christoforus Nalmpantis, Jane Dwivedi-Yu, Maksym Zhuravinskyi, Eric Hambro, and Roberta Railneau. 2024. Glore: When, where, and how to improve llm reasoning via global and local refinements. arXiv preprint arXiv:2402.10963.
  6. 6.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023a. Mistral 7b. arXiv preprint arXiv:2310.06825.
  7. 7.Shuyang Jiang, Yuhao Wang, and Yu Wang. 2023b. Selfevolve: A code evolution framework via large language models. arXiv preprint arXiv:2306.02907.
  8. 8.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles.
  9. 9.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. Let’s verify step by step. arXiv preprint arXiv:2305.20050.
  10. 10.Jiacheng Liu, Andrew Cohen, Ramakanth Pasunuru, Yejin Choi, Hannaneh Hajishirzi, and Asli Celikyilmaz. 2023. Don’t throw away your value model! making ppo even better via value-guided monte-carlo tree search decoding. arXiv e-prints, pages arXiv–2309.
  11. 11.Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
  12. 12.Trung Quoc Luong, Xinbo Zhang, Zhanming Jie, Peng Sun, Xiaoran Jin, and Hang Li. 2024. Reft: Reasoning with reinforced fine-tuning. arXiv preprint arXiv:2401.08967.
  13. 13.Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. 2024. Agentboard: An analytical evaluation board of multi-turn llm agents. arXiv preprint arXiv:2401.13178.
  14. 14.Qianli Ma, Haotian Zhou, Tingkai Liu, Jianbo Yuan, Pengfei Liu, Yang You, and Hongxia Yang. 2023. Let’s reward step by step: Step-level reward model as the navigators for reasoning. arXiv preprint arXiv:2310.10080.
  15. 15.AI Meta. 2024. Introducing meta llama 3: The most capable openly available llm to date. Meta AI.
  16. 16.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744.
  17. 17.Shuofei Qiao, Ningyu Zhang, Runnan Fang, Yujie Luo, Wangchunshu Zhou, Yuchen Eleanor Jiang, Chengfei Lv, and Huajun Chen. 2024. Autoact: Automatic agent learning from scratch via self-planning. arXiv preprint arXiv:2401.05268.
  18. 18.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36.
  19. 19.Toran Bruce Richards. 2023. Significant-gravitas/autogpt: An experimental open-source attempt to make gpt-4 fully autonomous. URL https://github. com/Significant-Gravitas/AutoGPT.
  20. 20.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  21. 21.Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large language model alignment: A survey. arXiv preprint arXiv:2309.15025.
  22. 22.Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36.
  23. 23.Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020. Alfworld: Aligning text and embodied environments for interactive learning. arXiv preprint arXiv:2010.03768.
  24. 24.Avi Singh, John D Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Peter J Liu, James Harrison, Jaehoon Lee, Kelvin Xu, Aaron Parisi, et al. 2023. Beyond human data: Scaling self-training for problem-solving with language models. arXiv preprint arXiv:2312.06585.
  25. 25.Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, Ye Tian, and Sujian Li. 2023. Restgpt: Connecting large language models with real-world applications via restful apis. arXiv preprint arXiv:2306.06624.
  26. 26.Yifan Song, Da Yin, Xiang Yue, Jie Huang, Sujian Li, and Bill Yuchen Lin. 2024. Trial and error: Exploration-based trajectory optimization for llm agents. arXiv preprint arXiv:2403.02502.
  27. 27.Zhengwei Tao, Ting-En Lin, Xiancai Chen, Hangyu Li, Yuchuan Wu, Yongbin Li, Zhi Jin, Fei Huang, Dacheng Tao, and Jingren Zhou. 2024. A survey on self-evolution of large language models. arXiv preprint arXiv:2404.14387.
  28. 28.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  29. 29.Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. 2022. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275.
  30. 30.Dennis Ulmer, Elman Mansimov, Kaixiang Lin, Justin Sun, Xibin Gao, and Yi Zhang. 2024. Bootstrapping llm-based task-oriented dialogue agents via self-talk. arXiv preprint arXiv:2401.05033.
  31. 31.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023a. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291.
  32. 32.Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Y Wu, and Zhifang Sui. 2023b. Math-shepherd: A label-free step-by-step verifier for llms in mathematical reasoning. arXiv preprint arXiv:2312.08935.
  33. 33.Xin Wang, Yudong Chen, and Wenwu Zhu. 2021. A survey on curriculum learning. IEEE transactions on pattern analysis and machine intelligence, 44(9):4555–4576.
  34. 34.Zihan Wang, Yunxuan Li, Yuexin Wu, Liangchen Luo, Le Hou, Hongkun Yu, and Jingbo Shang. 2024. Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision. arXiv preprint arXiv:2402.02658.
  35. 35.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  36. 36.Tianbao Xie, Fan Zhou, Zhoujun Cheng, Peng Shi, Luoxuan Weng, Yitao Liu, Toh Jing Hua, Junning Zhao, Qian Liu, Che Liu, et al. 2023. Openagents: An open platform for language agents in the wild. arXiv preprint arXiv:2310.10634.
  37. 37.John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. 2024. Intercode: Standardizing and benchmarking interactive coding with execution feedback. Advances in Neural Information Processing Systems, 36.
  38. 38.Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022a. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems, 35:20744–20757.
  39. 39.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022b. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
  40. 40.Da Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Chandu, Kai-Wei Chang, Yejin Choi, and Bill Yuchen Lin. 2023. Lumos: Learning agents with unified data, modular design, and open-source llms. arXiv preprint arXiv:2311.05657.
  41. 41.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv preprint arXiv:1809.08887.
  42. 42.Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Jia Deng, Boji Shan, Huimin Chen, Ruobing Xie, Yankai Lin, et al. 2024. Advancing llm reasoning generalists with preference trees. arXiv preprint arXiv:2404.02078.
  43. 43.Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Chuanqi Tan, and Chang Zhou. 2023. Scaling relationship on learning mathematical reasoning with large language models. arXiv preprint arXiv:2308.01825.
  44. 44.Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai. 2023. Self-taught optimizer (stop): Recursively self-improving code generation. arXiv preprint arXiv:2310.02304.
  45. 45.Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. 2023. Agenttuning: Enabling generalized agent abilities for llms. arXiv preprint arXiv:2310.12823.
  46. 46.Yuqi Zhu, Shuofei Qiao, Yixin Ou, Shumin Deng, Ningyu Zhang, Shiwei Lyu, Yue Shen, Lei Liang, Jinjie Gu, and Huajun Chen. 2024. Knowagent: Knowledge-augmented planning for llm-based agents. arXiv preprint arXiv:2403.03101.

Citation

MLA
Xiong, W., et al. “Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 1556–72, https://doi.org/10.18653/v1/2024.emnlp-main.93.
APA
Xiong, W., Song, Y., Zhao, X., Wu, W., Wang, X., Wang, K., Li, C., Peng, W., & (李素建), S. L. (2024). Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1556–1572. https://doi.org/10.18653/v1/2024.emnlp-main.93
Chicago
Xiong, W., Y. Song, X. Zhao, et al. 2024. “Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1556–72. https://doi.org/10.18653/v1/2024.emnlp-main.93.
Harvard
Xiong, W. et al. (2024) “Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1556–1572. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.93.
Vancouver
1. Xiong W, Song Y, Zhao X, Wu W, Wang X, Wang K, Li C, Peng W, (李素建) SL (2024) Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1556–1572

BibTeX

@inproceedings{xiong-etal-2024-watch,
    title = "Watch Every Step! {LLM} Agent Learning via Iterative Step-level Process Refinement",
    author = "Xiong, Weimin  and
      Song, Yifan  and
      Zhao, Xiutian  and
      Wu, Wenhao  and
      Wang, Xun  and
      Wang, Ke  and
      Li, Cheng  and
      Peng, Wei  and
      Li, Sujian",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.93/",
    doi = "10.18653/v1/2024.emnlp-main.93",
    pages = "1556--1572"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/