ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Yifei ZhouAndrea ZanetteJiayi PanSergey LevineAviral Kumar

article2024ICML236 citations

Introduces a hierarchical reinforcement learning framework that couples high-level off-policy value estimation across multi-turn interactions with low-level token-generation policy gradients, achieving a hundredfold improvement in sample efficiency when training language model agents on delayed-reward tasks.

Listen

Large language models are increasingly deployed as autonomous agents to handle complex, multi-step interactions such as dialogue, web navigation, and tool use. Standard reinforcement learning approaches for language models typically optimize rewards over a single turn, such as standard human feedback alignment, which prevents agents from planning ahead, gathering information, or attributing delayed success to earlier actions. Existing multi-turn methods face severe operational trade-offs: on-policy algorithms discard past interaction data and require prohibitively expensive real-time environment sampling, whereas off-policy algorithms that evaluate individual tokens struggle with long horizons and compound estimation errors over multiple conversational turns.

The article develops and evaluates Actor-Critic Framework with a Hierarchical Structure (ArCHer), a reinforcement learning framework designed to train language model agents efficiently across multi-turn interactions with delayed feedback. ArCHer introduces a two-level hierarchy that decouples utterance-level value estimation from token-level generation. At the high level, it applies off-policy temporal-difference learning across full utterances, enabling efficient sample reuse from past interactions without searching over an intractable action space. At the low level, it uses policy gradient optimization within each turn, using the high-level value estimate as the reward signal to guide token generation in a self-contained optimization loop.

The authors evaluated the framework across theoretical convergence criteria and empirical benchmarks comprising language games, sequential dialogue, and web shopping tasks. The evaluations compared ArCHer against standard on-policy policy optimization, filtered behavioral cloning, and utterance-ranking methods, utilizing base language models ranging up to seven billion parameters.

The findings demonstrate four critical outcomes. First, ArCHer achieves approximately a 100-fold improvement in sample efficiency compared to on-policy policy optimization baselines, reaching target performance with under 1,000 trajectories where standard baselines require upwards of 100,000. Second, the framework delivers superior final task performance, outperforming filtered imitation learning and utterance-ranking baselines across complex reasoning tasks, while a fine-tuned smaller model surpassed prompting techniques on proprietary large models in web navigation. Third, theoretical analysis proves that estimating advantages at the utterance level reduces statistical error accumulation by a factor proportional to the square root of utterance length compared to token-level estimators. Fourth, the architecture scales effectively with model capacity, showing faster policy improvement when scaled from smaller models to a seven-billion-parameter foundation model.

These results demonstrate that multi-turn reinforcement learning can be made computationally feasible and data-efficient without requiring massive online sampling budgets or brittle heuristic prompting. Organizations deploying language model agents can significantly reduce environment interaction costs and computational overhead while achieving reliable multi-turn goal pursuit. The approach allows practitioners to adapt existing single-turn alignment infrastructure directly to multi-turn tasks.

Decision-makers considering multi-turn agent training should prioritize hierarchical reinforcement learning designs over pure on-policy methods or simple demonstration cloning. When deploying agents in environments with long conversational turns, teams should incorporate token-level baseline value functions to maintain training stability. Furthermore, organizations should run initial offline or hybrid pilot implementations using logged interaction data before investing in full-scale online training environments.

The evaluations were conducted in simulated environments with programmed reward functions rather than direct human interactive trials, and the system still requires several thousand interactions to converge. Additionally, the authors note that free-form responses from imperfect auxiliary models can occasionally cause agents to fall into repetitive loops in out-of-distribution states. Confidence in the algorithmic efficiency gains is high based on the empirical and theoretical alignment, but performance in open-ended production environments with human users will require validation.

Cover for ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Abstract

A use case of large language models (LLMs) is in goal-directed decision-making tasks (or “agent” tasks), where an LLM needs to make intelligent decisions over a multi-turn interaction to accomplish a task. Reinforcement learning (RL) provides a general paradigm to address such agent tasks, but current RL methods for LLMs largely focus on optimizing single-turn rewards (e.g., PPO for RLHF). By construction, most single-turn RL methods cannot endow LLMs with the ability to perform credit assignment, or reason about their past actions, in multiple turns. How can we design effective and efficient multi-turn RL algorithms for LLMs? In this paper, we develop a framework for building multi-turn RL algorithms for fine-tuning LLMs, that preserves the flexibility of existing single-turn RL methods for LLMs, while accommodating multiple turns, long horizons, and delayed rewards. Our framework builds a hierarchical RL approach and runs two RL algorithms in parallel: a high-level off-policy value-based RL algorithm to aggregate reward over utterances, and a low-level policy gradient RL algorithm that utilizes this high-level value function to train a token policy within each turn. Our hierarchical framework, Actor-Critic Framework with a Hierarchical Structure (ArCHer), can also give rise to other RL methods. Empirically, we find that ArCHer significantly improves efficiency and performance on agent tasks, attaining a sample efficiency of about 100x over existing methods, while also improving with larger model capacity (upto the 7 billion scale).

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Actor-Critic Framework with a Hierarchical Structure (ArCHer)
  • 3.1. Language Generation as a Hierarchical MDP
  • 3.2. Preliminaries: Reinforcement Learning Definitions
  • 3.3. RL Algorithms in the Hierarchical Language MDP
  • 3.4. A Practical Instantiation of ArCHer for Sample-Efficient Online RL
  • 3.5. Other Practical Instantiations of ArCHer
  • 4. Theoretical Analysis
  • 5. Experiments
  • 5.1. Tasks and Environments
  • 5.2. Comparisons and Baseline Approaches
  • 5.3. Results: Sample Efficiency in the Online Setting
  • 5.4. Ablation Study: Importance of Off-Policy Data
  • 5.5. Ablation Study: Alternate Base Models for the High-Level Critic in ArCHer
  • 5.6. Ablation Study: Scaling the Base Model from 100M to 7B Parameters
  • 5.7. Alternate Practical Algorithms
  • 6. Discussion and Conclusion
  • Impact Statement
  • Acknowledgements
  • Author Contributions
  • References
  • Appendices
  • A. Environment and Dataset Details
  • A.1. Framework Summary and Practical Implementation Details
  • B. Offline Algorithm and Practical Considerations
  • C. Additional Baseline Details
  • C.1. Performance of PPO
  • C.2. Additional Reproduction Details for WebShop Experiment
  • D. Additional Experimental Results
  • D.1. Offline ArCHer with IQL and AWR.
  • D.2. TD-Learning v.s. MC Regression
  • D.3. Online IQL Critic Loss
  • E. Reward Hacking
  • F. Hyperparameters
  • G. Proof of Main Theorem
  • G.1. Equivalent Utterance and Token Level MDPs
  • G.2. Fitted Policy Evaluation Subroutine
  • G.3. Assumptions
  • G.4. Proof of Main Theorem
  • G.5. Technical Lemmas

Knowls

  1. Knowl 1 — ArCHer separates utterance-level learning from token generation

    model/method

    ArCHer models a multi-turn language agent as a hierarchy of two linked decision processes. At the high level, a state is the interaction history with the environment and an action is an entire variable-length utterance. At the low level, the agent generates that utterance one token at a time: each low-level state contains the high-level history plus the tokens already generated in the current turn, and each low-level action is the next token. In the practical version, the high-level process advances once per utterance, while the low-level process ends when the utterance ends, typically with an EOS token. An utterance-level value function supplies the terminal learning signal for token generation. This design applies off-policy temporal-difference learning to the high-level process and policy-gradient learning to the low-level process, avoiding token-by-token Bellman backups across long conversations and avoiding maximization over a large space of possible utterances.

  2. Knowl 2 — Online ArCHer learns utterance values from replayed interaction

    model/method

    In online ArCHer, the agent stores environment transitions (s,a,r,s′)(s,a,r,s') in a replay buffer D\mathcal D, where ss is an interaction history, aa is the generated utterance, rr is the reward, and s′s' is the next interaction history. The utterance-level action-value function Qθ(s,a)Q_\theta(s,a) is trained toward a one-step target formed from the observed reward and a delayed state-value estimate; the value function Vψ(s)V_\psi(s) is trained to predict the expected action value of utterances sampled from the current autoregressive token policy πϕ\pi_\phi. With high-level discount γ\gamma, the losses are

    JQ(θ)=E(s,a,r,s′)∼D[(Qθ(s,a)−r−γVψˉ(s′))2],JV(ψ)=Es∼D[Ea∼πϕ(⋅∣s)[(Vψ(s)−Qθˉ(s,a))2]].J_Q(\theta)=\mathbb E_{(s,a,r,s')\sim\mathcal D}\left[(Q_\theta(s,a)-r-\gamma V_{\bar\psi}(s'))^2\right], \qquad J_V(\psi)=\mathbb E_{s\sim\mathcal D}\left[\mathbb E_{a\sim\pi_\phi(\cdot\mid s)}[(V_\psi(s)-Q_{\bar\theta}(s,a))^2]\right].

    Here θ\theta and ψ\psi are the online critic parameters and θˉ,ψˉ\bar\theta,\bar\psi are delayed target parameters. The value target is estimated by sampling utterances from the current policy, while critic training reuses transitions from previous interactions. Target networks are updated by Polyak averaging. The practical implementation uses two independently trained QQ- and VV-heads, sharing a language-model encoder, and uses minima across the paired estimates to reduce overestimation.

  3. Knowl 3 — The utterance critic trains the token policy with a sequence-level advantage

    model/method

    ArCHer updates its autoregressive token policy πϕ\pi_\phi using policy gradients, but gives the same utterance-level advantage to every token in the generated response. For an interaction history ss, let a=(a1,…,aL)a=(a^1,\ldots,a^L) be an utterance of LL tokens, and define A(s,a)=Q(s,a)−V(s)A(s,a)=Q(s,a)-V(s). The online REINFORCE objective is

    Jϕ=Es∼D, a∼πϕ(⋅∣s)[A(s,a)∑i=1Llog⁡πϕ(ai∣s,a1:i−1)].J_\phi=\mathbb E_{s\sim\mathcal D,\,a\sim\pi_\phi(\cdot\mid s)}\left[ A(s,a)\sum_{i=1}^{L}\log\pi_\phi(a^i\mid s,a^{1:i-1})\right].

    The policy is optimized to maximize this objective; the utterance advantage acts as a terminal reward for token generation. For long or diverse utterances, the authors also use an optional token-level baseline V~η(s,a1:i−1)\widetilde V_\eta(s,a^{1:i-1}). It is fit by squared-error regression to the utterance advantage, and the actor then weights each token log-probability by A(s,a)−V~η(s,a1:i−1)A(s,a)-\widetilde V_\eta(s,a^{1:i-1}). This baseline can reduce policy-gradient variance, at the cost of training an additional value model.

  4. Knowl 4 — Offline ArCHer combines an IQL critic with an AWR actor

    model/method

    For a fixed dataset of interactions, ArCHer uses implicit Q-learning (IQL) for the utterance-level value update and advantage-weighted regression (AWR) for the token policy. IQL avoids explicitly maximizing over candidate utterances: it fits the state-value function to an expectile of the available action-value targets, keeping the backup tied to actions represented in the dataset. Its expectile parameter satisfies τ∈[0.5,1)\tau\in[0.5,1). The AWR actor maximizes the likelihood of dataset utterances, weighted by their estimated advantage:

    Jϕ=−E(s,a)∼D[exp⁡(βA(s,a))∑i=1Llog⁡πϕ(ai∣s,a1:i−1)],J_\phi=-\mathbb E_{(s,a)\sim\mathcal D}\left[\exp\bigl(\beta A(s,a)\bigr)\sum_{i=1}^{L}\log\pi_\phi(a^i\mid s,a^{1:i-1})\right],

    where β>0\beta>0 controls the trade-off between imitating the data-generating policy and favoring higher-advantage utterances. The paper notes that larger β\beta gives more aggressive reward maximization but can reduce stability. This combination is intended to avoid the out-of-distribution action overestimation that can result from directly optimizing critic predictions on offline data.

  5. Knowl 5 — Utterance-level critics have favorable estimation requirements and advantage-error bounds

    theoretical result

    The paper analyzes fitted policy evaluation under Bellman-completeness and density-ratio assumptions, comparing an utterance-level critic with a token-level critic on an equivalent task. The utterance-level critic requires weaker Bellman-completeness conditions: a function class that is Bellman complete at the token level is also complete at the utterance level. Under the same trajectory data, the density-ratio coverage requirement is identical at the two levels. Under these assumptions, with maximum utterance length LL and statistical error proportional to N−1/2N^{-1/2} for NN utterance-level transitions, the paper's main result gives a tighter worst-case bound on advantage-estimation error for the utterance-level critic than for the equivalent token-level critic. The authors characterize token-level error as larger by a factor proportional to γL\gamma\sqrt L, where γ\gamma is the discount factor, attributing the gap to error accumulation over tokens. The comparison concerns the stated fitted-policy-evaluation analysis and its assumptions; it is not an unconditional guarantee for arbitrary critics or data.

  6. Knowl 6 — Evaluation uses five multi-turn tasks with delayed or task-directed rewards

    experimental setup

    The experiments evaluate online training on five environments requiring interaction over multiple turns. Detective Game is an interactive fiction task with a 60-step timeout; its optimal solution takes 51 steps and yields a maximum reward of 360, with milestone rewards along the way. Twenty Questions asks an agent to identify one hidden word from 157 possibilities using up to 20 yes/no questions; an incorrect or non-terminating step costs −1-1, while a correct guess ends the episode with reward 0. Twenty Questions Subset uses 10 hidden words, creating a distribution shift from the full-task supervised initialization. Guess My City asks the agent to identify one of 100 hidden cities within 20 free-form question-and-answer turns, with the same reward structure as Twenty Questions. WebShop requires searching, selecting, configuring, and purchasing an item; it gives a dense similarity reward from 0 to 1 and times out after 10 interaction steps. Most experiments initialize methods from supervised instruction-tuned checkpoints trained on suboptimal data. The main policy is GPT-2 and the critic uses RoBERTa-base; the paper also evaluates a 7-billion-parameter Mistral policy. Main online comparisons report medians across three random seeds.

  7. Knowl 7 — ArCHer improves online sample efficiency and task performance

    empirical result

    Across the online comparisons with token-level PPO, filtered behavior cloning, and the utterance-level CHAI baseline, ArCHer steadily improves with additional interaction and achieves the strongest final performance on the four more challenging tasks: Twenty Questions Subset, Twenty Questions, Guess My City, and WebShop. On Detective Game, ArCHer matches the best prior approach. The sample-efficiency comparison is especially large on Twenty Questions: PPO requires more than 100,000 samples to reach an average return just above −17-17, whereas ArCHer reaches that return with fewer than 1,000 samples, which the paper reports as at least a 100×100\times improvement in sample efficiency. On WebShop, fine-tuning a GPT-2-base policy with ArCHer also exceeds the tested GPT-3.5 prompting baselines, including an expert-written prompt and ReAct. The experiments attribute PPO's weaker data efficiency to its need to collect fresh on-policy rollouts rather than reusing earlier interactions.

  8. Knowl 8 — Ablations support replay, token baselines, and scaling the actor

    empirical result

    Several ablations test components and model choices in ArCHer. On Guess My City, using the optional token-level baseline substantially improves performance over the version without it, consistent with lower variance being useful for the task's longer, more diverse utterances. A replay buffer containing only the most recent 48 rollouts produces unstable learning, while larger buffers are more stable; enlarging the buffer beyond a point has little further effect. On Twenty Questions, a RoBERTa critic learns somewhat faster initially than a GPT-2 critic, but their later learning curves are similar. On Twenty Questions Subset, replacing the 100-million-parameter GPT-2 actor with a zero-shot-capable Mistral 7B actor leads to much faster learning. These observations support the method's use of off-policy data and show that the framework can use different critic architectures and a larger policy model in the tested settings.

  9. Knowl 9 — Offline results favor IQL plus AWR over behavior cloning and unregularized policy gradients

    data/table

    The offline Twenty Questions comparison evaluates ArCHer variants and imitation-learning baselines using 1,280 trajectories across five random seeds for ArCHer and filtered behavior cloning. Return is the reported evaluation metric; higher (less negative) is better. The results show that the IQL critic paired with AWR performs best among the tested variants, while directly applying REINFORCE with the IQL critic collapses. Adding behavior-cloning regularization to that REINFORCE variant improves its return but does not match AWR. The IQL-plus-AWR and SARSA-plus-AWR variants both exceed unfiltered and filtered behavior cloning, while the reported SARSA variant is weaker than IQL plus AWR.

    Could not parse LaTeX table
  10. Knowl 10 — The evaluated method still needs substantial interaction and can fail under imperfect feedback

    limitation

    The paper identifies limited interaction efficiency as an unresolved issue: despite the reported gains, the methods still require thousands of environment interactions, and the experiments focus on tasks with computational rewards. The evaluation also exposes vulnerabilities in free-form simulated feedback. In Guess My City, the agent can exploit or become confused by imperfections in the oracle; the authors replace oracle responses containing the target city name with a fixed refusal, but example rollouts still show repetition and collapse in out-of-distribution states. In Twenty Questions, unsuccessful episodes can repeat the same question. These findings limit how broadly the reported task performance establishes reliable behavior with human feedback or in less controlled environments.

Coverage note — The paper's detailed derivations, proof-only lemmas, full hyperparameter grids, prompt transcripts, and individual example rollouts are omitted because they do not add standalone contributed knowledge beyond the stated method, guarantees, empirical comparisons, and limitations.

References

  1. 1.Abdulhai, M., White, I., Snell, C., Sun, C., Hong, J., Zhai, Y., Xu, K., and Levine, S. Lmrl gym: Benchmarks for multi-turn reinforcement learning with language models, 2023.
  2. 2.Bacon, P., Harb, J., and Precup, D. The option-critic architecture. CoRR, abs/1609.05140, 2016. URL http://arxiv.org/abs/1609.05140.
  3. 3.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, S., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022.
  4. 4.Budzianowski, P., Wen, T., Tseng, B., Casanueva, I., Ultes, S., Ramadan, O., and Gasic, M. Multiwoz - A large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. CoRR, abs/1810.00278, 2018. URL http://arxiv.org/abs/1810.00278.
  5. 5.Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D. Open problems and fundamental limitations of reinforcement learning from human feedback, 2023.
  6. 6.Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., and Yao, S. Fireact: Toward language agent fine-tuning. ArXiv, abs/2310.05915, 2023. URL https://api.semanticscholar.org/CorpusID:263829338.
  7. 7.Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences, 2023.
  8. 8.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models, 2022.
  9. 9.Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T. Raft: Reward ranked finetuning for generative foundation model alignment, 2023.
  10. 10.Foster, D. J., Krishnamurthy, A., Simchi-Levi, D., and Xu, Y. Offline reinforcement learning: Fundamental barriers for value function approximation. CoRR, abs/2111.10919, 2021. URL https://arxiv.org/abs/2111.10919.
  11. 11.Fujimoto, S. and Gu, S. S. A minimalist approach to offline reinforcement learning. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  12. 12.Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pp. 10835–10866. PMLR, 2023.
  13. 13.Ghosal, D., Shen, S., Majumder, N., Mihalcea, R., and Poria, S. Cicero: A dataset for contextualized commonsense inference in dialogues. In Annual Meeting of the Association for Computational Linguistics, 2022. URL https://api.semanticscholar.org/CorpusID:247762111.
  14. 14.Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokra, S., Fernando, N., Wu, B., Foley, R., ´ Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G. Improving alignment of dialogue agents via targeted human judgements, 2022.
  15. 15.Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023.
  16. 16.Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. CoRR, abs/1801.01290, 2018. URL http://arxiv.org/abs/1801.01290.
  17. 17.Hausknecht, M. J., Ammanabrolu, P., Cotˆ e, M., and Yuan, X. ´ Interactive fiction games: A colossal adventure. CoRR, abs/1909.05398, 2019. URL http://arxiv.org/abs/1909.05398.
  18. 18.Hong, J., Levine, S., and Dragan, A. Zero-shot goal-directed dialogue via rl on imagined conversations, 2023.
  19. 19.Jang, Y., Lee, J., and Kim, K.-E. GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=qaxhBG1UUaS.
  20. 20.Jaques, N., Shen, J. H., Ghandeharioun, A., Ferguson, C., Lapedriza, A., Jones, N., Gu, S. S., and Picard, R. W. ` Human-centric dialog training via offline reinforcement learning. CoRR, abs/2010.05848, 2020. URL https://arxiv.org/abs/2010.05848.
  21. 21.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mistral 7b, 2023.
  22. 22.Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C. L., Phang, J., Bowman, S. R., and Perez, E. Pretraining language models with human preferences, 2023.
  23. 23.Kostrikov, I., Nair, A., and Levine, S. Offline reinforcement learning with implicit q-learning, 2021.
  24. 24.Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. Stabilizing off-policy q-learning via bootstrapping error reduction. Advances in Neural Information Processing Systems, 32, 2019.
  25. 25.Kumar, A., Zhou, A., Tucker, G., and Levine, S. Conservative q-learning for offline reinforcement learning. CoRR, abs/2006.04779, 2020. URL https://arxiv.org/abs/2006.04779.
  26. 26.Lee, S., Zhu, Q., Takanobu, R., Li, X., Zhang, Y., Zhang, Z., Li, J., Peng, B., Li, X., Huang, M., and Gao, J. Convlab: Multi-domain end-to-end dialog system platform. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  27. 27.Li, Y., Choi, D. H., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Tom, Eccles, Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., de, C., d’Autume, M., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Alexey, Cherepanov, Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., de, N., Freitas, Kavukcuoglu, K., and Vinyals, O. Competition-level code generation with alphacode. Science, 378:1092 – 1097, 2022. URL https://api.semanticscholar.org/CorpusID:246527904.
  28. 28.Lin, X. V., Wang, C., Zettlemoyer, L., and Ernst, M. D. Nl2bash: A corpus and semantic parser for natural language interface to the linux operating system. CoRR, abs/1802.08979, 2018. URL http://arxiv.org/abs/1802.08979.
  29. 29.Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J. Agentbench: Evaluating llms as agents, 2023.
  30. 30.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692, 2019. URL http://arxiv.org/abs/1907.11692.
  31. 31.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A. Playing atari with deep reinforcement learning. CoRR, abs/1312.5602, 2013. URL http://arxiv.org/abs/1312.5602.
  32. 32.Nachum, O., Gu, S., Lee, H., and Levine, S. Data-efficient hierarchical reinforcement learning. CoRR, abs/1805.08296, 2018. URL http://arxiv.org/abs/1805.08296.
  33. 33.Nasiriany, S., Pong, V. H., Lin, S., and Levine, S. Planning with goal-conditioned policies. CoRR, abs/1911.08453, 2019. URL http://arxiv.org/abs/1911.08453.
  34. 34.Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A. The primacy bias in deep reinforcement learning. In International conference on machine learning, pp. 16828–16847. PMLR, 2022.
  35. 35.OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mely, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, ´ A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, P., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B. Gpt-4 technical report, 2023.
  36. 36.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L. E., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. J. Training language models to follow instructions with human feedback. ArXiv, abs/2203.02155, 2022. URL https://api.semanticscholar.org/CorpusID:246426909.
  37. 37.Park, S., Ghosh, D., Eysenbach, B., and Levine, S. Hiql: Offline goal-conditioned rl with latent states as actions, 2024.
  38. 38.Paulus, R., Xiong, C., and Socher, R. A deep reinforced model for abstractive summarization, 2017.
  39. 39.Peng, X. B., Kumar, A., Zhang, G., and Levine, S. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. CoRR, abs/1910.00177, 2019. URL http://arxiv.org/abs/1910.00177.
  40. 40.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019. URL https://api.semanticscholar.org/CorpusID:160025533.
  41. 41.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model, 2023.
  42. 42.Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization. 2022. URL https://arxiv.org/abs/2210.01241.
  43. 43.Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization, 2023.
  44. 44.Ranzato, M., Chopra, S., Auli, M., and Zaremba, W. Sequence level training with recurrent neural networks, 2015.
  45. 45.Schick, T., Dwivedi-Yu, J., Dess`ı, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language models can teach themselves to use tools, 2023.
  46. 46.Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P. High-dimensional continuous control using generalized advantage estimation. CoRR, abs/1506.02438, 2015. URL https://api.semanticscholar.org/CorpusID:3075448.
  47. 47.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017. URL http://arxiv.org/abs/1707.06347.
  48. 48.Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K. Dynamics-aware unsupervised discovery of skills. CoRR, abs/1907.01657, 2019. URL http://arxiv.org/abs/1907.01657.
  49. 49.Shinn, N., Cassano, F., Labash, B., Gopinath, A., Narasimhan, K., and Yao, S. Reflexion: Language agents with verbal reinforcement learning. 2023. URL https://api.semanticscholar.org/CorpusID:258833055.
  50. 50.Snell, C., Kostrikov, I., Su, Y., Yang, M., and Levine, S. Offline rl for natural language generation with implicit language q learning, 2023.
  51. 51.Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W. Hybrid rl: Using both offline and online data can make rl efficient, 2023.
  52. 52.Sutton, R. S. Between mdps and semi-mdps : Learning , planning , and representing knowledge at multiple temporal scales. 1998. URL https://api.semanticscholar.org/CorpusID:2191003.
  53. 53.Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In Solla, S., Leen, T., and Muller, K. (eds.), Advances in Neural Information Processing Systems, volume 12. MIT Press, 1999. URL https://proceedings.neurips.cc/paper_files/paper/1999/file/464d828b85b0bed98e80ade0a5c43b0f-Paper.pdf.
  54. 54.Szot, A., Schwarzer, M., Agrawal, H., Mazoure, B., Talbott, W., Metcalf, K., Mackraz, N., Hjelm, D., and Toshev, A. Large language models as generalizable policies for embodied tasks, 2023.
  55. 55.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models, 2023.
  56. 56.van Hasselt, H., Guez, A., and Silver, D. Deep reinforcement learning with double q-learning. CoRR, abs/1509.06461, 2015. URL http://arxiv.org/abs/1509.06461.
  57. 57.Verma, S., Fu, J., Yang, M., and Levine, S. Chai: A chatbot ai for task-oriented dialogue with offline reinforcement learning, 2022.
  58. 58.Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. Feudal networks for hierarchical reinforcement learning. CoRR, abs/1703.01161, 2017. URL http://arxiv.org/abs/1703.01161.
  59. 59.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L. J., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. ArXiv, abs/2305.16291, 2023. URL https://api.semanticscholar.org/CorpusID:258887849.
  60. 60.Williams, R. J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8:229–256, 2004. URL https://api.semanticscholar.org/CorpusID:19115634.
  61. 61.Wu, Y. and Hu, B. Learning to extract coherent summary via deep reinforcement learning, 2018.
  62. 62.Xie, T., Cheng, C., Jiang, N., Mineiro, P., and Agarwal, A. Bellman-consistent pessimism for offline reinforcement learning. CoRR, abs/2106.06926, 2021. URL https://arxiv.org/abs/2106.06926.
  63. 63.Yang, J., Prabhakar, A., Narasimhan, K., and Yao, S. Inter-code: Standardizing and benchmarking interactive coding with execution feedback, 2023a.
  64. 64.Yang, K., Swope, A. M., Gu, A., Chalamala, R., Song, P., Yu, S., Godil, S., Prenger, R., and Anandkumar, A. LeanDojo: Theorem Proving with Retrieval-Augmented Language Models. arXiv preprint arXiv:2306.15626, 2023b.
  65. 65.Yao, S., Chen, H., Yang, J., and Narasimhan, K. Webshop: Towards scalable real-world web interaction with grounded language agents, 2023a.
  66. 66.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models, 2023b.
  67. 67.Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F. Rrhf: Rank responses to align language models with human feedback without tears, 2023.
  68. 68.Zanette, A. When is realizability sufficient for off-policy reinforcement learning?, 2023.
  69. 69.Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., and Tang, J. Agenttuning: Enabling generalized agent abilities for llms, 2023.
  70. 70.Zhan, W., Huang, B., Huang, A., Jiang, N., and Lee, J. D. Offline reinforcement learning with realizability and single-policy concentrability, 2022.
  71. 71.Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Bisk, Y., Fried, D., Alon, U., and Neubig, G. Webarena: A realistic web environment for building autonomous agents. ArXiv, abs/2307.13854, 2023a. URL https://api.semanticscholar.org/CorpusID:260164780.
  72. 72.Zhou, Y., Sekhari, A., Song, Y., and Sun, W. Offline data enhanced on-policy policy gradient with provable guarantees, 2023b.
  73. 73.Zhu, Q., Zhang, Z., Fang, Y., Li, X., Takanobu, R., Li, J., Peng, B., Gao, J., Zhu, X., and Huang, M. Convlab-2: An open-source toolkit for building, evaluating, and diagnosing dialogue systems. CoRR, abs/2002.04793, 2020. URL https://arxiv.org/abs/2002.04793.
  74. 74.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P. F., and Irving, G. Fine-tuning language models from human preferences. CoRR, abs/1909.08593, 2019. URL http://arxiv.org/abs/1909.08593.

Citation

MLA
Zhou, Y., et al. “ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL”. arXiv, 2024, http://arxiv.org/abs/2402.19446v1.
APA
Zhou, Y., Zanette, A., Pan, J., Levine, S., & Kumar, A. (2024). ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL. arXiv. http://arxiv.org/abs/2402.19446v1
Chicago
Zhou, Y., A. Zanette, J. Pan, S. Levine, and A. Kumar. 2024. “ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL”. arXiv. http://arxiv.org/abs/2402.19446v1.
Harvard
Zhou, Y. et al. (2024) “ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.19446v1.
Vancouver
1. Zhou Y, Zanette A, Pan J, Levine S, Kumar A (2024) ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL. arXiv

BibTeX

@article{zhou2024archer,
  title = {ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL},
  author = {Zhou, Yifei and Zanette, Andrea and Pan, Jiayi and Levine, Sergey and Kumar, Aviral},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.19446v1},
  eprint = {2402.19446}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/