ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Yifei ZhouAndrea ZanetteJiayi PanSergey LevineAviral Kumar
Introduces a hierarchical reinforcement learning framework that couples high-level off-policy value estimation across multi-turn interactions with low-level token-generation policy gradients, achieving a hundredfold improvement in sample efficiency when training language model agents on delayed-reward tasks.
Large language models are increasingly deployed as autonomous agents to handle complex, multi-step interactions such as dialogue, web navigation, and tool use. Standard reinforcement learning approaches for language models typically optimize rewards over a single turn, such as standard human feedback alignment, which prevents agents from planning ahead, gathering information, or attributing delayed success to earlier actions. Existing multi-turn methods face severe operational trade-offs: on-policy algorithms discard past interaction data and require prohibitively expensive real-time environment sampling, whereas off-policy algorithms that evaluate individual tokens struggle with long horizons and compound estimation errors over multiple conversational turns.
The article develops and evaluates Actor-Critic Framework with a Hierarchical Structure (ArCHer), a reinforcement learning framework designed to train language model agents efficiently across multi-turn interactions with delayed feedback. ArCHer introduces a two-level hierarchy that decouples utterance-level value estimation from token-level generation. At the high level, it applies off-policy temporal-difference learning across full utterances, enabling efficient sample reuse from past interactions without searching over an intractable action space. At the low level, it uses policy gradient optimization within each turn, using the high-level value estimate as the reward signal to guide token generation in a self-contained optimization loop.
The authors evaluated the framework across theoretical convergence criteria and empirical benchmarks comprising language games, sequential dialogue, and web shopping tasks. The evaluations compared ArCHer against standard on-policy policy optimization, filtered behavioral cloning, and utterance-ranking methods, utilizing base language models ranging up to seven billion parameters.
The findings demonstrate four critical outcomes. First, ArCHer achieves approximately a 100-fold improvement in sample efficiency compared to on-policy policy optimization baselines, reaching target performance with under 1,000 trajectories where standard baselines require upwards of 100,000. Second, the framework delivers superior final task performance, outperforming filtered imitation learning and utterance-ranking baselines across complex reasoning tasks, while a fine-tuned smaller model surpassed prompting techniques on proprietary large models in web navigation. Third, theoretical analysis proves that estimating advantages at the utterance level reduces statistical error accumulation by a factor proportional to the square root of utterance length compared to token-level estimators. Fourth, the architecture scales effectively with model capacity, showing faster policy improvement when scaled from smaller models to a seven-billion-parameter foundation model.
These results demonstrate that multi-turn reinforcement learning can be made computationally feasible and data-efficient without requiring massive online sampling budgets or brittle heuristic prompting. Organizations deploying language model agents can significantly reduce environment interaction costs and computational overhead while achieving reliable multi-turn goal pursuit. The approach allows practitioners to adapt existing single-turn alignment infrastructure directly to multi-turn tasks.
Decision-makers considering multi-turn agent training should prioritize hierarchical reinforcement learning designs over pure on-policy methods or simple demonstration cloning. When deploying agents in environments with long conversational turns, teams should incorporate token-level baseline value functions to maintain training stability. Furthermore, organizations should run initial offline or hybrid pilot implementations using logged interaction data before investing in full-scale online training environments.
The evaluations were conducted in simulated environments with programmed reward functions rather than direct human interactive trials, and the system still requires several thousand interactions to converge. Additionally, the authors note that free-form responses from imperfect auxiliary models can occasionally cause agents to fall into repetitive loops in out-of-distribution states. Confidence in the algorithmic efficiency gains is high based on the empirical and theoretical alignment, but performance in open-ended production environments with human users will require validation.
- Paper: Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation, Tejas D. Kulkarni et al. (2016). Its two-level hierarchy of goal selection and low-level control supplies the hierarchical-RL foundation ArCHer adapts to multi-turn language-model training.
- Paper: TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents, Leitian Tao et al. (2026). TRACE continues ArCHer’s focus on long-horizon agent credit assignment by deriving turn-level rewards that make each interaction’s contribution to eventual success more explicit.
