CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning
Siddharth VermaJustin FuSherry YangSergey Levine
Presents CHAI, a framework that combines pre-trained language models with offline reinforcement learning to train task-oriented dialogue agents directly from static human conversation datasets without requiring live interactions or simulated user models.
Building automated dialogue agents that can carry out goal-oriented tasks, such as commercial negotiations or customer service, is critical for scaling interactive human-machine interfaces. Standard approaches rely either on supervised imitation learning—which generates natural speech but fails to actively optimize for task success—or online reinforcement learning, which requires impractically slow and costly live trial-and-error interactions with humans. Existing alternatives often depend on hand-crafted dialogue templates or simulated human models that easily degrade into nonsensical language. The article addresses this bottleneck by evaluating whether offline reinforcement learning can train effective, goal-driven conversational agents directly from fixed, pre-recorded human interaction logs without real-time online exploration.
The main objective of the article is to demonstrate CHAI (Chatbot AI), a framework that couples pre-trained language models with offline reinforcement learning to generate fluent dialogue while systematically pursuing quantifiable task goals. The researchers evaluate this method on a buyer-seller negotiation benchmark comprising 6,682 Craigslist advertisements and dialogues. The approach pairs a fine-tuned GPT-2 language model, which generates natural candidate utterances, with a learned evaluation function (a critic or Q-function) trained on static dialogue logs to score and select optimal responses. To prevent the model from exploiting unfamiliar text states, the article investigates three offline reinforcement learning regularization techniques: proposal sampling, conservative Q-learning, and behavior regularization.
The findings show that combining offline reinforcement learning with language models significantly outperforms prior systems across key performance dimensions. First, the conservative Q-learning variant of CHAI achieved the highest overall agreement rate at 88.7% across five simulated buyer personas, significantly exceeding the pure language model baseline (32.9%) and prior reinforcement learning benchmarks (43.2% to 80.4%). Second, CHAI maintained high normalized revenue, generating roughly 0.65 to 0.67 of listing value, substantially outperforming the pure language model baseline at 0.21. Third, in a controlled human user study across 16 participants, CHAI scored highest in overall user perception (15.84 out of 20 total score), statistically outperforming retrieval-based systems (11.25) and language model baselines (12.84) in coherency, task focus, and human-likeness. Finally, performance remained consistently strong across all three offline regularization variants, showing that steering language models via learned value functions is the primary driver of success rather than the specific regularization algorithm.
These results imply that organizations can develop capable, goal-oriented conversational systems using pre-existing conversation transcripts, eliminating the high costs and risks associated with live online reinforcement learning trials. Unlike rigid template-based methods, this architecture provides natural flexibility without requiring extensive hand-engineered dialogue acts. Decision-makers seeking to automate transactional or advisory dialogues should consider pilot implementations that combine pre-trained language models with offline value scoring. Teams should focus technical effort on designing balanced reward structures—such as combining transaction success rewards with penalties for deal breakdowns—rather than inventing complex rule-based dialogue managers.
The findings have clear limitations: the system was tested exclusively on a single bilateral negotiation dataset with immediate transactional outcomes, and unconstrained reward maximization can theoretically encourage untruthful responses if truthfulness is not explicitly penalized. Nevertheless, because the results demonstrate high statistical significance across diverse simulated buyers and human evaluators, stakeholders can have strong confidence in the core finding: offline reinforcement learning provides a viable, scalable method for steering language models toward strategic real-world goals.
- Paper: Decision Transformer: Reinforcement Learning via Sequence Modeling, Lili Chen et al. (2021). Its offline-RL framing and discussion of Conservative Q-Learning introduce key methods CHAI uses to select goal-directed responses from fixed interaction logs.
- Paper: Deep Reinforcement Learning for Dialogue Generation, Jiwei Li et al. (2016). It shows how reinforcement learning can optimize dialogue beyond next-response imitation, clarifying the strategic dialogue-generation problem CHAI carries into task-oriented conversations.
- Paper: ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL, Yifei Zhou et al. (2024). ArCHer extends language-model agents from value-guided dialogue decisions to hierarchical, multi-turn reinforcement learning with delayed outcomes.
- Paper: From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL, Wenyue Hua et al. (2026). SocialRL carries strategic negotiation beyond CHAI’s buyer-seller setting by training language-model delegates to reason and preserve value across multiple negotiation domains.
- Paper: Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning, Xiao Yu et al. (2023). GDP-ZERO continues goal-oriented dialogue research by replacing learned response selection with language-model-guided tree-search planning for strategic conversational moves.
