CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning

Siddharth VermaJustin FuSherry YangSergey Levine

article2022NAACL51 citations

Presents CHAI, a framework that combines pre-trained language models with offline reinforcement learning to train task-oriented dialogue agents directly from static human conversation datasets without requiring live interactions or simulated user models.

Listen

Building automated dialogue agents that can carry out goal-oriented tasks, such as commercial negotiations or customer service, is critical for scaling interactive human-machine interfaces. Standard approaches rely either on supervised imitation learning—which generates natural speech but fails to actively optimize for task success—or online reinforcement learning, which requires impractically slow and costly live trial-and-error interactions with humans. Existing alternatives often depend on hand-crafted dialogue templates or simulated human models that easily degrade into nonsensical language. The article addresses this bottleneck by evaluating whether offline reinforcement learning can train effective, goal-driven conversational agents directly from fixed, pre-recorded human interaction logs without real-time online exploration.

The main objective of the article is to demonstrate CHAI (Chatbot AI), a framework that couples pre-trained language models with offline reinforcement learning to generate fluent dialogue while systematically pursuing quantifiable task goals. The researchers evaluate this method on a buyer-seller negotiation benchmark comprising 6,682 Craigslist advertisements and dialogues. The approach pairs a fine-tuned GPT-2 language model, which generates natural candidate utterances, with a learned evaluation function (a critic or Q-function) trained on static dialogue logs to score and select optimal responses. To prevent the model from exploiting unfamiliar text states, the article investigates three offline reinforcement learning regularization techniques: proposal sampling, conservative Q-learning, and behavior regularization.

The findings show that combining offline reinforcement learning with language models significantly outperforms prior systems across key performance dimensions. First, the conservative Q-learning variant of CHAI achieved the highest overall agreement rate at 88.7% across five simulated buyer personas, significantly exceeding the pure language model baseline (32.9%) and prior reinforcement learning benchmarks (43.2% to 80.4%). Second, CHAI maintained high normalized revenue, generating roughly 0.65 to 0.67 of listing value, substantially outperforming the pure language model baseline at 0.21. Third, in a controlled human user study across 16 participants, CHAI scored highest in overall user perception (15.84 out of 20 total score), statistically outperforming retrieval-based systems (11.25) and language model baselines (12.84) in coherency, task focus, and human-likeness. Finally, performance remained consistently strong across all three offline regularization variants, showing that steering language models via learned value functions is the primary driver of success rather than the specific regularization algorithm.

These results imply that organizations can develop capable, goal-oriented conversational systems using pre-existing conversation transcripts, eliminating the high costs and risks associated with live online reinforcement learning trials. Unlike rigid template-based methods, this architecture provides natural flexibility without requiring extensive hand-engineered dialogue acts. Decision-makers seeking to automate transactional or advisory dialogues should consider pilot implementations that combine pre-trained language models with offline value scoring. Teams should focus technical effort on designing balanced reward structures—such as combining transaction success rewards with penalties for deal breakdowns—rather than inventing complex rule-based dialogue managers.

The findings have clear limitations: the system was tested exclusively on a single bilateral negotiation dataset with immediate transactional outcomes, and unconstrained reward maximization can theoretically encourage untruthful responses if truthfulness is not explicitly penalized. Nevertheless, because the results demonstrate high statistical significance across diverse simulated buyers and human evaluators, stakeholders can have strong confidence in the core finding: offline reinforcement learning provides a viable, scalable method for steering language models toward strategic real-world goals.

Cover for CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning

Abstract

Conventionally, generation of natural language for dialogue agents may be viewed as a statistical learning problem: determine the patterns in human-provided data and generate appropriate responses with similar statistical properties. However, dialogue can also be regarded as a goal directed process, where speakers attempt to accomplish a specific task. Reinforcement learning (RL) algorithms are designed specifically for solving such goal-directed problems, but the most direct way to apply RL – through trial-and-error learning in human conversations, – is costly. In this paper, we study how offline reinforcement learning can instead be used to train dialogue agents entirely using static datasets collected from human speakers. Our experiments show that recently developed offline RL methods can be combined with language models to yield realistic dialogue agents that better accomplish task goals.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 3.1 Negotiation via Task-Oriented Dialogue
  • 3.2 Reinforcement Learning Setup
  • 4 Offline Reinforcement Learning with Language Models
  • 4.1 Q-Learning with Language Models
  • 4.2 Dialogue Generation
  • 4.3 Architecture Details
  • 5 Experiments
  • 5.1 Simulated Evaluation
  • 5.2 Human User Study
  • 6 Discussion and Future Work
  • 7 Acknowledgements
  • References
  • A Appendix
  • A.1 Architecture Details
  • A.2 Experiment Details
  • A.3 Reward Ablation Study
  • A.3.1 User Study Parameters
  • A.4 Additional Qualitative Results
  • A.4.1 CHAI-prop
  • A.4.2 Retrieval
  • A.4.3 Language Model

Knowls

  1. Knowl 1 — CHAI combines a fine-tuned language model with offline Q-learning

    model/method

    CHAI trains a task-oriented dialogue agent from static, reward-labeled conversations by separating fluent language generation from goal-directed response selection. A GPT-2 medium language model is fine-tuned on CraigslistBargain dialogue transcripts, including each listing’s title and description; prices in transcripts are replaced by a special token so generated utterances can be paired with candidate prices. A learned critic assigns values to candidate responses, allowing the system to use offline reinforcement learning (RL) to steer a language model toward task goals without learning through live human interaction or a simulated human model.

    The critic represents dialogue history and candidate-response text using pooled GPT-2 attention embeddings, and combines these representations with normalized prices and one-hot action types. It is a two-layer feedforward network with 256 hidden units per layer and ReLU activations. The language model supplies candidate utterances; the critic learns which actions are valuable from offline transitions.

  2. Knowl 2 — CraigslistBargain is formulated as seller-side negotiation with explicit transaction rewards

    definition

    In the CraigslistBargain task, the CHAI agent is the seller and the environment is the buyer. The task contains 6,682 Craigslist listings and human dialogues. At each turn, a participant can send a message, make an offer, accept an offer, or reject an offer; accepting or rejecting ends the negotiation. The state contains the latest utterance, action type and normalized price, together with the listing context. The action contains an utterance, action type and normalized price; a price is a fraction of the listing price. For message actions, prices in text can be represented by a placeholder and replaced using the price component.

    The seller’s terminal reward is 1010 times the accepted sale price divided by the listing price. A negotiation ending in rejection instead incurs a reward of −20-20. Thus the reward favors higher sale prices while penalizing failed negotiations. The policy is trained to maximize expected discounted reward, with discount factor γ∈(0,1]\gamma\in(0,1].

  3. Knowl 3 — Proposal-sampled Bellman targets keep dialogue actions near the language-model distribution

    algorithm

    CHAI-prop trains a critic Qθ(s,a)Q_\theta(s,a) from offline transitions using a squared Bellman error, but restricts actions in the target calculation to candidates sampled from a language-model-based proposal distribution. For a state-action pair (s,a)(s,a), the target is

    Qtargetprop(s,a)=r(s,a)+γ Es′∼T(⋅∣s,a), {ai}i=1N∼μ(⋅∣s′)[max⁡1≤i≤NQˉ(s′,ai)].Q_{\mathrm{target}}^{\mathrm{prop}}(s,a)=r(s,a)+\gamma\,\mathbb{E}_{s'\sim T(\cdot\mid s,a),\,\{a_i\}_{i=1}^{N}\sim\mu(\cdot\mid s')}\left[\max_{1\leq i\leq N}\bar Q(s',a_i)\right].

    Here, ss and s′s' are dialogue states, aa is an action, r(s,a)r(s,a) is its reward, T(⋅∣s,a)T(\cdot\mid s,a) is the next-state distribution, γ\gamma is the discount factor, NN is the number of proposed actions, μ\mu is the proposal distribution, and Qˉ\bar Q is a target critic. The target critic tracks the learned critic through a soft update. Restricting the maximization to proposed actions is intended to reduce unreliable value estimates on arbitrary, out-of-distribution strings while favoring naturalistic responses.

    The proposal samples utterances from the fine-tuned language model. For computational efficiency, the implementation pre-generates five candidate utterances per dataset transition and resamples among them during training. Candidate prices are sampled uniformly from 70% to 100% of the previous offer. Action type is inferred from generated text: the special tokens offer, accept, and reject select those types; otherwise the action is a message.

  4. Knowl 4 — CHAI-CQL adds a conservative value penalty to proposal-sampled Q-learning

    model/method

    CHAI-CQL uses the same language-model proposals and proposal-sampled Bellman target as CHAI-prop, while adding the CQL(H) regularizer to discourage overly high values for actions unsupported by the offline data. Its objective is

    JCQL(θ)=E(s,a,r,s′)∼D[(Qθ(s,a)−Qtarget(s,a))2]+α Es∼D[log⁡∑a′∈Aexp⁡(Qθ(s,a′))−Ea∼D(⋅∣s)Qθ(s,a)].J^{\mathrm{CQL}}(\theta)=\mathbb{E}_{(s,a,r,s')\sim D}\left[(Q_\theta(s,a)-Q_{\mathrm{target}}(s,a))^2\right]+\alpha\,\mathbb{E}_{s\sim D}\left[\log\sum_{a'\in\mathcal A}\exp(Q_\theta(s,a'))-\mathbb{E}_{a\sim D(\cdot\mid s)}Q_\theta(s,a)\right].

    Here, DD is the offline transition dataset, ss and s′s' are states, aa and a′a' are actions, rr is reward, QθQ_\theta is the learned critic, QtargetQ_{\mathrm{target}} is its Bellman target, A\mathcal A is the action space, and α\alpha weights the regularizer. The regularizer raises values broadly across actions while subtracting the value of logged actions, thereby penalizing values assigned to actions not represented in the data.

  5. Knowl 5 — CHAI-BRAC regularizes candidate prices toward a learned behavior distribution

    model/method

    CHAI-BRAC replaces uniform price proposals with a learned conditional Gaussian price policy πϕ(aprice∣s)\pi_\phi(a_{\mathrm{price}}\mid s) and regularizes that policy toward a behavior prior πB(aprice∣s)\pi_B(a_{\mathrm{price}}\mid s). The prior is a univariate conditional Gaussian whose mean and standard deviation are linear functions of the previous offer. The learned policy maximizes candidate-action value while limiting divergence from the prior:

    max⁡πϕ  Es∼D, a′∼(πϕ,μ)[Q(s,a′)]−DKL ⁣(πϕ(⋅∣s),πB(⋅∣s)).\max_{\pi_\phi}\;\mathbb{E}_{s\sim D,\,a'\sim(\pi_\phi,\mu)}[Q(s,a')]-D_{\mathrm{KL}}\!\left(\pi_\phi(\cdot\mid s),\pi_B(\cdot\mid s)\right).

    Here, DD is the offline dataset, ss is a dialogue state, a′a' is a proposed action, QQ is the critic, μ\mu supplies utterances from the language model and action types uniformly, and DKLD_{\mathrm{KL}} is Kullback–Leibler divergence. The target critic is evaluated on actions sampled from the learned price policy together with the language-model utterance proposal. The price regularization is intended to avoid querying the critic at prices poorly supported by the dataset.

  6. Knowl 6 — CHAI selects among scored responses by softmax sampling

    algorithm

    At inference time, CHAI generates five candidate utterances from the fine-tuned language model and five candidate prices, forms their cross-product, and scores each candidate action with the critic. It then samples an action according to a softmax over the scores rather than always choosing the maximum-valued response:

    p(a∣s)=exp⁡(Q(s,a))∑a~∈C(s)exp⁡(Q(s,a~)),p(a\mid s)=\frac{\exp(Q(s,a))}{\sum_{\tilde a\in\mathcal C(s)}\exp(Q(s,\tilde a))},

    where ss is the current dialogue state, Q(s,a)Q(s,a) is the critic’s value for candidate action aa, and C(s)\mathcal C(s) is the set of candidate actions generated for that state. This selection method allows lower-valued candidates to be chosen occasionally, which the authors report increases response diversity. The five-by-five utterance-price combinations yield 25 candidates before action-type assignment and scoring.

  7. Knowl 7 — Simulated negotiations show high acceptance for CHAI-CQL and comparable revenue to the strongest baseline

    data/table

    The authors evaluated the three CHAI variants against five buyer agents: rule-based, stingy rule-based, and He et al. agents optimized for utility, fairness, or dialogue length. Each overall result is the mean across the five opponents; reported means and standard deviations are over 200 trials. Acceptance is the percentage of negotiations ending in acceptance, and revenue is the sale price divided by listing price, with rejected negotiations contributing zero revenue. CHAI-CQL achieved the highest mean acceptance rate among the CHAI variants and a mean revenue close to the strongest retrieval baseline on these aggregate measures.

    Method Acceptance rate (%) Normalized revenue
    CHAI-prop 81.9 0.65 ±\pm 0.34
    CHAI-CQL 88.7 0.67 ±\pm 0.29
    CHAI-BRAC 79.8 0.65 ±\pm 0.34
    Fine-tuned language model 32.9 0.21 ±\pm 0.31
    He et al. (2018), utility 42.4 0.42 ±\pm 0.49
    He et al. (2018), fairness 72.8 0.54 ±\pm 0.35
    He et al. (2018), length 80.4 0.66 ±\pm 0.36
    Lewis et al. (2017), RL 78.1 0.28 ±\pm 0.33
    Lewis et al. (2017), supervised 43.2 0.28 ±\pm 0.39

    The paper reports a chi-squared test comparing CHAI-CQL with the dialogue-length retrieval agent found a statistically significant acceptance-rate difference (p<1.96×10−9p<1.96\times10^{-9}). The revenue difference was not statistically significant according to the reported t-test (p<0.946p<0.946). The three CHAI variants performed comparably overall, while the fine-tuned language-model baseline had substantially lower acceptance and revenue.

  8. Knowl 8 — Human evaluators rated CHAI higher on task-relevant qualities than the retrieval and language-model baselines

    empirical result

    Sixteen participants each negotiated twice with each of three agents: CHAI-prop, the utility-optimized retrieval agent of He et al. (2018), and a fine-tuned language-model baseline. Participants rated fluency, coherency, on-topicness, and human-likeness on a five-point Likert scale; the table reports means and standard deviations over 32 trials per agent. CHAI-prop had the highest mean on coherency, on-topicness, human-likeness, and total score. Its fluency score was close to that of the language-model baseline. Repeated-measures ANOVA found an agent effect for every rating metric, with p<0.01p<0.01 in each case.

    Agent Fluency Coherency On-topic Human-likeness Total
    CHAI-prop 4.31 ±\pm 0.97 3.91 ±\pm 1.17 4.16 ±\pm 0.99 3.47 ±\pm 1.27 15.84 ±\pm 3.86
    He et al. (2018), utility 3.56 ±\pm 1.34 2.47 ±\pm 1.39 3.09 ±\pm 1.40 2.13 ±\pm 1.13 11.25 ±\pm 4.50
    Language model 4.06 ±\pm 1.11 2.66 ±\pm 1.36 3.63 ±\pm 1.18 2.50 ±\pm 1.10 12.84 ±\pm 3.66

    The paper reports that CHAI and the language model had similar fluency, consistent with their shared use of GPT-2, while CHAI’s offline-RL selection improved ratings of task alignment and dialogue quality relative to the language-model baseline.

  9. Knowl 9 — Negotiation behavior changes substantially with the reward design

    empirical result

    A simulated ablation evaluated five CHAI reward variants over 50 negotiations against a rule-based buyer. The final reward gives 1010 times the accepted sale price normalized by listing price and a −20-20 penalty for rejection. The penalty variant increases the rejection penalty; the accept variant gives +20+20 for acceptance and −20-20 for rejection regardless of price; the utility variant rewards only normalized sale price multiplied by 1010; and the fair variant rewards a midpoint between buyer and seller target prices. Prices offered, prices accepted, and revenue are fractions of listing price; reported price and revenue values include standard deviations where shown.

    Reward variant Accept rate Prices offered Prices accepted Revenue
    CHAI(accept) 0.74 0.80 ±\pm 0.15 0.74 ±\pm 0.13 0.55 ±\pm 0.35
    CHAI(fair) 0.76 0.80 ±\pm 0.15 0.75 ±\pm 0.13 0.57 ±\pm 0.34
    CHAI(penalty) 0.90 0.77 ±\pm 0.14 0.77 ±\pm 0.12 0.68 ±\pm 0.26
    CHAI(utility) 0.34 0.96 ±\pm 0.29 0.76 ±\pm 0.13 0.29 ±\pm 0.42
    CHAI(final) 0.66 0.84 ±\pm 0.14 0.87 ±\pm 0.07 0.51 ±\pm 0.38

    The results show a trade-off: acceptance-oriented rewards increased acceptance but tended to produce lower offers, whereas pure utility produced the highest offers and the lowest acceptance rate. The selected final reward balanced sale-price utility with a rejection penalty; the paper emphasizes that reward choice materially changes learned negotiation behavior.

  10. Knowl 10 — The demonstrated system is limited by reward specification and task-specific design

    limitation

    The paper demonstrates CHAI on one negotiation task, and notes that the same architecture—including its price-prediction component—may not transfer directly to other domains. The authors identify value-based response selection as more broadly applicable than the task-specific architecture. They also caution that reward maximization does not by itself ensure truthful responses: without reward terms for truthfulness, the agent has no objective reason to be truthful. Although the authors suggest that improved reward design could provide a direct way to influence such behavior, they acknowledge that designing suitable rewards is difficult.

Coverage note — Omitted the appendix’s illustrative conversation transcripts and the opponent-by-opponent breakdown of simulated results; the transcripts are examples rather than additional findings, and the aggregate comparison preserves the main quantitative conclusion.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977.
  2. 2.Layla El Asri, Jing He, and Kaheer Suleman. 2016. A sequence-to-sequence model for user simulation in spoken dialogue systems. arXiv preprint arXiv:1607.00070.
  3. 3.Chieh-Yang Chen, Pei-Hsin Wang, Shih-Chieh Chang, Da-Cheng Juan, Wei Wei, and Jia-Yu Pan. 2020. Air-concierge: Generating task-oriented dialogue via efficient large-scale knowledge retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 884–897.
  4. 4.Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan, and William Yang Wang. 2019. Semantically conditioned dialog response generation via hierarchical disentangled self-attention. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3696–3709.
  5. 5.Grace Chung. 2004. Developing a flexible spoken dialog system using simulation. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 63–70.
  6. 6.Ondˇrej Dušek and Filip Jurcicek. 2016. Sequence-to-sequence generation for spoken dialogue via deep syntax trees and strings. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 45–51.
  7. 7.Wieland Eckert, Esther Levin, and Roberto Pieraccini. 1997. User modeling for spoken dialogue system evaluation. In 1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings, pages 80–87. IEEE.
  8. 8.Mihail Eric and Christopher D Manning. 2017. A copy-augmented sequence-to-sequence architecture gives good performance on task-oriented dialogue. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 468–473.
  9. 9.Mehdi Fatemi, Layla El Asri, Hannes Schulz, Jing He, and Kaheer Suleman. 2016. Policy networks with two-stage training for dialogue systems. In Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 101–110.
  10. 10.Scott Fujimoto, David Meger, and Doina Precup. 2019. Off-policy deep reinforcement learning without exploration. In International Conference on Machine Learning, pages 2052–2062.
  11. 11.Jianfeng Gao, Michel Galley, and Lihong Li. 2018. Neural approaches to conversational ai. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, pages 1371–1374.
  12. 12.Milica Gašic, Filip Jur ´ cíˇ cek, Blaise Thomson, Kai Yu, ˇ and Steve Young. 2011. On-line policy optimisation of spoken dialogue systems via live interaction with human subjects. In 2011 IEEE Workshop on Automatic Speech Recognition & Understanding, pages 312–317. IEEE.
  13. 13.Kallirroi Georgila, James Henderson, and Oliver Lemon. 2006. User simulation for spoken dialogue systems: Learning and evaluation. In Ninth International Conference on Spoken Language Processing.
  14. 14.Kallirroi Georgila and David Traum. 2011. Reinforcement learning of argumentation dialogue policies in negotiation. In Twelfth Annual Conference of the International Speech Communication Association.
  15. 15.Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu. 2020. Emaq: Expected-max q-learning operator for simple yet effective offline and online rl. arXiv preprint arXiv:2007.11091.
  16. 16.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, pages 1861–1870. PMLR.
  17. 17.He He, Derek Chen, Anusha Balakrishnan, and Percy Liang. 2018. Decoupling strategy and generation in negotiation dialogues. In Conference on Empirical Methods in Natural Language Processing, pages 2333–2343. Association for Computational Linguistics.
  18. 18.Peter A Heeman. 2009. Representing the reinforcement learning state in a negotiation dialogue. In 2009 IEEE Workshop on Automatic Speech Recognition & Understanding, pages 450–455. IEEE.
  19. 19.James Henderson, Oliver Lemon, and Kallirroi Georgila. 2008. Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets. Computational Linguistics, 34(4):487–511.
  20. 20.Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020. A simple language model for task-oriented dialogue. In Advances in Neural Information Processing Systems.
  21. 21.Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456.
  22. 22.Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al. 2018. Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation. arXiv preprint arXiv:1806.10293.
  23. 23.Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine. 2019. Stabilizing off-policy q-learning via bootstrapping error reduction. In Advances in Neural Information Processing Systems.
  24. 24.Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems.
  25. 25.Esther Levin, Roberto Pieraccini, and Wieland Eckert. 2000. A stochastic model of human-machine interaction for learning dialog strategies. IEEE Transactions on speech and audio processing, 8(1):11–23.
  26. 26.Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643.
  27. 27.Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017. Deal or no deal? end-to-end learning of negotiation dialogues. In Conference on Empirical Methods in Natural Language Processing, pages 2443–2453.
  28. 28.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and William B Dolan. 2016a. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119.
  29. 29.Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016b. Deep reinforcement learning for dialogue generation. In Conference on Empirical Methods in Natural Language Processing.
  30. 30.Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016. Continuous control with deep reinforcement learning. International Conference on Learning Representations.
  31. 31.Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018. Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems. In Proceedings of NAACL-HLT, pages 2060–2069.
  32. 32.Hongyuan Mei, Mohit Bansal, and Matthew R Walter. 2017. Coherent dialogue with attention-based language models. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, pages 3252–3258. AAAI Press.
  33. 33.Tim Paek. 2006. Reinforcement learning for spoken dialogue systems: Comparing strengths and weaknesses for practical deployment. In Proc. Dialog-on-Dialog Workshop, Interspeech. Citeseer.
  34. 34.Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayan-deh, Lars Liden, and Jianfeng Gao. 2020. Soloist: Few-shot task-oriented dialog with a single pre-trained auto-regressive model. arXiv preprint arXiv:2005.05298.
  35. 35.Roberto Pieraccini, David Suendermann, Krishna Dayanidhi, and Jackson Liscombe. 2009. Are we there yet? research in commercial spoken dialog systems. In International Conference on Text, Speech and Dialogue, pages 3–13. Springer.
  36. 36.Olivier Pietquin, Matthieu Geist, Senthilkumar Chandramohan, and Hervé Frezza-Buet. 2011. Sample-efficient batch reinforcement learning for dialogue management optimization. ACM Transactions on Speech and Language Processing (TSLP), 7(3):1–21.
  37. 37.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018. Language models are unsupervised multitask learners. OpenAI Blog.
  38. 38.Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Shen, and Rosalind Picard. 2020. Hierarchical reinforcement learning for open-domain dialog. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8741–8748.
  39. 39.Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young. 2007. Agenda-based user simulation for bootstrapping a pomdp dialogue system. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Companion Volume, Short Papers, pages 149–152.
  40. 40.Iulian Serban, Tim Klinger, Gerald Tesauro, Kartik Talamadupula, Bowen Zhou, Yoshua Bengio, and Aaron Courville. 2017. Multiresolution recurrent neural networks: An application to dialogue response generation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 31.
  41. 41.Satinder Singh, Michael Kearns, Diane Litman, and Marilyn Walker. 1999. Reinforcement learning for spoken dialogue systems. Advances in neural information processing systems, 12:956–962.
  42. 42.Ronnie W Smith and D Richard Hipp. 1994. Spoken natural language dialog systems: A practical approach. Oxford University Press on Demand.
  43. 43.Pei-Hao Su, Paweł Budzianowski, Stefan Ultes, Milica Gasic, and Steve Young. 2017. Sample-efficient actor-critic reinforcement learning with supervised data for dialogue management. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 147–157.
  44. 44.Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016. Continuously learning neural dialogue management. arXiv preprint arXiv:1606.02689.
  45. 45.Pei-Hao Su, David Vandyke, Milica Gašic, Dongho ´ Kim, Nikola Mrkšic, Tsung-Hsien Wen, and Steve ´ Young. 2015. Learning from real users: Rating dialogue success with neural networks for reinforcement learning in spoken dialogue systems. In Sixteenth Annual Conference of the International Speech Communication Association.
  46. 46.Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, and Yunjie Gu. 2020. Modelling hierarchical structure between dialogue policy and natural language generator with option framework for task-oriented dialogue system. arXiv preprint arXiv:2006.06814.
  47. 47.Jason D Williams and Steve Young. 2007. Partially observable markov decision processes for spoken dialog systems. Computer Speech & Language, 21(2):393–422.
  48. 48.Qingyang Wu, Yichi Zhang, Yu Li, and Zhou Yu. 2021. Alternating recurrent dialog model with large-scale pre-trained language models. pages 1292–1301.
  49. 49.Yifan Wu, George Tucker, and Ofir Nachum. 2019. Behavior regularized offline reinforcement learning. arXiv preprint arXiv:1911.11361.
  50. 50.Denis Yarats and Mike Lewis. 2018. Hierarchical text generation and planning for strategic dialogue. In International Conference on Machine Learning, pages 5591–5599.
  51. 51.Steve Young, Milica Gašic, Blaise Thomson, and Jason D Williams. 2013. Pomdp-based statistical spoken dialog systems: A review. Proceedings of the IEEE, 101(5):1160–1179.
  52. 52.Zhou Yu, Ziyu Xu, Alan W Black, and Alexander Rudnicky. 2016. Strategy and policy learning for non-task-oriented conversational systems. In Proceedings of the 17th annual meeting of the special interest group on discourse and dialogue, pages 404–412.
  53. 53.Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020. Task-oriented dialog systems that consider multiple appropriate responses under the same context. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9604–9611.
  54. 54.Tiancheng Zhao, Kaige Xie, and Maxine Eskenazi. 2019. Rethinking action spaces for reinforcement learning in end-to-end dialog agents with latent variable models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1, pages 1208–1218.

Citation

MLA
Verma, S., et al. “CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 4471–91, https://doi.org/10.18653/v1/2022.naacl-main.332.
APA
Verma, S., Fu, J., Yang, S., & Levine, S. (2022). CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4471–4491. https://doi.org/10.18653/v1/2022.naacl-main.332
Chicago
Verma, S., J. Fu, S. Yang, and S. Levine. 2022. “CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4471–91. https://doi.org/10.18653/v1/2022.naacl-main.332.
Harvard
Verma, S. et al. (2022) “CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 4471–4491. Available at: https://doi.org/10.18653/v1/2022.naacl-main.332.
Vancouver
1. Verma S, Fu J, Yang S, Levine S (2022) CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 4471–4491

BibTeX

@inproceedings{verma-etal-2022-chai,
    title = "{CHAI}: A {CH}atbot {AI} for Task-Oriented Dialogue with Offline Reinforcement Learning",
    author = "Verma, Siddharth  and
      Fu, Justin  and
      Yang, Sherry  and
      Levine, Sergey",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.332/",
    doi = "10.18653/v1/2022.naacl-main.332",
    pages = "4471--4491"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/