Deep Reinforcement Learning for Dialogue Generation

Jiwei LiWill MonroeAlan RitterDan JurafskyMichel GalleyJianfeng Gao

article2016EMNLP1,431 citations

Proposes a deep reinforcement learning framework that trains neural dialogue agents through simulated multi-turn interactions, optimizing for long-term coherence and informativeness to overcome the repetitive, shortsighted traps of standard sequence-to-sequence models.

Listen

Standard neural conversational models rely on maximum likelihood estimation to predict responses one turn at a time. This single-turn focus creates major conversational failures: chatbots frequently default to repetitive, generic replies (such as "I don't know") and get trapped in repetitive loops that stall user engagement. The article demonstrates how integrating deep reinforcement learning into open-domain dialogue generation enables virtual agents to optimize for long-term conversational success rather than immediate, short-sighted word probabilities.

To achieve this, the authors created a reinforcement learning framework where two virtual agents simulate multi-turn conversations with each other. The models were pre-trained on a corpus of approximately 80 million subtitle dialogues before being optimized via policy gradient methods. The training used a composite reward function designed to encourage three key conversational qualities: informativity (penalizing repetitive responses), ease of answering (rewarding prompts that are easy to reply to and avoid dead-ends), and semantic coherence (preserving grammatical and contextual relevance).

The evaluation showed substantial improvements over standard sequence-to-sequence and mutual information baselines. First, the reinforcement learning model significantly extended conversation length, sustaining dialogues for an average of 4.48 turns compared to 2.68 turns for standard models and 3.40 turns for mutual information models before stalling. Second, vocabulary diversity rose markedly, with the unigram diversity score rising to 0.017 (compared to 0.0062 for standard models) and bigram diversity reaching 0.041 (compared to 0.015). Third, human judges strongly favored the reinforcement learning system, preferring its multi-turn conversation quality in 72% of pairwise comparisons (losing only 12%) and finding its individual turns easier to answer 52% of the time (losing 23%). General single-turn quality showed a modest win rate of 40% to 36%, reflecting that the system is optimized for multi-turn flow rather than single-turn word predictions.

These findings indicate that shifting optimization goals from short-term likelihood to long-term reward mechanisms directly resolves chatbot dead-ends and creates more interactive, engaging user experiences. For organizations deploying conversational artificial intelligence, this approach provides a viable pathway to improve user retention, reduce repetitive failure modes, and increase interactive capabilities without relying on brittle, handcrafted task-oriented scripts.

Organizations developing dialogue systems should incorporate reinforcement learning frameworks that reward forward-looking utterances such as follow-up questions. Future work should focus on moving from heuristic reward functions to human-in-the-loop feedback and expanding computational capacity to explore longer conversation horizons. However, leaders should note that the current model uses a limited dialogue history, which can occasionally cause longer-loop repetitions or topic drifting, and that heuristic reward metrics cannot fully capture all nuances of natural human conversation.

arXiv: 1606.01541
Cover for Deep Reinforcement Learning for Dialogue Generation

Abstract

Recent neural models of dialogue generation offer great promise for generating responses for conversational agents, but tend to be shortsighted, predicting utterances one at a time while ignoring their influence on future outcomes. Modeling the future direction of a dialogue is crucial to generating coherent, interesting dialogues, a need which led traditional NLP models of dialogue to draw on reinforcement learning. In this paper, we show how to integrate these goals, applying deep reinforcement learning to model future reward in chatbot dialogue. The model simulates dialogues between two virtual agents, using policy gradient methods to reward sequences that display three useful conversational properties: informativity (non-repetitive turns), coherence, and ease of answering (related to forward-looking function). We evaluate our model on diversity, length as well as with human judges, showing that the proposed algorithm generates more interactive responses and manages to foster a more sustained conversation in dialogue simulation. This work marks a first step towards learning a neural conversational model based on the long-term success of dialogues.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Reinforcement Learning for Open-Domain Dialogue
  • 3.1 Action
  • 3.2 State
  • 3.3 Policy
  • 3.4 Reward
  • 4 Simulation
  • 4.1 Supervised Learning
  • 4.2 Mutual Information
  • 4.3 Dialogue Simulation between Two Agents
  • 4.4 Curriculum Learning
  • 5 Experimental Results
  • 5.1 Dataset
  • 5.2 Automatic Evaluation
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Reinforcement Learning Formulation for Open-Domain Dialogue Generation

    model/method

    The open-domain dialogue generation framework models conversation as a sequential decision-making process between two virtual agents that take turns generating utterances. Let p1,q1,p2,q2,…,pi,qip_1, q_1, p_2, q_2, \dots, p_i, q_i denote the alternating sequence of dialogue turns produced by the first agent (pp) and the second agent (qq).

    At turn i+1i+1, the acting agent generates an action a=pi+1a = p_{i+1} (or qi+1q_{i+1}), where aa is a natural language utterance. The action space is infinite because utterances can be of arbitrary length.

    The state sis_i available to an agent is defined by the preceding two dialogue turns, [pi,qi][p_i, q_i]. The state is mapped into a continuous vector representation by encoding the concatenation of pip_i and qiq_i using a long short-term memory (LSTM) network.

    The policy pRL(a∣pi,qi)p_\text{RL}(a \mid p_i, q_i) is parameterized as an LSTM encoder-decoder recurrent neural network that outputs a probability distribution over target utterances conditioned on the encoded state. A stochastic policy formulation is employed to allow optimization via policy gradient methods.

  2. Knowl 2 — Composite Multi-Objective Reward Function for Dialogue Evaluation

    equation

    To evaluate the quality of a generated response and guide policy gradient optimization, the immediate scalar reward r(a,[pi,qi])r(a, [p_i, q_i]) observed after generating utterance aa in response to dialogue state [pi,qi][p_i, q_i] is defined as a weighted sum of three sub-rewards:

    r(a,[pi,qi])=λ1r1+λ2r2+λ3r3r(a, [p_i, q_i]) = \lambda_1 r_1 + \lambda_2 r_2 + \lambda_3 r_3

    where:

    • r1r_1 is the ease of answering reward, which measures forward-looking capacity by penalizing the likelihood of eliciting dull responses.
    • r2r_2 is the information flow reward, which penalizes semantic similarity between consecutive turns generated by the same speaker.
    • r3r_3 is the semantic coherence reward, which evaluates the mutual information between the action aa and the context history.
    • λ1,λ2,λ3≥0\lambda_1, \lambda_2, \lambda_3 \ge 0 are non-negative weighting hyperparameters satisfying λ1+λ2+λ3=1\lambda_1 + \lambda_2 + \lambda_3 = 1. The system uses λ1=0.25\lambda_1 = 0.25, λ2=0.25\lambda_2 = 0.25, and λ3=0.5\lambda_3 = 0.5.
  3. Knowl 3 — Ease of Answering Dialogue Reward

    equation

    The ease of answering reward r1r_1 measures how easily a conversational partner can respond to an utterance aa by penalizing utterances that lead to generic, conversation-terminating responses (e.g., "I don't know what you are talking about").

    Let S\mathbb{S} denote a predefined set of dull phrases with cardinality NS=∣S∣N_\mathbb{S} = |\mathbb{S}| (NS=8N_\mathbb{S} = 8), and let NsN_s denote the number of word tokens in phrase s∈Ss \in \mathbb{S}. The reward is defined as the negative log-likelihood of generating phrases in S\mathbb{S} given aa under a sequence-to-sequence model:

    r1=−1NS∑s∈S1Nslog⁡pseq2seq(s∣a)r_1 = - \frac{1}{N_\mathbb{S}} \sum_{s \in \mathbb{S}} \frac{1}{N_s} \log p_\text{seq2seq}(s \mid a)

    where pseq2seq(s∣a)p_\text{seq2seq}(s \mid a) is the conditional probability of ss given aa computed by an MLE-trained sequence-to-sequence model. Normalizing by NsN_s prevents bias from target phrase length. Minimizing the probability of dull responses yields higher rewards.

  4. Knowl 4 — Information Flow Dialogue Reward

    equation

    The information flow reward r2r_2 encourages diversity across dialogue turns and penalizes repetitive output from the same conversational agent.

    Let hpih_{p_i} and hpi+1h_{p_{i+1}} denote vector representations obtained from the hidden state of an LSTM encoder for two consecutive turns pip_i and pi+1p_{i+1} produced by the same agent. The reward is defined as the negative log cosine similarity between these representations:

    r2=−log⁡cos⁡(hpi,hpi+1)=−log⁡(hpi⋅hpi+1∥hpi∥ ∥hpi+1∥)r_2 = -\log \cos(h_{p_i}, h_{p_{i+1}}) = -\log \left( \frac{h_{p_i} \cdot h_{p_{i+1}}}{\|h_{p_i}\| \, \|h_{p_{i+1}}\|} \right)

    where ⋅\cdot is the vector dot product and ∥⋅∥\|\cdot\| is the Euclidean norm. If consecutive turns are semantically redundant (high cosine similarity), r2r_2 is penalized; if consecutive turns introduce novel semantic content (lower cosine similarity), r2r_2 increases.

  5. Knowl 5 — Semantic Coherence Dialogue Reward

    equation

    The semantic coherence reward r3r_3 measures the relevance and grammatical appropriateness of an action aa relative to the dialogue context [pi,qi][p_i, q_i] using mutual information:

    r3=1Nalog⁡pseq2seq(a∣qi,pi)+1Nqilog⁡pseq2seqbackward(qi∣a)r_3 = \frac{1}{N_a} \log p_\text{seq2seq}(a \mid q_i, p_i) + \frac{1}{N_{q_i}} \log p_\text{seq2seq}^\text{backward}(q_i \mid a)

    where:

    • pseq2seq(a∣qi,pi)p_\text{seq2seq}(a \mid q_i, p_i) denotes the forward probability of generating candidate response aa given dialogue turns pip_i and qiq_i, computed using a standard sequence-to-sequence model.
    • pseq2seqbackward(qi∣a)p_\text{seq2seq}^\text{backward}(q_i \mid a) denotes the backward probability of reconstructing the prior turn qiq_i given response aa, computed using a sequence-to-sequence model trained on reversed source-target pairs.
    • NaN_a and NqiN_{q_i} denote the token counts of aa and qiq_i, respectively, normalizing the objective against length discrepancies.
  6. Knowl 6 — Policy Gradient Optimization for Multi-Turn Dialogue Simulation

    algorithm

    The policy parameters θ\theta of an encoder-decoder network pRLp_\text{RL} are optimized to maximize the total expected cumulative future reward across multi-turn interactions between two agents:

    JRL(θ)=EpRL(a1:T)[∑i=1TR(ai,[pi,qi])]J_\text{RL}(\theta) = \mathbb{E}_{p_\text{RL}(a_{1:T})} \left[ \sum_{i=1}^T R(a_i, [p_i, q_i]) \right]

    The policy gradient is estimated via the REINFORCE likelihood ratio approach with curriculum learning across simulation turns:

    Input: Dataset of starting messages D, initial policy parameters theta, max turns T_max = 5, candidate beam size K = 5
    Output: Optimized policy parameters theta
    for turn_depth = 2 to T_max do
        for each input message m in D do
            Initialize turn 1 state s_1 = m
            for turn i = 1 to turn_depth do
                Sample K candidate responses from p_RL(. | s_i)
                Select response a_i from candidates
                Compute immediate reward R(a_i, s_i) = lambda_1 * r_1 + lambda_2 * r_2 + lambda_3 * r_3
                Update next state s_{i+1} with [s_i, a_i]
            end for
            Compute total dialogue reward R_total = sum_{i=1}^{turn_depth} R(a_i, s_i)
            Compute gradient g = sum_{i=1}^{turn_depth} grad_theta log p_RL(a_i | s_i) * R_total
            Update parameters theta = theta + alpha * g
        end for
    end for
  7. Knowl 7 — Maximum Mutual Information Policy Pre-Training via Reinforcement Learning

    model/method

    Initializing reinforcement learning directly from standard maximum likelihood estimation (MLE) sequence-to-sequence models results in generic responses during initial simulations. To provide a diverse starting policy, the encoder-decoder model pRLp_\text{RL} is pre-trained to optimize the maximum mutual information (MMI) objective before multi-turn simulation.

    For a state [pi,qi][p_i, q_i] and candidate sample a^∼pRL\hat{a} \sim p_\text{RL}, the mutual information score is defined as:

    m(a^,[pi,qi])=1Na^log⁡pseq2seq(a^∣pi,qi)+1Nqilog⁡pseq2seqbackward(qi∣a^)m(\hat{a}, [p_i, q_i]) = \frac{1}{N_{\hat{a}}} \log p_\text{seq2seq}(\hat{a} \mid p_i, q_i) + \frac{1}{N_{q_i}} \log p_\text{seq2seq}^\text{backward}(q_i \mid \hat{a})

    where pseq2seqp_\text{seq2seq} and pseq2seqbackwardp_\text{seq2seq}^\text{backward} are pre-trained forward and backward sequence-to-sequence models.

    The policy gradient update incorporates a baseline neural estimator b(a^,[pi,qi])b(\hat{a}, [p_i, q_i]) to reduce variance:

    ∇θJ(θ)=∇θlog⁡pRL(a^∣[pi,qi])[m(a^,[pi,qi])−b]\nabla_\theta J(\theta) = \nabla_\theta \log p_\text{RL}(\hat{a} \mid [p_i, q_i]) \left[ m(\hat{a}, [p_i, q_i]) - b \right]

    A token-level curriculum is applied: for target sequences of length TT, MLE loss is computed for the first LL tokens and the reinforcement learning policy gradient for the remaining T−LT - L tokens, with LL annealed to 0 over training.

  8. Knowl 8 — Simulated Dialogue Length and N-Gram Diversity Improvements

    data/table

    The performance of the reinforcement learning (RL) dialogue generation system was evaluated against a standard sequence-to-sequence model and a mutual information (MMI) reranking model on dialogue length and response diversity. Dialogue length is the average number of turns simulated before an agent produces a generic response (matching an 8-phrase list) or generates an utterance with greater than 80% word overlap with the prior turn (capped at 8 turns). Lexical diversity is measured as the type-token ratio of distinct unigrams and distinct bigrams.

    Model # of Simulated Turns Distinct Unigrams Distinct Bigrams
    SEQ2SEQ 2.68 0.0062 0.015
    Mutual Information 3.40 0.0110 0.031
    Reinforcement Learning 4.48 0.0170 0.041

    The RL model sustained conversations for an average of 4.48 turns before loop or dull termination, compared to 2.68 turns for vanilla SEQ2SEQ and 3.40 for MMI. RL also achieved higher lexical diversity (0.017 unigram and 0.041 bigram type-token ratios).

  9. Knowl 9 — Pairwise Human Evaluation of Dialogue Quality and Ease of Answering

    data/table

    Crowdsourced human annotators (3 judges per sample) conducted pairwise comparisons between the reinforcement learning (RL) model and the mutual information (MMI) baseline across three settings:

    1. Single-turn general quality (500 random input messages).
    2. Single-turn ease to answer (500 random input messages).
    3. Multi-turn general quality across 5 simulated dialogue turns (200 simulated conversations).
    Setting RL-win RL-lose Tie
    Single-turn general quality 0.40 0.36 0.24
    Single-turn ease to answer 0.52 0.23 0.25
    Multi-turn general quality 0.72 0.12 0.16

    In single-turn general quality, the RL model won 40% of comparisons and lost 36%. In single-turn ease to answer, RL responses were judged easier to respond to in 52% of comparisons (vs. 23% losses). In multi-turn evaluation, dialogues generated by the RL system outperformed the MMI baseline in 72% of comparisons (vs. 12% losses).

  10. Knowl 10 — Repetition Cycles and Search Space Limitations in Simulated Dialogue RL

    limitation

    Error analysis of the two-agent reinforcement learning dialogue simulation identified key limitations:

    1. Multi-turn cyclical loops: Penalizing semantic similarity between immediately consecutive utterances from the same agent (pip_i and pi+1p_{i+1}) prevents 1-turn immediate repetition, but agents can still enter cyclical loops of period greater than one (e.g., alternating "What's your name? / Daniel / How old are you? / Twelve" repeatedly). This occurs because the policy state is restricted to the previous two turns [pi,qi][p_i, q_i].
    2. Relevance versus diversity trade-off: Heavily penalizing generic responses occasionally leads agents to shift to unrelated topics, highlighting sensitivity to the reward component weights (λ1,λ2,λ3\lambda_1, \lambda_2, \lambda_3).
    3. Incomplete reward heuristics: Handcrafted heuristic rewards do not capture all nuances of human dialogue quality.
    4. Exponential branching: Exploring full multi-turn candidate trajectories is computationally bounded, restricting simulations to a small beam of candidates (5 candidates per turn) and short rollouts (up to 5 turns).

Coverage note — None was omitted; the extracted knowls fully cover the reinforcement learning framework, all three reward definitions and composite objective, policy gradient and curriculum algorithms, pre-training steps, empirical evaluations (length, diversity, human judgments), and structural limitations.

References

  1. 1.V. M. Aleksandrov, V. I. Sysoyev, and V. V. Shemeneva. 1968. Stochastic optimization. Engineering Cybernetics, 5:11–16.
  2. 2.Jens Allwood, Joakim Nivre, and Elisabeth Ahlsén. 1992. On the semantics and pragmatics of linguistic feedback. Journal of Semantics, 9:1–26.
  3. 3.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proc. of ICLR.
  4. 4.Rafael E Banchs and Haizhou Li. 2012. IRIS: a chat-oriented dialogue system based on the vector space model. In Proceedings of the ACL 2012 System Demonstrations, pages 37–42.
  5. 5.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48. ACM.
  6. 6.SRK Branavan, David Silver, and Regina Barzilay. 2011. Learning to win by reading manuals in a monte-carlo framework. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, pages 268–277.
  7. 7.Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015. deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets. In Proc. of ACL-IJCNLP, pages 445–450, Beijing, China, July.
  8. 8.Milica Gašić, Catherine Breslin, Matthew Henderson, Dongho Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young. 2013a. Pomdp-based dialogue manager adaptation to extended domains. In Proceedings of SIGDIAL.
  9. 9.Milica Gašić, Catherine Breslin, Mike Henderson, Dongkyu Kim, Martin Szummer, Blaise Thomson, Pirros Tsiakoulis, and Steve Young. 2013b. On-line policy optimisation of bayesian spoken dialogue systems via human interaction. In Proceedings of ICASSP 2013, pages 8367–8371. IEEE.
  10. 10.Milica Gašić, Dongho Kim, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, and Steve Young. 2014. Incremental on-line adaptation of pomdp-based dialogue managers to extended domains. In Proceedings on InterSpeech.
  11. 11.Peter W Glynn. 1990. Likelihood ratio gradient estimation for stochastic systems. Communications of the ACM, 33(10):75–84.
  12. 12.Ji He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, and Mari Ostendorf. 2016. Deep reinforcement learning with a natural language action space. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1621–1630, Berlin, Germany, August.
  13. 13.Esther Levin, Roberto Pieraccini, and Wieland Eckert. 1997. Learning dialogue strategies within the markov decision process framework. In Automatic Speech Recognition and Understanding, 1997. Proceedings., 1997 IEEE Workshop on, pages 72–79. IEEE.
  14. 14.Esther Levin, Roberto Pieraccini, and Wieland Eckert. 2000. A stochastic model of human-machine interaction for learning dialog strategies. IEEE Transactions on Speech and Audio Processing, 8(1):11–23.
  15. 15.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016a. A diversity-promoting objective function for neural conversation models. In Proc. of NAACL-HLT.
  16. 16.Jiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, and Bill Dolan. 2016b. A persona-based neural conversation model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 994–1003, Berlin, Germany, August.
  17. 17.Chia-Wei Liu, Ryan Lowe, Iulian V Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. arXiv preprint arXiv:1603.08023.
  18. 18.Yi Luan, Yangfeng Ji, and Mari Ostendorf. 2016. LSTM based conversation models. arXiv preprint arXiv:1603.09457.
  19. 19.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing Atari with deep reinforcement learning. NIPS Deep Learning Workshop.
  20. 20.Karthik Narasimhan, Tejas Kulkarni, and Regina Barzilay. 2015. Language understanding for text-based games using deep reinforcement learning. arXiv preprint arXiv:1506.08941.
  21. 21.Lasguido Nio, Sakriani Sakti, Graham Neubig, Tomoki Toda, Mirna Adriani, and Satoshi Nakamura. 2014. Developing non-goal dialog system based on examples of drama television. In Natural Interaction with Robots, Knowbots and Smartphones, pages 355–361. Springer.
  22. 22.Alice H Oh and Alexander I Rudnicky. 2000. Stochastic language generation for spoken dialogue systems. In Proceedings of the 2000 ANLP/NAACL Workshop on Conversational systems-Volume 3, pages 27–32.
  23. 23.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics, pages 311–318.
  24. 24.Roberto Pieraccini, David Suendermann, Krishna Dayanidhi, and Jackson Liscombe. 2009. Are we there yet? Research in commercial spoken dialog systems. In Text, Speech and Dialogue, pages 3–13. Springer.
  25. 25.Marc'Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015. Sequence level training with recurrent neural networks. arXiv preprint arXiv:1511.06732.
  26. 26.Adwait Ratnaparkhi. 2002. Trainable approaches to surface natural language generation and their application to conversational dialog systems. Computer Speech & Language, 16(3):435–455.
  27. 27.Alan Ritter, Colin Cherry, and William B Dolan. 2011. Data-driven response generation in social media. In Proceedings of EMNLP 2011, pages 583–593.
  28. 28.Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young. 2006. A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies. The knowledge engineering review, 21(02):97–126.
  29. 29.Emanuel A. Schegloff and Harvey Sacks. 1973. Opening up closings. Semiotica, 8(4):289–327.
  30. 30.Iulian V Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau. 2016. Building end-to-end dialogue systems using generative hierarchical neural network models. In Proceedings of AAAI, February.
  31. 31.Lifeng Shang, Zhengdong Lu, and Hang Li. 2015. Neural responding machine for short-text conversation. In Proceedings of ACL-IJCNLP, pages 1577–1586.
  32. 32.David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. 2016. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489.
  33. 33.Satinder P Singh, Michael J Kearns, Diane J Litman, and Marilyn A Walker. 1999. Reinforcement learning for spoken dialogue systems. In Nips, pages 956–962.
  34. 34.Satinder Singh, Michael Kearns, Diane J Litman, Marilyn A Walker, et al. 2000. Empirical evaluation of a reinforcement learning spoken dialogue system. In AAAI/IAAI, pages 645–651.
  35. 35.Satinder Singh, Diane Litman, Michael Kearns, and Marilyn Walker. 2002. Optimizing dialogue management with reinforcement learning: Experiments with the njfun system. Journal of Artificial Intelligence Research, pages 105–133.
  36. 36.Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Meg Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015. A neural network approach to context-sensitive generation of conversational responses. In Proceedings of NAACL-HLT.
  37. 37.Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016. Continuously learning neural dialogue management. arxiv.
  38. 38.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Advances in neural information processing systems, pages 3104–3112.
  39. 39.Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al. 1999. Policy gradient methods for reinforcement learning with function approximation. In NIPS, volume 99, pages 1057–1063.
  40. 40.Oriol Vinyals and Quoc Le. 2015. A neural conversational model. In Proceedings of ICML Deep Learning Workshop.
  41. 41.Adam Vogel and Dan Jurafsky. 2010. Learning to follow navigational directions. In Proceedings of ACL 2010, pages 806–814.
  42. 42.Marilyn A Walker, Rashmi Prasad, and Amanda Stent. 2003. A trainable generator for recommendations in multimodal dialog. In Proceeedings of INTERSPEECH 2003.
  43. 43.Marilyn A. Walker. 2000. An application of reinforcement learning to dialogue strategy selection in a spoken dialogue system for email. Journal of Artificial Intelligence Research, pages 387–416.
  44. 44.Tsung-Hsien Wen, Milica Gasic, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015. Semantically conditioned LSTM-based natural language generation for spoken dialogue systems. In Proceedings of EMNLP, pages 1711–1721, Lisbon, Portugal.
  45. 45.Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Lina M Rojas-Barahona, Pei-Hao Su, Stefan Ultes, David Vandyke, and Steve Young. 2016. A network-based end-to-end trainable task-oriented dialogue system. arXiv preprint arXiv:1604.04562.
  46. 46.Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3-4):229–256.
  47. 47.Zhen Xu, Bingquan Liu, Baoxun Wang, Chengjie Sun, and Xiaolong Wang. 2016. Incorporating loose-structured knowledge into LSTM with recall gate for conversation modeling. arXiv preprint arXiv:1605.05110.
  48. 48.Kaisheng Yao, Geoffrey Zweig, and Baolin Peng. 2015. Attention with intention for a neural network conversation model. In NIPS workshop on Machine Learning for Spoken Language Understanding and Interaction.
  49. 49.Steve Young, Milica Gašić, Simon Keizer, François Mairesse, Jost Schatzmann, Blaise Thomson, and Kai Yu. 2010. The hidden information state model: A practical framework for pomdp-based spoken dialogue management. Computer Speech & Language, 24(2):150–174.
  50. 50.Steve Young, Milica Gasic, Blaise Thomson, and Jason D Williams. 2013. Pomdp-based statistical spoken dialog systems: A review. Proceedings of the IEEE, 101(5):1160–1179.
  51. 51.Wojciech Zaremba and Ilya Sutskever. 2015. Reinforcement learning neural Turing machines. arXiv preprint arXiv:1505.00521.

Citation

MLA
Li, J., et al. “Deep Reinforcement Learning for Dialogue Generation”. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016, pp. 1192–202, https://doi.org/10.18653/v1/D16-1127.
APA
Li, J., Monroe, W., Ritter, A., Jurafsky, D., Galley, M., & Gao, J. (2016). Deep Reinforcement Learning for Dialogue Generation. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1192–1202. https://doi.org/10.18653/v1/D16-1127
Chicago
Li, J., W. Monroe, A. Ritter, D. Jurafsky, M. Galley, and J. Gao. 2016. “Deep Reinforcement Learning for Dialogue Generation”. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 1192–1202. https://doi.org/10.18653/v1/D16-1127.
Harvard
Li, J. et al. (2016) “Deep Reinforcement Learning for Dialogue Generation”, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1192–1202. Available at: https://doi.org/10.18653/v1/D16-1127.
Vancouver
1. Li J, Monroe W, Ritter A, Jurafsky D, Galley M, Gao J (2016) Deep Reinforcement Learning for Dialogue Generation. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1192–1202

BibTeX

@inproceedings{li-etal-2016-deep,
    title = "Deep Reinforcement Learning for Dialogue Generation",
    author = "Li, Jiwei  and
      Monroe, Will  and
      Ritter, Alan  and
      Jurafsky, Dan  and
      Galley, Michel  and
      Gao, Jianfeng",
    editor = "Su, Jian  and
      Duh, Kevin  and
      Carreras, Xavier",
    booktitle = "Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2016",
    address = "Austin, Texas",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/D16-1127/",
    doi = "10.18653/v1/D16-1127",
    pages = "1192--1202"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/