How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation

Chia-Wei LiuRyan LoweIulian SerbanMichael NoseworthyLaurent CharlinJoelle Pineau

article2016EMNLP1,410 citations

Demonstrates that widely used machine translation metrics correlate poorly or not at all with human judgments when evaluating dialogue response generation, exposing critical flaws in standard conversational benchmarks.

Listen

Organizations are increasingly deploying conversational artificial intelligence and chatbots using data-driven models that learn without explicit task completion signals. Evaluating the response quality of these systems automatically and accurately is critical to speed up development and avoid prohibitively expensive manual review. In response to this need, practitioners have widely adopted standard automated metrics from related language domains, comparing system outputs directly against single human reference responses.

The article systematically evaluates whether standard automated metrics accurately reflect human judgments of response quality in open-ended conversation systems. Specifically, it assesses both word-matching metrics borrowed from translation and summarization as well as semantic similarity metrics derived from word embeddings.

To test metric validity, the authors evaluated multiple conversation models, spanning retrieval-based systems and generative neural architectures. These systems were tested across two contrasting datasets: a technical support domain and an informal social media dialogue setting. Outputs from these models, along with human-authored references, were scored by automated metrics and independently rated for quality by human annotators in a structured study.

The analysis reveals that current automated metrics correlate either very weakly or not at all with human judgments. On the non-technical social media corpus, word-overlap and embedding-based metrics showed only minor positive correlations with human scores. On the technical support corpus, none of the automated metrics exhibited any statistically significant correlation with human ratings. Higher-order word matching metrics suffered heavily from sparsity, frequently producing scores near zero because valid conversational replies often use entirely different words than a single reference target. Furthermore, while automated metrics could reliably differentiate between baseline algorithms and advanced architectures on paper, this separation did not reflect actual alignment with human assessment.

These findings indicate that relying on current automated metrics introduces severe risk into development pipelines. Teams that use these scores to benchmark or select dialogue models may deploy low-quality systems while discarding genuinely effective solutions. Dialogue inherently allows a wide variety of appropriate responses for any given context; therefore, metrics that depend solely on exact word matches or basic semantic averages against a single reference answer fail to capture conversational validity.

Decision-makers and engineering teams must immediately move away from using standard word-overlap metrics like BLEU or basic embedding averages as primary benchmarks for conversational systems. Until reliable automated methods emerge, human evaluation remains essential for validating unsupervised dialogue quality. Future research should prioritize developing context-aware evaluation models and advanced sentence-level representations rather than relying on superficial word similarity.

Confidence in these findings is high for open-domain dialogue, supported by strong agreement levels among human raters across diverse models. However, the study evaluated scenarios with only a single reference response per context. Constrained dialogue systems operating under strictly defined task scripts or systems evaluated with multiple reference answers might exhibit slightly different metric behavior, though caution remains strongly advised across all unconstrained settings.

  • Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). It addresses the specific weaknesses of word-overlap metrics demonstrated in the source by developing BLEURT, a robust learned evaluation metric pre-trained to align closely with human quality judgments.
  • Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). It advances dialogue and NLG evaluation beyond surface-level metrics by using large language models with chain-of-thought prompting to achieve significantly higher correlation with human assessments.
  • Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). It extends conversational evaluation into the modern LLM era by establishing benchmarks like MT-Bench and Chatbot Arena to evaluate open-ended multi-turn dialogue using LLM-as-a-judge approaches.
  • Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey systematically explores modern model-based evaluation frameworks that overcome the failure modes of unsupervised overlap metrics highlighted in the source.
  • Paper: COMET: A Neural Framework for MT Evaluation, Ricardo Rei et al. (2020). It proposes a neural pretrained evaluation framework that moves past exact word matching to provide human-correlated, semantic-level evaluation for text generation.
  • Paper: DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation, Yizhe Zhang et al. (2019). It evaluates large-scale pretrained dialogue generation systems using both automatic metrics and human evaluations, directly operationalizing the source's recommendations for rigorous dialogue assessment.
  • Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). It implements a scalable crowdsourced human-preference platform to address the long-standing challenge of evaluating open-domain conversational models reliably without ground-truth targets.
Cover for How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation

Abstract

We investigate evaluation metrics for dialogue response generation systems where supervised labels, such as task completion, are not available. Recent works in response generation have adopted metrics from machine translation to compare a model's generated response to a single target response. We show that these metrics correlate very weakly with human judgements in the non-technical Twitter domain, and not at all in the technical Ubuntu domain. We provide quantitative and qualitative results highlighting specific weaknesses in existing metrics, and provide recommendations for future development of better automatic evaluation metrics for dialogue systems.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Evaluation Metrics
  • 3.1 Word Overlap-based Metrics
  • 3.2 Embedding-based Metrics
  • 4 Dialogue Response Generation Models
  • 4.1 Retrieval Models
  • 4.2 Generative Models
  • 4.3 Conclusions from an Incomplete Analysis
  • 5 Human Correlation Analysis
  • 6 Discussion
  • References

Knowls

  1. Knowl 1 — Correlation Between Automatic Evaluation Metrics and Human Judgements in Dialogue Systems

    empirical result

    Automatic evaluation metrics adapted from machine translation and text summarization (including BLEU-1 to BLEU-4, METEOR, and ROUGE-L) as well as word embedding similarity metrics (Embedding Average, Vector Extrema, and Greedy Matching) exhibit weak to no correlation with human ratings of response quality in unsupervised dialogue systems.

    In an empirical study comparing metric ratings to human adequacy ratings (on a 1 to 5 scale):

    • In the open-domain, casual chit-chat Twitter dataset, metrics show only weak positive correlation with human judgements (maximum Pearson r=0.3874r = 0.3874 and Spearman ρ=0.3576\rho = 0.3576 for smoothed BLEU-2; embedding metrics achieve Spearman correlations between 0.210.21 and 0.230.23).
    • In the technical, task-oriented Ubuntu Dialogue Corpus, no metric significantly correlates with human judgements (all Pearson and Spearman correlation coefficients are near zero or negative, ranging from ρ=−0.1387\rho = -0.1387 to ρ=0.1218\rho = 0.1218, none statistically significant at p<0.05p < 0.05).

    In contrast, the correlation between two randomly split halves of human judges is high (Spearman ρ=0.9476\rho = 0.9476 on Twitter and ρ=0.9550\rho = 0.9550 on Ubuntu, p<0.01p < 0.01).

  2. Knowl 2 — Correlation Analysis Table of Automatic Metrics vs. Human Judgements

    data/table

    The table below reports the Pearson (rr) and Spearman (ρ\rho) correlation coefficients, along with their associated pp-values, between automatic evaluation metrics and human adequacy scores for proposed dialogue responses on the Twitter Corpus and the Ubuntu Dialogue Corpus.

    Twitter Ubuntu
    Metric Spearman p-value Pearson p-value Spearman p-value Pearson p-value
    Greedy Matching 0.2119 0.034 0.1994 0.047 0.05276 0.60 0.02049 0.84
    Embedding Average 0.2259 0.024 0.1971 0.049 -0.1387 0.17 -0.1631 0.10
    Vector Extrema 0.2103 0.036 0.1842 0.067 0.09243 0.36 -0.002903 0.98
    METEOR 0.1887 0.06 0.1927 0.055 0.06314 0.53 0.1419 0.16
    BLEU-1 0.1665 0.098 0.1288 0.20 -0.02552 0.80 0.01929 0.85
    BLEU-2 0.3576 < 0.01 0.3874 < 0.01 0.03819 0.71 0.0586 0.56
    BLEU-3 0.3423 < 0.01 0.1443 0.15 0.0878 0.38 0.1116 0.27
    BLEU-4 0.3417 < 0.01 0.1392 0.17 0.1218 0.23 0.1132 0.26
    ROUGE 0.1235 0.22 0.09714 0.34 0.05405 0.5933 0.06401 0.53
    Human 0.9476 < 0.01 1.0 0.0 0.9550 < 0.01 1.0 0.0

    The human baseline correlation is calculated by randomly dividing human evaluators into two independent groups. While human-to-human agreement is very strong (ρ>0.94\rho > 0.94 across both datasets), automatic metrics show weak correlation on Twitter and fail to achieve statistically significant correlation with human judgement on Ubuntu.

  3. Knowl 3 — Embedding-Based Evaluation Metrics for Dialogue Responses

    model/method

    Embedding-based metrics assess semantic similarity between a ground truth response rr and a proposed model response r^\hat{r} using distributed word representations (Word2Vec vectors ew∈Rde_w \in \mathbb{R}^d for each token ww, trained on an independent corpus). Three standard formulations are used:

    1. Greedy Matching: For every token w∈rw \in r, find the token w^∈r^\hat{w} \in \hat{r} with the maximum cosine similarity: G(r,r^)=∑w∈rmax⁡w^∈r^cos⁡(ew,ew^)∣r∣G(r, \hat{r}) = \frac{\sum_{w \in r} \max_{\hat{w} \in \hat{r}} \cos(e_w, e_{\hat{w}})}{|r|} Because this measure is asymmetric, the final Greedy Matching score GM\mathrm{GM} is the symmetric average: GM(r,r^)=G(r,r^)+G(r^,r)2\mathrm{GM}(r, \hat{r}) = \frac{G(r, \hat{r}) + G(\hat{r}, r)}{2}

    2. Embedding Average: The sentence-level embedding eˉr\bar{e}_r is the normalized mean of constituent word vectors: eˉr=∑w∈rew∥∑w′∈rew′∥\bar{e}_r = \frac{\sum_{w \in r} e_w}{\left\|\sum_{w' \in r} e_{w'}\right\|} The Embedding Average metric score EA\mathrm{EA} is the cosine similarity between the sentence embeddings: EA=cos⁡(eˉr,eˉr^)\mathrm{EA} = \cos(\bar{e}_r, \bar{e}_{\hat{r}})

    3. Vector Extrema: For each dimension dd in the embedding space, the sentence representation erde_{rd} takes the coordinate value of maximum absolute magnitude across all word vectors in the sentence: erd={max⁡w∈rewdif max⁡w∈rewd>∣min⁡w′∈rew′d∣min⁡w∈rewdotherwisee_{rd} = \begin{cases} \max_{w \in r} e_{wd} & \text{if } \max_{w \in r} e_{wd} > \left|\min_{w' \in r} e_{w'd}\right| \\ \min_{w \in r} e_{wd} & \text{otherwise} \end{cases} The Vector Extrema score is defined as the cosine similarity cos⁡(er,er^)\cos(e_r, e_{\hat{r}}). This heuristic prioritizes informative words over frequent common words.

  4. Knowl 4 — Word Overlap-Based Evaluation Metrics for Dialogue Response Generation

    model/method

    Word overlap metrics compute surface similarity between a candidate dialogue response r^\hat{r} and a single ground truth response rr:

    • BLEU-N: Computes modified nn-gram precision Pn(r,r^)P_n(r, \hat{r}): Pn(r,r^)=∑kmin⁡(h(k,r),h(k,r^))∑kh(k,r^)P_n(r, \hat{r}) = \frac{\sum_k \min(h(k, r), h(k, \hat{r}))}{\sum_k h(k, \hat{r})} where h(k,r)h(k, r) is the count of nn-gram kk in response rr. The sentence-level BLEU-N score combines precisions up to length NN (with uniform weights βn=1/N\beta_n = 1/N) and a brevity penalty b(r,r^)b(r, \hat{r}): BLEU-N=b(r,r^)exp⁡(∑n=1Nβnlog⁡Pn(r,r^))\text{BLEU-N} = b(r, \hat{r}) \exp\left(\sum_{n=1}^N \beta_n \log P_n(r, \hat{r})\right) Sentence-level smoothing is applied to prevent zero scores when higher-order nn-gram overlaps are absent.

    • METEOR: Constructs an explicit word alignment between r^\hat{r} and rr using exact token matching, WordNet synonyms, stemmed tokens, and paraphrases, computing the harmonic mean of alignment precision and recall with a penalty for fragmentation.

    • ROUGE-L: Computes the F-measure based on the Longest Common Subsequence (LCS) between candidate r^\hat{r} and reference rr, capturing non-consecutive in-order word co-occurrences.

  5. Knowl 5 — Model Discrimination by Metrics Does Not Imply Correlation with Human Judgement

    empirical result

    Automatic evaluation metrics can clearly differentiate between baseline models and state-of-the-art models with high statistical significance, yet simultaneously fail to correlate with human judgements of response quality.

    When evaluating response retrieval models (Response-TFIDF, Context-TFIDF, Dual Encoder) and generative models (LSTM language model, Hierarchical Recurrent Encoder-Decoder / HRED) on vector-based metrics (Embedding Averaging, Greedy Matching, Vector Extrema):

    • On the Ubuntu Dialogue Corpus, the Dual Encoder (DE) scores 0.650±0.0030.650 \pm 0.003 on Embedding Averaging compared to 0.536±0.0030.536 \pm 0.003 for R-TFIDF, and HRED scores 0.580±0.0030.580 \pm 0.003 compared to 0.130±0.0030.130 \pm 0.003 for the baseline LSTM.
    • On the Twitter Corpus, DE achieves 0.597±0.0020.597 \pm 0.002 versus 0.483±0.0020.483 \pm 0.002 for R-TFIDF, and HRED achieves 0.599±0.0020.599 \pm 0.002 versus 0.593±0.0020.593 \pm 0.002 for LSTM.

    Although these metrics cleanly separate baseline models from neural models, human correlation tests show that these same embedding metrics exhibit no statistically significant correlation with human ratings on Ubuntu and only weak correlation on Twitter. Thus, a metric's capacity to separate sophisticated models from baselines is insufficient evidence of its validity for dialogue evaluation.

  6. Knowl 6 — Experimental Setup for Human Evaluation of Dialogue Response Quality

    experimental setup

    Human evaluation of dialogue response quality was conducted using the following protocol:

    • Participants: 25 volunteer judges rated responses on an adequacy scale from 1 (not appropriate or sensible given context) to 5 (very reasonable).
    • Inter-Rater Filtering: Evaluators were screened using Cohen's kappa (κ\kappa) agreement. 23 evaluators achieved κ>0.2\kappa > 0.2 (median κ≈0.55\kappa \approx 0.55, indicating moderate to strong agreement) and were retained, while 2 participants with κ<0.2\kappa < 0.2 were excluded.
    • Corpora and Items: Evaluators assessed 100 questions per dataset across two corpora:
      1. The Ubuntu Dialogue Corpus: A domain-specific technical assistance corpus with multi-turn conversations.
      2. The Twitter Corpus: An open-domain social media conversation dataset.
    • Response Diversity: Each test context (20 contexts per dataset) was paired with 5 candidate responses representing different quality levels: a randomly sampled response from elsewhere in the test set, responses selected/generated by TF-IDF, Dual Encoder (DE), Hierarchical Recurrent Encoder-Decoder (HRED), and a ground truth human response.
  7. Knowl 7 — Sensitivity of Evaluation Metrics to Response Length Discrepancy

    empirical result

    Automatic word-overlap metrics are sensitive to discrepancies in sentence length between the ground truth response and the generated response, whereas human quality ratings and embedding average metrics are robust to length differences.

    In an analysis on the Twitter dataset dividing candidate-reference pairs by absolute word length difference Δw=∣∣r∣−∣r^∣∣\Delta w = ||r| - |\hat{r}||:

    Metric Mean Score (Δw≤6\Delta w \le 6, n=47n=47) Mean Score (Δw≥6\Delta w \ge 6, n=53n=53) p-value
    BLEU-1 0.1724 0.1009 < 0.01
    BLEU-2 0.0744 0.04176 < 0.01
    Embedding Average 0.6587 0.6246 0.25
    METEOR 0.2386 0.2073 < 0.01
    Human Score 2.66 2.57 0.73

    BLEU-1, BLEU-2, and METEOR assign significantly lower scores when the length discrepancy Δw≥6\Delta w \ge 6 (p<0.01p < 0.01), whereas the human adequacy ratings show no statistically significant shift (2.662.66 vs. 2.572.57, p=0.73p = 0.73), and Embedding Average shows no significant change (0.65870.6587 vs. 0.62460.6246, p=0.25p = 0.25).

  8. Knowl 8 — Effect of Stopwords and Punctuation Removal on BLEU Correlation

    empirical result

    Stripping stopwords and punctuation from dialogue responses weakens the correlation between BLEU scores and human adequacy ratings. On the Twitter dataset:

    • For BLEU-1, removing stopwords and punctuation yields a Spearman correlation of ρ=0.1580\rho = 0.1580 (p=0.12p = 0.12) and Pearson correlation of r=0.2074r = 0.2074 (p=0.038p = 0.038), compared to ρ=0.1665\rho = 0.1665 (p=0.098p = 0.098) and r=0.1288r = 0.1288 (p=0.20p = 0.20) on unfiltered text.
    • For BLEU-2, removing stopwords and punctuation drops the Spearman correlation from ρ=0.3576\rho = 0.3576 (p<0.01p < 0.01) to ρ=0.2030\rho = 0.2030 (p=0.043p = 0.043), and the Pearson correlation drops from r=0.3874r = 0.3874 (p<0.01p < 0.01) to r=0.1300r = 0.1300 (p=0.20p = 0.20).

    This reduction indicates that a substantial portion of BLEU's weak correlation in dialogue settings is driven by matching function words and punctuation marks rather than response semantics or conversational adequacy.

  9. Knowl 9 — Sparsity and Behavior of BLEU-3 and BLEU-4 in Dialogue Evaluation

    empirical result

    In single-reference dialogue response evaluation, higher-order n-gram metrics BLEU-3 and BLEU-4 suffer from extreme sparsity. For BLEU-4, only 4 out of 100 test response pairs obtain a raw score greater than 10−910^{-9}.

    Because higher-order n-gram matches rarely occur between a valid candidate response and a single ground truth reference, smoothed BLEU-3 and BLEU-4 scores function essentially as scaled, noisy versions of BLEU-2, relying almost entirely on the smoothing weight assigned to unigram and bigram overlaps. Consequently, BLEU-2 is preferable to BLEU-3 or BLEU-4 when evaluating dialogue responses with single references, though none correlate strongly with human judgements.

  10. Knowl 10 — Failure Modes of Overlap and Embedding Metrics in Dialogue Response Evaluation

    limitation

    Automatic word-overlap and embedding-based metrics fail to reflect human judgements of dialogue response quality due to two primary failure modes:

    1. Intrinsic Response Diversity (False Negatives): In conversational dialogue, many semantically distinct and valid responses exist for any given context. When a proposed response is reasonable and appropriate (e.g., replying with an expression of gratitude using different words from the reference), but shares no lexical overlap with the single ground truth response, word-overlap metrics (such as BLEU and METEOR) assign near-zero scores and embedding metrics assign low similarity scores because they cannot distinguish salient keywords from non-salient tokens.
    2. Context-Insensitive Semantic Matching (False Positives): Embedding-based metrics evaluate similarity by aggregating word vectors directly between the ground truth and proposed responses without conditioning on conversational context. When a proposed response contains common tokens (e.g., pronouns like 'i') and semantically proximate words (e.g., 'happy' and 'welcome') that match the reference response, embedding metrics assign high similarity scores even when the proposed response is completely inappropriate or contradicts the conversational context.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.R. Artstein, S. Gandhe, J. Gerten, A. Leuski, and D. Traum. 2009. Semi-formal evaluation of conversational characters. In Languages: From Formal to Natural, pages 22–35. Springer.
  2. 2.S. Banerjee and A. Lavie. 2005. METEOR: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization.
  3. 3.O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, et al. 2014. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 12–58. Association for Computational Linguistics Baltimore, MD, USA.
  4. 4.A. Cahill. 2009. Correlating human and automatic evaluation of a german surface realiser. In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers, pages 97–100. Association for Computational Linguistics.
  5. 5.C. Callison-Burch, M. Osborne, and P. Koehn. 2006. Re-evaluation the role of bleu in machine translation research. In EACL, volume 6, pages 249–256.
  6. 6.C. Callison-Burch, P. Koehn, C. Monz, K. Peterson, M. Przybocki, and O. F. Zaidan. 2010. Findings of the 2010 joint workshop on statistical machine translation and metrics for machine translation. In Proceedings of the Joint Fifth Workshop on Statistical Machine Translation and MetricsMATR, pages 17–53. Association for Computational Linguistics.
  7. 7.C. Callison-Burch, P. Koehn, C. Monz, and O. F. Zaidan. 2011. Findings of the 2011 workshop on statistical machine translation. In Proceedings of the Sixth Workshop on Statistical Machine Translation, pages 22–64. Association for Computational Linguistics.
  8. 8.B. Chen and C. Cherry. 2014. A systematic comparison of smoothing techniques for sentence-level bleu. ACL 2014, page 362.
  9. 9.J. Cohen. 1968. Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological bulletin, 70(4):213.
  10. 10.D. Espinosa, R. Rajkumar, M. White, and S. Berleant. 2010. Further meta-evaluation of broad-coverage surface realization. In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 564–574. Association for Computational Linguistics.
  11. 11.P. W. Foltz, W. Kintsch, and T. K. Landauer. 1998. The measurement of textual coherence with latent semantic analysis. Discourse processes, 25(2-3):285–307.
  12. 12.G. Forgues, J. Pineau, J.-M. Larcheveque, and R. Tremblay. 2014. Bootstrapping dialog systems with word embeddings.
  13. 13.M. Galley, C. Brockett, A. Sordoni, Y. Ji, M. Auli, C. Quirk, M. l, J. Gao, and B. Dolan. 2015a. deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets. In Proceedings of the Annual Meeting of the Association for Computational Linguistics and the International Joint Conference on Natural Language Processing (Short Papers).
  14. 14.M. Galley, C. Brockett, A. Sordoni, Y. Ji, M. Auli, C. Quirk, M. Mitchell, J. Gao, and B. Dolan. 2015b. deltableu: A discriminative metric for generation tasks with intrinsically diverse targets. arXiv preprint arXiv:1506.06863.
  15. 15.Y. Graham, N. Mathur, and T. Baldwin. 2015. Accurate evaluation of segment-level machine translation metrics. In Proc. of NAACL-HLT, pages 1183–1191. Citeseer.
  16. 16.A. Graves. 2013. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850.
  17. 17.S. Hochreiter and J. Schmidhuber. 1997. Long short-term memory. Neural Computation, 9(8):1735–1780.
  18. 18.E. Hovy. 1999. Toward finely differentiated evaluation metrics for machine translation. In Proceedings of the Eagles Workshop on Standards and Evaluation.
  19. 19.K. Jokinen and M. McTear. 2009. Spoken Dialogue Systems. Morgan Claypool.
  20. 20.C. Kamm. 1995. User interfaces for voice applications. Proceedings of the National Academy of Sciences, 92(22):10031–10037.
  21. 21.R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler. 2015. Skip-thought vectors. In Advances in Neural Information Processing Systems, pages 3276–3284.
  22. 22.T. K. Landauer and S. T. Dumais. 1997. A solution to plato’s problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge. Psychological review, 104(2):211.
  23. 23.N. Lasguido, S. Sakti, G. Neubig, T. Tomoki, and S. Nakamura. 2014. Utilizing human-to-human conversation examples for a multi domain chat-oriented dialog system. IEICE TRANSACTIONS on Information and Systems, 97(6):1497–1505.
  24. 24.J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan. 2015. A diversity-promoting objective function for neural conversation models. arXiv preprint arXiv:1510.03055.
  25. 25.J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan. 2016. A persona-based neural conversation model. arXiv preprint arXiv:1603.06155.
  26. 26.C.-Y. Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out: Proceedings of the ACL-04 workshop, volume 8.
  27. 27.R. Lowe, N. Pow, I. V. Serban, and J. Pineau. 2015. The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems. In SIGDIAL.
  28. 28.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119.
  29. 29.J. Mitchell and M. Lapata. 2008. Vector-based models of semantic composition. In ACL, pages 236–244.
  30. 30.S. Möller, R. Englert, K. Engelbrecht, V. Hafner, A. Jameson, A. Oulasvirta, A. Raake, and N. Reithinger. 2006. MeMo: towards automatic usability evaluation of spoken dialogue services by user error simulations. In INTERSPEECH.
  31. 31.K. Papineni, S. Roukos, T. Ward, and W. Zhu. 2002a. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on Association for Computational Linguistics (ACL).
  32. 32.K. Papineni, S. Roukos, T. Ward, J. Henderson, and F. Reeder. 2002b. Corpus-based comprehensive and diagnostic MT evaluation: Initial Arabic, Chinese, French, and Spanish results. In Proceedings of the second international conference on Human Language Technology Research, pages 132–137.
  33. 33.E. Reiter and A. Belz. 2009. An investigation into the validity of some metrics for automatically evaluating natural language generation systems. Computational Linguistics, 35(4):529–558.
  34. 34.A. Ritter, C. Cherry, and B. Dolan. 2010. Unsupervised modeling of twitter conversations. In North American Chapter of the Association for Computational Linguistics (NAACL).
  35. 35.A. Ritter, C. Cherry, and W. B. Dolan. 2011. Data-driven response generation in social media. In Proceedings of the conference on empirical methods in natural language processing, pages 583–593. Association for Computational Linguistics.
  36. 36.V. Rus and M. Lintean. 2012. A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics. In Proceedings of the Seventh Workshop on Building Educational Applications Using NLP, pages 157–162, Stroudsburg, PA, USA. Association for Computational Linguistics.
  37. 37.J. Schatzmann, K. Georgila, and S. Young. 2005. Quantitative evaluation of user simulation techniques for spoken dialogue systems. In 6th Special Interest Group on Discourse and Dialogue (SIGDIAL).
  38. 38.I. V. Serban, A. Sordoni, Y. Bengio, A. Courville, and J. Pineau. 2015. Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Networks. In AAAI Conference on Artificial Intelligence.
  39. 39.I. V. Serban, A. Sordoni, R. Lowe, L. Charlin, J. Pineau, A. Courville, and Y. Bengio. 2016. A hierarchical latent variable encoder-decoder model for generating dialogues. arXiv preprint arXiv:1605.06069.
  40. 40.A. Sordoni, M. Galley, M. Auli, C. Brockett, Y. Ji, M. Mitchell, J. Nie, J. Gao, and B. Dolan. 2015. A neural network approach to context-sensitive generation of conversational responses. In Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT 2015).
  41. 41.A. Stent, M. Marge, and M. Singhai. 2005. Evaluating evaluation methods for generation in the presence of variation. In International Conference on Intelligent Text Processing and Computational Linguistics, pages 341–351. Springer.
  42. 42.O. Vinyals and Q. Le. 2015. A neural conversational model. arXiv preprint arXiv:1506.05869.
  43. 43.M. Walker, D. Litman, C. Kamm, and A. Abella. 1997. Paradise: A framework for evaluating spoken dialogue agents. In Proceedings of the eighth conference on European chapter of the Association for Computational Linguistics, pages 271–280. ACL.
  44. 44.T.-H. Wen, M. Gasic, N. Mrksic, P.-H. Su, D. Vandyke, and S. Young. 2015. Semantically conditioned lstm-based natural language generation for spoken dialogue systems. arXiv preprint arXiv:1508.01745.
  45. 45.J. Wieting, M. Bansal, K. Gimpel, and K. Livescu. 2015. Towards universal paraphrastic sentence embeddings. CoRR, abs/1511.08198.

Citation

MLA
Liu, C.-W., et al. “How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation”. arXiv, 2016, http://arxiv.org/abs/1603.08023v2.
APA
Liu, C.-W., Lowe, R., Serban, I. V., Noseworthy, M., Charlin, L., & Pineau, J. (2016). How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. arXiv. http://arxiv.org/abs/1603.08023v2
Chicago
Liu, C.-W., R. Lowe, I. V. Serban, M. Noseworthy, L. Charlin, and J. Pineau. 2016. “How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation”. arXiv. http://arxiv.org/abs/1603.08023v2.
Harvard
Liu, C.-W. et al. (2016) “How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1603.08023v2.
Vancouver
1. Liu C-W, Lowe R, Serban IV, Noseworthy M, Charlin L, Pineau J (2016) How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. arXiv

BibTeX

@article{liu2016how,
  title = {How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation},
  author = {Liu, Chia-Wei and Lowe, Ryan and Serban, Iulian V. and Noseworthy, Michael and Charlin, Laurent and Pineau, Joelle},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1603.08023v2},
  eprint = {1603.08023}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/