DialFact: A Benchmark for Fact-Checking in Dialogue

Prakhar GuptaChien-Sheng WuWenhao LiuCaiming Xiong

article2022ACL87 citations

Introduces a benchmark of over 22,000 annotated conversational claims to advance fact verification in dialogue, revealing why standard non-dialogue models fail and offering a data-efficient training strategy to address conversational challenges like coreference and colloquialisms.

Listen

Online misinformation and inaccurate statements generated by automated conversational systems present serious risks, particularly during public crises. While automated fact-checking has advanced for formal texts like news and encyclopedia articles, verifying factual correctness in everyday conversations remains challenging. Conversational statements are typically informal, sparse in factual details, full of colloquialisms and personal opinions, and heavily reliant on prior conversational context to resolve references.

The main objective of the article is to establish the first dedicated benchmark dataset and evaluation pipeline for conversational fact-checking, and to demonstrate practical modeling techniques that improve fact verification performance in dialogue settings.

To address this challenge, the authors created DIALFACT, a benchmark consisting of 22,245 conversational claims derived from the Wizard of Wikipedia dialogue dataset and paired with Wikipedia evidence. The dataset includes both human-written and machine-generated responses. Claims were generated through rule-based mutations, language model text infilling, and open-domain dialogue generation, followed by rigorous multi-round annotations and quality screening on Amazon Mechanical Turk. The fact-checking pipeline was broken into three sequential tasks: detecting verifiable factual claims versus personal opinions, retrieving relevant evidence from Wikipedia, and classifying whether the evidence supports, refutes, or provides insufficient information to judge the claim.

The investigation revealed four key findings. First, existing fact-checking models trained on traditional, non-conversational data perform poorly on conversational claims. Second, providing dialogue context substantially improves evidence retrieval; incorporating conversation history raised document retrieval recall from 60.8% to 75.0% for web-based search and from 44.7% to 58.8% for dense neural retrieval. Third, the authors' weakly supervised training approach, named Aug-WoW, outperformed all standard baselines across all testing conditions, achieving 69.2% claim verification accuracy with ground-truth evidence and approximately 51.5% with automatically retrieved evidence. Fourth, end-to-end performance drops significantly (by roughly 18 percentage points) when models rely on retrieved evidence rather than perfect reference text, underscoring that retrieval is a critical performance bottleneck.

These findings demonstrate that conversational fact-checking cannot be solved simply by repurposing standard fact-checking tools. Because conversational systems are prone to confusing mere topical keyword overlap with factual consistency, deploying them without dialogue-specific retrieval and verification mechanisms creates substantial risks of undetected errors or false alarms. Automated dialogue agents operating in customer service, public information, or advisory capacities require tailored architectures to remain reliable.

Organizations developing or deploying conversational agents should integrate dialogue context into their knowledge retrieval pipelines rather than evaluating claims in isolation. Practitioners should adopt weakly supervised data augmentation methods—such as entity swapping, negation, and masked generation—to train conversational verification tools cost-effectively. Further development should prioritize improving the precision of initial evidence retrieval and enhancing models' ability to distinguish between subjective opinions and unsupported factual assertions.

The conclusions should be interpreted within certain boundaries. The evaluation relies exclusively on English Wikipedia as the knowledge source and centers around open-domain conversational topics. While human validation confirmed solid data consistency, borderline cases between subjective personal opinions and factual claims remain a source of modest uncertainty. Overall, confidence is high that tailored conversational modeling is necessary for effective dialogue verification, but additional research is required before fully autonomous systems can be deployed reliably across broad, specialized enterprise domains.

Gupta et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for DialFact: A Benchmark for Fact-Checking in Dialogue

Abstract

Fact-checking is an essential tool to mitigate the spread of misinformation and disinformation. We introduce the task of fact-checking in dialogue, which is a relatively unexplored area. We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia. There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information. We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue. We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Task Background
  • 4 Dataset Construction and Annotation
  • 4.1 Automatically Generated Claims
  • 4.1.1 Methods for claim generation
  • 4.1.2 Evidence set creation
  • 4.1.3 Claim and Evidence Annotation
  • 4.2 Human Written Claims
  • 4.3 Dataset Statistics
  • 4.4 Quality Control
  • 5 Experiments
  • 5.1 Verifiable Claim Detection
  • 5.2 Evidence Retrieval
  • 5.2.1 Document Retrieval
  • 5.2.2 Evidence Sentence Selection
  • 5.3 Claim Verification
  • 5.3.1 Baselines
  • 5.3.2 Results
  • 5.3.3 Discussion
  • 6 Conclusion
  • Ethical Considerations & Broader Impact
  • References
  • A Supplementary Results
  • B Implementation Details
  • C AMT Instructions

Knowls

  1. Knowl 1 — Dialogue fact-checking is a three-stage task with four response outcomes

    definition

    DIALFACT treats the final response in a dialogue as the claim to check, using the preceding conversation as context and Wikipedia as the background knowledge source. The task consists of deciding whether the response contains verifiable factual information, retrieving evidence relevant to it, and classifying it as SUPPORTED, REFUTED, or NOT ENOUGH INFORMATION (NEI). A response is NON-VERIFIABLE if it contains no factual information that can be checked against the source, including personal opinions or personal information; such responses are assigned NEI. A VERIFIABLE response contains at least one checkable fact and can receive any of the three verification labels. SUPPORTED means its factual content is valid in light of the evidence; REFUTED means factual content is invalid in light of the evidence; NEI means the available evidence cannot establish either truth or falsity.

  2. Knowl 2 — Synthetic dialogue claims are generated and selected for diversity and difficulty

    algorithm

    DIALFACT automatically creates claims for Wizard of Wikipedia dialogue contexts by modifying original responses or generating new ones. For REFUTED claims, the methods include 42 rule-based negation transformations; entity substitutions that replace an entity with one of the same type drawn from the conversation or associated Wikipedia articles, restricted to entities present in the original knowledge sentence; sense-based entity substitutions using sense2vec; and adjective replacement with WordNet antonyms, excluding emotion-related adjectives. For REFUTED or NEI claims, a mask-and-fill method masks salient response tokens with a Neutrality Masker and infills them using a T5-base model trained on Wizard of Wikipedia to reconstruct responses from dialogue context and evidence. Infilling is conditioned on empty evidence, randomly selected articles associated with the original response, or articles about sense2vec-related entities. A separate generation method fine-tunes Blenderbot on Wizard of Wikipedia and samples with temperature 1.5 and nucleus probability 0.9, using the same evidence conditions to encourage unexpected or contextually mismatched content.

    For each context, the candidate responses from these methods are filtered to remove responses with fewer than three words different from another candidate and responses whose GPT-2 perplexity exceeds 1.1 times the candidate-set average. Four existing dialogue inference and fact-checking models then score the candidates. The entropy of each model’s predicted label distribution is summed across models, and the four highest-scoring candidates—those estimated to be most confusing to classify—are retained for that context.

  3. Knowl 3 — Evidence collection and crowd annotation produce labeled dialogue claims

    model/method

    For each claim, DIALFACT creates candidate evidence by collecting named entities and noun phrases from the claim, dialogue context, original Wizard of Wikipedia response, and titles of the Wikipedia articles associated with that response. The MediaWiki API is used to find relevant pages; the first ten sentences of each page are ranked with spaCy word2vec similarity and BM25, and the non-overlapping top-ten results from both rankings are combined. The original Wizard of Wikipedia knowledge sentence is added if absent.

    Crowd workers see the dialogue context, response, and candidate evidence. They label the response VERIFIABLE or NON-VERIFIABLE, select evidence, search Wikipedia and add evidence when needed, and assign SUPPORTED, REFUTED, or NEI; NON-VERIFIABLE responses are assigned NEI. For automatically generated claims, the first annotation round also allows workers to edit claims for contextual appropriateness or flag incoherent claims; 5% were removed as incoherent. Two additional annotations are collected, the majority of three labels is used, and evidence is the union selected across rounds. For human-written claims, workers first write VERIFIABLE responses conditioned on the dialogue and evidence for a pre-specified target label. Subsequent labeling rounds check that target label; claims that fail to match it after the third round are dropped, amounting to about 7% of human-written claims.

  4. Knowl 4 — DIALFACT contains 22,245 claims with moderate label agreement

    data/table

    DIALFACT has 10,436 validation claims and 11,809 test claims, for 22,245 total. The validation set covers 3,738 dialogue contexts and averages 2.8 claims per context; the test set covers 3,760 contexts and averages 3.1 claims per context. Average claim length is 20.0 tokens in validation and 22.0 in test; average evidence count is 1.1 and 1.3 per claim, respectively. The counts below are ordered as SUPPORTED, REFUTED, NEI-Factual, and NEI-Personal.

    Validation: automatically generated claims = 1,686, 1,047, 150, 1,745 (total 4,628); human-written claims = 1,656, 2,316, 1,836, 0 (total 5,808); combined totals = 3,342, 3,363, 1,986, 1,745 (total 10,436).

    Test: automatically generated claims = 2,446, 1,195, 1,278, 1,305 (total 6,224); human-written claims = 1,493, 2,740, 1,268, 84 (total 5,585); combined totals = 3,939, 3,935, 2,546, 1,389 (total 11,809).

    In additional annotations of 1,200 claims, Krippendorff’s alpha for the three-way category label was 0.68 for human-written claims and 0.58 for generated claims. Agreement on VERIFIABLE versus NON-VERIFIABLE was lower, with alpha 0.49; annotators sometimes disagreed on whether evaluative statements were opinions or checkable claims. Analysis of test-set bigrams found no obvious category bias from negation phrases such as do not or is not: prominent REFUTED bigrams were mostly topical, and top bigrams overlapped across categories.

  5. Knowl 5 — Aug-WoW trains claim verification from weakly supervised synthetic data

    model/method

    Because DIALFACT is intended as a validation and test benchmark, the authors do not train the proposed verifier on its labeled claims. Instead, Aug-WoW is trained on automatically constructed Wizard of Wikipedia data. Its 46,934 SUPPORTED examples are original response–knowledge-sentence pairs after NON-VERIFIABLE responses are filtered with the lexical-overlap baseline. It adds 38,895 REFUTED examples produced with negation and substitution, and selects 40,000 NEI examples from two constructions: replacing a context–claim–evidence triplet’s evidence with unrelated evidence, or generating a response conditioned on random evidence. The synthetic data is used to fine-tune the Colloquial claim-verification model. The verifier takes the last two dialogue turns and claim as input; evidence sentences are concatenated for evidence-based BERT model inputs. At evaluation, the model is tested with oracle evidence or retrieved evidence, rather than being trained on DIALFACT labels.

  6. Knowl 6 — Dialogue context improves Wikipedia document and evidence retrieval

    empirical result

    Document retrieval was evaluated by recall on the DIALFACT test set, counting a document as relevant when it contains a gold evidence sentence. WikiAPI extracts entities, retrieves up to three Wikipedia pages per entity, and uses the first five KILT paragraphs; its context version adds the last two dialogue turns to the claim. DPR retrieves the top 100 documents, with an original question-answering model and Wizard of Wikipedia fine-tuned versions using either the claim alone or claim plus context. Document recall was: DPR-original 40.3%; DPR-WoWft-claimonly 44.7%; DPR-WoWft-ctx 58.8%; Wiki-claimonly 60.8%; Wiki-ctx 75.0%. Thus, adding dialogue context improved both retrieval families, and WikiAPI with context performed best.

    For evidence sentence selection, a BERT-base ranker scores the union of sentences from retrieved documents, trained on Wizard of Wikipedia context-response data with random and hard negatives drawn from articles shown to the wizard. Sentences scoring above zero are eligible for the top-five evidence set. Test recall@5 for claim-only versus context-conditioned selection was 67.1% versus 69.3% with DPR-WoWft-ctx documents, and 70.1% versus 75.4% with Wiki-ctx documents. The gains show that context helps not only find documents but also identify relevant sentences within them.

  7. Knowl 7 — Aug-WoW outperforms existing baselines on three-way claim verification

    empirical result

    Claim verification was evaluated on the DIALFACT test set with NON-VERIFIABLE claims included as NEI. Accuracy and macro F1 are percentages, reported as accuracy/macro F1 for oracle evidence, Wiki-retrieved evidence, and DPR-retrieved evidence, respectively. Oracle evidence is gold evidence; Wiki evidence uses Wiki-ctx document retrieval with context-aware sentence selection; DPR evidence uses DPR-WoWft-ctx retrieval with context-aware sentence selection. At most five evidence sentences are supplied.

    DNLI: 43.3/35.4, 39.1/31.5, 38.4/29.5. DECODE: 37.8/30.3, 35.3/25.3, 34.5/22.5. VitaminC: 57.6/56.1, 46.2/44.7, 45.9/44.2. CorefBert-Colloquial: 61.4/60.0, 47.6/45.2, 46.4/41.1. Colloquial: 63.5/62.8, 48.1/46.3, 48.7/46.4. Aug-WoW: 69.2/69.0, 51.6/51.3, 51.5/50.2.

    Aug-WoW is the strongest model in each evidence condition, but every model performs worse with retrieved evidence than with oracle evidence. Even with gold evidence, no model exceeds 70% accuracy, indicating that both evidence retrieval and claim reasoning remain difficult.

  8. Knowl 8 — Verifiable-claim detection remains difficult for the non-verifiable class

    empirical result

    On the DIALFACT test set, the verifiable-claim detector predicts VERIFIABLE versus NON-VERIFIABLE using a threshold selected on validation data. The lexical baseline uses maximum word overlap between a claim and its evidence after punctuation and stopword removal; DNLI uses the probability of the neutral class; Lexical+DNLI sums the baseline scores. Reported accuracy, VERIFIABLE F1, and NON-VERIFIABLE F1 were: Random, 50.0%, 64.2%, 19.2%; Lexical, 79.4%, 88.1%, 33.8%; DNLI, 82.1%, 89.9%, 37.1%; Lexical+DNLI, 82.8%, 90.2%, 39.1%. Combining the signals gives the best overall result, but all baselines have low F1 for identifying NON-VERIFIABLE responses.

  9. Knowl 9 — Verifier ablations expose sensitivity to context, evidence quality, and claim type

    empirical result

    For Aug-WoW, removing dialogue context yields 68.1% accuracy and 68.1% macro F1 with oracle evidence, 52.4% and 52.3% with Wiki evidence, and 52.4% and 51.3% with DPR evidence. The full model scores 69.2%/69.0%, 51.6%/51.3%, and 51.5%/50.2% in those conditions. Replacing the base encoder with BERT-large raises oracle performance to 70.9%/70.9% but lowers retrieved-evidence performance to 45.8%/44.6% with Wiki and 43.5%/39.1% with DPR. A claim-only Aug-WoW model without evidence scores 33.2% accuracy and 28.9% macro F1, showing that claim text alone does not provide strong test performance.

    With oracle evidence, Aug-WoW scores 63.9% accuracy and 60.7% macro F1 on automatically generated claims, compared with 74.2% and 74.0% on human-written claims. The authors attribute the generated-claim difficulty partly to selecting those claims for high predicted confusion. Their error analysis finds that models can treat lexical overlap as support while missing semantic mismatch, fail at commonsense reasoning, and be misled by colloquial wording or local negation. NEI is the weakest class, with confusion particularly between NEI and REFUTED.

  10. Knowl 10 — DIALFACT is a Wikipedia-grounded benchmark rather than a universal fact-checker

    limitation

    DIALFACT is built from Wizard of Wikipedia dialogues and uses Wikipedia as its background corpus, so its coverage does not establish performance in other domains or against other knowledge sources. The authors caution that labels may still be incorrect or biased despite quality control; evidence collection and annotation can miss relevant support or refutation. Moderate category agreement and lower agreement on verifiability further indicate annotation ambiguity. The benchmark should therefore not be treated as a universal fact-checking tool, and systems evaluated on it may not be reliable when transferred to domains or resources outside its scope.

Coverage note — The paper’s supplemental two-way verification results and implementation minutiae for individual baseline models were omitted because they add less to reconstructing the benchmark, its main methods, and its principal findings.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977.
  2. 2.Rami Aly, Zhijiang Guo, Michael Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. Feverous: Fact extraction and verification over unstructured and structured information. arXiv preprint arXiv:2106.05707.
  3. 3.Pepa Atanasova, Dustin Wright, and Isabelle Augenstein. 2020. Generating label cohesive and well-formed adversarial claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3168–3177, Online. Association for Computational Linguistics.
  4. 4.Giannis Bekoulis, Christina Papagiannopoulou, and Nikos Deligiannis. 2021. A review on fact extraction and verification. ACM Comput. Surv., 55(1).
  5. 5.David DeVault and Matthew Stone. 2007. Managing ambiguities across utterances in dialogue. In Proceedings of the 11th Workshop on the Semantics and Pragmatics of Dialogue - Full Papers, Roverto, Italy. SEMDIAL.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019. Wizard of wikipedia: Knowledge-powered conversational agents. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  8. 8.Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2021. Evaluating groundedness in dialogue systems: The begin benchmark. arXiv preprint arXiv:2105.00071.
  9. 9.Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021. A survey of data augmentation approaches for NLP. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 968–988, Online. Association for Computational Linguistics.
  10. 10.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018. AllenNLP: A deep semantic natural language processing platform. In Proceedings of Workshop for NLP Open Source Software (NLP-OSS), pages 1–6, Melbourne, Australia. Association for Computational Linguistics.
  11. 11.Sarik Ghazarian, Zixi Liu, Akash S M, Ralph Weischedel, Aram Galstyan, and Nanyun Peng. 2021. Plot-guided adversarial example construction for evaluating open-domain story generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4334–4344, Online. Association for Computational Linguistics.
  12. 12.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2021. A survey on automated fact-checking. arXiv preprint arXiv:2108.11896.
  13. 13.Prakhar Gupta, Yulia Tsvetkov, and Jeffrey Bigham. 2021. Synthesizing adversarial negative responses for robust response ranking and evaluation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3867–3883, Online. Association for Computational Linguistics.
  14. 14.Andreas Hanselowski, Hao Zhang, Zile Li, Daniil Sorokin, Benjamin Schiller, Claudia Schulz, and Iryna Gurevych. 2018. UKP-athene: Multi-sentence textual entailment for claim verification. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER), pages 103–108, Brussels, Belgium. Association for Computational Linguistics.
  15. 15.Christopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona Diab, and Smaranda Muresan. 2020. DeSePtion: Dual sequence prediction and adversarial examples for improved fact-checking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8593–8606, Online. Association for Computational Linguistics.
  16. 16.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  17. 17.Matthew Honnibal and Ines Montani. 2017. spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. To appear.
  18. 18.Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. Q2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering. arXiv preprint arXiv:2104.08202.
  19. 19.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. HoVer: A dataset for many-hop fact extraction and claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441–3460, Online. Association for Computational Linguistics.
  20. 20.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781, Online. Association for Computational Linguistics.
  21. 21.Byeongchang Kim, Hyunwoo Kim, Seokhee Hong, and Gunhee Kim. 2021. How robust are fact checking systems on colloquial claims? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1535–1548, Online. Association for Computational Linguistics.
  22. 22.Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2021. Internet-augmented dialogue generation. arXiv preprint arXiv:2107.07566.
  23. 23.Zhenghao Liu, Chenyan Xiong, Maosong Sun, and Zhiyuan Liu. 2020. Fine-grained fact verification with kernel graph attention network. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7342–7351, Online. Association for Computational Linguistics.
  24. 24.George A Miller. 1998. WordNet: An electronic lexical database. MIT press.
  25. 25.Preslav Nakov, Giovanni Da San Martino, Tamer Elsayed, Alberto Barrón-Cedeño, Rubén Míguez, Shaden Shaar, Firoj Alam, Fatima Haouari, Maram Hasanain, Nikolay Babulkov, Alex Nikolov, Gautam Kishore Shahi, Julia Maria Struß, and Thomas Mandl. 2021. The clef-2021 checkthat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In Advances in Information Retrieval, pages 639–649, Cham. Springer International Publishing.
  26. 26.Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, and Jason Weston. 2021. I like fish, especially dolphins: Addressing contradictions in dialogue modeling. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1699–1713, Online. Association for Computational Linguistics.
  27. 27.Piotr Niewinski, Maria Pszona, and Maria Janicka. 2019. GEM: Generative enhanced model for adversarial attacks. In Proceedings of the Second Workshop on Fact Extraction and VERification (FEVER), pages 20–26, Hong Kong, China. Association for Computational Linguistics.
  28. 28.Jeppe Nørregaard and Leon Derczynski. 2021. DanFEVER: claim verification dataset for Danish. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 422–428, Reykjavik, Iceland (Online). Linköping University Electronic Press, Sweden.
  29. 29.Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021. KILT: a benchmark for knowledge intensive language tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online. Association for Computational Linguistics.
  30. 30.Libo Qin, Tianbao Xie, Shijue Huang, Qiguang Chen, Xiao Xu, and Wanxiang Che. 2021. Don’t be contradicted with anything! CI-ToD: Towards benchmarking consistency for task-oriented dialogue system. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2357–2367, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  33. 33.Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021. Increasing faithfulness in knowledge-grounded dialogue with controllable features. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 704–718, Online. Association for Computational Linguistics.
  34. 34.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325, Online. Association for Computational Linguistics.
  35. 35.Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. COVID-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2116–2129, Online. Association for Computational Linguistics.
  36. 36.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Winogrande: An adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106.
  37. 37.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get your vitamin C! robust fact verification with contrastive evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 624–643, Online. Association for Computational Linguistics.
  38. 38.Tal Schuster, Darsh Shah, Yun Jie Serene Yeo, Daniel Roberto Filizzola Ortiz, Enrico Santus, and Regina Barzilay. 2019. Towards debiasing fact verification models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3419–3425, Hong Kong, China. Association for Computational Linguistics.
  39. 39.Darsh J Shah, Tal Schuster, and Regina Barzilay. 2020. Automatic fact-guided sentence modification. In Association for the Advancement of Artificial Intelligence (AAAI).
  40. 40.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567.
  41. 41.Dominik Stammbach and Guenter Neumann. 2019. Team DOMLIN: Exploiting evidence enhancement for the FEVER shared task. In Proceedings of the Second Workshop on Fact Extraction and VERification (FEVER), pages 105–109, Hong Kong, China. Association for Computational Linguistics.
  42. 42.James Thorne and Andreas Vlachos. 2021. Evidence-based factual error correction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3298–3309, Online. Association for Computational Linguistics.
  43. 43.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  44. 44.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2019. Evaluating adversarial attacks against multiple fact verification systems. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2944–2953, Hong Kong, China. Association for Computational Linguistics.
  45. 45.Andrew Trask, Phil Michalak, and John Liu. 2015. sense2vec-a fast and accurate method for word sense disambiguation in neural word embeddings. arXiv preprint arXiv:1511.06388.
  46. 46.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online. Association for Computational Linguistics.
  47. 47.William Yang Wang. 2017. “liar, liar pants on fire”: A new benchmark dataset for fake news detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 422–426, Vancouver, Canada. Association for Computational Linguistics.
  48. 48.Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019. Dialogue natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3731–3741, Florence, Italy. Association for Computational Linguistics.
  49. 49.Jianshu Chen Wenhu Chen, Hongmin Wang, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact : A large-scale dataset for table-based fact verification. In International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia.
  50. 50.Wenquan Wu, Zhen Guo, Xiangyang Zhou, Hua Wu, Xiyuan Zhang, Rongzhong Lian, and Haifeng Wang. 2019. Proactive human-machine conversation with explicit conversation goal. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3794–3804, Florence, Italy. Association for Computational Linguistics.
  51. 51.Jing Xu, Arthur Szlam, and Jason Weston. 2021. Beyond goldfish memory: Long-term open-domain conversation. arXiv preprint arXiv:2107.07567.
  52. 52.Deming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu, Peng Li, Maosong Sun, and Zhiyuan Liu. 2020. Coreferential Reasoning Learning for Language Representation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7170–7186, Online. Association for Computational Linguistics.

Citation

MLA
Gupta, P., et al. “DialFact: A Benchmark for Fact-Checking in Dialogue”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3785–801, https://doi.org/10.18653/v1/2022.acl-long.263.
APA
Gupta, P., Wu, C.-S., Liu, W., & Xiong, C. (2022). DialFact: A Benchmark for Fact-Checking in Dialogue. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3785–3801. https://doi.org/10.18653/v1/2022.acl-long.263
Chicago
Gupta, P., C.-S. Wu, W. Liu, and C. Xiong. 2022. “DialFact: A Benchmark for Fact-Checking in Dialogue”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3785–3801. https://doi.org/10.18653/v1/2022.acl-long.263.
Harvard
Gupta, P. et al. (2022) “DialFact: A Benchmark for Fact-Checking in Dialogue”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3785–3801. Available at: https://doi.org/10.18653/v1/2022.acl-long.263.
Vancouver
1. Gupta P, Wu C-S, Liu W, Xiong C (2022) DialFact: A Benchmark for Fact-Checking in Dialogue. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3785–3801

BibTeX

@inproceedings{gupta-etal-2022-dialfact,
    title = "{D}ial{F}act: A Benchmark for Fact-Checking in Dialogue",
    author = "Gupta, Prakhar  and
      Wu, Chien-Sheng  and
      Liu, Wenhao  and
      Xiong, Caiming",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.263/",
    doi = "10.18653/v1/2022.acl-long.263",
    pages = "3785--3801"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/