DialFact: A Benchmark for Fact-Checking in Dialogue
Prakhar GuptaChien-Sheng WuWenhao LiuCaiming Xiong
Introduces a benchmark of over 22,000 annotated conversational claims to advance fact verification in dialogue, revealing why standard non-dialogue models fail and offering a data-efficient training strategy to address conversational challenges like coreference and colloquialisms.
Online misinformation and inaccurate statements generated by automated conversational systems present serious risks, particularly during public crises. While automated fact-checking has advanced for formal texts like news and encyclopedia articles, verifying factual correctness in everyday conversations remains challenging. Conversational statements are typically informal, sparse in factual details, full of colloquialisms and personal opinions, and heavily reliant on prior conversational context to resolve references.
The main objective of the article is to establish the first dedicated benchmark dataset and evaluation pipeline for conversational fact-checking, and to demonstrate practical modeling techniques that improve fact verification performance in dialogue settings.
To address this challenge, the authors created DIALFACT, a benchmark consisting of 22,245 conversational claims derived from the Wizard of Wikipedia dialogue dataset and paired with Wikipedia evidence. The dataset includes both human-written and machine-generated responses. Claims were generated through rule-based mutations, language model text infilling, and open-domain dialogue generation, followed by rigorous multi-round annotations and quality screening on Amazon Mechanical Turk. The fact-checking pipeline was broken into three sequential tasks: detecting verifiable factual claims versus personal opinions, retrieving relevant evidence from Wikipedia, and classifying whether the evidence supports, refutes, or provides insufficient information to judge the claim.
The investigation revealed four key findings. First, existing fact-checking models trained on traditional, non-conversational data perform poorly on conversational claims. Second, providing dialogue context substantially improves evidence retrieval; incorporating conversation history raised document retrieval recall from 60.8% to 75.0% for web-based search and from 44.7% to 58.8% for dense neural retrieval. Third, the authors' weakly supervised training approach, named Aug-WoW, outperformed all standard baselines across all testing conditions, achieving 69.2% claim verification accuracy with ground-truth evidence and approximately 51.5% with automatically retrieved evidence. Fourth, end-to-end performance drops significantly (by roughly 18 percentage points) when models rely on retrieved evidence rather than perfect reference text, underscoring that retrieval is a critical performance bottleneck.
These findings demonstrate that conversational fact-checking cannot be solved simply by repurposing standard fact-checking tools. Because conversational systems are prone to confusing mere topical keyword overlap with factual consistency, deploying them without dialogue-specific retrieval and verification mechanisms creates substantial risks of undetected errors or false alarms. Automated dialogue agents operating in customer service, public information, or advisory capacities require tailored architectures to remain reliable.
Organizations developing or deploying conversational agents should integrate dialogue context into their knowledge retrieval pipelines rather than evaluating claims in isolation. Practitioners should adopt weakly supervised data augmentation methods—such as entity swapping, negation, and masked generation—to train conversational verification tools cost-effectively. Further development should prioritize improving the precision of initial evidence retrieval and enhancing models' ability to distinguish between subjective opinions and unsupported factual assertions.
The conclusions should be interpreted within certain boundaries. The evaluation relies exclusively on English Wikipedia as the knowledge source and centers around open-domain conversational topics. While human validation confirmed solid data consistency, borderline cases between subjective personal opinions and factual claims remain a source of modest uncertainty. Overall, confidence is high that tailored conversational modeling is necessary for effective dialogue verification, but additional research is required before fully autonomous systems can be deployed reliably across broad, specialized enterprise domains.
- Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). FEVER establishes the Wikipedia-based retrieval, evidence selection, and three-way claim-verification pipeline that DialFact adapts to conversational claims.
No sufficiently relevant recommendations were found.
