topic
verification tools
Verification tools are software programs used to mathematically check that hardware and software systems satisfy their specified requirements and correctness properties. Rooted in formal logic and theoretical computer science, these tools employ techniques such as model checking, automated theorem proving, and static analysis to evaluate system designs, execution models, or source code. Rather than relying solely on empirical testing, verification tools systematically analyze system behaviors and state spaces to detect subtle design flaws, logical inconsistencies, and vulnerabilities, making them essential for ensuring reliability, correctness, and security in safety-critical applications.
2 items

Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation
Max Glockner, Yufang Hou, Iryna Gurevych
Why you should read this
Reveals that automated fact-checking models fail on real-world misinformation because they rely on leaked counter-evidence from post-hoc reports rather than disproving the underlying reasoning behind novel claims.
The task of misinformation detection has tremendous potential to make a significant contribution to society. However, current automatic fact-checking systems ignore an important aspect of fact-checking: the lack of counter-evidence. In this paper, we analyze the prevalence of counter-evidence in real-world fact-checking datasets. We find that counter-evidence is almost always present in the evidence used by human fact-checkers, but is missing in the evidence retrieved by current automatic fact-checking systems. We argue that this discrepancy is a major reason for the lack of robustness of current systems. To address this issue, we propose a new task setting that requires the system to identify the lack of counter-evidence and to abstain from making a prediction in such cases. We show that this is a challenging task for current systems, and that our proposed method can improve the robustness of fact-checking systems.
Added
2026-10-03

DialFact: A Benchmark for Fact-Checking in Dialogue
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, Caiming Xiong
Why you should read this
Introduces a benchmark of over 22,000 annotated conversational claims to advance fact verification in dialogue, revealing why standard non-dialogue models fail and offering a data-efficient training strategy to address conversational challenges like coreference and colloquialisms.
Fact-checking is an essential tool to mitigate the spread of misinformation and disinformation. We introduce the task of fact-checking in dialogue, which is a relatively unexplored area. We construct DIALFACT, a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia. There are three sub-tasks in DIALFACT: 1) Verifiable claim detection task distinguishes whether a response carries verifiable factual information; 2) Evidence retrieval task retrieves the most relevant Wikipedia snippets as evidence; 3) Claim verification task predicts a dialogue response to be supported, refuted, or not enough information. We found that existing fact-checking models trained on non-dialogue data like FEVER (Thorne et al., 2018) fail to perform well on our task, and thus, we propose a simple yet data-efficient solution to effectively improve fact-checking performance in dialogue. We point out unique challenges in DIALFACT such as handling the colloquialisms, coreferences and retrieval ambiguities in the error analysis to shed light on future research in this direction.
Added
2026-10-02
