Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation

Max GlocknerYufang HouIryna Gurevych

article2022EMNLP61 citations

Reveals that automated fact-checking models fail on real-world misinformation because they rely on leaked counter-evidence from post-hoc reports rather than disproving the underlying reasoning behind novel claims.

Listen

Misinformation frequently emerges during times of high public uncertainty when credible information is scarce, leading to significant real-world harm across public health, politics, and social stability. Natural language processing systems have been developed to automate the fact-checking pipeline, but these systems typically rely on finding direct counter-evidence in trusted databases to disprove false claims. The article evaluates whether standard automated fact-checking methods can realistically refute novel, real-world misinformation in the same way professional human fact-checkers do.

To assess this, the researchers compared automated fact-checking task formulations against human journalistic practices through an analysis of 100 verified misinformation claims from PolitiFact and Snopes. They established two critical criteria for realistic evaluation: evidence must be sufficient to justify a verdict and must remain unleaked, meaning it does not incorporate reports produced only after the claim was already fact-checked. The authors surveyed 16 established fact-checking datasets and conducted empirical experiments using a large-scale dataset, MultiFC, training deep learning language models on claim-evidence pairs to observe their reliance on leaked text.

The findings show that current automated systems operate on fundamentally unrealistic assumptions. In the human verification study, professional journalists used direct global counter-evidence for only about 26.7% of claims. For roughly 65.3% of claims, journalists refuted misinformation by identifying the underlying premise or claimant's original source and disproving that specific rationale. Furthermore, the dataset survey revealed that not a single existing dataset containing real-world misinformation satisfies both the sufficient and unleaked evidence criteria. In MultiFC, approximately 69.7% of misinformation claims were found to contain leaked evidence from post-verification articles. When tested on PolitiFact data, language models experienced severe performance drops when deprived of leaked cues, with accuracy falling from 57.6% on leaked evidence to 25.8% on unleaked evidence.

These results demonstrate that existing automated fact-checking models do not genuinely reason over evidence to debunk novel false claims. Instead, they exploit statistical shortcuts and leaked conclusions from existing journalistic fact-checks, rendering them ineffective for newly emerging rumors where no prior debunking exists. Relying on current benchmarks introduces substantial operational risk, as stakeholders might deploy systems that appear highly accurate in testing but fail entirely when confronted with zero-day misinformation campaigns.

The article recommends restructuring natural language processing fact-checking pipelines to emulate human journalistic methodology. Rather than assuming global counter-evidence exists, automated systems should incorporate automated provenance detection, context tracking across online platforms, and the identification of logical fallacies. Stakeholders should avoid deploying automated verification tools for novel claims until evaluation benchmarks eliminate data leakage and support source-tracing workflows.

The main limitations of the study include its focus on English-language political and news claims from two prominent fact-checking platforms, which may not capture all forms of digital misinformation, such as multi-modal images or non-English content. While confidence in the technical findings regarding benchmark leakage and current model shortcomings is very high, further research and improved dataset designs are required before automated systems can reliably verify novel misinformation in production environments.

Glockner et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation

Abstract

The task of misinformation detection has tremendous potential to make a significant contribution to society. However, current automatic fact-checking systems ignore an important aspect of fact-checking: the lack of counter-evidence. In this paper, we analyze the prevalence of counter-evidence in real-world fact-checking datasets. We find that counter-evidence is almost always present in the evidence used by human fact-checkers, but is missing in the evidence retrieved by current automatic fact-checking systems. We argue that this discrepancy is a major reason for the lack of robustness of current systems. To address this issue, we propose a new task setting that requires the system to identify the lack of counter-evidence and to abstain from making a prediction in such cases. We show that this is a challenging task for current systems, and that our proposed method can improve the robustness of fact-checking systems.

Table of Contents

  • 1 Introduction
  • 2 How Humans Fact-check
  • 3 Can FCNLP Help Human Verification?
  • 3.1 Human Verification Strategies
  • 3.2 NLP Fact Verification
  • 3.3 Human and NLP Comparison
  • 4 NLP Fact-Checking Datasets
  • 5 A Case Study of Leaked Evidence
  • 5.1 Quantification of Leaked Evidence
  • 5.2 Impact on Trained Systems
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Human Misinformation Verification Examples
  • B Leaked Evidence Analysis
  • B.1 Misinformation Labels
  • B.2 Automatic Identification of Leaked Evidence
  • B.3 Manual Guidelines
  • B.3.1 Leaked Evidence Snippets
  • B.3.2 Stance of Evidence Snippets
  • C Experiments on MULTIFC
  • C.1 Training details
  • C.2 Performance per Label
  • C.3 Evaluation on Identical Claims with Different Evidence
  • C.4 Comparison with a Claim-Only Baseline

Knowls

  1. Knowl 1 — Claim-content evidence retrieval excludes most source-dependent refutations

    model/method

    Evidence-based NLP fact-checking systems in the FCNLP framework retrieve evidence from a claim’s semantic content and use the evidence’s stance toward the claim to predict a verdict. This design assumes that the claim itself can lead to global counter-evidence: evidence that contradicts the claim without establishing where the claim came from or what reasoning produced it. Many misinformation claims instead require local counter-evidence tied to the claimant’s source or reasoning. Claim content alone cannot establish that two different accounts describe the same event, or that a particular document was the basis for a claim. Establishing that connection from the claim’s falsity would be circular: the connection is needed to refute the claim, but recognizing it may depend on already knowing the claim is false. Thus, for this class of systems, evidence retrieval based only on claim content cannot supply the source guarantee needed for local refutation and is limited to global counter-evidence. In the authors’ sample, global counter-evidence was the primary strategy for 20 of 75 applicable claims (26.7%); those 75 were drawn from a 100-claim audit, so the 20 global-counter-evidence cases were 20% of the full sample.

  2. Knowl 2 — Source guarantee and context availability enable verification without global counter-evidence

    definition

    A source guarantee is assurance that identified evidence either is the claimant’s reason for a claim or refers to that reason. This connection lets a fact-checker evaluate the claim’s underlying reasoning, including whether it misinterprets a source, relies on an invalid source, or draws an unsupported conclusion. Once the connection is established, the evidence need not directly contradict the claim: showing that the claim does not follow from its source, is speculative, or relies on a non-credible source can be enough to refute it.

    Context availability means access to the claim’s original environment sufficient to understand it unambiguously and, when needed, trace the claim and its sources across platforms. It is a logical precondition for establishing the source guarantee.

  3. Knowl 3 — Human fact-checkers primarily use source-dependent strategies on misinformation

    empirical result

    The authors manually studied 100 false or mostly false claims, randomly sampling 50 from PolitiFact and 50 from Snopes. For each organization, the sample included 25 claims represented in MULTIFC and 25 claims from 2020–2021. Twenty-five claims were excluded because verification required, for example, identifying scam pages or imposter content, or multimodal reasoning; the strategy analysis therefore covers 75 claims.

    Fact-checkers used global counter-evidence for 20 claims (26.7% of the 75), local counter-evidence tied to the claimant’s reason for 35 (46.7%), a non-credible source argument tied to that reason for 14 (18.7%), a no-evidence assertion for 5 (6.7%), and another strategy for 1 (1.3%). The local-counter-evidence and non-credible-source strategies together relied on a source guarantee in 49 of 75 cases (65.3%). The study therefore found that fact-checkers commonly refute misinformation by examining its source or reasoning, rather than by finding direct contradiction. Separately, a crawl of 20,274 PolitiFact claims from 2007–2021 showed that false verdicts became more prevalent after 2016, with fewer than 10% of claims selected for fact-checking in 2021 judged correct.

  4. Knowl 4 — Realistic fact-checking evidence must be sufficient and unleaked

    definition

    For a dataset to test evidence-based systems on real-world misinformation, its evidence must meet two requirements. Sufficient evidence must enable a human to justify the verdict. Unleaked evidence must not contain information that became available only after the claim was verified. Evidence taken from a fact-checking report, or from later information that depends on the verification, violates the unleaked requirement. The requirements are complementary: evidence can be useful enough to support a verdict yet be leaked, or be available independently of a fact-check yet fail to justify the verdict.

  5. Knowl 5 — Existing datasets with real-world misinformation do not satisfy both evidence requirements

    data/table

    The survey assessed whether evidence in fact-checking datasets was unleaked, sufficient for a human to justify the verdict, and annotated for its stance toward the claim. None of the surveyed datasets containing real-world misinformation satisfied both the sufficient and unleaked criteria.

    • LIARPLUS and POLITIHOP use fact-checking-article evidence. The survey rates that evidence sufficient and stance-annotated, but leaked.
    • CLIMATEFEVER and HEALTHVER use external web evidence. The evidence is rated unleaked and stance-annotated, but insufficient.
    • UKP-SNOPES, PUBHEALTH, and WATCLAIMCHECK are rated unleaked but insufficient. UKP-SNOPES includes stance annotations; PUBHEALTH and WATCLAIMCHECK do not. In UKP-SNOPES, the annotated evidence often has a stance that conflicts with the verdict, and annotators found no stance for 45.5% of selected evidence snippets.
    • Baly et al. (2018) is rated sufficient and stance-annotated, but leaked. MULTIFC and X-FACT are rated neither sufficient nor unleaked, and lack stance annotations.

    The survey also considered six datasets that do not provide real-world misinformation claims: SCIFACT, COVID-FACT, WIKIFACTCHECK, and FM2 use generated claims; Thorne et al. (2021) and FAVIQ use paraphrased user queries. Their construction ensures or presupposes evidence availability, so they do not resolve the evidence problems identified for real-world misinformation.

  6. Knowl 6 — Most MULTIFC misinformation claims have automatically detectable leaked evidence

    empirical result

    The authors analyzed 16,244 MULTIFC claims labeled as misinformation. They marked a snippet as leaked when its source URL matched a fact-checking-organization pattern or its title or text matched a regular expression associated with fact-checking or debunking language. This pattern-based method identified leaked evidence for 11,267 claims (69.7%): URL patterns identified 8,999 claims (55.6%), and phrase patterns identified 9,656 (59.7%). These categories overlap.

    To check the detections, the authors manually assessed 230 automatically flagged snippets from 100 claims. They judged 83.9% of the snippets to be leaked, and 97 of the 100 claims had at least one leaked snippet. The results show that leakage is common in MULTIFC and that it can be detected in many claims through evidence-source and text cues.

  7. Knowl 7 — A manual audit finds further leakage and widespread evidence insufficiency

    empirical result

    A separate manual audit examined evidence for 100 MULTIFC misinformation claims for which the automatic pattern method found no leaked snippet. The manual audit found additional leaked evidence for 32 claims; 68 had no leaked evidence identified. Across these claims, 37 had evidence with no discernible stance and 63 had no refuting evidence. Fifteen claims had unleaked evidence that refuted at least part of the claim; in 10 of those cases, leaked evidence was also present, leaving only 5 with refuting evidence and no detected leaked evidence. These categories overlap.

    The audit also found evidence supporting 40 claims: for 35, the supporting evidence was itself misinformation; for the other 5, later evidence made the claim accurate, reflecting a change in what was known after the claim was made. The findings illustrate why a single evidence snippet may not suffice and why available snippets can be non-refuting, temporally misleading, or leaked.

  8. Knowl 8 — BERT experiments compare verdict prediction with leaked and unleaked MULTIFC evidence

    experimental setup

    The authors fine-tuned BERT-base-uncased to predict veracity labels for MULTIFC claims from Snopes and PolitiFact. Evidence snippets were concatenated in their supplied order, separated by semicolons, and truncated at 512 tokens. The models received one of four inputs: snippet text only, snippet titles only, complete snippets (title and text), or the claim together with the complete snippets. The claim and evidence were separated by a BERT [SEP] token; a linear layer over the [CLS] representation predicted the label.

    Models were trained for 5 epochs with a learning rate of 2×10−52\times10^{-5} and batch size 16. The model with the highest development-set F1, checked after each epoch, was selected. Other parameters were left at their defaults, no hyperparameters were tuned, and results were averaged over three random seeds (1, 2, and 3). The test data were split according to whether the claim had automatically detected leaked evidence: Snopes had 1,014 test claims (482 leaked, 532 unleaked), and PolitiFact had 2,717 (2,111 leaked, 606 unleaked).

  9. Knowl 9 — Verdict models perform substantially better with leaked evidence on PolitiFact

    empirical result

    For each input configuration, the authors report F1-macro/F1-micro scores averaged over three runs, comparing all test claims with subsets containing leaked or unleaked evidence. The final column is the difference in F1-micro between the leaked and unleaked subsets. Scores are percentages; “text” uses snippet bodies, “title” uses snippet titles, “snippets” uses both, and “full” adds the claim to the complete snippets.

    On PolitiFact, scores were consistently much higher for claims with leaked evidence: the F1-micro advantage ranged from 12.7 to 36.2 points, and the full-input model scored 59.7/59.3 on leaked claims versus 25.6/25.9 on unleaked claims. On Snopes, the F1-micro advantage ranged from 11.3 to 13.9 points, while F1-macro was higher on the unleaked subset. The authors attribute this difference to Snopes’ label imbalance toward “false”: leaked evidence encouraged predictions of the majority label, harming performance on other labels. Evaluating the same claims with leaked versus unleaked evidence also supported the finding that models rely on leaked evidence; leaked snippets strongly cued “false,” while the PolitiFact model’s performance nearly doubled with leaked evidence.

    DatasetInputAll claims (macro/micro)Leaked (macro/micro)Unleaked (macro/micro)Micro-F1 difference
    SnopesText29.4/60.426.3/66.330.1/55.0+11.3
    SnopesTitle27.3/57.823.8/64.628.2/51.5+13.1
    SnopesSnippets30.5/60.528.7/67.630.2/53.7+13.9
    SnopesFull32.7/62.730.6/68.733.0/57.2+11.5
    PolitiFactText35.5/34.538.0/37.224.1/24.5+12.7
    PolitiFactTitle48.0/47.455.1/54.621.1/21.6+33.0
    PolitiFactSnippets52.0/51.359.7/59.222.1/23.0+36.2
    PolitiFactFull52.6/51.959.7/59.325.6/25.9+33.4
  10. Knowl 10 — Findings are limited to selected textual claims and two English-language fact-checkers

    limitation

    The analyses concern textual misinformation claims that professional fact-checkers selected as important to verify, not all misinformation or multimodal claims; 25 claims requiring scam-page, imposter-content, or multimodal analysis were excluded from the strategy audit. The human analysis and experiments draw on PolitiFact and Snopes, so their claim-selection practices and domain, language, and geographic biases limit generalization. Fact-checking organizations can use different verification strategies in other settings. The leakage estimates depend on the organizations, time period, and evidence available in MULTIFC, as well as the pattern-based detection method; the authors did not evaluate different evidence-collection strategies or the influence of language, domain, or popularity. The experiments also restrict labels to veracity-scale categories and exclude labels such as “misleading.”

Coverage note — Omitted the appendix’s individual claim examples and detailed per-label score tables; they illustrate or diagnose the reported mechanisms but do not add a separate load-bearing finding.

References

  1. 1.Sahar Abdelnabi, Rakibul Hasan, and Mario Fritz. 2022. Open-Domain, Content-based, Multi-modal Fact-checking of Out-of-Context Images via Online Resources. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14940–14949.
  2. 2.Hamidreza Aghababaeian, Lara Hamdanieh, and Abbas Ostadtaghizadeh. 2020. Alcohol intake in an attempt to fight COVID-19: A medical myth in Iran. Alcohol, 88:29–32.
  3. 3.Tariq Alhindi, Savvas Petridis, and Smaranda Muresan. 2018. Where is your evidence: Improving fact-checking by justification modeling. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER), pages 85–90, Brussels, Belgium. Association for Computational Linguistics.
  4. 4.Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  5. 5.Phoebe Arnold. 2020. The challenges of online fact checking: how technology can (and can’t) help. Technical report, FullFact.
  6. 6.Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4685–4697, Hong Kong, China. Association for Computational Linguistics.
  7. 7.Ramy Baly, Mitra Mohtarami, James Glass, Lluís Màrquez, Alessandro Moschitti, and Preslav Nakov. 2018. Integrating Stance Detection and Fact Checking in a Unified Corpus. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 21–27, New Orleans, Louisiana. Association for Computational Linguistics.
  8. 8.Brooke Borel. 2016. The Chicago guide to fact-checking. University of Chicago Press.
  9. 9.Alexandre Bovet and Hernán A Makse. 2019. Influence of fake news in Twitter during the 2016 US presidential election. Nature communications, 10(1):1–14.
  10. 10.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  11. 11.Steve Buttry. 2014. Verification fundamentals: Rules to live by. Verification Handbook: A Definitive Guide to Verifying Digital Content for Emergency Coverage, pages 15–23.
  12. 12.Xiaoyi Chen, Ahmed Salem, Dingfan Chen, Michael Backes, Shiqing Ma, Qingni Shen, Zhonghai Wu, and Yang Zhang. 2021. BadNL: Backdoor Attacks against NLP Models with Semantic-preserving Improvements. In Annual Computer Security Applications Conference, pages 554–569.
  13. 13.John Cook. 2020. Deconstructing Climate Science Denial. In Edward Elgar Research Handbook in Communicating Climate Change. Edward Elgar Publishing.
  14. 14.Limeng Cui and Dongwon Lee. 2020. CoAID: COVID-19 Healthcare Misinformation Dataset. arXiv preprint arXiv:2006.00885.
  15. 15.Giovanni Da San Martino, Stefano Cresci, Alberto Barrón-Cedeño, Seunghak Yu, Roberto Di Pietro, and Preslav Nakov. 2021. A survey on computational propaganda detection. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 4826–4832.
  16. 16.Giovanni Da San Martino, Seunghak Yu, Alberto Barrón-Cedeño, Rostislav Petrov, and Preslav Nakov. 2019. Fine-Grained Analysis of Propaganda in News Article. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5636–5646, Hong Kong, China. Association for Computational Linguistics.
  17. 17.Sajad Dadgar and Mehdi Ghatee. 2021. Checkovid: A COVID-19 misinformation detection system on Twitter using network and content mining perspectives. arXiv preprint arXiv:2107.09768.
  18. 18.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims. In Tackling Climate Change with Machine Learning workshop at NeurIPS.
  20. 20.Julian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, and Jordan Boyd-Graber. 2021. Fool me twice: Entailment from Wikipedia gamification. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 352–365, Online. Association for Computational Linguistics.
  21. 21.William Ferreira and Andreas Vlachos. 2016. Emergent: a novel data-set for stance classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1163–1168, San Diego, California. Association for Computational Linguistics.
  22. 22.Marc Fisher, John Woodrow Cox, and Peter Hermann. 2016. Pizzagate: From rumor, to hashtag, to gunfire in DC. Washington Post, 6:8410–8415.
  23. 23.FullFact. 2020. Framework for information incidents. Technical report, FullFact.
  24. 24.Michael Golebiewski and Danah Boyd. 2019. Data voids: Where missing data can easily be exploited. Technical report, Data & Society Research Institute.
  25. 25.Lucas Graves. 2018. Understanding the Promise and Limits of Automated Fact-Checking. In Reuters Institute for the Study of Journalism (Reuters Institute for the Study of Journalism Factsheets). Reuters Institute for the Study of Journalism.
  26. 26.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A Survey on Automated Fact-Checking. Transactions of the Association for Computational Linguistics, 10:178–206.
  27. 27.Ashim Gupta and Vivek Srikumar. 2021. X-Fact: A New Benchmark Dataset for Multilingual Fact Checking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 675–682, Online. Association for Computational Linguistics.
  28. 28.Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. DialFact: A Benchmark for Fact-Checking in Dialogue. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3785–3801, Dublin, Ireland. Association for Computational Linguistics.
  29. 29.Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein. 2018. The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1930–1940, New Orleans, Louisiana. Association for Computational Linguistics.
  30. 30.Andreas Hanselowski, Christian Stab, Claudia Schulz, Zile Li, and Iryna Gurevych. 2019. A Richly Annotated Corpus for Different Tasks in Automated Fact-Checking. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 493–503, Hong Kong, China. Association for Computational Linguistics.
  31. 31.Casper Hansen, Christian Hansen, and Lucas Chaves Lima. 2021. Automatic Fake News Detection: Are Models Learning to Reason? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 80–86, Online. Association for Computational Linguistics.
  32. 32.Momchil Hardalov, Arnav Arora, Preslav Nakov, and Isabelle Augenstein. 2022a. A Survey on Stance Detection for Mis- and Disinformation Identification. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1259–1277, Seattle, United States. Association for Computational Linguistics.
  33. 33.Momchil Hardalov, Anton Chernyavskiy, Ivan Koychev, Dmitry Ilvovsky, and Preslav Nakov. 2022b. CrowdChecked: Detecting Previously Fact-Checked Claims in Social Media. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, page (to appear), online. Association for Computational Linguistics.
  34. 34.Tamanna Hossain, Robert L. Logan IV, Arjuna Ugarte, Yoshitomo Matsubara, Sean Young, and Sameer Singh. 2020. COVIDLies: Detecting COVID-19 Misinformation on Social Media. In Proceedings of the 1st Workshop on NLP for COVID-19 (Part 2) at EMNLP 2020, Online. Association for Computational Linguistics.
  35. 35.Kung-Hsiang Huang, Kathleen McKeown, Preslav Nakov, Yejin Choi, and Heng Ji. 2022. Faking Fake News for Real Fake News Detection: Propaganda-loaded Training Data Generation. arXiv preprint arXiv:2203.05386.
  36. 36.Md Rafiqul Islam, Shaowu Liu, Xianzhi Wang, and Guandong Xu. 2020. Deep learning for misinformation detection on online social networks: a survey and new perspectives. Social Network Analysis and Mining, 10(1):1–20.
  37. 37.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. HoVer: A Dataset for Many-Hop Fact Extraction And Claim Verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441–3460, Online. Association for Computational Linguistics.
  38. 38.Zhijing Jin, Abhinav Lalwani, Tejas Vaidhya, Xiaoyu Shen, Yiwen Ding, Zhiheng Lyu, Mrinmaya Sachan, Rada Mihalcea, and Bernhard Schölkopf. 2022. Logical Fallacy Detection. arXiv preprint arXiv:2202.13758.
  39. 39.Kashif Khan, Ruizhe Wang, and Pascal Poupart. 2022. WatClaimCheck: A new dataset for claim entailment and inference. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1293–1304, Dublin, Ireland. Association for Computational Linguistics.
  40. 40.Neema Kotonya and Francesca Toni. 2020a. Explainable Automated Fact-Checking: A Survey. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5430–5443, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  41. 41.Neema Kotonya and Francesca Toni. 2020b. Explainable Automated Fact-Checking for Public Health Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7740–7754, Online. Association for Computational Linguistics.
  42. 42.Dilek Küçük and Fazli Can. 2020. Stance Detection: A Survey. ACM Computing Surveys (CSUR), 53(1).
  43. 43.Stephan Lewandowsky, John Cook, Ullrich Ecker, Dolores Albarracin, Michelle Amazeen, P. Kendou, D. Lombardi, E. Newman, G. Pennycook, E. Porter, D. Rand, D. Rapp, J. Reifler, J. Roozenbeek, P. Schmid, C. Seifert, G. Sinatra, B. Swire-Thompson, S. van der Linden, E. Vraga, T. Wood, and M. Zaragoza. 2020. The Debunking Handbook 2020. OpenBU.
  44. 44.Yichuan Li, Bohan Jiang, Kai Shu, and Huan Liu. 2020. Toward A Multilingual and Multimodal Data Repository for COVID-19 Disinformation. In 2020 IEEE International Conference on Big Data (Big Data), pages 4325–4330.
  45. 45.Alejandro Martín, Javier Huertas-Tato, Álvaro Huertas-García, Guillermo Villar-Rodríguez, and David Camacho. 2022. FacTeR-Check: Semi-automated fact-checking through semantic similarity and natural language inference. Knowledge-Based Systems, 251:109265.
  46. 46.Elena Musi and Andrea Rocci. 2022. Staying Up to Date with Fact and Reason Checking: An Argumentative Analysis of Outdated News. In Steve Oswald, Marcin Lewinski, Sara Greco, and Serena Villata, editors, The Pandemic of Argumentation, pages 311–330. Springer International Publishing, Cham.
  47. 47.Preslav Nakov, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barrón-Cedeño, Paolo Papotti, Shaden Shaar, and Giovanni Da San Martino. 2021. Automated Fact-Checking for Assisting Human Fact-Checkers. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 4551–4558. International Joint Conferences on Artificial Intelligence Organization.
  48. 48.Sakari Nieminen and Valtteri Sankari. 2021. Checking PolitiFact’s Fact-Checks. Journalism Studies, 22(3):358–378.
  49. 49.Ray Oshikawa, Jing Qian, and William Yang Wang. 2020. A Survey on Natural Language Processing for Fake News Detection. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 6086–6093, Marseille, France. European Language Resources Association.
  50. 50.Wojciech Ostrowski, Arnav Arora, Pepa Atanasova, and Isabelle Augenstein. 2021. Multi-Hop Fact Checking of Political Claims. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 3892–3898. International Joint Conferences on Artificial Intelligence Organization.
  51. 51.Jungsoo Park, Sewon Min, Jaewoo Kang, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. FaVIQ: FAct Verification from Information-seeking Questions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5154–5166, Dublin, Ireland. Association for Computational Linguistics.
  52. 52.Parth Patwa, Shivam Sharma, Srinivas Pykl, Vineeth Guptha, Gitanjali Kumari, Md Shad Akhtar, Asif Ekbal, Amitava Das, and Tanmoy Chakraborty. 2021. Fighting an infodemic: Covid-19 fake news dataset. In International Workshop on Combating Online Hostile Posts in Regional Languages during Emergency Situation, pages 21–29. Springer.
  53. 53.Dean Pomerleau and Delip Rao. 2017. Fake News Challenge. http://www.fakenewschallenge.org/.
  54. 54.Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. 2017. Truth of Varying Shades: Analyzing Language in Fake News and Political Fact-Checking. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2931–2937, Copenhagen, Denmark. Association for Computational Linguistics.
  55. 55.Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 Pandemic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2116–2129, Online. Association for Computational Linguistics.
  56. 56.Mourad Sarrouti, Asma Ben Abacha, Yassine Mrabet, and Dina Demner-Fushman. 2021. Evidence-based Fact-Checking of Health-related Claims. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3499–3512, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  57. 57.Aalok Sathe, Salar Ather, Tuan Manh Le, Nathan Perry, and Joonsuk Park. 2020. Automated fact-checking of claims from Wikipedia. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 6874–6882, Marseille, France. European Language Resources Association.
  58. 58.Tal Schuster, Darsh Shah, Yun Jie Serene Yeo, Daniel Roberto Filizzola Ortiz, Enrico Santus, and Regina Barzilay. 2019. Towards Debiasing Fact Verification Models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3419–3425, Hong Kong, China. Association for Computational Linguistics.
  59. 59.Shaden Shaar, Nikolay Babulkov, Giovanni Da San Martino, and Preslav Nakov. 2020. That is a Known Lie: Detecting Previously Fact-Checked Claims. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3607–3618, Online. Association for Computational Linguistics.
  60. 60.Tommy Shane and Pedro Noel. 2020. Data deficits: why we need to monitor the demand and supply of information in real time. Technical report, First Draft.
  61. 61.Craig Silverman. 2014. Verification handbook: An ultimate guideline on digital age sourcing for emergency coverage. European Journalism Centre.
  62. 62.Craig Silverman. 2016. Verification handbook: Additional Materials. European Journalism Centre.
  63. 63.Felix Simon, Philip N. Howard, and Rasmus Kleis Nielsen. 2020. Types, sources, and claims of COVID-19 misinformation. Technical report, Reuters Institute for the Study of Journalism.
  64. 64.James Thorne, Max Glockner, Gisela Vallejo, Andreas Vlachos, and Iryna Gurevych. 2021. Evidence-based Verification for Real World Information Needs. arXiv preprint arXiv:2104.00640.
  65. 65.James Thorne and Andreas Vlachos. 2018. Automated Fact Checking: Task Formulations, Methods and Future Directions. In Proceedings of the 27th International Conference on Computational Linguistics, pages 3346–3359, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  66. 66.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a Large-scale Dataset for Fact Extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  67. 67.Shaydanay Urbani. 2020. Verifying Online Information. Technical report, First Draft.
  68. 68.Sander van der Linden. 2022. Misinformation: susceptibility, spread, and interventions to immunize the public. Nature Medicine, 28(3):460–467.
  69. 69.Otávio Vinhas and Marco Bastos. 2022. Fact-Checking Misinformation: Eight Notes on Consensus Reality. Journalism Studies, 23(4):448–468.
  70. 70.Andreas Vlachos and Sebastian Riedel. 2014. Fact Checking: Task definition and dataset construction. In Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science, pages 18–22, Baltimore, MD, USA. Association for Computational Linguistics.
  71. 71.Nguyen Vo and Kyumin Lee. 2020. Where Are the Facts? Searching for Fact-checked Information to Alleviate the Spread of Fake News. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7717–7731, Online. Association for Computational Linguistics.
  72. 72.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online. Association for Computational Linguistics.
  73. 73.William Yang Wang. 2017. “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 422–426, Vancouver, Canada. Association for Computational Linguistics.
  74. 74.Claire Wardle et al. 2017. Fake news. it’s complicated. Technical report, First Draft.
  75. 75.Maxwell Weinzierl and Sanda Harabagiu. 2022. VaccineLies: A natural language resource for learning to recognize misinformation about the COVID-19 and HPV vaccines. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6967–6975, Marseille, France. European Language Resources Association.
  76. 76.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-Art Natural Language Processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  77. 77.John Zarocostas. 2020. How to fight an infodemic. The Lancet, 395(10225):676.
  78. 78.Yi Zhang, Zachary Ives, and Dan Roth. 2020. “Who said it, and Why?” Provenance for Natural Language Claims. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4416–4426, Online. Association for Computational Linguistics.
  79. 79.Yi Zhang, Zachary Ives, and Dan Roth. 2021. What is Your Article Based On? Inferring Fine-grained Provenance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5894–5903, Online. Association for Computational Linguistics.
  80. 80.Xinyi Zhou and Reza Zafarani. 2020. A survey of fake news: Fundamental theories, detection methods, and opportunities. ACM Computing Surveys (CSUR), 53(5):1–40.
  81. 81.Dimitrina Zlatkova, Preslav Nakov, and Ivan Koychev. 2019. Fact-Checking Meets Fauxtography: Verifying Claims About Images. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2099–2108, Hong Kong, China. Association for Computational Linguistics.
  82. 82.Arkaitz Zubiaga, Ahmet Aker, Kalina Bontcheva, Maria Liakata, and Rob Procter. 2018. Detection and Resolution of Rumours in Social Media: A Survey. ACM Computing Surveys (CSUR), 51(2).
  83. 83.Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, and Michal Lukasik. 2016. Stance Classification in Rumours as a Sequential Task Exploiting the Tree Structure of Social Media Conversations. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 2438–2448, Osaka, Japan. The COLING 2016 Organizing Committee.

Citation

MLA
Glockner, M., et al. “Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 5916–36, https://doi.org/10.18653/v1/2022.emnlp-main.397.
APA
Glockner, M., Hou, Y., & Gurevych, I. (2022). Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5916–5936. https://doi.org/10.18653/v1/2022.emnlp-main.397
Chicago
Glockner, M., Y. Hou, and I. Gurevych. 2022. “Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5916–36. https://doi.org/10.18653/v1/2022.emnlp-main.397.
Harvard
Glockner, M., Hou, Y. and Gurevych, I. (2022) “Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 5916–5936. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.397.
Vancouver
1. Glockner M, Hou Y, Gurevych I (2022) Missing Counter-Evidence Renders NLP Fact-Checking Unrealistic for Misinformation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 5916–5936

BibTeX

@inproceedings{glockner-etal-2022-missing,
    title = "Missing Counter-Evidence Renders {NLP} Fact-Checking Unrealistic for Misinformation",
    author = "Glockner, Max  and
      Hou, Yufang  and
      Gurevych, Iryna",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.397/",
    doi = "10.18653/v1/2022.emnlp-main.397",
    pages = "5916--5936"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/