Generating Literal and Implied Subquestions to Fact-check Complex Claims

Jifan ChenAniruddh SriramEunsol ChoiGreg Durrett

article2022EMNLP84 citations

Introduces ClaimDecomp, a benchmark for breaking down complex political claims into literal and implied yes-no subquestions, proving that question decomposition improves evidence retrieval and provides explainable step-by-step veracity judgments.

Listen

Modern political discourse frequently involves complex claims that blend partial truths with nuanced context, intent, and subtle misrepresentations. Automated fact-checking systems often struggle to handle these complexities, typically producing isolated judgments such as "half-true" that fail to explain why a statement is misleading or which specific aspects are inaccurate. Providing transparent, step-by-step reasoning is essential for building public trust, preventing misinformation, and enabling users to independently evaluate evidence.

The article demonstrates that complex political claims can be broken down into a concise set of binary yes-or-no subquestions that cover both explicit assertions and implicit contextual nuances. It evaluates whether language models can automatically generate these decompositions to improve evidence retrieval and accurately assess overall claim veracity.

To conduct this evaluation, the researchers created CLAIMDECOMP, a benchmark dataset of 1,200 complex claims from PolitiFact comprising 6,555 annotated subquestions. Trained annotators analyzed claims alongside professional fact-checkers' justifications to reverse-engineer the core binary questions required for verification. The researchers then fine-tuned sequence-to-sequence neural models to generate these questions directly from the claims and tested their downstream utility in natural language inference and evidence retrieval tasks.

The analysis produced three primary findings. First, the fine-tuned model successfully generated subquestions that covered 58% of reference questions, capturing 74% of literal subquestions but only 18% of implied subquestions. Second, human evaluations confirmed that these decomposed binary questions were significantly more helpful and relevant for determining veracity than broader existing approaches, scoring 3.60 versus 2.88 on a five-point scale. Third, applying natural language inference models to the decomposed subquestions notably improved evidence retrieval from verification documents, achieving an F1 score of up to 59.6 compared to 36.9 when using the full claim alone.

These findings indicate that claim decomposition provides a viable framework for explainable artificial intelligence in automated fact-checking. Instead of relying on opaque verdict classifications, verification pipelines can break down complex statements into transparent, verifiable components. This approach significantly enhances the retrieval of critical background context and allows final veracity ratings to be derived directly from the answers to individual subquestions.

For future development, the article recommends exploring iterative human-in-the-loop workflows where retrieved background context assists the model in generating higher-quality implicit subquestions. Additionally, developers should focus on creating question salience models to weight the relative importance of individual subquestions before making automated decisions.

Confidence in these findings is strong within the studied scope, but key limitations apply. Generating implied subquestions without access to external evidence remains challenging, and the dataset is limited to English-language, United States-focused political claims. Real-world deployment will require cautious validation, as retrieving evidence across the open web involves handling noisy or untrustworthy sources.

Cover for Generating Literal and Implied Subquestions to Fact-check Complex Claims

Abstract

Verifying political claims is a challenging task, as politicians can use various tactics to subtly misrepresent the facts for their agenda. Existing automatic fact-checking systems fall short here, and their predictions like “half-true” are not very useful in isolation, since it is unclear which parts of a claim are true or false. In this work, we focus on decomposing a complex claim into a comprehensive set of yes-no subquestions whose answers influence the veracity of the claim. We present ClaimDecomp, a dataset of decompositions for over 1000 claims. Given a claim and its verification paragraph written by fact-checkers, our trained annotators write subquestions covering both explicit propositions of the original claim and its implicit facets, such as additional political context that changes our view of the claim’s veracity. We study whether state-of-the-art pre-trained models can learn to generate such subquestions. Our experiments show that these models generate reasonable questions, but predicting implied subquestions based only on the claim (without consulting other evidence) remains challenging. Nevertheless, we show that predicted subquestions can help identify relevant evidence to fact-check the full claim and derive the veracity through their answers, suggesting that claim decomposition can be a useful piece of a fact-checking pipeline.¹

Table of Contents

  • 1 Introduction
  • 2 Motivation and Task
  • 3 Dataset Collection
  • Dataset statistics and inter-annotator agreement
  • 4 Automatic Claim Decomposition
  • 5 Analyzing Decomposition Annotations
  • 5.1 Subquestion Type Analysis
  • 5.2 Comparison to QABriefs
  • 5.3 Deriving the Veracity of Claims from Decomposed Questions
  • 6 Evidence Retrieval with Decomposition
  • 7 Related Work
  • 8 Conclusion
  • 9 Limitations
  • Domain limitations and lack of representation
  • Acknowledgments
  • References
  • A Question Annotation Workflow
  • A.1 Workflow
  • A.2 Annotation Interface
  • B Evidence Annotation Interface
  • C User Study Interface
  • D Inter-annotator Agreement
  • E Automatic Claim Decomposition Evaluation
  • F Training Details for Question Generation
  • G GPT-3 for Question Conversion
  • H Qualitative Analysis of Generated Questions
  • I More examples of QABriefs
  • J Datasheet for CLAIMDECOMP
  • J.1 Motivation for Datasheet Creation
  • J.2 Dataset Composition
  • J.3 Data Collection Process
  • J.4 Data Preprocessing
  • J.5 Dataset Distribution
  • J.6 Legal and Ethical Considerations

Knowls

  1. Knowl 1 — CLAIMDECOMP dataset of decomposed political claims

    data/table

    CLAIMDECOMP contains two independently annotated decompositions for 1,200 complex English political claims from PolitiFact, totaling 6,555 yes/no subquestions. The source collection began with 6,859 claims sampled from the top 50 PolitiFact pages for each of its six truth labels: pants on fire, false, barely true, half-true, mostly true, and true. The authors retained claims with at most three verbs and an available fact-checker justification paragraph, producing 1,494 candidate complex claims before annotation.

    For each claim, annotators received the claim, its speaker/date/venue context, and the professional fact-checker's justification. They wrote questions, assigned each answer as yes, no, or unknown, and marked whether the question was grounded in the claim or the justification and which text supported it. Most questions were grounded in the justification. The recommended split statistics are: train, 800 claims, 33.4 tokens per claim, 2.7 questions per single annotation, and answer proportions of 48.9% yes, 45.3% no, and 5.8% unknown; validation, 200 claims, 33.8 tokens, 2.7 questions, and 48.3%, 44.8%, and 6.9%; and test, 200 claims, 33.2 tokens, 2.7 questions, and 45.8%, 43.1%, and 11.1%. A 50-claim subset of validation, called Validation-sub, has 33.7 tokens per claim, 2.9 questions per annotation, and answer proportions of 45.2% yes, 47.8% no, and 7.0% unknown.

  2. Knowl 2 — Complex claim decomposition task

    definition

    Given a political claim cc, its context oo consisting of the speaker, date, and venue, and a target number NN of questions, the task is to generate a set of yes/no subquestions q={q1,…,qN}q=\{q_1,\ldots,q_N\}. A useful decomposition should be comprehensive enough that its answers support judging the veracity of the original claim, concise enough to avoid repeated or minor correlated questions, relevant enough that an answer changes the reader's belief about the claim, and fluent and clear. Subquestions are interpreted relative to the claim and its context and therefore are not required to be self-contained in isolation.

    For evaluation, the generator is required to produce the same number of questions as the reference annotation, which controls concision. Recall is then the fraction of reference questions covered by generated questions. Coverage is judged manually using contextual semantic equivalence rather than exact wording; for example, questions about whether in-person voting is safer than mail voting and whether mail voting has a greater fraud risk can count as equivalent when they target the same issue.

  3. Knowl 3 — Literal and implied reasoning in decompositions

    empirical result

    The annotations distinguish two post-hoc categories. A literal question can be posed from the surface information in the claim alone, whereas an implied question requires additional knowledge or reasoning to formulate. In a manual analysis of 285 questions from 100 development claims, each claim contained an average of 2.15 literal and 1.02 implied questions. Literal questions had ROUGE-1, ROUGE-2, and ROUGE-L precision of 0.56, 0.30, and 0.47 against their claims; implied questions had substantially lower values of 0.28, 0.09, and 0.22, showing that implied questions are less lexically recoverable from the claim.

    Among 50 analyzed implied questions, 38.8% required domain knowledge, 37.6% required broader context, 16.5% unpacked implicit meaning or speaker intent, and 7.1% checked statistical rigor such as the difference between a raw count and a per-capita rate. Thus, the dataset targets more than extraction of explicit propositions: roughly one question per claim asks about information that is not directly stated.

    The two independent annotations for 50 sampled claims had Fleiss' kappa of 0.52. On average, 18.4% of questions in one annotation set had no semantically matching question in the other set; the corresponding unmatched rates were 26.1% for the larger question set and 8.5% for the smaller set. The authors judged the questions generally concise, fluent, clear, and grammatical.

  4. Knowl 4 — T5-based subquestion generation models

    model/method

    The authors fine-tune T5-3B to generate a specified number of subquestions from a claim. QG-MULTIPLE treats the entire decomposition as one sequence: the input contains the desired count NN and claim cc, and the output concatenates q1,…,qNq_1,\ldots,q_N with separator tokens. It is trained for 20 epochs with batch size 8, maximum input and output lengths of 128 and 256 tokens, AdamW, initial learning rate 3×10−53\times10^{-5}, and DeepSpeed memory optimization; inference uses beam search with beam size 5.

    QG-NUCLEUS instead creates one training pair (c,qi)(c,q_i) for every annotated question and samples multiple questions independently at inference. It is trained for 10 epochs with batch size 16, maximum input and output lengths of 128 tokens, AdamW, initial learning rate 3×10−53\times10^{-5}, and DeepSpeed. Inference uses nucleus sampling with p=0.95p=0.95 and top-kk sampling with k=50k=50, followed by exact-string duplicate removal. Both systems generate the reference number of questions. Oracle variants append the fact-checker's justification paragraph to the claim; these are reported as QG-MULTIPLE-JUSTIFY and QG-NUCLEUS-JUSTIFY.

  5. Knowl 5 — Generation recovers explicit questions better than implied questions

    empirical result

    Manual evaluation on the 50-claim Validation-sub subset, containing 146 evaluated questions, measured recall of semantically equivalent reference questions. Generating all questions as one sequence was more effective than independently sampling them, and the gap was especially large for implied questions.

    Model R-all R-literal R-implied
    QG-MULTIPLE 0.58 0.74 0.18
    QG-NUCLEUS 0.43 0.59 0.11
    QG-MULTIPLE-JUSTIFY 0.81 0.95 0.50
    QG-NUCLEUS-JUSTIFY 0.52 0.72 0.18

    Without the justification, QG-MULTIPLE recalls 74% of literal questions but only 18% of implied questions. Supplying the justification in the oracle setting raises these values to 95% and 50%, respectively, demonstrating that the main difficulty is access to the contextual information needed for implied reasoning rather than only question fluency. QG-MULTIPLE is the strongest non-oracle generator and recovers 58% of all reference questions.

  6. Knowl 6 — Question answers provide a simple claim-veracity estimator

    equation

    For a claim with NN subquestions, let ai∈{0,1}a_i\in\{0,1\} be the answer to question qiq_i, where ai=1a_i=1 denotes yes and ai=0a_i=0 denotes no. The question-aggregation estimator assigns the claim a continuous veracity score equal to the fraction of yes answers:

    v^=1N∑i=1N1[ai=1].\hat v=\frac{1}{N}\sum_{i=1}^{N}\mathbb{1}[a_i=1].

    The score is mapped to PolitiFact's six labels by assigning pants on fire, false, barely true, half-true, mostly true, and true to the intervals [0,1/6)[0,1/6), [1/6,2/6)[1/6,2/6), [2/6,3/6)[2/6,3/6), [3/6,4/6)[3/6,4/6), [4/6,5/6)[4/6,5/6), and [5/6,1][5/6,1], respectively. On the development data, this unweighted baseline achieved macro-F1 0.30, micro-F1 0.29, and mean absolute error 1.05 across the six labels, compared with macro-F1 0.16, micro-F1 0.18, and mean absolute error 1.68 for random label assignment, and macro-F1 0.06, micro-F1 0.23, and mean absolute error 1.31 for always predicting the most frequent label. Removing questions judged unrelated to the claim improves the estimator to macro-F1 0.46, micro-F1 0.45, and mean absolute error 0.73. The result demonstrates promise but does not establish that all questions have equal importance or that their answers always correlate in the same direction with claim truth.

  7. Knowl 7 — NLI retrieval from question-derived hypotheses

    algorithm

    The evidence-retrieval proof of concept operates on a PolitiFact article divided into MM paragraphs p1,…,pMp_1,\ldots,p_M and a decomposition of NN questions q1,…,qNq_1,\ldots,q_N. GPT-3 converts each question into a declarative hypothesis hjh_j and also generates its negation; ten demonstrations make this conversion error rate less than 5% in the authors' check. For an NLI model, the entailment score for paragraph pip_i and hypothesis hjh_j is sij=P(Entailment∣pi,hj)s_{ij}=P(\mathrm{Entailment}\mid p_i,h_j).

    Input: Article paragraphs p1,...,pMp_1, ..., p_M, subquestions q1,...,qNq_1, ..., q_N, NLI model, retrieval count KK
    Output: Top-KK evidence paragraphs for support and refutation
    Convert every qjq_j into a statement hjh_j and a negated statement hˉj\bar h_j
    For each paragraph pip_i:
        Compute sijs_{ij} for every positive hypothesis hjh_j
        Set pi′=max⁡jsijp'_i = \max_j s_{ij}
    Select the KK paragraphs with largest pi′p'_i as support evidence
    Repeat the scoring and top-KK selection using the negated hypotheses hˉj\bar h_j
    Merge support and refutation evidence, remove duplicate paragraphs, and retain the final top-KK paragraphs by score
    Return the merged evidence set

    The procedure is applied with RoBERTa-based NLI models trained on MNLI, NQ-NLI, or DocNLI. The value of KK is set to the number of article paragraphs that human annotators marked as either supporting or refuting at least one subquestion. BM25 retrieval replaces the NLI score for the lexical baseline. NLI models trained on NQ-NLI and DocNLI provide entailment versus non-entailment rather than a dedicated contradiction class, so the authors use the positive and negated hypothesis passes rather than treating non-entailment as refutation.

  8. Knowl 8 — Decomposed questions improve evidence retrieval

    empirical result

    On 50 Validation-sub claims, human annotators labeled every paragraph in each full PolitiFact article as context, support, or refutation for each subquestion. Articles averaged 12.4 paragraphs per claim; aggregated across subquestions, 68.8% of paragraph labels were context, 12.0% support, and 19.2% refutation. The retrieval methods were evaluated by paragraph-level F1, comparing the original claim with either predicted or gold annotated decompositions.

    Model Predicted decomposition Gold decomposition Original claim
    MNLI 41.0 48.8 35.2
    NQ-NLI 38.8 34.5 40.9
    DocNLI 44.7 59.6 36.9
    BM25 36.2 47.5 39.2

    A random label-distribution baseline obtains F1 24.9 and the human comparison obtains 69.0. Decomposition improves retrieval over the original claim for every listed model in the gold-question condition and for all predicted-question systems except DocNLI and BM25. DocNLI with gold questions reaches 59.6 F1, substantially above the original-claim score of 36.9 and closer to human performance, indicating that implied and detailed subquestion information retrieves evidence missed by the claim alone.

  9. Knowl 9 — CLAIMDECOMP questions are more useful than QABriefs questions

    empirical result

    The authors compare their binary, claim-focused decompositions with QABriefs, a prior question-generation dataset whose questions are broader and can include background that is not directly needed to verify the claim. In a user study over 42 claims, five ratings were collected for each question set, yielding 210 ratings per method. Participants rated how accurately they could judge a claim after learning the answers to the questions on a 1-to-5 scale.

    CLAIMDECOMP questions received mean 3.60 with standard deviation 1.19, while QABriefs questions received mean 2.88 with standard deviation 1.20. The mean difference was 0.72, with a 95% confidence interval of 0.48 to 0.97 and a reported pp-value of at most 0.0001. The result supports the design choice of asking concise yes/no questions directly tied to the claim's veracity rather than collecting loosely related explanatory background.

  10. Knowl 10 — Scope and limitations of the decomposition pipeline

    limitation

    The study does not build a full real-world fact-checking system. Its retrieval experiment searches only within the known PolitiFact justification article, whereas web-scale retrieval can encounter inaccessible statistics, untrustworthy sources, and information that requires structured-table parsing. The question generator and retriever also depend on one another: high-quality implied questions need evidence context, while useful retrieval benefits from those questions, so a strict one-way pipeline is inadequate and the authors propose iterative or human-in-the-loop interaction.

    Automatic string-similarity metrics are unreliable for evaluating decompositions because small wording changes can alter question semantics while producing very high lexical or embedding similarity; the authors therefore rely primarily on manual semantic recall and downstream retrieval. The simple veracity aggregator further assumes equal question importance, can be distorted by irrelevant questions, and can fail when a question is phrased with an inverse polarity. Finally, the dataset contains only English, largely US-centric political claims from PolitiFact, so the findings should not be generalized without qualification to other languages, countries, or domains.

Coverage note — Fine-grained annotation-interface, worker-payment, licensing, datasheet, and automatic metric tables were omitted because they document collection and release procedures rather than adding load-bearing scientific content beyond the dataset, model, retrieval, and limitation results summarized here.

References

  1. 1.Naser Ahmadi, Joohyung Lee, Paolo Papotti, and Mohammed Saeed. 2019. Explainable fact checking with probabilistic answer set programming. In Conference on Truth and Trust Online.
  2. 2.Chris Alberti, Daniel Andor, Emily Pitler, Jacob Devlin, and Michael Collins. 2019. Synthetic QA corpora generation with roundtrip consistency. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6168–6173, Florence, Italy. Association for Computational Linguistics.
  3. 3.Tariq Alhindi, Savvas Petridis, and Smaranda Muresan. 2018. Where is your evidence: Improving fact-checking by justification modeling. In Proceedings of the First Workshop on Fact Extraction and VERification (FEVER), pages 85–90, Brussels, Belgium. Association for Computational Linguistics.
  4. 4.Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  5. 5.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. Generating fact checking explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7352–7364, Online. Association for Computational Linguistics.
  6. 6.Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. MultiFC: A real-world multi-domain dataset for evidence-based fact checking of claims. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4685–4697, Hong Kong, China. Association for Computational Linguistics.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  8. 8.Jifan Chen, Eunsol Choi, and Greg Durrett. 2021. Can NLI models verify QA systems’ predictions? In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3841–3854, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  9. 9.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2019. TabFact: A Large-scale Dataset for Table-based Fact Verification. In International Conference on Learning Representations.
  10. 10.Eunsol Choi, Jennimaria Palomaki, Matthew Lamm, Tom Kwiatkowski, Dipanjan Das, and Michael Collins. 2021. Decontextualization: Making sentences stand-alone. Transactions of the Association for Computational Linguistics, 9:447–461.
  11. 11.Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and psychological measurement, 20(1):37–46.
  12. 12.Xinya Du, Junru Shao, and Claire Cardie. 2017. Learning to ask: Neural question generation for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1342–1352, Vancouver, Canada. Association for Computational Linguistics.
  13. 13.Nan Duan, Duyu Tang, Peng Chen, and Ming Zhou. 2017. Question generation for question answering. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 866–874, Copenhagen, Denmark. Association for Computational Linguistics.
  14. 14.Esin Durmus, He He, and Mona Diab. 2020. FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5055–5070, Online. Association for Computational Linguistics.
  15. 15.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  16. 16.Angela Fan, Aleksandra Piktus, Fabio Petroni, Guillaume Wenzek, Marzieh Saeidi, Andreas Vlachos, Antoine Bordes, and Sebastian Riedel. 2020. Generating fact checking briefs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7147–7161, Online. Association for Computational Linguistics.
  17. 17.William Ferreira and Andreas Vlachos. 2016. Emergent: a novel data-set for stance classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1163–1168, San Diego, California. Association for Computational Linguistics.
  18. 18.Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378.
  19. 19.Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen, Franziska Roesner, Eunsol Choi, and Yejin Choi. 2022. Misinfo reaction frames: Reasoning about readers’ reactions to news headlines. In ACL.
  20. 20.Mohamed H Gad-Elrab, Daria Stepanova, Jacopo Urbani, and Gerhard Weikum. 2019. Exfakt: A framework for explaining facts over knowledge graphs and text. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining, pages 87–95.
  21. 21.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178–206.
  22. 22.Ashim Gupta and Vivek Srikumar. 2021. X-fact: A new benchmark dataset for multilingual fact checking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 675–682, Online. Association for Computational Linguistics.
  23. 23.Luheng He, Mike Lewis, and Luke Zettlemoyer. 2015. Question-answer driven semantic role labeling: Using natural language to annotate natural language. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 643–653, Lisbon, Portugal. Association for Computational Linguistics.
  24. 24.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In International Conference on Learning Representations.
  25. 25.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. In Proceedings of the International Conference on Learning Representations (ICLR).
  26. 26.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. HoVer: A dataset for many-hop fact extraction and claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441–3460, Online. Association for Computational Linguistics.
  27. 27.Ryo Kamoi, Tanya Goyal, and Greg Durrett. 2022. Shortcomings of Question Answering Based Factuality Frameworks for Error Localization. In arXiv.
  28. 28.Ayal Klein, Jonathan Mamou, Valentina Pyatkin, Daniela Stepanov, Hangfeng He, Dan Roth, Luke Zettlemoyer, and Ido Dagan. 2020. QANom: Question-answer driven SRL for nominalizations. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3069–3083, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  29. 29.Neema Kotonya and Francesca Toni. 2020. Explainable automated fact-checking for public health claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7740–7754, Online. Association for Computational Linguistics.
  30. 30.Harold W Kuhn. 1955. The Hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97.
  31. 31.Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10:163–177.
  32. 32.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. In arXiv.
  33. 33.Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In International Conference on Learning Representations.
  34. 34.Yi-Ju Lu and Cheng-Te Li. 2020. GCAN: Graph-aware co-attention networks for explainable fake news detection on social media. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 505–514, Online. Association for Computational Linguistics.
  35. 35.Bodhisattwa Prasad Majumder, Sudha Rao, Michel Galley, and Julian McAuley. 2021. Ask what’s missing and what’s useful: Improving clarification question generation using global knowledge. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4300–4312, Online. Association for Computational Linguistics.
  36. 36.Preslav Nakov, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barr’on-Cedeno, Paolo Papotti, Shaden Shaar, and Giovanni Da San Martino. 2021. Automated fact-checking for assisting human fact-checkers. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI).
  37. 37.Wojciech Ostrowski, Arnav Arora, Pepa Atanasova, and Isabelle Augenstein. 2021. Multi-hop fact checking of political claims. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI).
  38. 38.Kashyap Popat, Subhabrata Mukherjee, Jannik Strötgen, and Gerhard Weikum. 2017. Where the truth lies: Explaining the credibility of emerging claims on the web and social media. In Proceedings of the 26th International Conference on World Wide Web Companion, pages 1003–1012. International World Wide Web Conferences Steering Committee.
  39. 39.Kashyap Popat, Subhabrata Mukherjee, Andrew Yates, and Gerhard Weikum. 2018. DeClarE: Debunking fake news and false claims using evidence-aware deep learning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 22–32, Brussels, Belgium. Association for Computational Linguistics.
  40. 40.Valentina Pyatkin, Ayal Klein, Reut Tsarfaty, and Ido Dagan. 2020. QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2804–2819, Online. Association for Computational Linguistics.
  41. 41.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  42. 42.Sudha Rao and Hal Daumé III. 2018. Learning to ask good questions: Ranking clarification questions using neural expected value of perfect information. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2737–2746, Melbourne, Australia. Association for Computational Linguistics.
  43. 43.Hannah Rashkin, Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and Yejin Choi. 2017. Truth of varying shades: Analyzing language in fake news and political fact-checking. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2931–2937, Copenhagen, Denmark. Association for Computational Linguistics.
  44. 44.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506.
  45. 45.Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. COVID-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2116–2129, Online. Association for Computational Linguistics.
  46. 46.Mrinmaya Sachan and Eric Xing. 2018. Self-training for jointly learning to ask and answer questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 629–640, New Orleans, Louisiana. Association for Computational Linguistics.
  47. 47.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 624–643, Online. Association for Computational Linguistics.
  48. 48.Kai Shu, Limeng Cui, Suhang Wang, Dongwon Lee, and Huan Liu. 2019. defend: Explainable fake news detection. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pages 395–405.
  49. 49.Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Unsupervised commonsense question answering with self-talk. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4615–4629, Online. Association for Computational Linguistics.
  50. 50.Prakhar Singh, Anubrata Das, Junyi Jessy Li, and Matthew Lease. 2021. The Case for Claim Difficulty Assessment in Automatic Fact Checking. arXiv ePrint 2109.09689.
  51. 51.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.
  52. 52.Sebastian Tschiatschek, Adish Singla, Manuel Gomez-Rodriguez, Arpit Merchant, and Andreas Krause. 2018. Fake news detection in social networks via crowd signals. In The Web Conference, Alternate Track on Journalism, Misinformation, and Fact-checking.
  53. 53.Andreas Vlachos and Sebastian Riedel. 2014. Fact checking: Task definition and dataset construction. In Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science, pages 18–22, Baltimore, MD, USA. Association for Computational Linguistics.
  54. 54.Svitlana Volkova, Kyle Shaffer, Jin Yea Jang, and Nathan Hodas. 2017. Separating facts from fiction: Linguistic models to classify suspicious and trusted news posts on Twitter. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 647–653, Vancouver, Canada. Association for Computational Linguistics.
  55. 55.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online. Association for Computational Linguistics.
  56. 56.Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and answering questions to evaluate the factual consistency of summaries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5008–5020, Online. Association for Computational Linguistics.
  57. 57.William Yang Wang. 2017. “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 422–426, Vancouver, Canada. Association for Computational Linguistics.
  58. 58.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  59. 59.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  60. 60.Fan Yang, Shiva K Pentyala, Sina Mohseni, Mengnan Du, Hao Yuan, Rhema Linder, Eric D Ragan, Shuiwang Ji, and Xia Hu. 2019. Xfake: Explainable fake news detector with visualizations. In The World Wide Web Conference, pages 3600–3604.
  61. 61.Wenpeng Yin, Dragomir Radev, and Caiming Xiong. 2021. DocNLI: A large-scale dataset for document-level natural language inference. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4913–4922, Online. Association for Computational Linguistics.
  62. 62.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations (ICLR).

Citation

MLA
Chen, J., et al. “Generating Literal and Implied Subquestions to Fact-check Complex Claims”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3495–516, https://doi.org/10.18653/v1/2022.emnlp-main.229.
APA
Chen, J., Sriram, A., Choi, E., & Durrett, G. (2022). Generating Literal and Implied Subquestions to Fact-check Complex Claims. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3495–3516. https://doi.org/10.18653/v1/2022.emnlp-main.229
Chicago
Chen, J., A. Sriram, E. Choi, and G. Durrett. 2022. “Generating Literal and Implied Subquestions to Fact-check Complex Claims”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3495–3516. https://doi.org/10.18653/v1/2022.emnlp-main.229.
Harvard
Chen, J. et al. (2022) “Generating Literal and Implied Subquestions to Fact-check Complex Claims”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3495–3516. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.229.
Vancouver
1. Chen J, Sriram A, Choi E, Durrett G (2022) Generating Literal and Implied Subquestions to Fact-check Complex Claims. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3495–3516

BibTeX

@inproceedings{chen-etal-2022-generating,
    title = "Generating Literal and Implied Subquestions to Fact-check Complex Claims",
    author = "Chen, Jifan  and
      Sriram, Aniruddh  and
      Choi, Eunsol  and
      Durrett, Greg",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.229/",
    doi = "10.18653/v1/2022.emnlp-main.229",
    pages = "3495--3516"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/