Generating Literal and Implied Subquestions to Fact-check Complex Claims
Jifan ChenAniruddh SriramEunsol ChoiGreg Durrett
Introduces ClaimDecomp, a benchmark for breaking down complex political claims into literal and implied yes-no subquestions, proving that question decomposition improves evidence retrieval and provides explainable step-by-step veracity judgments.
Modern political discourse frequently involves complex claims that blend partial truths with nuanced context, intent, and subtle misrepresentations. Automated fact-checking systems often struggle to handle these complexities, typically producing isolated judgments such as "half-true" that fail to explain why a statement is misleading or which specific aspects are inaccurate. Providing transparent, step-by-step reasoning is essential for building public trust, preventing misinformation, and enabling users to independently evaluate evidence.
The article demonstrates that complex political claims can be broken down into a concise set of binary yes-or-no subquestions that cover both explicit assertions and implicit contextual nuances. It evaluates whether language models can automatically generate these decompositions to improve evidence retrieval and accurately assess overall claim veracity.
To conduct this evaluation, the researchers created CLAIMDECOMP, a benchmark dataset of 1,200 complex claims from PolitiFact comprising 6,555 annotated subquestions. Trained annotators analyzed claims alongside professional fact-checkers' justifications to reverse-engineer the core binary questions required for verification. The researchers then fine-tuned sequence-to-sequence neural models to generate these questions directly from the claims and tested their downstream utility in natural language inference and evidence retrieval tasks.
The analysis produced three primary findings. First, the fine-tuned model successfully generated subquestions that covered 58% of reference questions, capturing 74% of literal subquestions but only 18% of implied subquestions. Second, human evaluations confirmed that these decomposed binary questions were significantly more helpful and relevant for determining veracity than broader existing approaches, scoring 3.60 versus 2.88 on a five-point scale. Third, applying natural language inference models to the decomposed subquestions notably improved evidence retrieval from verification documents, achieving an F1 score of up to 59.6 compared to 36.9 when using the full claim alone.
These findings indicate that claim decomposition provides a viable framework for explainable artificial intelligence in automated fact-checking. Instead of relying on opaque verdict classifications, verification pipelines can break down complex statements into transparent, verifiable components. This approach significantly enhances the retrieval of critical background context and allows final veracity ratings to be derived directly from the answers to individual subquestions.
For future development, the article recommends exploring iterative human-in-the-loop workflows where retrieved background context assists the model in generating higher-quality implicit subquestions. Additionally, developers should focus on creating question salience models to weight the relative importance of individual subquestions before making automated decisions.
Confidence in these findings is strong within the studied scope, but key limitations apply. Generating implied subquestions without access to external evidence remains challenging, and the dataset is limited to English-language, United States-focused political claims. Real-world deployment will require cautious validation, as retrieving evidence across the open web involves handling noisy or untrustworthy sources.
- Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). This foundational benchmark formalized the pipeline of evidence retrieval and textual entailment for automated claim verification, which the source extends using fine-grained subquestion decomposition.
- Paper: “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection, William Yang Wang (2017). Understanding this foundational benchmark of real-world PolitiFact statements clarifies the nature of nuanced political claims and sets the stage for decomposing such statements into verifiable components.
- Paper: BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions, Christopher Clark et al. (2019). Examining the complex inference requirements of naturally occurring binary questions helps explain why decomposing claims into yes/no subquestions is an effective verification mechanism.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). Building on the idea of breaking down claims to evaluate veracity, this work operationalizes atomic fact decomposition to systematically assess precision in long-form generation.
- Paper: LM vs LM: Detecting Factual Errors via Cross Examination, Roi Cohen et al. (2023). This work continues the paradigm of verifiable multi-step questioning by using cross-examination dialogues to uncover factual errors without external references.
- Paper: Enabling Large Language Models to Generate Text with Citations, Tianyu Gao et al. (2023). This paper extends verifiable generation pipelines by developing automated evaluation benchmarks for retrieving evidence and generating text with precise citations.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). This study critically investigates the faithfulness of step-by-step model rationales, offering important cautionary insights for pipelines relying on generated sub-reasoning for explainability.
