Fact-Checking Complex Claims with Program-Guided Reasoning
Liangming PanXiaobao WuXinyuan LuAnh Tuan LuuWilliam Yang WangMin-Yen KanPreslav Nakov
Proposes a program-guided framework that decomposes complex claims into executable reasoning steps handled by specialized tools, delivering interpretable and data-efficient automated fact-checking that outperforms competitive baselines across multiple evidence settings.
Automated fact-checking faces significant challenges when evaluating complex real-world claims that require gathering multiple pieces of evidence and applying multi-step reasoning. Existing approaches often require extensive labeled training data or operate as opaque systems that fail to provide transparent explanations for their verdicts, limiting their reliability and adoption in operational environments.
The article demonstrates and evaluates a novel framework called Program-Guided Fact-Checking (PROGRAMFC). This framework aims to verify complex claims efficiently by decomposing them into structured, step-by-step reasoning programs that execute specialized sub-tasks using minimal training data.
The approach uses a code-pretrained large language model (Codex) prompted with only 20 demonstration examples to generate Python-like reasoning programs. These programs are executed by delegating sub-tasks—such as answering intermediate questions, verifying simple claims, and executing boolean logic—to specialized modules powered by models such as FLAN-T5. The authors evaluated the system across two benchmark datasets requiring multi-step reasoning (HOVER and FEVEROUS-S) under various operational settings, including gold-standard evidence, open-corpus web retrieval, and closed-book internal model knowledge, benchmarking performance against seven few-shot baselines.
The findings show that PROGRAMFC outperforms all seven baseline models across seven of the eight main evaluation settings. Its relative advantage increases substantially with reasoning depth, outperforming baselines by an average of 10.38% on two-hop, 11.37% on three-hop, and 14.77% on four-hop claims. On complex four-hop claims, the system's performance dropped by only 11.7%, compared to a 21.7% drop observed in the best-performing fine-tuned baseline. Furthermore, the framework enables smaller sub-task models (such as an 80-million parameter model) to achieve performance comparable to an end-to-end model with 11 billion parameters, representing a 137-fold efficiency gain in solver capacity. Iterative, program-guided retrieval also improved gold evidence recall by up to 37.1% over standard one-step retrieval methods, while aggregating multiple reasoning paths via majority voting further improved overall accuracy.
These results imply that structured neuro-symbolic decomposition is superior to end-to-end language modeling for complex verification tasks. By decoupling high-level planning from sub-task execution, organizations can deploy smaller, lower-cost sub-models without sacrificing performance. Additionally, the generated reasoning programs provide transparent, human-auditable logic that aids human fact-checkers, reduces the risk of unexplainable errors, and lowers data curation costs.
For practical implementation, organizations should adopt modular, program-guided architectures for automated verification workflows and implement multi-path aggregation to increase decision confidence. However, because the system requires multiple sequential model calls, decision-makers must account for an approximate 4-to-5-fold increase in inference latency compared to direct end-to-end models. Future work should focus on piloting these methods in production settings, optimizing computational efficiency, and expanding capabilities to handle implicit reasoning, fake news detection, and multi-modal claims.
Confidence in these findings is high for explicitly structured, multi-step textual claims evaluated against reference corpora. Readers should exercise caution regarding real-world claims requiring implicit commonsense reasoning, where the program generator showed higher error rates, as well as closed-book scenarios where large language models perform only slightly better than random guessing without external reference data.
- Paper: PAL: Program-aided Language Models, Luyu Gao et al. (2023). This paper establishes the foundational Program-Aided Language Models (PAL) paradigm of generating and executing intermediate code to solve reasoning tasks, which directly inspires program-guided claim verification.
- Paper: Generating Literal and Implied Subquestions to Fact-check Complex Claims, Jifan Chen et al. (2022). This work introduces the decomposition of complex fact-checking claims into structured subquestions, directly motivating the sub-task decomposition pipeline used in PROGRAMFC.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). This paper demonstrates how to disentangle natural language reasoning from computation via executable program generation, providing essential background for program-guided execution.
- Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). This work establishes the standard benchmark and pipeline for multi-step evidence extraction and claim verification that PROGRAMFC builds upon.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This seminal work introduces chain-of-thought prompting in large language models, which forms the underlying in-context reasoning mechanism structured into programs by PROGRAMFC.
- Paper: Language Models of Code are Few-Shot Commonsense Learners, Aman Madaan et al. (2022). This paper demonstrates that code-structured language models excel at structured reasoning and sub-task organization over standard text generation.
- Paper: Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations, Jaehun Jung et al. (2022). This study introduces recursive explanation generation and logical consistency solving to verify complex claims without extensive supervised data.
- Paper: SciAgent: Tool-augmented Language Models for Scientific Reasoning, Yubo Ma et al. (2024). This paper extends program-guided and tool-augmented reasoning frameworks to complex multi-step scientific problem solving with dedicated function libraries.
- Paper: Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting, Xinyan Guan et al. (2024). This work advances multi-step verification and hallucination reduction by autonomously querying and retrofitting intermediate reasoning steps against structured knowledge bases.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). This comprehensive survey contextualizes program-guided verification within the broader landscape of evaluation benchmarks, tool-based mitigation, and factuality frameworks in large language models.
- Paper: FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation, Sewon Min et al. (2023). This paper applies atomic fact decomposition and automated evidence verification to evaluate fine-grained factual precision in long-form language model generation.
- Paper: Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future, Zheng Chu et al. (2024). This survey provides a systematic categorization of intermediate reasoning paradigms, including programmatic and external tool-augmented reasoning structures.
