Fact-Checking Complex Claims with Program-Guided Reasoning

Liangming PanXiaobao WuXinyuan LuAnh Tuan LuuWilliam Yang WangMin-Yen KanPreslav Nakov

article2023ACL194 citations

Proposes a program-guided framework that decomposes complex claims into executable reasoning steps handled by specialized tools, delivering interpretable and data-efficient automated fact-checking that outperforms competitive baselines across multiple evidence settings.

Listen

Automated fact-checking faces significant challenges when evaluating complex real-world claims that require gathering multiple pieces of evidence and applying multi-step reasoning. Existing approaches often require extensive labeled training data or operate as opaque systems that fail to provide transparent explanations for their verdicts, limiting their reliability and adoption in operational environments.

The article demonstrates and evaluates a novel framework called Program-Guided Fact-Checking (PROGRAMFC). This framework aims to verify complex claims efficiently by decomposing them into structured, step-by-step reasoning programs that execute specialized sub-tasks using minimal training data.

The approach uses a code-pretrained large language model (Codex) prompted with only 20 demonstration examples to generate Python-like reasoning programs. These programs are executed by delegating sub-tasks—such as answering intermediate questions, verifying simple claims, and executing boolean logic—to specialized modules powered by models such as FLAN-T5. The authors evaluated the system across two benchmark datasets requiring multi-step reasoning (HOVER and FEVEROUS-S) under various operational settings, including gold-standard evidence, open-corpus web retrieval, and closed-book internal model knowledge, benchmarking performance against seven few-shot baselines.

The findings show that PROGRAMFC outperforms all seven baseline models across seven of the eight main evaluation settings. Its relative advantage increases substantially with reasoning depth, outperforming baselines by an average of 10.38% on two-hop, 11.37% on three-hop, and 14.77% on four-hop claims. On complex four-hop claims, the system's performance dropped by only 11.7%, compared to a 21.7% drop observed in the best-performing fine-tuned baseline. Furthermore, the framework enables smaller sub-task models (such as an 80-million parameter model) to achieve performance comparable to an end-to-end model with 11 billion parameters, representing a 137-fold efficiency gain in solver capacity. Iterative, program-guided retrieval also improved gold evidence recall by up to 37.1% over standard one-step retrieval methods, while aggregating multiple reasoning paths via majority voting further improved overall accuracy.

These results imply that structured neuro-symbolic decomposition is superior to end-to-end language modeling for complex verification tasks. By decoupling high-level planning from sub-task execution, organizations can deploy smaller, lower-cost sub-models without sacrificing performance. Additionally, the generated reasoning programs provide transparent, human-auditable logic that aids human fact-checkers, reduces the risk of unexplainable errors, and lowers data curation costs.

For practical implementation, organizations should adopt modular, program-guided architectures for automated verification workflows and implement multi-path aggregation to increase decision confidence. However, because the system requires multiple sequential model calls, decision-makers must account for an approximate 4-to-5-fold increase in inference latency compared to direct end-to-end models. Future work should focus on piloting these methods in production settings, optimizing computational efficiency, and expanding capabilities to handle implicit reasoning, fake news detection, and multi-modal claims.

Confidence in these findings is high for explicitly structured, multi-step textual claims evaluated against reference corpora. Readers should exercise caution regarding real-world claims requiring implicit commonsense reasoning, where the program generator showed higher error rates, as well as closed-book scenarios where large language models perform only slightly better than random guessing without external reference data.

Cover for Fact-Checking Complex Claims with Program-Guided Reasoning

Abstract

Fact-checking real-world claims often requires collecting multiple pieces of evidence and applying complex multi-step reasoning. In this paper, we present Program-Guided Fact-Checking (PROGRAMFC), a novel fact-checking model that decomposes complex claims into simpler sub-tasks that can be solved using a shared library of specialized functions. We first leverage the in-context learning ability of large language models to generate reasoning programs to guide the verification process. Afterward, we execute the program by delegating each sub-task to the corresponding sub-task handler. This process makes our model both explanatory and data-efficient, providing clear explanations of its reasoning process and requiring minimal training data. We evaluate PROGRAMFC on two challenging fact-checking datasets and show that it outperforms seven fact-checking baselines across different settings of evidence availability, with explicit output programs that benefit human debugging.1

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 PROGRAMFC
  • 3.1 Problem Formulation
  • 3.2 Program-Guided Reasoning
  • 3.3 Reasoning Program Generation
  • 3.4 Sub-Task Functions
  • 4 Experiments
  • 4.1 Main Results
  • 4.2 How Does the Reasoning Program Help?
  • 4.3 Interpretability of Reasoning Programs
  • 4.4 Closed-Book Fact-Checking
  • 5 Conclusion and Future Work
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Implementation Details about the Baselines
  • A.1 Pre-trained Models
  • A.2 FC/NLI Fine-Tuned Models
  • A.3 In-Context Learning Models
  • B Examples of Generated Reasoning Programs
  • C Error Analysis for Reasoning Programs
  • D Program Generation Prompts
  • E Prompts for Closed-Book Fact-Checking

Knowls

  1. Knowl 1 — Program-Guided Fact-Checking (PROGRAMFC) Architecture

    model/method

    Program-Guided Fact-Checking (PROGRAMFC) is a neuro-symbolic fact-checking framework designed to verify complex, multi-hop claims by decomposing them into structured, executable reasoning programs.

    Given an input claim CC, the framework operates in two distinct stages:

    1. Program Generation: A code-pretrained large language model (specifically Codex, code-davinci-002) acts as a planner P\mathcal{P}. Using in-context learning with a small set of K=20K = 20 demonstration exemplars, the planner maps CC into a domain-specific reasoning program P=[S1,S2,…,Sn]P = [S_1, S_2, \dots, S_n]. Each step SiS_i is a triple (fi,Ai,Vi)(f_i, A_i, V_i), where fi∈Ff_i \in \mathcal{F} denotes a specialized sub-task function from a predefined library F\mathcal{F}, AiA_i is an input argument (a question, a simple claim, or a Boolean expression), and ViV_i is a variable storing the return value fi(Ai)f_i(A_i). Arguments in later steps can dynamically reference previous variables {V1,…,Vi−1}\{V_1, \dots, V_{i-1}\} via string interpolation.

    2. Program Execution: An interpreter executes the sequence of steps S1,…,SnS_1, \dots, S_n by dispatching each argument AiA_i (with resolved variable bindings) to the appropriate sub-task handler fif_i. The final step SnS_n returns a Boolean variable Vn∈{TRUE,FALSE}V_n \in \{\text{TRUE}, \text{FALSE}\}, which serves as the predicted veracity label of the claim CC, while the executed program provides a step-by-step explanatory trace.

  2. Knowl 2 — Sub-Task Function Library and Prompting Specification

    model/method

    The execution engine of PROGRAMFC relies on a shared library F={QUESTION,VERIFY,PREDICT}\mathcal{F} = \{\text{QUESTION}, \text{VERIFY}, \text{PREDICT}\} of three specialized sub-task modules:

    1. QUESTION(Q)\text{QUESTION}(Q): A question-answering sub-task function that takes a natural language question QQ as input and returns an answer string AA. It is implemented using an instruction-tuned FLAN-T5 model. The input prompting format depends on the evidence setting:

      • Closed-book setting: Q: <Question>? The answer is:
      • Gold-evidence or open-book setting: <Evidence> Q: <Question>? The answer is:
    2. VERIFY(C′)\text{VERIFY}(C'): A claim verification sub-task function that takes a simple, single-statement claim C′C' and returns a Boolean label in {TRUE,FALSE}\{\text{TRUE}, \text{FALSE}\}. It is implemented using FLAN-T5 with the prompt format: <Evidence> Q: Is it true that <Claim>'? True or False? The answer is:

    3. PREDICT(L)\text{PREDICT}(L): A symbolic logical reasoning function that takes a logical expression LL composed of standard Boolean operators (AND,OR,NOT\text{AND}, \text{OR}, \text{NOT}) applied over return variables from previous steps (e.g., V1∧V2V_1 \land V_2) and evaluates the truth value, returning the final verdict for the claim.

  3. Knowl 3 — Program-Guided Fact-Checking with Reasoning Path Aggregation

    algorithm

    The overall inference procedure for PROGRAMFC generates multiple candidate reasoning programs via sampling and determines the final claim veracity through majority voting over the executed programs.

    Input: Claim CC, knowledge corpus or gold evidence documents K\mathcal{K}, candidate count NN, prompt demonstration set D\mathcal{D} (K=20K=20 examples)
    Output: Predicted veracity label Y∈{TRUE,FALSE}Y \in \{\text{TRUE}, \text{FALSE}\}, program execution traces
    Initialize empty list of predicted labels Y←[]\mathcal{Y} \leftarrow []
    for j←1j \leftarrow 1 to NN do
        Prompt LLM planner with D\mathcal{D} and CC using temperature T=0.7T=0.7 to sample reasoning program Pj=[S1,…,Sn]P_j = [S_1, \dots, S_n]
        Initialize variable environment V←{}\mathcal{V} \leftarrow \{\}
        for each step Si=(fi,Ai,Vi)S_i = (f_i, A_i, V_i) in PjP_j do
            Resolve all variable placeholders in argument AiA_i using V\mathcal{V}
            if setting is open-book then
                Retrieve relevant evidence EiE_i from K\mathcal{K} using BM25 with query AiA_i
            else if setting is gold-evidence then
                Ei←KE_i \leftarrow \mathcal{K}
            else
                Ei←∅E_i \leftarrow \emptyset
            end if
            Execute sub-task function: vi←fi(Ai,Ei)v_i \leftarrow f_i(A_i, E_i)
            Store result in variable environment: V[Vi]←vi\mathcal{V}[V_i] \leftarrow v_i
        end for
        Append final variable value Vn∈{TRUE,FALSE}V_n \in \{\text{TRUE}, \text{FALSE}\} to Y\mathcal{Y}
    end for
    Return Y←majority_vote(Y)Y \leftarrow \text{majority\_vote}(\mathcal{Y})
  4. Knowl 4 — Few-Shot Fact-Checking Performance on Multi-Hop Benchmarks

    data/table

    The table below presents Macro-F1 scores on the validation sets of HOVER (divided into 2-hop, 3-hop, and 4-hop subsets) and FEVEROUS-S (sentence-only subset of FEVEROUS with 2,962 claims) in few-shot settings (K=20K = 20 in-domain demonstration examples), evaluated under both Gold evidence and Open-book (BM25 top-10 retrieved paragraphs) settings.

    Model HOVER (2-hop) HOVER (3-hop) HOVER (4-hop) FEVEROUS-S
    Gold Open Gold Open Gold Open Gold Open
    BERT-FC 53.40 50.68 50.90 49.86 50.86 48.57 74.71 51.67
    LisT5 56.15 52.56 53.76 51.89 51.67 50.46 77.88 54.15
    RoBERTa-NLI 74.62 63.62 62.23 53.99 57.98 52.40 88.28 57.80
    DeBERTaV3-NLI 77.22 68.72 65.98 60.76 60.49 56.00 91.98 58.81
    MULTIVERS 68.86 60.17 59.87 52.55 55.67 51.86 86.03 56.61
    Codex 70.63 65.07 66.46 56.63 63.49 57.27 89.77 62.58
    FLAN-T5 73.69 69.02 65.66 60.23 58.08 55.42 90.81 63.73
    ProgramFC (N=1N=1) 74.10 69.36 66.13 60.63 65.69 59.16 91.77 67.80
    ProgramFC (N=5N=5) 75.65 70.30 68.48 63.43 66.75 57.74 92.69 68.06

    PROGRAMFC achieves the best performance across 7 of the 8 evaluations. As claim complexity increases on HOVER, the performance drop for DeBERTaV3-NLI between 2-hop and 4-hop is 21.7% (77.22 to 60.49 in Gold), whereas PROGRAMFC experiences only an 11.7% drop (75.65 to 66.75). Decomposing claims with PROGRAMFC (N=1N=1) improves Macro-F1 over direct FLAN-T5 prediction by 6.0% (Gold) and 4.5% (Open) on average, reaching a 14.9% gain on 4-hop Gold evidence. Multi-program aggregation (N=5N=5) improves performance over single-program execution (N=1N=1) by an average of 1.5%.

  5. Knowl 5 — Effect of Sub-Task Model Size on Program-Guided Fact-Checking

    empirical result

    When comparing PROGRAMFC against direct end-to-end claim verification using FLAN-T5 across five model sizes—small (80M parameters), base (250M), large (780M), XL (3B), and XXL (11B)—under the gold evidence setting on HOVER:

    • The end-to-end FLAN-T5 model exhibits sharp performance degradation as model capacity decreases, especially on complex multi-hop claims (e.g., dropping on 4-hop claims from 63.39 F1 at 11B down to 48.59 F1 at 80M).
    • PROGRAMFC exhibits strong resilience to smaller sub-task solvers because the high-level reasoning program decomposes the problem into simple, single-step tasks. On HOVER 4-hop claims, PROGRAMFC utilizing FLAN-T5-small (80M) achieves an F1 score of 62.46, comparable to the end-to-end FLAN-T5-XXL (11B) model (63.39 F1) despite having 137×137\times fewer parameters in the execution engine.
  6. Knowl 6 — Iterative Evidence Retrieval via Reasoning Programs

    empirical result

    In the open-book fact-checking setting, PROGRAMFC performs iterative, step-by-step BM25 retrieval by dispatching intermediate questions and sub-claims generated during program execution as search queries, rather than performing a single-step retrieval on the raw input claim.

    Evaluating retrieval recall of gold paragraphs among the top-10 retrieved paragraphs (Recall@10):

    • HOVER 2-hop: Iterative retrieval achieves 77.13% vs. 73.18% for one-step retrieval.
    • HOVER 3-hop: Iterative retrieval achieves 59.17% vs. 51.33% for one-step retrieval.
    • HOVER 4-hop: Iterative retrieval achieves 49.93% vs. 36.43% for one-step retrieval (a relative gain of 37.1%).
    • FEVEROUS-S: Iterative retrieval achieves 85.65% vs. 76.25% for one-step retrieval.

    The iterative approach excels on higher-hop claims because crucial entities and connecting facts (e.g., intermediate bridging entities) are not present in the original claim text and only emerge during step-by-step program execution.

  7. Knowl 7 — Closed-Book Multi-Hop Fact-Checking Performance

    data/table

    In the closed-book setting where models have no access to external evidence corpus (K=∅K = \emptyset) and must rely entirely on internal parametric knowledge, Macro-F1 scores across HOVER and FEVEROUS-S are evaluated as follows:

    Model HOVER FEVEROUS
    2-hop 3-hop 4-hop
    InstructGPT (text-davinci-002)
    – Direct Prompting 56.51 51.75 49.68 60.13
    – Zero-Shot CoT (ZS-CoT) 50.30 52.30 51.58 54.78
    – Few-Shot CoT 57.20 53.66 51.83 61.05
    – Self-Ask 51.54 51.47 52.45 56.82
    Codex (code-davinci-002) 55.57 53.42 45.59 57.85
    FLAN-T5 (XXL 11B) 48.27 52.11 51.13 55.16
    ProgramFC 54.27 54.18 52.88 59.66

    All evaluated models achieve scores only slightly above random guessing (Macro-F1 around 50--61%), demonstrating that relying purely on parametric memory in LLMs is insufficient for verifying complex multi-hop claims. Step-by-step reasoning via few-shot Chain-of-Thought (CoT) improves performance over direct prompting by an average of 2.7 points. PROGRAMFC outperforms all baselines on HOVER 3-hop (54.18) and 4-hop (52.88) claims, as structured symbolic programs provide greater stability for long reasoning chains compared to free-form natural language explanations.

  8. Knowl 8 — Error Taxonomy and Structural Complexity of Reasoning Programs

    data/table

    Human error analysis was conducted on 300 randomly sampled claims incorrectly predicted by PROGRAMFC (100 each from HOVER 2-hop, 3-hop, and 4-hop datasets). Errors were classified into three main categories: Syntactic errors (grammar/parsing failures), Semantic errors (Token/argument errors, Structure errors, and Subtask call errors), and Incorrect execution (correct program, but sub-task handler returned an incorrect result).

    Error Category Proportion (%)
    2-hop 3-hop 4-hop
    Syntax error 0% 0% 0%
    Semantic error 29% 38% 77%
    – Token (incorrect/missing argument) 8% 20% 18%
    – Structure (incorrect program logic/flow) 19% 13% 57%
    – Subtask (wrong function call type) 2% 5% 2%
    Incorrect execution (handler failed on valid program) 71% 62% 23%

    Codex achieves 0% syntax errors under 20-shot prompting. For simpler 2-hop claims, the majority (71%) of model failures stem from sub-task execution errors. For complex 4-hop claims, the error distribution shifts dramatically toward semantic structural errors (57%), reflecting the increased difficulty of generating multi-step reasoning plans for deep compositionality.

  9. Knowl 9 — Limitations of Program-Guided Fact-Checking

    limitation

    PROGRAMFC has two primary limitations:

    1. Implicit Multi-Step Reasoning: The model relies on claims where the reasoning decomposition is syntactically explicit from the surface form. For claims requiring implicit reasoning (e.g., verifying "Aristotle couldn't have used a laptop", which requires unstated intermediate world knowledge regarding Aristotle's lifespan and the invention date of laptops), the LLM planner struggles to generate valid reasoning programs without prior commonsense grounding.

    2. Computational Overhead: Generating reasoning programs through LLMs and subsequently invoking multiple sequential sub-task model handlers results in an actual inference latency approximately 4×4\times to 5×5\times higher than direct end-to-end fact-checking models (such as single-pass FLAN-T5).

Coverage note — Specific full-length prompt listings from Appendices D and E were omitted as concrete formatting details rather than distinct conceptual contributions.

References

  1. 1.Naser Ahmadi, Joohyung Lee, Paolo Papotti, and Mohammed Saeed. 2019. Explainable fact checking with probabilistic answer set programming. In Proceedings of the Truth and Trust Online Conference (TTO), London, UK.
  2. 2.Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks, Online.
  3. 3.Rami Aly and Andreas Vlachos. 2022. Natural logicguided autoregressive multi-hop document retrieval for fact verification. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6123–6135, Abu Dhabi, United Arab Emirates.
  4. 4.Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020. Generating fact checking explanations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7352–7364, Online.
  5. 5.Isabelle Augenstein, Christina Lioma, Dongsheng Wang, Lucas Chaves Lima, Casper Hansen, Christian Hansen, and Jakob Grue Simonsen. 2019. MultiFC: A real-world multi-domain dataset for evidencebased fact checking of claims. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4685–4697, Hong Kong, China.
  6. 6.Giorgio Barnabò, Federico Siciliano, Carlos Castillo, Stefano Leonardi, Preslav Nakov, Giovanni Da San Martino, and Fabrizio Silvestri. 2022. FbMultiLingMisinfo: Challenging large-scale multilingual benchmark for misinformation detection. In Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8, Padova, Italy.
  7. 7.Giorgio Barnabò, Federico Siciliano, Carlos Castillo, Stefano Leonardi, Preslav Nakov, Giovanni Da San Martino, and Fabrizio Silvestri. 2023. Deep active learning for misinformation detection using geometric deep learning. Online Social Networks and Media, 33:100244.
  8. 8.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. ArXiv preprint, abs/2004.05150.
  9. 9.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 632–642, Lisbon, Portugal.
  10. 10.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), Online.
  11. 11.Jifan Chen, Aniruddh Sriram, Eunsol Choi, and Greg Durrett. 2022a. Generating literal and implied subquestions to fact-check complex claims. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3495–3516, Abu Dhabi, United Arab Emirates.
  12. 12.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. ArXiv preprint, abs/2107.03374.
  13. 13.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2022b. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. CoRR, abs/2211.12588.
  14. 14.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. 2022. Binding language models in symbolic languages. CoRR, abs/2210.02875.
  15. 15.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. CoRR, abs/2210.11416.
  16. 16.Limeng Cui, Kai Shu, Suhang Wang, Dongwon Lee, and Huan Liu. 2019. dEFEND: A system for explainable fake news detection. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM), pages 2961–2964, Beijing, China.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 4171–4186, Minneapolis, Minnesota, USA.
  18. 18.Mohamed H. Gad-Elrab, Daria Stepanova, Jacopo Urbani, and Gerhard Weikum. 2019. Exfakt: A framework for explaining facts over knowledge graphs and text. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (WSDM), pages 87–95, Melbourne, Australia.
  19. 19.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022. PAL: program-aided language models. CoRR, abs/2211.10435.
  20. 20.Max Glockner, Yufang Hou, and Iryna Gurevych. 2022. Missing counter-evidence renders NLP fact-checking unrealistic for misinformation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5916–5936, Abu Dhabi, United Arab Emirates.
  21. 21.Zhijiang Guo, Michael Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178–206.
  22. 22.Ashim Gupta and Vivek Srikumar. 2021. X-Fact: A new benchmark dataset for multilingual fact checking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), pages 675–682, Online.
  23. 23.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradientdisentangled embedding sharing. ArXiv preprint, abs/2111.09543.
  24. 24.Kelvin Jiang, Ronak Pradeep, and Jimmy Lin. 2021. Exploring listwise evidence reasoning with T5 for fact verification. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), pages 402–410, Online.
  25. 25.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. HoVer: A dataset for many-hop fact extraction and claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441–3460, Online.
  26. 26.Shailza Jolly, Pepa Atanasova, and Isabelle Augenstein. 2022. Generating fluent fact checking explanations with unsupervised post-editing. Information, 13(10):500.
  27. 27.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. CoRR, abs/2205.11916.
  28. 28.Neema Kotonya and Francesca Toni. 2020. Explainable automated fact-checking for public health claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7740–7754, Online.
  29. 29.Amrith Krishna, Sebastian Riedel, and Andreas Vlachos. 2022. ProoFVer: Natural logic theorem proving for fact verification. Transactions of the Association for Computational Linguistics (TACL), 10:1013–1030.
  30. 30.Nayeon Lee, Yejin Bang, Andrea Madotto, and Pascale Fung. 2021. Towards few-shot fact-checking via perplexity. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 1971–1981, Online.
  31. 31.Nayeon Lee, Belinda Z. Li, Sinong Wang, Wen-tau Yih, Hao Ma, and Madian Khabsa. 2020. Language models as fact checkers? In Proceedings of the Third Workshop on Fact Extraction and VERification (FEVER), pages 36–41, Online.
  32. 32.Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), pages 2356–2362, Online.
  33. 33.Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022. WANLI: Worker and AI collaboration for natural language inference dataset creation. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 6826–6847, Abu Dhabi, United Arab Emirates.
  34. 34.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. ArXiv preprint, abs/1907.11692.
  35. 35.Zhenghao Liu, Chenyan Xiong, Maosong Sun, and Zhiyuan Liu. 2020. Fine-grained fact verification with kernel graph attention network. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7342–7351, Online.
  36. 36.Yi-Ju Lu and Cheng-Te Li. 2020. GCAN: Graph-aware co-attention networks for explainable fake news detection on social media. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 505–514, Online.
  37. 37.Gregoire Mialon, Roberto Dessı, Maria Lomeli, Christoforos Nalmpantis, Ramakanth Pasunuru, Roberta Raileanu, Baptiste Roziere, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. 2023. Augmented language models: a survey. CoRR, abs/2302.07842.
  38. 38.Preslav Nakov, Alberto Barron-CedeȀno, Giovanni Da San Martino, Firoj Alam, Julia Maria Stru, Thomas Mandl, Ruben Mıguez, Tommaso Caselli, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Gautam Kishore Shahi, Hamdy Mubarak, Alex Nikolov, Nikolay Babulkov, Yavuz Selim Kartal, and Javier Beltran. 2022. The CLEF-2022 CheckThat! lab on fighting the COVID19 infodemic and fake news detection. In Proceedings of the 44th European Conference on IR Research: Advances in Information Retrieval (ECIR), pages 416–428, Berlin, Heidelberg.
  39. 39.Preslav Nakov, David Corney, Maram Hasanain, Firoj Alam, Tamer Elsayed, Alberto Barron-CedeȀno, Paolo Papotti, Shaden Shaar, and Giovanni Da San Martino. 2021a. Automated fact-checking for assisting human fact-checkers. In Proceedings of the Joint Conference on Artificial Intelligence (IJCAI), pages 4551–4558, Online.
  40. 40.Preslav Nakov, Giovanni Da San Martino, Tamer Elsayed, Alberto Barron-CedeȀno, Ruben Mıguez, Shaden Shaar, Firoj Alam, Fatima Haouari, Maram Hasanain, Nikolay Babulkov, Alex Nikolov, Gautam Kishore Shahi, Julia Maria Stru, and Thomas Mandl. 2021b. The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously factchecked claims, and fake news. In Proceedings of the 43rd European Conference on Information Retrieval (ECIR), pages 639–649, Lucca, Italy.
  41. 41.Van-Hoang Nguyen, Kazunari Sugiyama, Preslav Nakov, and Min-Yen Kan. 2020. FANG: leveraging social context for fake news detection using graph representation. In Proceedings of the 29th ACM International Conference on Information and Knowledge Management (CIKM), pages 1165–1174.
  42. 42.Yixin Nie, Haonan Chen, and Mohit Bansal. 2019. Combining fact extraction and verification with neural semantic matching networks. In Proceedings of the 33rd AAAI Conference on Artificial Intelligence (AAAI), pages 6859–6866, Honolulu, Hawaii, USA.
  43. 43.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4885–4901, Online.
  44. 44.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. CoRR, abs/2203.02155.
  45. 45.Liangming Pan, Wenhu Chen, Wenhan Xiong, MinYen Kan, and William Yang Wang. 2021. Zero-shot fact verification by claim generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), pages 476–483, Online.
  46. 46.Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, and Samuel R. Bowman. 2021. Does putting a linguist in the loop improve NLU data collection? In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4886–4901, Punta Cana, Dominican Republic.
  47. 47.Kashyap Popat, Subhabrata Mukherjee, Jannik Strotgen, and Gerhard Weikum. 2017. Where the truth lies: Explaining the credibility of emerging claims on the web and social media. In Proceedngs of the International World Wide Web Conference (WWW), pages 1003–1012.
  48. 48.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. CoRR, abs/2210.03350.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  50. 50.Stephen E. Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4):333–389.
  51. 51.Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. COVID-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), pages 2116–2129, Online.
  52. 52.Aalok Sathe, Salar Ather, Tuan Manh Le, Nathan Perry, and Joonsuk Park. 2020. Automated fact-checking of claims from Wikipedia. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC), pages 6874–6882, Marseille, France.
  53. 53.Timo Schick, Jane Dwivedi-Yu, Roberto Dessı, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. CoRR, abs/2302.04761.
  54. 54.Tal Schuster, Adam Fisch, and Regina Barzilay. 2021. Get your vitamin C! robust fact verification with contrastive evidence. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 624–643, Online.
  55. 55.Amir Soleimani, Christof Monz, and Marcel Worring. 2020. BERT for evidence retrieval and claim verification. In Advances in Information Retrieval (ECIR), volume 12036, pages 359–366.
  56. 56.James Thorne and Andreas Vlachos. 2018. Automated fact checking: Task formulations, methods and future directions. In Proceedings of the 27th International Conference on Computational Linguistics (COLING), pages 3346–3359, Santa Fe, New Mexico, USA.
  57. 57.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 809–819, New Orleans, Louisiana.
  58. 58.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems (NeurIPS), pages 5998–6008, Long Beach, California, USA.
  59. 59.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550, Online.
  60. 60.David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Iz Beltagy, Lucy Lu Wang, and Hannaneh Hajishirzi. 2022a. SciFact-open: Towards open-domain scientific claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 4719–4734, Abu Dhabi, United Arab Emirates.
  61. 61.David Wadden, Kyle Lo, Lucy Wang, Arman Cohan, Iz Beltagy, and Hannaneh Hajishirzi. 2022b. MultiVerS: Improving scientific claim verification with weak supervision and full-document context. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 61–76, Seattle, Washington, USA.
  62. 62.William Yang Wang. 2017. ‘‘Liar, liar pants on fire’’: A new benchmark dataset for fake news detection. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL), pages 422–426, Vancouver, Canada.
  63. 63.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, and Denny Zhou. 2022. Selfconsistency improves chain of thought reasoning in language models. CoRR, abs/2203.11171.
  64. 64.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv preprint, abs/2201.11903.
  65. 65.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACLHLT), pages 1112–1122, New Orleans, Louisiana, USA.
  66. 66.Dustin Wright, David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Isabelle Augenstein, and Lucy Wang. 2022. Generating scientific claims for zero-shot scientific fact checking. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2448–2460, Dublin, Ireland.
  67. 67.Fan Yang, Shiva K. Pentyala, Sina Mohseni, Mengnan Du, Hao Yuan, Rhema Linder, Eric D. Ragan, Shuiwang Ji, and Xia (Ben) Hu. 2019. XFake: Explainable fake news detector with visualizations. In Proceedings of the The World Wide Web Conference (WWW), pages 3600–3604, San Francisco, California, USA.
  68. 68.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2369–2380, Brussels, Belgium.
  69. 69.Wanjun Zhong, Jingjing Xu, Duyu Tang, Zenan Xu, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin. 2020. Reasoning over semantic-level graph for fact checking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 6170–6180, Online.
  70. 70.Jie Zhou, Xu Han, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun. 2019. GEAR: Graph-based evidence aggregating and reasoning for fact verification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 892–901, Florence, Italy.

Citation

MLA
Pan, L., et al. “Fact-Checking Complex Claims with Program-Guided Reasoning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 6981–7004, https://doi.org/10.18653/v1/2023.acl-long.386.
APA
Pan, L., Wu, X., Lu, X., Tuan, L. A., Wang, W. Y., Kan, M.-Y., & Nakov, P. (2023). Fact-Checking Complex Claims with Program-Guided Reasoning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6981–7004. https://doi.org/10.18653/v1/2023.acl-long.386
Chicago
Pan, L., X. Wu, X. Lu, et al. 2023. “Fact-Checking Complex Claims with Program-Guided Reasoning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6981–7004. https://doi.org/10.18653/v1/2023.acl-long.386.
Harvard
Pan, L. et al. (2023) “Fact-Checking Complex Claims with Program-Guided Reasoning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6981–7004. Available at: https://doi.org/10.18653/v1/2023.acl-long.386.
Vancouver
1. Pan L, Wu X, Lu X, Tuan LA, Wang WY, Kan M-Y, Nakov P (2023) Fact-Checking Complex Claims with Program-Guided Reasoning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6981–7004

BibTeX

@inproceedings{pan-etal-2023-fact,
    title = "Fact-Checking Complex Claims with Program-Guided Reasoning",
    author = "Pan, Liangming  and
      Wu, Xiaobao  and
      Lu, Xinyuan  and
      Luu, Anh Tuan  and
      Wang, William Yang  and
      Kan, Min-Yen  and
      Nakov, Preslav",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.386/",
    doi = "10.18653/v1/2023.acl-long.386",
    pages = "6981--7004"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/