FOLIO: Natural Language Reasoning with First-Order Logic

Simeng HanHailey SchoelkopfYilun ZhaoZhenting QiMartin RiddellWenfei ZhouJames CoadyDavid PengYujie QiaoLuke Benson

article2024EMNLP234 citations

Introduces FOLIO, an expert-annotated benchmark with parallel first-order logic formulas and automated validity verification, exposing major deductive reasoning limitations in leading large language models like GPT-4.

Listen

Artificial intelligence systems increasingly handle complex tasks, yet existing evaluation benchmarks fail to adequately measure their logical reasoning abilities in isolation. Many prior datasets rely on highly synthetic, repetitive language with limited vocabulary and shallow reasoning chains, or they confound pure logic with commonsense knowledge. The article introduces FOLIO, an expert-annotated benchmark designed to rigorously evaluate deductive natural language reasoning and formal translation into first-order logic.

The article evaluates both supervised language models and state-of-the-art large language models across natural language reasoning and natural language-to-logic translation. The dataset comprises 1,430 conclusions across 487 multi-premise stories created through open-ended writing based on Wikipedia articles and structured syllogistic templates. Expert annotators paired every natural language premise and conclusion with mathematically sound first-order logic formulas, which were computationally verified using an automated theorem prover to guarantee ground-truth logical consistency.

Evaluation reveals that complex deductive reasoning remains a critical vulnerability for top-tier models. Under standard few-shot prompting, GPT-4 achieved an accuracy of 64.2%, trailing expert human performance of 95.98% by nearly 32 percentage points. Error analysis demonstrates that model accuracy deteriorates steeply as reasoning depth increases: GPT-4 fell from 75.43% on shorter open-ended stories to 53.10% on complex template-based stories requiring 5 to 8 inferential steps. An evaluation of model failures showed that 65% stemmed from constructing faulty reasoning paths, while 25% were due to erroneous deduction steps. In translation experiments, while models generated syntactically valid logic expressions over 93% of the time, their execution accuracy remained low at 56% to 64%.

These findings indicate that general-purpose language models cannot yet reliably serve as autonomous deductive decision-makers in zero-shot or standard few-shot configurations. In high-stakes applications such as legal compliance, medical protocol verification, or automated policy enforcement, relying on unassisted models introduces substantial risk of logical failure. However, specialized neuro-symbolic methods that pair language models with external logic solvers—such as Logic-LM and DetermLR—boosted accuracy to over 77%, demonstrating that hybrid architectures significantly mitigate deductive errors.

Organizations developing logic-dependent AI workflows should prioritize hybrid neuro-symbolic architectures over pure natural language prompting. Future technical work should focus on scaling up high-quality verified reasoning corpora, improving intermediate chain-of-thought planning to prevent faulty paths, and creating more nuanced translation evaluation metrics. Confidence in these findings is high due to rigorous formal verification by an inference engine, though users should note that the dataset's size reflects a deliberate prioritization of expert quality over massive scale.

Abstract

Large language models (LLMs) have achieved remarkable performance on a variety of natural language understanding tasks. However, existing benchmarks are inadequate in measuring the complex logical reasoning capabilities of a model. We present FOLIO, a human-annotated, logically complex and diverse dataset for reasoning in natural language (NL), equipped with first-order logic (FOL) annotations. FOLIO consists of 1,430 examples (unique conclusions), each paired with one of 487 sets of premises used to deductively reason for the validity of each conclusion. The logical correctness of the premises and conclusions is ensured by their FOL annotations, which are automatically verified by an FOL inference engine. In addition to the main NL reasoning task, NL-FOL pairs in FOLIO constitute a new NL-FOL translation dataset. Our experiments on FOLIO systematically evaluate the FOL reasoning ability of supervised fine-tuning on medium-sized language models. For both NL reasoning and NL-FOL translation, we benchmark multiple state-of-the-art language models. Our results show that a subset of FOLIO presents a challenge for one of the most capable Large Language Model (LLM) publicly available, GPT-4.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Datasets for reasoning from text
  • 2.2 Reasoning using large language models
  • 3 FOLIO Corpus Construction
  • 3.1 Example collection
  • 3.2 Quality control for NL sentences
  • 3.3 Quality control for FOL formulas
  • 3.4 NL-FOL alignment review
  • 3.5 FOL verification
  • 3.6 Dataset statistics
  • 4 Task Definition
  • 4.1 Natural language reasoning with first-order logic
  • 4.2 NL-FOL translation
  • 5 Experiments
  • 5.1 Experimental setup
  • 5.2 Models
  • 5.3 Main results
  • 6 Error Analysis
  • 6.1 Human performance
  • 7 Conclusion
  • 8 Limitations
  • References
  • A Annotator Selection
  • B HybLogic Template Example
  • C Factuality and Bias Elimination Protocol
  • D Language Quality Control
  • E First-Order Logic
  • E.1 First-Order Logic VS Natural Language
  • E.2 FOL definition
  • E.3 FOL modeling conventions
  • F FOL Annotation Protocol
  • G FOL Inference Engine
  • H Distribution of Readability
  • I Case study
  • J Model Performance Analysis

Knowls

  1. Knowl 1 — FOLIO pairs natural-language deduction stories with first-order logic

    definition

    FOLIO is a human-annotated dataset for deciding whether conclusions follow from natural-language premises by first-order logic (FOL). It contains 487 stories, with 2,407 premise pairs and 1,435 conclusion pairs; each natural-language premise and conclusion is paired with an annotated FOL formula. Conclusions are labeled True, False, or Unknown. The paired representations support both deductive reasoning from natural-language stories and translation between natural language and FOL.

  2. Knowl 2 — Stories are collected through free-form and syllogism-template annotation

    model/method

    FOLIO combines two human data-collection approaches. In WikiLogic, annotators use randomly selected Wikipedia articles as topic inspiration, then write plausible stories from scratch rather than copying article text or applying fixed templates. In HybLogic, annotators first receive logically valid story templates formed by chaining valid syllogisms—using one syllogism’s conclusion as the next premise—and adding operators such as conjunction, disjunction, and implication. Annotators instantiate the templates with real-world entities, categories, and natural-language wording. The dataset contains 304 WikiLogic stories and 183 HybLogic stories.

  3. Knowl 3 — Expert review and automated proving are used to control annotation quality

    model/method

    FOLIO’s stories and FOL annotations were written and reviewed by annotators with strong English proficiency and formal FOL or semantic-parsing preparation; specialists in language quality reviewed the natural language, while FOL experts reviewed the formulas. Quality control covered naturalness, grammar, ambiguity, factuality, and bias. Screening found at least one factuality or bias issue in 39.2% of stories, which were then handled under a rewriting protocol. FOL annotations were guided to preserve sentence meaning and structure, while avoiding unnecessary semantic decomposition. Reviewers also aligned natural-language and FOL meanings and added missing commonsense premises when needed for the intended deduction. Finally, a parser converted the annotated formulas into the format required by an inference engine, which checked formula syntax and the consistency of conclusion labels with the premises.

  4. Knowl 4 — FOLIO defines story-level reasoning and NL-to-FOL translation tasks

    definition

    In the reasoning task, the input is a natural-language story containing multiple premises and multiple conclusions; the output is a True, False, or Unknown label for each conclusion, based on deduction from the premises. Each story also has a parallel FOL representation. In the NL-to-FOL translation task, a system translates the entire story—including its premises and conclusions—into FOL formulas that are logically and semantically equivalent to the natural-language sentences and preserve the conclusions’ truth values. Translation is evaluated with two measures: Syntactic Validity (SynV), which is 1 only when every formula in a story passes the syntax check and 0 otherwise; and inference-engine execution accuracy (ExcAcc), the accuracy of conclusion labels returned when the translated premises and conclusions are run through the prover.

  5. Knowl 5 — FOLIO combines a broad vocabulary with varied logical structures and reasoning depths

    data/table

    The dataset contains 4,351 words of natural-language vocabulary and 76 distinct sentence-level abstract syntax trees (ASTs), indicating varied logical forms. Its 487 stories contain 2,407 premise pairs and 1,435 conclusion pairs; natural-language sentences average 9.86 words. The reasoning-depth distribution has a mode of four, and 28.7% of examples require at least five reasoning steps. The dataset’s two collection methods differ in story structure: most WikiLogic conclusions depend on one to five premises, while HybLogic conclusions generally depend on five to eight.

  6. Knowl 6 — Benchmark evaluation uses story-disjoint splits and several prompting conditions

    experimental setup

    FOLIO is split by story into 1,001 training, 203 validation, and 226 test examples, so test stories are unseen during training. Logical reasoning is scored by accuracy; NL-to-FOL translation uses SynV and ExcAcc. Fine-tuning experiments include BERT-base/large, RoBERTa-base/large, and Flan-T5-Large. Prompting experiments include zero-shot and eight-shot settings; GPT-4 is also tested with chain-of-thought, chain-of-thought with self-consistency, and tree-of-thought prompting. Logic-LM, LINC, and DetermLR are evaluated with GPT-4 as the base model. Reported prompting results are averaged over five randomly sampled sets of training examples.

  7. Knowl 7 — Logical-reasoning accuracy varies substantially across model and prompting methods

    empirical result

    On the FOLIO test set, the majority-label baseline is 38.5% and random prediction is 33.3%. Flan-T5-Large has the highest reported accuracy among the fine-tuned models at 65.9%. Eight-shot prompting with GPT-4 reaches 64.2% without an added reasoning strategy, rising to 70.0% with tree-of-thought prompting. The highest reported results come from GPT-4-based logical-reasoning systems: Logic-LM reaches 78.1% and DetermLR reaches 77.5%. The table reports accuracy for the listed test-set conditions; model sizes are unavailable for the proprietary GPT models.

    Condition Model Size Accuracy (%)
    Baseline Majority – 38.5
    Baseline Random – 33.3
    Fine-tuning BERT-base 110M 56.8
    Fine-tuning BERT-large 340M 59.0
    Fine-tuning RoBERTa-base 110M 56.8
    Fine-tuning RoBERTa-large 340M 62.1
    Fine-tuning Flan-T5-Large 783M 65.9
    Zero-shot GPT-3.5-Turbo – 53.1
    Zero-shot GPT-4 – 61.3
    Eight-shot LLaMA-13B 13B 33.6
    Eight-shot LLaMA-70B 70B 44.0
    Eight-shot LLaMA-70B, chain-of-thought 70B 47.8
    Eight-shot LLaMA-70B, tree-of-thought 70B 48.4
    Eight-shot text-davinci-002 – 49.5
    Eight-shot GPT-3.5-Turbo – 58.3
    Eight-shot GPT-4 – 64.2
    Eight-shot GPT-4, chain-of-thought – 68.9
    Eight-shot GPT-4, chain-of-thought with self-consistency – 69.5
    Eight-shot GPT-4, tree-of-thought – 70.0
    GPT-4-based logical method Logic-LM – 78.1
    GPT-4-based logical method LINC – 73.1
    GPT-4-based logical method DetermLR – 77.5
  8. Knowl 8 — Language models often produce valid FOL syntax without a faithful translation

    empirical result

    In NL-to-FOL translation, eight-shot GPT-4 achieves 93.9% SynV but only 63.8% ExcAcc; eight-shot GPT-3.5-Turbo achieves 93.3% SynV and 56.0% ExcAcc. Thus, in these conditions, syntactically valid output is more common than output whose prover-derived conclusion labels match the dataset labels. Zero-shot results are lower on both measures. SynV is a story-level all-formulas-valid score, whereas ExcAcc evaluates the conclusions produced by executing the translated story.

    Model Prompting SynV (%) ExcAcc (%) SynV (%) ExcAcc (%)
    Zero-shot Eight-shot
    GPT-3.5-Turbo 68.4 50.4 93.3 56.0
    GPT-4 86.1 51.7 93.9 63.8
  9. Knowl 9 — Performance is lower on deeper reasoning stories, especially for prompted LLMs on HybLogic

    empirical result

    Accuracy is higher on examples with zero to three reasoning depths than on examples with four to seven depths for few-shot GPT-3.5 and GPT-4; RoBERTa’s fine-tuned performance also favors the shallower group, but by a smaller margin. Source-specific test results show that GPT-3.5-Turbo and GPT-4 perform much worse on HybLogic than on WikiLogic for natural-language prompting, while RoBERTa-large is slightly more accurate on HybLogic. NL-to-FOL execution accuracy shows the reverse pattern for both prompted models: accuracy is higher on HybLogic. The authors suggest that deeper reasoning contributes to the prompted-model gap and that HybLogic’s story-level complexity can coexist with simpler sentence-level translations; these are proposed explanations rather than experimentally isolated causes.

    Evaluation Model WikiLogic accuracy (%) HybLogic accuracy (%)
    Fine-tuning RoBERTa-large 60.71 63.48
    NL prompting GPT-3.5-Turbo 68.88 47.70
    NL prompting GPT-4 75.43 53.10
    NL-FOL ExcAcc GPT-3.5-Turbo 45.17 61.82
    NL-FOL ExcAcc GPT-4 59.12 67.93
  10. Knowl 10 — Expert annotators outperform non-experts, and GPT-4 errors mainly involve faulty reasoning paths

    empirical result

    On the FOLIO test set, FOL-familiar expert annotators achieve 95.98% accuracy, compared with 61.82% for native-English-speaking non-experts who had not studied FOL; the reported expert-to-GPT-4 accuracy gap is 31.82 percentage points. A human review of GPT-4’s incorrect truth-value predictions categorizes 65% as faulty reasoning paths, 25% as errors within derivation steps, 5% as difficulty understanding complex conclusion syntax, and 5% as spurious shortcuts based on commonsense reasoning. These observations indicate that incorrect answers were most often associated with constructing or carrying out the required deductions, rather than merely producing an invalid output format.

Coverage note — Secondary premise-order shuffling and NL-only versus FOL-only versus combined-input prompting analyses are omitted because they refine robustness and prompting conclusions but are not needed to reconstruct the dataset, task definitions, or principal benchmark findings.

References

  1. 1.David Arps, Jan Kels, Florian Krämer, Yunus Renz, Regina Stodden, and Wiebke Petersen. 2022. HHU-plexity at text complexity DE challenge 2022. In Proceedings of the GermEval 2022 Workshop on Text Complexity Assessment of German Text, pages 27–32, Potsdam, Germany. Association for Computational Linguistics.
  2. 2.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint.
  3. 3.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal. Association for Computational Linguistics.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang. 2020. HybridQA: A dataset of multi-hop question answering over tabular and textual data. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1026–1036, Online. Association for Computational Linguistics.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020. Transformers as soft reasoners over language. CoRR, abs/2002.05867.
  8. 8.Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2021. Transformers as soft reasoners over language. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 3882–3890.
  9. 9.Robin Cooper, Dick Crouch, Jan Van Eijck, Chris Fox, Johan Van Genabith, Jan Jaspars, Hans Kamp, David Milward, Manfred Pinkal, Massimo Poesio, et al. 1996. Using the framework. Technical report, Technical Report LRE 62-051 D-16, The FraCaS Consortium.
  10. 10.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022. Selection-inference: Exploiting large language models for interpretable logical reasoning. arXiv preprint arXiv:2205.09712.
  11. 11.Edgar Dale and Jeanne S. Chall. 1948. A formula for predicting readability. Educational Research Bulletin, 27(1):11–28.
  12. 12.Edgar Dale and Jeanne S. Chall. 1995. Readability Revisited: The New Dale-Chall Readability Formula. Brookline Books.
  13. 13.Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. 2022. Language models show human-like content effects on reasoning. arXiv preprint arXiv:2207.07051.
  14. 14.Donald Davidson. 2001. 105The Logical Form of Action Sentences. In Essays on Actions and Events. Oxford University Press.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-dickstein, Kevin Murphy, and Charles Sutton. 2022. Language model cascades. arXiv preprint.
  17. 17.Weinan He, Canming Huang, Yongmei Liu, and Xiaodan Zhu. 2021. WinoLogic: A zero-shot logic-based diagnostic dataset for Winograd Schema Challenge. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3779–3789, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Akira Kawabata and Saku Sugawara. 2023. Evaluating the rationale understanding of critical reasoning in logical reading comprehension. Preprint, arXiv:2311.18353.
  19. 19.Mehran Kazemi, Quan Yuan, Deepti Bhatia, Najoung Kim, Xin Xu, Vaiva Imbrasaite, and Deepak Ramachandran. 2023. Boardgameqa: A dataset for natural language reasoning with contradictory information. Preprint, arXiv:2306.07934.
  20. 20.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  21. 21.Anne Lehman. 1973. Two sets of perfect syllogisms. Notre Dame Journal of Formal Logic, 14(3):425 – 429.
  22. 22.Sarah-Jane Leslie. 2008. Generics: Cognition and Acquisition. The Philosophical Review, 117(1):1–47.
  23. 23.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  24. 24.Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2021. Logiqa: a challenge dataset for machine reading comprehension with logical reasoning. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 3622–3628.
  25. 25.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Ro{bert}a: A robustly optimized {bert} pretraining approach. arXiv preprint arXiv:1907.11692.
  26. 26.W. McCune. 2005–2010. Prover9 and mace4. http://www.cs.unm.edu/~mccune/prover9/.
  27. 27.Ansong Ni, Pengcheng Yin, Yilun Zhao, Martin Riddell, Troy Feng, Rui Shen, Stephen Yin, Ye Liu, Semih Yavuz, Caiming Xiong, Shafiq Joty, Yingbo Zhou, Dragomir Radev, and Arman Cohan. 2023. L2ceval: Evaluating language-to-code generation capabilities of large language models. Preprint, arXiv:2309.17446.
  28. 28.Tobias Nipkow, Lawrence C. Paulson, and Markus Wenzel. 2002. Isabelle/Hol a Proof Assistant for Higher-Order Logic. Springer.
  29. 29.Theo Olausson, Alex Gu, Ben Lipkin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy. 2023. LINC: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5153–5176, Singapore. Association for Computational Linguistics.
  30. 30.OpenAI, Josh Achiam, and Others. 2023. Gpt-4 technical report. Preprint, arXiv:2303.08774.
  31. 31.Liangming Pan, Alon Albalak, Xinyi Wang, and William Wang. 2023. Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3806–3824, Singapore. Association for Computational Linguistics.
  32. 32.Terence Parsons. 1990. Events in the Semantics of English. MIT Press, Cambridge, MA, USA.
  33. 33.Stuart Russell and Peter Norvig. 2010. Artificial Intelligence: A Modern Approach, 3 edition. Prentice Hall.
  34. 34.Mohammed Saeed, Naser Ahmadi, Preslav Nakov, and Paolo Papotti. 2021. RuleBERT: Teaching soft rules to pre-trained language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1460–1476, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  35. 35.Abulhair Saparov and He He. 2023. Language models can (kind of) reason: A systematic formal analysis of chain-of-thought. In International Conference on Learning Representations.
  36. 36.Hrituraj Singh, Milan Aggrawal, and Balaji Krishnamurthy. 2020. Exploring neural models for parsing natural language into first-order logic. arXiv preprint arXiv:2002.06544.
  37. 37.Pranaydeep Singh, Luna De Bruyne, Orphée De Clercq, and Els Lefever. 2023. Misery loves complexity: Exploring linguistic complexity in the context of emotion detection. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12871–12880, Singapore. Association for Computational Linguistics.
  38. 38.Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton. 2019. CLUTRR: A diagnostic benchmark for inductive reasoning from text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4506–4515, Hong Kong, China. Association for Computational Linguistics.
  39. 39.Aarohi Srivastava, Abhinav Rastogi, and +447 Authors. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Preprint, arXiv:2206.04615.
  40. 40.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  41. 41.Hongda Sun, Weikai Xu, Wei Liu, Jian Luan, Bin Wang, Shuo Shang, Ji-Rong Wen, and Rui Yan. 2023. From indeterminacy to determinacy: Augmenting logical reasoning capabilities with large language models. Preprint, arXiv:2310.18659.
  42. 42.G. Sutcliffe. 2017. The TPTP Problem Library and Associated Infrastructure. From CNF to TH0, TPTP v6.4.0. Journal of Automated Reasoning, 59(4):483–502.
  43. 43.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3621–3634, Online. Association for Computational Linguistics.
  44. 44.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota. Association for Computational Linguistics.
  45. 45.Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, and Jonathan Berant. 2020. Leap-of-thought: Teaching pre-trained models to systematically reason over implicit knowledge. Advances in Neural Information Processing Systems, 33:20227–20237.
  46. 46.Jidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao, Hao He, and Yaohui Jin. 2021. Diagnosing the first-order logical reasoning ability through LogicNLI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3738–3747, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. Preprint, arXiv:2302.13971.
  48. 48.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019a. Superglue: A stickier benchmark for general-purpose language understanding systems. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  49. 49.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019b. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32.
  50. 50.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  51. 51.Jason Wei, Kelly Finn, Emma Templeton, Thalia Wheatley, and Soroush Vosoughi. 2021. Linguistic complexity loss in text-based therapy. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4450–4459, Online. Association for Computational Linguistics.
  52. 52.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  53. 53.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  54. 54.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  55. 55.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. In Thirty-seventh Conference on Neural Information Processing Systems.
  56. 56.Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020. Reclor: A reading comprehension dataset requiring logical reasoning. In International Conference on Learning Representations.
  57. 57.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022. Star: Bootstrapping reasoning with reasoning. arXiv preprint.

Citation

MLA
Han, S., et al. “FOLIO: Natural Language Reasoning with First-Order Logic”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 22017–31, https://doi.org/10.18653/v1/2024.emnlp-main.1229.
APA
Han, S., Schoelkopf, H., Zhao, Y., Qi, Z., Riddell, M., Zhou, W., Coady, J., Peng, D., Qiao, Y., Benson, L., Sun, L., Wardle-Solano, A., Szabó, H., Zubova, E., Burtell, M., Fan, J., Liu, Y., Wong, B., Sailor, M., … Radev, D. (2024). FOLIO: Natural Language Reasoning with First-Order Logic. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 22017–22031. https://doi.org/10.18653/v1/2024.emnlp-main.1229
Chicago
Han, S., H. Schoelkopf, Y. Zhao, et al. 2024. “FOLIO: Natural Language Reasoning with First-Order Logic”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 22017–31. https://doi.org/10.18653/v1/2024.emnlp-main.1229.
Harvard
Han, S. et al. (2024) “FOLIO: Natural Language Reasoning with First-Order Logic”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 22017–22031. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1229.
Vancouver
1. Han S, Schoelkopf H, Zhao Y, et al (2024) FOLIO: Natural Language Reasoning with First-Order Logic. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 22017–22031

BibTeX

@inproceedings{han-etal-2024-folio,
    title = "{FOLIO}: Natural Language Reasoning with First-Order Logic",
    author = "Han, Simeng  and
      Schoelkopf, Hailey  and
      Zhao, Yilun  and
      Qi, Zhenting  and
      Riddell, Martin  and
      Zhou, Wenfei  and
      Coady, James  and
      Peng, David  and
      Qiao, Yujie  and
      Benson, Luke  and
      Sun, Lucy  and
      Wardle-Solano, Alexander  and
      Szab{\'o}, Hannah  and
      Zubova, Ekaterina  and
      Burtell, Matthew  and
      Fan, Jonathan  and
      Liu, Yixin  and
      Wong, Brian  and
      Sailor, Malcolm  and
      Ni, Ansong  and
      Nan, Linyong  and
      Kasai, Jungo  and
      Yu, Tao  and
      Zhang, Rui  and
      Fabbri, Alexander  and
      Kryscinski, Wojciech Maciej  and
      Yavuz, Semih  and
      Liu, Ye  and
      Lin, Xi Victoria  and
      Joty, Shafiq  and
      Zhou, Yingbo  and
      Xiong, Caiming  and
      Ying, Rex  and
      Cohan, Arman  and
      Radev, Dragomir",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1229/",
    doi = "10.18653/v1/2024.emnlp-main.1229",
    pages = "22017--22031"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/