FOLIO: Natural Language Reasoning with First-Order Logic
Simeng HanHailey SchoelkopfYilun ZhaoZhenting QiMartin RiddellWenfei ZhouJames CoadyDavid PengYujie QiaoLuke Benson
Introduces FOLIO, an expert-annotated benchmark with parallel first-order logic formulas and automated validity verification, exposing major deductive reasoning limitations in leading large language models like GPT-4.
Artificial intelligence systems increasingly handle complex tasks, yet existing evaluation benchmarks fail to adequately measure their logical reasoning abilities in isolation. Many prior datasets rely on highly synthetic, repetitive language with limited vocabulary and shallow reasoning chains, or they confound pure logic with commonsense knowledge. The article introduces FOLIO, an expert-annotated benchmark designed to rigorously evaluate deductive natural language reasoning and formal translation into first-order logic.
The article evaluates both supervised language models and state-of-the-art large language models across natural language reasoning and natural language-to-logic translation. The dataset comprises 1,430 conclusions across 487 multi-premise stories created through open-ended writing based on Wikipedia articles and structured syllogistic templates. Expert annotators paired every natural language premise and conclusion with mathematically sound first-order logic formulas, which were computationally verified using an automated theorem prover to guarantee ground-truth logical consistency.
Evaluation reveals that complex deductive reasoning remains a critical vulnerability for top-tier models. Under standard few-shot prompting, GPT-4 achieved an accuracy of 64.2%, trailing expert human performance of 95.98% by nearly 32 percentage points. Error analysis demonstrates that model accuracy deteriorates steeply as reasoning depth increases: GPT-4 fell from 75.43% on shorter open-ended stories to 53.10% on complex template-based stories requiring 5 to 8 inferential steps. An evaluation of model failures showed that 65% stemmed from constructing faulty reasoning paths, while 25% were due to erroneous deduction steps. In translation experiments, while models generated syntactically valid logic expressions over 93% of the time, their execution accuracy remained low at 56% to 64%.
These findings indicate that general-purpose language models cannot yet reliably serve as autonomous deductive decision-makers in zero-shot or standard few-shot configurations. In high-stakes applications such as legal compliance, medical protocol verification, or automated policy enforcement, relying on unassisted models introduces substantial risk of logical failure. However, specialized neuro-symbolic methods that pair language models with external logic solvers—such as Logic-LM and DetermLR—boosted accuracy to over 77%, demonstrating that hybrid architectures significantly mitigate deductive errors.
Organizations developing logic-dependent AI workflows should prioritize hybrid neuro-symbolic architectures over pure natural language prompting. Future technical work should focus on scaling up high-quality verified reasoning corpora, improving intermediate chain-of-thought planning to prevent faulty paths, and creating more nuanced translation evaluation metrics. Confidence in these findings is high due to rigorous formal verification by an inference engine, though users should note that the dataset's size reflects a deliberate prioritization of expert quality over massive scale.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
