BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information
Mehran KazemiQuan YuanDeepti BhatiaNajoung KimXin XuVaiva ImbrasaiteDeepak Ramachandran
Introduces BoardgameQA, a benchmark for evaluating language models on multi-hop defeasible reasoning with contradictory rules and missing background knowledge, revealing that state-of-the-art models struggle to resolve conflicting information even after fine-tuning.
Real-world automated reasoning requires systems to handle conflicting and incomplete information, such as contradictory data from multiple web sources or general rules that have explicit exceptions. While large language models have demonstrated strong reasoning capabilities on consistent inputs, existing benchmarks assume the input data is fully coherent and complete. To evaluate how models manage realistic reasoning challenges, the article introduces BoardgameQA, a synthetic benchmark designed to measure multi-hop defeasible reasoning—a framework where conflicts between contradictory rules are resolved using source preferences and background knowledge.
The article systematically evaluates various model architectures and learning paradigms across controlled reasoning scenarios. The evaluation includes encoder-only, encoder-decoder, and decoder-only language models tested under fine-tuning, soft prompt-tuning, and few-shot in-context learning with chain-of-thought prompting. The BoardgameQA benchmark isolates specific reasoning dimensions by programmatically adjusting reasoning depth, the frequency and type of rule contradictions, the amount of required unstated commonsense knowledge, and the presence of distracting facts.
The experiments show that current language models struggle significantly when reasoning with contradictory inputs. Few-shot models fail to perform conflict resolution out-of-the-box, exhibiting performance near random chance across multi-hop scenarios. While fine-tuning and prompt-tuning improve accuracy on simple one-hop tasks, model performance degrades sharply as reasoning depth increases. Additionally, evaluation of intermediate reasoning steps reveals that models frequently produce invalid or hallucinated proofs even when predicting the correct final label. Model accuracy drops monotonically as the frequency of conflicts and distracting facts increases, and smaller models suffer severe performance declines when required background information is omitted.
These findings indicate that current language models are brittle when deployed in environments containing noisy, conflicting, or incomplete knowledge. Organizations relying on language models for high-stakes decision-making, such as automated policy analysis, legal reasoning, or intelligence retrieval, face substantial operational risks if systems accept contradictory inputs without specialized reasoning mechanisms. Standard prompting techniques and basic model scaling are insufficient to guarantee dependable conflict resolution.
To address these limitations, development teams should avoid relying on off-the-shelf few-shot models for tasks involving conflicting evidence. Organizations should invest in modular reasoning frameworks that explicitly track source preferences, fine-tune models on formal proof generation, or integrate symbolic solvers to ensure logical faithfulness. The primary limitations of the study include its focus on binary contradiction within a structured board game domain using rule-based preferences. Future work should evaluate non-binary conflicts, broader logical rule structures, and complex real-world document environments to build more resilient reasoning systems.
- Paper: Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations, Jaehun Jung et al. (2022). This paper establishes formal methods for resolving contradictory model explanations and enforcing logical consistency via satisfiability solvers, providing direct foundational context for non-monotonic and defeasible reasoning benchmarks in LLMs.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). This study analyzes how language models struggle with multi-step formal logical deductions and proof planning, providing necessary background on LLM rule-following limitations before evaluating defeasible reasoning over contradictory rules.
- Paper: Large Language Models Can Be Easily Distracted by Irrelevant Context, Freda Shi et al. (2023). This work demonstrates how distracting, conflicting, or irrelevant textual context breaks multi-step language model reasoning, laying important groundwork for evaluating reasoning under conflicting information sources.
- Paper: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge, Alon Talmor et al. (2019). This benchmark introduced the paradigm of evaluating question answering requiring implicit background commonsense knowledge, which BoardgameQA builds upon to reflect downstream application reasoning.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). This paper introduces the standard self-consistency sampling methodology used to evaluate and elicit multi-path reasoning in large language models.
- Paper: What Evidence Do Language Models Find Convincing?, Alexander Wan et al. (2024). This paper directly extends the investigation of reasoning with contradictory information by assessing how LLMs resolve conflicting evidence in subjective and retrieval-augmented settings.
- Paper: Exploring the Potential of Large Language Models in Computational Argumentation, Guizhen Chen et al. (2024). This work applies and evaluates language models in computational argumentation and counter-argument generation, broadening the study of defeasible reasoning with competing claims to full natural language debate.
- Paper: Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate, Tian Liang et al. (2024). This work develops a multi-agent debate framework to resolve conflicting hypotheses and reasoning traps, offering an interactive system to mitigate the contradictory reasoning failures benchmarked in BoardgameQA.
- Paper: Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training, Feiteng Fang et al. (2024). This paper investigates methods to improve model robustness when handling counterfactual and noisy conflicting context during retrieval-augmented reasoning.
