BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information

Mehran KazemiQuan YuanDeepti BhatiaNajoung KimXin XuVaiva ImbrasaiteDeepak Ramachandran

article2023NeurIPS78 citationsOutstanding Paper Award

Introduces BoardgameQA, a benchmark for evaluating language models on multi-hop defeasible reasoning with contradictory rules and missing background knowledge, revealing that state-of-the-art models struggle to resolve conflicting information even after fine-tuning.

Listen

Real-world automated reasoning requires systems to handle conflicting and incomplete information, such as contradictory data from multiple web sources or general rules that have explicit exceptions. While large language models have demonstrated strong reasoning capabilities on consistent inputs, existing benchmarks assume the input data is fully coherent and complete. To evaluate how models manage realistic reasoning challenges, the article introduces BoardgameQA, a synthetic benchmark designed to measure multi-hop defeasible reasoning—a framework where conflicts between contradictory rules are resolved using source preferences and background knowledge.

The article systematically evaluates various model architectures and learning paradigms across controlled reasoning scenarios. The evaluation includes encoder-only, encoder-decoder, and decoder-only language models tested under fine-tuning, soft prompt-tuning, and few-shot in-context learning with chain-of-thought prompting. The BoardgameQA benchmark isolates specific reasoning dimensions by programmatically adjusting reasoning depth, the frequency and type of rule contradictions, the amount of required unstated commonsense knowledge, and the presence of distracting facts.

The experiments show that current language models struggle significantly when reasoning with contradictory inputs. Few-shot models fail to perform conflict resolution out-of-the-box, exhibiting performance near random chance across multi-hop scenarios. While fine-tuning and prompt-tuning improve accuracy on simple one-hop tasks, model performance degrades sharply as reasoning depth increases. Additionally, evaluation of intermediate reasoning steps reveals that models frequently produce invalid or hallucinated proofs even when predicting the correct final label. Model accuracy drops monotonically as the frequency of conflicts and distracting facts increases, and smaller models suffer severe performance declines when required background information is omitted.

These findings indicate that current language models are brittle when deployed in environments containing noisy, conflicting, or incomplete knowledge. Organizations relying on language models for high-stakes decision-making, such as automated policy analysis, legal reasoning, or intelligence retrieval, face substantial operational risks if systems accept contradictory inputs without specialized reasoning mechanisms. Standard prompting techniques and basic model scaling are insufficient to guarantee dependable conflict resolution.

To address these limitations, development teams should avoid relying on off-the-shelf few-shot models for tasks involving conflicting evidence. Organizations should invest in modular reasoning frameworks that explicitly track source preferences, fine-tune models on formal proof generation, or integrate symbolic solvers to ensure logical faithfulness. The primary limitations of the study include its focus on binary contradiction within a structured board game domain using rule-based preferences. Future work should evaluate non-binary conflicts, broader logical rule structures, and complex real-world document environments to build more resilient reasoning systems.

Cover for BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information

Abstract

Automated reasoning with unstructured natural text is a key requirement for many potential applications of NLP and for developing robust AI systems. Recently, Language Models (LMs) have demonstrated complex reasoning capacities even without any finetuning. However, existing evaluation for automated reasoning assumes access to a consistent and coherent set of information over which models reason. When reasoning in the real-world, the available information is frequently inconsistent or contradictory, and therefore models need to be equipped with a strategy to resolve such conflicts when they arise. One widely-applicable way of resolving conflicts is to impose preferences over information sources (e.g., based on source credibility or information recency) and adopt the source with higher preference. In this paper, we formulate the problem of reasoning with contradictory information guided by preferences over sources as the classical problem of defeasible reasoning, and develop a dataset called BoardgameQA for measuring the reasoning capacity of LMs in this setting. BoardgameQA also incorporates reasoning with implicit background knowledge, to better reflect reasoning problems in downstream applications. We benchmark various LMs on BoardgameQA and the results reveal a significant gap in the reasoning capacity of state-of-the-art LMs on this problem, showing that reasoning with conflicting information does not surface out-of-the-box in LMs. While performance can be improved with finetuning, it nevertheless remains poor.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background and Notation
  • 4 The BoardgameQA Dataset
  • 5 Experiments
  • 5.1 Can LMs Reason with Contradictory Inputs?
  • 5.2 Does Correct Label Prediction Mean Correct Proof?
  • 5.3 Do Conflicts Make Reasoning More Difficult?
  • 5.4 Which Conflict Type is More Difficult to Resolve?
  • 5.5 Does Information Incompleteness Make Reasoning More Difficult?
  • 5.6 Do Distractors Make Reasoning More Difficult?
  • 6 Limitations
  • 7 Conclusion
  • Acknowledgements
  • References
  • A More Experimental Results and Analysis
  • B Experimental Details
  • C BoardgameQA Details
  • C.1 Consistency of the Dataset
  • C.2 Incomplete Information
  • C.3 Entities, Predicates, and Templates
  • C.4 Sample Proofs

Knowls

  1. Knowl 1 — Formalization of Defeasible Theory and Conflict Types

    definition

    A defeasible theory is defined as a tuple T(d)=(F,R,O)\mathcal{T}^{(d)} = (\mathcal{F}, \mathcal{R}, \mathcal{O}), where:

    • F\mathcal{F} is a set of positive or negative facts describing the environment state. The initial facts in F\mathcal{F} are internally consistent and hold priority over any derived facts.
    • R={r1,…,r∣R∣}\mathcal{R} = \{r_1, \dots, r_{|\mathcal{R}|}\} is a set of defeasible rules of the form r:rb→rhr: r_b \rightarrow r_h, where rbr_b is the body (antecedent) and rhr_h is the head (consequent). Rules hold defeasibly, meaning a derived conclusion can be defeated by contrary evidence from a rule with higher priority.
    • O={rt1>rt2,… }\mathcal{O} = \{r_{t1} > r_{t2}, \dots\} is a set of pairwise relative priority preferences over the rules in R\mathcal{R}.

    The entailment relation T(d)⊨f\mathcal{T}^{(d)} \models f indicates that fact ff can be derived from the theory after resolving all conflicts according to O\mathcal{O}. For two rules with contradictory heads r:rb→zr: r_b \rightarrow z and r′:rb′→!zr': r'_b \rightarrow !z, conflict resolution proceeds according to two distinct conflict types:

    • Type 1 Conflict: Rule rr has strictly higher priority than r′r' ((r>r′)∈O(r > r') \in \mathcal{O}) and rbr_b is provable. Here, zz is entailed without needing to evaluate whether the body rb′r'_b of the conflicting rule is provable.
    • Type 2 Conflict: Rule rr has strictly lower priority than r′r' ((r′>r)∈O(r' > r) \in \mathcal{O}), rbr_b is provable, and rb′r'_b cannot be proved. In this case, establishing that rb′r'_b fails to hold is required to entail zz.
  2. Knowl 2 — Backward Theory Generation Algorithm for BoardgameQA

    algorithm

    BoardgameQA constructs defeasible reasoning theories using a recursive backward-chaining generation procedure with parameterized control over reasoning depth dd, probability of conflict pConfp_{\text{Conf}}, probability of Type 1 conflict pConfType1p_{\text{ConfType1}}, and probability of incomplete information pMissInfop_{\text{MissInfo}}.

    GenerateTheory(q, d):
        Input: Target question qq, proof depth dd
        Output: Theory elements added to global sets F\mathcal{F}, R\mathcal{R}, and O\mathcal{O}
        if d==0d == 0 then
            Add qq to F\mathcal{F}
        else
            Q,r=SampleRuleAndSubq(q)Q, r = \text{SampleRuleAndSubq}(q)
            Add rr to R\mathcal{R}
            if CoinFlip(pConf)==Conflict\text{CoinFlip}(p_{\text{Conf}}) == \text{Conflict} then
                Q′,r′=SampleRuleAndSubq(!q)Q', r' = \text{SampleRuleAndSubq}(!q)
                Add r′r' to R\mathcal{R}
                if CoinFlip(pConfType1)==Type1\text{CoinFlip}(p_{\text{ConfType1}}) == \text{Type1} then
                    Q=Q∪SubSample(Q′)Q = Q \cup \text{SubSample}(Q')
                    Add (r>r′)(r > r') to O\mathcal{O}
                else
                    Q=Q∪RemoveOneSubquestion(Q′)Q = Q \cup \text{RemoveOneSubquestion}(Q')
                    Add (r′>r)(r' > r) to O\mathcal{O}
            for each sub-question qi∈Qq_i \in Q do
                GenerateTheory(qiq_i, d−1d - 1)
    SampleRuleAndSubq(q):
        Input: Goal question qq
        Output: Sub-questions QQ, Rule rr
        Sample a rule template rr whose head unifies with qq
        if rr uses incomplete background knowledge (Type 5) then
            Sample Q={q1,…,qn}Q = \{q_1, \dots, q_n\} and implicit knowledge Q^={q1′,…,qm′}\hat{Q} = \{q'_1, \dots, q'_m\} such that qq is derived from Q^\hat{Q} via rr, and Q^\hat{Q} is derived from QQ
        else
            Sample Q={q1,…,qn}Q = \{q_1, \dots, q_n\} such that qq is derived from QQ and rr
        return Q,rQ, r

    To preserve defeasible consistency and eliminate circular dependencies, recursive branches for conjunctions sample from mutually disjoint entity sets. Negative (disproved) instances are created by negating the target question of a proved theory, while unknown instances are constructed by perturbing facts, predicates, rule signs, or preference orders and validating undecidability via an external defeasible logic solver.

  3. Knowl 3 — Rule Templates and Incomplete Information Framework in BoardgameQA

    model/method

    BoardgameQA uses six formal rule templates to generate multi-hop logic scenarios:

    1. Universal implication: ∀X:(X,p1,e1)⇒(X,p2,e2)\forall X : (X, p_1, e_1) \Rightarrow (X, p_2, e_2)
    2. Universal conjunctive implication: ∀X:(X,p1,e1)∧(X,p2,e2)⇒(X,p3,e3)\forall X : (X, p_1, e_1) \wedge (X, p_2, e_2) \Rightarrow (X, p_3, e_3)
    3. Entity-specific implication: (e1,p1,e2)⇒(e2,p2,e3)(e_1, p_1, e_2) \Rightarrow (e_2, p_2, e_3)
    4. Entity-specific conjunctive implication: (e1,p1,e2)∧(e3,p2,e2)⇒(e2,p3,e4)(e_1, p_1, e_2) \wedge (e_3, p_2, e_2) \Rightarrow (e_2, p_3, e_4)
    5. Incomplete information implication: (e1,p^,e^)⇒(e1,p2,e2)(e_1, \hat{p}, \hat{e}) \Rightarrow (e_1, p_2, e_2), where p^\hat{p} or e^\hat{e} is absent from explicit theory facts and must be bridged using implicit commonsense or world knowledge.
    6. Existential implication: ∃X(X,p1,e1)⇒(e2,p2,e3)\exists X (X, p_1, e_1) \Rightarrow (e_2, p_2, e_3)

    For incomplete information rules (template 5, sampled with probability pMissInfop_{\text{MissInfo}}), the required bridging reasoning spans multiple categories:

    • Time Conversion: Converting units of age (days, weeks, months, years) to evaluate comparative numeric conditions.
    • Orthography: Comparing lexical features, such as matching the initial letters of entity names.
    • Numeric Comparisons & Money: Calculating arithmetic sums (e.g., combining asset values of multiple players) and checking inequality thresholds.
    • Lexical/Textual Entailment: Recognizing semantic synonymy/implication pairs (e.g., assassinated the mayor ⇒\Rightarrow killed the mayor).
    • World Knowledge & Geography: Identifying geographic inclusions (e.g., in Montreal ⇒\Rightarrow in Canada).
    • Event Times: Placing historical events relative to specific dates (e.g., movie release date relative to the 1969 moon landing).
    • Part Of / Jobs: Mapping specific professions to broader industries (e.g., nurse ⇒\Rightarrow healthcare).
    • Affordance: Associating objects with functional physical attributes (e.g., has a knife ⇒\Rightarrow has a sharp object).
    • 3D Volume & Spatial Fitting: Determining if 3D dimensions of balls or notebooks fit within target box dimensions.
  4. Knowl 4 — Experimental Benchmark Configuration and Evaluation Metrics

    experimental setup

    The BoardgameQA evaluation setup tests language models across three labels: {proved,disproved,unknown}\{\text{proved}, \text{disproved}, \text{unknown}\}, with balanced dataset splits of 1,000 training, 500 validation, and 1,000 test examples (majority class baseline ≈33.3%\approx 33.3\%). Disjoint entity and predicate sets are used across train and test splits to test generalization.

    Evaluated models and paradigms include:

    • BERT-Large: Finetuned with a classification head to directly predict the label without generating reasoning proofs.
    • T5 1.1 XXL: Finetuned to generate the explicit step-by-step natural language proof before outputting the label.
    • PaLM 62B & PaLM 540B: Evaluated under few-shot in-context learning with Chain-of-Thought (CoT) prompts containing one demonstration per label.
    • FLAN-PaLM 540B: Instruction-finetuned model evaluated under few-shot CoT.
    • Prompt-Tuned PaLM 62B: Parameter-efficient soft prompt tuning with a prompt length of 100 tokens, frozen model weights, trained with batch size 8 and learning rate 0.1 for up to 50k steps.

    Evaluation metrics consist of:

    • Label Accuracy: Multi-class classification accuracy.
    • Rule F1: Precision and recall of the extracted rule identifiers used in the model's generated proof relative to the gold proof.
    • Conflict F1: F1 score of rule pairs involved in conflict resolutions correctly cited in the generated proof.
    • Overall Proof Accuracy: Exact logical correctness verified by manual inspection of 50 sampled instances per model on depth 2.
  5. Knowl 5 — Reasoning Degradation Across Reasoning Depths in Defeasible Settings

    empirical result

    Benchmarking models on default BoardgameQA settings (pConf=0.5p_{\text{Conf}} = 0.5, pConfType1=0.5p_{\text{ConfType1}} = 0.5, pMissInfo=0.5p_{\text{MissInfo}} = 0.5) reveals severe performance degradation as reasoning depth increases from 1 to 3 hops:

    • Finetuned T5 1.1 XXL: Achieves ≈80%\approx 80\% accuracy at Depth 1, dropping to ≈64%\approx 64\% at Depth 2, and ≈50%\approx 50\% at Depth 3.
    • Prompt-Tuned PaLM 62B: Reaches ≈78%\approx 78\% accuracy at Depth 1, dropping to ≈68%\approx 68\% at Depth 2, and ≈58%\approx 58\% at Depth 3.
    • Finetuned BERT-Large: Achieves ≈65%\approx 65\% accuracy at Depth 1, falling to ≈36%\approx 36\% at Depth 2, and ≈33%\approx 33\% (majority baseline) at Depth 3.
    • Few-Shot Models (PaLM 62B, PaLM 540B, FLAN-PaLM 540B): Struggle across all depths, scoring between 34%34\% and 55%55\% accuracy, showing that out-of-the-box in-context learning fails at defeasible multi-hop reasoning.

    In confusion matrix analyses of proof-generating models on unknown instances, models at Depth 1 correctly output unknown on over 260 of 333 cases, but at Depth 2 and Depth 3 they almost never predict unknown (e.g., prompt-tuned PaLM 62B predicts unknown for only 2 instances at Depth 2), instead hallucinating spurious derivations.

  6. Knowl 6 — Disparity Between Label Accuracy and Proof Faithfulness

    empirical result

    On Depth 2 examples where the target label is correctly predicted as proved or disproved, evaluation of intermediate proofs reveals that correct label predictions frequently stem from invalid reasoning steps:

    • Rule Identification (Rule F1): All proof-generating models achieve relatively high overlap with gold derivation rules: finetuned T5 XXL ≈80%\approx 80\%, prompt-tuned PaLM 62B ≈78%\approx 78\%, and few-shot PaLM 540B ≈76%\approx 76\%.
    • Conflict Resolution (Conflict F1): Few-shot PaLM 540B fails to model rule preferences, scoring only ≈20%\approx 20\% Conflict F1. Finetuned T5 XXL and prompt-tuned PaLM 62B achieve ≈72%\approx 72\% and ≈68%\approx 68\% Conflict F1, respectively.
    • Overall Proof Accuracy (Manual Validation, N=50N=50): Prompt-tuned PaLM 62B produces fully valid proofs in ≈64%\approx 64\% of correct label predictions, T5 XXL in ≈42%\approx 42\%, and few-shot PaLM 540B in only ≈18%\approx 18\%.

    Dominant proof failure modes include: hallucinating or reversing rule priority preferences, failing to verify all conjuncts in multi-antecedent rules, starting proof chains from irrevelant distractor facts, and asserting false premises to force a proof.

  7. Knowl 7 — Impact of Conflict Frequency on Reasoning Accuracy

    empirical result

    Varying the probability of conflict pConf∈{0.0,0.2,0.5,0.8}p_{\text{Conf}} \in \{0.0, 0.2, 0.5, 0.8\} (corresponding to NoConflict, LowConflict, MediumConflict, and HighConflict splits) causes monotonic performance degradation across all evaluated model architectures at Depth 2:

    • Finetuned T5 XXL: Decreases from ≈80%\approx 80\% accuracy (NoConflict) to ≈73%\approx 73\% (LowConflict), ≈64%\approx 64\% (MediumConflict), and ≈60%\approx 60\% (HighConflict).
    • Prompt-Tuned PaLM 62B: Decreases from ≈74%\approx 74\% (NoConflict) to ≈70%\approx 70\% (LowConflict), ≈68%\approx 68\% (MediumConflict), and ≈59%\approx 59\% (HighConflict).
    • Finetuned BERT-Large: Achieves above-random performance on NoConflict (≈45%\approx 45\%) and LowConflict (≈42%\approx 42\%), but degrades to the random majority baseline (≈33%\approx 33\%) on MediumConflict and HighConflict.
    • Few-Shot PaLM 540B: Declines from ≈58%\approx 58\% (NoConflict) down to ≈46%\approx 46\% (HighConflict).
  8. Knowl 8 — Comparative Reasoning Difficulty of Type 1 vs Type 2 Conflicts

    empirical result

    Experiments varying pConfType1∈{0.2,0.5,0.8}p_{\text{ConfType1}} \in \{0.2, 0.5, 0.8\} (representing Mostly Type 2, Balanced, and Mostly Type 1 conflicts) demonstrate that Type 2 conflicts are systematically more difficult than Type 1 conflicts:

    • In Type 1 conflicts (r>r′r > r'), a model can derive the conclusion zz from rule rr while disregarding whether the lower-priority defeating rule body rb′r'_b activates.
    • In Type 2 conflicts (r′>rr' > r), rule r′r' has strictly higher priority, so entailing zz requires the model to actively prove that the antecedent rb′r'_b fails.

    All models achieve higher accuracy on datasets composed mostly of Type 1 conflicts compared to Type 2 conflicts (e.g., T5 XXL achieves ≈70%\approx 70\% on Mostly Type 1 vs. ≈66%\approx 66\% on Mostly Type 2). Furthermore, tuned models perform best when theories are skewed toward a single conflict type rather than a 50/50 mixture of both types.

  9. Knowl 9 — Effects of Incomplete Information and Distractor Rules on Reasoning

    empirical result

    Evaluating model sensitivity to incomplete external knowledge (pMissInfo∈{0.2,0.5,0.8}p_{\text{MissInfo}} \in \{0.2, 0.5, 0.8\}) and distractor items (0, 1, or 2 distractor facts per step) demonstrates distinct failure profiles:

    • Incomplete Knowledge Sensitivity: Smaller finetuned models degrade substantially as required external knowledge increases; finetuned T5 XXL accuracy drops from ≈80%\approx 80\% (KnowledgeLight) to ≈68%\approx 68\% (KnowledgeMedium) and ≈58%\approx 58\% (KnowledgeHeavy), and BERT-Large drops from ≈45%\approx 45\% to ≈34%\approx 34\%. Conversely, prompt-tuned PaLM 62B and few-shot PaLM 540B exhibit virtually flat accuracy across knowledge levels, reflecting stronger latent world knowledge in larger pre-trained models.
    • Distractor Sensitivity: Adding 1 distractor fact maintains or slightly improves tuned model performance (likely by reducing spurious statistical shortcuts). However, adding 2 distractors (ManyDistractors) causes significant performance drops across all models (e.g., T5 XXL drops from ≈70%\approx 70\% to ≈54%\approx 54\%, prompt-tuned PaLM 62B drops from ≈68%\approx 68\% to ≈58%\approx 58\%, and few-shot PaLM 540B drops monotonically).
  10. Knowl 10 — Scope and Formal Limitations of the BoardgameQA Benchmark

    limitation

    BoardgameQA exhibits several defined structural boundaries in its reasoning setup:

    1. Deductive Entailment and Modus Ponens Scope: Inference is strictly formulated as 3-way deductive classification (label∈{proved,disproved,unknown}\text{label} \in \{\text{proved}, \text{disproved}, \text{unknown}\}) applying the modus ponens rule; it does not cover proof by contradiction, disjunction elimination, or open-ended entity question answering.
    2. Binary Contradictions: Rule conflicts are strictly binary oppositions between an atom zz and its negation !z!z, excluding non-binary semantic mutual exclusions (e.g., asserting mutually exclusive geographic locations).
    3. Rule-Level Preferences Only: Preference orders O\mathcal{O} are defined exclusively over rule pairs (r>r′)(r > r'), leaving out priority conflicts among factual assertions.
    4. Context Length Boundedness: The benchmark assumes all state facts and defeasible rules fit entirely within a single model input prompt.

Coverage note — All substantial contributions—including the defeasible theory formulation, conflict classification taxonomy, dataset construction algorithms, incomplete knowledge categories, benchmark evaluations, proof faithfulness metrics, sensitivity analyses, and formal limitations—are covered in the knowls.

References

  1. 1.Allaway, E., Hwang, J. D., Bhagavatula, C., McKeown, K., Downey, D., and Choi, Y. Penguins don’t fly: Reasoning about generics through instantiations and exceptions. arXiv preprint arXiv:2205.11658, 2022.
  2. 2.Arabshahi, F., Lee, J., Gawarecki, M., Mazaitis, K., Azaria, A., and Mitchell, T. Conversational neuro-symbolic commonsense reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 4902–4911, 2021.
  3. 3.BehnamGhader, P., Miret, S., and Reddy, S. Can retriever-augmented language models reason? the blame game between the retriever and the language model. arXiv preprint arXiv:2212.09146, 2022.
  4. 4.Betz, G., Voigt, C., and Richardson, K. Critical thinking for language models. In Proceedings of the 14th International Conference on Computational Semantics (IWCS), pp. 63–75, Groningen, The Netherlands (online), June 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.iwcs-1.7.
  5. 5.Bhagavatula, C., Bras, R. L., Malaviya, C., Sakaguchi, K., Holtzman, A., Rashkin, H., Downey, D., Yih, S. W.-t., and Choi, Y. Abductive commonsense reasoning. arXiv preprint arXiv:1908.05739, 2019.
  6. 6.Bhakthavatsalam, S., Anastasiades, C., and Clark, P. Genericskb: A knowledge base of generic statements. arXiv preprint arXiv:2005.00660, 2020.
  7. 7.Billi, M., Calegari, R., Contissa, G., Lagioia, F., Pisano, G., Sartor, G., and Sartor, G. Argumentation and defeasible reasoning in the law. J, 4(4):897–914, 2021.
  8. 8.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  9. 9.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. PaLM: Scaling language modeling with pathways. arXiv:2204.02311, 2022. URL https://arxiv.org/abs/2204.02311.
  10. 10.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  11. 11.Clark, P., Tafjord, O., and Richardson, K. Transformers as soft reasoners over language. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI’20, 2021. ISBN 9780999241165. URL https://dl.acm.org/doi/abs/10.5555/3491440.3491977.
  12. 12.Creswell, A., Shanahan, M., and Higgins, I. Selection-inference: Exploiting large language models for interpretable logical reasoning. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=3Pf3Wg6o-A4.
  13. 13.Dalvi, B., Jansen, P., Tafjord, O., Xie, Z., Smith, H., Pipatanangkura, L., and Clark, P. Explaining answers with entailment trees. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7358–7370, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.585. URL https://aclanthology.org/2021.emnlp-main.585.
  14. 14.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  15. 15.Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361, 2021.
  16. 16.Gómez, S. A., Chesnevar, C. I., and Simari, G. R. Defeasible reasoning in web-based forms through argumentation. International Journal of Information Technology & Decision Making, 7(01):71–101, 2008.
  17. 17.Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. Retrieval augmented language model pre-training. In International conference on machine learning, pp. 3929–3938. PMLR, 2020.
  18. 18.Han, S., Schoelkopf, H., Zhao, Y., Qi, Z., Riddell, M., Benson, L., Sun, L., Zubova, E., Qiao, Y., Burtell, M., et al. FOLIO: Natural language reasoning with first-order logic. arXiv:2209.00840, 2022. URL https://arxiv.org/abs/2209.00840.
  19. 19.Hecham, A., Bisquert, P., and Croitoru, M. On a flexible representation for defeasible reasoning variants. In AAMAS 2018-17th International Conference on Autonomous Agents and MultiAgent Systems, number AAMAS’18, pp. 1123–1131, 2018.
  20. 20.Hewitt, C. Planner: A language for proving theorems in robots. In Proceedings of the 1st International Joint Conference on Artificial Intelligence, IJCAI’69, pp. 295–301, San Francisco, CA, USA, 1969. Morgan Kaufmann Publishers Inc.
  21. 21.Hu, S., Luo, Y., Wang, H., Cheng, X., Liu, Z., and Sun, M. Won’t get fooled again: Answering questions with false premises. In ACL, 2023.
  22. 22.Katz, U., Geva, M., and Berant, J. Inferring implicit relations in complex questions with language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 2548–2566, 2022.
  23. 23.Kazemi, S. M., Kim, N., Bhatia, D., Xu, X., and Ramachandran, D. Lambada: Backward chaining for automated reasoning in natural language. In ACL, 2023.
  24. 24.Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A. Decomposed prompting: A modular approach for solving complex tasks. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=_nGgzQjzaRy.
  25. 25.Kim, N. and Schuster, S. Entity tracking in language models. arXiv preprint arXiv:2305.02363, 2023.
  26. 26.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  27. 27.Madaan, A., Tandon, N., Rajagopal, D., Clark, P., Yang, Y., and Hovy, E. Think about it! improving defeasible reasoning by first modeling the question scenario. arXiv preprint arXiv:2110.12349, 2021.
  28. 28.Maher, M. J., Tachmazidis, I., Antoniou, G., Wade, S., and Cheng, L. Rethinking defeasible reasoning: A scalable approach. Theory and Practice of Logic Programming, 20(4):552–586, 2020.
  29. 29.McCarthy, J. Programs with common sense. In Proceedings of the Teddington Conference on the Mechanization of Thought Processes, pp. 75–91, London, 1959. Her Majesty’s Stationary Office. URL http://www-formal.stanford.edu/jmc/mcc59.html.
  30. 30.Min, S., Zettlemoyer, L., Hajishirzi, H., et al. Crepe: Open-domain question answering with false presuppositions. In ACL, pp. arXiv–2211, 2023.
  31. 31.Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A. Show your work: Scratchpads for intermediate computation with language models. In Deep Learning for Code Workshop, 2022. URL https://openreview.net/forum?id=HBlx2idbkbq.
  32. 32.Pan, L., Albalak, A., Wang, X., and Wang, W. Y. Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning. arXiv preprint arXiv:2305.12295, 2023.
  33. 33.Pollock, J. L. Defeasible reasoning. Cognitive science, 11(4):481–518, 1987.
  34. 34.Poole, D. A logical framework for default reasoning. Artificial intelligence, 36(1):27–47, 1988.
  35. 35.Qiao, S., Ou, Y., Zhang, N., Chen, X., Yao, Y., Deng, S., Tan, C., Huang, F., and Chen, H. Reasoning with language model prompting: A survey. arXiv preprint arXiv:2212.09597, 2022.
  36. 36.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  37. 37.Reiter, R. Nonmonotonic reasoning. In Exploring artificial intelligence, pp. 439–481. Elsevier, 1988.
  38. 38.Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189, 2022. URL https://arxiv.org/abs/2203.17189.
  39. 39.Rudinger, R., Shwartz, V., Hwang, J. D., Bhagavatula, C., Forbes, M., Le Bras, R., Smith, N. A., and Choi, Y. Thinking like a skeptic: Defeasible inference in natural language. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 4661–4675, 2020.
  40. 40.Saeed, M., Ahmadi, N., Nakov, P., and Papotti, P. RuleBERT: Teaching soft rules to pre-trained language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1460–1476, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.110. URL https://aclanthology.org/2021.emnlp-main.110.
  41. 41.Saparov, A. and He, H. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=qFVVBzXxR2V.
  42. 42.Saparov, A., Yuanzhe Pang, R., Padmakumar, V., Joshi, N., Kazemi, S. M., Kim, N., and He, H. Testing the general deductive reasoning capacity of large language models using ood examples. arxiv preprint arXiv:2305.15269, 2023.
  43. 43.Sartor, G. Defeasibility in legal reasoning. Springer, 1995.
  44. 44.Shoenfield, J. Mathematical Logic. Taylor & Francis, 2001. ISBN 9781568811352. URL https://books.google.com/books?id=a9zuAAAAMAAJ.
  45. 45.Sinha, K., Sodhani, S., Dong, J., Pineau, J., and Hamilton, W. L. Clutrr: A diagnostic benchmark for inductive reasoning from text. arXiv preprint arXiv:1908.06177, 2019.
  46. 46.Sprague, Z., Bostrom, K., Chaudhuri, S., and Durrett, G. Natural language deduction with incomplete information. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 8230–8258, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.564.
  47. 47.Sun, H., Cohen, W. W., and Salakhutdinov, R. Conditionalqa: A complex reading comprehension dataset with conditional answers. arXiv preprint arXiv:2110.06884, 2021.
  48. 48.Tafjord, O., Dalvi, B., and Clark, P. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 3621–3634, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.317. URL https://aclanthology.org/2021.findings-acl.317.
  49. 49.Talmor, A., Herzig, J., Lourie, N., and Berant, J. Commonsenseqa: A question answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937, 2018.
  50. 50.Talmor, A., Tafjord, O., Clark, P., Goldberg, Y., and Berant, J. Leap-of-thought: Teaching pre-trained models to systematically reason over implicit knowledge. Advances in Neural Information Processing Systems, 33:20227–20237, 2020.
  51. 51.Wang, B., Deng, X., and Sun, H. Iteratively prompt pre-trained language models for chain of thought. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 2714–2730, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.174.
  52. 52.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 24824–24837. Curran Associates, Inc., 2022.
  53. 53.Weston, J., Bordes, A., Chopra, S., Rush, A. M., Van Merriënboer, B., Joulin, A., and Mikolov, T. Towards ai-complete question answering: A set of prerequisite toy tasks. arXiv preprint arXiv:1502.05698, 2015.
  54. 54.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601, 2023.
  55. 55.Ye, A., Cui, C., Shi, T., and Riedl, M. O. Neural story planning. arXiv preprint arXiv:2212.08718, 2022.
  56. 56.Yu, F., Zhang, H., and Wang, B. Nature language reasoning, a survey. arXiv preprint arXiv:2303.14725, 2023.
  57. 57.Yuan, Q., Kazemi, M., Xu, X., Noble, I., Imbrasaite, V., and Ramachandran, D. Tasklama: Probing the complex task understanding of language models. arXiv preprint arXiv:2308.15299, 2023.
  58. 58.Zelikman, E., Wu, Y., Mu, J., and Goodman, N. STaR: Bootstrapping reasoning with reasoning. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp. 15476–15488. Curran Associates, Inc., 2022.
  59. 59.Zhang, H., Li, L. H., Meng, T., Chang, K.-W., and Broeck, G. V. d. On the paradox of learning to reason from data. arXiv:2205.11502, 2022. URL https://arxiv.org/abs/2205.11502.
  60. 60.Zhang, H., Li, Z., Huang, J., Naik, M., and Xing, E. Improved logical reasoning of language models via differentiable symbolic programming. In First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022, 2022. URL https://openreview.net/forum?id=8lNy3QCaxHX.
  61. 61.Zhong, W., Wang, S., Tang, D., Xu, Z., Guo, D., Wang, J., Yin, J., Zhou, M., and Duan, N. AR-LSAT: Investigating analytical reasoning of text. arXiv preprint arXiv:2104.06598, 2021.

Citation

MLA
Kazemi, M., et al. “BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information”. Advances in Neural Information Processing Systems 36, 2023, pp. 39052–74, https://doi.org/10.52202/075280-1697.
APA
Kazemi, M., Yuan, Q., Bhatia, D., Kim, N., Xu, X., Imbrasaite, V., & Ramachandran, D. (2023). BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information. Advances in Neural Information Processing Systems 36, 39052–39074. https://doi.org/10.52202/075280-1697
Chicago
Kazemi, M., Q. Yuan, D. Bhatia, et al. 2023. “BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information”. Advances in Neural Information Processing Systems 36, 39052–74. https://doi.org/10.52202/075280-1697.
Harvard
Kazemi, M. et al. (2023) “BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information”, Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp. 39052–39074. Available at: https://doi.org/10.52202/075280-1697.
Vancouver
1. Kazemi M, Yuan Q, Bhatia D, Kim N, Xu X, Imbrasaite V, Ramachandran D (2023) BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information. In: Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp 39052–39074

BibTeX

@inproceedings{Kazemi_2023, series={NeurIPS 2023}, title={BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information}, url={http://dx.doi.org/10.52202/075280-1697}, DOI={10.52202/075280-1697}, booktitle={Advances in Neural Information Processing Systems 36}, publisher={Neural Information Processing Systems Foundation, Inc. (NeurIPS)}, author={Kazemi, Mehran and Yuan, Quan and Bhatia, Deepti and Kim, Najoung and Xu, Xin and Imbrasaite, Vaiva and Ramachandran, Deepak}, year={2023}, pages={39052–39074}, collection={NeurIPS 2023} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/