Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations

Jaehun JungLianhui QinSean WelleckFaeze BrahmanChandra BhagavatulaRonan Le BrasYejin Choi

article2022EMNLP231 citations

Proposes an unsupervised prompting method that recursively generates trees of abductive explanations and resolves their logical inconsistencies with a satisfiability solver, improving commonsense question-answering accuracy by up to 20% over standard prompting baselines.

Listen

Large language models often struggle with logical consistency and commonsense reasoning when answering complex questions. While prompting models to generate step-by-step explanations has shown promise, these generated rationales are frequently noisy, factually inaccurate, or self-contradictory, ultimately misleading the model's final conclusions.

The article introduces and evaluates "Maieutic Prompting," a novel unsupervised inference method designed to derive accurate, logically consistent answers from potentially unreliable model-generated explanations. Inspired by the Socratic method, the approach aims to eliminate contradictory hypotheses and ensure robust reasoning without requiring labeled task training data.

The approach constructs a tree of explanations by prompting a language model to abductively explain both possible outcomes (true and false) and recursively evaluating deeper reasoning paths. Branches are pruned unless they reach logically integral propositions—statements where the model reliably distinguishes between a assertion and its negation. The logical relationships and beliefs across all generated statements are then converted into logical constraints and solved using a weighted maximum satisfiability solver, which identifies the most coherent set of true statements.

The evaluation yields several key findings across benchmark datasets including Com2Sense, CSQA 2.0, and CREAK. First, Maieutic Prompting improves accuracy by up to 20% over state-of-the-art prompting techniques such as Chain of Thought and Self-Consistency. Second, as a fully unsupervised method using an off-the-shelf model, it performs competitively with, and in several cases outperforms, large supervised fine-tuned models with billions of parameters. Third, the framework demonstrates significantly higher resilience to semantic perturbations, such as paired opposite statements, and shows greater stability against variations in prompt wording and order. Finally, expert human evaluations confirm that the extracted explanations provide high grammatical quality, factual relevance, and interpretable decision rationales.

These results demonstrate that neuro-symbolic reasoning can substantially mitigate the unreliability of generative AI without costly dataset curation or model fine-tuning. For decision-makers, this offers an effective way to deploy general-purpose language models in high-stakes reasoning and verification environments while reducing the risk of silent logical failures and enhancing auditability.

Organizations seeking to improve the reliability of automated reasoning systems should adopt multi-depth explanation verification and symbolic constraint solving rather than relying on raw single-hop model outputs. Future development should focus on extending this framework from binary true/false verification to multi-choice and open-ended question formats, as well as modeling shared knowledge graphs across multiple related questions.

Key limitations include computational overhead, as recursive tree expansion requires multiple model queries, though this is partially mitigated by depth-adaptive pruning. Additionally, the method's effectiveness relies on an external logical solver or natural language inference verifier to map inter-statement relationships, and current empirical validations remain bounded to statement-verification tasks.

Cover for Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations

Abstract

Pre-trained language models (LMs) struggle with consistent reasoning; recently, prompting LMs to generate explanations that self-guide the inference has emerged as a promising direction to amend this. However, these approaches are fundamentally bounded by the correctness of explanations, which themselves are often noisy and inconsistent. In this work, we develop MAIEUTIC PROMPTING, which aims to infer a correct answer to a question even from the unreliable generations of LM. MAIEUTIC PROMPTING induces a tree of explanations abductively (e.g. X is true, because . . .) and recursively, then frames the inference as a satisfiability problem over these explanations and their logical relations. We test MAIEUTIC PROMPTING for true/false QA on three challenging benchmarks that require complex commonsense reasoning. MAIEUTIC PROMPTING achieves up to 20% better accuracy than state-of-the-art prompting methods, and as a fully unsupervised approach, performs competitively with supervised models. We also show that MAIEUTIC PROMPTING improves robustness in inference while providing interpretable rationales.

Table of Contents

  • 1 Introduction
  • 2 Problem Setup and Background
  • 3 Maieutic Prompting
  • 3.1 Maieutic Tree Generation
  • 3.1.1 Abductive Explanation Generation
  • 3.1.2 Depth-wise Knowledge Spanning
  • 3.1.3 When to Stop Generating
  • 3.2 Defining the Relations
  • 3.3 Inference
  • 3.4 Verifier Model
  • 4 Experiments
  • 4.2 Robustness Analysis
  • 4.3 Ablation Study
  • 4.4 Human Evaluation
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Tree Generation Algorithm
  • B Dataset Details
  • C Multi-hop Reasoning on StrategyQA
  • D Inference Examples

Knowls

  1. Knowl 1 — Maieutic Prompting Framework

    model/method

    Maieutic Prompting is an unsupervised, few-shot reasoning framework designed to infer the truth value of a statement Q∈{True,False}Q \in \{\text{True}, \text{False}\} by generating a structured, multi-depth tree of explanations (a maieutic tree) and resolving contradictions using symbolic optimization.

    The framework proceeds in three stages:

    1. Abductive and Recursive Tree Generation: Given a query statement QQ, a pre-trained language model generates abductive rationales for both candidate answers (QQ is True because ETE_T; QQ is False because EFE_F). To validate these 1-hop explanations, the model recursively prompts itself with its own generated explanations as new questions, expanding the tree depth-wise until it reaches logically integral propositions.
    2. Relation Quantification: The logical relationships among generated propositions are formalized into unary belief weights (the model's confidence in individual propositions) and binary consistency constraints (logical implications between parents and children or across arbitrary proposition pairs via an NLI verifier).
    3. Symbolic Inference via MAX-SAT: The truth assignments for all propositions and the root question QQ are jointly determined by solving a weighted Maximum Satisfiability (MAX-SAT) problem that maximizes the aggregate weight of satisfied belief and consistency clauses.
  2. Knowl 2 — Logical Integrity of Language Model Propositions

    definition

    A proposition EE is defined as logically integral with respect to a language model pLMp_{\text{LM}} and in-context demonstration examples CC if the model consistently assigns opposite truth values to EE and its logical negation ¬E\neg E.

    Formally, the boolean indicator function integral(E)∈{0,1}\text{integral}(E) \in \{0, 1\} is defined as:

    integral(E)=I(Condition 1∨Condition 2)\text{integral}(E) = \mathbb{I}\left(\text{Condition 1} \lor \text{Condition 2}\right)

    where the conditions are:

    1. Logically Integral / True: arg⁡max⁡A∈{T,F}pLM(A∣E,C)=Tandarg⁡max⁡A∈{T,F}pLM(A∣¬E,C)=F\arg\max_{A \in \{T, F\}} p_{\text{LM}}(A \mid E, C) = T \quad \text{and} \quad \arg\max_{A \in \{T, F\}} p_{\text{LM}}(A \mid \neg E, C) = F
    2. Logically Integral / False: arg⁡max⁡A∈{T,F}pLM(A∣E,C)=Fandarg⁡max⁡A∈{T,F}pLM(A∣¬E,C)=T\arg\max_{A \in \{T, F\}} p_{\text{LM}}(A \mid E, C) = F \quad \text{and} \quad \arg\max_{A \in \{T, F\}} p_{\text{LM}}(A \mid \neg E, C) = T

    Propositions that receive identical truth-value predictions for both EE and ¬E\neg E fail the test (integral(E)=0\text{integral}(E) = 0), signaling that the language model is inconsistent under negation and that the proposition should not be relied upon without further recursive validation.

  3. Knowl 3 — Maieutic Tree Generation and Pruning Algorithm

    algorithm

    The maieutic tree generation procedure constructs a tree T\mathcal{T} of recursive explanations up to a maximum depth DD and prunes non-integral leaf branches.

    Input: Statement QQ, maximum tree depth DD
    Output: Maieutic tree T\mathcal{T}
    T←init(Q)\mathcal{T} \leftarrow \text{init}(Q)
    S0←{Q}S_0 \leftarrow \{Q\}
    for d=1d = 1 to DD do
        Sd←∅S_d \leftarrow \emptyset
        for E∈Sd−1E \in S_{d-1} do
            if integral(E)=0\text{integral}(E) = 0 then
                Sd←Sd∪abduction(E)S_d \leftarrow S_d \cup \text{abduction}(E)
            end if
        end for
        T.add(Sd)\mathcal{T}.\text{add}(S_d)
    end for
    V←{E∈leaf(T)∣integral(E)=0}V \leftarrow \{E \in \text{leaf}(\mathcal{T}) \mid \text{integral}(E) = 0\}
    while V≠∅V \neq \emptyset do
        T.remove(V)\mathcal{T}.\text{remove}(V)
        V←{E∈leaf(T)∣integral(E)=0}V \leftarrow \{E \in \text{leaf}(\mathcal{T}) \mid \text{integral}(E) = 0\}
    end while
    return T\mathcal{T}

    Operational Details:

    • Abduction Function: abduction(E)=(ET,EF)\text{abduction}(E) = (E_T, E_F), where EA∼pLM(E∣Q=E,A,C)E_A \sim p_{\text{LM}}(E \mid Q=E, A, C) for A∈{T,F}A \in \{T, F\}.
    • Depth-Adaptive Decoding: For depth 1, nucleus sampling (p=1.0p = 1.0) generates multiple explanations per answer (e.g., 3 ETE_T and 3 EFE_F). For depth 2, greedy decoding generates 1 ETE_T and 1 EFE_F per parent node, bounding the tree size to at most 18 nodes excluding the root QQ.
    • Stopping and Pruning Criterion: Recursive expansion halts along any branch as soon as a logically integral proposition (integral(E)=1\text{integral}(E) = 1) is reached. After reaching maximum depth D=2D = 2, all remaining non-integral leaves and their unsupported ancestral sub-branches are pruned.
  4. Knowl 4 — MAX-SAT Objective for Joint Truth-Value Inference

    equation

    Given a maieutic tree T\mathcal{T}, the truth assignment to all propositions E∈node(T)E \in \text{node}(\mathcal{T}) and the root statement QQ is computed by maximizing the sum of weights of satisfied unary belief clauses Cblf\mathcal{C}_{\text{blf}} and binary consistency/implication clauses Ccon\mathcal{C}_{\text{con}}:

    max⁡v∈{0,1}∣node(T)∣∑c∈Cblf∪Cconwc⋅I(c=True)\max_{\mathbf{v} \in \{0, 1\}^{|\text{node}(\mathcal{T})|}} \sum_{c \in \mathcal{C}_{\text{blf}} \cup \mathcal{C}_{\text{con}}} w_c \cdot \mathbb{I}(c = \text{True})

    where:

    • Unary Belief Clauses (defined for each leaf node E∈leaf(T)E \in \text{leaf}(\mathcal{T})): cblf={Eif E is logically integral / True¬Eif E is logically integral / Falsec_{\text{blf}} = \begin{cases} E & \text{if } E \text{ is logically integral / True} \\ \neg E & \text{if } E \text{ is logically integral / False} \end{cases} with weight wcblf=wEw_{c_{\text{blf}}} = w_E.
    • Binary Consistency Clauses (defined for each directed edge (Eparent,Echild)(E_{\text{parent}}, E_{\text{child}}) in T\mathcal{T} where EchildE_{\text{child}} rationalizes label AA for EparentE_{\text{parent}}): ccon={Echild→Eparentif A=TrueEchild→¬Eparentif A=Falsec_{\text{con}} = \begin{cases} E_{\text{child}} \to E_{\text{parent}} & \text{if } A = \text{True} \\ E_{\text{child}} \to \neg E_{\text{parent}} & \text{if } A = \text{False} \end{cases} with weight wccon=wEchild,Eparent,Aw_{c_{\text{con}}} = w_{E_{\text{child}}, E_{\text{parent}}, A}.
  5. Knowl 5 — Belief and Generation-Consistency Weight Metrics

    equation

    To assign weights to clauses in the weighted MAX-SAT program, belief and consistency are quantified via language model token probabilities:

    1. Propositional Belief Weight (wEw_E) measures the model's confidence in proposition EE by comparing the probability assigned to label True\text{True} given EE versus ¬E\neg E: wE:=pLM(T∣E,C)−pLM(T∣¬E,C)pLM(T∣E,C)+pLM(T∣¬E,C)w_E := \frac{p_{\text{LM}}(T \mid E, C) - p_{\text{LM}}(T \mid \neg E, C)}{p_{\text{LM}}(T \mid E, C) + p_{\text{LM}}(T \mid \neg E, C)} where wE∈[−1,1]w_E \in [-1, 1] (with positive values indicating belief in EE, negative in ¬E\neg E).

    2. Likelihood-Based Consistency Weight (wE,Q,Aw_{E, Q, A}) measures whether explanation EE is more probable under answer hypothesis AA than under ¬A\neg A: wE,Q,A:=pLM(E∣Q,A,C)pLM(E∣Q,A,C)+pLM(E∣Q,¬A,C)w_{E, Q, A} := \frac{p_{\text{LM}}(E \mid Q, A, C)}{p_{\text{LM}}(E \mid Q, A, C) + p_{\text{LM}}(E \mid Q, \neg A, C)} where wE,Q,A∈[0,1]w_{E, Q, A} \in [0, 1].

  6. Knowl 6 — Cross-Branch NLI Verifier for Inter-Node Relations

    model/method

    To capture logical relationships across different branches of the maieutic tree (beyond direct parent-child likelihood edges), a Natural Language Inference (NLI) model is used as an auxiliary verifier.

    For every distinct pair of propositions (E1,E2)∈node(T)2(E_1, E_2) \in \text{node}(\mathcal{T})^2 (E1≠E2E_1 \neq E_2), the verifier predicts an NLI label, which is mapped into propositional clauses CNLI\mathcal{C}_{\text{NLI}}: Entail(E1,E2)  ⟹  E1→E2\text{Entail}(E_1, E_2) \implies E_1 \to E_2 Contradict(E1,E2)  ⟹  E1→¬E2\text{Contradict}(E_1, E_2) \implies E_1 \to \neg E_2

    Each NLI clause is assigned a fixed weight of w=1.0w = 1.0, replacing the edge-only consistency set Ccon\mathcal{C}_{\text{con}} with CNLI\mathcal{C}_{\text{NLI}} in the MAX-SAT optimization problem. In standard configurations, a RoBERTa model fine-tuned on MNLI (achieving 90.2% accuracy on the MNLI dev set) is used as the verifier.

  7. Knowl 7 — Benchmark Accuracy on Commonsense and Fact Verification Datasets

    data/table

    Maieutic Prompting was evaluated using GPT-3 (text-davinci-001, 6-shot) against standard prompting, explanation-based prompting baselines (Chain of Thought, Self-Consistency with N=20N=20, Generated Knowledge Prompting with N=20N=20), and supervised models on three binary QA benchmarks: Com2Sense, CSQA 2.0, and CREAK.

    Dataset Com2Sense CSQA 2.0 CREAK
    Model dev test pairwise dev test dev test contrast
    Supervised
    RoBERTa-large 62.8 59.4 33.3 - - 80.6 80.3 61.5
    T5-large 62.8 60.6 41.8 53.8 54.6 - - -
    T5-3B 73.2 - - - 60.2 85.6 85.1 70.0
    UnifiedQA-3B 75.1 71.3 51.3 - - - - -
    T5-11B 77.2 - - 68.5 67.8 89.5 - 75.2
    Unicorn-11B - - - 69.9 70.2 - - -
    Prompting (GPT-3)
    Standard 58.1 - - 54.1 - 60.3 - 55.2
    Chain of Thought 61.6 - - 59.6 - 64.8 - 59.4
    Self Consistency 61.4 - - 60.8 - 70.5 - 64.8
    GKP 61.8 - - 59.7 - 75.4 - 68.2
    Maieutic Prompting 72.5 75.0 68.7 69.5 68.3 85.2 85.3 77.4

    Maieutic Prompting outperforms all prompting baselines across all benchmarks (up to +13.4% on CSQA 2.0 and +14.7% on CREAK dev over Chain of Thought). It also matches or exceeds competitive billion-scale supervised models (e.g., outperforming UnifiedQA-3B test accuracy on Com2Sense and T5-11B test accuracy on CSQA 2.0).

  8. Knowl 8 — Ablation of Maieutic Prompting Components and Tree Topologies

    data/table

    Ablation experiments conducted on the Com2Sense development set analyze the impact of abductive generation, decoding strategies, consistency formulations, and tree dimensions.

    Model Variant Accuracy (%)
    Non-abductive generation 68.4
    All greedy decoding (no depth-adaptive) 67.2
    All nucleus sampling (no depth-adaptive) 72.0
    Likelihood-based consistency 65.6
    Maieutic Prompting (Full) 72.5
    Dimension 1 2 3 5 10
    Depth 61.3 72.5 72.4 - -
    Width 62.4 66.5 72.5 71.5 72.1

    Key Findings:

    • Replacing abductive generation with non-abductive explanations degrades accuracy by 4.1%.
    • Using NLI-based consistency provides a 6.9% improvement over likelihood-based consistency.
    • Depth-adaptive decoding (sampling at depth 1, greedy at depth 2) outperforms purely greedy decoding (67.2%) while remaining more computationally efficient than pure nucleus sampling.
    • Scaling tree depth beyond 2 or width beyond 3 yields diminishing returns due to topic drift at deeper hops and knowledge redundancy at higher widths.
  9. Knowl 9 — Robustness to Semantic Contrast and Prompt Sensitivity

    empirical result

    Maieutic Prompting demonstrates enhanced robustness compared to other prompting baselines under semantic perturbations and few-shot prompt variations:

    1. Semantic Perturbations (Pairwise & Contrast Sets): On the Com2Sense test set (which evaluates complementary sentence pairs with opposite ground-truth labels), Maieutic Prompting achieves 68.7% pairwise accuracy, outperforming UnifiedQA-3B (51.3%) and RoBERTa-large (33.3%). On the CREAK contrast set, it achieves 77.4% accuracy, exceeding T5-11B (75.2%) and all prompting baselines (55.2% to 68.2%).
    2. Prompt Example & Order Sensitivity: When evaluated over 3 distinct sets of few-shot demonstration examples, Maieutic Prompting achieves an accuracy of 72.34%±σ72.34\% \pm \sigma on Com2Sense dev, maintaining higher mean and lower variance than Standard (57.88%), Chain of Thought (60.67%), and Self-Consistency (61.20%). Across 5 prompt order permutations, it achieves 71.68% mean accuracy, demonstrating stability against prompt ordering.
  10. Knowl 10 — Multi-Hop Reasoning on StrategyQA via Question Decomposition

    empirical result

    On the StrategyQA development set (a benchmark requiring implicit multi-hop reasoning), Maieutic Prompting was tested in both standard zero-decomposition and multi-hop decomposition settings (where the question is first decomposed into 2-3 sub-questions):

    • Standard Setup (no explicit decomposition):
      • Standard Prompting: 56.3%
      • Chain of Thought: 58.2%
      • Maieutic Prompting: 60.7%
    • Multi-Hop Setup (with sub-question decomposition):
      • Chain of Thought (Multi-hop): 57.9%
      • Maieutic Prompting (Multi-hop): 61.4%

    Maieutic Prompting consistently outperforms standard and Chain of Thought prompting in both settings.

  11. Knowl 11 — Limitations of Maieutic Prompting

    limitation

    The authors identify two key limitations of the Maieutic Prompting method:

    1. Task Format Restriction: The framework is formulated and evaluated strictly on binary (True/False) statement validation tasks. Applying it to multi-choice QA or open-ended generation requires adapting choices into individual binary propositions and aggregating scores (e.g., via satisfied MAX-SAT clause weights).
    2. Isolated Tree Representation: Knowledge relations are modeled strictly within the tree constructed for a single question. The current framework does not support cross-tree relations where knowledge generated for one question can be transferred or composed to validate propositions for another question.

Coverage note — No substantial contributed material was omitted; the knowls cover the full algorithm, mathematical formulations of integrity/weights/MAX-SAT, NLI verifier, main benchmark results, ablations, robustness studies, multi-hop experiments, and limitations.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977.
  2. 2.Roberto Battiti. 2009. Maximum satisfiability problem, pages 2035–2041. Springer US, Boston, MA.
  3. 3.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, and Yejin Choi. 2019. Abductive commonsense reasoning. arXiv preprint arXiv:1908.05739.
  4. 4.Kaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, and Greg Durrett. 2021. Flexible generation of natural language deductions. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6266–6278.
  5. 5.Faeze Brahman, Vered Shwartz, Rachel Rudinger, and Yejin Choi. 2021. Learning to rationalize for nonmonotonic reasoning with distant supervision. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 12592–12601. AAAI Press.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  7. 7.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. Advances in Neural Information Processing Systems, 31.
  8. 8.Howard Chen, Jacqueline He, Karthik Narasimhan, and Danqi Chen. 2022. Can rationalization improve robustness? arXiv preprint arXiv:2204.11790.
  9. 9.Xinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou, Dawn Song, and Quoc V Le. 2019. Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension. In International Conference on Learning Representations.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  11. 11.Bhavana Dalvi, Oyvind Tafjord, and Peter Clark. 2022. Towards teachable reasoning systems. arXiv preprint arXiv:2204.13074.
  12. 12.Mengnan Du, Ninghao Liu, and Xia Hu. 2019. Techniques for interpretable machine learning. Communications of the ACM, 63(1):68–77.
  13. 13.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transactions of the Association for Computational Linguistics, 9:1012–1031.
  14. 14.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  15. 15.Robert G. M. Hausmann and Kurt VanLehn. 2007. Explaining self-explaining: A contrast between content and generation. In AIED.
  16. 16.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In International Conference on Learning Representations.
  17. 17.Ari Holtzman, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7038–7051.
  18. 18.Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021. Contrastive explanations for model interpretability. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 1597–1611.
  19. 19.Nora Kassner and Hinrich Schütze. 2020. Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7811–7818, Online. Association for Computational Linguistics.
  20. 20.Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark. 2021. Beliefbank: Adding memory to a pre-trained language model for a systematic notion of belief. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8849–8861.
  21. 21.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. Unifiedqa: Crossing format boundaries with a single qa system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907.
  22. 22.Jannik Kossen, Neil Band, Clare Lyle, Aidan N Gomez, Thomas Rainforth, and Yarin Gal. 2021. Selfattention between datapoints: Going beyond individual input-output pairs in deep learning. Advances in Neural Information Processing Systems, 34:28742–28756.
  23. 23.Klaus Krippendorff. 2007. Computing krippendorff’s alpha-reliability. annenberg school for communication departmental paper 43.
  24. 24.Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill. 2022. Can language models learn from explanations in context? arXiv preprint arXiv:2204.02329.
  25. 25.Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras, Yejin Choi, and Hannaneh Hajishirzi. 2021. Generated knowledge prompting for commonsense reasoning. arXiv preprint arXiv:2110.08387.
  26. 26.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  27. 27.Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. Unicorn on rainbow: A universal commonsense reasoning model on a new multitask benchmark. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13480–13488.
  28. 28.Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021a. Neurologic decoding:(un) supervised neural text generation with predicate logic constraints. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4299.
  29. 29.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021b. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786.
  30. 30.Pasquale Minervini and Sebastian Riedel. 2018. Adversarially regularising neural nli models to integrate logical background knowledge. arXiv preprint arXiv:1808.08609.
  31. 31.António Morgado, Carmine Dodaro, and Joao Marques-Silva. 2014. Core-guided maxsat with soft cardinality constraints. In International Conference on Principles and Practice of Constraint Programming, pages 564–573. Springer.
  32. 32.Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020. Wt5?! training text-to-text models to explain their predictions. arXiv preprint arXiv:2004.14546.
  33. 33.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021a. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114.
  34. 34.Maxwell Nye, Michael Tessler, Josh Tenenbaum, and Brenden M Lake. 2021b. Improving coherence and consistency in neural sequence models with dualsystem, neuro-symbolic reasoning. Advances in Neural Information Processing Systems, 34.
  35. 35.Yasumasa Onoe, Michael J.Q. Zhang, Eunsol Choi, and Greg Durrett. 2021. Creak: A dataset for commonsense reasoning over entity knowledge. OpenReview.
  36. 36.Charles Sanders Peirce. 1974. Collected papers of charles sanders peirce, volume 5. Harvard University Press.
  37. 37.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  38. 38.Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019. Explain yourself! leveraging language models for commonsense reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4932–4942.
  39. 39.Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. Unsupervised commonsense question answering with self-talk. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4615–4629.
  40. 40.Shikhar Singh, Nuan Wen, Yu Hou, Pegah Alipoormolabashi, Te-lin Wu, Xuezhe Ma, and Nanyun Peng. 2021. COM2SENSE: A commonsense reasoning benchmark with complementary sentences. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 883–898, Online. Association for Computational Linguistics.
  41. 41.Alon Talmor, Ori Yoran, Ronan Le Bras, Chandra Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant. 2021. CommonsenseQA 2.0: Exposing the limits of AI through gamification. In Thirty-fifth Conference on Neural Information Processing Systems.
  42. 42.Gregory Vlastos. 1991. Socrates, ironist and moral philosopher, volume 50. Cornell University Press.
  43. 43.Haohan Wang, Da Sun, and Eric P Xing. 2019. What if we simply swap the two text fragments? a straightforward yet effective way to test the robustness of methods to confounding signals in nature language inference tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 7136–7143.
  44. 44.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  45. 45.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  46. 46.Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, and Kyunghyun Cho. 2020. Consistency of a recurrent language model with respect to incomplete decoding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5553–5568.
  47. 47.Sarah Wiegreffe and Ana Marasović. 2021. Teach me to explain: A review of datasets for explainable nlp. arXiv preprint arXiv:2102.12060.
  48. 48.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122.
  49. 49.Xi Ye and Greg Durrett. 2022. The unreliability of explanations in few-shot in-context learning. arXiv preprint arxiv:2205.03401.
  50. 50.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.

Citation

MLA
Jung, J., et al. “Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 1266–79, https://doi.org/10.18653/v1/2022.emnlp-main.82.
APA
Jung, J., Qin, L., Welleck, S., Brahman, F., Bhagavatula, C., Bras, R. L., & Choi, Y. (2022). Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1266–1279. https://doi.org/10.18653/v1/2022.emnlp-main.82
Chicago
Jung, J., L. Qin, S. Welleck, et al. 2022. “Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1266–79. https://doi.org/10.18653/v1/2022.emnlp-main.82.
Harvard
Jung, J. et al. (2022) “Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1266–1279. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.82.
Vancouver
1. Jung J, Qin L, Welleck S, Brahman F, Bhagavatula C, Bras RL, Choi Y (2022) Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1266–1279

BibTeX

@inproceedings{jung-etal-2022-maieutic,
    title = "Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations",
    author = "Jung, Jaehun  and
      Qin, Lianhui  and
      Welleck, Sean  and
      Brahman, Faeze  and
      Bhagavatula, Chandra  and
      Le Bras, Ronan  and
      Choi, Yejin",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.82/",
    doi = "10.18653/v1/2022.emnlp-main.82",
    pages = "1266--1279"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/