Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations
Jaehun JungLianhui QinSean WelleckFaeze BrahmanChandra BhagavatulaRonan Le BrasYejin Choi
Proposes an unsupervised prompting method that recursively generates trees of abductive explanations and resolves their logical inconsistencies with a satisfiability solver, improving commonsense question-answering accuracy by up to 20% over standard prompting baselines.
Large language models often struggle with logical consistency and commonsense reasoning when answering complex questions. While prompting models to generate step-by-step explanations has shown promise, these generated rationales are frequently noisy, factually inaccurate, or self-contradictory, ultimately misleading the model's final conclusions.
The article introduces and evaluates "Maieutic Prompting," a novel unsupervised inference method designed to derive accurate, logically consistent answers from potentially unreliable model-generated explanations. Inspired by the Socratic method, the approach aims to eliminate contradictory hypotheses and ensure robust reasoning without requiring labeled task training data.
The approach constructs a tree of explanations by prompting a language model to abductively explain both possible outcomes (true and false) and recursively evaluating deeper reasoning paths. Branches are pruned unless they reach logically integral propositions—statements where the model reliably distinguishes between a assertion and its negation. The logical relationships and beliefs across all generated statements are then converted into logical constraints and solved using a weighted maximum satisfiability solver, which identifies the most coherent set of true statements.
The evaluation yields several key findings across benchmark datasets including Com2Sense, CSQA 2.0, and CREAK. First, Maieutic Prompting improves accuracy by up to 20% over state-of-the-art prompting techniques such as Chain of Thought and Self-Consistency. Second, as a fully unsupervised method using an off-the-shelf model, it performs competitively with, and in several cases outperforms, large supervised fine-tuned models with billions of parameters. Third, the framework demonstrates significantly higher resilience to semantic perturbations, such as paired opposite statements, and shows greater stability against variations in prompt wording and order. Finally, expert human evaluations confirm that the extracted explanations provide high grammatical quality, factual relevance, and interpretable decision rationales.
These results demonstrate that neuro-symbolic reasoning can substantially mitigate the unreliability of generative AI without costly dataset curation or model fine-tuning. For decision-makers, this offers an effective way to deploy general-purpose language models in high-stakes reasoning and verification environments while reducing the risk of silent logical failures and enhancing auditability.
Organizations seeking to improve the reliability of automated reasoning systems should adopt multi-depth explanation verification and symbolic constraint solving rather than relying on raw single-hop model outputs. Future development should focus on extending this framework from binary true/false verification to multi-choice and open-ended question formats, as well as modeling shared knowledge graphs across multiple related questions.
Key limitations include computational overhead, as recursive tree expansion requires multiple model queries, though this is partially mitigated by depth-adaptive pruning. Additionally, the method's effectiveness relies on an external logical solver or natural language inference verifier to map inter-statement relationships, and current empirical validations remain bounded to statement-verification tasks.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It introduces foundational chain-of-thought prompting, which Maieutic Prompting directly builds upon and modifies to correct unreliable intermediate rationales.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). It establishes self-consistency sampling over model reasoning paths, providing an essential baseline and concept of logical consistency that Maieutic Prompting extends via formal constraint satisfaction.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). It establishes zero-shot step-by-step reasoning in language models, setting the stage for unsupervised recursive explanation generation.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). It introduces hierarchical problem decomposition via sequential prompting, a core prerequisite conceptualized recursively in Maieutic Prompting's explanation trees.
- Paper: Discovering Latent Knowledge in Language Models Without Supervision, Collin Burns et al. (2023). It demonstrates unsupervised extraction of latent knowledge by enforcing logical consistency between assertions and their negations, directly informing Maieutic Prompting's premise verification.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). It generalizes tree-structured prompting methods like Maieutic Prompting into arbitrary graph topologies supporting aggregation, feedback loops, and dynamic network pruning.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). It provides a rigorous formal diagnostic framework to evaluate the local validity versus global planning failures of step-by-step reasoning trees.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It investigates the unfaithfulness and rationalization flaws in generated explanations that recursive, constraint-based methods seek to address.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). It builds upon the formal verification of explanation steps by introducing a reference-free framework to evaluate intermediate reasoning correctness and informativeness.
- Paper: Improving Factuality and Reasoning in Language Models through Multiagent Debate, Yilun Du et al. (2023). It extends Socratic self-consistency and abductive inquiry into multi-agent debate frameworks for iterative error correction.
- Paper: Consistency Analysis of ChatGPT, Myeongjun Jang et al. (2023). It provides a systematic consistency benchmark analyzing negation and symmetric reasoning vulnerabilities in language models.
