Successive Prompting for Decomposing Complex Questions
Dheeru DuaShivanshu GuptaSameer SinghMatt Gardner
Proposes an iterative prompting method that decouples question decomposition from question answering to enable step-specific in-context learning and synthetic bootstrapping for multi-step reasoning.
Organizations increasingly rely on automated language systems to analyze unstructured texts and answer complex questions that require multi-step reasoning. However, standard language models struggle to perform these complex reasoning tasks accurately when limited training data is available. Existing approaches often force a single model to generate all intermediate reasoning steps and the final answer in one continuous pass, which limits flexibility and prevents the integration of specialized tools for tasks like arithmetic.
The article evaluates "Successive Prompting," a framework designed to iteratively break complex questions into simple question-answer pairs until reaching a final solution. The objective is to demonstrate that decoupling the question decomposition process from the question answering process improves multi-step reasoning performance in data-constrained environments.
To test this approach, the researchers conducted experiments on the DROP reading comprehension benchmark using a challenging few-shot setup with only 300 annotated examples. They evaluated both prompt-based prompting on a large language model and fine-tuned modular models. To address the scarcity of intermediate training data, the authors developed a synthetic data generator using structured Wikipedia tables, producing roughly 141,000 multi-step reasoning questions covering 10 distinct operations.
The analysis revealed several key findings. First, the best fine-tuned Successive Prompting model achieved a 50.2% test accuracy score, outperforming the previous state-of-the-art benchmark by an absolute 5.1 percentage points under the same limited supervision. Second, pre-training models on out-of-domain synthetic data universally improved performance across all baseline architectures, boosting one baseline by nearly 20 percentage points. Third, when using few-shot in-context learning without fine-tuning, Successive Prompting outperformed single-pass Chain-of-Thought prompting by 4.3 percentage points on the development set. Finally, delegating mathematical sub-tasks to a deterministic symbolic calculator provided an immediate 1.5 percentage point performance boost over relying on the language model alone.
These findings indicate that modular, iterative architectures provide greater accuracy and interpretability than monolithic, single-pass language models for complex analytical tasks. By isolating individual reasoning steps, organizations can target training data to specific weaknesses and plug in deterministic computational engines for arithmetic operations, thereby reducing the risk of calculation errors. However, decision-makers should note that Successive Prompting increases operational computational costs and latency, as answering a single question requires multiple sequential model queries rather than a single pass.
For practical implementation, organizations developing complex question-answering pipelines should adopt modular frameworks that separate reasoning breakdown from factual answering and integrate external symbolic engines for mathematical tasks. Teams should also utilize synthetic data generation to bootstrap capabilities before collecting costly human annotations. Further development is recommended to extend synthetic data generators to cover implicit, causal, and common-sense reasoning patterns, which were identified as key sources of remaining model errors.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Its foundational account of chain-of-thought prompting clarifies the step-by-step demonstrations that Successive Prompting decomposes into separate reasoning steps.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Its analysis of multi-hop question answering and self-ask establishes the compositionality problem that motivates decomposing complex questions.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Its least-to-most method provides a direct precedent for decomposing hard questions into sequential subproblems and passing intermediate answers forward.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). Its iterative generation and filtering of reasoning traces provides context for Successive Prompting’s use of synthetic supervision to bootstrap intermediate reasoning.
- Paper: Is a Question Decomposition Unit All We Need?, Pruthvi Patel et al. (2022). Its human-in-the-loop question decomposition illustrates the manual intermediate supervision that Successive Prompting seeks to generate synthetically.
- Paper: Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models, Zhihong Shao et al. (2023). It extends synthetic reasoning supervision by having models generate and select chain-of-thought demonstrations, building on the source’s strategy for reducing costly manual data creation.
