Is a Question Decomposition Unit All We Need?
Pruthvi PatelSwaroop MishraMihir ParmarChitta Baral
Demonstrates that breaking complex reasoning questions into simpler human-annotated sub-questions boosts accuracy by up to 29% across diverse benchmarks without retraining or scaling language models.
As natural language processing benchmarks grow increasingly complex, the conventional response has been to build larger, more resource-intensive language models. This standard approach carries significant financial costs, lengthy development timelines, and considerable environmental impacts. To address this sustainability challenge, the article evaluates whether modifying input data—specifically by breaking down complex questions into simpler sub-questions—can substantially improve the performance of existing models on unseen, challenging tasks without requiring continuous model scaling.
The article demonstrates the effectiveness of Human-in-the-loop Question Decomposition across eight diverse datasets covering reading comprehension, mathematical reasoning, fact-based multiple choice, and strategic reasoning. The evaluation examined both a large-scale model (GPT-3) and smaller, fine-tuned models (RoBERTa-base variants), coupled with a symbolic calculation unit for arithmetic operations. For each dataset, 50 randomly sampled instances were manually decomposed into sequential sub-questions based on human intuition, and model accuracy was measured against original, single-prompt baselines using standard evaluation metrics such as F1-score, Exact Match, and Rouge-L.
The findings show that human-guided question decomposition substantially improves model accuracy across all task domains. GPT-3 achieved an average performance gain of approximately 24% in F1-score across evaluated categories, while RoBERTa-based models improved by roughly 29%. Qualitative inspection revealed that decomposition corrected more than 60% of the errors made on original questions. Performance improvements were especially pronounced on mathematical reasoning tasks, where models could focus on text extraction while offloading numerical operations to a symbolic calculator. In contrast, initial attempts to automate the decomposition process using GPT-3 and fine-tuned smaller models struggled, frequently producing flawed reasoning chains and incorrect arithmetic operations.
These results indicate that structured data modification and human-in-the-loop workflows offer a viable, cost-effective alternative to perpetually training larger models. Organizations can achieve state-of-the-art reasoning performance with smaller, existing models by aligning task structures with model strengths. However, the analysis also revealed a key operational risk: decomposition chains are vulnerable to cascading errors, where an incorrect answer or an incomplete entity retrieval in an early sub-question causes all downstream steps to fail.
Senior decision-makers should consider human-in-the-loop decomposition as an immediate strategy to enhance model accuracy on complex reasoning and analytical tasks. However, relying on fully automated question decomposition is not yet recommended given current error rates in automated chain generation. Organizations should invest in hybrid workflows where human oversight guides question breakdown and validates intermediate steps, while prioritizing research into robust automated decomposition techniques.
Confidence in these findings is high regarding the manual decomposition methodology, but caution is warranted due to the relatively small evaluation sample of 50 instances per dataset. Additionally, the approach faces natural boundaries: inherently simple questions cannot be easily broken down, and questions with multiple valid intermediate answers can unintentionally divert sequential reasoning paths.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the foundational paradigm of chain-of-thought prompting for multi-step reasoning in language models that the source adapts through explicit question decomposition.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Presents least-to-most prompting, establishing the core problem-decomposition strategy of breaking complex questions into sequential subproblems that the source directly tests and analyzes.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates step-by-step elicitation of latent reasoning in large language models, providing the essential zero-shot reasoning baseline that question decomposition seeks to improve upon.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Introduces GPT-3 and the in-context few-shot learning paradigm, which serves as the primary large-model testbed evaluated in the source.
- Paper: RoBERTa: A Robustly Optimized BERT Pretraining Approach, Yinhan Liu et al. (2019). Details the architecture and pre-training of RoBERTa, establishing the baseline encoder model family used in the source's smaller-scale decomposition experiments.
- Paper: ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering, Zhiyu Chen et al. (2022). Introduces sequential numerical reasoning chains over text and tables, formulating the hybrid question answering and arithmetic decomposition challenges explored in the source.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Formalizes the compositionality gap and introduces automated sub-question decomposition via self-ask prompting, directly addressing the manual decomposition limitations highlighted in the source.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Extends the goal of enabling smaller models to solve complex reasoning problems by distilling multi-step reasoning capabilities from large teacher models into compact student architectures.
- Paper: Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes, Cheng-Yu Hsieh et al. (2023). Builds upon reasoning decomposition to train compact language models on step-by-step rationales, matching large-model performance with fewer parameters and training examples.
- Paper: Branch-Solve-Merge Improves Large Language Model Evaluation and Generation, Swarnadeep Saha et al. (2024). Generalizes sequential sub-question decomposition into a modular parallel framework that branches, solves, and merges sub-tasks for complex generation and evaluation.
- Paper: DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction, Mohammadreza Pourreza et al. (2023). Applies sub-task decomposition and intermediate reasoning chains specifically to complex natural language-to-SQL translation with self-correction mechanisms.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Enhances multi-step reasoning pipelines by sampling multiple reasoning paths and taking a majority vote to mitigate the cascading errors noted in the source.
- Paper: Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication, Zhangyue Yin et al. (2023). Extends single-model step-by-step reasoning by allowing multiple language models to communicate and critique intermediate steps collaboratively.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Provides a systematic formal evaluation of why step-by-step reasoning chains suffer from myopic local decisions and global planning failures.
