Is a Question Decomposition Unit All We Need?

Pruthvi PatelSwaroop MishraMihir ParmarChitta Baral

article2022EMNLP66 citations

Demonstrates that breaking complex reasoning questions into simpler human-annotated sub-questions boosts accuracy by up to 29% across diverse benchmarks without retraining or scaling language models.

Listen

As natural language processing benchmarks grow increasingly complex, the conventional response has been to build larger, more resource-intensive language models. This standard approach carries significant financial costs, lengthy development timelines, and considerable environmental impacts. To address this sustainability challenge, the article evaluates whether modifying input data—specifically by breaking down complex questions into simpler sub-questions—can substantially improve the performance of existing models on unseen, challenging tasks without requiring continuous model scaling.

The article demonstrates the effectiveness of Human-in-the-loop Question Decomposition across eight diverse datasets covering reading comprehension, mathematical reasoning, fact-based multiple choice, and strategic reasoning. The evaluation examined both a large-scale model (GPT-3) and smaller, fine-tuned models (RoBERTa-base variants), coupled with a symbolic calculation unit for arithmetic operations. For each dataset, 50 randomly sampled instances were manually decomposed into sequential sub-questions based on human intuition, and model accuracy was measured against original, single-prompt baselines using standard evaluation metrics such as F1-score, Exact Match, and Rouge-L.

The findings show that human-guided question decomposition substantially improves model accuracy across all task domains. GPT-3 achieved an average performance gain of approximately 24% in F1-score across evaluated categories, while RoBERTa-based models improved by roughly 29%. Qualitative inspection revealed that decomposition corrected more than 60% of the errors made on original questions. Performance improvements were especially pronounced on mathematical reasoning tasks, where models could focus on text extraction while offloading numerical operations to a symbolic calculator. In contrast, initial attempts to automate the decomposition process using GPT-3 and fine-tuned smaller models struggled, frequently producing flawed reasoning chains and incorrect arithmetic operations.

These results indicate that structured data modification and human-in-the-loop workflows offer a viable, cost-effective alternative to perpetually training larger models. Organizations can achieve state-of-the-art reasoning performance with smaller, existing models by aligning task structures with model strengths. However, the analysis also revealed a key operational risk: decomposition chains are vulnerable to cascading errors, where an incorrect answer or an incomplete entity retrieval in an early sub-question causes all downstream steps to fail.

Senior decision-makers should consider human-in-the-loop decomposition as an immediate strategy to enhance model accuracy on complex reasoning and analytical tasks. However, relying on fully automated question decomposition is not yet recommended given current error rates in automated chain generation. Organizations should invest in hybrid workflows where human oversight guides question breakdown and validates intermediate steps, while prioritizing research into robust automated decomposition techniques.

Confidence in these findings is high regarding the manual decomposition methodology, but caution is warranted due to the relatively small evaluation sample of 50 instances per dataset. Additionally, the approach faces natural boundaries: inherently simple questions cannot be easily broken down, and questions with multiple valid intermediate answers can unintentionally divert sequential reasoning paths.

Cover for Is a Question Decomposition Unit All We Need?

Abstract

Large Language Models (LMs) have achieved state-of-the-art performance on many Natural Language Processing (NLP) benchmarks. With the growing number of new benchmarks, we build bigger and more complex LMs. However, building new LMs may not be an ideal option owing to the cost, time and environmental impact associated with it. We explore an alternative route: can we modify data by expressing it in terms of the model’s strengths, so that a question becomes easier for models to answer? We investigate if humans can decompose a hard question into a set of simpler questions that are relatively easier for models to solve. We analyze a range of datasets involving various forms of reasoning and find that it is indeed possible to significantly improve model performance (24% for GPT3 and 29% for RoBERTa-SQuAD along with a symbolic calculator) via decomposition. Our approach provides a viable option to involve people in NLP research in a meaningful way. Our findings indicate that Human-in-the-loop Question Decomposition (HQD) can potentially provide an alternate path to building large LMs1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methods
  • 3.1 Datasets
  • 3.2 Decomposition Process
  • 4 Experimental Setup
  • 5 Results and Analysis
  • 5.1 Experimental Results
  • 5.2 Analysis
  • Customized Question Decomposition for Each
  • Effect of Number of Decompositions on Results
  • 6 Conclusion
  • Limitations
  • References
  • A Prompts
  • A.1 HOTPOTQA, DROP, BREAK
  • A.2 MATHQA
  • A.3 SVAMP
  • A.4 StrategyQA
  • A.5 QASC
  • A.6 MultiRC
  • B Error Examples
  • C Examples, Results and Details for Automation
  • D Results

Knowls

  1. Knowl 1 — Human-in-the-Loop Question Decomposition Framework

    model/method

    Human-in-the-Loop Question Decomposition (HQD) reformulates complex reasoning questions into a sequence of simpler, granular sub-questions that align with the basic competencies of language models, such as single-hop extractive question answering.

    Formally, each decomposed example DiD_i in a dataset DD is represented as a 5-tuple: Di={Ci,Qi,Qd,Ai,Ad}D_i = \{C_i, Q_i, Q_d, A_i, A_d\} where:

    • CiC_i is the set of context paragraphs,
    • QiQ_i is the original complex question,
    • Qd={q1,q2,…,qk}Q_d = \{q_1, q_2, \dots, q_k\} is an ordered sequence of decomposed sub-questions,
    • AiA_i is the original ground-truth answer,
    • Ad={a1,a2,…,ak}A_d = \{a_1, a_2, \dots, a_k\} is the sequence of ground-truth answers corresponding to each decomposed sub-question in QdQ_d.

    During inference, sub-questions are answered sequentially. Intermediate answers aja_j can be referenced in subsequent sub-questions qmq_m (m>jm > j) via positional place-holders (such as #j\#j), allowing subsequent model executions or symbolic modules to resolve dependencies step by step.

  2. Knowl 2 — Symbolic Calculator Interface for Discrete and Arithmetic Decomposition

    model/method

    For questions requiring discrete arithmetic or logical operations, question decomposition interfaces with an external symbolic execution unit. The final sub-question or arithmetic step is expressed in the format: {O}!#m1!#m2!…!#mn\{O\}!\#m_1!\#m_2!\dots!\#m_n where:

    • O∈{summation,difference,division,multiplication,greater,lesser,power,concat,return,remainder}O \in \{\text{summation}, \text{difference}, \text{division}, \text{multiplication}, \text{greater}, \text{lesser}, \text{power}, \text{concat}, \text{return}, \text{remainder}\} denotes the operation to be executed,
    • #mj\#m_j represents the extracted numerical or textual answer from the jj-th preceding sub-question,
    • ! serves as the delimiter separating the operation and operand references.

    This delegates deterministic arithmetic calculation to a symbolic calculator rather than relying on language models to generate numerical answers directly.

  3. Knowl 3 — Experimental Benchmark Setup Across Diverse Reasoning Tasks

    experimental setup

    To evaluate human-in-the-loop question decomposition across diverse reasoning domains, 50 instances are randomly sampled from each of eight datasets across four reasoning categories:

    1. Reading Comprehension (RC): HotpotQA (multi-hop RC), DROP (multi-hop RC with discrete reasoning), MultiRC (multi-sentence RC), and BREAK (question understanding benchmark).
    2. Mathematical Reasoning (MATH): MathQA (operation-based math problems) and SVAMP (context-based math word problems).
    3. Multi-Choice QA (MC): QASC (fact-based multi-choice science questions).
    4. Strategy Reasoning (SR): StrategyQA (implicit multi-step strategic reasoning).

    Evaluated Models:

    • GPT-3: Evaluated via few-shot instructional prompts on original questions and decomposed questions.
    • RoBERTa-base (fine-tuned on SQuAD 2.0): Used for extractive sub-questions on MultiRC, BREAK, HotpotQA, and DROP, and operand extraction on SVAMP (paired with a symbolic calculator).
    • RoBERTa-base (fine-tuned on BoolQ): Used to evaluate original and final boolean questions in StrategyQA, while RoBERTa-SQuAD handles intermediate extractive decompositions.

    Evaluation Metrics: Exact Match (EM), F1-score, and Rouge-L.

  4. Knowl 4 — Performance Gains of GPT-3 via Question Decomposition

    empirical result

    Evaluating GPT-3 on manually decomposed questions across eight datasets yields an overall average F1-score improvement of approximately 24%24\% over standard baseline prompting:

    • Reading Comprehension (HotpotQA, DROP, MultiRC, BREAK): Average F1 improves from 60.10%60.10\% (baseline) to 81.97%81.97\% (with decomposition).
    • Mathematical Reasoning (MathQA, SVAMP): Average F1 improves from 46.45%46.45\% (baseline) to 80.62%80.62\% (with decomposition).
    • Multi-Choice QA (QASC): Average F1 improves from 75.23%75.23\% (baseline) to 89.52%89.52\% (with decomposition).
    • Strategy Reasoning (StrategyQA): Average F1 improves from 63.15%63.15\% (baseline) to 84.61%84.61\% (with decomposition).

    Mathematical reasoning tasks demonstrate the largest gains, driven by decoupling operand extraction from symbolic computation.

  5. Knowl 5 — Out-of-Domain Performance Gains on RoBERTa via Question Decomposition

    empirical result

    Applying human question decomposition allows lightweight models (RoBERTa-base fine-tuned on SQuAD 2.0 and BoolQ) to solve unseen, complex reasoning benchmarks without task-specific fine-tuning. Combining RoBERTa-SQuAD with an external symbolic calculator achieves an average F1-score improvement of ∼29%\sim 29\% over direct baseline answering on original questions:

    • Reading Comprehension (HotpotQA, DROP, MultiRC, BREAK): Average F1 rises from 23.68%23.68\% to 47.65%47.65\%, and EM rises from 20.26%20.26\% to 43.96%43.96\%.
    • Mathematical Reasoning (SVAMP): F1 and EM rise from 2.00%2.00\% (baseline) to 58.00%58.00\% (decomposed operand extraction + symbolic calculator).
    • Strategy Reasoning (StrategyQA): Using RoBERTa-SQuAD for intermediate sub-questions and RoBERTa-BoolQ for the boolean root question increases F1 and EM from 47.36%47.36\% to 55.26%55.26\%.
  6. Knowl 6 — GPT-3 Baseline vs. Decomposition Performance Across Benchmarks

    data/table

    Evaluating GPT-3 across reading comprehension, mathematical reasoning, strategic reasoning, and multi-choice QA tasks demonstrates consistent gains across F1, Exact Match (EM), and Rouge-L when applying human-in-the-loop question decomposition compared to baseline direct prompting on 50 sampled instances per dataset:

    Dataset F1 EM Rouge-L
    Baseline Decompose Baseline Decompose Baseline Decompose
    HotpotQA 71.97 78.53 70.00 76.00 73.33 79.93
    DROP 52.97 78.16 46.87 75.86 46.72 77.66
    MultiRC 64.39 80.74 33.33 55.55 61.24 77.31
    BREAK 66.81 84.54 58.00 74.00 62.30 78.56
    RC Average 60.10 81.97 52.64 76.26 59.35 81.10
    MATH 31.10 82.50 27.44 82.22 23.40 80.85
    SVAMP 61.80 78.75 58.88 77.50 55.00 77.50
    MATH Average 46.45 80.62 43.16 79.86 39.20 79.17
    StrategyQA 63.15 84.61 63.15 84.61 63.15 84.61
    QASC 75.23 89.52 75.23 89.52 71.40 85.71
    SR/MC Average 69.19 87.06 69.19 87.06 67.27 85.16

    The largest absolute gains occur on tasks requiring intermediate arithmetic derivations (MATH F1 increases by +51.40+51.40 points; DROP F1 increases by +25.19+25.19 points).

  7. Knowl 7 — RoBERTa-Based Model Performance with Question Decomposition

    data/table

    Evaluating RoBERTa-base fine-tuned on SQuAD 2.0 (and BoolQ for StrategyQA) on decomposed questions demonstrates substantial out-of-domain transfer on complex reasoning benchmarks compared to baseline single-question answering:

    Dataset F1 EM Rouge-L
    Baseline Decompose Baseline Decompose Baseline Decompose
    HotpotQA 32.14 49.50 26.00 42.00 33.33 50.72
    DROP 25.56 66.14 25.00 62.50 25.56 66.14
    MultiRC 45.74 48.10 24.44 28.88 44.83 46.95
    BREAK 24.60 36.17 18.00 28.00 24.31 35.50
    RC Average 23.68 47.65 20.26 43.96 28.76 50.74
    StrategyQA 47.36 55.26 47.36 55.26 47.36 55.26
    SVAMP 2.00 58.00 2.00 58.00 2.00 58.00
    Non-RC Average 24.68 56.63 24.68 56.63 24.68 56.63

    For SVAMP, RoBERTa-SQuAD is used to extract numerical quantities specified in sub-questions, and a symbolic calculator evaluates the arithmetic expression, transforming a near-zero baseline (2.00%2.00\% F1) into 58.00%58.00\% F1.

  8. Knowl 8 — Error Correction and Failure Modes in Sequential Decomposition Chains

    empirical result

    Manual qualitative inspection indicates that question decomposition corrects more than 60%60\% of the errors made by language models on original complex questions across both GPT-3 and RoBERTa architectures.

    However, sequential question decomposition exhibits distinct failure modes:

    1. Error Cascades/Propagation: If an intermediate sub-question produces an incorrect answer (such as incorrect span selection or omitting candidates during multi-entity retrieval), the error propagates to all subsequent dependent sub-questions, causing the final answer to fail.
    2. Context-Independent Under-Specification: In datasets without passage context (such as QASC), converting multi-choice options into independent yes/no decomposition sub-questions can produce open-ended questions that language models answer inaccurately due to lack of grounding.
    3. Multiple Valid Intermediates: When an intermediate question admits multiple valid answers (such as naming one of several founders of an organization), the model may pick a valid entity that differs from the one assumed in gold subsequent references, resulting in metric failure.
  9. Knowl 9 — Impact of Sub-Question Decomposition Count on Task Accuracy

    empirical result

    Decomposing a complex question into an increasing number of sub-questions (from 2 up to 5 steps) presents a trade-off:

    • Increasing decomposition steps reduces single-step reasoning complexity by breaking multi-operation logic into atomic single-hop questions.
    • Increasing decomposition steps simultaneously increases the cumulative probability of cascading failure, where an error in any intermediate step corrupts all downstream calculations.

    Empirical evaluation on HotpotQA and SVAMP shows no monotonic correlation between the number of decomposition steps and final task performance. Relative F1 score improvements fluctuate across different chain lengths rather than increasing strictly with the number of sub-questions.

  10. Knowl 10 — Degradation Under Automated Question Decomposition

    empirical result

    Automating question decomposition using GPT-3 few-shot generation or fine-tuned sequence-to-sequence models (BART-base) fails to match human decomposition quality and degrades downstream accuracy below baseline performance on reading comprehension benchmarks:

    Dataset F1 EM Rouge-L
    Baseline Automated Decomp. Baseline Automated Decomp. Baseline Automated Decomp.
    HotpotQA 32.68 14.12 29.50 11.47 33.29 14.00
    DROP 22.80 3.77 21.69 3.77 23.40 3.76
    SVAMP 7.40 17.35 7.40 17.35 7.40 17.35
    Average 20.96 11.74 19.53 10.86 21.36 11.70

    Key failure causes in automated decomposition include:

    • Arithmetic Incoherence: GPT-3 frequently outputs invalid arithmetic operator syntax or mismatched operands in the final sub-question for DROP and SVAMP.
    • Context Length Limitations: In long-context tasks (such as HotpotQA), prompt length constraints limit the number of in-context demonstration examples.
    • Overfitting on Small Datasets: Fine-tuning BART-base on small sets of human decompositions results in severe overfitting and non-coherent decomposition outputs.
  11. Knowl 11 — Limitations of Question Decomposition

    limitation

    The question decomposition methodology has several core limitations:

    1. Atomic Non-Decomposable Questions: Certain questions cannot be naturally decomposed into simpler sub-questions (e.g., direct factual queries such as 'Which country is New York in?').
    2. Decomposition Decision Boundary: Determining a priori which questions require decomposition versus direct single-step answering remains an unsolved challenge.
    3. Ambiguity and Multi-Valued Steps: If an intermediate sub-question possesses multiple valid answers, branching ambiguity can steer subsequent sub-questions toward divergent or incorrect final answers.
    4. Reliance on Human Annotation: High performance currently depends heavily on human intuition for formulating decompositions, as fully automated decomposition models struggle with coherence and arithmetic formalisms.

Coverage note — Detailed prompt templates for individual datasets in Appendix A were omitted as repetitive few-shot formatting examples, retaining the core symbolic decomposition syntax and representative prompt structures.

References

  1. 1.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. ArXiv, abs/1905.13319.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  3. 3.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044.
  4. 4.Dheeru Dua, Yizhong Wang Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. CoRR, abs/1903.00161.
  5. 5.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  6. 6.Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. 2019. Towards complex text-to-sql in cross-domain database with intermediate representation. arXiv preprint arXiv:1905.08205.
  7. 7.Nitish Gupta and Mike Lewis. 2018. Neural compositional denotational semantics for question answering. arXiv preprint arXiv:1808.09942.
  8. 8.Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1821–1831, Vancouver, Canada. Association for Computational Linguistics.
  9. 9.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface:a challenge set for reading comprehension over multiple sentences. In NAACL.
  10. 10.Tushar Khot, Peter Clark, Michal Guerquin, Peter Alexander Jansen, and Ashish Sabharwal. 2020a. Qasc: A dataset for question answering via sentence composition. In AAAI.
  11. 11.Tushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2020b. Text modular networks: Learning to decompose tasks in the language of existing models. CoRR, abs/2009.00751.
  12. 12.Kirby Kuznia, Swaroop Mishra, Mihir Parmar, and Chitta Baral. 2022. Less is more: Summary of long instructions is better for program synthesis. arXiv preprint arXiv:2203.08597.
  13. 13.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880.
  14. 14.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  15. 15.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  16. 16.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  17. 17.Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2019. Multi-hop reading comprehension through question decomposition and rescoring. arXiv preprint arXiv:1906.02916.
  18. 18.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021a. Reframing instructional prompts to gptk’s language. ACL Findings.
  19. 19.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021b. Cross-task generalization via natural language crowdsourcing instructions. ACL.
  20. 20.Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, and Ashwin Kalyan. 2022. Numglue: A suite of fundamental yet challenging mathematical reasoning tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3505–3523.
  21. 21.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  22. 22.Mihir Parmar, Swaroop Mishra, Mirali Purohit, Man Luo, M Hassan Murad, and Chitta Baral. 2022. In-BoXBART: Get Instructions into Biomedical Multi-Task Learning. NAACL 2022 Findings.
  23. 23.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? arXiv preprint arXiv:2103.07191.
  24. 24.Ethan Perez, Patrick Lewis, Wen-tau Yih, Kyunghyun Cho, and Douwe Kiela. 2020. Unsupervised question decomposition for question answering. arXiv preprint arXiv:2002.09758.
  25. 25.Ravsehaj Singh Puri, Swaroop Mishra, Mihir Parmar, and Chitta Baral. 2022. How many data samples is an additional instruction worth? arXiv preprint arXiv:2203.09161.
  26. 26.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  27. 27.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822.
  28. 28.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021a. Multitask prompted training enables zero-shot task generalization.
  29. 29.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021b. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  30. 30.Alon Talmor and Jonathan Berant. 2018. The web as a knowledge-base for answering complex questions. arXiv preprint arXiv:1803.06643.
  31. 31.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  32. 32.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  33. 33.Tomer Wolfson, Mor Geva, Ankit Gupta, Matt Gardner, Yoav Goldberg, Daniel Deutch, and Jonathan Berant. 2020. Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics, 8:183–198.
  34. 34.Kaige Xie, Sarah Wiegreffe, and Mark Riedl. 2022. Calibrating trust of multi-hop question answering systems with decompositional probes.
  35. 35.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. CoRR, abs/1809.09600.

Citation

MLA
Patel, P., et al. “Is a Question Decomposition Unit All We Need?”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 4553–69, https://doi.org/10.18653/v1/2022.emnlp-main.302.
APA
Patel, P., Mishra, S., Parmar, M., & Baral, C. (2022). Is a Question Decomposition Unit All We Need?. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4553–4569. https://doi.org/10.18653/v1/2022.emnlp-main.302
Chicago
Patel, P., S. Mishra, M. Parmar, and C. Baral. 2022. “Is a Question Decomposition Unit All We Need?”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4553–69. https://doi.org/10.18653/v1/2022.emnlp-main.302.
Harvard
Patel, P. et al. (2022) “Is a Question Decomposition Unit All We Need?”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4553–4569. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.302.
Vancouver
1. Patel P, Mishra S, Parmar M, Baral C (2022) Is a Question Decomposition Unit All We Need?. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4553–4569

BibTeX

@inproceedings{patel-etal-2022-question,
    title = "Is a Question Decomposition Unit All We Need?",
    author = "Patel, Pruthvi  and
      Mishra, Swaroop  and
      Parmar, Mihir  and
      Baral, Chitta",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.302/",
    doi = "10.18653/v1/2022.emnlp-main.302",
    pages = "4553--4569"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/