Successive Prompting for Decomposing Complex Questions

Dheeru DuaShivanshu GuptaSameer SinghMatt Gardner

article2022EMNLP157 citations

Proposes an iterative prompting method that decouples question decomposition from question answering to enable step-specific in-context learning and synthetic bootstrapping for multi-step reasoning.

Listen

Organizations increasingly rely on automated language systems to analyze unstructured texts and answer complex questions that require multi-step reasoning. However, standard language models struggle to perform these complex reasoning tasks accurately when limited training data is available. Existing approaches often force a single model to generate all intermediate reasoning steps and the final answer in one continuous pass, which limits flexibility and prevents the integration of specialized tools for tasks like arithmetic.

The article evaluates "Successive Prompting," a framework designed to iteratively break complex questions into simple question-answer pairs until reaching a final solution. The objective is to demonstrate that decoupling the question decomposition process from the question answering process improves multi-step reasoning performance in data-constrained environments.

To test this approach, the researchers conducted experiments on the DROP reading comprehension benchmark using a challenging few-shot setup with only 300 annotated examples. They evaluated both prompt-based prompting on a large language model and fine-tuned modular models. To address the scarcity of intermediate training data, the authors developed a synthetic data generator using structured Wikipedia tables, producing roughly 141,000 multi-step reasoning questions covering 10 distinct operations.

The analysis revealed several key findings. First, the best fine-tuned Successive Prompting model achieved a 50.2% test accuracy score, outperforming the previous state-of-the-art benchmark by an absolute 5.1 percentage points under the same limited supervision. Second, pre-training models on out-of-domain synthetic data universally improved performance across all baseline architectures, boosting one baseline by nearly 20 percentage points. Third, when using few-shot in-context learning without fine-tuning, Successive Prompting outperformed single-pass Chain-of-Thought prompting by 4.3 percentage points on the development set. Finally, delegating mathematical sub-tasks to a deterministic symbolic calculator provided an immediate 1.5 percentage point performance boost over relying on the language model alone.

These findings indicate that modular, iterative architectures provide greater accuracy and interpretability than monolithic, single-pass language models for complex analytical tasks. By isolating individual reasoning steps, organizations can target training data to specific weaknesses and plug in deterministic computational engines for arithmetic operations, thereby reducing the risk of calculation errors. However, decision-makers should note that Successive Prompting increases operational computational costs and latency, as answering a single question requires multiple sequential model queries rather than a single pass.

For practical implementation, organizations developing complex question-answering pipelines should adopt modular frameworks that separate reasoning breakdown from factual answering and integrate external symbolic engines for mathematical tasks. Teams should also utilize synthetic data generation to bootstrap capabilities before collecting costly human annotations. Further development is recommended to extend synthetic data generators to cover implicit, causal, and common-sense reasoning patterns, which were identified as key sources of remaining model errors.

arXiv: 2212.04092
Cover for Successive Prompting for Decomposing Complex Questions

Abstract

Answering complex questions that require making latent decisions is a challenging task, especially when limited supervision is available. Recent works leverage the capabilities of large language models (LMs) to perform complex question answering in a few-shot setting by demonstrating how to output intermediate rationalizations while solving the complex question in a single pass. We introduce “Successive Prompting”, where we iteratively break down a complex task into a simple task, solve it, and then repeat the process until we get the final solution. Successive prompting decouples the supervision for decomposing complex questions from the supervision for answering simple questions, allowing us to (1) have multiple opportunities to query in-context examples at each reasoning step (2) learn question decomposition separately from question answering, including using synthetic data, and (3) use bespoke (fine-tuned) components for reasoning steps where a large LM does not perform well. The intermediate supervision is typically manually written, which can be expensive to collect. We introduce a way to generate a synthetic dataset which can be used to bootstrap a model’s ability to decompose and answer intermediate questions. Our best model (with successive prompting) achieves an improvement of ~5% absolute F1 on a few-shot version of the DROP dataset when compared with a state-of-the-art model with the same supervision.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Decomposing Complex Questions
  • 2.1 Successive prompting
  • 2.2 Training paradigm
  • 3 Synthetic Dataset
  • 4 Experiments and Results
  • 4.1 In-context Learning
  • 4.2 Model Fine-tuning
  • 4.3 In-context vs Fine-Tuning
  • 4.4 Qualitative Examples
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • Limitations
  • Ethics Statement
  • References
  • A Appendix
  • A.1 Control codes for Model Fine-tuning
  • A.2 Synthetic Dataset Statistics

Knowls

  1. Knowl 1 — Successive prompting interleaves decomposition and answering

    model/method

    Successive prompting answers a complex question about a passage through alternating question-decomposition (QD) and question-answering (QA) steps. At each QD step, a model uses the original question, passage, and previously asked and answered subquestions to produce the next simple question. A separate QA step answers that question from the passage. The process repeats until QD emits a termination marker indicating that no more questions are needed, together with the final answer. The intermediate state is represented as simple question-answer pairs rather than declarative reasoning sentences, and QD and QA are separate model queries.

  2. Knowl 2 — Separate retrieval indexes tailor demonstrations to each step

    model/method

    For in-context successive prompting, the method maintains a QD index and a QA index. The QD index stores partially decomposed training examples, including the next question appropriate at each step; it is queried using the held-out complex question and current step. The QA index stores simple question-answer pairs from the training decompositions and is queried using the newly generated simple question. This allows QD and QA to retrieve demonstrations suited to their different tasks, including QA examples whose simple questions resemble a subquestion but not the original complex question. Feeding the accumulated question-answer history back to QD also supports decompositions that need to refer to earlier intermediate results.

  3. Knowl 3 — Synthetic decompositions generated from Wikipedia tables

    model/method

    The paper generates supervision by converting semi-structured English Wikipedia tables into paragraph contexts and using curated templates to create questions and decompositions. Questions based on individual columns provide first-order operations; combinations of columns and operations create higher-order reasoning examples. The generated operations include COUNT, TOP(k), BOTTOM(k), FILTER, SUM, COMPARISON, DIFFERENCE, NEGATION, GATHER, and INTERSECTION. Arithmetic steps are represented in natural language when a language model will solve them and symbolically when a separate reasoning engine will execute them. The generator produced approximately 141K complex questions, yielding 525K QD examples and 257K QA examples.

  4. Knowl 4 — Fine-tuning shares a model across QD and QA

    model/method

    The fine-tuned successive-prompting system uses a T5-large UnifiedQA sequence-to-sequence model with shared parameters for QD and QA, trained as separate tasks in a multitask setup. Training begins with two epochs of cross-entropy training, followed by an auxiliary contrastive-estimation loss that favors an intermediate subquestion leading to a correct subanswer over one that does not; up to three negative samples are used per step. The learning rate is 5e-5 and the maximum input length is 768 tokens. To balance synthetic reasoning types, each epoch samples 80,000 instances, allocating them according to each type’s performance drop from the preceding epoch on held-out synthetic data; the first epoch samples in proportion to the original type counts. At inference, beam search uses size 5, and QD and QA alternate until the end-of-decomposition marker is produced or 10 steps have been reached.

  5. Knowl 5 — A symbolic module executes parseable arithmetic questions

    model/method

    Successive prompting can route generated simple questions to a deterministic symbolic QA module instead of asking a language model to perform every operation. The module parses the question to identify a supported mathematical operation and its arguments, then executes that operation; the paper gives counting, difference, and sorting as examples of operations that can challenge language models. If the question cannot be parsed as a mathematical operation, the system uses a language model to answer it. This makes QA modular while leaving QD responsible for generating the intermediate questions.

  6. Knowl 6 — Fine-tuned successive prompting leads DROP dev results

    empirical result

    On the DROP development set, the fine-tuned successive-prompting system achieved 49.8 F1 with synthetic-data training and no in-domain DROP supervision, and 51.3 F1 when trained with the 300 DROP examples and their decompositions. For comparison, TASE with synthetic data and 300 DROP decompositions scored 45.9 F1, a 5.4-point difference from successive prompting under the corresponding supervision setting. In the reported test evaluation, the best successive-prompting model scored 50.2 F1 versus 45.1 F1 for TASE with synthetic data and comparable decomposition supervision, a 5.1-point difference. The broader development comparison also showed gains from synthetic data: UnifiedQA rose from 27.2 to 32.6 F1 with QA-only DROP supervision and from 24.5 to 26.6 in the zero-shot setting; PReasM rose from 37.5 to 38.1 and from 24.9 to 30.2, respectively; and TASE rose from 27.6 to 45.9 with decomposition supervision and from 26.1 to 44.1 with QA-only supervision. DROP is a reading-comprehension benchmark requiring discrete reasoning over passages.

  7. Knowl 7 — In-context successive prompting improves DROP F1

    empirical result

    With GPT-J (6B), six in-context examples, and DROP development-set evaluation, successive prompting outperformed chain-of-thought prompting in all three supervision conditions. The F1 scores for synthetic-only, DROP-only, and combined synthetic-plus-DROP examples were, respectively: standard prompting, 22.7, 23.8, and 24.9; chain-of-thought, 25.3, 26.2, and 27.6; successive prompting without the calculator, 27.2, 29.3, and 29.9; and successive prompting with the calculator, 28.8, 30.8, and 31.9. Thus, with the calculator, successive prompting gained 3.5 F1 over chain-of-thought with synthetic-only examples, 4.6 with DROP-only examples, and 4.3 with both sources. The best development-selected model scored 30.6 F1 on the test set. Replacing the symbolic calculator with language-model arithmetic reduced F1 by 1.5 in the reported ablation.

  8. Knowl 8 — QA capability is a bottleneck in the QD–QA ablation

    empirical result

    An ablation on DROP compared in-context and fine-tuned QD and QA modules. With in-context QD, changing QA from in-context to fine-tuned increased F1 from 30.8 to 40.3. With fine-tuned QD, the corresponding scores were 31.4 with in-context QA and 51.3 with fine-tuned QA. These results show that improving QA alone produced a large gain, whereas fine-tuning QD while retaining in-context QA produced only a small gain over using both in-context modules. The paper attributes this pattern to final-answer performance depending on whether QA can correctly answer the decompositions that QD generates.

  9. Knowl 9 — Construction of the 300-example DROP supervision set

    experimental setup

    To select a compact, broadly representative set of DROP training examples, the authors embedded the questions with a sentence-embedding model trained on QQP and found the 50 nearest-neighbor questions for each training question using cosine similarity. They formed a question-neighbor graph and applied a vertex-cover algorithm to select 300 questions intended to cover most of the training data. The selected complex questions were manually annotated with simple question-answer decompositions in the same format as the synthetic examples. For synthetic in-context demonstrations, examples were sampled uniformly across reasoning types.

  10. Knowl 10 — Manual audit identifies QA, decomposition, and coverage errors

    empirical result

    A manual analysis sampled 50 correct and 50 incorrect development-set predictions for each of the in-context and fine-tuned settings. Among the sampled correct predictions, the QD stage was judged accurate in 88% of in-context cases and 96% of fine-tuned cases; some apparently incorrect decompositions still yielded correct answers when the original question itself was directly answerable. Among sampled incorrect predictions, the reported causes were incorrect QA answers in 40% of in-context cases and 22% of fine-tuned cases, incorrect next-question predictions in 30% of both settings, and reasoning types not covered by the synthetic annotations in 28% and 46%, respectively. These are proportions from the manually inspected samples, not estimates over all model predictions.

  11. Knowl 11 — Decomposition coverage and repeated queries limit applicability

    limitation

    The approach requires some decomposition supervision, which may be difficult or impossible to obtain, and generating useful synthetic data can be challenging in domains without readily parsed structure. The paper’s generator covers many DROP reasoning types but not all complex reasoning, including commonsense and causal questions. Decomposition granularity also depends on the underlying model: a question that a model can answer directly need not be broken down, and changing model capabilities can make a chosen decomposition strategy less suitable. Finally, successive prompting requires multiple language-model queries for one complex question, increasing computation compared with answering in a single query.

Coverage note — The appendix’s individual synthetic decomposition templates and per-reasoning-type example counts are omitted because they instantiate the table-based generator already described rather than add a separate main result.

References

  1. 1.Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don’t just assume; look and answer: Overcoming priors for visual question answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 4971–4980. IEEE Computer Society.
  2. 2.Daniel Andor, Luheng He, Kenton Lee, and Emily Pitler. 2019. Giving BERT a calculator: Finding operations and arguments with reading comprehension. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5947–5952, Hong Kong, China. Association for Computational Linguistics.
  3. 3.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  4. 4.Xinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou, Dawn Song, and Quoc V. Le. 2020. Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  5. 5.Yanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis, and He He. 2022. Meta-learning via language model in-context tuning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 719–730, Dublin, Ireland. Association for Computational Linguistics.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. ArXiv preprint, abs/2204.02311.
  7. 7.Dheeru Dua, Pradeep Dasigi, Sameer Singh, and Matt Gardner. 2021. Learning with instance bundles for reading comprehension. ArXiv preprint, abs/2104.08735.
  8. 8.Dheeru Dua, Sameer Singh, and Matt Gardner. 2020. Benefits of intermediate annotations in reading comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5627–5634, Online. Association for Computational Linguistics.
  9. 9.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2368–2378, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Ananth Gottumukkala, Dheeru Dua, Sameer Singh, and Matt Gardner. 2020. Dynamic sampling strategies for multi-task reading comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 920–924, Online. Association for Computational Linguistics.
  11. 11.Nitish Gupta, Kevin Lin, Dan Roth, Sameer Singh, and Matt Gardner. 2020. Neural module networks for reasoning over text. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  12. 12.Minghao Hu, Yuxing Peng, Zhen Huang, and Dongsheng Li. 2019. A multi-type multi-span network for reading comprehension that requires discrete reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1596–1606, Hong Kong, China. Association for Computational Linguistics.
  13. 13.Yue Jin, Tianqing Zheng, Chao Gao, and Guoqiang Xu. 2021. Mtmsn: Multi-task and multi-modal sequence network for facial action unit and expression recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3597–3602.
  14. 14.Ehud Karpas, Omri Abend, Yonatan Belinkov, Barak Lenz, Opher Lieber, Nir Ratner, Yoav Shoham, Hofit Bata, Yoav Levine, Kevin Leyton-Brown, et al. 2022. Mrkl systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning. ArXiv preprint, abs/2205.00445.
  15. 15.Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020. UNIFIEDQA: Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907, Online. Association for Computational Linguistics.
  16. 16.Tushar Khot, Daniel Khashabi, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2021. Text modular networks: Learning to decompose tasks in the language of existing models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1264–1279, Online. Association for Computational Linguistics.
  17. 17.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. ArXiv preprint, abs/2210.02406.
  18. 18.Xiaofei Ma, Cicero Nogueira dos Santos, and Andrew O. Arnold. 2021. Contrastive fine-tuning improves robustness for neural rankers. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 570–582, Online. Association for Computational Linguistics.
  19. 19.Ana Marasovic, Iz Beltagy, Doug Downey, and Matthew E Peters. 2021. Few-shot self-rationalization with natural language prompts. ArXiv preprint, abs/2111.08284.
  20. 20.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021. Show your work: Scratchpads for intermediate computation with language models. ArXiv preprint, abs/2112.00114.
  21. 21.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. Advances in Neural Information Processing Systems, 34.
  22. 22.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. ArXiv preprint, abs/2210.03350.
  23. 23.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  24. 24.Dheeraj Rajagopal, Siamak Shakeri, Cicero Nogueira dos Santos, Eduard Hovy, and Chung-Ching Chang. 2022. Counterfactual data augmentation improves factuality of abstractive summarization. ArXiv preprint, abs/2205.12416.
  25. 25.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  26. 26.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2021. Learning to retrieve prompts for in-context learning. ArXiv preprint, abs/2112.08633.
  27. 27.Timo Schick. 2022. Few-shot learning with language models: Learning from instructions and contexts. Ph.D. thesis, lmu.
  28. 28.Elad Segal, Avia Efrat, Mor Shoham, Amir Globerson, and Jonathan Berant. 2020. A simple and effective model for answering multi-span questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3074–3080, Online. Association for Computational Linguistics.
  29. 29.Noah A. Smith and Jason Eisner. 2005. Contrastive estimation: Training log-linear models on unlabeled data. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), pages 354–362, Ann Arbor, Michigan. Association for Computational Linguistics.
  30. 30.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv preprint, abs/2201.11903.
  31. 31.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  32. 32.Ori Yoran, Alon Talmor, and Jonathan Berant. 2021. Turning tables: Generating examples from semi-structured tables for endowing language models with reasoning skills. ArXiv preprint, abs/2107.07261.
  33. 33.Eric Zelikman, Yuhuai Wu, and Noah D Goodman. 2022. Star: Bootstrapping reasoning with reasoning. ArXiv preprint, abs/2203.14465.
  34. 34.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. ArXiv preprint, abs/2205.10625.

Citation

MLA
Dua, D., et al. “Successive Prompting for Decomposing Complex Questions”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 1251–65, https://doi.org/10.18653/v1/2022.emnlp-main.81.
APA
Dua, D., Gupta, S., Singh, S., & Gardner, M. (2022). Successive Prompting for Decomposing Complex Questions. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1251–1265. https://doi.org/10.18653/v1/2022.emnlp-main.81
Chicago
Dua, D., S. Gupta, S. Singh, and M. Gardner. 2022. “Successive Prompting for Decomposing Complex Questions”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1251–65. https://doi.org/10.18653/v1/2022.emnlp-main.81.
Harvard
Dua, D. et al. (2022) “Successive Prompting for Decomposing Complex Questions”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1251–1265. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.81.
Vancouver
1. Dua D, Gupta S, Singh S, Gardner M (2022) Successive Prompting for Decomposing Complex Questions. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1251–1265

BibTeX

@inproceedings{dua-etal-2022-successive,
    title = "Successive Prompting for Decomposing Complex Questions",
    author = "Dua, Dheeru  and
      Gupta, Shivanshu  and
      Singh, Sameer  and
      Gardner, Matt",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.81/",
    doi = "10.18653/v1/2022.emnlp-main.81",
    pages = "1251--1265"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/