Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Gangwoo KimSungdong KimByeongguk JeonJoonsuk ParkJaewoo Kang

article2023EMNLP63 citations

Proposes Tree of Clarifications, a framework that recursively builds a tree of disambiguated questions guided by retrieved external knowledge and self-verification pruning to generate comprehensive long-form answers to ambiguous open-domain questions.

Listen

In real-world applications, user queries posed to search and question-answering systems are often ambiguous and open to multiple interpretations. Standard systems either fail to cover these nuances, demand time-consuming back-and-forth clarifications from users, or rely on language models prone to factual inaccuracies and hallucinations when answering broad questions.

The article demonstrates a novel framework called Tree of Clarifications (TOC), designed to automatically identify multiple interpretations of ambiguous questions and synthesize comprehensive, factually grounded long-form answers without requiring user intervention.

The approach combines large language models with external retrieval mechanisms in a recursive tree structure. Starting from an ambiguous question, the system retrieves external evidence using web search and dense passage retrieval (Bing and ColBERT), recursively branches the question into specific interpretations, prunes irrelevant or fact-distorting paths using an automated self-verification step, and aggregates the verified answers into a detailed response. The method was evaluated on the benchmark dataset ASQA (containing over 6,000 ambiguous questions) using a five-shot prompting setup rather than costly full-model retraining.

The findings show that TOC establishes a new state of the art in long-form ambiguous question answering. Operating with only five demonstration examples, TOC achieved a factual correctness score (Disambig-F1) of 33.7, outperforming fully supervised baseline models trained on the entire dataset (26.4) by 7.3 points and standard few-shot baselines (25.0) by 8.7 points. The self-verification pruning component raised the accuracy of intermediate answers from 40.9 to 59.3, effectively filtering out spurious disambiguations. In addition, combining dual retrieval sources with reranking improved the retrieval answer coverage across interpretations to 80.1%.

These results indicate that organizations can deliver highly accurate, comprehensive responses to complex, multi-faceted queries without expensive full-scale model training. Grounding multi-path reasoning in external factual retrieval significantly mitigates hallucination risks while preserving high response quality across diverse user intents.

For practical implementation, technical teams should consider adopting retrieval-augmented tree reasoning architectures for knowledge-intensive search and automated assistance workflows. To balance performance and operating costs, deployments should implement structured stopping rules (such as node caps and search-depth limits) to control query latency and API expenses.

Confidence in the reported benchmark gains is high, though readers should note certain limitations. The evaluation relies primarily on a single benchmark dataset and a single primary language model backbone (GPT-3). Furthermore, recursive tree exploration requires multiple model queries per question—up to 20 calls in worst-case paths—which increases computational overhead compared to direct single-pass generation. Further testing on diverse domain datasets and smaller, open-source models is recommended before broad enterprise rollout.

arXiv: 2310.14696
Cover for Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Abstract

Questions in open-domain question answering are often ambiguous, allowing multiple interpretations. One approach to handling them is to identify all possible interpretations of the ambiguous question (AQ) and to generate a long-form answer addressing them all, as suggested by Stelmakh et al. (2022). While it provides a comprehensive response without bothering the user for clarification, considering multiple dimensions of ambiguity and gathering corresponding knowledge remains a challenge. To cope with the challenge, we propose a novel framework, Tree of Clarifications (ToC): It recursively constructs a tree of disambiguations for the AQ—via few-shot prompting leveraging external knowledge—and uses it to generate a long-form answer. ToC outperforms existing baselines on ASQA in a few-shot setup across all metrics, while surpassing fully-supervised baselines trained on the whole training set in terms of Disambig-F1 and Disambig-ROUGE. Code is available at github.com/gankim/tree-of-clarifications.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Tree of Clarifications
  • 3.1 Retrieval-Augmented Clarification (RAC)
  • 3.2 Tree Structure (TS)
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Experimental Results
  • 5 Discussion
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Experimental Setup Details
  • A.1 Ambiguous QA Datasets
  • A.2 Evaluation Metrics
  • A.3 Implementation Details
  • B Additional Experiment
  • B.1 Intrinsic Evaluation for Retrieval Systems
  • C Qualitative Analysis
  • C.1 Prompt Format
  • C.2 Question Clarification
  • C.3 Self Verification
  • C.4 Answer Generation

Knowls

  1. Knowl 1 — Retrieval-grounded recursive framework for answering ambiguous questions

    model/method

    TREE OF CLARIFICATIONS (TOC) addresses open-domain questions that allow multiple interpretations by producing a single long-form answer that covers plausible interpretations. Given an ambiguous question (AQ), TOC retrieves supporting passages, uses a large language model (LLM) to generate disambiguated question–answer pairs, and recursively expands those pairs into a tree of clarifications. After removing unhelpful nodes, it uses the retained disambiguations and their evidence to generate a comprehensive answer. The tree is intended to help the LLM explore distinct dimensions of ambiguity, while retrieval supplies external knowledge for identifying interpretations and answers.

  2. Knowl 2 — Retrieval-Augmented Clarification (RAC)

    model/method

    Retrieval-Augmented Clarification (RAC) grounds the generation of disambiguated questions (DQs) and answers in passages relevant to the current question. It retrieves Wikipedia passages using both ColBERT and Bing search, combining the results into a set of more than 200 passages. ColBERT uses an off-the-shelf model pretrained on MS MARCO. A SentenceBERT model pretrained on MS MARCO reranks the passages, and the top five are supplied to the LLM. For few-shot prompting, the system selects five training examples by nearest-neighbor search over question representations produced by MiniLM, with similarity search implemented using Faiss. The LLM is prompted with the current question, retrieved passages, and examples to produce DQs and corresponding short answers. During recursive tree expansion, retrieval and reranking are performed again for the current query.

  3. Knowl 3 — Tree search over clarification dimensions

    model/method

    TOC represents the original ambiguous question as the root of a clarification tree. Each expansion applies RAC to a question node; generated DQ–answer pairs become child nodes, allowing later expansions to focus on distinct interpretations. The default traversal is breadth-first search (BFS), chosen to explore broad interpretations before extending deeper paths. In the reported experiments, exploration stops after reaching 10 valid nodes, reaching the configured maximum depth, or failing to expand for three consecutive attempts; an expansion typically produces two to five DQs. A node is valid if it is retained after self-verification pruning.

  4. Knowl 4 — Self-verification pruning of clarification nodes

    model/method

    TOC prunes a generated DQ node when its answer does not remain factually coherent with the original ambiguous question. For each target node, the system selects the most relevant passage from passages containing the node’s proposed answer, then prompts the LLM with the root question, that passage, and the proposed answer. The LLM returns True or False for whether the proposed answer could answer the original question; nodes judged unable to do so are discarded. This is intended to remove factually supported but out-of-scope clarifications, such as a question about who hosted the 2018 World Cup when the original asks who will host the 2022 World Cup.

  5. Knowl 5 — Aggregation and long-form answer generation

    model/method

    After constructing the clarification tree, TOC generates a long-form answer from the original ambiguous question, retained DQs and their answers, and supporting passages. It selects up to 10 valid disambiguations in breadth-first order and five answer-containing passages, prioritizing passages that support valid nodes. If pruning leaves too few nodes, TOC reverses pruning decisions in breadth-first order, beginning with nodes nearer the root, until enough disambiguations are available. The answer-generation prompt asks for a detailed response that explains the question’s multiple interpretations and incorporates the corresponding answers.

  6. Knowl 6 — ASQA evaluation setup and metrics

    experimental setup

    TOC and its baselines were evaluated on ASQA, a long-form open-domain question-answering benchmark containing 6,316 ambiguous questions derived from AmbigNQ. Its train, development, and test splits contain 4,353, 948, and 1,015 questions, respectively; the reported main results are on the development set. The experimental TOC backbone was GPT-3 text-davinci-002, with five dynamically selected in-context examples, five retrieved passages for clarification, a 300-token maximum generation length, and top-p set to 1.0.

    The evaluation reports Disambig-F1 (D-F1), which measures factual correctness of answers to the benchmark’s disambiguated questions; ROUGE-L (R-L), which measures lexical overlap with reference long-form answers; and Disambiguation-ROUGE (DR), the geometric mean of D-F1 and R-L. The benchmark provides two reference answers, and the maximum ROUGE-L score is used. Intrinsic evaluations also use Answer-F1 for answer accuracy on a single DQ.

  7. Knowl 7 — TOC improves few-shot ambiguous-question answering

    empirical result

    On the ASQA development set, the following results compare fully supervised and five-shot baselines with TOC. D-F1 evaluates factoid answer correctness, R-L evaluates long-form lexical overlap, and DR is their geometric mean; scores are reported as in the paper.

    Model D-F1 R-L DR
    T5-Large, closed-book 7.4 33.5 15.7
    T5-Large with JPR 26.4 43.0 33.7
    PaLM with soft prompt tuning 27.8 37.4 32.1
    PaLM, 5-shot 25.3 34.5 29.6
    GPT-3, 5-shot 25.0 31.8 28.2
    GPT-3 + RAC 31.1 39.6 35.1
    GPT-3 + RAC + tree structure 32.4 40.0 36.0
    GPT-3 + RAC + tree structure with pruning 33.7 39.7 36.6

    Adding the tree structure to RAC raises D-F1 from 31.1 to 32.4 and DR from 35.1 to 36.0. Adding self-verification pruning raises D-F1 to 33.7 and DR to 36.6. The pruned TOC model exceeds the best five-shot baseline (PaLM) by 8.4 D-F1 points and 7.0 DR points, and exceeds the fully supervised T5-Large with JPR by 7.3 D-F1 points and 2.9 DR points. Its R-L score, 39.7, is below the 43.0 of T5-Large with JPR.

  8. Knowl 8 — Retrieval contributes to QA quality, and the retrievers are complementary

    empirical result

    Ablations on the ASQA development set show that retrieval and the intermediate disambiguations contribute to GPT-3 + RAC performance. D-F1 measures factual correctness of answers to disambiguated questions, R-L measures long-form answer overlap, and DR is their geometric mean. Removing disambiguations from the few-shot examples reduces R-L from 39.6 to 37.3; removing Bing reduces D-F1 from 31.1 to 28.5; removing retrieval systems reduces D-F1 to 25.6.

    Model or ablation D-F1 R-L DR
    GPT-3 baseline 24.2 36.0 29.5
    GPT-3 with RAC 31.1 39.6 35.1
    Without disambiguations in few-shot examples 30.5 37.3 33.7
    Without Bing Search Engine 28.5 37.4 32.7
    Without retrieval systems 25.6 35.1 30.0

    A separate intrinsic retrieval evaluation randomly sampled 100 ASQA examples. Answer coverage at kk (AC@kk) is the proportion of reference DQ answers found in the top kk retrieved passages. Combining ColBERT and Bing with reranking gives the highest reported coverage at all tested cutoffs, indicating that the retrieval sources are complementary.

    Retrieval system AC@10 AC@30 AC@100
    ColBERTv2 56.4 68.4 73.4
    ColBERTv2 with reranker 56.8 69.0 73.4
    Bing Search Engine 43.3 58.3 73.5
    Bing Search Engine with reranker 62.7 68.0 72.8
    Combined with reranker 64.2 77.4 80.1
  9. Knowl 9 — Self-verification improves retained-node answer accuracy

    empirical result

    On the generated disambiguations evaluated with Answer-F1, self-verification pruning substantially increases the accuracy of retained answers, whereas deduplication alone does not. Without pruning, the system produces 12,838 DQs with Answer-F1 40.9. With deduplication, it retains 10,598 DQs and obtains Answer-F1 40.1. With self-verification, it retains 4,239 DQs and obtains Answer-F1 59.3, an increase of 18.4 points over the no-pruning result. Answer-F1 measures F1 accuracy of answers to a single DQ.

  10. Knowl 10 — Limitations and unresolved issues

    limitation

    The paper evaluates TOC only on ASQA and does not demonstrate generalization across different LLM architectures or model sizes. Its iterative prompting incurs non-negligible cost, although the reported setup uses fewer than 20 LLM calls per question. TOC does not explicitly classify whether an input is ambiguous and may generate duplicate or irrelevant DQs for questions that cannot be further disambiguated; the authors suggest treating failed expansion or pruning of every candidate as evidence that a question may be unambiguous. A pilot using chain-of-thought prompting did not improve performance. The authors identify more effective pruning and improved answer-sentence reranking as possible directions for further gains.

Coverage note — Individual prompt exemplars and qualitative success/failure cases are omitted because they illustrate the included clarification and pruning procedures rather than add independent findings.

References

  1. 1.Reinald Kim Amplayo, Kellie Webster, Michael Collins, Dipanjan Das, and Shashi Narayan. 2023. Query refinement prompts for closed-book long-form question answering. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.
  2. 2.Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al. 2016. Ms marco: A human generated machine reading comprehension dataset. 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain.
  3. 3.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870–1879.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  5. 5.Jeremy R Cole, Michael JQ Zhang, Daniel Gillick, Julian Martin Eisenschlos, Bhuwan Dhingra, and Jacob Eisenstein. 2023. Selectively answering ambiguous questions. arXiv preprint arXiv:2305.14613.
  6. 6.Yifan Gao, Henghui Zhu, Patrick Ng, Cicero dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O Arnold, and Bing Xiang. 2021. Answering ambiguous questions through generative evidence fusion and round-trip prediction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3263–3276.
  7. 7.Siddhant Garg, Thuy Vu, and Alessandro Moschitti. 2020. Tanda: Transfer and adapt pre-trained transformer models for answer sentence selection. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7780–7788.
  8. 8.Meiqi Guo, Mingda Zhang, Siva Reddy, and Malihe Alikhani. 2021. Abg-coqa: Clarifying ambiguity in conversational question answering. In 3rd Conference on Automated Knowledge Base Construction.
  9. 9.Gautier Izacard and Édouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880.
  10. 10.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7(3):535–547.
  11. 11.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221.
  12. 12.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769–6781.
  13. 13.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp. arXiv preprint arXiv:2212.14024.
  14. 14.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, pages 39–48.
  15. 15.Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021. Hurdles to progress in long-form question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4940–4957.
  16. 16.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2022. Clam: Selective clarification for ambiguous questions with large language models. arXiv preprint arXiv:2212.07769.
  17. 17.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  18. 18.Ivano Lauriola and Alessandro Moschitti. 2021. Answer sentence selection using local and global context in transformer models. In European Conference on Information Retrieval, pages 298–312. Springer.
  19. 19.Dongryeol Lee, Segwang Kim, Minwoo Lee, Hwanhee Lee, Joonsuk Park, Sang-Woo Lee, and Kyomin Jung. 2023. Asking clarification questions to handle ambiguity in open-domain qa. arXiv preprint arXiv:2305.13808.
  20. 20.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  21. 21.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  22. 22.Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, and Hannaneh Hajishirzi. 2021. Joint passage ranking for diverse multi-answer retrieval. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6997–7008.
  23. 23.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. Ambigqa: Answering ambiguous open-domain questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5783–5797.
  24. 24.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  25. 25.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  26. 26.Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know what you don’t know: Unanswerable questions for squad. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784–789.
  27. 27.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992.
  28. 28.Zhihong Shao and Minlie Huang. 2022. Answering open-domain multi-answer questions via a recall-then-verify framework. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1825–1838.
  29. 29.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: Factoid questions meet long-form answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8273–8288, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  30. 30.Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. Advances in Neural Information Processing Systems, 33:5776–5788.
  31. 31.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  32. 32.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601.

Citation

MLA
Kim, G., et al. “Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 996–1009, https://doi.org/10.18653/v1/2023.emnlp-main.63.
APA
Kim, G., Kim, S., Jeon, B., Park, J., & Kang, J. (2023). Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 996–1009. https://doi.org/10.18653/v1/2023.emnlp-main.63
Chicago
Kim, G., S. Kim, B. Jeon, J. Park, and J. Kang. 2023. “Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 996–1009. https://doi.org/10.18653/v1/2023.emnlp-main.63.
Harvard
Kim, G. et al. (2023) “Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 996–1009. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.63.
Vancouver
1. Kim G, Kim S, Jeon B, Park J, Kang J (2023) Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 996–1009

BibTeX

@inproceedings{kim-etal-2023-tree,
    title = "Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models",
    author = "Kim, Gangwoo  and
      Kim, Sungdong  and
      Jeon, Byeongguk  and
      Park, Joonsuk  and
      Kang, Jaewoo",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.63/",
    doi = "10.18653/v1/2023.emnlp-main.63",
    pages = "996--1009"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/