Aligning Language Models to Explicitly Handle Ambiguity

Hyuhng Joon KimYouna KimCheonbok ParkJunyeob KimChoonghyun ParkKang Min YooSang-goo LeeTaeuk Kim

article2024EMNLP79 citations

Proposes an alignment framework that enables language models to detect query ambiguity based on their own internal knowledge and proactively ask clarifying questions without degrading performance on unambiguous inputs.

Listen

When users interact with large language models, their queries frequently lack precision or omit key details. While a human might ask for clarification, current artificial intelligence models typically select an arbitrary interpretation and provide a single answer. In high-stakes fields such as healthcare and law, misinterpreting an ambiguous query introduces severe operational, safety, and compliance risks. Standard language models fail to address this issue because they are not explicitly trained to ask for clarification, and their recognition of ambiguity depends entirely on their internal knowledge scope.

The article introduces and evaluates a four-stage training framework called Alignment with Perceived Ambiguity (APA). This approach teaches models to explicitly identify ambiguous inputs and request clarification by evaluating ambiguity from the model's own perspective, rather than relying strictly on pre-defined human labels.

The authors evaluated their approach using open-source models (LLaMA-2 and Mistral) across five question-answering datasets, including AmbigQA, SituatedQA, and three newly introduced benchmarks (AmbigTriviaQA, AmbigWebQuestions, and AmbigFreebaseQA). The framework operates by first identifying queries that a model fails to answer correctly, prompting the model to add clarifying details, and measuring the resulting drop in model uncertainty (information gain). Queries with high information gain are classified as ambiguous from the model's perspective. The model is then fine-tuned to answer clear questions directly while generating explicit clarification requests for queries it perceives as ambiguous.

Across all test benchmarks, APA consistently outperformed standard prompting methods, uncertainty sampling techniques, and models fine-tuned on full human-annotated datasets. Prompting alone proved ineffective, either failing to catch ambiguous queries entirely or excessively triggering clarification requests for clear questions. Models trained with APA achieved up to a 6-point improvement in ambiguity detection accuracy while maintaining high accuracy on clear queries. Furthermore, APA required only 13% to 32% of the training data used by full-dataset fine-tuning, demonstrating that training on model-perceived ambiguity is more effective than training on much larger volumes of static, human-labeled data.

These findings indicate that aligning models to their own internal knowledge boundaries substantially improves reliability and safety. By requesting clarification only when necessary, models reduce the risk of confident but incorrect assumptions without degrading everyday performance. The efficiency of the data selection process also highlights that high-quality, model-tailored data curation can reduce training compute costs and time compared to traditional large-scale fine-tuning.

Organizations deploying conversational systems should transition away from simple ambiguity-prompting strategies and adopt training pipelines that leverage model uncertainty and perceived ambiguity. System designers should evaluate whether to use fixed clarification templates, which offer slightly higher detection stability, or generated clarifications, which provide a more tailored user experience. Because the study focused on single-turn, short-form questions using small-to-medium open-source models, further validation is recommended. Organizations should conduct pilot tests before applying this framework to multi-turn conversational agents, long-form reasoning tasks, and proprietary commercial models.

arXiv: 2404.11972heyjoonkim/APA
Cover for Aligning Language Models to Explicitly Handle Ambiguity

Abstract

In interactions between users and language model agents, user utterances frequently exhibit ellipsis (omission of words or phrases) or imprecision (lack of exactness) to prioritize efficiency. This can lead to varying interpretations of the same input based on different assumptions or background knowledge. It is thus crucial for agents to adeptly handle the inherent ambiguity in queries to ensure reliability. However, even state-of-the-art large language models (LLMs) still face challenges in such scenarios, primarily due to the following hurdles: (1) LLMs are not explicitly trained to deal with ambiguous utterances; (2) the degree of ambiguity perceived by the LLMs may vary depending on the possessed knowledge. To address these issues, we propose Alignment with Perceived Ambiguity (APA), a novel pipeline that aligns LLMs to manage ambiguous queries by leveraging their own assessment of ambiguity (i.e., perceived ambiguity). Experimental results on question-answering datasets demonstrate that APA empowers LLMs to explicitly detect and manage ambiguous queries while retaining the ability to answer clear questions. Furthermore, our finding proves that APA excels beyond training with gold-standard labels, especially in out-of-distribution scenarios. The data and code are available at https://github.com/heyjoonkim/APA.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Initial Prediction Assessment
  • 3.2 Perceived Ambiguity Detection
  • 3.3 Response Construction
  • 3.4 Supervised Fine-Tuning (SFT)
  • 4 Experimental Setting
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Evaluation Metrics
  • 4.4 Implementation Details
  • 5 Experimental Results
  • 6 Ablation Study
  • 6.1 Analysis on Sample-level Misalignment
  • 6.2 The Effect of Threshold Values
  • 6.3 Impact of INFOGAIN for Data Selection
  • 6.4 Case Study
  • 7 Conclusion
  • Limitations
  • Acknowledgement
  • References
  • A Implementations Details
  • A.1 Pipeline Details
  • A.2 Training Details
  • B Dataset Overview
  • B.1 Dataset Details
  • B.2 Dataset Construction Details
  • C Baseline Details
  • D Evaluation Details
  • D.1 Unambiguous Query Evaluation
  • D.2 Ambiguous Query Evaluation
  • E Details of Ablation Experiments
  • E.1 Details of Sample-level Misalignment Analysis
  • E.2 Details of Threshold Ablation
  • E.3 Details of Data Selection Ablation
  • F Additional Case Studies
  • F.1 Failure Cases Before Alignment
  • F.2 Case Study of Disambiguations
  • F.3 Failure Cases of Clarification Request Generation

Knowls

  1. Knowl 1 — Ambiguity is handled by asking the user to resolve underspecification

    definition

    For this work, an ambiguous query is an expression with multiple possible denotations, and the same wording may be interpreted differently depending on a model’s knowledge. The target behavior differs by query type: answer an unambiguous query, but request clarification for an ambiguous one. The approach assumes that the user is best placed to specify their intended meaning; choosing one valid answer arbitrarily may miss that intent, while presenting every possible answer may be impractical.

  2. Knowl 2 — APA aligns a model using its own perceived ambiguity

    model/method

    Alignment with Perceived Ambiguity (APA) is a four-stage supervised alignment pipeline for question answering. First, the model’s current outputs are checked against the training labels, separating correctly handled examples from errors. Second, each error is prompted for a self-disambiguation: the model either adds details using its intrinsic knowledge or repeats a clear query. The change in predictive uncertainty between the original and disambiguated query is used to select queries the model perceives as ambiguous. Third, APA obtains a clarification-request target for each selected query, either from a fixed set of generic requests or by generating a request specific to the ambiguity. Fourth, the model is fine-tuned on these selected ambiguity examples together with examples it already handled correctly, so it learns both to clarify perceived ambiguity and to retain answer behavior on clear queries.

  3. Knowl 3 — Information gain estimates model-perceived ambiguity

    equation

    APA prompts a model to produce a disambiguated version of a query and measures whether the model is more certain about that version. For a token sequence zz of length NzN_z, let VV be the model vocabulary and let pz,i(v)p_{z,i}(v) be the model’s probability for vocabulary token vv at position ii, conditioned on the preceding context and the inference prompt. The token entropy and average sequence entropy are

    Hz,i=−∑v∈Vpz,i(v)log⁡pz,i(v),Hz=1Nz∑i=1NzHz,i.H_{z,i}=-\sum_{v\in V}p_{z,i}(v)\log p_{z,i}(v),\qquad H_z=\frac{1}{N_z}\sum_{i=1}^{N_z}H_{z,i}.

    For an input query xx and its model-generated disambiguation x^disambig\hat{x}_{\mathrm{disambig}}, APA defines information gain as

    INFOGAINx,x^disambig=Hx−Hx^disambig.\mathrm{INFOGAIN}_{x,\hat{x}_{\mathrm{disambig}}}=H_x-H_{\hat{x}_{\mathrm{disambig}}}.

    A query is selected as perceived ambiguous when this value exceeds threshold ϵ\epsilon. The experiments use ϵ=0.1\epsilon=0.1 by default. A large positive value indicates that the model became more certain after adding specificity; a small value indicates that the disambiguation did not appreciably reduce its uncertainty.

  4. Knowl 4 — APA trains on balanced perceived-ambiguity and correct-answer examples

    model/method

    Let DcorrectD_{\mathrm{correct}} contain query–target pairs the original model handled correctly, and let DambigD_{\mathrm{ambig}} contain queries selected by information gain, paired with clarification-request targets. APA balances the two sets to equal size: if DcorrectD_{\mathrm{correct}} is larger, it randomly downsamples that set; if DambigD_{\mathrm{ambig}} is larger, it retains the ambiguous examples with the largest information gain. The resulting training set contains both the answer targets for clear queries and clarification targets for selected ambiguous queries.

    For model parameters θ\theta, query xx, target token sequence yy, and fixed inference template t(x)t(x), training minimizes token-level negative log likelihood:

    min⁡θ∑(x,y)∈D∑i=1∣y∣−log⁡Mθ(yi∣y<i,t(x)).\min_{\theta}\sum_{(x,y)\in D}\sum_{i=1}^{|y|}-\log M_{\theta}(y_i\mid y_{<i},t(x)).

    APA has two target-generation variants. APAFIXED randomly assigns one of several generic clarification phrases to each selected query. APAGEN uses the query and its disambiguation to generate a clarification request that names the factor needing specification.

  5. Knowl 5 — Three additional datasets test ambiguity outside the training domain

    data/table

    The paper constructs AmbigTriviaQA, AmbigWebQuestions, and AmbigFreebaseQA from the corresponding question-answer datasets. GPT-4o is prompted to introduce ambiguity by broadening a noun phrase or omitting crucial details while leaving the rest of the question unchanged. GPT-4o then verifies whether the altered question has multiple valid answers, and human annotators review the verified examples to reduce reliance on model-only validation. Each resulting dataset has 500 ambiguous and 500 unambiguous question-answer pairs.

    The complete evaluation set also includes SituatedQA’s geographical and temporal ambiguity splits. AmbigQA is the in-domain training and validation dataset; the other datasets are held-out tests.

    Dataset Train unambiguous Validation/test unambiguous Validation/test ambiguous
    AmbigQA 5,287 830 1,172
    SituatedQA-Geo – 506 129
    SituatedQA-Temp – 2,795 876
    AmbigTriviaQA – 500 500
    AmbigWebQuestions – 500 500
    AmbigFreebaseQA – 500 500
  6. Knowl 6 — Evaluation measures both answering clear queries and detecting ambiguous ones

    experimental setup

    The experiments use Llama 2 7B, Llama 2 13B, and Mistral 7B, with QLoRA fine-tuning and results averaged over three random seeds. Training uses AdamW with batch size 32; learning rates are selected from {10−3,5×10−4,10−4}\{10^{-3},5\times10^{-4},10^{-4}\} and epoch counts from {1,2,3}\{1,2,3\} using validation performance. QLoRA uses rank r=4r=4 and α=16\alpha=16.

    Evaluation distinguishes five outcomes: (1) a clarification request for an ambiguous query, (2) a non-clarification response to an ambiguous query, (3) a correct answer to an unambiguous query, (4) an incorrect answer to an unambiguous query, and (5) an unnecessary clarification request for an unambiguous query. Unambiguous answers are judged by ROUGE-L against the valid answers, with scores above 0.3 counted as correct. Ambiguous-query clarification is detected using ambiguity-related terms in the output.

    The paper reports unambiguous prediction F1 (F1uF1_u) and ambiguity detection F1 (F1aF1_a). Using the outcome numbers above, F1uF1_u is the harmonic mean of precision 3/(2+3+4)3/(2+3+4) and recall 3/(3+4+5)3/(3+4+5); F1aF1_a is the harmonic mean of precision 1/(1+5)1/(1+5) and recall 1/(1+2)1/(1+2). Thus, the measures account respectively for answer behavior on clear queries and clarification behavior on ambiguous queries, while penalizing unnecessary clarification.

  7. Knowl 7 — APA achieves the strongest ambiguity-detection F1 on every tested model–dataset pair

    empirical result

    Across the three tested models and five held-out datasets, the best APA variant has higher F1aF1_a than the strongest non-APA trained baseline in every model–dataset pair. Scores below are on the paper’s reported 0–100 F1 scale and are means over three seeds. The comparison baseline is the highest-scoring among FULL-SET, SUBSETRAND, and SUBSETENT; the APA column gives the better of APAFIXED and APAGEN.

    Model Dataset Best baseline Best APA APA variant
    Llama 2 7B SituatedQA-Geo 41.45 43.10 APAGEN
    Llama 2 7B SituatedQA-Temp 36.92 41.89 APAGEN
    Llama 2 7B AmbigTriviaQA 74.98 75.74 APAGEN
    Llama 2 7B AmbigWebQuestions 75.86 77.07 APAFIXED
    Llama 2 7B AmbigFreebaseQA 83.89 84.90 APAGEN
    Mistral 7B SituatedQA-Geo 41.28 42.07 APAGEN
    Mistral 7B SituatedQA-Temp 37.01 43.95 APAFIXED
    Mistral 7B AmbigTriviaQA 76.38 83.48 APAFIXED
    Mistral 7B AmbigWebQuestions 74.72 81.07 APAFIXED
    Mistral 7B AmbigFreebaseQA 84.67 90.12 APAFIXED
    Llama 2 13B SituatedQA-Geo 38.39 40.23 APAFIXED
    Llama 2 13B SituatedQA-Temp 37.84 42.18 APAFIXED
    Llama 2 13B AmbigTriviaQA 76.82 80.99 APAFIXED
    Llama 2 13B AmbigWebQuestions 77.43 79.22 APAFIXED
    Llama 2 13B AmbigFreebaseQA 83.00 87.61 APAGEN

    The APA advantage over the strongest trained baseline ranges from 0.79 to 7.10 F1aF1_a points. Both APA variants also improve F1uF1_u over direct prompting in the tested settings, indicating that the gain in clarification behavior is not simply a shift to clarifying every query. APAFIXED is usually stronger than APAGEN, which the authors attribute to the extra difficulty of generating ambiguity-specific requests.

  8. Knowl 8 — Selecting examples by perceived ambiguity outperforms selection restricted to gold ambiguous queries

    data/table

    This selection ablation uses Llama 2 7B and reports mean F1aF1_a with standard deviation over three seeds. RAND randomly chooses ground-truth ambiguous training queries; MAX chooses the ground-truth ambiguous queries with the largest information gain; MIN chooses those with the smallest information gain. APA instead selects the largest-information-gain queries regardless of their ground-truth ambiguity, so it can include queries labeled unambiguous in the source data.

    Selection SituatedQA-Geo SituatedQA-Temp AmbigTriviaQA AmbigFreebaseQA
    RAND 39.31 (1.28) 38.34 (0.44) 72.05 (0.58) 81.28 (1.88)
    MIN 34.95 (1.71) 36.03 (0.90) 70.30 (1.50) 79.19 (2.02)
    MAX 40.96 (0.71) 39.33 (0.88) 73.95 (1.03) 82.23 (1.31)
    APA 43.10 (0.39) 41.89 (2.02) 75.74 (1.52) 84.90 (0.40)

    APA has the highest score on all four datasets. MAX also exceeds RAND, supporting information-gain-based selection even when selection is restricted to gold ambiguous examples. MIN is lowest in every dataset, consistent with training low-information-gain examples as ambiguous being a poor match to the model’s perceived ambiguity.

  9. Knowl 9 — APA’s ambiguity-detection advantage persists across selection thresholds while limiting unnecessary clarification

    empirical result

    The threshold ablation varies ϵ\epsilon over {0.1,0.3,0.5,0.7,0.9}\{0.1,0.3,0.5,0.7,0.9\}. In the reported Llama 2 7B setup, the number of training samples is 3,088 at thresholds 0.1 and 0.3, 1,860 at 0.5, 886 at 0.7, and 396 at 0.9. As the threshold rises, fewer examples remain eligible; performance generally declines with the reduced training data, but APA continues to outperform the compared baselines in F1aF1_a across the tested thresholds.

    The paper also measures Misaligned Clarification Request Rate (MCR): among unambiguous queries the base model answered correctly before alignment, it is the proportion that an aligned model turns into incorrect clarification requests. Lower MCR means better retention of previously successful answer behavior. APA has the lowest MCR among the trained methods across the reported datasets, indicating that its ambiguity alignment causes fewer such regressions.

  10. Knowl 10 — The evaluation covers short-form, single-query QA and a limited set of model and alignment families

    limitation

    The study is limited to short-form question answering with a single query and does not evaluate ambiguity that depends on conversational context or long-form generation. The experiments cover Llama 2 and Mistral rather than a broad range of architectures and model scales. APA is evaluated as supervised fine-tuning; the paper does not establish whether the same approach or results carry over to preference-based alignment methods such as RLHF or DPO.

Coverage note — The paper’s case studies of individual disambiguation and clarification-generation failures were not made separate knowls because they illustrate, rather than add generalizable results beyond, the pipeline and evaluation findings captured here.

References

  1. 1.Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Xiaochun Cao, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, and Saeid Nahavandi. 2021. A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion, 76:243–297.
  2. 2.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  3. 3.Alfonso Amayuelas, Liangming Pan, Wenhu Chen, and William Wang. 2023. Knowledge of knowledge: Exploring known-unknowns uncertainty with large language models. Preprint, arXiv:2305.13712.
  4. 4.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022. Constitutional ai: Harmlessness from ai feedback. Preprint, arXiv:2212.08073.
  5. 5.Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1533–1544, Seattle, Washington, USA. Association for Computational Linguistics.
  6. 6.Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Singh Bedi, and Mengdi Wang. 2024. Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences. Preprint, arXiv:2402.08925.
  7. 7.Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2024. Alpagasus: Training a better alpaca model with fewer data. In The Twelfth International Conference on Learning Representations.
  8. 8.Jonathan H Choi. 2024. Measuring clarity in legal text. U. Chi. L. Rev., 91:1.
  9. 9.Xu Chu, Ihab F. Ilyas, Sanjay Krishnan, and Jiannan Wang. 2016. Data cleaning: Overview and emerging challenges. In Proceedings of the 2016 International Conference on Management of Data, SIGMOD ’16, page 2201–2206, New York, NY, USA. Association for Computing Machinery.
  10. 10.Jeremy Cole, Michael Zhang, Daniel Gillick, Julian Eisenschlos, Bhuwan Dhingra, and Jacob Eisenstein. 2023. Selectively answering ambiguous questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 530–543, Singapore. Association for Computational Linguistics.
  11. 11.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. Preprint, arXiv:2305.14314.
  12. 12.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3029–3051, Singapore. Association for Computational Linguistics.
  13. 13.Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, KaShun SHUM, and Tong Zhang. 2023. RAFT: Reward ranked finetuning for generative foundation model alignment. Transactions on Machine Learning Research.
  14. 14.Romina Etezadi and Mehrnoush Shamsfard. 2023. The state of the art in open domain complex question answering: a survey. Applied Intelligence, 53(4):4124–4144.
  15. 15.H.A. Gleason. 1963. Linguistics and English Grammar. H.A. Gleason jr.
  16. 16.Meiqi Guo, Mingda Zhang, Siva Reddy, and Malihe Alikhani. 2021. Abg-coQA: Clarifying ambiguity in conversational question answering. In 3rd Conference on Automated Knowledge Base Construction.
  17. 17.Benjamin M Gyori, Charles Tapley Hoyt, and Albert Steppi. 2022. Gilda: biomedical entity text normalization with machine-learned disambiguation as a service. Bioinformatics Advances, 2(1):vbac034.
  18. 18.Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2023a. Beavertails: Towards improved safety alignment of LLM via a human-preference dataset. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  19. 19.Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao. 2023b. Ai alignment: A comprehensive survey. Preprint, arXiv:2310.19852.
  20. 20.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. Preprint, arXiv:2310.06825.
  21. 21.Kelvin Jiang, Dekun Wu, and Hui Jiang. 2019. FreebaseQA: A new factoid QA data set matching trivia-style question-answer pairs with Freebase. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 318–323, Minneapolis, Minnesota. Association for Computational Linguistics.
  22. 22.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  23. 23.Daniel Jurafsky. 1996. A probabilistic model of lexical and syntactic access and disambiguation. Cognitive science, 20(2):137–194.
  24. 24.Gangwoo Kim, Sungdong Kim, Byeongguk Jeon, Joonsuk Park, and Jaewoo Kang. 2023a. Tree of clarifications: Answering ambiguous questions with retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 996–1009, Singapore. Association for Computational Linguistics.
  25. 25.Najoung Kim, Phu Mon Htut, Samuel R. Bowman, and Jackson Petty. 2023b. (QA)2: Question answering with questionable assumptions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8466–8487, Toronto, Canada. Association for Computational Linguistics.
  26. 26.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment. In Advances in Neural Information Processing Systems, volume 36, pages 47669–47681. Curran Associates, Inc.
  27. 27.Sushant Kumar, Sumit Datta, Vishakha Singh, Sanjay Kumar Singh, and Ritesh Sharma. 2024. Opportunities and challenges in data-centric ai. IEEE Access.
  28. 28.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  29. 29.Dongryeol Lee, Segwang Kim, Minwoo Lee, Hwanhee Lee, Joonsuk Park, Sang-Woo Lee, and Kyomin Jung. 2023. Asking clarification questions to handle ambiguity in open-domain QA. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11526–11544, Singapore. Association for Computational Linguistics.
  30. 30.Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. 2018. Scalable agent alignment via reward modeling: a research direction. CoRR, abs/1811.07871.
  31. 31.Chin-Yew Lin and Franz Josef Och. 2004. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 605–612, Barcelona, Spain.
  32. 32.Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah Smith, and Yejin Choi. 2023. We’re afraid language models aren’t modeling ambiguity. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 790–807, Singapore. Association for Computational Linguistics.
  33. 33.Wei Liu, Weihao Zeng, Keqing He, Yong Jiang, and Junxian He. 2024a. What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning. In The Twelfth International Conference on Learning Representations.
  34. 34.Zixuan Liu, Xiaolin Sun, and Zizhan Zheng. 2024b. Enhancing llm safety via constrained direct preference optimization. Preprint, arXiv:2403.02475.
  35. 35.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  36. 36.Donald G Mackay and Thomas G Bever. 1967. In search of ambiguity. Perception & Psychophysics, 2:193–200.
  37. 37.A. Majeed and S. Hwang. 2023. Data-centric artificial intelligence, preprocessing, and the quest for transformative artificial intelligence systems development. Computer, 56(05):109–115.
  38. 38.Andrey Malinin and Mark Gales. 2021. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations.
  39. 39.Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. 2022. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft.
  40. 40.Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. AmbigQA: Answering ambiguous open-domain questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5783–5797, Online. Association for Computational Linguistics.
  41. 41.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730–27744.
  42. 42.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  43. 43.Jonathan Pilault, Xavier Garcia, Arthur Bražinskas, and Orhan Firat. 2023. Interactive-chain-prompting: Ambiguity resolution for crosslingual conditional generation with interaction. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 455–483, Nusa Dua, Bali. Association for Computational Linguistics.
  44. 44.Massimo Poesio and Ron Artstein. 2005. The reliability of anaphoric annotation, reconsidered: Taking ambiguity into account. In Proceedings of the Workshop on Frontiers in Corpus Annotations II: Pie in the Sky, pages 76–83, Ann Arbor, Michigan. Association for Computational Linguistics.
  45. 45.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems.
  46. 46.Sanford Schane. 2002. Ambiguity and misunderstanding in the law. T. Jefferson L. Rev., 25:167.
  47. 47.Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: Factoid questions meet long-form answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8273–8288, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  48. 48.Mark Stevenson and Yikun Guo. 2010. Disambiguation in the biomedical domain: the role of ambiguity type. Journal of biomedical informatics, 43(6):972–981.
  49. 49.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  50. 50.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  51. 51.Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn. 2024. Fine-tuning language models for factuality. In The Twelfth International Conference on Learning Representations.
  52. 52.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.
  53. 53.Thomas Wasow, Amy Perfors, and David Beaver. 2005. The puzzle of ambiguity. Morphology and the web of grammar: Essays in memory of Steven G. Lapointe, pages 265–282.
  54. 54.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  55. 55.Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024. Less: Selecting influential data for targeted instruction tuning. Preprint, arXiv:2402.04333.
  56. 56.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024. WizardLM: Empowering large pre-trained language models to follow complex instructions. In The Twelfth International Conference on Learning Representations.
  57. 57.Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023. Alignment for honesty. arXiv preprint arXiv:2312.07000.
  58. 58.Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang. 2023. Do large language models know what they don’t know? In Findings of the Association for Computational Linguistics: ACL 2023, pages 8653–8665, Toronto, Canada. Association for Computational Linguistics.
  59. 59.Yuewei Yuan, Chaitanya Malaviya, and Mark Yatskar. 2023. AmbiCoref: Evaluating human and model sensitivity to ambiguous coreference. In Findings of the Association for Computational Linguistics: EACL 2023, pages 1023–1030, Dubrovnik, Croatia. Association for Computational Linguistics.
  60. 60.Lingxi Zhang, Jing Zhang, Xirui Ke, Haoyang Li, Xinmei Huang, Zhonghui Shao, Shulin Cao, and Xin Lv. 2023. A survey on complex factual question answering. AI Open, 4:1–12.
  61. 61.Michael Zhang and Eunsol Choi. 2021. SituatedQA: Incorporating extra-linguistic contexts into QA. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7371–7387, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  62. 62.Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36.

Citation

MLA
Kim, H. J., et al. “Aligning Language Models to Explicitly Handle Ambiguity”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 1989–2007, https://doi.org/10.18653/v1/2024.emnlp-main.119.
APA
Kim, H. J., Kim, Y., Park, C., Kim, J., Park, C., Yoo, K. M., Lee, S.-. goo ., & Kim, T. (2024). Aligning Language Models to Explicitly Handle Ambiguity. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1989–2007. https://doi.org/10.18653/v1/2024.emnlp-main.119
Chicago
Kim, H. J., Y. Kim, C. Park, et al. 2024. “Aligning Language Models to Explicitly Handle Ambiguity”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1989–2007. https://doi.org/10.18653/v1/2024.emnlp-main.119.
Harvard
Kim, H.J. et al. (2024) “Aligning Language Models to Explicitly Handle Ambiguity”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1989–2007. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.119.
Vancouver
1. Kim HJ, Kim Y, Park C, Kim J, Park C, Yoo KM, Lee S-goo, Kim T (2024) Aligning Language Models to Explicitly Handle Ambiguity. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1989–2007

BibTeX

@inproceedings{kim-etal-2024-aligning,
    title = "Aligning Language Models to Explicitly Handle Ambiguity",
    author = "Kim, Hyuhng Joon  and
      Kim, Youna  and
      Park, Cheonbok  and
      Kim, Junyeob  and
      Park, Choonghyun  and
      Yoo, Kang Min  and
      Lee, Sang-goo  and
      Kim, Taeuk",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.119/",
    doi = "10.18653/v1/2024.emnlp-main.119",
    pages = "1989--2007"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/