MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

Yavuz Faruk BakmanDuygu Nur YaldizBaturalp BuyukatesChenyang TaoDimitrios DimitriadisSalman Avestimehr

article2024ACL65 citations

Presents Meaning-Aware Response Scoring (MARS), a framework that weights each token by its semantic contribution to the answer rather than applying uniform length normalization, substantially improving uncertainty estimation and error detection across multiple large language models and question-answering benchmarks.

Listen

Generative large language models frequently produce incorrect or misleading answers. When organizations deploy these models in high-stakes environments, such as medical advice or professional decision-making, deploying unreliable responses poses severe operational, safety, and reputational risks. Estimating output uncertainty helps determine when an automated response should be trusted or escalated to human review. However, existing uncertainty estimation methods rely on length-normalized scoring, a technique that treats every word in a generated response equally regardless of whether it provides the core answer or merely syntactic filler.

The article introduces and evaluates Meaning-Aware Response Scoring (MARS), a new scoring method designed to improve uncertainty estimation by weighting words based on their semantic contribution to answering the prompt. Rather than dividing token probabilities uniformly by sentence length, the approach identifies phrase boundaries and uses a compact 110-million-parameter neural network to calculate how much removing each phrase alters the response's factual meaning in context. The researchers evaluated this technique across five open-source language models—including Llama-2 (7B and 13B variants), Mistral-7B, and Falcon-7B—tested on three general question-answering benchmarks (TriviaQA, Natural Questions, and WebQA) and a specialized medical dataset.

The findings show that integrating MARS universally improves the performance of all primary probability-based uncertainty estimation techniques. When measuring the ability to distinguish correct from incorrect answers, adding MARS increased predictive accuracy by up to 5.8 points for basic confidence scoring and up to 6.24 points for standard entropy methods. Grouping words into phrases proved substantially more effective than evaluating tokens individually, as it preserves critical contextual relationships. Furthermore, MARS delivers these gains with minimal computational overhead, adding only about 0.8% to 1.5% to the memory and computation footprint during inference because it operates in a single forward pass.

These results demonstrate that organizations can significantly enhance the safety and reliability of generative AI pipelines without incurring prohibitive computational costs or latency delays. By accurately isolating the most informative keywords in an answer, systems can better flag potential hallucinations before they reach end users. While MARS improved uncertainty estimation on medical questions, overall accuracy scores remained lower in the medical domain than on general knowledge tests, illustrating that complex, multi-sentence domain-specific responses introduce additional uncertainty challenges.

Decision-makers should consider adopting meaning-aware weighting in place of standard length normalization within existing risk-management and automated evaluation workflows. Prior to deploying such methods in specialized or safety-critical fields like healthcare and law, organizations should conduct domain-specific pilot testing and validation. Users should remain cautious, as these techniques improve error detection but do not achieve absolute accuracy or eliminate underlying model biases, and they have thus far been validated primarily on English, closed-ended question-answering tasks.

Cover for MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

Abstract

Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. However, their tendency to produce inaccurate or misleading outputs poses a potential risk, particularly in high-stakes environments. Therefore, estimating the correctness of generative LLM outputs is an important task for enhanced reliability. Uncertainty Estimation (UE) in generative LLMs is an evolving domain, where SOTA probability-based methods commonly employ length-normalized scoring. In this work, we propose Meaning-Aware Response Scoring (MARS) as an alternative to length-normalized scoring for UE methods. MARS is a novel scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. We demonstrate that integrating MARS into UE methods results in a universal and significant improvement in UE performance. We conduct experiments using three distinct closed-book question-answering datasets across five popular pre-trained LLMs. Lastly, we validate the efficacy of MARS on a Medical QA dataset. Code can be found here.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Bayesian View to Estimate Uncertainty
  • 2.2 Uncertainty Estimation (UE) of Auto-Regressive Generative Models
  • 2.3 Length-Normalized Scoring
  • 2.4 Entropy-Based UE for Generative LLMs
  • 3 Method
  • 3.1 Key Intuition
  • 3.2 Meaning-Aware Response Scoring
  • 3.3 Importance Function Design
  • 4 Understanding Generative LLM Probabilities from a Classification Perspective
  • 5 Experiments
  • 5.1 Experimental Design
  • 5.2 Main Results
  • 5.3 Ablation Studies
  • 5.4 Effect of Sampling Hyperparameters
  • 5.5 UE in Medical QA Dataset
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics Statement
  • References
  • A Related Works
  • A.1 Discussion of the Differences with TokenSAR
  • B Training of BERT-like Model for Importance Function
  • B.1 Dividing a Sentence to Phrases
  • B.2 Pseudocode of the Importance Function Algorithm
  • C Experimental Details

Knowls

  1. Knowl 1 — MARS replaces uniform length normalization with meaning-weighted token scores

    equation

    Meaning-Aware Response Scoring (MARS) scores a generated response by raising each token’s conditional probability to a weight that combines sequence-length normalization with that token’s semantic importance:

    Pˉ(s∣x,θ)=∏l=1LP(sl∣s<l,x;θ)w(s,x,L,l),w(s,x,L,l)=12L+u(s,x,l)2.\bar{P}(s\mid x,\theta)=\prod_{l=1}^{L}P(s_l\mid s_{<l},x;\theta)^{w(s,x,L,l)},\qquad w(s,x,L,l)=\frac{1}{2L}+\frac{u(s,x,l)}{2}.

    Here, xx is the question, s=(s1,…,sL)s=(s_1,\ldots,s_L) is the generated token sequence of length LL, s<ls_{<l} is its prefix before position ll, and θ\theta denotes the generative language model parameters. The conditional probability is the model’s probability of generating token sls_l after that prefix and question. The importance function u(s,x,l)u(s,x,l) assigns a value in [0,1][0,1] to each token and is normalized so that ∑l=1Lu(s,x,l)=1\sum_{l=1}^{L}u(s,x,l)=1; consequently, the weights ww also sum to one. MARS is an auxiliary scoring function, not a claim that the score is a probability distribution. It replaces length-normalized sequence scores in probability-based uncertainty estimators, including confidence, entropy, and semantic entropy.

  2. Knowl 2 — Token importance is estimated by measuring phrase-masking effects on answer correctness

    model/method

    MARS assigns importance according to how much removing a phrase from an answer changes its judged correctness in the question context. For each question xx and generated answer ss, first divide the answer into phrases. For a phrase hkh_k, mask its tokens to form s∖hks_{\setminus h_k}, then give a BERT Matching (BEM) answer evaluator the question, the original answer ss as the reference answer, and the masked answer. If the evaluator returns correctness score ok∈[0,1]o_k\in[0,1], the phrase’s preliminary importance is 1−ok1-o_k: a large drop in judged correctness indicates an important phrase. Divide that phrase score equally among its tokens, then apply softmax normalization with temperature τ=0.01\tau=0.01 to obtain token importance coefficients whose total is one. The method uses phrases rather than masking individual tokens so that semantically linked tokens are evaluated together.

  3. Knowl 3 — A 110-million-parameter model predicts phrase spans and importance in one pass

    model/method

    To avoid running a separate BEM evaluation for every phrase, the authors train a BERT-like model with 110 million parameters to predict phrase boundaries and token-importance scores jointly in one forward pass over the question and answer. It adapts BERT-base-uncased by removing the final layer and adding two independent fully connected heads: a two-logit head for phrase-boundary labels (“Begin Phrase” and “Inside Phrase”) and a one-logit-per-token head for importance. Training targets use Flair phrase-chunking labels and BEM-derived importance coefficients. The training data comprises 69,192 TriviaQA training questions and 87,925 Natural Questions training questions; four 7B models generate answers to those questions. The model is trained for one epoch with learning rate 5×10−55\times10^{-5} and batch size 32, using an equally weighted combination of cross-entropy phrase-chunking loss and negative log-likelihood importance loss. The authors report that this model adds approximately 1.5% computational and memory overhead for 7B LLMs and 0.8% for 13B LLMs.

  4. Knowl 4 — A semantic-outcome view relates MARS to existing generative uncertainty scores

    model/method

    The paper frames answer uncertainty through a random variable YY representing the meaning of a generated response in its question context. A meaning function g(s,x)g(s,x) maps response ss and question xx to that outcome; a well-calibrated distribution over YY should assign greater probability to meanings corresponding to more likely correct answers. Under this view, length-normalized scoring treats distinct response strings as distinct outcomes, while semantic entropy groups strings that express the same meaning and sums their scores within a meaning cluster. MARS offers another scoring rule for constructing such outcome scores: it replaces uniform length-normalized token weighting with weights that emphasize tokens important to correctness. The claim that this may improve calibration is the paper’s rationale for MARS, not a guarantee that calibration is achieved.

  5. Knowl 5 — Main evaluation covers three closed-book QA datasets and five open LLMs

    experimental setup

    The main evaluation uses TriviaQA, Natural Questions (NaturalQA), and WebQA, with five models: Llama2-7b, Llama2-7b-chat, Mistral-7b, Falcon-7b, and Llama2-13b. The evaluation subsets contain 8,000 TriviaQA validation examples, 3,610 Natural Questions validation examples, and 6,642 WebQA examples formed by combining its training and test splits. Responses use a shared two-shot prompt and greedy decoding (num_beams=1\texttt{num\_beams}=1), with a period set as the end-of-sequence token. Entropy-based methods use five sampled responses at temperature 0.5. The compared estimators are negative length-normalized confidence, entropy, and semantic entropy, each evaluated both with length-normalized scoring and with MARS in its place. GPT-3.5-turbo judges response correctness; AUROC measures how well uncertainty scores distinguish correct from incorrect answers, with higher scores indicating better discrimination. Reported results are from a single run.

  6. Knowl 6 — MARS improves all three estimators across the main QA evaluation

    data/table

    The following are the reported AUROC scores for the baseline estimators and their MARS versions. Each dataset block compares the same five LLMs; higher AUROC indicates better prediction of answer correctness. MARS improves every listed baseline/model/dataset combination. The reported maximum gains are 5.8 AUROC points for confidence, 6.24 for entropy, and 1.51 for semantic entropy.

    DatasetMethodLlama2-7bLlama2-7b-chatMistral-7bFalcon-7bLlama2-13b
    TriviaQAConfidence70.1870.4072.5568.4768.19
    TriviaQAEntropy69.7069.9472.5769.1069.04
    TriviaQASemantic entropy (SE)81.1076.1982.1776.7879.49
    TriviaQAConfidence + MARS75.0674.2377.9772.9573.99
    TriviaQAEntropy + MARS75.9473.8278.5172.8774.95
    TriviaQASE + MARS82.2277.6783.6377.4881.00
    NaturalQAConfidence68.5665.9869.5463.7868.56
    NaturalQAEntropy67.0865.2368.0563.2868.34
    NaturalQASE72.4768.6675.1270.4173.56
    NaturalQAConfidence + MARS69.8167.8671.3668.3070.88
    NaturalQAEntropy + MARS69.3267.4170.7167.5170.63
    NaturalQASE + MARS72.7569.4375.5071.2473.89
    WebQAConfidence64.7664.0665.6666.5662.60
    WebQAEntropy64.0463.8264.1565.9862.11
    WebQASE69.4467.1169.5173.1667.31
    WebQAConfidence + MARS66.0464.4867.1668.2664.23
    WebQAEntropy + MARS65.8364.6965.7668.4464.02
    WebQASE + MARS69.8867.2769.8673.5767.75
  7. Knowl 7 — Phrase-level importance outperforms token masking, and equal phrase allocation works well

    empirical result

    A TriviaQA ablation compares token-level versus phrase-level importance assignment, and three ways to distribute a phrase’s importance among its tokens: equally, only to the most uncertain token (“Max”), or only to the least uncertain token (“Min”). Phrase-level assignment achieves higher AUROC than token-level assignment for every estimator and both tested models. Equal allocation and allocation to the most uncertain token perform similarly, whereas allocation to the least uncertain token is weaker.

    Importance granularityMethod + MARSLlama2-7bMistral-7b
    TokenConfidence72.5375.31
    TokenEntropy74.4677.58
    TokenSE81.5583.25
    PhraseConfidence75.0677.97
    PhraseEntropy75.9478.51
    PhraseSE82.2283.63
    Phrase-token allocationConfidence + MARS (Llama2-7b / Mistral-7b)Entropy + MARS (Llama2-7b / Mistral-7b)SE + MARS (Llama2-7b / Mistral-7b)
    Min69.92 / 72.2070.56 / 72.7581.67 / 82.33
    Max75.13 / 77.7377.11 / 79.2282.07 / 83.62
    Equal75.06 / 77.9775.94 / 78.5182.22 / 83.63
  8. Knowl 8 — MARS gains persist across sampling temperatures and sample counts

    empirical result

    Sensitivity experiments on TriviaQA evaluate entropy and semantic entropy, with and without MARS, for Llama2-13b and Mistral-7b under multiple sampling temperatures and numbers of sampled responses. The authors report that MARS improves the sampling-based estimators across the tested temperatures and that its performance remains stable as the number of samples changes. Its advantage is more apparent when fewer responses are sampled, where estimating entropy from samples is more difficult. The experiments also expose a resource trade-off: increasing sample count costs more, and semantic entropy additionally requires natural-language-inference model passes to cluster responses.

  9. Knowl 9 — A curated medical-QA test also shows improvements, but AUROC remains modest

    data/table

    For a medical-domain evaluation, the authors curate 415 objectively answerable questions from the MedMCQA multiple-choice dataset with medical-professional input, removing the multiple-choice format. They evaluate AdaptLLM Medicine-Chat (LLaMA-2-Chat-7B) and use GPT-4 to assess response validity. MARS improves all three probability-based estimators, although the resulting AUROCs are modest compared with the general-knowledge QA results. The authors suggest that medical answers often need longer, more complex explanations than the short answers common in general-knowledge QA, which may hinder these estimators.

    MethodMedicine-Chat-7b AUROC
    Confidence62.41
    Entropy59.58
    SE62.89
    Confidence + MARS62.89
    Entropy + MARS60.33
    SE + MARS64.48
  10. Knowl 10 — Evaluation scope and uncertainty-function supervision limit the claims

    limitation

    The study focuses on English, closed-ended question answering with objective ground-truth answers; it does not establish performance on open-ended answers, other languages, or other specialized domains. The importance function is trained using automatically generated BEM-based importance targets rather than human-assigned importance labels, which the authors identify as a potential improvement direction. The medical experiment also indicates that gains can coexist with low absolute AUROC, and the authors caution that probability-based uncertainty methods do not attain perfect accuracy and may carry model biases into uncertainty estimates.

Coverage note — Deliberately omitted the auxiliary model’s illustrative sample outputs and exact train/validation loss values because they document implementation but add no distinct method or finding; background and related-work material is also excluded.

References

  1. 1.Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1638–1649, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  2. 2.Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. The Falcon series of open language models.
  3. 3.Neil Band, Tim G. J. Rudner, Qixuan Feng, Angelos Filos, Zachary Nado, Michael W Dusenberry, Ghassen Jerfel, Dustin Tran, and Yarin Gal. 2021. Benchmarking Bayesian deep learning on diabetic retinopathy detection tasks. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  4. 4.Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, tomahto. beyond token-level answer equivalence for question answering evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 291–305.
  5. 5.Yingshan Chang, Guihong Cao, Mridu Narang, Jianfeng Gao, Hisami Suzuki, and Yonatan Bisk. 2022. WebQA: Multihop and Multimodal QA. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16474–16483.
  6. 6.Jiuhai Chen and Jonas Mueller. 2023. Quantifying uncertainty in answers from any language model and enhancing their trustworthiness.
  7. 7.Daixuan Cheng, Shaohan Huang, and Furu Wei. 2023. Adapting large language models via reading comprehension.
  8. 8.Roi Cohen, May Hamri, Mor Geva, and Amir Globerson. 2023. LM vs LM: Detecting factual errors via cross examination. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12621–12640, Singapore. Association for Computational Linguistics.
  9. 9.Shrey Desai and Greg Durrett. 2020. Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 295–302, Online. Association for Computational Linguistics.
  10. 10.Jinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, and Kaidi Xu. 2024. Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models.
  11. 11.Marina Fomicheva, Shuo Sun, Lisa Yankovskaya, Frédéric Blain, Francisco Guzmán, Mark Fishel, Nikolaos Aletras, Vishrav Chaudhary, and Lucia Specia. 2020. Unsupervised quality estimation for neural machine translation. Transactions of the Association for Computational Linguistics, 8:539–555.
  12. 12.Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of The 33rd International Conference on Machine Learning, volume 48, pages 1050–1059.
  13. 13.Tilmann Gneiting and Adrian E. Raftery. 2007. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102:359 – 378.
  14. 14.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On calibration of modern neural networks.
  15. 15.Mengting Hu, Zhen Zhang, Shiwan Zhao, Minlie Huang, and Bingzhe Wu. 2023. Uncertainty in natural language processing: Sources, quantification, and applications.
  16. 16.Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. 2023. Benchmarking large language models as AI research agents.
  17. 17.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b.
  18. 18.Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021. How can we know when language models know? On the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962–977.
  19. 19.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14).
  20. 20.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  21. 21.Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022. Language models (mostly) know what they know.
  22. 22.Neema Kotonya and Francesca Toni. 2020. Explainable automated fact-checking for public health claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7740–7754, Online. Association for Computational Linguistics.
  23. 23.Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations.
  24. 24.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  25. 25.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the 31st International Conference on Neural Information Processing Systems, page 6405–6416.
  26. 26.Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023. Generating with confidence: Uncertainty quantification for black-box large language models.
  27. 27.Andrey Malinin and Mark Gales. 2021. Uncertainty estimation in autoregressive structured prediction. In International Conference on Learning Representations.
  28. 28.OpenAI. 2023. GPT-4 Technical Report.
  29. 29.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. MedMCQA: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Proceedings of the Conference on Health, Inference, and Learning, volume 174 of Proceedings of Machine Learning Research, pages 248–260.
  30. 30.Yichen Shen, Zhilu Zhang, Mert R Sabuncu, and Lin Sun. 2021. Real-time uncertainty estimation in computer vision via uncertainty-aware distribution distillation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 707–716.
  31. 31.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  32. 32.Artem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev, Manvel Avetisian, and Leonid Zhukov. 2022. Uncertainty estimation of transformer predictions for misclassification detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8237–8252, Dublin, Ireland. Association for Computational Linguistics.
  33. 33.Tim Z. Xiao, Aidan N. Gomez, and Yarin Gal. 2020. Wat zei je? detecting out-of-distribution translations with variational transformers. CoRR, abs/2006.08344.
  34. 34.Yijun Xiao and William Yang Wang. 2019. Quantifying uncertainties in natural language processing tasks. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):7322–7329.
  35. 35.Yuxin Xiao, Paul Pu Liang, Umang Bhatt, Willie Neiswanger, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022. Uncertainty quantification with pre-trained language models: A large-scale empirical analysis. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 7273–7284, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  36. 36.Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, Jie Zhou, Siming Chen, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023. A comprehensive capability analysis of GPT-3 and GPT-3.5 series models.
  37. 37.Ming Zhu, Aman Ahuja, Da-Cheng Juan, Wei Wei, and Chandan K. Reddy. 2020. Question answering with long multiple-span answers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3840–3849, Online. Association for Computational Linguistics.
  38. 38.Ming Zhu, Aman Ahuja, Wei Wei, and Chandan K. Reddy. 2019. A hierarchical attention retrieval model for healthcare question answering. In The World Wide Web Conference, WWW ’19, page 2472–2482, New York, NY, USA. Association for Computing Machinery.

Citation

MLA
Bakman, Y. F., et al. “MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7752–67, https://doi.org/10.18653/v1/2024.acl-long.419.
APA
Bakman, Y. F., Yaldiz, D. N., Buyukates, B., Tao, C., Dimitriadis, D., & Avestimehr, S. (2024). MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7752–7767. https://doi.org/10.18653/v1/2024.acl-long.419
Chicago
Bakman, Y. F., D. N. Yaldiz, B. Buyukates, C. Tao, D. Dimitriadis, and S. Avestimehr. 2024. “MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7752–67. https://doi.org/10.18653/v1/2024.acl-long.419.
Harvard
Bakman, Y.F. et al. (2024) “MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7752–7767. Available at: https://doi.org/10.18653/v1/2024.acl-long.419.
Vancouver
1. Bakman YF, Yaldiz DN, Buyukates B, Tao C, Dimitriadis D, Avestimehr S (2024) MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7752–7767

BibTeX

@inproceedings{bakman-etal-2024-mars,
    title = "{MARS}: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative {LLM}s",
    author = "Bakman, Yavuz Faruk  and
      Yaldiz, Duygu Nur  and
      Buyukates, Baturalp  and
      Tao, Chenyang  and
      Dimitriadis, Dimitrios  and
      Avestimehr, Salman",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.419/",
    doi = "10.18653/v1/2024.acl-long.419",
    pages = "7752--7767"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/