Built independently by an author, for readers. Read the story and support ChapterPal

keyword

generative LLMs

Generative large language models are artificial intelligence systems trained on massive volumes of text data to generate coherent, contextually relevant natural language and other sequence-based outputs in response to user prompts. Typically built on deep transformer architectures, these models operate autoregressively by estimating probability distributions across vocabularies to predict subsequent tokens in a sequence. They are capable of executing diverse language-based tasks, including question answering, open-ended conversation, summarization, translation, and code synthesis. Because their outputs are derived from probabilistic patterns learned during training rather than explicit knowledge retrieval or reasoning, they can occasionally produce factually incorrect or ungrounded information, making output evaluation, alignment, and uncertainty estimation essential for reliable deployment.

1 item

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs

Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr

OrganizationsAmazonUniversity of Southern California

Why you should read this

Presents Meaning-Aware Response Scoring (MARS), a framework that weights each token by its semantic contribution to the answer rather than applying uniform length normalization, substantially improving uncertainty estimation and error detection across multiple large language models and question-answering benchmarks.

Generative Large Language Models (LLMs) are widely utilized for their excellence in various tasks. However, their tendency to produce inaccurate or misleading outputs poses a potential risk, particularly in high-stakes environments. Therefore, estimating the correctness of generative LLM outputs is an important task for enhanced reliability. Uncertainty Estimation (UE) in generative LLMs is an evolving domain, where SOTA probability-based methods commonly employ length-normalized scoring. In this work, we propose Meaning-Aware Response Scoring (MARS) as an alternative to length-normalized scoring for UE methods. MARS is a novel scoring function that considers the semantic contribution of each token in the generated sequence in the context of the question. We demonstrate that integrating MARS into UE methods results in a universal and significant improvement in UE performance. We conduct experiments using three distinct closed-book question-answering datasets across five popular pre-trained LLMs. Lastly, we validate the efficacy of MARS on a Medical QA dataset. Code can be found here.

Added

2026-10-04