MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
Yavuz Faruk BakmanDuygu Nur YaldizBaturalp BuyukatesChenyang TaoDimitrios DimitriadisSalman Avestimehr
Presents Meaning-Aware Response Scoring (MARS), a framework that weights each token by its semantic contribution to the answer rather than applying uniform length normalization, substantially improving uncertainty estimation and error detection across multiple large language models and question-answering benchmarks.
Generative large language models frequently produce incorrect or misleading answers. When organizations deploy these models in high-stakes environments, such as medical advice or professional decision-making, deploying unreliable responses poses severe operational, safety, and reputational risks. Estimating output uncertainty helps determine when an automated response should be trusted or escalated to human review. However, existing uncertainty estimation methods rely on length-normalized scoring, a technique that treats every word in a generated response equally regardless of whether it provides the core answer or merely syntactic filler.
The article introduces and evaluates Meaning-Aware Response Scoring (MARS), a new scoring method designed to improve uncertainty estimation by weighting words based on their semantic contribution to answering the prompt. Rather than dividing token probabilities uniformly by sentence length, the approach identifies phrase boundaries and uses a compact 110-million-parameter neural network to calculate how much removing each phrase alters the response's factual meaning in context. The researchers evaluated this technique across five open-source language models—including Llama-2 (7B and 13B variants), Mistral-7B, and Falcon-7B—tested on three general question-answering benchmarks (TriviaQA, Natural Questions, and WebQA) and a specialized medical dataset.
The findings show that integrating MARS universally improves the performance of all primary probability-based uncertainty estimation techniques. When measuring the ability to distinguish correct from incorrect answers, adding MARS increased predictive accuracy by up to 5.8 points for basic confidence scoring and up to 6.24 points for standard entropy methods. Grouping words into phrases proved substantially more effective than evaluating tokens individually, as it preserves critical contextual relationships. Furthermore, MARS delivers these gains with minimal computational overhead, adding only about 0.8% to 1.5% to the memory and computation footprint during inference because it operates in a single forward pass.
These results demonstrate that organizations can significantly enhance the safety and reliability of generative AI pipelines without incurring prohibitive computational costs or latency delays. By accurately isolating the most informative keywords in an answer, systems can better flag potential hallucinations before they reach end users. While MARS improved uncertainty estimation on medical questions, overall accuracy scores remained lower in the medical domain than on general knowledge tests, illustrating that complex, multi-sentence domain-specific responses introduce additional uncertainty challenges.
Decision-makers should consider adopting meaning-aware weighting in place of standard length normalization within existing risk-management and automated evaluation workflows. Prior to deploying such methods in specialized or safety-critical fields like healthcare and law, organizations should conduct domain-specific pilot testing and validation. Users should remain cautious, as these techniques improve error detection but do not achieve absolute accuracy or eliminate underlying model biases, and they have thus far been validated primarily on English, closed-ended question-answering tasks.
- Paper: Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus, Tianhang Zhang et al. (2023). Its reference-free hallucination detector weights informative keywords and word probabilities, providing the closest prior approach to MARS’s meaning-sensitive response scoring.
- Paper: A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks, Dan Hendrycks et al. (2017). This establishes maximum softmax probability as a baseline for detecting model errors, a basic probability-based confidence signal that MARS reweights.
- Paper: LUQ: Long-text Uncertainty Quantification for LLMs, Caiqi Zhang et al. (2024). LUQ extends uncertainty estimation to long-form outputs by scoring sentence-level consistency, a natural next step after MARS’s meaning-aware scoring of answers.
- Paper: Language Models with Conformal Factuality Guarantees, Christopher Mohri et al. (2024). Conformal factuality applies uncertainty scoring at the claim level to selectively remove unreliable content, extending MARS’s error-detection goal into controlled factual guarantees.
- Paper: Rationale-Guided Retrieval Augmented Generation for Medical Question Answering, Jiwoong Sohn et al. (2025). RAG2 uses model-uncertainty signals to filter evidence in medical question answering, extending uncertainty estimation into a practical reliability intervention.
