All claims are equal, but some claims are more equal than others: Importance-sensitive factuality evaluation of LLM generations

Miriam WannerLeif AzzopardiPaul ThomasSoham DanBen Van DurmeNick Craswell

article2025arXiv8 citations

Introduces the VITALERRORS benchmark and VITAL metrics to evaluate large language model factuality by weighting claim importance, exposing critical errors in core information that uniform scoring methods overlook.

Listen

As large language models (LLMs) are increasingly integrated into enterprise applications, verifying the factual accuracy of their outputs is critical to preventing misinformation and maintaining user trust. Standard automated evaluation frameworks typically break model outputs into atomic subclaims and measure factual precision and recall by scoring each subclaim equally. However, this equal-weighting approach creates a significant blind spot: an answer that gets peripheral background details correct can still receive a high factual score even if the core answer is entirely wrong or missing. The article addresses this evaluation flaw by proposing an importance-sensitive evaluation methodology.

The main objective of the article is to demonstrate that conventional factuality metrics fail to detect critical factual errors in LLM generations and to introduce an importance-weighted evaluation framework, termed VITAL, that reliably penalizes missing or incorrect key information.

To conduct this evaluation, the researchers constructed VITALERRORS, a benchmark dataset comprising 6,733 queries drawn from six established question-answering and reasoning datasets covering both open-ended queries (such as biographies) and single-answer queries. Using advanced language models, the authors generated baseline responses alongside minimally altered adversarial variants where the single most important piece of information was either omitted or falsified. They then evaluated these responses using standard factuality metrics (FActScore and Nugget Recall) and compared the results against their proposed VITAL metrics, which prompt an evaluator model to categorize decomposed claims and information nuggets by query importance (vital, okay, or less important).

The analysis yielded several key findings. First, existing factuality metrics proved highly insensitive to critical mistakes; for example, standard precision scores dropped by only 5% to 7% when the core answer was falsified in single-answer queries because the overwhelming volume of correct background claims masked the error. Second, the proposed vital precision metric correctly exposed these failures, showing a substantial drop of roughly 24% to 25% on falsified single-answer responses compared to normal outputs. Third, response-level binary metrics—which flag whether any vital claim is false or omitted—proved to be the most effective error detectors, identifying incorrect vital information in nearly 90% of falsified single-answer responses compared to only 45% in normal responses. Finally, detecting omissions of vital information remains more difficult than detecting explicit factual falsehoods across all metric types.

These findings imply that organizations relying on standard fact-checking metrics are exposed to operational and compliance risks by overestimating the reliability of LLM outputs. In enterprise and high-stakes settings, a single wrong key figure or omitted instruction can render an entire response harmful despite high aggregate precision scores. Incorporating importance weighting aligns automated evaluation more closely with human judgment, where errors in core answers cause disproportionate reputational and operational damage.

The article recommends that organizations developing and auditing LLM systems adopt importance-weighted or response-level vital metrics rather than unweighted factual precision. Evaluators should also benchmark their fact-checking pipelines against adversarial test suites like VITALERRORS to identify whether their evaluation systems are masking severe errors. For production deployment, teams should carefully consider weighting schemes tailored to user preferences and specific task domains.

Confidence in these findings is supported by the consistent results across a large sample of 6,733 diverse queries. However, decision-makers should note certain limitations: the VITAL framework requires an extra model step to rank claim importance, adding computational cost and processing latency. Furthermore, because the framework relies on language models to judge importance and verify facts against reference texts, it inherits the underlying models' potential biases and is constrained by the quality and coverage of the retrieved grounding documents.

arXiv: 2510.07083

No sufficiently relevant recommendations were found.

Cover for All claims are equal, but some claims are more equal than others: Importance-sensitive factuality evaluation of LLM generations

Abstract

Existing methods for evaluating the factuality of large language model (LLM) responses treat all claims as equally important. This results in misleading evaluations when vital information is missing or incorrect as it receives the same weight as peripheral details, raising the question: how can we reliably detect such differences when there are errors in key information? Current approaches that measure factuality tend to be insensitive to omitted or false key information. To investigate this lack of sensitivity, we construct VITALERRORS, a benchmark of 6,733 queries with minimally altered LLM responses designed to omit or falsify key information. Using this dataset, we demonstrate the insensitivities of existing evaluation metrics to key information errors. To address this gap, we introduce VITAL, a set of metrics that provide greater sensitivity in measuring the factuality of responses by incorporating the relevance and importance of claims with respect to the query. Our analysis demonstrates that VITAL metrics more reliably detect errors in key information than previous methods. Our dataset, metrics, and analysis provide a foundation for more accurate and robust assessment of LLM factuality.

Table of Contents

  • 1 Introduction
  • 2 Metrics
  • 2.1 Factual Precision
  • 2.2 Nugget Recall
  • 2.3 Comparison of Metrics
  • 3 Methods
  • 3.1 Vital Metrics
  • 3.2 VitalErrors Dataset
  • 3.3 Experiments
  • 4 Results
  • 5 Analysis
  • 5.1 Vital Metrics
  • 5.2 VitalErrors
  • 6 Conclusion
  • References
  • A Prompts, Model Details, Compute
  • B Results
  • B.1 Cumulative Precision
  • B.2 Linear Decay Weighting

Knowls

  1. Knowl 1 — VITAL ranks and labels claims by query importance

    model/method

    VITAL starts with subclaims decomposed from an answer and asks an LLM to rank them by how directly they address the query, independently of whether they are factually correct. It assigns each subclaim one of three labels: vital for essential claims addressing the query’s core, okay for useful but nonessential support, and less important for background or tangential material. The ranking places these categories in that order, with claims ordered by importance within each category. Separating query importance from factual correctness means a false claim that directly answers the question can still be labeled vital; correctness is assessed separately.

  2. Knowl 2 — VITAL decomposition-level precision and recall score only vital content

    equation

    Let V^\widehat{V} be the set of vital subclaims extracted from a response and let VV be the set of vital nuggets generated for its query. A subclaim is supported if it is judged correct against trusted sources; a nugget is supported if the response supports it. VITAL precision and recall are the supported fractions within their respective vital sets: VITALPREC=∣{c∈V^:c is supported}∣/∣V^∣\mathrm{VITALPREC}=|\{c\in\widehat{V}:c\text{ is supported}\}|/|\widehat{V}| and VITALREC=∣{n∈V:n is supported by the response}∣/∣V∣\mathrm{VITALREC}=|\{n\in V:n\text{ is supported by the response}\}|/|V|. These scores exclude non-vital claims or nuggets rather than weighting all information in the response equally.

  3. Knowl 3 — VITAL response-level scores flag any vital error

    equation

    VITAL also reports binary response-level error indicators. For a response with vital subclaims V^\widehat{V} and its query’s vital nuggets VV, VITALRLP=1\mathrm{VITALRLP}=1 if at least one vital subclaim is unsupported, and 00 otherwise; VITALRLR=1\mathrm{VITALRLR}=1 if at least one vital nugget is unsupported by the response, and 00 otherwise. In aggregate, the paper reports the percentage of responses flagged, so lower values indicate fewer responses with a vital error. The page-2 comparison graphic illustrates the distinction: conventional precision and recall can remain high when a long answer omits or falsifies the planet-count answer, while the vital-error indicators distinguish those cases from peripheral errors.

  4. Knowl 4 — VITALERRORS pairs normal answers with minimal key-information attacks

    experimental setup

    VITALERRORS contains 6,733 queries drawn from six datasets, with both open-ended and single-answer questions. For each query, the authors generated a long-form normal LLM response and adversarial missing and wrong variants. The missing variant removes the information most important to answering the query; the wrong variant changes a key piece of information to an incorrect one, generally in one sentence, while leaving the rest of the response unchanged. This design produces minimally altered answers whose central answer is omitted or falsified, enabling evaluation of whether a metric detects key-information errors.

  5. Knowl 5 — The benchmark spans three open-ended and three single-answer datasets

    data/table

    The 6,733 VITALERRORS queries comprise 3,733 open-ended queries and 3,000 single-answer queries. The open-ended portion uses 500 FACTSCORE Bios queries, 2,484 WILDHALLUCINATIONS queries from its Cultural & Entertainment and Geographic subsets, and 749 BRIGHT queries across Biology, Earth Science, Economics, Psychology, Robotics, Stack Overflow, and Sustainable Living. The single-answer portion uses 1,000-query subsets of HotpotQA, Natural Questions (NQ), and TriviaQA. The collection therefore tests key-information errors across both broad prompts with many plausible answers and questions with a specific answer expected.

  6. Knowl 6 — Evaluation uses GPT-4o for claim verification, nugget scoring, and importance ranking

    experimental setup

    The authors compare FACTSCORE, NUGGETRECALL, and VITAL on normal, missing, and wrong responses, reporting results separately for open-ended and single-answer queries. FACTSCORE uses GPT-4o to generate atomic facts and evaluate them; nugget generation, importance labeling, and response evaluation use AutoNuggetizer with GPT-4o and its All Strict recall metric. VITAL adds GPT-4o ranking and labeling of response subclaims by query importance. Normal and adversarial answer generation used GPT-4o at temperature 0.2 with a 2,000-token maximum; importance ranking used temperature 0.2 with a 4,000-token maximum. The reported compute was 15 GPU-hours for response generation and 450 GPU-hours for evaluating with FACTSCORE, nuggets, and VITAL.

  7. Knowl 7 — VITAL vital-only scores separate key errors more clearly than standard precision and recall

    empirical result

    The reported percentages for normal, missing, and wrong responses show that standard metrics often change little after key information is omitted or falsified, while vital-only metrics can show larger differences. For open-ended queries, FACTSCORE precision is 83.62%, 83.58%, and 78.84%; VITALPREC is 82.69%, 82.88%, and 75.40%. NUGGETRECALL is 24.35%, 18.61%, and 23.23%; VITALREC is 40.90%, 29.91%, and 37.89%. For single-answer queries, FACTSCORE precision is 82.58%, 82.75%, and 76.63%; VITALPREC is 72.81%, 72.22%, and 48.73%. NUGGETRECALL is 27.71%, 19.13%, and 23.49%; VITALREC is 52.44%, 27.45%, and 36.70%. In particular, single-answer VITALPREC falls by 24.08 percentage points from normal to wrong, compared with a 5.95-point fall in FACTSCORE. The larger single-answer VITALREC declines also make omissions more apparent than standard nugget recall does.

  8. Knowl 8 — Any-vital-error indicators detect wrong answers especially well for single-answer queries

    empirical result

    The response-level metrics report the percentage of responses flagged for at least one vital error, so lower values are preferable. For open-ended queries, VITALRLP is 52.70% for normal, 45.48% for missing, and 73.79% for wrong responses; VITALRLR is 87.22%, 91.18%, and 88.58%, respectively. For single-answer queries, VITALRLP is 45.03%, 42.17%, and 89.53%; VITALRLR is 54.30%, 68.10%, and 68.47%. Thus VITALRLP flags wrong responses much more often than normal responses, particularly for single-answer questions. VITALRLR separates normal from missing and wrong responses by about 14–15 percentage points for single-answer questions, but its open-ended rates differ by less than five points.

  9. Knowl 9 — Cumulative precision reveals how a wrong key claim is diluted in a long answer

    empirical result

    The page-8 cumulative-precision plot tracks FACTSCORE precision as successive subclaims are added for single-answer responses; the authors report the same trend for open-ended responses. Wrong responses begin with a low cumulative score—described in the text as around 60%—because the falsified key claim occurs early, then gain precision as later, mostly supported claims enter the average. Normal and missing responses retain high cumulative precision across the response. This shows why a whole-response precision score can obscure an incorrect answer when the surrounding long response contains many accurate details.

  10. Knowl 10 — Open-ended queries make any-error evaluation harder than single-answer queries

    empirical result

    The paper attributes weaker separation on open-ended questions partly to their diffuse information needs: many different details can be relevant, and it may be unclear which is most important. Nugget recall also depends on the documents used to create nuggets, so a valid response may not match that particular evidence pool. Open-ended responses consequently have more vital claims and nuggets, creating more chances for an any-error metric to flag at least one unsupported item. For single-answer questions, the expected answer and its supporting nuggets are more constrained, making missing or wrong key information easier to distinguish.

  11. Knowl 11 — VITAL inherits LLM, grounding, and scalability limitations

    limitation

    VITAL relies on LLMs for decomposition, importance ranking, and verification, so biases and inconsistencies in those judgments can affect its scores. Like other factuality metrics, it depends on the quality of the reference material; the study assumes its paired sources are factual and relevant, while poor retrieved evidence could cause false positives or false negatives, especially in domains with sparse references. Importance ranking adds an extra LLM call per response and may limit scalability. The benchmark’s automated perturbations also may not represent the full diversity of real-world factual errors, and the authors caution against treating it as an exhaustive reliability measure or using it alone in high-stakes settings.

Coverage note — The mean subclaim and nugget count breakdown and the supplementary linear-decay weighting results are omitted: the former is descriptive rather than load-bearing, and the latter assumes every higher-ranked item is strictly more important, an assumption the authors explicitly regard as too strong for the proposed VITAL metrics.

References

  1. 1.Zahra Abbasiantaeb, Simon Lupart, Leif Azzopardi, Jeffrey Dalton, and Mohammad Aliannejadi. 2025. Conversational gold: Evaluating personalized conversational search system using gold nuggets. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’25, page 3455–3465, New York, NY, USA. Association for Computing Machinery.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, page 610–623, New York, NY, USA. Association for Computing Machinery.
  3. 3.Arie Cattan, Paul Roit, Shiyue Zhang, David Wan, Roee Aharoni, Idan Szpektor, Mohit Bansal, and Ido Dagan. 2024. Localizing factual inconsistencies in attributable text generation. ArXiv, abs/2410.07473.
  4. 4.Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, and Kai-Wei Chang. 2022. On measures of biases and harms in NLP. In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, pages 246–267, Online only. Association for Computational Linguistics.
  5. 5.Ron Eliav, Arie Cattan, Eran Hirsch, Shahaf Bassan, Elias Stengel-Eskin, Mohit Bansal, and Ido Dagan. 2025. Clatter: Comprehensive entailment reasoning for hallucination detection. Preprint, arXiv:2506.05243.
  6. 6.Ameya Godbole and Robin Jia. 2025. Verify with caution: The pitfalls of relying on imperfect factuality metrics. In Findings of the Association for Computational Linguistics: ACL 2025, pages 22889–22912, Vienna, Austria. Association for Computational Linguistics.
  7. 7.Anisha Gunjal and Greg Durrett. 2024. Molecular facts: Desiderata for decontextualization in LLM fact verification. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 3751–3768, Miami, Florida, USA. Association for Computational Linguistics.
  8. 8.Heyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, and Mark Dredze. 2025. Medscore: Factuality evaluation of free-form medical answers. Preprint, arXiv:2505.18452.
  9. 9.Zhengping Jiang, Jingyu Zhang, Nathaniel Weir, Seth Ebner, Miriam Wanner, Kate Sanders, Daniel Khashabi, Anqi Liu, and Benjamin Van Durme. 2025. Core: Robust factual precision with informative subclaim identification. In Findings of the Association for Computational Linguistics: ACL 2025, pages 19833–19856, Vienna, Austria. Association for Computational Linguistics.
  10. 10.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. arXiv e-prints, arXiv:1705.03551.
  11. 11.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics.
  12. 12.Jimmy Lin and Dina Demner-Fushman. 2006. Will pyramids built of nuggets topple over? In Proceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 383–390, New York City, USA. Association for Computational Linguistics.
  13. 13.Xin Liu, Lechen Zhang, Sheza Munir, Yiyang Gu, and Lu Wang. 2025. Verifact: Enhancing long-form factuality evaluation with refined fact extraction and reference facts. Preprint, arXiv:2505.09701.
  14. 14.Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076–12100, Singapore. Association for Computational Linguistics.
  15. 15.Ani Nenkova and Rebecca Passonneau. 2004. Evaluating content selection in summarization: The pyramid method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 145–152, Boston, Massachusetts, USA. Association for Computational Linguistics.
  16. 16.OpenAI, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, and 400 others. 2024. Gpt-4o system card. Preprint, arXiv:2410.21276.
  17. 17.Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024. Initial nugget evaluation results for the TREC 2024 RAG Track with the AutoNuggetizer Framework. arXiv:2411.09607.
  18. 18.Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025. The great nugget recall: Automating fact extraction and rag evaluation with large language models. Preprint, arXiv:2504.15068.
  19. 19.Chris Samarinas, Alexander Krubner, Alireza Salemi, Youngwoo Kim, and Hamed Zamani. 2025. Beyond factual accuracy: Evaluating coverage of diverse factual information in long-form text generation. Preprint, arXiv:2501.03545.
  20. 20.Yixiao Song, Yekyung Kim, and Mohit Iyyer. 2024. VeriScore: Evaluating the factuality of verifiable claims in long-form text generation. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 9447–9474, Miami, Florida, USA. Association for Computational Linguistics.
  21. 21.Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Quan Shi, Zachary S Siegel, Michael Tang, Ruoxi Sun, Jinsung Yoon, Sercan O Arik, Danqi Chen, and Tao Yu. 2024. Bright: A realistic and challenging benchmark for reasoning-intensive retrieval.
  22. 22.Ellen M. Voorhees. 2003. Overview of the TREC 2003 question answering track. In Proceedings of The Twelfth Text REtrieval Conference, TREC 2003, Gaithersburg, Maryland, USA, November 18-21, 2003, volume 500-255 of NIST Special Publication, pages 54–68. National Institute of Standards and Technology (NIST).
  23. 23.Miriam Wanner, Benjamin Van Durme, and Mark Dredze. 2024a. Dndscore: Decontextualization and decomposition for factuality verification in long-form text generation. Preprint, arXiv:2412.13175.
  24. 24.Miriam Wanner, Seth Ebner, Zhengping Jiang, Mark Dredze, and Benjamin Van Durme. 2024b. A closer look at claim decomposition. In Proceedings of the 13th Joint Conference on Lexical and Computational Semantics (*SEM 2024), pages 153–175, Mexico City, Mexico. Association for Computational Linguistics.
  25. 25.Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, and Quoc V. Le. 2024. Long-form factuality in large language models. In Advances in Neural Information Processing Systems, volume 37, pages 80756–80827. Curran Associates, Inc.
  26. 26.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP).
  27. 27.James Xu Zhao, Jimmy Z.j. Liu, Bryan Hooi, and See-Kiong Ng. 2025. How does response length affect long-form factuality. In Findings of the Association for Computational Linguistics: ACL 2025, pages 3102–3125, Vienna, Austria. Association for Computational Linguistics.
  28. 28.Wenting Zhao, Tanya Goyal, Yu Ying Chiu, Liwei Jiang, Benjamin Newman, Abhilasha Ravichander, Khyathi Chandu, Ronan Le Bras, Claire Cardie, Yuntian Deng, and Yejin Choi. 2024. Wildhallucinations: Evaluating long-form factuality in llms with real-world entity queries. Preprint, arXiv:2407.17468.

Citation

MLA
Wanner, M., et al. “All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations”. arXiv, 2025, http://arxiv.org/abs/2510.07083v1.
APA
Wanner, M., Azzopardi, L., Thomas, P., Dan, S., Durme, B. V., & Craswell, N. (2025). All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations. arXiv. http://arxiv.org/abs/2510.07083v1
Chicago
Wanner, M., L. Azzopardi, P. Thomas, S. Dan, B. V. Durme, and N. Craswell. 2025. “All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations”. arXiv. http://arxiv.org/abs/2510.07083v1.
Harvard
Wanner, M. et al. (2025) “All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2510.07083v1.
Vancouver
1. Wanner M, Azzopardi L, Thomas P, Dan S, Durme BV, Craswell N (2025) All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations. arXiv

BibTeX

@article{wanner2025all,
  title = {All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations},
  author = {Wanner, Miriam and Azzopardi, Leif and Thomas, Paul and Dan, Soham and Durme, Benjamin Van and Craswell, Nick},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2510.07083v1},
  eprint = {2510.07083}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/