Med-HALT: Medical Domain Hallucination Test for Large Language Models
Ankit PalLogesh Kumar UmapathiMalaikannan Sankarasubbu
Introduces Med-HALT, a multinational benchmark featuring reasoning and memory-based tests to evaluate and mitigate factual errors and hallucinations in large language models applied to healthcare.
Large language models show significant promise across healthcare workflows, but they frequently hallucinate by generating plausible, confident, and factually incorrect statements. In clinical contexts, these errors pose critical patient safety and liability risks. The article introduces Med-HALT, an open evaluation benchmark designed to measure and evaluate hallucination tendencies across leading artificial intelligence models within the medical domain.
The benchmark evaluates models using a multinational dataset of 18,866 reasoning samples sourced from medical licensing and entrance examinations across the United States, Spain, India, and Taiwan, alongside 4,916 life sciences records from the PubMed biomedical archive. The evaluation framework divides tests into two core areas: reasoning hallucination tests, which examine susceptibility to false confidence, handling fake or nonsensical questions, and identifying "none of the above" conditions; and memory hallucination tests, which evaluate the factual retrieval of biomedical literature identifiers, abstracts, and titles.
The evaluation showed that all tested models exhibited substantial hallucination rates, with open-access foundation models generally outperforming proprietary commercial variants. For reasoning tests, Meta's Llama-2 70B achieved the highest overall average accuracy at 72.33%, compared to 54.46% for OpenAI's Text-Davinci and 44.48% for GPT-3.5. On the False Confidence Test, every evaluated model struggled, with top accuracy capping at only 42.21%. For biomedical information retrieval tasks, Falcon 40B performed best with an overall average accuracy of 30.36%, while GPT-3.5 achieved only 19.96%. Furthermore, the analysis found that instruction tuning and reinforcement learning often degraded hallucination resistance, while prompt framing and few-shot examples improved accuracy up to a plateau of roughly three examples.
These findings demonstrate that current state-of-the-art language models remain fragile, brittle, and prone to severe factual errors when handling clinical reasoning and biomedical information recall. Consequently, deploying these systems directly in unsupervised healthcare workflows introduces severe clinical, safety, and compliance risks. Deployers should exercise extreme caution and must not rely on zero-shot or unassisted model outputs for clinical decision-making.
Organizations evaluating medical language models should implement structured, specific prompt engineering, leverage few-shot examples, and explore external knowledge integration mechanisms rather than relying solely on internal model memory. Future development should incorporate broader task evaluations, explore advanced mitigation techniques, and assess newer models such as GPT-4, which were excluded from this study due to financial constraints.
- Paper: What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams, Di Jin et al. (2020). It introduces MEDQA, the foundational multinational medical licensing exam dataset that Med-HALT directly builds upon and adapts to evaluate medical domain hallucinations.
- Paper: Large language models encode clinical knowledge, Karan Singhal et al. (2022). It establishes the MultiMedQA clinical evaluation benchmark and demonstrates how large language models encode medical knowledge, providing essential context for hallucination risks in healthcare.
- Paper: Survey of Hallucination in Natural Language Generation, Ziwei Ji et al. (2022). It offers the standard foundational taxonomy and definitions for intrinsic versus extrinsic hallucinations in natural language generation that underpin domain-specific hallucination benchmarks.
- Paper: Capabilities of GPT-4 on Medical Challenge Problems, Harsha Nori et al. (2023). It documents the baseline performance and error patterns of advanced LLMs like GPT-4 on medical licensing examinations, framing the problem space Med-HALT aims to test.
- Paper: TruthfulQA: Measuring How Models Mimic Human Falsehoods, Stephanie C. Lin et al. (2022). It pioneered standardized benchmark methodologies for measuring factual inaccuracies and imitative falsehoods in LLMs across diverse topics.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). It details the architecture and training of the Llama 2 model family, which serves as one of the primary evaluated open-weight LLMs in the Med-HALT benchmark.
- Paper: Rationale-Guided Retrieval Augmented Generation for Medical Question Answering, Jiwoong Sohn et al. (2025). It develops a rationale-guided retrieval-augmented generation framework specifically engineered to mitigate the medical hallucinations and reasoning failures diagnosed by benchmarks like Med-HALT.
- Paper: MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding, Yuxin Zuo et al. (2025). It advances beyond text-based medical licensing examinations by introducing a benchmark for expert-level multimodal clinical reasoning and diagnostic evaluation.
- Paper: RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models, Cheng Niu et al. (2024). It extends hallucination benchmarking into standard retrieval-augmented generation pipelines across multiple domains, including fine-grained word-level analysis.
- Paper: Fine-Tuning Language Models for Factuality, Katherine Tian et al. (2024). It applies preference optimization fine-tuning to systematically reduce factual hallucination rates in language models on complex downstream tasks such as medical question-answering.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). It broadens domain-specific truthfulness evaluations into a comprehensive, unified trustworthiness benchmark spanning safety, truthfulness, and reliability across mainstream LLMs.
