Med-HALT: Medical Domain Hallucination Test for Large Language Models

Ankit PalLogesh Kumar UmapathiMalaikannan Sankarasubbu

article2023CoNLL297 citations

Introduces Med-HALT, a multinational benchmark featuring reasoning and memory-based tests to evaluate and mitigate factual errors and hallucinations in large language models applied to healthcare.

Listen

Large language models show significant promise across healthcare workflows, but they frequently hallucinate by generating plausible, confident, and factually incorrect statements. In clinical contexts, these errors pose critical patient safety and liability risks. The article introduces Med-HALT, an open evaluation benchmark designed to measure and evaluate hallucination tendencies across leading artificial intelligence models within the medical domain.

The benchmark evaluates models using a multinational dataset of 18,866 reasoning samples sourced from medical licensing and entrance examinations across the United States, Spain, India, and Taiwan, alongside 4,916 life sciences records from the PubMed biomedical archive. The evaluation framework divides tests into two core areas: reasoning hallucination tests, which examine susceptibility to false confidence, handling fake or nonsensical questions, and identifying "none of the above" conditions; and memory hallucination tests, which evaluate the factual retrieval of biomedical literature identifiers, abstracts, and titles.

The evaluation showed that all tested models exhibited substantial hallucination rates, with open-access foundation models generally outperforming proprietary commercial variants. For reasoning tests, Meta's Llama-2 70B achieved the highest overall average accuracy at 72.33%, compared to 54.46% for OpenAI's Text-Davinci and 44.48% for GPT-3.5. On the False Confidence Test, every evaluated model struggled, with top accuracy capping at only 42.21%. For biomedical information retrieval tasks, Falcon 40B performed best with an overall average accuracy of 30.36%, while GPT-3.5 achieved only 19.96%. Furthermore, the analysis found that instruction tuning and reinforcement learning often degraded hallucination resistance, while prompt framing and few-shot examples improved accuracy up to a plateau of roughly three examples.

These findings demonstrate that current state-of-the-art language models remain fragile, brittle, and prone to severe factual errors when handling clinical reasoning and biomedical information recall. Consequently, deploying these systems directly in unsupervised healthcare workflows introduces severe clinical, safety, and compliance risks. Deployers should exercise extreme caution and must not rely on zero-shot or unassisted model outputs for clinical decision-making.

Organizations evaluating medical language models should implement structured, specific prompt engineering, leverage few-shot examples, and explore external knowledge integration mechanisms rather than relying solely on internal model memory. Future development should incorporate broader task evaluations, explore advanced mitigation techniques, and assess newer models such as GPT-4, which were excluded from this study due to financial constraints.

Cover for Med-HALT: Medical Domain Hallucination Test for Large Language Models

Abstract

This research paper focuses on the challenges posed by hallucinations in large language models (LLMs), particularly in the context of the medical domain. Hallucination, wherein these models generate plausible yet unverified or incorrect information, can have serious consequences in healthcare applications. We propose a new benchmark and dataset, Med-HALT (Medical Domain Hallucination Test), designed specifically to evaluate and reduce hallucinations. Med-HALT provides a diverse multinational dataset derived from medical examinations across various countries and includes multiple innovative testing modalities. Med-HALT includes two categories of tests reasoning and memory-based hallucination tests, designed to assess LLMs’ problem-solving and information retrieval abilities.

Our study evaluated leading LLMs, including Text Davinci, GPT-3.5, LlaMa-2, MPT, and Falcon, revealing significant differences in their performance. The paper provides detailed insights into the dataset, promoting transparency and reproducibility. Through this work, we aim to contribute to the development of safer and more reliable language models in healthcare. Our benchmark can be found at med-halt.github.io

Table of Contents

  • 1 Introduction
  • 1.1 Task Definition
  • 2 Datasets Statistics
  • 3 Types of Hallucination Evaluated
  • 3.1 Reasoning Hallucination Tests (RHTs)
  • 3.2 Memory Hallucination Tests (MHTs)
  • 4 Data Analysis
  • 4.1 Subject and Topic Analysis
  • 4.2 Exam Types Analysis
  • 4.3 Difficulty and Diversity of Questions
  • 5 Experiments
  • 5.1 Baseline Models
  • 5.2 Implementation Details
  • 5.3 Evaluation matrices
  • 6 Results
  • 6.1 Effect of Instruction tuning
  • 7 Exploratory Analysis
  • 7.1 Effect of Temperature parameter
  • 7.2 Impact of number of few-shot examples
  • 7.3 Sensitivity to Prompt Framing
  • 7.4 Repetition Experiments
  • 7.5 Brittleness of LLMs
  • 8 Conclusion
  • References
  • A Med-HALT Selection Criteria
  • A.1 Difficulty and Diversity of Questions
  • B Parsing Output and Handling Exceptions
  • Acknowledgements
  • Limitations & Future Scope

Knowls

  1. Knowl 1 — Med-HALT benchmark and task formulation

    definition

    Med-HALT (Medical Domain Hallucination Test) is a benchmark for evaluating hallucination in medical large language models through two task families: Reasoning Hallucination Tests (RHTs) and Memory Hallucination Tests (MHTs).

    An RHT instance consists of a medical question QQ and candidate options O={O1,…,On}O=\{O_1,\ldots,O_n\}. The evaluated model produces a binary selection vector y=(y1,…,yn)y=(y_1,\ldots,y_n), where yi∈{0,1}y_i\in\{0,1\} indicates whether option OiO_i is selected. The task measures hallucination through the model’s correctness when reasoning over medical questions, including questions containing fake or deliberately unanswerable content.

    An MHT instance consists of input biomedical information DD—such as an abstract, PubMed identifier, article title, or article link—and a requested item II to retrieve. The model receives a binary score y∈{0,1}y\in\{0,1\}, where y=1y=1 means that the requested information was retrieved correctly. MHTs therefore assess whether a model can recall biomedical information without inventing incorrect or incomplete facts.

  2. Knowl 2 — Multinational medical dataset composition

    data/table

    Med-HALT combines multiple-choice medical examinations from India, Spain, Taiwan, and the United States with a separate PubMed retrieval collection. The examination questions cover many medical subjects and are used to construct the reasoning tests; the PubMed records are used to construct the four memory-retrieval tests. The five examination sources contain 18,866 questions in total, while each PubMed retrieval task contains 4,916 samples. The paper reports that nearly 95% of subjects contain more than 50 topics and approximately 70% contain more than 100 topics.

    Examination sourceCountryQuestionsVocabularyMaximum question tokensMaximum answer tokensAverage question tokensAverage answer tokens
    AIIMS PGIndia666013508939111.7319.34
    NEET PGIndia285575111358611.5418.91
    Exámenes de residencia médicaSpain40681383226436321.6437.28
    TWMLETaiwan28011288517218527.7737.70
    USMILEUnited States248221074526154117.8723.42

    The PubMed portion contains 4,916 title records and 4,916 abstract records. For the title representation, the vocabulary size is 8,776, the maximum document length is 37 tokens, and the average document length is 5 tokens. For the abstract representation, the vocabulary size is 61,323, the maximum document length is 661 tokens, and the average document length is 8 tokens.

    The seven benchmark datasets contain 18,866 False Confidence Test samples, 1,858 Fake Questions Test samples, 18,866 None-of-the-Above Test samples, and 4,916 samples for each of PMID-to-Title, Title-to-PubMed-Link, Abstract-to-PubMed-Link, and PubMed-Link-to-Title retrieval.

  3. Knowl 3 — Reasoning hallucination test modalities

    model/method

    Med-HALT evaluates reasoning hallucination with three complementary procedures.

    • False Confidence Test (FCT): The model receives a medical multiple-choice question, its options, and a proposed answer. It must state whether the proposed answer is correct, provide the correct answer, explain why it is correct or incorrect, and explain why the remaining options are wrong. This tests whether a model expresses unjustified certainty when evaluating a potentially incorrect answer.
    • None of the Above Test (NOTA): The correct answer in a multiple-choice question is replaced by a None of the Above option. The model must identify the correct option and justify the choice. This tests whether the model can reject plausible but incorrect alternatives rather than forcing a conventional answer.
    • Fake Questions Test (FQT): The model receives fabricated or nonsensical medical questions and must recognize that they cannot be answered as ordinary medical questions. The fake questions were produced through a hybrid process in which some were written by human experts and the remainder were generated with GPT-3.5.

    The tests require structured responses containing the selected answer, answer index, and explanations; FCT additionally requires an explicit correctness judgment for the proposed answer.

  4. Knowl 4 — Memory hallucination test modalities

    model/method

    Med-HALT evaluates factual retrieval from model memory with four PubMed-based transformations.

    • PMID-to-Title: input a PubMed identifier and output the corresponding article title.
    • Title-to-Link: input a PubMed article title and output its PubMed URL.
    • Abstract-to-Link: input an article abstract and output the corresponding PubMed URL.
    • Link-to-Title: input a PubMed URL and output the corresponding article title.

    Each retrieval prompt also asks the model to indicate whether the article exists. A retrieval is correct only when the returned existence judgment and requested title or URL agree with the reference record. The four transformations test different forms of factual association: identifier-to-title mapping, title-to-source mapping, abstract-based article identification, and source-to-title recall.

  5. Knowl 5 — Evaluation protocol and pointwise metric

    experimental setup

    The evaluated models were OpenAI Text-Davinci, GPT-3.5 Turbo, Meta Llama-2 models, Falcon 40B models, and MPT 7B models, including base and instruction-tuned or chat variants where available. Models were evaluated in their default configurations without task-specific fine-tuning or hyperparameter optimization.

    OpenAI models were queried through the Azure OpenAI ChatGPT API with temperature 0.70.7, top-p=1.0p=1.0, frequency penalty 00, and a fixed generation-token limit. Open-source models were run with PyTorch and Hugging Face Text Generation Inference on a Quadro RTX 8000 GPU with 48 GB of VRAM, using temperature 0.60.6 and top-p=0.95p=0.95.

    Accuracy is the fraction of samples for which the model’s answer is correct. The pointwise score assigns +1+1 point to a correct answer and −0.25-0.25 points to an incorrect answer. For NN samples, true labels yiy_i, predicted labels y^i\hat y_i, indicator function I(⋅)I(\cdot), correct-answer reward Pc=1P_c=1, and incorrect-answer penalty Pw=−0.25P_w=-0.25, the reported average score is

    S=1N∑i=1N[I(yi=y^i)Pc+I(yi≠y^i)Pw].S=\frac{1}{N}\sum_{i=1}^{N}\left[I(y_i=\hat y_i)P_c+I(y_i\ne\hat y_i)P_w\right].

    The exploratory analyses used a randomly selected 30% of questions from each examination dataset and PubMed articles, stratified by country, examination type, and question difficulty.

  6. Knowl 6 — Reasoning hallucination performance

    data/table

    The reasoning results compare accuracy and pointwise score on FCT, FQT, and NOTA. The strongest overall reasoning model was Llama-2 70B, although no model achieved a high accuracy on FCT. Falcon 40B and Falcon 40B-instruct were especially strong on fake-question detection. Llama-2 70B obtained the best reported NOTA accuracy and score.

    ModelFCT accuracyFCT scoreFQT accuracyFQT scoreNOTA accuracyNOTA scoreAverage accuracyAverage score
    GPT-3.534.1533.3771.6411.9927.6418.0144.4821.12
    Text-Davinci16.76-7.6482.7214.5763.89103.5154.4636.81
    Llama-2 70B42.2152.3797.2617.9477.53188.6672.3386.32
    Llama-2 70B Chat13.34-15.705.49-3.3714.96-11.8811.26-10.32
    Falcon 40B18.66-3.1799.8918.5658.7291.3159.0935.57
    Falcon 40B-instruct1.11-44.5599.3518.4355.6984.1752.0519.35
    Llama-2 13B1.72-43.189.4516.1374.38128.2555.1833.76
    Llama-2-13B-chat7.95-28.4221.480.3433.4331.6720.951.20
    Llama-2-7B0.45-46.1258.728.9969.49116.7142.8926.53
    Llama-2-7B-chat0.42-46.1721.960.4631.1026.1917.83-6.51
    MPT 7B0.85-45.1548.496.6219.88-0.2823.07-12.94
    MPT 7B instruct0.17-46.7622.550.5924.3410.3415.69-11.94

    The results show large differences between base and instruction-tuned variants. In particular, instruction-tuned Llama-based models performed substantially worse than their base counterparts on these reasoning hallucination tests.

  7. Knowl 7 — Memory hallucination performance

    data/table

    The memory results evaluate four retrieval transformations: PMID-to-Title, Title-to-PubMed-Link, Abstract-to-PubMed-Link, and PubMed-Link-to-Title. Falcon 40B and Falcon 40B-instruct achieved the strongest retrieval accuracies among the evaluated models on the first three transformations, while retrieval from a PubMed link to a title was difficult for almost every model. The paper characterizes Falcon as the strongest family overall for information retrieval and reports an average accuracy of 42.46% across the retrieval tasks, while the displayed aggregate row for Falcon 40B is 30.36%.

    ModelPMID-to-Title accuracyPMID-to-Title scoreTitle-to-Link accuracyTitle-to-Link scoreAbstract-to-Link accuracyAbstract-to-Link scoreLink-to-Title accuracyLink-to-Title scoreDisplayed average accuracyDisplayed average score
    GPT-3.50.29-12.1239.1011.7440.4512.570.02-12.2819.96-0.02
    Text-Davinci0.02-12.2838.5311.3940.4412.560.00-12.2919.75-0.15
    Llama-2 70B0.12-12.2214.79-3.2017.21-1.720.02-12.288.04-7.36
    Llama-2 70B Chat0.81-11.7932.877.9017.90-1.290.61-11.9213.05-4.27
    Falcon 40B40.4612.5740.4612.5740.4612.570.06-12.2530.366.37
    Falcon 40B-instruct40.4612.5740.4612.5740.4412.560.08-12.7530.366.24
    Llama-2 13B0.53-11.9710.56-5.804.70-9.4023.722.299.88-6.22
    Llama-2-13B-chat1.38-11.4438.8511.5938.3211.261.73-11.2320.070.04
    Llama-2-7B0.00-12.293.72-10.000.26-12.130.00-12.291.0-11.68
    Llama-2-7B-chat0.00-12.2930.926.7112.80-4.430.00-12.2910.93-5.57
    MPT 7B20.080.0540.4612.5740.0312.310.00-12.2925.143.16
    MPT 7B instruct0.04-12.2738.2411.2140.4612.570.00-12.2919.69-0.19

    The low scores, especially for PubMed-Link-to-Title, indicate substantial remaining hallucination and retrieval errors across all model families.

  8. Knowl 8 — Question diversity and sensitivity analyses

    empirical result

    A manual analysis of a 30% sample classified the reasoning demands in Med-HALT as factual recall (31.6%), diagnosis (22.6%), question logic (9.1%), explanation or description (8.3%), fact-based reasoning (8.1%), natural-language inference (7.6%), multihop reasoning (6.4%), exclusion of distractors (1.9%), mathematical reasoning (1.8%), fill-in-the-blank questions (1.4%), and comparison questions (1.3%).

    The exploratory experiments found that:

    • Changing GPT-3.5’s temperature produced only minimal accuracy variation over the tested values, suggesting that temperature alone had little effect on medical hallucination in this evaluation.
    • GPT-3.5 accuracy was 7.31% with zero few-shot examples and increased as examples were added, but the gains diminished beyond three examples.
    • Prompt framing affected accuracy. The three tested prompt variants produced accuracies of 24.44%, 22.97%, and 25.48%, respectively; the authors conclude that explicit, directed prompts are important even though the individual variants were not monotonically ordered.
    • Repeated GPT-3.5 runs on the same sampled questions were relatively stable but not identical: the highest accuracy was 28.52% on the fourth attempt, while the lowest was approximately 27.87% on the second and fifth attempts.
    • Instruction tuning and reinforcement learning from human feedback sometimes reduced hallucination control, with the clearest degradation observed among Llama-based models and a weaker effect for OpenAI and Falcon models.
    • Small changes to prompting or decoding settings could turn previously correct outputs into hallucinated answers, which the authors describe as brittleness.
  9. Knowl 9 — Structured-output parsing and exception handling

    model/method

    Med-HALT evaluation relies on valid JSON responses so that model answers can be scored automatically. The processing procedure first attempts direct parsing with the Promptify module. If a response contains conflicting single and double quotation marks, a regular-expression escaping step modifies the string before parsing. Outputs that remain malformed after these steps are not repeatedly regenerated; they are counted as failures to follow the requested format.

    This procedure distinguishes format-following failures from ordinary answer errors and permits an instruction-compliance rate to be measured. The reported malformed-output rates were:

    ModelFCTFQTNOTAPMID-to-TitleTitle-to-LinkAbstract-to-LinkLink-to-Title
    GPT-3.52.24%3.19%1.28%2.42%2.03%1.97%1.06%
    Text-Davinci1.31%2.24%0.8%1.60%1.76%1.93%0.4%
    Falcon 40B0000000
    Falcon 40B-instruct0000000
    Llama-2 7B0.04%00.01%0000
    Llama-2 7B-chat0000000
    Llama-2 13B0.01%000000
    Llama-2 70B0000000
    Llama-2 70B-chat41.1%024.92%0000

    The results show that most models rarely violated the output format, but Llama-2 70B-chat produced exceptionally high exception rates on FCT and NOTA.

  10. Knowl 10 — Scope limitations and future directions

    limitation

    The evaluation covers reasoning and biomedical information retrieval, so its results do not establish how the models behave on every medical task, such as open-ended clinical dialogue, diagnosis generation, treatment planning, or multimodal healthcare tasks. The study does not isolate the effects of model architecture, training-data diversity, or task type.

    The work observes that instruction tuning can worsen hallucination control but does not evaluate mitigation strategies such as external knowledge retrieval or hallucination-specific training objectives. Temperature sensitivity was studied primarily with GPT-3.5, so the observed behavior may not generalize to other models. Financial constraints prevented evaluation of GPT-4. The authors identify broader model coverage, external knowledge, targeted training objectives, and further analysis of decoding-parameter interactions as future work.

Coverage note — Illustrative hallucination examples, full prompt-variant listings, subject-by-subject plots, and appendix question examples were omitted because they elaborate the benchmark without adding load-bearing methods or aggregate findings.

References

  1. 1.Monica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim, and David Sontag. 2022. Large language models are few-shot clinical information extractors.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback.
  3. 3.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. ArXiv, abs/2302.04023.
  4. 4.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  5. 5.Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Loser, Daniel Truhn, and Keno K Bressem. 2023. Medalpaca–an open-source collection of medical conversational ai models and training data. arXiv preprint arXiv:2304.08247.
  6. 6.Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. CoRR, abs/1904.09751.
  7. 7.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Wenliang Dai, Andrea Madotto, and Pascale Fung. 2022. Survey of hallucination in natural language generation. ACM Computing Surveys, 55:1 – 38.
  8. 8.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. ArXiv, abs/2009.13081.
  9. 9.Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jianyun Nie, and Ji rong Wen. 2023a. Halueval: A large-scale hallucination evaluation benchmark for large language models. ArXiv, abs/2305.11747.
  10. 10.Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang. 2023b. Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge. Cureus, 15(6).
  11. 11.Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan. 2021. A token-level reference-free hallucination detection benchmark for free-form text generation. arXiv preprint arXiv:2104.08704.
  12. 12.MosaicML. 2023. Introducing mpt-30b: Raising the bar for open-source foundation models. Accessed: 2023-06-22.
  13. 13.Ankit Pal. 2022. Promptify: Structured output from llms. https://github.com/promptslab/Promptify. Prompt-Engineering components for NLP tasks in Python.
  14. 14.Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Proceedings of the Conference on Health, Inference, and Learning, volume 174 of Proceedings of Machine Learning Research, pages 248–260. PMLR.
  15. 15.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. ArXiv, abs/1912.01703.
  16. 16.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023a. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only.
  17. 17.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023b. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116.
  18. 18.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  19. 19.Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. CoRR, abs/2104.07567.
  20. 20.K. Singhal, Shekoofeh Azizi, and Tao Tu. 2022. Large language models encode clinical knowledge. ArXiv, abs/2212.13138.
  21. 21.K. Singhal, Tao Tu, and Juraj Gottweis. 2023. Towards expert-level medical question answering with large language models. ArXiv, abs/2305.09617.
  22. 22.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models.
  23. 23.David Vilares and Carlos Gomez Rodríguez. 2019. Head-qa: A healthcare dataset for complex reasoning. ArXiv, abs/1906.04701.
  24. 24.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions.
  25. 25.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners. CoRR, abs/2109.01652.
  26. 26.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771.

Citation

MLA
Pal, A., et al. “Med-HALT: Medical Domain Hallucination Test for Large Language Models”. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 2023, pp. 314–34, https://doi.org/10.18653/v1/2023.conll-1.21.
APA
Pal, A., Umapathi, L. K., & Sankarasubbu, M. (2023). Med-HALT: Medical Domain Hallucination Test for Large Language Models. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 314–334. https://doi.org/10.18653/v1/2023.conll-1.21
Chicago
Pal, A., L. K. Umapathi, and M. Sankarasubbu. 2023. “Med-HALT: Medical Domain Hallucination Test for Large Language Models”. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), 314–34. https://doi.org/10.18653/v1/2023.conll-1.21.
Harvard
Pal, A., Umapathi, L.K. and Sankarasubbu, M. (2023) “Med-HALT: Medical Domain Hallucination Test for Large Language Models”, Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). Association for Computational Linguistics, pp. 314–334. Available at: https://doi.org/10.18653/v1/2023.conll-1.21.
Vancouver
1. Pal A, Umapathi LK, Sankarasubbu M (2023) Med-HALT: Medical Domain Hallucination Test for Large Language Models. In: Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL). Association for Computational Linguistics, pp 314–334

BibTeX

@inproceedings{pal-etal-2023-med,
    title = "{M}ed-{HALT}: Medical Domain Hallucination Test for Large Language Models",
    author = "Pal, Ankit  and
      Umapathi, Logesh Kumar  and
      Sankarasubbu, Malaikannan",
    editor = "Jiang, Jing  and
      Reitter, David  and
      Deng, Shumin",
    booktitle = "Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL)",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.conll-1.21/",
    doi = "10.18653/v1/2023.conll-1.21",
    pages = "314--334"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/