Large language models are few-shot clinical information extractors

Monica AgrawalStefan HegselmannHunter LangYoon KimDavid A. Sontag

article2022EMNLP467 citations

Demonstrates that general-domain large language models can perform complex few-shot clinical information extraction tasks, such as span identification and relation extraction, significantly outperforming existing baselines without domain-specific training.

Listen

Clinical notes contain critical patient information, yet unlocking this data remains difficult because medical text relies heavily on ambiguous abbreviations, nonstandard phrasing, and sensitive patient records that cannot be broadly shared. Building conventional machine learning systems for clinical text typically demands labor-intensive domain engineering, brittle rule sets, and large volumes of expert-annotated training data. The article evaluates whether modern, domain-agnostic large language models can perform accurate information extraction from clinical records using zero-shot and few-shot prompting, and it demonstrates how to transform the natural-language outputs of these models into structured clinical variables.

To conduct this evaluation, the researchers tested large language models across five diverse clinical tasks: abbreviation sense disambiguation, biomedical evidence extraction from clinical trial abstracts, pronoun coreference resolution, medication status extraction, and medication attribute relation extraction. Because strict data use agreements prevent sending most existing clinical benchmarks to third-party model programming interfaces, the authors created and publicly released three new benchmark datasets by re-annotating snippets from the public Clinical Acronym Sense Inventory. The approach couples guided prompting—supplying formatting instructions and minimal demonstration examples—with simple software post-processors called resolvers to map unstructured text generations directly into structured labels.

The investigation produced several key findings. First, guided large language models consistently outperformed strong baseline systems across multiple tasks without task-specific architecture tuning. In medication extraction, a one-shot guided model achieved 0.90 recall and 0.92 precision, substantially exceeding a standard biomedical rule-based baseline at 0.73 recall and 0.67 precision. Second, for clinical trial arm identification, the zero-shot language model achieved 0.85 abstract-level accuracy, identifying the correct study arms in 17 of 20 trials, compared to only 0.35 accuracy for a supervised biomedical language model that required thousands of training examples. Third, for abbreviation disambiguation, zero-shot prompting reached 0.86 accuracy and 0.69 macro F1 score on the acronym inventory, outperforming specialized baselines. Furthermore, when these outputs were used as weak supervision to train a smaller biomedical model, performance rose to 0.90 accuracy and transferred effectively to an independent critical care dataset. Fourth, the experiments revealed that including even a single formatted example reduced post-processing code complexity from dozens of lines to fewer than ten lines of code, and that the structural format of the example mattered more than whether its label was factually correct.

These findings indicate that general-purpose large language models can dramatically lower the engineering cost and annotation burden required to extract complex clinical information. Practitioners do not necessarily need massive labeled datasets or brittle custom rules to achieve high extraction accuracy. Moreover, distilling large language model outputs into smaller, locally hosted models provides a viable pathway to overcome patient privacy constraints and recurring external computing costs.

Organizations exploring automated clinical extraction should adopt guided prompt formatting to simplify data pipelines and minimize engineering overhead. Teams can leverage large language models to generate weak labels on de-identified public data and distill those capabilities into smaller, internally hosted models that safeguard patient privacy. Before operational deployment in high-stakes clinical workflows, institutions should validate systems locally across diverse patient demographics, clinical specialties, and note types, while implementing verification steps to prevent models from generating answers when target entities are absent.

Confidence in the reported extraction capabilities is high for English-language notes within the evaluated scopes, supported by rigorous manual adjudication and baseline comparisons. However, readers should note that the evaluation relied primarily on dictated snippets from one multi-hospital system and published biomedical abstracts. Additional validation across broader international healthcare systems, diverse clinical documentation practices, and non-English clinical texts remains necessary.

Cover for Large language models are few-shot clinical information extractors

Abstract

A long-running goal of the clinical NLP community is the extraction of important variables trapped in clinical notes. However, roadblocks have included dataset shift from the general domain and a lack of public clinical corpora and annotations. In this work, we show that large language models, such as InstructGPT, perform well at zero- and few-shot information extraction from clinical text despite not being trained specifically for the clinical domain. Whereas text classification and generation performance have already been studied extensively in such models, here we additionally demonstrate how to leverage them to tackle a diverse set of NLP tasks which require more structured outputs, including span identification, token-level sequence classification, and relation extraction. Further, due to the dearth of available data to evaluate these systems, we introduce new datasets for benchmarking few-shot clinical information extraction based on a manual re-annotation of the CASI dataset for new tasks. On the clinical extraction tasks we studied, the GPT-3 systems significantly outperform existing zero- and few-shot baselines.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Prompt-Based Learning
  • 2.2 Pretrained LMs for Clinical NLP
  • 3 Methods
  • 3.1 Predicting Structured Outputs with LLMs
  • 3.2 Dataset Annotation
  • 4 Clinical Sense Disambiguation
  • 5 Biomedical Evidence Extraction
  • 6 Coreference Resolution
  • 7 Medication Extraction
  • 7.1 Recognition + Status Classification
  • 7.2 Recognition + Relation Extraction
  • 8 Conclusion
  • References
  • A Prompts and Sample GPT-3 Outputs
  • A.1 Clinical Sense Disambiguation
  • A.2 Biomedical Evidence Extraction
  • A.3 Coreference Resolution
  • A.4 Medication Status Extraction
  • A.5 Medication Attribute Extraction
  • B Annotation Process
  • B.1 Biomedical Evidence Extraction
  • B.2 Coreference Resolution
  • B.3 Medication Status Extraction
  • B.4 Medication Attribute Extraction
  • C Additional Experimental details
  • C.1 Clinical Sense Disambiguation
  • C.2 Biomedical Evidence Extraction
  • C.3 Coreference Resolution
  • C.4 Medication + Status Extraction
  • C.5 Medication + Relation Extraction
  • D Experimental Cost

Knowls

  1. Knowl 1 — Structured Information Extraction Framework via Guided Prompting and Resolvers

    model/method

    In settings where large language models (LLMs) are accessible only via text queries without access to model gradients or token probabilities, structured extraction maps natural language text to a structured target space O\mathcal{O} (such as sequence labels {0,1}∣xi∣\{0, 1\}^{|x_i|}, lists of entities, or relational key-value attributes).

    Let xix_i denote the input clinical text string and aia_i optional side information (e.g., an abbreviation to expand). A prompt template pj(x,a)p_j(x, a) converts inputs into a prompt string. The language model f:Σ∗→Σ∗f: \Sigma^* \to \Sigma^* produces an output string f(pj(x,a))f(p_j(x, a)). A resolver function R(x,a,f(pj(x,a)))R(x, a, f(p_j(x, a))) maps the input and raw LLM text generation to the task-specific output space O\mathcal{O}.

    Guided prompt design is used to simplify the resolver RR into minimal code logic (often fewer than 10 lines of code). Guided prompting consists of:

    1. Providing a 1-shot demonstration that exemplifies the desired structural syntax (e.g., bulleted lists, key-value pairs, or sequence tags), where the example output adheres to the structural template regardless of whether the demonstration label is factually correct.
    2. Including explicit instructional guidance directing the LLM to format its response identically to the demonstration.

    This structural guidance constrains the LLM to output concise, un-embellished extractions, allowing the resolver RR to perform simple line-splitting or regex parsing rather than complex natural language interpretation.

  2. Knowl 2 — Weakly Supervised Distillation of LLM Extractions to Task-Specific Clinical Models

    model/method

    Direct deployment of commercial large language models in healthcare is often prevented by patient data privacy agreements (prohibiting data transmission to external APIs) and computational inference costs. To resolve this, an LLM combined with a resolver is employed as a weak supervision labeling function rather than an end-user classifier.

    The distillation procedure operates as follows:

    1. The LLM-plus-resolver system generates pseudolabels over an unannotated, publicly accessible clinical dataset (such as the Clinical Acronym Sense Inventory, CASI).
    2. Candidate pseudolabels are filtered using quality heuristics: retaining examples where string overlap with candidate options meets a threshold (e.g., ≥5\ge 5 characters) and applying subset selection (such as the cut statistic to retain a high-quality subset, e.g., the top 75% cleanest pseudolabels).
    3. A smaller, domain-specific transformer (such as PubMedBERT) is fine-tuned on the filtered weakly labeled dataset using multiple-choice or sequence classification loss objectives.
    4. The distilled, task-specific model is deployed locally or transferred zero-shot to private clinical databases (such as MIMIC-III) without transmitting private clinical records to external API endpoints.
  3. Knowl 3 — Clinical Sense Disambiguation and Cross-Dataset Transfer Performance

    data/table

    Clinical sense disambiguation expands ambiguous abbreviations in medical notes. Zero-shot GPT-3 edit mode (text-davinci-edit-001, greedy decoding) prompted with Expand the abbreviation: {abbr} was evaluated on a 41-acronym subset of the Clinical Acronym Sense Inventory (CASI; 18,164 notes) and transferred to the MIMIC-III Reverse Substitution dataset (8,912 test examples) via PubMedBERT distillation.

    Algorithm CASI Acc. CASI Macro F1 MIMIC Acc. MIMIC Macro F1
    Random 0.31 0.23 0.32 0.28
    Most Common 0.79 0.28 0.51 0.23
    Clinical BioBERT (Adams et al., 2020) 0.42 0.23 0.40 0.33
    ELMo (Adams et al., 2020) 0.55 0.38 0.58 0.53
    Latent Meaning Cells (Adams et al., 2020) 0.71 0.51 0.74 0.69
    GPT-3 edit + R: 0-shot 0.86 0.69 * *
    GPT-3 edit + R: 0-shot + distillation 0.90 0.76 0.78 0.69

    Zero-shot GPT-3 edit with a character-overlap resolver outperforms specialized zero-shot latent variable models (Latent Meaning Cells) on CASI (0.86 vs 0.71 accuracy). Distilling GPT-3 pseudolabels into PubMedBERT improves CASI accuracy to 0.90 (0.76 macro F1) and, when transferred to MIMIC-III, achieves 0.78 accuracy and 0.69 macro F1, matching or exceeding in-domain pre-trained baselines.

  4. Knowl 4 — Biomedical Evidence Extraction: Proxy Sequence Labeling vs. Clinical Trial Arm Identification

    data/table

    In evidence-based medicine, intervention and control extraction from randomized trial abstracts is evaluated both as a token-level sequence labeling proxy task (187 test abstracts from EBM-NLP) and as abstract-level arm identification (20 manually verified test abstracts). InstructGPT (text-davinci-002, greedy decoding) is evaluated zero-shot against fully supervised models trained on 4,800 abstracts.

    Algorithm Token-level F1 Abstract-level Accuracy
    PubMedBERT-CRF (Supervised, 4,800 abstracts) 0.69 0.35
    LSTM-CRF (Supervised, 4,800 abstracts) 0.65 *
    GPT-3 + R: 0-shot 0.61 0.85

    While supervised PubMedBERT-CRF achieves a higher token-level proxy F1 score (0.69 vs 0.61) due to dataset-specific token boundary conventions, zero-shot GPT-3 with a line-splitting resolver achieves 0.85 accuracy at identifying the true clinical trial arms (correctly identifying 17 of 20 abstracts). Even under theoretical assumptions of oracle span splitting and oracle coreference resolution, supervised PubMedBERT-CRF fails on 13 of 20 abstracts (0.35 accuracy) due to missing control arms or failing to consolidate synonymous intervention entities.

  5. Knowl 5 — Medication Extraction and Status Classification

    data/table

    Medication extraction involves identifying medication names and classifying their status as active, discontinued (including temporarily on hold), or neither (such as allergies or hypothetical treatments). Evaluation was conducted on 105 narrative CASI snippets containing 340 medication-status instances.

    Medication Extraction Algorithm Micro Recall Micro Precision
    ScispaCy (UMLS entity linker) 0.73 0.67
    GPT-3 + R (32 LOC resolver, 0-shot) 0.87 0.83
    GPT-3 + R (8 LOC resolver, 1-shot) 0.90 ±\pm 0.01 0.92 ±\pm 0.01
    Status Classification Algorithm Conditional Accuracy Conditional Macro F1
    T-Few (20-shot, fine-tuned T0-11B) 0.86 0.57
    GPT-3 + R (32 LOC resolver, 0-shot) 0.85 0.69
    GPT-3 + R (8 LOC resolver, 1-shot) 0.89 ±\pm 0.01 0.62 ±\pm 0.04
    GPT-3 + R (8 LOC resolver, 1-shot + added classes) 0.88 ±\pm 0.02 0.71 ±\pm 0.03
    GPT-3 + R (8 LOC resolver, 1-shot with shuffled classes) 0.88 ±\pm 0.01 0.66 ±\pm 0.03

    For medication recognition, 1-shot InstructGPT outperforms the ScispaCy UMLS linking baseline by +17% recall and +25% precision. For status classification evaluated across the subset of 241 medications extracted by all GPT-3 prompts, augmenting the 1-shot demonstration so that all three status classes are represented improves macro F1 from 0.62 to 0.71, outperforming 20-shot T-Few fine-tuning.

  6. Knowl 6 — Medication Attribute and Relation Extraction Across Output Framings

    data/table

    Medication attribute extraction extracts medications and up to five associated relational attributes (Dosage, Route, Frequency, Reason, Duration). Evaluated across 105 CASI snippets (313 medications, 533 attributes) using three task formulations: token-level sequence labeling, phrase-level sequence labeling, and end-to-end relation extraction. Baseline models (PubMedBERT+CRF and Shi & Lin relation extraction) were trained on the 2009 i2b2 medication challenge dataset and transferred to the test domain.

    Subtask Algorithm Medication Dosage Route Frequency Reason Duration
    Token-level PubMedBERT + CRF (Sup.) 0.82 0.92 0.77 0.76 0.35 0.57
    Token-level GPT-3 + R: 1-shot 0.85 0.92 0.87 0.91 0.38 0.52
    Phrase-level PubMedBERT + CRF (Sup.) 0.73 0.78 0.71 0.41 0.22 0.30
    Phrase-level GPT-3 + R: 1-shot 0.75 0.82 0.81 0.87 0.21 0.25
    Relation Extr. PubMedBERT + CRF + Shi Lin * 0.67 0.65 0.36 0.19 0.21
    Relation Extr. GPT-3 + R: 1-shot * 0.80 0.63 0.60 0.34 0.16

    1-shot GPT-3 + resolver outperforms supervised i2b2 baselines across most entities, notably achieving +0.24 higher F1 on Frequency in relation extraction (0.60 vs 0.36) and +0.13 on Dosage (0.80 vs 0.67). Supervised pipelines degrade in end-to-end relation extraction due to compounding errors from preliminary span extraction.

  7. Knowl 7 — Clinical Pronoun Coreference Resolution via Guided Prompting

    data/table

    Pronoun coreference resolution in clinical records maps a queried pronoun to its antecedent noun phrase span without token overlap. Evaluated on 100 non-trivial pronoun-antecedent pairs from CASI using tokenized macro unigram precision and recall.

    Algorithm Recall Precision
    Longdoc baseline (Toshniwal et al., 2020, 2021) 0.73 0.60
    GPT-3 + R (50 LOC resolver): 0-shot unguided 0.78 0.58
    GPT-3 + R (1 LOC resolver): 1-shot guided (incorrect demonstration) 0.76 ±\pm 0.02 0.78 ±\pm 0.04
    GPT-3 + R (1 LOC resolver): 1-shot guided (correct demonstration) 0.75 ±\pm 0.04 0.77 ±\pm 0.04

    Guided 1-shot InstructGPT prompts increase precision from 0.58 to 0.77-0.78 compared to 0-shot unguided prompting, while reducing resolver complexity from 50 lines of code (handling conversational explanations and paraphrases) to a single line of code (stripping quotes). The longdoc baseline, trained on multiple coreference datasets, achieves 0.73 recall and 0.60 precision.

  8. Knowl 8 — Public Clinical NLP Benchmark Suite via CASI and EBM-NLP Re-Annotation

    experimental setup

    To enable clinical NLP benchmarking under external API data use restrictions, four public datasets were created via manual annotation by two domain experts (combining clinical NLP and medical backgrounds) with adjudication via PRAnCER:

    1. Pronoun Coreference Resolution: 105 CASI snippets (5 validation, 100 evaluation) where personal pronouns are linked to full antecedent noun phrases. Personal pronouns appearing before any secondary person was introduced were filtered to ensure non-triviality.
    2. Medication Status Extraction: 105 CASI snippets enriched for treatment changes (filtered via keywords: discont, adverse, side effect, switch, dosage), yielding 340 medication mentions labeled as active, discontinued, or neither.
    3. Medication Attribute and Relation Extraction: 105 CASI snippets annotated following modified 2009 i2b2 challenge guidelines, covering 313 medications and 533 modifier attributes (dosage, route, frequency, reason, duration) formatted as token-level sequences, phrase-level chunks, and end-to-end relations.
    4. Biomedical Evidence Arm Identification: 20 randomized clinical trial abstracts from the EBM-NLP test set manually annotated for distinct intervention and control trial arms with 100% inter-annotator agreement.
  9. Knowl 9 — Demonstration Invariance to Label Correctness vs. Class Space Coverage in Few-Shot Prompts

    empirical result

    Experiments assessing in-context demonstrations for structured clinical extraction demonstrate two properties:

    1. Invariance to Demonstration Label Accuracy: In clinical span extraction (e.g., coreference resolution), replacing the correct ground-truth antecedent in the 1-shot demonstration with a random incorrect noun phrase yields equivalent extraction performance (0.76 unigram recall, 0.78 unigram precision for incorrect vs. 0.75 recall, 0.77 precision for correct). The demonstration serves primarily to enforce output syntax, quote constraints, and direct span extraction rather than providing factual supervision.
    2. Sensitivity to Label Space Coverage: In multi-class modifier classification (e.g., medication status), the LLM rarely predicts classes absent from the demonstration prompt. When the rare neither status is absent from the 1-shot demonstration, conditional macro F1 drops to 0.62. Artificially injecting or shuffling synthetic demonstrations to ensure every candidate label appears in the prompt recovers macro F1 to 0.71, even if the injected demonstration labels are permuted.
  10. Knowl 10 — Elimination Bias in Negative Entity Extraction and Prompt Structuring Strategies

    limitation

    Large language models demonstrate an extraction bias toward returning non-trivial outputs when queried for specific subsets of clinical entities, even when none exist in the source note. For example, prompting an LLM with Create a bulleted list of discontinued medications, if any on a note containing only active medications causes the model to extract the active medications as discontinued rather than returning an empty list.

    To mitigate this structural hallucination bias, clinical prompts must be designed using:

    1. Joint categorical extraction: Requesting all entity categories and their respective modifier statuses simultaneously within a single prompt (e.g., Create a bulleted list of medications and whether they are active, discontinued, or neither), allowing negative/discontinued entities to be correctly distinguished from active ones.
    2. Prompt chaining: Employing a multi-step prompt sequence that first verifies whether an entity category is present before requesting span extraction.
    3. Full sequence tagging: Framing the extraction task as an exhaustive token- or phrase-level labeling problem.

Coverage note — None was omitted; all key contributions, newly annotated benchmark datasets, model-distillation methods, task evaluations, and prompting limitations were fully covered.

References

  1. 1.Griffin Adams, Mert Ketenci, Shreyas Bhave, Adler J. Perotte, and Noemie Elhadad. 2020. Zero-shot clinical acronym expansion via latent meaning cells. In Machine Learning for Health Workshop, ML4H@NeurIPS 2020, Virtual Event, 11 December 2020, volume 136 of Proceedings of Machine Learning Research, pages 12–40. PMLR.
  2. 2.Emily Alsentzer, John R Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew McDermott. 2019. Publicly available clinical bert embeddings. arXiv preprint arXiv:1904.03323.
  3. 3.Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, et al. 2018. Construction of the literature graph in semantic scholar. arXiv preprint arXiv:1805.02262.
  4. 4.Hilda Bastian, Paul Glasziou, and Iain Chalmers. 2010. Seventy-five trials and eleven systematic reviews a day: how will we ever keep up? PLoS medicine, 7(9):e1000326.
  5. 5.Olivier Bodenreider. 2004. The unified medical language system (umls): integrating biomedical terminology. Nucleic acids research, 32(suppl_1):D267–D270.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Wendy W Chapman, Will Bridewell, Paul Hanbury, Gregory F Cooper, and Bruce G Buchanan. 2001. A simple algorithm for identifying negated findings and diseases in discharge summaries. Journal of biomedical informatics, 34(5):301–310.
  8. 8.Geeticka Chauhan, Ruizhi Liao, William Wells, Jacob Andreas, Xin Wang, Seth Berkowitz, Steven Horng, Peter Szolovits, and Polina Golland. 2020. Joint modeling of chest radiographs and radiology reports for pulmonary edema assessment. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 529–539. Springer.
  9. 9.Irene Y Chen, Emily Alsentzer, Hyesun Park, Richard Thomas, Babina Gosangi, Rahul Gujrathi, and Bharti Khurana. 2020. Intimate partner violence and injury prediction from radiology reports. In BIOCOMPUTING 2021: Proceedings of the Pacific Symposium, pages 55–66. World Scientific.
  10. 10.Leyang Cui, Yu Wu, Jian Liu, Sen Yang, and Yue Zhang. 2021. Template-based named entity recognition using bart. arXiv preprint arXiv:2106.01760.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  12. 12.Elisa Ferracane, Iain Marshall, Byron C Wallace, and Katrin Erk. 2016. Leveraging coreference to identify arms in medical abstracts: An experimental study. In Proceedings of the Seventh International Workshop on Health Text Mining and Information Analysis, pages 86–95.
  13. 13.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23.
  14. 14.Bernal Jimenez Gutierrez, Nikolas McNeal, Clay Washington, You Chen, Lang Li, Huan Sun, and Yu Su. 2022. Thinking about gpt-3 in-context learning for biomedical ie? think again. arXiv preprint arXiv:2203.08410.
  15. 15.Sam Henry, Kevin Buchan, Michele Filannino, Amber Stubbs, and Ozlem Uzuner. 2020. 2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records. Journal of the American Medical Informatics Association, 27(1):3–12.
  16. 16.Or Honovich, Uri Shaham, Samuel R Bowman, and Omer Levy. 2022. Instruction induction: From few examples to natural language task descriptions. arXiv preprint arXiv:2205.10782.
  17. 17.Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. 2019. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 590–597.
  18. 18.Min Jiang, Yukun Chen, Mei Liu, S Trent Rosenbloom, Subramani Mani, Joshua C Denny, and Hua Xu. 2011. A study of machine-learning-based approaches to extract clinical entities and their assertions from discharge summaries. Journal of the American Medical Informatics Association, 18(5):601–606.
  19. 19.Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. 2019. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, 6(1):1–8.
  20. 20.Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9.
  21. 21.Kundan Krishna, Amy Pavel, Benjamin Schloss, Jeffrey P Bigham, and Zachary C Lipton. 2021. Extracting structured data from physician-patient conversations by predicting noteworthy utterances. In Explainable AI in Healthcare and Medicine, pages 155–169. Springer.
  22. 22.Ivy Fenton Kuhn. 2007. Abbreviations and acronyms in healthcare: when shorter isn’t sweeter. Pediatric nursing, 33(5).
  23. 23.Hunter Lang, Monica Agrawal, Yoon Kim, and David Sontag. 2022a. Co-training improves prompt-based learning for large language models. arXiv preprint arXiv:2202.00828.
  24. 24.Hunter Lang, Aravindan Vijayaraghavan, and David Sontag. 2022b. Training subset selection for weak supervision. arXiv preprint arXiv:2206.02914.
  25. 25.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
  26. 26.Kenton Lee, Luheng He, Mike Lewis, and Luke Zettlemoyer. 2017. End-to-end neural coreference resolution. arXiv preprint arXiv:1707.07045.
  27. 27.Ariel Levy, Monica Agrawal, Arvind Satyanarayan, and David Sontag. 2021. Assessing the impact of automated suggestions on decision making: Domain experts mediate model errors but take less initiative. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–13.
  28. 28.Xiaoya Li, Jingrong Feng, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2019a. A unified mrc framework for named entity recognition. arXiv preprint arXiv:1910.11476.
  29. 29.Xiaoya Li, Fan Yin, Zijun Sun, Xiayu Li, Arianna Yuan, Duo Chai, Mingxin Zhou, and Jiwei Li. 2019b. Entity-relation extraction as multi-turn question answering. arXiv preprint arXiv:1905.05529.
  30. 30.Andy T Liu, Wei Xiao, Henghui Zhu, Dejiao Zhang, Shang-Wen Li, and Andrew Arnold. 2022a. Qaner: Prompting question answering models for few-shot named entity recognition. arXiv preprint arXiv:2203.01543.
  31. 31.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022b. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. arXiv preprint arXiv:2205.05638.
  32. 32.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  33. 33.Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In International Conference on Learning Representations.
  34. 34.Yen-Fu Luo, Sam Henry, Yanshan Wang, Feichen Shen, Ozlem Uzuner, and Anna Rumshisky. 2020. The 2019 n2c2/umass lowell shared task on clinical concept normalization. Journal of the American Medical Informatics Association, 27(10):1529–e1.
  35. 35.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems, 26.
  36. 36.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  37. 37.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021. Reframing instructional prompts to gptk’s language. arXiv preprint arXiv:2109.07830.
  38. 38.Sungrim Moon, Serguei Pakhomov, Nathan Liu, James O Ryan, and Genevieve B Melton. 2014. A sense inventory for clinical abbreviations and acronyms created using clinical notes and medical dictionary resources. Journal of the American Medical Informatics Association, 21(2):299–307.
  39. 39.Milad Moradi, Kathrin Blagec, Florian Haberl, and Matthias Samwald. 2021. Gpt-3 models are poor few-shot learners in the biomedical domain. arXiv preprint arXiv:2109.02555.
  40. 40.Danielle L Mowery, Brett R South, Lee Christensen, Jianwei Leng, Laura-Maria Peltonen, Sanna Salantera, Hanna Suominen, David Martinez, Sumithra Velupillai, Noemie Elhadad, et al. 2016. Normalizing acronyms and abbreviations to aid patient understanding of clinical texts: Share/clef ehealth challenge 2013, task 2. Journal of biomedical semantics, 7(1):1–13.
  41. 41.Shawn N Murphy, Griffin Weber, Michael Mendis, Vivian Gainer, Henry C Chueh, Susanne Churchill, and Isaac Kohane. 2010. Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2). Journal of the American Medical Informatics Association, 17(2):124–130.
  42. 42.Mark Neumann, Daniel King, Iz Beltagy, and Waleed Ammar. 2019. ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing. In Proceedings of the 18th BioNLP Workshop and Shared Task, pages 319–327, Florence, Italy. Association for Computational Linguistics.
  43. 43.Aurelie Neveol, Hercules Dalianis, Sumithra Velupillai, Guergana Savova, and Pierre Zweigenbaum. 2018. Clinical natural language processing in languages other than english: opportunities and challenges. Journal of biomedical semantics, 9(1):1–13.
  44. 44.Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain J Marshall, Ani Nenkova, and Byron C Wallace. 2018. A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature. In Proceedings of the conference. Association for Computational Linguistics. Meeting, volume 2018, page 197. NIH Public Access.
  45. 45.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  46. 46.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2227–2237, New Orleans, Louisiana. Association for Computational Linguistics.
  47. 47.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D Manning. 2020. Stanza: A python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 101–108.
  48. 48.Alexander Ratner, Stephen H Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Re. 2017. Snorkel: Rapid training data creation with weak supervision. In Proceedings of the VLDB Endowment. International Conference on Very Large Data Bases, volume 11, page 269. NIH Public Access.
  49. 49.Kirk Roberts. 2016. Assessing the corpus size vs. similarity trade-off for word embeddings in clinical nlp. In Proceedings of the Clinical Natural Language Processing Workshop (ClinicalNLP), pages 54–63.
  50. 50.David L Sackett. 1997. Evidence-based medicine. In Seminars in perinatology, volume 21, pages 3–5. Elsevier.
  51. 51.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  52. 52.Timo Schick and Hinrich Schutze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269.
  53. 53.Peng Shi and Jimmy Lin. 2019. Simple bert models for relation extraction and semantic role labeling. arXiv preprint arXiv:1904.05255.
  54. 54.Lotan Shilo and Gila Shilo. 2018. Analysis of abbreviations used by residents in admission notes and discharge summaries. QJM: An International Journal of Medicine, 111(3):179–183.
  55. 55.Yuqi Si and Kirk Roberts. 2019. Deep patient representation of clinical notes via multi-task learning for mortality prediction. AMIA Summits on Translational Science Proceedings, 2019:779.
  56. 56.Marta Skreta, Aryan Arbabi, Jixuan Wang, Erik Drysdale, Jacob Kelly, Devin Singh, and Michael Brudno. 2021. Automatically disambiguating medical acronyms with ontology-aware deep learning. Nature communications, 12(1):1–10.
  57. 57.Ryan Smith, Jason A Fries, Braden Hancock, and Stephen H Bach. 2022. Language models in the loop: Incorporating prompting into weak supervision. arXiv preprint arXiv:2205.02318.
  58. 58.Sunghwan Sohn, Cheryl Clark, Scott R Halgrim, Sean P Murphy, Christopher G Chute, and Hongfang Liu. 2014. Medxn: an open source medication extraction and normalization tool for clinical text. Journal of the American Medical Informatics Association, 21(5):858–865.
  59. 59.Shubham Toshniwal, Sam Wiseman, Allyson Ettinger, Karen Livescu, and Kevin Gimpel. 2020. Learning to ignore: Long document coreference with bounded memory neural networks. arXiv preprint arXiv:2010.02807.
  60. 60.Shubham Toshniwal, Patrick Xia, Sam Wiseman, Karen Livescu, and Kevin Gimpel. 2021. On generalization in coreference resolution. arXiv preprint arXiv:2109.09667.
  61. 61.Ozlem Uzuner, Andreea Bodnari, Shuying Shen, Tyler Forbush, John Pestian, and Brett R South. 2012. Evaluating the state of the art in coreference resolution for electronic medical records. Journal of the American Medical Informatics Association, 19(5):786–791.
  62. 62.Ozlem Uzuner, Imre Solti, and Eithon Cadag. 2010. Extracting medication information from clinical text. Journal of the American Medical Informatics Association, 17(5):514–518.
  63. 63.Mathias Verbeke, Vincent Van Asch, Roser Morante, Paolo Frasconi, Walter Daelemans, and Luc De Raedt. 2012. A statistical relational learning approach to identifying evidence based medicine categories. In Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, pages 579–589.
  64. 64.Shuohang Wang, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021. Want to reduce labeling cost? gpt-3 can help. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4195–4205.
  65. 65.Yanshan Wang, Liwei Wang, Majid Rastegar-Mojarad, Sungrim Moon, Feichen Shen, Naveed Afzal, Sijia Liu, Yuqun Zeng, Saeed Mehrabi, Sunghwan Sohn, et al. 2018. Clinical information extraction applications: a literature review. Journal of biomedical informatics, 77:34–49.
  66. 66.Marc Weeber, James G Mork, and Alan R Aronson. 2001. Developing a test collection for biomedical word sense disambiguation. In Proceedings of the AMIA Symposium, page 746. American Medical Informatics Association.
  67. 67.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Fine-tuned language models are zero-shot learners. arXiv:2109.01652.
  68. 68.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, et al. 2019. Huggingface’s transformers: State-of-the-art natural language processing. arXiv preprint arXiv:1910.03771.
  69. 69.Stephen Wu, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiong Wang, Qiang Wei, Yang Xiang, et al. 2020. Deep learning in clinical natural language processing: a methodical review. Journal of the American Medical Informatics Association, 27(3):457–470.
  70. 70.Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In CHI Conference on Human Factors in Computing Systems, pages 1–22.
  71. 71.Yonghui Wu, Jun Xu, Yaoyun Zhang, and Hua Xu. 2015. Clinical abbreviation disambiguation using neural word embeddings. In Proceedings of BioNLP 15, pages 171–176.
  72. 72.Fei Xia and Meliha Yetisgen-Yildiz. 2012. Clinical corpus annotation: challenges and strategies. In Proceedings of the Third Workshop on Building and Evaluating Resources for Biomedical Text Mining (BioTxtM’2012) in conjunction with the International Conference on Language Resources and Evaluation (LREC), Istanbul, Turkey.
  73. 73.Xiaohan Yang, Eduardo Peynetti, Vasco Meerman, and Chris Tanner. 2022. What gpt knows about who is who. arXiv preprint arXiv:2205.07407.
  74. 74.Jieyu Zhang, Yue Yu, Yinghao Li, Yujing Wang, Yaming Yang, Mao Yang, and Alexander Ratner. 2021. Wrench: A comprehensive benchmark for weak supervision. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  75. 75.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  76. 76.Zachariah Zhang, Jingshu Liu, and Narges Razavian. 2020. Bert-xml: Large scale automated icd coding using bert pretraining. arXiv preprint arXiv:2006.03685.
  77. 77.Jiaping Zheng, Wendy W Chapman, Rebecca S Crowley, and Guergana K Savova. 2011. Coreference resolution: A review of general methodologies and applications in the clinical domain. Journal of biomedical informatics, 44(6):1113–1122.
  78. 78.Pierre Zweigenbaum, Dina Demner-Fushman, Hong Yu, and Kevin B Cohen. 2007. Frontiers of biomedical text mining: current progress. Briefings in bioinformatics, 8(5):358–375.

Citation

MLA
Agrawal, M., et al. “Large Language Models Are Few-shot Clinical Information Extractors”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 1998–2022, https://doi.org/10.18653/v1/2022.emnlp-main.130.
APA
Agrawal, M., Hegselmann, S., Lang, H., Kim, Y., & Sontag, D. (2022). Large language models are few-shot clinical information extractors. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1998–2022. https://doi.org/10.18653/v1/2022.emnlp-main.130
Chicago
Agrawal, M., S. Hegselmann, H. Lang, Y. Kim, and D. Sontag. 2022. “Large Language Models Are Few-shot Clinical Information Extractors”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1998–2022. https://doi.org/10.18653/v1/2022.emnlp-main.130.
Harvard
Agrawal, M. et al. (2022) “Large language models are few-shot clinical information extractors”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1998–2022. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.130.
Vancouver
1. Agrawal M, Hegselmann S, Lang H, Kim Y, Sontag D (2022) Large language models are few-shot clinical information extractors. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1998–2022

BibTeX

@inproceedings{agrawal-etal-2022-large,
    title = "Large language models are few-shot clinical information extractors",
    author = "Agrawal, Monica  and
      Hegselmann, Stefan  and
      Lang, Hunter  and
      Kim, Yoon  and
      Sontag, David",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.130/",
    doi = "10.18653/v1/2022.emnlp-main.130",
    pages = "1998--2022"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/