Large language models are few-shot clinical information extractors
Monica AgrawalStefan HegselmannHunter LangYoon KimDavid A. Sontag
Demonstrates that general-domain large language models can perform complex few-shot clinical information extraction tasks, such as span identification and relation extraction, significantly outperforming existing baselines without domain-specific training.
Clinical notes contain critical patient information, yet unlocking this data remains difficult because medical text relies heavily on ambiguous abbreviations, nonstandard phrasing, and sensitive patient records that cannot be broadly shared. Building conventional machine learning systems for clinical text typically demands labor-intensive domain engineering, brittle rule sets, and large volumes of expert-annotated training data. The article evaluates whether modern, domain-agnostic large language models can perform accurate information extraction from clinical records using zero-shot and few-shot prompting, and it demonstrates how to transform the natural-language outputs of these models into structured clinical variables.
To conduct this evaluation, the researchers tested large language models across five diverse clinical tasks: abbreviation sense disambiguation, biomedical evidence extraction from clinical trial abstracts, pronoun coreference resolution, medication status extraction, and medication attribute relation extraction. Because strict data use agreements prevent sending most existing clinical benchmarks to third-party model programming interfaces, the authors created and publicly released three new benchmark datasets by re-annotating snippets from the public Clinical Acronym Sense Inventory. The approach couples guided prompting—supplying formatting instructions and minimal demonstration examples—with simple software post-processors called resolvers to map unstructured text generations directly into structured labels.
The investigation produced several key findings. First, guided large language models consistently outperformed strong baseline systems across multiple tasks without task-specific architecture tuning. In medication extraction, a one-shot guided model achieved 0.90 recall and 0.92 precision, substantially exceeding a standard biomedical rule-based baseline at 0.73 recall and 0.67 precision. Second, for clinical trial arm identification, the zero-shot language model achieved 0.85 abstract-level accuracy, identifying the correct study arms in 17 of 20 trials, compared to only 0.35 accuracy for a supervised biomedical language model that required thousands of training examples. Third, for abbreviation disambiguation, zero-shot prompting reached 0.86 accuracy and 0.69 macro F1 score on the acronym inventory, outperforming specialized baselines. Furthermore, when these outputs were used as weak supervision to train a smaller biomedical model, performance rose to 0.90 accuracy and transferred effectively to an independent critical care dataset. Fourth, the experiments revealed that including even a single formatted example reduced post-processing code complexity from dozens of lines to fewer than ten lines of code, and that the structural format of the example mattered more than whether its label was factually correct.
These findings indicate that general-purpose large language models can dramatically lower the engineering cost and annotation burden required to extract complex clinical information. Practitioners do not necessarily need massive labeled datasets or brittle custom rules to achieve high extraction accuracy. Moreover, distilling large language model outputs into smaller, locally hosted models provides a viable pathway to overcome patient privacy constraints and recurring external computing costs.
Organizations exploring automated clinical extraction should adopt guided prompt formatting to simplify data pipelines and minimize engineering overhead. Teams can leverage large language models to generate weak labels on de-identified public data and distill those capabilities into smaller, internally hosted models that safeguard patient privacy. Before operational deployment in high-stakes clinical workflows, institutions should validate systems locally across diverse patient demographics, clinical specialties, and note types, while implementing verification steps to prevent models from generating answers when target entities are absent.
Confidence in the reported extraction capabilities is high for English-language notes within the evaluated scopes, supported by rigorous manual adjudication and baseline comparisons. However, readers should note that the evaluation relied primarily on dictated snippets from one multi-hospital system and published biomedical abstracts. Additional validation across broader international healthcare systems, diverse clinical documentation practices, and non-English clinical texts remains necessary.
- Paper: Publicly Available Clinical BERT Embeddings, Emily Alsentzer et al. (2019). It establishes standard clinical language modeling and named entity recognition baselines using adapted BERT representations on clinical notes, setting the foundational domain context that few-shot LLM approaches aim to replace.
- Paper: ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission, Kexin Huang et al. (2019). It demonstrates the domain-specific challenges and representation modeling requirements of clinical notes from MIMIC datasets that motivate zero- and few-shot information extraction methods.
- Paper: Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing, Yu Gu et al. (2020). It introduces standard domain-specific biomedical NLP benchmarks (BLURB) and highlights traditional fine-tuning methodologies against which LLM few-shot extractors are evaluated.
- Paper: BioBERT: a pre-trained biomedical language representation model for biomedical text mining, Jinhyuk Lee et al. (2019). It details how pre-trained transformers can extract biomedical entities and relations, providing essential background on traditional supervised extraction before few-shot general LLMs.
- Paper: Language Models are Unsupervised Multitask Learners, Alec Radford et al. (2019). It introduces the core paradigm of using large autoregressive language models as zero-shot and few-shot multitask learners, which the target paper adapts to structured clinical information extraction.
- Paper: Revisiting Relation Extraction in the era of Large Language Models, Somin Wadhwa et al. (2023). It builds directly on the application of LLMs for biomedical relation extraction by analyzing evaluation pitfalls and distilling chain-of-thought capabilities into smaller open-source models.
- Paper: GPT-RE: In-context Learning for Relation Extraction using Large Language Models, Zhen Wan et al. (2023). It extends LLM-based in-context relation extraction by introducing task-aware demonstration selection and label-induced reasoning logic to solve structural extraction errors.
- Paper: Empirical Study of Zero-Shot NER with ChatGPT, Tingyu Xie et al. (2023). It advances zero-shot structured named entity recognition with chat-tuned LLMs across specialized domains using multi-step decomposition and external linguistic prompting.
- Paper: Capabilities of GPT-4 on Medical Challenge Problems, Harsha Nori et al. (2023). It examines the next generation of foundation models (GPT-4) on complex medical problem-solving and clinical benchmarks without specialized fine-tuning.
- Paper: Prompting Language Models for Linguistic Structure, Terra Blevins et al. (2023). It provides a generalized exploration of structured prompting to elicit token-level sequences and entity extraction from autoregressive LLMs without fine-tuning.
- Paper: GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer, Urchade Zaratiana et al. (2024). It presents a lightweight, non-autoregressive alternative for open-vocabulary and zero-shot entity extraction across domains, addressing the computational bottlenecks of massive LLMs.
