LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Chunyuan LiCliff WongSheng ZhangNaoto UsuyamaHaotian LiuJianwei YangTristan NaumannHoifung PoonJianfeng Gao

article2023NeurIPS1,584 citations

Presents LLaVA-Med, a multimodal conversational assistant trained in under fifteen hours using curriculum learning and GPT-4-generated instruction tuning from PubMed Central to accurately answer open-ended questions about biomedical images.

Listen

Conversational artificial intelligence has shown major promise in supporting healthcare workflows, yet existing models have largely been restricted to processing text alone. While general-domain multimodal assistants can interpret standard web images, they routinely fail, hallucinate, or avoid responding when presented with complex biomedical imagery such as X-rays, computed tomography scans, and pathology slides. At the same time, traditional biomedical visual question answering systems are typically trained as narrow classification models that select from a fixed list of answers, making them unsuitable for dynamic, open-ended clinical dialogue. To bridge this gap, the article set out to develop and evaluate a cost-effective, end-to-end multimodal conversational assistant capable of answering open-ended visual questions across biomedical domains.

The researchers developed an automated pipeline to generate training data without manual annotation. Using a large public repository of biomedical literature, they extracted figure-caption pairs along with surrounding text and prompted a state-of-the-art language model to generate multi-turn conversational instructions. They then adapted a general-domain vision-language model using a two-stage curriculum learning method. In the first stage, the model aligned biomedical concepts by learning to match 600,000 biomedical images to descriptive captions while language weights remained frozen. In the second stage, the model underwent full instruction-tuning using 60,000 multi-turn conversational examples across five major imaging modalities. The team evaluated the resulting system on conversational benchmarks using automated scoring and on three established biomedical question-answering benchmark datasets.

The evaluation revealed several key findings regarding model efficiency and performance. First, the entire two-stage training process was completed in less than 15 hours using eight standard graphics processing units, demonstrating a highly resource-efficient training pathway. Second, in open-ended conversational evaluations, the adapted model achieved a relative score of 50.2 percent against a text-informed upper baseline, significantly outperforming the general-domain base model's score of 36.1 percent. Third, incorporating surrounding article text as external context during data generation measurably boosted conversational performance over captions alone. Finally, after fine-tuning on downstream benchmark datasets, the model surpassed existing state-of-the-art methods on closed-ended visual questions across multiple benchmarks, while maintaining competitive generative performance on open-ended questions.

These findings indicate that general-domain multimodal artificial intelligence can be effectively and economically specialized for high-value vertical domains without requiring massive proprietary datasets or prolonged training timelines. By framing visual question answering as open-ended text generation rather than rigid classification, the system offers a more flexible foundation for clinical decision support, medical education, and diagnostic research. Furthermore, the model exhibited zero-shot multilingual understanding by correctly answering questions in languages absent from the instruction training set.

Organizations planning to develop domain-specific visual assistants should adopt this two-stage curriculum strategy and leverage automated instruction generation from scientific literature to reduce data collection costs. Before deploying such assistants in live healthcare environments, stakeholders must conduct extensive safety pilots and human-in-the-loop clinical evaluations. Decision-makers should note that the model is still prone to hallucinations and exhibits limitations in deep multi-step clinical reasoning. The system is intended as an assistive research tool rather than an autonomous diagnostic device, and strong governance is required prior to operational adoption.

  • Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). LLaVA provides the foundational vision-language instruction-tuning architecture and methodology that LLaVA-Med adapts and scales to the biomedical domain.
  • Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). Improved Baselines with Visual Instruction Tuning establishes critical architectural and training refinements that LLaVA-Med builds upon for robust multimodal performance.
  • Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision directly extends LLaVA-Med by scaling the unified multimodal framework to handle complex video streams, multi-image sequences, and massive task transfers.
Cover for LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

Abstract

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vision-language models still lack sophistication in understanding and conversing about biomedical images. In this paper, we propose a cost-efficient approach for training a vision-language conversational assistant that can answer open-ended research questions of biomedical images. The key idea is to leverage a large-scale, broad-coverage biomedical figure-caption dataset extracted from PubMed Central, use GPT-4 to self-instruct open-ended instruction-following data from the captions, and then fine-tune a large general-domain vision-language model using a novel curriculum learning method. Specifically, the model first learns to align biomedical vocabulary using the figure-caption pairs as is, then learns to master open-ended conversational semantics using GPT-4 generated instruction-following data, broadly mimicking how a layperson gradually acquires biomedical knowledge. This enables us to train a Large Language and Vision Assistant for BioMedicine (LLaVA-Med) in less than 15 hours (with eight A100s). LLaVA-Med exhibits excellent multimodal conversational capability and can follow open-ended instruction to assist with inquiries about a biomedical image. On three standard biomedical visual question answering datasets, LLaVA-Med outperforms previous supervised state-of-the-art on certain metrics. To facilitate biomedical multimodal research, we will release our instruction-following data and the LLaVA-Med model.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Biomedical Visual Instruction-Following Data
  • 4 Adapting Multimodal Conversational Models to the Biomedical Domain
  • 5 Experiments
  • 5.1 Biomedical Visual Chatbot
  • 5.2 Performance on Established Benchmarks
  • 6 Conclusions
  • References
  • A Data
  • B Prompts

Knowls

  1. Knowl 1 — LLaVA-Med Architecture and Two-Stage Curriculum Adaptation

    model/method

    LLaVA-Med adapts a general-domain vision-language architecture (LLaVA) to the biomedical domain through a two-stage curriculum learning pipeline:

    1. Architecture: A vision encoder (either general-domain CLIP ViT-L/14 or domain-specific BioMedCLIP) extracts visual features from an input biomedical image. A trainable linear projection matrix maps visual tokens into the embedding space of a causal large language model (e.g., a 7B or 13B parameter LLaMA/Vicuna/LLaVA model).

    2. Stage 1: Biomedical Concept Feature Alignment: The objective is to align biomedical visual concepts with corresponding language token embeddings without altering pre-trained general linguistic capacities. Both the visual encoder and the causal language model are kept frozen, and only the linear projection layer is trained on 600,000 biomedical image-caption pairs extracted from PubMed Central (PMC-15M) for 1 epoch. The training objective is standard next-token language modeling of the caption given the image and a generic captioning prompt.

    3. Stage 2: End-to-End Instruction-Tuning: The visual encoder remains frozen, while both the linear projection layer and the causal language model are jointly fine-tuned on 60,000 multi-turn biomedical visual instruction-following conversations for 3 epochs. This step teaches the model to engage in multi-turn conversational reasoning and open-ended question answering about biomedical visual inputs.

  2. Knowl 2 — Biomedical Multimodal Instruction-Following Data Generation via GPT-4 and Citances

    model/method

    To train a multimodal biomedical conversational assistant without manual visual annotations, a self-instruct data curation pipeline leverages text-only GPT-4 alongside PubMed Central (PMC-15M) figure data:

    1. Data Sampling: Filter PMC-15M to select single-panel images representing five major biomedical imaging modalities: chest X-ray (CXR), computed tomography (CT), magnetic resonance imaging (MRI), histopathology, and gross pathology (sampling 60,000 total image-caption instances).

    2. Context Enrichment (Inline Mentions / Citances): Because figure captions in scientific literature are often too brief to support rich dialogue generation, the pipeline extracts the in-text sentences citing the figure (termed citances or inline mentions) from the full-text PubMed paper. The combined caption and citances serve as comprehensive background knowledge for GPT-4.

    3. Few-Shot Self-Instruct Prompting: Language-only GPT-4 is prompted with few-shot demonstration conversations to generate 2 to 3 dialogue turns per image. GPT-4 is instructed to adopt the persona of a visual assistant that can inspect the image directly, to avoid meta-linguistic references (such as "caption", "context", or specific bibliographic details), and to include safety disclaimers avoiding medical advice or diagnostic confirmation.

  3. Knowl 3 — Prompt Stratification for Biomedical Concept Alignment Data

    model/method

    For Stage 1 biomedical concept alignment, 600,000 image-caption pairs from PMC-15M are transformed into single-turn instruction-following instances formatted as:

    Human:Xq Xv<STOP>\n Assistant:Xc<STOP>\n\text{Human}: X_q\ X_v\text{<STOP>}\backslash\text{n } \text{Assistant}: X_c\text{<STOP>}\backslash\text{n}

    where XvX_v is the visual token sequence, XcX_c is the target caption text, and XqX_q is a sampled prompt instructing the model to describe the image.

    To ensure syntactic diversity matching caption verbosity, XqX_q is selected from two distinct instruction pools based on caption length:

    • If the caption length is under 30 words (which comprises approximately 25% of PMC-15M pairs), XqX_q is sampled from a set of concise description instructions (e.g., "Describe the image concisely.", "Offer a succinct explanation of the picture presented.").
    • If the caption length is 30 words or more, XqX_q is sampled from a set of detailed description instructions (e.g., "Describe the following image in detail", "Analyze the image in a comprehensive and detailed manner").
  4. Knowl 4 — Biomedical Visual Chatbot Evaluation Benchmark and Results

    data/table

    To evaluate open-ended multimodal conversation capability, a benchmark of 193 novel questions was created from 50 unseen PMC-15M image-caption pairs across two question types: multi-turn conversation (143 questions) and detailed image description (50 questions), covering five modalities (CXR, MRI, Histology, Gross pathology, CT). Responses from candidate models and language-only GPT-4 (supplied with ground-truth captions and citances as an upper bound) are scored by GPT-4 on a 1–10 scale across helpfulness, relevance, accuracy, and detail. Relative scores normalized against GPT-4 performance are reported:

    Model Question Types Domains Overall
    Conversation Description CXR MRI Histology Gross CT (193)
    (143) (50) (37) (38) (44) (34) (40)
    LLaVA 39.4 26.2 41.6 33.4 38.4 32.9 33.4 36.1
    LLaVA-Med (Stage 1 only) 22.6 25.2 25.8 19.0 24.8 24.7 22.2 23.3
    LLaVA-Med (10K) 42.4 32.5 46.1 36.7 43.5 34.7 37.5 39.9
    LLaVA-Med (60K) 53.7 36.9 57.3 39.8 49.8 47.4 52.4 49.4
    LLaVA-Med (60K-IM) 55.1 36.4 56.2 40.4 52.7 51.8 50.1 50.2

    Stage 1 alignment alone degrades chat versatility (23.3% overall), whereas Stage 2 instruction tuning with inline-mention-enriched data (60K-IM) achieves 50.2% relative to GPT-4.

  5. Knowl 5 — Supervised Fine-Tuning Performance on Biomedical VQA Benchmarks

    data/table

    LLaVA-Med was evaluated on three established medical visual question answering datasets: VQA-RAD (radiology), SLAKE (English radiology subset), and PathVQA (pathology). For closed-set questions (e.g., yes/no), performance is measured by classification accuracy (%). For open-set questions, LLaVA-Med uses free-form text generation evaluated via recall (%) of ground-truth tokens in the generated response, whereas prior baseline methods typically formulated open-set VQA as closed-vocabulary classification over training set candidate answers:

    Method VQA-RAD SLAKE PathVQA
    Open Closed Open Closed Open Closed
    Supervised Fine-Tuned
    LLaVA (General Domain) 50.00 65.07 78.18 63.22 7.74 63.20
    LLaVA-Med (From LLaVA init.) 61.52 84.19 83.08 85.34 37.95 91.21
    LLaVA-Med (From Vicuna init.) 64.39 81.98 84.71 83.17 38.87 91.65
    LLaVA-Med (BioMedCLIP vision) 64.75 83.09 87.11 86.78 39.60 91.09
    Literature Baselines
    VL Encoder–Decoder 71.49 82.47 – – 71.49 85.61
    Q2ATransformer 79.19 81.20 – – 54.85 88.85
    Prefix T. Medical LM – – 84.30 82.01 40.00 87.00
    PubMedCLIP 60.10 80.00 78.40 82.50 – –
    BiomedCLIP 67.60 79.80 82.05 89.70 – –
    M2I2 66.50 83.50 74.70 91.10 36.30 88.00

    Fine-tuned LLaVA-Med sets new state-of-the-art results on closed-set questions on VQA-RAD (84.19%) and PathVQA (91.65%).

  6. Knowl 6 — Ablations on Instruction Data Volume, Citance Context, Model Scale, and Vision Encoders

    data/table

    Ablation studies on VQA-RAD, SLAKE, and PathVQA demonstrate the impact of training epochs across Stage 1, Stage 2, downstream fine-tuning (FT), data selection (10K vs. 60K vs. 60K-IM with inline mentions), model parameter size (7B vs. 13B LM), and vision backbone (general-domain CLIP vs. BioMedCLIP):

    Config (Vision / LM) Data S1 S2 FT VQA-RAD SLAKE PathVQA Avg
    (Op / Cl) (Op / Cl) (Op / Cl)
    CLIP / 7B (Base LLaVA) 0 0 0 0 20.74 / 59.19 26.82 / 50.24 8.74 / 45.65 35.23
    CLIP / 7B 0 1 0 0 15.27 / 12.50 18.55 / 13.46 6.26 / 13.51 13.26
    CLIP / 7B 10K 1 3 0 25.79 / 57.35 31.50 / 51.68 8.49 / 59.66 39.08
    CLIP / 7B 60K 1 3 0 29.67 / 60.29 35.53 / 53.85 11.76 / 53.20 40.72
    CLIP / 7B 60K-IM 1 3 0 28.23 / 61.40 39.17 / 52.16 12.30 / 54.05 41.22
    CLIP / 7B 60K-IM 1 3 3 55.50 / 66.54 80.57 / 64.18 35.88 / 89.15 65.30
    CLIP / 7B 60K-IM 1 3 9 66.26 / 80.88 82.30 / 84.86 37.59 / 91.54 73.90
    CLIP / 7B 60K-IM 1 3 15 61.53 / 84.19 83.08 / 85.34 37.95 / 91.21 73.88
    CLIP / 13B 60K-IM 1 3 0 31.66 / 61.40 37.71 / 49.76 11.34 / 49.63 40.25
    CLIP / 13B 60K-IM 1 3 9 64.58 / 77.94 84.97 / 85.58 38.82 / 92.39 74.05
    BioMedCLIP / 7B 60K-IM 1 3 0 37.84 / 60.66 39.73 / 54.33 11.65 / 49.07 42.21
    BioMedCLIP / 7B 60K-IM 1 3 9 64.75 / 83.09 87.11 / 86.78 39.60 / 91.09 75.40

    Key takeaways:

    1. Zero-shot performance increases with larger instruction datasets and inclusion of inline mentions (60K-IM reaches 41.22% average vs. 35.23% for base LLaVA).
    2. Downstream fine-tuning converges around 9–15 epochs.
    3. Initializing with domain-specific BioMedCLIP yields the highest overall fine-tuned average score (75.40%).
  7. Knowl 7 — LLaVA-Med Computational Training Time and Efficiency

    data/table

    LLaVA-Med domain adaptation is designed for high computational efficiency, training completely in less than 15 hours on eight 40GB NVIDIA A100 GPUs using a batch size of 128:

    Stage 1 Stage 2 VQA-RAD SLAKE PathVQA
    1 ep 3 ep Dataset 1 ep 3 ep 1 ep 3 ep 1 ep 3 ep 1 ep 3 ep
    6.8h 19.4h 10K 0.6h 1.8h 0.3h 0.6h 0.6h 1.0h 1.0h 2.5h
    60K 2.6h 8.0h

    The standard training recipe (1 epoch of Stage 1 concept alignment on 600K pairs for 6.8 hours, followed by 3 epochs of Stage 2 instruction tuning on 60K pairs for 8.0 hours) totals 14.8 GPU hours on a single 8-GPU node.

  8. Knowl 8 — Zero-Shot Cross-Lingual Transfer on Biomedical Visual Questions

    empirical result

    Although LLaVA-Med was trained exclusively on English biomedical instruction-following data, it exhibits zero-shot cross-lingual transfer on non-English biomedical visual questions. When presented with Chinese visual queries on the SLAKE dataset (e.g., asking for the imaging modality "这张图片的成像方式是什么?" or the MR sequence type "这张图片展示的是核磁共振的哪种类型?"), LLaVA-Med accurately comprehends the medical intent and returns correct diagnostic answers in English (e.g., identifying abdominal computed tomography in the portal phase or T1-weighted MRI nodular hyperintensity). This capability is attributed to multilingual semantic transfer preserved from the underlying LLaMA/Vicuna base language model.

  9. Knowl 9 — Clinical Reasoning Depth and Hallucination Limitations in LLaVA-Med

    limitation

    Despite strong performance in multi-turn conversation and closed-set benchmark evaluations, LLaVA-Med has two principal limitations:

    1. Hallucination and Factuality: Like general-domain large multimodal models, LLaVA-Med can hallucinate biomedical findings or misidentify fine-grained pathological features when visual cues are ambiguous.
    2. Depth of Complex Diagnostic Reasoning: The model exhibits weaker performance on open-set questions requiring complex multi-step clinical reasoning without constrained answer candidates. Consequently, the model is not intended to provide clinical diagnostic advice or replace qualified healthcare practitioners.

Coverage note — No substantial contributed material was omitted. All primary methodological stages, data generation pipelines, empirical benchmarks (visual chat, VQA-RAD, SLAKE, PathVQA), ablation studies, training costs, cross-lingual findings, and limitations are covered.

References

  1. 1.Clinical Camel. https://wanglab.ml/clinical_camel.html, 2023.
  2. 2.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  3. 3.Malek Ayoub, Megan Quamme, Abdul-Rahman K Abdel-Reheem, Poe Lwin, and Megan K Quamme. Covid or not covid? a great mimicker behind the smoke screen. Cureus, 13(11), 2021.
  4. 4.Bappy Basak, Alexander Haragan, Michael Shackcloth, and Joyce Thekkudan. Chondromyxoid fibroma of the rib: A rare benign tumor with potential for local recurrence. Cureus, 13(10), 2021.
  5. 5.Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Laila Bashmal, and Mansour Zuair. Vision–language model for visual question answering in medical imagery. Bioengineering, 2023.
  6. 6.Anchit Bharat, Nikita Jain, Belaal Sheikh, Hafiz Jeelani, and Maryna Shayuk. Vaping-induced lung injury: An uncharted territory. Cureus, 12, 07 2020.
  7. 7.Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. Making the most of text semantics to improve biomedical vision–language processing. In ECCV. Springer, 2022.
  8. 8.Sedigheh Eslami, Christoph Meinel, and Gerard De Melo. Pubmedclip: How much does clip benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1151–1163, 2023.
  9. 9.Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao, et al. Vision–language pre-training: Basics, recent advances, and future trends. Foundations and Trends® in Computer Graphics and Vision, 2022.
  10. 10.Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1–23, 2021.
  11. 11.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. Don’t stop pretraining: Adapt language models to domains and tasks. arXiv preprint arXiv:2004.10964, 2020.
  12. 12.Tianyu Han, Lisa C Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K Bressem. Medalpaca–an open-source collection of medical conversational ai models and training data. arXiv preprint arXiv:2304.08247, 2023.
  13. 13.Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286, 2020.
  14. 14.Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv preprint arXiv:1904.05342, 2019.
  15. 15.Alistair EW Johnson, Tom J Pollard, Seth J Berkowitz, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Roger G Mark, and Steven Horng. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, page 317, 2019.
  16. 16.Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 2018.
  17. 17.Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240, 2020.
  18. 18.Peter Lee, Sebastien Bubeck, and Joseph Petro. Benefits, limits, and risks of gpt-4 as an ai chatbot for medicine. New England Journal of Medicine, 388(13):1233–1239, 2023.
  19. 19.Peter Lee, Carey Goldberg, and Isaac Kohane. The ai revolution in medicine: Gpt-4 and beyond. 2023.
  20. 20.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. NeurIPS, 2020.
  21. 21.Chunyuan Li, Haotian Liu, Liunian Harold Li, Pengchuan Zhang, Jyoti Aneja, Jianwei Yang, Ping Jin, Houdong Hu, Zicheng Liu, Yong Jae Lee, and Jianfeng Gao. ELEVATER: A benchmark and toolkit for evaluating language-augmented visual models. In NeurIPS Track on Datasets and Benchmarks, 2022.
  22. 22.Pengfei Li, Gang Liu, Lin Tan, Jinying Liao, and Shenjun Zhong. Self-supervised vision–language pretraining for medical visual question answering. arXiv preprint arXiv:2211.13594, 2022.
  23. 23.Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In International Symposium on Biomedical Imaging (ISBI). IEEE, 2021.
  24. 24.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023.
  25. 25.Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, and Chunyuan Li. Learning customized visual models with retrieval-augmented knowledge. arXiv preprint arXiv:2301.07094, 2023.
  26. 26.Yunyi Liu, Zhanyu Wang, Dong Xu, and Luping Zhou. Q2atransformer: Improving medical vqa via an answer querying decoder. arXiv preprint arXiv:2304.01611, 2023.
  27. 27.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 2022.
  28. 28.Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 2022.
  29. 29.Hassan Mirmohammad Sadeghi, Abbas Karimi, Samira Derakhshan, Pouyan Aminishakib, and Kiarash Parchami. Conventional osteosarcoma of the mandible: Report of a rare case. Clinical Case Reports, 9(9):e04843, 2021.
  30. 30.Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375, 2023.
  31. 31.OpenAI. ChatGPT. https://openai.com/blog/chatgpt/, 2022.
  32. 32.OpenAI. GPT-4 technical report. https://arxiv.org/abs/2303.08774, 2023.
  33. 33.Kyriakos A Papavasiliou, Dimitrios Stamiris, Stavros Stamiris, Antonia Bintoudi, and Eleftherios Tsiridis. Quadratus femoris partial tear secondary to occult ischiofemoral impingement. Journal of Orthopaedic Case Reports, 11(9):7, 2021.
  34. 34.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. Instruction tuning with GPT-4. arXiv preprint arXiv:2304.03277, 2023.
  35. 35.Roger Kevin Pringle and Lawrence H Wyatt. The appropriate use of radiography in clinical practice: a report of two cases of biomechanical versus malignant spine pain. Chiropractic & Osteopathy, 14(1):1–8, 2006.
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021.
  37. 37.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 2019.
  38. 38.George Shih, Carol C Wu, Safwan S Halabi, Marc D Kohli, Luciano M Prevedello, Tessa S Cook, Arjun Sharma, Judith K Amorosa, Veronica Arteaga, Maya Galperin-Aizenberg, et al. Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia. Radiology: Artificial Intelligence, 2019.
  39. 39.Chang Shu, Baian Chen, Fangyu Liu, Zihao Fu, Ehsan Shareghi, and Nigel Collier. Visual med-alpaca: A parameter-efficient biomedical llm with visual capabilities. 2023.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  41. 41.Tom van Sonsbeek, Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Cees GM Snoek, and Marcel Worring. Open-ended medical visual question answering through prefix tuning of language models. arXiv preprint arXiv:2303.05977, 2023.
  42. 42.A Venigalla, J Frankle, and M Carbin. BiomedLM: a domain-specific large language model for biomedical text. MosaicML. Accessed: Dec, 23, 2022.
  43. 43.Vicuna. Vicuna: An open-source chatbot impressing GPT-4 with 90%* chatgpt quality. https://vicuna.lmsys.org/, 2023.
  44. 44.Haochun Wang, Chi Liu, Nuwa Xi, Zewen Qiang, Sendong Zhao, Bing Qin, and Ting Liu. Huatuo: Tuning llama model with chinese medical knowledge, 2023.
  45. 45.Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-llama: Further finetuning llama on medical papers. arXiv preprint arXiv:2304.14454, 2023.
  46. 46.Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Qian Wang, and Dinggang Shen. Doctorglm: Fine-tuning your chinese doctor is not a herculean task. arXiv preprint arXiv:2304.01097, 2023.
  47. 47.Li Yunxiang, Li Zihan, Zhang Kai, Dan Ruilong, and Zhang You. Chatdoctor: A medical chat model fine-tuned on llama model using medical domain knowledge. arXiv preprint arXiv:2303.14070, 2023.
  48. 48.Mansoor Zafar, Abdul Wahab Paracha, Muteeb Ashraf, Tila Muhammad, Mark Whitehead, Muhammad Toqeer, and Abdul Paracha. Delayed spontaneous regression of metastatic gastric cancer: A case report of a rare finding. Cureus, 13(12), 2021.
  49. 49.Sheng Zhang, Yanbo Xu, Naoto Usuyama, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, et al. Large-scale domain-specific pretraining for biomedical vision-language processing. arXiv preprint arXiv:2303.00915, 2023.

Citation

MLA
Li, C., et al. “LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day”. arXiv, 2023, http://arxiv.org/abs/2306.00890v1.
APA
Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., & Gao, J. (2023). LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. arXiv. http://arxiv.org/abs/2306.00890v1
Chicago
Li, C., C. Wong, S. Zhang, et al. 2023. “LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day”. arXiv. http://arxiv.org/abs/2306.00890v1.
Harvard
Li, C. et al. (2023) “LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.00890v1.
Vancouver
1. Li C, Wong C, Zhang S, Usuyama N, Liu H, Yang J, Naumann T, Poon H, Gao J (2023) LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. arXiv

BibTeX

@article{li2023llava,
  title = {LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day},
  author = {Li, Chunyuan and Wong, Cliff and Zhang, Sheng and Usuyama, Naoto and Liu, Haotian and Yang, Jianwei and Naumann, Tristan and Poon, Hoifung and Gao, Jianfeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.00890v1},
  eprint = {2306.00890}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission