HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

Tianwei LinWenqiao ZhangSijing LiYuqian YuanBinhe YuHaoyuan LiWanggui HeHao JiangMengze LiXiaohui Song

article2025ICML137 citations

Presents HealthGPT, a medical large vision-language model that unifies multimodal comprehension and image generation tasks within a single autoregressive framework using a heterogeneous low-rank adaptation technique to prevent task interference.

Listen

Artificial intelligence applications in healthcare have made notable strides using vision-language foundation models, yet current systems typically remain specialized for single functions. Most medical models focus strictly on visual comprehension, such as answering clinical questions or drafting textual reports, and lack the generative capabilities required to create or enhance medical imagery. Meanwhile, general-domain unified models that attempt both tasks often struggle in healthcare settings due to the scarcity of high-quality medical data and the inherent conflict between abstract comprehension tasks and detail-heavy generation tasks.

The article demonstrates and evaluates HealthGPT, a medical large vision-language model designed to unify clinical visual understanding and image generation within a single autoregressive framework. The system aims to bridge the gap between diagnostic reasoning and medical image synthesis using parameter-efficient fine-tuning on limited domain-specific data.

The researchers developed a three-part technical framework evaluated across multiple open medical benchmarks. First, they introduced a parameter-efficient fine-tuning technique called Heterogeneous Low-Rank Adaptation, which uses dynamic expert routing and matrix block merging to separate comprehension and generation knowledge into distinct plugins while keeping base models frozen. Second, a hierarchical visual perception mechanism routes fine-grained visual details from early neural network layers to generation tasks and abstract features from deeper layers to comprehension tasks. Third, the authors introduced a three-stage training strategy combined with a curated dataset named VL-Health, which contains over 1.5 million instruction-following samples across 11 imaging modalities, such as computed tomography, magnetic resonance imaging, and X-rays.

Evaluation results show that HealthGPT consistently outperforms both medical-specific models and general-purpose multimodal models across several critical benchmarks. In medical visual question answering across seven modalities on the OmniMedVQA benchmark, the compact 3.8-billion parameter HealthGPT-M3 scored 68.5, substantially higher than the specialized medical model HuatuoGPT-Vision (50.0) and the open-world model Llama-3.2 (63.2). Larger variants scaled this performance further, with the 32-billion parameter model achieving a benchmark score of 77.2. For generation tasks, HealthGPT surpassed dedicated standalone tools in cross-modality translation—such as converting computed tomography scans to magnetic resonance imaging—and achieved state-of-the-art detail restoration in four-times image super-resolution. Furthermore, its custom adaptation structure operated without the linear training slowdowns found in standard mixture-of-experts approaches, running approximately 33% faster with four experts.

These findings suggest that healthcare organizations can deploy a single consolidated artificial intelligence model rather than maintaining isolated systems for diagnosis, reporting, and image processing. Unifying these tasks within a shared model reduces computational training costs, lowers infrastructure maintenance demands, and minimizes the risk of catastrophic forgetting when handling diverse clinical imaging tasks. Blinded evaluations by five practicing clinicians further validated these results, selecting HealthGPT's responses as superior more often than competing models in open-ended diagnostic reasoning.

Decision-makers and research teams should consider adopting decoupled adaptation architectures when deploying multimodal systems in data-constrained enterprise environments. While HealthGPT establishes a strong baseline for unified medical models, prospective adopters should conduct domain-specific clinical validation pilots before deploying it into active diagnostic pipelines, particularly for critical imaging conversions. Future work should focus on expanding the model's text-to-image reporting capabilities and exploring broader generative capabilities across larger model sizes.

arXiv: 2502.09838

No sufficiently relevant recommendations were found.

Cover for HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. HealthGPT
  • 4.1. Unified Autoregressive Generation
  • 4.2. Hierarchical Visual Perception
  • 4.3. Heterogeneous Knowledge Adaptation
  • 4.4. Training Pipeline
  • 5. Experiments
  • 5.1. Data and Experimental Setup
  • 5.2. Main Experiments
  • 5.3. In-Depth Study
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • Appendix
  • A. Implementation Details
  • A.1. Model Details
  • A.2. Training Details
  • A.3. VL-Health
  • A.3.1. DATA STATISTICS
  • A.3.2. DATA FORMAT
  • B. Analysis of Heterogeneous Low-Rank Adaptation
  • C. Supplemental Experimental Results
  • C.1. OmniMedVQA Benchmark
  • C.2. Stability Analysis of Number of Experts
  • C.3. Impact of Heterogeneous Knowledge Fusion on Performance
  • C.4. Human Evaluation
  • C.5. Reconstruction Performance
  • C.6. Case Study

Knowls

  1. Knowl 1 — HealthGPT unifies medical comprehension and image generation in one autoregressive model

    model/method

    HealthGPT uses a single large language model (LLM) to process text and medical-image features and to generate either text or images. For comprehension, it autoregressively produces text response tokens conditioned on the input image and text. For image generation, it begins with a special image-start token, predicts a sequence of discrete VQGAN image-code tokens, and ends with an image-end token; a VQGAN decoder converts the predicted codes into an image. The model therefore represents both textual responses and image outputs as next-token prediction, rather than delegating image generation to a separate external generator. HealthGPT expands the LLM vocabulary with 8,192 VQGAN code tokens. Its implementations use CLIP-L/14 for visual encoding and are based on Phi-3-mini (3.8B parameters), Phi-4 (14B), or Qwen2.5-32B (32B).

  2. Knowl 2 — H-LoRA separates task knowledge and combines routed experts without per-expert matrix products

    model/method

    Heterogeneous Low-Rank Adaptation (H-LoRA) assigns separate low-rank adaptation plugins to comprehension and generation, with a hard task-type selection choosing the relevant plugin. Within a plugin, a router produces expert weights for each token. For an input hidden-state matrix x∈RN×dinx\in\mathbb{R}^{N\times d_{\mathrm{in}}} containing NN tokens, expert ii has matrices Ai∈Rdin×rA_i\in\mathbb{R}^{d_{\mathrm{in}}\times r} and Bi∈Rr×doutB_i\in\mathbb{R}^{r\times d_{\mathrm{out}}}, where rr is the rank and din,doutd_{\mathrm{in}},d_{\mathrm{out}} are the input and output widths. The model concatenates the AiA_i matrices along their output dimension to form Acat∈Rdin×rkA_{\mathrm{cat}}\in\mathbb{R}^{d_{\mathrm{in}}\times rk} and the BiB_i matrices along their input dimension to form Bcat∈Rrk×doutB_{\mathrm{cat}}\in\mathbb{R}^{rk\times d_{\mathrm{out}}}, where kk is the number of experts. If W∈RN×kW\in\mathbb{R}^{N\times k} contains router weights, each weight is replicated rr times and scaled by the paper's factor αk/r\alpha k/r, giving Wexpanded∈RN×rkW_{\mathrm{expanded}}\in\mathbb{R}^{N\times rk}; α\alpha is the adapter scaling parameter. The adapter contribution is (xAcat⊙Wexpanded)Bcat(xA_{\mathrm{cat}}\odot W_{\mathrm{expanded}})B_{\mathrm{cat}}, where ⊙\odot is elementwise multiplication. This contribution is added to the frozen base linear transformation xW0xW_0, with W0∈Rdin×doutW_0\in\mathbb{R}^{d_{\mathrm{in}}\times d_{\mathrm{out}}}. In the paper's equal-cost operation-count analysis, H-LoRA has added overhead O(6)O(6) per adapted fully connected layer, independent of kk, compared with O(5k+1)O(5k+1) for MoELoRA. With four experts, measured training time was 1.00× for H-LoRA and 1.49× for MoELoRA. H-LoRA also performed better on most reported comprehension and all reported generation metrics; for example, OmniMedVQA scores were 68.50 for H-LoRA, 65.10 for ordinary LoRA, and 64.90 for MoELoRA. H-LoRA did not lead on every individual metric: its SLAKE all-question score was 56.4, below LoRA's 57.2.

  3. Knowl 3 — Hierarchical visual perception supplies different ViT features to comprehension and generation

    model/method

    HealthGPT divides the visual features produced by the layers of a vision transformer (ViT) into concrete-grained features from shallower layers and abstract-grained features from deeper layers. The shallower features retain more global visual detail and are selected for image-generation tasks; the deeper features encode more abstract semantics and are selected for comprehension tasks. The selected image features are aligned by visual adapters and concatenated with text features before entering the LLM. In the reported implementation, CLIP-L/14 features from its second layer provide concrete-grained inputs, while features from its penultimate layer provide abstract-grained inputs. The authors' ablation found faster convergence for comprehension with abstract-grained inputs and better generation performance with concrete-grained inputs.

  4. Knowl 4 — Three-stage learning adapts plugins separately before unifying and instruction-tuning them

    model/method

    HealthGPT is trained in three stages to limit conflicts between medical comprehension and image generation. In Stage 1, comprehension training updates an abstract-feature visual adapter while the LLM and its H-LoRA plugin remain frozen. Generation training updates a concrete-feature visual adapter and the generation H-LoRA plugin while keeping the LLM frozen; the LLM vocabulary is extended with VQGAN image-code tokens, and image-code supervision teaches reconstruction. In Stage 2, all task-specific H-LoRA plugins are frozen and a small mixed-task dataset is used to tune the shared word-embedding layer and output head, aligning the plugins with shared token representations. In Stage 3, task-specific instruction data is used to adapt the H-LoRA and visual-adapter modules for downstream comprehension and generation, while the shared embedding layer and output head remain fixed.

  5. Knowl 5 — VL-Health combines medical comprehension and image-generation data across modalities

    data/table

    VL-Health is the training dataset assembled for unified medical vision-language comprehension and generation. It contains 765,802 additional visual question-answering samples and 783,045 generation samples, in addition to LLaVA-558k and PubMedVision-PT data used for alignment. Comprehension sources include VQA-RAD, SLAKE, PathVQA, MIMIC-CXR-VQA, LLaVA-Med, and PubMedVision; the collection also incorporates open-world instruction data to retain general instruction-following ability. Generation data supports image reconstruction, super-resolution, modality conversion, and report-to-chest-X-ray generation, drawing on LLaVA-558k, IXI, SynthRAD2023, and chest-X-ray images paired with reports. The dataset covers 11 medical imaging modalities, including CT, MRI, X-ray, microscopy, OCT, ultrasound, and fundus photography. Processing includes standardizing visual question answering as open-ended or single-choice instruction-response samples, excluding multi-image samples, and preparing generation images through operations such as slicing, registration, augmentation, and normalization.

  6. Knowl 6 — HealthGPT achieves the highest reported average comprehension scores among compared unified models

    empirical result

    On the paper's medical comprehension evaluation, HealthGPT's reported average score increased with model size: HealthGPT-M3 (3.8B parameters) scored 61.3, HealthGPT-L14 (14B) scored 66.4, and HealthGPT-XL32 (32B) scored 71.1. The comparison averages combine VQA-RAD, SLAKE, and PathVQA close/all scores with MMMU-Med and OmniMedVQA (OMVQA). In that metric order, HealthGPT-M3 scored 73.7/55.9, 74.6/56.4, and 78.7/39.7, followed by 43.3 on MMMU-Med and 68.5 on OMVQA; HealthGPT-L14 scored 77.7/58.3, 76.4/64.5, and 85.9/44.4, followed by 49.2 and 74.4; HealthGPT-XL32 scored 78.1/60.5, 83.7/68.2, and 92.4/50.9, followed by 58.0 and 77.2. The reported averages for Llama-3.2 (11B), HuatuoGPT-Vision (7B), and the unified model Emu3 (8B) were 54.7, 50.7, and 47.2, respectively. These results show that the 3.8B HealthGPT-M3 exceeded the compared unified models on the reported average, while the larger HealthGPT versions scored higher still.

  7. Knowl 7 — HealthGPT performs strongly on CT–MRI modality conversion, though not on every metric

    empirical result

    HealthGPT was evaluated on four paired-image conversion tasks: CT-to-MRI and MRI-to-CT for brain and pelvis images. Metrics are structural similarity (SSIM; higher is better), peak signal-to-noise ratio (PSNR; higher is better), and mean squared error (MSE; lower is better). In the order SSIM, PSNR, MSE, HealthGPT-M3 scored 79.38, 33.03, 33.48 for CT-to-MRI brain; 71.81, 31.83, 43.45 for CT-to-MRI pelvis; 85.06, 34.40, 25.49 for MRI-to-CT brain; and 84.23, 34.29, 27.99 for MRI-to-CT pelvis. HealthGPT-L14 scored 79.73, 33.10, 32.96; 71.92, 31.87, 43.09; 85.31, 34.29, 26.20; and 84.96, 34.14, 28.13 on those respective tasks. The compared baselines trained separate models for each conversion task, whereas HealthGPT learned the four tasks in one training process. HealthGPT scores exceeded the strongest listed baseline on many task-metric combinations, including CT-to-MRI pelvis MSE and MRI-to-CT pelvis SSIM, but BBDM's MRI-to-CT brain SSIM of 86.40 exceeded both HealthGPT versions.

  8. Knowl 8 — HealthGPT improves the reported 4× MRI super-resolution metrics

    empirical result

    On 4× super-resolution using the IXI dataset, HealthGPT-M3 achieved SSIM 78.19, PSNR 32.76, MSE 34.47, and LPIPS 12.02. HealthGPT-L14 achieved 77.94, 32.71, 35.19, and 12.43, respectively. The paper compares these results with SRGAN, DASR, Real-ESRGAN, LIIF, and BSRGAN; among the listed methods, the best baseline values were SSIM 73.27 (LIIF), PSNR 32.34 (DASR), MSE 38.25 (DASR), and LPIPS 19.17 (DASR). Thus HealthGPT-M3 had the strongest reported value on all four metrics, with lower MSE and LPIPS indicating reduced pixel error and perceptual distance.

  9. Knowl 9 — Medical-image reconstruction results are strong but not uniformly better than Unified-IO 2

    empirical result

    For image reconstruction, HealthGPT-M3 was evaluated on CT brain, CT pelvis, MRI brain, and MRI pelvis images using SSIM (higher is better), PSNR (higher is better), and MSE (lower is better). Its scores in that metric order were 91.73, 36.42, 15.46 for CT brain; 94.26, 37.30, 12.53 for CT pelvis; 88.76, 33.97, 27.05 for MRI brain; and 84.40, 33.11, 32.62 for MRI pelvis. HealthGPT-M3 exceeded Unified-IO 2 on all three metrics for CT brain and CT pelvis. On MRI brain it had higher SSIM but lower PSNR and higher MSE; on MRI pelvis it scored below Unified-IO 2 on all three metrics. Its scores were substantially higher than SEED-X on all four image groups and all three metrics. The results therefore support reconstruction capability while showing that performance relative to Unified-IO 2 depends on the image modality and region.

  10. Knowl 10 — Three-stage training mitigates mixed-task degradation on most reported measures

    empirical result

    The paper compared its H-LoRA-based three-stage strategy with direct mixed training. For comprehension, the three-stage strategy versus mixed training scored 72.5/55.2 versus 56.6/37.9 on VQA-RAD close/all, 77.9/59.6 versus 45.0/32.9 on SLAKE close/all, and 79.7/49.0 versus 65.7/33.6 on PathVQA close/all. MMMU-Med was the exception: three-stage training scored 42.7, below mixed training's 44.0. OMVQA scores were 68.5 versus 48.9. The reported generation scores for CT brain, CT pelvis, MRI brain, and MRI pelvis were 70.84, 72.99, 65.26, and 61.33 with three-stage training, compared with 65.64, 62.75, 56.61, and 50.77 with mixed training. The comparison shows broad improvements from task-decoupled learning, but not a gain on every comprehension benchmark.

Coverage note — The exploratory report-to-CXR study, clinician response-ranking evaluation, and qualitative case studies are omitted because they are supporting demonstrations rather than central method or benchmark results.

References

  1. 1.Abdin, M., Aneja, J., Awadalla, H., Awadallah, A., Awan, A. A., Bach, N., Bahree, A., Bakhtiari, A., Bao, J., Behl, H., et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  2. 2.Bae, S., Kyung, D., Ryu, J., Cho, E., Lee, G., Kweon, S., Oh, J., JI, L., Chang, E., Kim, T., et al. Mimic-ext-mimic-cxr-vqa: A complex, diverse, and large-scale visual question answering dataset for chest x-ray images, 2024.
  3. 3.Bansal, H., Israel, D., Zhao, S., Li, S., Nguyen, T., and Grover, A. Medmax: Mixed-modal instruction tuning for training biomedical assistants, 2025. URL https://arxiv.org/abs/2412.12661.
  4. 4.Chen, J., Gui, C., Ouyang, R., Gao, A., Chen, S., Chen, G. H., Wang, X., Cai, Z., Ji, K., Wan, X., and Wang, B. Towards injecting medical visual knowledge into multimodal llms at scale. In Conference on Empirical Methods in Natural Language Processing, 2024a.
  5. 5.Chen, X., Wu, Z., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., and Ruan, C. Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811, 2025.
  6. 6.Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences, 67(12):220101, 2024b.
  7. 7.Chern, E., Su, J., Ma, Y., and Liu, P. Anole: An open, autoregressive, native large multimodal models for interleaved image-text generation. arXiv preprint arXiv:2407.06135, 2024.
  8. 8.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. C. H. Instructblip: Towards general-purpose vision-language models with instruction tuning. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023.
  9. 9.Davies, R. L., Royston, P. A., Leung, M. S., Haider, M. E. A. M. J., Barkhof, S. G. A. L., and B., P. E. T. M. The ixi dataset, 2014. URL https://brain-development.org/ixi-dataset/. Accessed: 2025-01-30.
  10. 10.Ding, N., Qin, Y., Yang, G., Wei, F., Yang, Z., Su, Y., Hu, S., Chen, Y., Chan, C.-M., Chen, W., et al. Parameter-efficient fine-tuning of large-scale pre-trained language models. Nature Machine Intelligence, 5(3):220–235, 2023.
  11. 11.Dong, R., Han, C., Peng, Y., Qi, Z., Ge, Z., Yang, J., Zhao, L., Sun, J., Zhou, H., Wei, H., Kong, X., Zhang, X., Ma, K., and Yi, L. Dreamllm: Synergistic multimodal comprehension and creation. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=y01KGvd9Bw.
  12. 12.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  13. 13.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 12873–12883, 2021.
  14. 14.Ge, Y., Ge, Y., Zeng, Z., Wang, X., and Shan, Y. Planting a seed of vision in large language model. arXiv preprint arXiv:2307.08041, 2023.
  15. 15.Ge, Y., Zhao, S., Zhu, J., Ge, Y., Yi, K., Song, L., Li, C., Ding, X., and Shan, Y. Seed-x: Multimodal models with unified multi-granularity comprehension and generation. arXiv preprint arXiv:2404.14396, 2024.
  16. 16.He, X., Zhang, Y., Mou, L., Xing, E., and Xie, P. Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286, 2020.
  17. 17.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3, 2022.
  18. 18.Hu, Y., Li, T., Lu, Q., Shao, W., He, J., Qiao, Y., and Luo, P. Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22170–22183, 2024.
  19. 19.Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation. Advances in Neural Information Processing Systems, 36:78723–78747, 2023.
  20. 20.Huang, Z., He, W., Long, Q., Wang, Y., Li, H., Yu, Z., Shu, F., Chan, L., Jiang, H., Gan, L., et al. T2i-factualbench: Benchmarking the factuality of text-to-image models with knowledge-intensive concepts. arXiv preprint arXiv:2412.04300, 2024.
  21. 21.Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1125–1134, 2017.
  22. 22.Johnson, A. E., Pollard, T. J., Greenbaum, N. R., Lungren, M. P., Deng, C.-y., Peng, Y., Lu, Z., Mark, R. G., Berkowitz, S. J., and Horng, S. Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs. arXiv preprint arXiv:1901.07042, 2019.
  23. 23.Lau, J. J., Gayen, S., Ben Abacha, A., and Demner-Fushman, D. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5(1):1–10, 2018.
  24. 24.Lee, T., Yasunaga, M., Meng, C., Mai, Y., Park, J. S., Gupta, A., Zhang, Y., Narayanan, D., Teufel, H., Bellagente, M., et al. Holistic evaluation of text-to-image models. Advances in Neural Information Processing Systems, 36:69981–70011, 2023.
  25. 25.Li, B., Xue, K., Liu, B., and Lai, Y.-K. Bbdm: Image-to-image translation with brownian bridge diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern Recognition, pp. 1952–1961, 2023a.
  26. 26.Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., and Gao, J. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36, 2024a.
  27. 27.Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., and Gao, J. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36, 2024b.
  28. 28.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp. 19730–19742. PMLR, 2023b.
  29. 29.Li, S., Lin, T., Lin, L., Zhang, W., Liu, J., Yang, X., Li, J., He, Y., Song, X., Xiao, J., et al. Eyecaregpt: Boosting comprehensive ophthalmology understanding with tailored dataset, benchmark and model. arXiv preprint arXiv:2504.13650, 2025.
  30. 30.Lin, T., Liu, J., Zhang, W., Li, Z., Dai, Y., Li, H., Yu, Z., He, W., Li, J., Jiang, H., et al. Teamlora: Boosting low-rank adaptation with expert collaboration and competition. arXiv preprint arXiv:2408.09856, 2024.
  31. 31.Liu, B., Zhan, L.-M., Xu, L., Ma, L., Yang, Y., and Wu, X.-M. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pp. 1650–1654. IEEE, 2021.
  32. 32.Liu, D., Zhao, S., Zhuo, L., Lin, W., Qiao, Y., Li, H., and Gao, P. Lumina-mgpt: Illuminate flexible photorealistic text-to-image generation with multimodal generative pretraining. arXiv preprint arXiv:2408.02657, 2024a.
  33. 33.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. In NeurIPS, 2023.
  34. 34.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26296–26306, 2024b.
  35. 35.Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J. Llava-next: Improved reasoning, ocr, and world knowledge. https://llava-vl.github.io/blog/2024-01-30-llava-next/, 2024c.
  36. 36.Liu, Q., Wu, X., Zhao, X., Zhu, Y., Xu, D., Tian, F., and Zheng, Y. When moe meets llms: Parameter efficient fine-tuning for multi-task medical applications. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1104–1114, 2024d.
  37. 37.Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., Jiao, J., and Liu, Y. Vmamba: Visual state space model. Advances in neural information processing systems, 37:103031–103063, 2024e.
  38. 38.Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A. Unified-io: A unified model for vision, language, and multi-modal tasks. In The Eleventh International Conference on Learning Representations, 2022.
  39. 39.Lu, J., Clark, C., Lee, S., Zhang, Z., Khosla, S., Marten, R., Hoiem, D., and Kembhavi, A. Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 26439–26455, 2024.
  40. 40.Luo, T., Lei, J., Lei, F., Liu, W., He, S., Zhao, J., and Liu, K. Moelora: Contrastive learning guided mixture of experts on parameter-efficient fine-tuning for large language models. arXiv preprint arXiv:2402.12851, 2024a.
  41. 41.Luo, Y., Zhang, J., Fan, S., Yang, K., Hong, M., Wu, Y., Qiao, M., and Nie, Z. Biomedgpt: An open multimodal large language model for biomedicine. IEEE Journal of Biomedical and Health Informatics, 2024b.
  42. 42.Masoudnia, S. and Ebrahimpour, R. Mixture of experts: a literature survey. Artificial Intelligence Review, 42:275–293, 2014.
  43. 43.Moor, M., Huang, Q., Wu, S., Yasunaga, M., Dalmia, Y., Leskovec, J., Zakka, C., Reis, E. P., and Rajpurkar, P. Med-flamingo: a multimodal medical few-shot learner. In Machine Learning for Health (ML4H), pp. 353–367. PMLR, 2023.
  44. 44.Nath, V., Li, W., Yang, D., Myronenko, A., Zheng, M., Lu, Y., Liu, Z., Yin, H., Law, Y. M., Tang, Y., et al. Vila-m3: Enhancing vision-language models with medical expert knowledge. arXiv preprint arXiv:2411.12915, 2024.
  45. 45.Ong, W., Zhu, L., Zhang, W., Kuah, T., Lim, D. S. W., Low, X. Z., Thian, Y. L., Teo, E. C., Tan, J. H., Kumar, N., et al. Application of artificial intelligence methods for imaging of spinal metastasis. Cancers, 14(16):4025, 2022.
  46. 46.OpenAI. Gpt-4v(ision) system card. https://cdn.openai.com/papers/GPTV_System_Card.pdf, 2023.
  47. 47.Pan, K., Tang, S., Li, J., Fan, Z., Chow, W., Yan, S., Chua, T.-S., Zhuang, Y., and Zhang, H. Auto-encoding morph-tokens for multimodal llm. In Proceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024.
  48. 48.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  49. 49.Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al. Large language models encode clinical knowledge. Nature, 620(7972):172–180, 2023.
  50. 50.Team, C. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024a.
  51. 51.Team, Q. Qwen2.5: A party of foundation models, September 2024b. URL https://qwenlm.github.io/blog/qwen2.5/.
  52. 52.Thawakar, O., Shaker, A. M., Mullappilly, S. S., Cholakkal, H., Anwer, R. M., Khan, S. S., Laaksonen, J., and Khan, F. S. Xraygpt: Chest radiographs summarization using large medical vision-language models. In Workshop on Biomedical Natural Language Processing, 2023.
  53. 53.Thummerer, A., van der Bijl, E., Galapon Jr, A., Verhoeff, J. J., Langendijk, J. A., Both, S., van den Berg, C. N. A., and Maspero, M. Synthrad2023 grand challenge dataset: Generating synthetic ct for radiotherapy. Medical physics, 50(7):4664–4674, 2023.
  54. 54.Tian, D., Jiang, S., Zhang, L., Lu, X., and Xu, Y. The role of large language models in medical image processing: a narrative review. Quantitative Imaging in Medicine and Surgery, 14(1):1108, 2023.
  55. 55.Tong, S., Fan, D., Zhu, J., Xiong, Y., Chen, X., Sinha, K., Rabbat, M., LeCun, Y., Xie, S., and Liu, Z. Metamorph: Multimodal understanding and generation via instruction tuning. arXiv preprint arXiv:2412.14164, 2024.
  56. 56.Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P.-C., Carroll, A., Lau, C., Tanno, R., Ktena, I., et al. Towards generalist biomedical ai. NEJM AI, 1(3):AIoa2300138, 2024.
  57. 57.Vig, J. A multiscale visualization of attention in the transformer model. arXiv preprint arXiv:1906.05714, 2019.
  58. 58.Wang, X., Zhang, X., Luo, Z., Sun, Q., Cui, Y., Wang, J., Zhang, F., Wang, Y., Li, Z., Yu, Q., et al. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024a.
  59. 59.Wang, Z., Wu, Z., Agarwal, D., and Sun, J. Medclip: Contrastive learning from unpaired medical images and text. Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing, 2022:3876–3887, 2022.
  60. 60.Wang, Z., Zhang, L., Wang, L., and Zhang, Z. Soft masked mamba diffusion model for ct to mri conversion. arXiv preprint arXiv:2406.15910, 2024b.
  61. 61.Wu, C., Zhang, X., Zhang, Y., Wang, Y., and Xie, W. Towards generalist foundation model for radiology by leveraging web-scale 2d3d medical data, 2023. URL https://arxiv.org/abs/2308.02463.
  62. 62.Wu, C., Chen, X., Wu, Z., Ma, Y., Liu, X., Pan, Z., Liu, W., Xie, Z., Yu, X., Ruan, C., and Luo, P. Janus: Decoupling visual encoding for unified multimodal understanding and generation, 2024a. URL https://arxiv.org/abs/2410.13848.
  63. 63.Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S. Next-gpt: Any-to-any multimodal llm. In Forty-first International Conference on Machine Learning, 2024b.
  64. 64.Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z. Show-o: One single transformer to unify multimodal understanding and generation. arXiv preprint arXiv:2408.12528, 2024.
  65. 65.Xie, Y., Zhou, C., Gao, L., Wu, J., Li, X., Zhou, H., Liu, S., Xing, L., Zou, J., Xie, C., and Zhou, Y. Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=IwgmgidYPS.
  66. 66.Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., et al. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652, 2024.
  67. 67.Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of CVPR, 2024.
  68. 68.Zhang, W., Zhu, L., Hallinan, J., Zhang, S., Makmur, A., Cai, Q., and Ooi, B. C. Boostmis: Boosting medical image semi-supervised learning with adaptive pseudo labeling and informative active annotation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 20666–20676, 2022.
  69. 69.Zhang, W., Lin, T., Liu, J., Shu, F., Li, H., Zhang, L., Wangg u i, H., Zhou, H., Lv, Z., Jiang, H., et al. Hyperllava: Dynamic visual and language expert tuning for multimodal large language models. arXiv preprint arXiv:2403.13447, 2024a.
  70. 70.Zhang, W., Lv, Z., Zhou, H., Liu, J.-W., Li, J., Li, M., Li, Y., Zhang, D., Zhuang, Y., and Tang, S. Revisiting the domain shift and sample uncertainty in multi-source active domain transfer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16751–16761, 2024b.
  71. 71.Zhou, H., Liu, F., Gu, B., Zou, X., Huang, J., Wu, J., Li, Y., Chen, S. S., Zhou, P., Liu, J., et al. A survey of large language models in medicine: Progress, application, and challenge. arXiv preprint arXiv:2311.05112, 2023.
  72. 72.Zhu, J.-Y., Park, T., Isola, P., and Efros, A. A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, pp. 2223–2232, 2017.

Citation

MLA
Lin, T., et al. “HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation”. arXiv, 2025, https://doi.org/10.48550/arxiv.2502.09838.
APA
Lin, T., Zhang, W., Li, S., Yuan, Y., Yu, B., Li, H., He, W., Jiang, H., Li, M., Song, X., Tang, S., Xiao, J., Lin, H., Zhuang, Y., & Ooi, B. C. (2025). HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation. arXiv. https://doi.org/10.48550/arxiv.2502.09838
Chicago
Lin, T., W. Zhang, S. Li, et al. 2025. “HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2502.09838.
Harvard
Lin, T. et al. (2025) “HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation”. arXiv. Available at: https://doi.org/10.48550/arxiv.2502.09838.
Vancouver
1. Lin T, Zhang W, Li S, et al (2025) HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation. https://doi.org/10.48550/arxiv.2502.09838

BibTeX

@misc{https://doi.org/10.48550/arxiv.2502.09838,
  doi = {10.48550/ARXIV.2502.09838},
  url = {https://arxiv.org/abs/2502.09838},
  author = {Lin, Tianwei and Zhang, Wenqiao and Li, Sijing and Yuan, Yuqian and Yu, Binhe and Li, Haoyuan and He, Wanggui and Jiang, Hao and Li, Mengze and Song, Xiaohui and Tang, Siliang and Xiao, Jun and Lin, Hui and Zhuang, Yueting and Ooi, Beng Chin},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation},
  publisher = {arXiv},
  year = {2025},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/