HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
Tianwei LinWenqiao ZhangSijing LiYuqian YuanBinhe YuHaoyuan LiWanggui HeHao JiangMengze LiXiaohui Song
Presents HealthGPT, a medical large vision-language model that unifies multimodal comprehension and image generation tasks within a single autoregressive framework using a heterogeneous low-rank adaptation technique to prevent task interference.
Artificial intelligence applications in healthcare have made notable strides using vision-language foundation models, yet current systems typically remain specialized for single functions. Most medical models focus strictly on visual comprehension, such as answering clinical questions or drafting textual reports, and lack the generative capabilities required to create or enhance medical imagery. Meanwhile, general-domain unified models that attempt both tasks often struggle in healthcare settings due to the scarcity of high-quality medical data and the inherent conflict between abstract comprehension tasks and detail-heavy generation tasks.
The article demonstrates and evaluates HealthGPT, a medical large vision-language model designed to unify clinical visual understanding and image generation within a single autoregressive framework. The system aims to bridge the gap between diagnostic reasoning and medical image synthesis using parameter-efficient fine-tuning on limited domain-specific data.
The researchers developed a three-part technical framework evaluated across multiple open medical benchmarks. First, they introduced a parameter-efficient fine-tuning technique called Heterogeneous Low-Rank Adaptation, which uses dynamic expert routing and matrix block merging to separate comprehension and generation knowledge into distinct plugins while keeping base models frozen. Second, a hierarchical visual perception mechanism routes fine-grained visual details from early neural network layers to generation tasks and abstract features from deeper layers to comprehension tasks. Third, the authors introduced a three-stage training strategy combined with a curated dataset named VL-Health, which contains over 1.5 million instruction-following samples across 11 imaging modalities, such as computed tomography, magnetic resonance imaging, and X-rays.
Evaluation results show that HealthGPT consistently outperforms both medical-specific models and general-purpose multimodal models across several critical benchmarks. In medical visual question answering across seven modalities on the OmniMedVQA benchmark, the compact 3.8-billion parameter HealthGPT-M3 scored 68.5, substantially higher than the specialized medical model HuatuoGPT-Vision (50.0) and the open-world model Llama-3.2 (63.2). Larger variants scaled this performance further, with the 32-billion parameter model achieving a benchmark score of 77.2. For generation tasks, HealthGPT surpassed dedicated standalone tools in cross-modality translation—such as converting computed tomography scans to magnetic resonance imaging—and achieved state-of-the-art detail restoration in four-times image super-resolution. Furthermore, its custom adaptation structure operated without the linear training slowdowns found in standard mixture-of-experts approaches, running approximately 33% faster with four experts.
These findings suggest that healthcare organizations can deploy a single consolidated artificial intelligence model rather than maintaining isolated systems for diagnosis, reporting, and image processing. Unifying these tasks within a shared model reduces computational training costs, lowers infrastructure maintenance demands, and minimizes the risk of catastrophic forgetting when handling diverse clinical imaging tasks. Blinded evaluations by five practicing clinicians further validated these results, selecting HealthGPT's responses as superior more often than competing models in open-ended diagnostic reasoning.
Decision-makers and research teams should consider adopting decoupled adaptation architectures when deploying multimodal systems in data-constrained enterprise environments. While HealthGPT establishes a strong baseline for unified medical models, prospective adopters should conduct domain-specific clinical validation pilots before deploying it into active diagnostic pipelines, particularly for critical imaging conversions. Future work should focus on expanding the model's text-to-image reporting capabilities and exploring broader generative capabilities across larger model sizes.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o establishes the general single-transformer approach to combining multimodal understanding and image generation that HealthGPT adapts to medical data and tasks.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). LLaVA-Med provides a medical vision-language assistant baseline for understanding how HealthGPT extends domain-specific multimodal comprehension toward unified image generation.
- Paper: Zero-Shot Text-to-Image Generation, Aditya Ramesh et al. (2021). The autoregressive text-to-image design in this chapter clarifies the generation paradigm HealthGPT incorporates alongside medical visual comprehension.
No sufficiently relevant recommendations were found.
