Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning
Fuying WangYuyin ZhouShujun WangVarut VardhanabhutiLequan Yu
Proposes a cross-modal alignment framework that simultaneously aligns medical images and radiology reports across pathological region, instance, and disease levels to improve downstream classification, detection, and segmentation tasks under low-data regimes.
Training deep learning systems for medical imaging typically demands extensive manual annotation by clinical experts, which is costly, slow, and hard to scale. While recent approaches attempt to learn representations directly from paired medical images and unstructured radiology reports, existing methods only align information at a single resolution—either globally across entire images or locally across isolated patches. This incomplete alignment discards vital cross-modal relationships and limits how well artificial intelligence models transfer to diverse clinical workflows.
The article introduces and evaluates the Multi-Granularity Cross-modal Alignment framework, an automated approach designed to learn generalizable medical visual representations from paired radiology reports. The system explicitly captures natural correspondences across three complementary semantic tiers: instance-level (matching full images to whole reports), token-level (aligning localized image regions to specific descriptive phrases via bidirectional cross-attention), and disease-level (grouping high-level clinical prototypes using cross-modal clustering).
The framework was pre-trained on approximately 217,000 frontal chest radiograph and report pairs from the MIMIC-CXR database. Evaluated across seven downstream medical benchmark datasets covering classification, localized object detection, and pixel-level semantic segmentation, the model demonstrated substantial data efficiency and diagnostic accuracy. Across three diagnostic classification benchmarks, the framework achieved top performance, showing an accuracy improvement of up to 8.3 percentage points on COVID-19 detection when using only 1% of labeled training data. In localized object detection tasks, it outperformed leading methods across all data sampling levels, delivering a 12.9% mean average precision on the RSNA Pneumonia benchmark with 1% training data. For semantic segmentation, the approach improved Dice overlap scores by up to 12.3 percentage points over previous multimodal models under constrained 1% annotation settings.
These findings indicate that integrating multi-level cross-modal alignment allows healthcare machine learning models to reach high diagnostic accuracy with minimal human annotation. By significantly reducing the volume of manually labeled cases required to train robust diagnostic tools, organizations can accelerate clinical deployment timelines, lower engineering costs, and reduce specialist annotation burdens. Direct comparisons also demonstrated that generic vision-language models pre-trained on natural images transfer poorly to medical tasks, underscoring the operational necessity of domain-specific medical pre-training.
Healthcare technology leaders and clinical deployment teams should prioritize multi-tier pre-training frameworks over single-level architectures when developing diagnostic vision tools. Organizations should pilot this framework on target clinical workflows where labeled data is scarce, while maintaining standard compliance reviews to audit source datasets for patient privacy and demographic biases before practical deployment. While the framework demonstrates robust statistical stability across repeated trials, its evaluation remains focused on chest X-rays without assessing retrieval tasks. Future research should expand the framework toward unified generative models and broader imaging modalities.
- Paper: CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison, Jeremy Irvin et al. (2019). It introduces standard radiology NLP labeling methodologies and chest radiograph benchmarks that motivate paired image-report extraction and diagnostic baseline evaluations.
- Paper: ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases, Xiaosong Wang et al. (2017). It pioneered large-scale text mining of radiology reports to supervise multi-label thoracic disease classification and weakly-supervised localization.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It defines the core contrastive representation learning principles across complementary views that multi-modal medical alignment architectures adapt.
- Paper: Transfusion: Understanding Transfer Learning for Medical Imaging, Maithra Raghu et al. (2019). It provides critical empirical analysis showing why natural-image transfer learning fails to generalize effectively to clinical radiology tasks.
- Paper: Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports Generation, Bin Yan et al. (2022). It establishes domain-specific vision-language pre-training over MIMIC-CXR reports to connect clinical terminology directly with radiograph features.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). It extends biomedical vision-language representation learning into conversational, instruction-tuned multimodal clinical assistants.
- Paper: Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale, Junying Chen et al. (2024). It advances medical vision-language modeling by scaling up cross-modal image-text instruction datasets from biomedical literature to power multimodal LLMs.
- Paper: EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images, Seongsu Bae et al. (2023). It applies chest radiograph vision-language reasoning to complex, joint question-answering over multimodal electronic health records.
- Paper: Best of Both Worlds: Multimodal Contrastive Learning with Tabular and Imaging Data, Paul Hager et al. (2023). It broadens multimodal contrastive pre-training from image-text pairs to the integration of medical imaging and structured clinical tabular data.
- Paper: ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification, Jiangbo Shi et al. (2024). It generalizes multi-scale vision-language alignment concepts to gigapixel whole slide images in computational pathology.
- Paper: SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger, Yuting Gao et al. (2024). It refines fine-grained cross-modal contrastive alignment techniques by softening one-to-one pairing constraints to better capture overlapping semantics.
- Paper: MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding, Yuxin Zuo et al. (2025). It provides a comprehensive expert-level benchmark to evaluate advanced multimodal clinical reasoning beyond standard classification and segmentation.
