ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification
Jiangbo ShiChen LiTieliang GongYefeng ZhengHuazhu Fu
Proposes a dual-scale vision-language multiple instance learning framework that uses frozen LLM prompts and lightweight decoders to transfer vision-language models to gigapixel whole slide image classification with minimal labeled data.
Pathological review of whole slide images serves as the clinical standard for cancer diagnosis and subtyping. However, digitizing these giga-pixel images creates immense data volumes that make pixel-level annotation practically unattainable. Traditional automated methods rely strictly on weakly supervised multiple instance learning, which requires substantial quantities of labeled training slides that are often unavailable for rare diseases. While adapting pre-trained vision-language models can incorporate useful diagnostic language priors, existing approaches depend heavily on generic class labels or expensive, resource-intensive collection of millions of domain-specific image-text pairs.
The article introduces and evaluates ViLa-MIL, a dual-scale vision-language multiple instance learning framework designed to efficiently classify whole slide images with limited labeled data. The system aims to bridge diagnostic pathology knowledge and computational vision by generating descriptive text prompts across image scales and efficiently aligning them with visual features.
To achieve this, the authors developed a multi-scale prompting pipeline leveraging a frozen large language model to describe diagnostic criteria at low resolution, representing global tissue architecture, and high resolution, representing cellular details. The approach incorporates two lightweight neural decoders: a prototype-guided patch decoder that groups similar image patches to summarize slide-level visual context, and a context-guided text decoder that refines text features using local and global visual cues. The framework was evaluated under a strict few-shot regime using 16 training slides per category across three multi-center cancer datasets covering renal cell carcinoma and lung cancer.
The experimental findings show that ViLa-MIL consistently outperforms leading multiple instance learning baselines under few-shot conditions. Specifically, the framework improved the area under the receiver operating characteristic curve by 1.7% to 7.2%, the F1 score by 2.1% to 7.3%, and overall classification accuracy across all benchmark datasets. In domain-shift tests between different hospital centers, ViLa-MIL maintained superior robustness, exceeding baseline cross-center diagnostic accuracy by 5.5% in area under the curve. Visual analyses confirmed that the method generates well-separated class clusters and accurately localizes tumor regions within slides.
These findings demonstrate that integrating structured, scale-appropriate text descriptions eliminates the need for massive pre-training on paired medical datasets, significantly reducing computational overhead and annotation costs. The framework enables high diagnostic accuracy even when data is scarce, providing a scalable pathway for developing automated diagnostic tools for rare cancers and heterogeneous clinical settings.
Healthcare technology leaders and research teams should consider adopting multi-scale vision-language prompting to build cost-effective diagnostic support pipelines. Future efforts should focus on deploying the framework across broader tumor types and exploring more capable language models to further enhance diagnostic descriptions. While current tests were restricted to renal and lung cancer datasets under controlled 16-shot conditions, the consistent gains across independent evaluation runs provide high confidence in the method's effectiveness for low-resource computational pathology.
- Paper: Data-efficient and weakly supervised computational pathology on whole-slide images, Ming Y. Lu et al. (2020). Introduces the clustering-constrained attention multiple instance learning framework for whole slide images that serves as a core methodological foundation and baseline for ViLa-MIL.
- Paper: Attention-based Deep Multiple Instance Learning, Maximilian Ilse et al. (2018). Establishes the foundational attention-based deep multiple instance learning paradigm used for aggregating patch embeddings into slide-level representations.
- Paper: LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day, Chunyuan Li et al. (2023). Pioneers the adaptation of vision-language foundation models to biomedical domains using curriculum instruction tuning.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). Provides a comprehensive taxonomy of vision-language architectures and prompt-tuning strategies leveraged for downstream visual tasks.
- Paper: DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations, Ximeng Sun et al. (2022). Develops dual context optimization for adapting frozen vision-language models with limited annotations, informing prompt-based classification in few-shot regimes.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). Addresses aligning local visual image patches with text representations in contrastive vision-language learning.
- Paper: Benchmarking Self-Supervised Learning on Diverse Pathology Datasets, Mingu Kang et al. (2023). Systematically benchmarks self-supervised feature extraction backbones on diverse histopathology datasets for multiple instance learning pipelines.
- Paper: VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge, Vishwesh Nath et al. (2025). Extends the integration of medical domain knowledge into vision-language architectures by incorporating specialized expert clinical guidance across broader multimodal diagnostic tasks.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Advances multi-scale visual reasoning by formulating a chain-of-thought framework that iteratively focuses on localized high-resolution regions for complex image interpretation.
- Paper: Few-Shot Object Detection with Foundation Models, Guangxing Han et al. (2024). Generalizes the few-shot foundation model adaptation strategy using frozen visual backbones and in-context language prompts to object detection.
