ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification

Jiangbo ShiChen LiTieliang GongYefeng ZhengHuazhu Fu

article2024CVPR52 citations

Proposes a dual-scale vision-language multiple instance learning framework that uses frozen LLM prompts and lightweight decoders to transfer vision-language models to gigapixel whole slide image classification with minimal labeled data.

Listen

Pathological review of whole slide images serves as the clinical standard for cancer diagnosis and subtyping. However, digitizing these giga-pixel images creates immense data volumes that make pixel-level annotation practically unattainable. Traditional automated methods rely strictly on weakly supervised multiple instance learning, which requires substantial quantities of labeled training slides that are often unavailable for rare diseases. While adapting pre-trained vision-language models can incorporate useful diagnostic language priors, existing approaches depend heavily on generic class labels or expensive, resource-intensive collection of millions of domain-specific image-text pairs.

The article introduces and evaluates ViLa-MIL, a dual-scale vision-language multiple instance learning framework designed to efficiently classify whole slide images with limited labeled data. The system aims to bridge diagnostic pathology knowledge and computational vision by generating descriptive text prompts across image scales and efficiently aligning them with visual features.

To achieve this, the authors developed a multi-scale prompting pipeline leveraging a frozen large language model to describe diagnostic criteria at low resolution, representing global tissue architecture, and high resolution, representing cellular details. The approach incorporates two lightweight neural decoders: a prototype-guided patch decoder that groups similar image patches to summarize slide-level visual context, and a context-guided text decoder that refines text features using local and global visual cues. The framework was evaluated under a strict few-shot regime using 16 training slides per category across three multi-center cancer datasets covering renal cell carcinoma and lung cancer.

The experimental findings show that ViLa-MIL consistently outperforms leading multiple instance learning baselines under few-shot conditions. Specifically, the framework improved the area under the receiver operating characteristic curve by 1.7% to 7.2%, the F1 score by 2.1% to 7.3%, and overall classification accuracy across all benchmark datasets. In domain-shift tests between different hospital centers, ViLa-MIL maintained superior robustness, exceeding baseline cross-center diagnostic accuracy by 5.5% in area under the curve. Visual analyses confirmed that the method generates well-separated class clusters and accurately localizes tumor regions within slides.

These findings demonstrate that integrating structured, scale-appropriate text descriptions eliminates the need for massive pre-training on paired medical datasets, significantly reducing computational overhead and annotation costs. The framework enables high diagnostic accuracy even when data is scarce, providing a scalable pathway for developing automated diagnostic tools for rare cancers and heterogeneous clinical settings.

Healthcare technology leaders and research teams should consider adopting multi-scale vision-language prompting to build cost-effective diagnostic support pipelines. Future efforts should focus on deploying the framework across broader tumor types and exploring more capable language models to further enhance diagnostic descriptions. While current tests were restricted to renal and lung cancer datasets under controlled 16-shot conditions, the consistent gains across independent evaluation runs provide high confidence in the method's effectiveness for low-resource computational pathology.

Cover for ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification

Abstract

Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a substantial number of bag-level labels and solely learn from the original slides, which are easily affected by variations in data distribution. Recently, vision language model (VLM)-based methods introduced the language prior by pre-training on large-scale pathological image-text pairs. However, the previous text prompt lacks the consideration of pathological prior knowledge, therefore does not substantially boost the model’s performance. Moreover, the collection of such pairs and the pre-training process are very time-consuming and source-intensive. To solve the above problems, we propose a dual-scale vision-language multiple instance learning (ViLa-MIL) framework for whole slide image classification. Specifically, we propose a dual-scale visual descriptive text prompt based on the frozen large language model (LLM) to boost the performance of VLM effectively. To transfer the VLM to process WSI efficiently, for the image branch, we propose a prototype-guided patch decoder to aggregate the patch features progressively by grouping similar patches into the same prototype; for the text branch, we introduce a context-guided text decoder to enhance the text features by incorporating the multi-granular image contexts. Extensive studies on three multi-cancer and multi-center subtyping datasets demonstrate the superiority of ViLa-MIL.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. Revisiting CLIP
  • 3.2. Dual-scale Visual Descriptive Text Prompt
  • 3.3. Prototype-guided Patch Decoder
  • 3.4. Context-guided Text Decoder
  • 3.5. Training Strategy
  • 4. Experiments
  • 4.1. Settings
  • 4.2. Comparisons with State-of-the-Art
  • 4.3. Generalization on Domain Shift
  • 4.4. Ablation Studies
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Dual-Scale Vision-Language Multiple Instance Learning (ViLa-MIL) Framework

    model/method

    ViLa-MIL is a weakly supervised vision-language framework designed for whole slide image (WSI) classification under few-shot conditions. It operates across two magnification scales: low resolution (e.g., 5×5\times) capturing macroscopic architectural context (such as papillary or solid growth patterns) and high resolution (e.g., 10×10\times) capturing fine cytological details (such as nuclear pleomorphism and cytoplasm clarity).

    Given a WSI bag W={Wl,Wh}W = \{W_l, W_h\} at low (ll) and high (hh) magnifications, non-overlapping patches I={Il∈RNl×N0×N0×3,Ih∈RNh×N0×N0×3}I = \{I_l \in \mathbb{R}^{N_l \times N_0 \times N_0 \times 3}, I_h \in \mathbb{R}^{N_h \times N_0 \times N_0 \times 3}\} (with patch size N0×N0N_0 \times N_0, e.g., 256×256256 \times 256) are extracted and mapped into visual feature embeddings Hl∈RNl×dH_l \in \mathbb{R}^{N_l \times d} and Hh∈RNh×dH_h \in \mathbb{R}^{N_h \times d} using a frozen vision-language image encoder EI(⋅)E_I(\cdot) (e.g., ResNet-50 CLIP, d=1024d=1024).

    For the language branch, a frozen Large Language Model (LLM, e.g., GPT-3.5) generates scale-specific visual descriptive text prompts for all CC candidate categories, concatenated with MM learnable prefix context vectors. Initial scale-specific text representations Dl,Dh∈RC×dD_l, D_h \in \mathbb{R}^{C \times d} are extracted using the frozen CLIP text encoder ET(⋅)E_T(\cdot).

    To adapt the vision-language model efficiently without fine-tuning the heavy backbones, ViLa-MIL introduces two trainable modules:

    1. A Prototype-Guided Patch Decoder that compresses the NkN_k patch features into slide-level feature embeddings Sk∈RdS_k \in \mathbb{R}^d for scale k∈{l,h}k \in \{l, h\} via cross-attention with learnable prototype vectors followed by attention-weighted pooling.
    2. A Context-Guided Text Decoder that refines category text embeddings DkD_k into Dk′∈RC×dD'_k \in \mathbb{R}^{C \times d} by cross-attending to multi-granular visual context (the union of prototype and patch features).

    Final classification probabilities are computed by combining cosine similarities across both scales, optimized end-to-end using standard cross-entropy loss.

  2. Knowl 2 — Dual-Scale Visual Descriptive Text Prompt Generation

    model/method

    To inject domain-specific pathological prior knowledge into vision-language whole slide image (WSI) classification, ViLa-MIL constructs scale-differentiated visual descriptive prompts using a frozen Large Language Model (LLM). Pathologists evaluate tissue hierarchically by first identifying macroscopic structures at low magnification and then zooming into high magnification for cytological and nuclear details.

    An LLM is queried using the structured prompt template: "What are the visually descriptive characteristics of {class name} at low and high resolution in the whole slide image?"\text{"What are the visually descriptive characteristics of } \{\text{class name}\} \text{ at low and high resolution in the whole slide image?"}

    For a target class (e.g., renal cell carcinoma subtypes such as clear cell CCRCC, papillary PRCC, and chromophobe CRCC), this query yields two complementary visual descriptions:

    • Low-scale descriptive text: focuses on macro tissue architecture (e.g., "papillary growth pattern", "well-circumscribed borders", "solid mass").
    • High-scale descriptive text: focuses on cellular micro-features (e.g., "nuclei arranged in layers", "heterogeneous cytoplasm", "clear cytoplasm and round nuclei").

    To bridge the domain gap between pre-trained natural language representations and pathology, MM learnable continuous prefix vectors [Vk]1,[Vk]2,…,[Vk]M[V_k]_1, [V_k]_2, \dots, [V_k]_M (where each vector has dimension matching word token embeddings) are prepended to the generated descriptions, forming the full dual-scale text prompts TlT_l and ThT_h: Tl=[Vl]1[Vl]2…[Vl]M[Low-scale Text Prompt]T_l = [V_l]_1 [V_l]_2 \dots [V_l]_M [\text{Low-scale Text Prompt}] Th=[Vh]1[Vh]2…[Vh]M[High-scale Text Prompt]T_h = [V_h]_1 [V_h]_2 \dots [V_h]_M [\text{High-scale Text Prompt}]

  3. Knowl 3 — Prototype-Guided Patch Decoder

    model/method

    The prototype-guided patch decoder aggregates an arbitrary number NkN_k of patch embeddings Hk∈RNk×dH_k \in \mathbb{R}^{N_k \times d} (where k∈{l,h}k \in \{l, h\} indexes low and high magnifications) into a fixed slide-level visual embedding Sk∈RdS_k \in \mathbb{R}^d.

    A set of NpN_p learnable prototype vectors Pr∈RNp×dPr \in \mathbb{R}^{N_p \times d} is randomly initialized to guide the aggregation by grouping semantically similar patches into shared prototype clusters. A cross-attention layer treats the prototypes as queries Qk=PrQ_k = Pr and all patch features as keys Kk=HkK_k = H_k and values Vk=HkV_k = H_k: Prk=LayerNorm(Softmax(QkKk⊤d)Vk)+PrPr_k = \mathrm{LayerNorm}\left(\mathrm{Softmax}\left(\frac{Q_k K_k^\top}{\sqrt{d}}\right) V_k\right) + Pr where LayerNorm(⋅)\mathrm{LayerNorm}(\cdot) denotes layer normalization, Softmax(⋅)\mathrm{Softmax}(\cdot) operates along the key dimension, and Prk∈RNp×dPr_k \in \mathbb{R}^{N_p \times d} contains the context-enriched prototype features.

    To obtain the final slide-level representation SkS_k, the updated prototypes are aggregated via parameterized gated attention pooling: Prk,i′=WaPrk,iPr'_{k,i} = W_a Pr_{k,i} Ak,i=exp⁡(Wb⊤tanh⁡(WvPr′k,i⊤))∑j=1Npexp⁡(Wb⊤tanh⁡(WvPr′k,j⊤))A_{k,i} = \frac{\exp\left(W_b^\top \tanh(W_v {Pr'}_{k,i}^\top)\right)}{\sum_{j=1}^{N_p} \exp\left(W_b^\top \tanh(W_v {Pr'}_{k,j}^\top)\right)} Sk=Wc∑i=1NpAk,iPrk,i′S_k = W_c \sum_{i=1}^{N_p} A_{k,i} Pr'_{k,i} where Prk,i∈RdPr_{k,i} \in \mathbb{R}^d is the ii-th prototype vector in PrkPr_k; Wa,Wc,Wv∈Rd×dW_a, W_c, W_v \in \mathbb{R}^{d \times d} are trainable weight matrices; Wb∈Rd×1W_b \in \mathbb{R}^{d \times 1} is a trainable projection vector; tanh⁡(⋅)\tanh(\cdot) is the hyperbolic tangent activation; and Ak,i∈[0,1]A_{k,i} \in [0, 1] is the normalized attention weight for prototype ii with ∑i=1NpAk,i=1\sum_{i=1}^{N_p} A_{k,i} = 1.

  4. Knowl 4 — Context-Guided Text Decoder

    model/method

    The context-guided text decoder refines the text prompt features by conditioning them on multi-granular visual context extracted from the whole slide image, closing the cross-modal representation gap.

    For scale k∈{l,h}k \in \{l, h\}, the initial text features Dk∈RC×dD_k \in \mathbb{R}^{C \times d} are obtained by passing the prompt sequences TkT_k of all CC candidate categories through the frozen CLIP text encoder ET(⋅)E_T(\cdot). Multi-granular visual contextual information is formed by concatenating the global prototype features Prk∈RNp×dPr_k \in \mathbb{R}^{N_p \times d} from the prototype-guided patch decoder and the local patch features Hk∈RNk×dH_k \in \mathbb{R}^{N_k \times d} to build the key and value matrices: Kt,k=[Prk;Hk]∈R(Np+Nk)×dK_{t,k} = [Pr_k; H_k] \in \mathbb{R}^{(N_p + N_k) \times d} Vt,k=[Prk;Hk]∈R(Np+Nk)×dV_{t,k} = [Pr_k; H_k] \in \mathbb{R}^{(N_p + N_k) \times d}

    Using the text features as queries Qt,k=Dk∈RC×dQ_{t,k} = D_k \in \mathbb{R}^{C \times d}, the refined text representations Dk′∈RC×dD'_k \in \mathbb{R}^{C \times d} are computed via cross-attention with a residual connection: Dk′=Softmax(Qt,kKt,k⊤d)Vt,k+DkD'_k = \mathrm{Softmax}\left(\frac{Q_{t,k} K_{t,k}^\top}{\sqrt{d}}\right) V_{t,k} + D_k where Dk′D'_k aligns category textual semantics with the slide-specific morphological characteristics observed at scale kk.

  5. Knowl 5 — Dual-Scale Vision-Language Classification Objective

    equation

    For a whole slide image evaluated at low (ll) and high (hh) magnifications, let Sk∈RdS_k \in \mathbb{R}^d denote the slide-level image embedding and Dk,i′∈RdD'_{k,i} \in \mathbb{R}^d denote the refined text feature embedding for class i∈{1,…,C}i \in \{1, \dots, C\} at scale k∈{l,h}k \in \{l, h\}.

    The predicted class probability PiP_i for class ii is calculated by weighting the scale-specific softmax cosine similarity distributions: Pi=∑k∈{l,h}αkexp⁡(cos⁡(Sk,Dk,i′))∑j=1Cexp⁡(cos⁡(Sk,Dk,j′))P_i = \sum_{k \in \{l, h\}} \alpha_k \frac{\exp\left(\cos(S_k, D'_{k,i})\right)}{\sum_{j=1}^C \exp\left(\cos(S_k, D'_{k,j})\right)} where cos⁡(u,v)=u⊤v∥u∥2∥v∥2\cos(u, v) = \frac{u^\top v}{\|u\|_2 \|v\|_2} denotes the cosine similarity, and αl,αh≥0\alpha_l, \alpha_h \ge 0 are non-negative hyperparameters controlling the relative importance of each magnification scale (set to αl=αh=1\alpha_l = \alpha_h = 1).

    The entire model is trained end-to-end using the slide-level multi-class cross-entropy loss: L=CE(P,GT)=−∑i=1CGTilog⁡Pi\mathcal{L} = \mathrm{CE}(P, GT) = -\sum_{i=1}^C GT_i \log P_i where GT∈{0,1}CGT \in \{0, 1\}^C is the one-hot encoded ground-truth slide-level diagnostic label.

  6. Knowl 6 — Few-Shot Classification Performance on Multi-Cancer WSI Datasets

    data/table

    The performance of ViLa-MIL and baseline multiple instance learning (MIL) methods was evaluated under a 16-shot setting across three WSI subtyping benchmarks: TIHD-RCC (in-house renal cell carcinoma subtyping), TCGA-RCC (TCGA renal cell carcinoma subtyping: CCRCC vs PRCC vs CRCC), and TCGA-Lung (TCGA lung cancer subtyping: LUAD vs LUSC). All models used a 4:3:3 train/validation/test split, with 16 randomly selected training samples per class, averaged over 5 runs using frozen CLIP ResNet-50 patch embeddings (d=1024d=1024).

    Method TIHD-RCC TCGA-RCC TCGA-Lung
    AUC F1 ACC AUC F1 ACC AUC F1 ACC
    Max-pooling 52.5±4.952.5 \pm 4.9 32.2±3.532.2 \pm 3.5 37.5±2.337.5 \pm 2.3 67.4±4.967.4 \pm 4.9 46.7±11.646.7 \pm 11.6 54.1±4.854.1 \pm 4.8 53.0±6.053.0 \pm 6.0 45.8±8.945.8 \pm 8.9 53.3±3.453.3 \pm 3.4
    Mean-pooling 67.2±4.667.2 \pm 4.6 42.4±10.142.4 \pm 10.1 47.9±3.947.9 \pm 3.9 83.3±6.083.3 \pm 6.0 60.9±8.560.9 \pm 8.5 62.3±7.462.3 \pm 7.4 67.4±7.267.4 \pm 7.2 61.1±5.561.1 \pm 5.5 61.9±5.561.9 \pm 5.5
    ABMIL 65.6±9.665.6 \pm 9.6 46.9±7.146.9 \pm 7.1 47.9±6.747.9 \pm 6.7 83.6±3.183.6 \pm 3.1 64.4±4.264.4 \pm 4.2 65.7±4.765.7 \pm 4.7 60.5±15.960.5 \pm 15.9 56.8±11.856.8 \pm 11.8 61.2±6.161.2 \pm 6.1
    CLAM-SB 68.9±14.468.9 \pm 14.4 50.5±16.350.5 \pm 16.3 53.1±13.853.1 \pm 13.8 90.1±2.290.1 \pm 2.2 75.3±7.475.3 \pm 7.4 77.6±7.077.6 \pm 7.0 66.7±13.666.7 \pm 13.6 59.9±13.859.9 \pm 13.8 64.0±7.764.0 \pm 7.7
    CLAM-MB 71.8±9.871.8 \pm 9.8 54.8±9.554.8 \pm 9.5 55.4±10.255.4 \pm 10.2 90.9±4.190.9 \pm 4.1 76.2±4.476.2 \pm 4.4 78.6±4.978.6 \pm 4.9 68.8±12.568.8 \pm 12.5 60.3±11.160.3 \pm 11.1 63.0±9.363.0 \pm 9.3
    TransMIL 71.2±4.271.2 \pm 4.2 53.9±3.953.9 \pm 3.9 54.9±4.154.9 \pm 4.1 89.4±5.689.4 \pm 5.6 73.0±7.873.0 \pm 7.8 75.3±7.275.3 \pm 7.2 64.2±8.564.2 \pm 8.5 57.5±6.457.5 \pm 6.4 59.7±5.459.7 \pm 5.4
    DSMIL 69.9±9.469.9 \pm 9.4 49.9±6.949.9 \pm 6.9 50.0±7.450.0 \pm 7.4 87.6±4.587.6 \pm 4.5 71.5±6.671.5 \pm 6.6 72.8±6.472.8 \pm 6.4 67.9±8.067.9 \pm 8.0 61.0±7.061.0 \pm 7.0 61.3±7.061.3 \pm 7.0
    GTMIL 77.1±3.577.1 \pm 3.5 61.4±4.361.4 \pm 4.3 62.2±4.362.2 \pm 4.3 88.1±13.388.1 \pm 13.3 71.1±15.771.1 \pm 15.7 76.1±12.976.1 \pm 12.9 66.0±15.366.0 \pm 15.3 61.1±12.361.1 \pm 12.3 63.8±9.963.8 \pm 9.9
    DTMIL 70.3±10.370.3 \pm 10.3 55.6±9.655.6 \pm 9.6 55.8±9.955.8 \pm 9.9 90.0±4.690.0 \pm 4.6 74.4±5.374.4 \pm 5.3 76.8±5.276.8 \pm 5.2 67.5±10.367.5 \pm 10.3 57.3±11.357.3 \pm 11.3 66.6±7.566.6 \pm 7.5
    IBMIL 71.5±6.371.5 \pm 6.3 57.2±7.657.2 \pm 7.6 56.4±4.856.4 \pm 4.8 90.5±4.190.5 \pm 4.1 75.1±5.275.1 \pm 5.2 77.2±4.277.2 \pm 4.2 69.2±7.469.2 \pm 7.4 57.4±8.357.4 \pm 8.3 66.9±6.566.9 \pm 6.5
    ViLa-MIL Low 79.5±3.679.5 \pm 3.6 61.2±7.161.2 \pm 7.1 63.1±3.063.1 \pm 3.0 90.9±2.990.9 \pm 2.9 77.3±4.077.3 \pm 4.0 79.5±3.779.5 \pm 3.7 71.9±6.271.9 \pm 6.2 64.1±7.864.1 \pm 7.8 65.8±5.865.8 \pm 5.8
    ViLa-MIL High 83.0±5.683.0 \pm 5.6 65.3±6.465.3 \pm 6.4 66.2±6.866.2 \pm 6.8 92.0±1.392.0 \pm 1.3 77.8±3.777.8 \pm 3.7 80.0±3.080.0 \pm 3.0 72.3±7.672.3 \pm 7.6 66.6±6.566.6 \pm 6.5 66.9±6.366.9 \pm 6.3
    ViLa-MIL 84.3±4.6\mathbf{84.3 \pm 4.6} 68.7±7.3\mathbf{68.7 \pm 7.3} 68.8±7.3\mathbf{68.8 \pm 7.3} 92.6±3.0\mathbf{92.6 \pm 3.0} 78.3±6.9\mathbf{78.3 \pm 6.9} 80.3±6.2\mathbf{80.3 \pm 6.2} 74.7±3.5\mathbf{74.7 \pm 3.5} 67.0±4.9\mathbf{67.0 \pm 4.9} 67.7±4.4\mathbf{67.7 \pm 4.4}

    ViLa-MIL achieves the highest performance across all datasets and metrics, outperforming the best existing MIL baseline by 1.7–7.2% in AUC, 2.1–7.3% in F1 score, and 0.8–6.6% in accuracy. Even the single-scale variants (ViLa-MIL Low and High) outperform purely visual MIL methods.

  7. Knowl 7 — Cross-Center Domain Generalization Between TIHD-RCC and TCGA-RCC

    data/table

    Cross-center evaluation was performed between TIHD-RCC and TCGA-RCC under the 16-shot setting to assess model generalization across different institutions and staining protocols. Models were trained on one dataset and evaluated on the test split of the other.

    Method TIHD →\to TCGA TCGA →\to TIHD
    AUC F1 ACC AUC F1 ACC
    Mean 61.9±9.261.9 \pm 9.2 31.9±7.731.9 \pm 7.7 24.6±3.624.6 \pm 3.6 69.1±5.969.1 \pm 5.9 43.9±7.143.9 \pm 7.1 34.0±6.534.0 \pm 6.5
    ABMIL 61.6±7.061.6 \pm 7.0 26.9±2.726.9 \pm 2.7 24.7±3.324.7 \pm 3.3 66.8±6.166.8 \pm 6.1 43.1±6.943.1 \pm 6.9 34.1±8.634.1 \pm 8.6
    CLAM-MB 77.1±6.977.1 \pm 6.9 36.1±2.036.1 \pm 2.0 35.3±2.635.3 \pm 2.6 68.2±8.768.2 \pm 8.7 46.9±8.246.9 \pm 8.2 35.0±6.635.0 \pm 6.6
    TransMIL 71.1±5.971.1 \pm 5.9 50.3±6.350.3 \pm 6.3 47.3±4.847.3 \pm 4.8 63.2±7.963.2 \pm 7.9 44.6±7.544.6 \pm 7.5 30.9±8.430.9 \pm 8.4
    DSMIL 74.3±7.374.3 \pm 7.3 24.6±4.624.6 \pm 4.6 21.9±6.721.9 \pm 6.7 68.1±6.068.1 \pm 6.0 42.1±6.042.1 \pm 6.0 32.3±6.332.3 \pm 6.3
    GTMIL 76.7±9.276.7 \pm 9.2 46.7±4.846.7 \pm 4.8 43.1±8.643.1 \pm 8.6 72.7±9.772.7 \pm 9.7 51.0±7.451.0 \pm 7.4 37.4±5.737.4 \pm 5.7
    DTMIL 77.9±5.677.9 \pm 5.6 30.0±6.830.0 \pm 6.8 30.9±5.730.9 \pm 5.7 69.8±8.369.8 \pm 8.3 38.0±6.638.0 \pm 6.6 35.4±7.035.4 \pm 7.0
    IBMIL 78.2±4.778.2 \pm 4.7 37.0±6.537.0 \pm 6.5 36.9±4.636.9 \pm 4.6 70.1±7.470.1 \pm 7.4 47.0±6.447.0 \pm 6.4 35.5±7.235.5 \pm 7.2
    ViLa-MIL 83.4±3.3\mathbf{83.4 \pm 3.3} 51.7±5.6\mathbf{51.7 \pm 5.6} 50.3±5.8\mathbf{50.3 \pm 5.8} 75.4±4.9\mathbf{75.4 \pm 4.9} 51.4±4.7\mathbf{51.4 \pm 4.7} 38.3±5.4\mathbf{38.3 \pm 5.4}

    Performance across all models drops when evaluated on out-of-domain data. However, ViLa-MIL improves AUC by 5.5% on TIHD→TCGA\text{TIHD} \to \text{TCGA} (83.4% vs. 78.2% for IBMIL) and by 2.7% on TCGA→TIHD\text{TCGA} \to \text{TIHD} (75.4% vs. 72.7% for GTMIL). The descriptive text prompt encodes invariant diagnostic criteria that provide robustness against scanner and staining distribution shifts.

  8. Knowl 8 — Ablation Study of ViLa-MIL Architectural Components

    data/table

    An ablation study on the TIHD-RCC dataset under the 16-shot setting assesses the incremental impact of each component in ViLa-MIL: the prototype-guided patch decoder, the dual-scale visual descriptive prompt, and the context-guided text decoder.

    Method AUC (%) F1 (%) ACC (%)
    ABMIL + Low-scale 76.8±3.176.8 \pm 3.1 57.2±3.957.2 \pm 3.9 61.4±2.861.4 \pm 2.8
    ABMIL + High-scale 79.3±3.279.3 \pm 3.2 62.7±3.662.7 \pm 3.6 63.3±3.163.3 \pm 3.1
    Patch Decoder + Low-scale 79.4±2.179.4 \pm 2.1 60.8±5.160.8 \pm 5.1 62.8±3.962.8 \pm 3.9
    Patch Decoder + High-scale 82.9±2.682.9 \pm 2.6 64.8±5.164.8 \pm 5.1 65.6±4.765.6 \pm 4.7
    Patch Decoder + Dual-scale 83.6±2.783.6 \pm 2.7 67.8±4.567.8 \pm 4.5 68.3±4.168.3 \pm 4.1
    ViLa-MIL (Full Framework) 84.3±4.6\mathbf{84.3 \pm 4.6} 68.7±7.3\mathbf{68.7 \pm 7.3} 68.8±7.3\mathbf{68.8 \pm 7.3}

    Key takeaways:

    1. Replacing ABMIL attention pooling with the prototype-guided patch decoder improves low-scale AUC from 76.8% to 79.4% (+2.6%) and high-scale AUC from 79.3% to 82.9% (+3.6%).
    2. Integrating dual-scale features and prompts lifts AUC from 82.9% to 83.6% and F1 from 64.8% to 67.8% (+3.0%).
    3. Introducing the context-guided text decoder further improves performance to 84.3% AUC and 68.7% F1.
  9. Knowl 9 — Comparison of Prompting Strategies and Large Language Model Backbones

    data/table

    The effectiveness of prompt construction strategies and choice of frozen Large Language Model (LLM) was evaluated on the TIHD-RCC dataset under the 16-shot setting.

    Prompting Strategy AUC (%) F1 (%) ACC (%)
    Class-name-replacement (Ä WSI of {class name}") 66.4±2.566.4 \pm 2.5 40.0±4.440.0 \pm 4.4 44.3±3.044.3 \pm 3.0
    Diagnostic Guideline (WHO tumor classification text) 78.0±2.878.0 \pm 2.8 59.5±2.259.5 \pm 2.2 61.1±1.261.1 \pm 1.2
    Large Language Model (single-scale GPT-3.5) 80.5±1.580.5 \pm 1.5 63.3±4.563.3 \pm 4.5 64.6±4.864.6 \pm 4.8
    ViLa-MIL (dual-scale descriptive prompt) 84.3±4.6\mathbf{84.3 \pm 4.6} 68.7±7.3\mathbf{68.7 \pm 7.3} 68.8±7.3\mathbf{68.8 \pm 7.3}
    LLM Backbone AUC (%) F1 (%) ACC (%)
    PaLM-2 83.5±3.683.5 \pm 3.6 65.8±3.665.8 \pm 3.6 66.3±3.966.3 \pm 3.9
    LLaMA-2 81.5±3.381.5 \pm 3.3 66.2±7.366.2 \pm 7.3 66.8±3.566.8 \pm 3.5
    GPT-3.5 84.3±4.684.3 \pm 4.6 68.7±7.3\mathbf{68.7 \pm 7.3} 68.8±7.368.8 \pm 7.3
    GPT-4 85.6±2.5\mathbf{85.6 \pm 2.5} 68.0±4.568.0 \pm 4.5 69.2±4.8\mathbf{69.2 \pm 4.8}

    Descriptive visual texts substantially outperform simple class-name substitution (+17.9% AUC). Dual-scale descriptive prompting improves upon single-scale LLM prompting by +3.8% AUC and +5.4% F1. ViLa-MIL is robust across open and closed LLM backbones (PaLM-2, LLaMA-2, GPT-3.5, GPT-4), with GPT-4 yielding the highest AUC (85.6%).

  10. Knowl 10 — Comparison of Slide Feature Aggregation Mechanisms

    data/table

    The prototype-guided patch decoder was compared against common patch feature aggregation mechanisms on the TIHD-RCC dataset under the 16-shot setting, holding the rest of the ViLa-MIL pipeline constant.

    Patch Feature Aggregator AUC (%) F1 (%) ACC (%)
    Mean Pooling 80.8±4.580.8 \pm 4.5 66.6±5.266.6 \pm 5.2 67.6±4.967.6 \pm 4.9
    Attention-based Pooling 81.1±2.981.1 \pm 2.9 62.9±7.562.9 \pm 7.5 64.9±4.464.9 \pm 4.4
    Self-attention-based Pooling 80.1±3.280.1 \pm 3.2 62.3±4.262.3 \pm 4.2 62.8±3.962.8 \pm 3.9
    ViLa-MIL (Prototype-Guided Patch Decoder) 84.3±4.6\mathbf{84.3 \pm 4.6} 68.7±7.3\mathbf{68.7 \pm 7.3} 68.8±7.3\mathbf{68.8 \pm 7.3}

    The prototype-guided patch decoder outperforms mean pooling, attention pooling, and self-attention pooling by 3.2–4.2% in AUC, 2.1–6.4% in F1 score, and 1.2–6.0% in accuracy. Grouping patch features into NpN_p learnable prototypes before final slide aggregation provides a more effective representation for alignment with textual features.

Coverage note — No substantial contributed material was omitted. Qualitative visual analysis (t-SNE embedding plots and spatial cancer localization heatmaps) is addressed as contextual evidence in the empirical knowls rather than as standalone knowls.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020.
  2. 2.Tsai Hor Chan, Fernando Julio Cendra, Lan Ma, Guosheng Yin, and Lequan Yu. Histopathology whole slide image analysis with heterogeneous graph representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15661–15670, 2023.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. PaLM: Scaling language modeling with Pathways. arXiv preprint arXiv:2204.02311, 2022.
  4. 4.Miao Cui and David Y Zhang. Artificial intelligence and computational pathology. Laboratory Investigation, 101(4):412–422, 2021.
  5. 5.Jevgenij Gamper and Nasir Rajpoot. Multiple instance captioning: Learning representations from histopathology textbooks and articles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16549–16559, 2021.
  6. 6.Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. CLIP-Adapter: Better vision-language models with feature adapters. International Journal of Computer Vision, pages 1–15, 2023.
  7. 7.Zixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai, Yiwen Guo, and Wangmeng Zuo. Texts as images in prompt tuning for multi-label image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2808–2817, 2023.
  8. 8.Cheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao, Wenguan Wang, Siyuan Qi, and Dongfang Liu. E2VPT\text{E}^2\text{VPT}: An effective and efficient approach for visual prompt tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 17491–17502, 2023.
  9. 9.Wentai Hou, Lequan Yu, Chengxuan Lin, Helong Huang, Rongshan Yu, Jing Qin, and Liansheng Wang. H2\text{H}^2-MIL: Exploring hierarchical representation with heterogeneous multiple instance learning for whole slide image analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 933–941, 2022.
  10. 10.Yanyan Huang, Weiqin Zhao, Shujun Wang, Yu Fu, Yuming Jiang, and Lequan Yu. ConSlide: Asynchronous hierarchical interaction Transformer with breakup-reorganize rehearsal for continual whole slide image analysis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 21349–21360, 2023.
  11. 11.Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. A visual-language foundation model for pathology image analysis using medical Twitter. Nature Medicine, pages 1–10, 2023.
  12. 12.Wisdom Ikezogwo, Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, and Linda Shapiro. Quilt-1M: One million image-text pairs for histopathology. Advances in Neural Information Processing Systems, 36, 2024.
  13. 13.Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learning. In Proceedings of the International Conference on Machine Learning, pages 2127–2136, 2018.
  14. 14.Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan. MaPle: Multi-modal prompt learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19113–19122, 2023.
  15. 15.Bin Li, Yin Li, and Kevin W Eliceiri. Dual-stream multiple instance learning network for whole slide image classification with self-supervised contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14318–14328, 2021.
  16. 16.Honglin Li, Chenglu Zhu, Yunlong Zhang, Yuxuan Sun, Zhongyi Shui, Wenwei Kuang, Sunyi Zheng, and Lin Yang. Task-specific fine-tuning via variational information bottleneck for weakly-supervised pathology whole slide image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7454–7463, 2023.
  17. 17.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR, 2022.
  18. 18.Kailu Li, Ziniu Qian, Yingnan Han, I Eric, Chao Chang, Bingzheng Wei, Maode Lai, Jing Liao, Yubo Fan, and Yan Xu. Weakly supervised histopathology image segmentation with self-attention. Medical Image Analysis, 86:102791, 2023.
  19. 19.Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He. Scaling language-image pre-training via masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23390–23400, 2023.
  20. 20.Tiancheng Lin, Zhimiao Yu, Hongyu Hu, Yi Xu, and Chang-Wen Chen. Interventional bag multi-instance learning on whole-slide pathological images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19830–19839, 2023.
  21. 21.Yuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu, Ke Li, Binbin Lin, Haifeng Liu, and Xiaofei He. CLIP is also an efficient segmenter: A text-driven approach for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15305–15314, 2023.
  22. 22.Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guillaume Jaume, Igor Odintsov, Andrew Zhang, Long Phi Le, et al. Towards a visual-language foundation model for computational pathology. arXiv preprint arXiv:2307.12914, 2023.
  23. 23.Ming Y Lu, Bowen Chen, Andrew Zhang, Drew FK Williamson, Richard J Chen, Tong Ding, Long Phi Le, Yung-Sung Chuang, and Faisal Mahmood. Visual language pre-trained multiple instance zero-shot transfer for histopathology images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19764–19775, 2023.
  24. 24.Ming Y Lu, Drew FK Williamson, Tiffany Y Chen, Richard J Chen, Matteo Barbieri, and Faisal Mahmood. Data-efficient and weakly supervised computational pathology on whole-slide images. Nature Biomedical Engineering, 5(6):555–570, 2021.
  25. 25.Oded Maron and Tom'{a}s Lozano-P'{e}rez. A framework for multiple-instance learning. Advances in Neural Information Processing Systems, 10:570–576, 1997.
  26. 26.Ramin Nakhli, Puria Azadi Moghadam, Haoyang Mi, Hossein Farahani, Alexander Baras, Blake Gilks, and Ali Bashashati. Sparse multi-modal graph Transformer with shared-context processing for representation learning of giga-pixel images. pages 11547–11557, 2023.
  27. 27.Muhammad Khalid Khan Niazi, Anil V Parwani, and Metin N Gurcan. Digital pathology and artificial intelligence. The Lancet Oncology, 20(5):e253–e261, 2019.
  28. 28.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  29. 29.Linhao Qu, Kexue Fu, Manning Wang, Zhijian Song, et al. The rise of ai language pathologists: Exploring two-level prompt learning for few-shot weakly-supervised whole slide image classification. Advances in Neural Information Processing Systems, 36, 2024.
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763, 2021.
  31. 31.Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. DenseCLIP: Language-guided dense prediction with context-aware prompting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18082–18091, 2022.
  32. 32.Jeongun Ryu, Aaron Valero Puche, JaeWoong Shin, Seonwook Park, Biagio Brattoli, Jinhee Lee, Wonkyung Jung, Soo Ick Cho, Kyunghyun Paeng, Chan-Young Ock, et al. OCELOT: Overlapped cell on tissue dataset for histopathology. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23902–23912, 2023.
  33. 33.Zhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang, Jian Zhang, Xiangyang Ji, et al. TransMIL: Transformer based correlated multiple instance learning for whole slide image classification. Advances in Neural Information Processing Systems, 34:2136–2147, 2021.
  34. 34.Jiangbo Shi, Lufei Tang, Zeyu Gao, Yang Li, Chunbao Wang, Tieliang Gong, Chen Li, and Huazhu Fu. MG-Trans: Multi-scale graph transformer with information bottleneck for whole slide image classification. IEEE Transactions on Medical Imaging, 42(12):3871–3883, 2023.
  35. 35.Jiangbo Shi, Lufei Tang, Yang Li, Xianli Zhang, Zeyu Gao, Yefeng Zheng, Chunbao Wang, Tieliang Gong, and Chen Li. A structure-aware hierarchical graph-based multiple instance learning framework for pT staging in histopathological image. IEEE Transactions on Medical Imaging, 42(10):3000–3011, 2023.
  36. 36.Andrew H Song, Guillaume Jaume, Drew FK Williamson, Ming Y Lu, Anurag Vaidya, Tiffany R Miller, and Faisal Mahmood. Artificial intelligence for digital and computational pathology. Nature Reviews Bioengineering, pages 1–20, 2023.
  37. 37.Wenhao Tang, Sheng Huang, Xiaoxian Zhang, Fengtao Zhou, Yi Zhang, and Bo Liu. Multiple instance learning framework with masked hard instance mining for whole slide image classification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4078–4087, 2023.
  38. 38.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth'{e}e Lacroix, Baptiste Rozi`{e}re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  39. 39.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-SNE. Journal of Machine Learning Research, 9(11):2579–2605, 2008.
  40. 40.Xiaoshi Wu, Feng Zhu, Rui Zhao, and Hongsheng Li. CORA: Adapting CLIP for open-vocabulary detection with region prompting and anchor pre-matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7031–7040, 2023.
  41. 41.Gang Xu, Zhigang Song, Zhuo Sun, Calvin Ku, Zhe Yang, Cancheng Liu, Shuhao Wang, Jianpeng Ma, and Wei Xu. CAMEL: A weakly supervised learning framework for histopathology image segmentation. In Proceedings of the IEEE/CVF International Conference on computer vision, pages 10682–10691, 2019.
  42. 42.Tao Yu, Zhihe Lu, Xin Jin, Zhibo Chen, and Xinchao Wang. Task residual for tuning vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10899–10909, 2023.
  43. 43.Hongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao, Xiaoyun Yang, Sarah E Coupland, and Yalin Zheng. DTFD-MIL: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18802–18812, 2022.
  44. 44.Sheng Zhang, Yanbo Xu, Naoto Usuyama, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, et al. Large-scale domain-specific pretraining for biomedical vision-language processing. arXiv preprint arXiv:2303.00915, 2023.
  45. 45.Yunkun Zhang, Jin Gao, Mu Zhou, Xiaosong Wang, Yu Qiao, Shaoting Zhang, and Dequan Wang. Text-guided foundation model adaptation for pathological image classification. In International Conference on Medical Image Computing and Computer Assisted Intervention, pages 272–282. Springer, 2023.
  46. 46.Yi Zheng, Rushin H. Gindra, Emily J. Green, Eric J. Burks, Margrit Betke, Jennifer E. Beane, and Vijaya B. Kolachalama. A graph-Transformer for whole slide image classification. IEEE Transactions on Medical Imaging, 41(11):3003–3015, 2022.
  47. 47.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. RegionCLIP: Region-based language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16793–16803, 2022.
  48. 48.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022.

Citation

MLA
Shi, J., et al. “ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification”. arXiv, 2025, http://arxiv.org/abs/2502.08391v1.
APA
Shi, J., Li, C., Gong, T., Zheng, Y., & Fu, H. (2025). ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification. arXiv. http://arxiv.org/abs/2502.08391v1
Chicago
Shi, J., C. Li, T. Gong, Y. Zheng, and H. Fu. 2025. “ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification”. arXiv. http://arxiv.org/abs/2502.08391v1.
Harvard
Shi, J. et al. (2025) “ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.08391v1.
Vancouver
1. Shi J, Li C, Gong T, Zheng Y, Fu H (2025) ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification. arXiv

BibTeX

@article{shi2025vila,
  title = {ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification},
  author = {Shi, Jiangbo and Li, Chen and Gong, Tieliang and Zheng, Yefeng and Fu, Huazhu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.08391v1},
  eprint = {2502.08391}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE