Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning

Fuying WangYuyin ZhouShujun WangVarut VardhanabhutiLequan Yu

article2022NeurIPS297 citations

Proposes a cross-modal alignment framework that simultaneously aligns medical images and radiology reports across pathological region, instance, and disease levels to improve downstream classification, detection, and segmentation tasks under low-data regimes.

Listen

Training deep learning systems for medical imaging typically demands extensive manual annotation by clinical experts, which is costly, slow, and hard to scale. While recent approaches attempt to learn representations directly from paired medical images and unstructured radiology reports, existing methods only align information at a single resolution—either globally across entire images or locally across isolated patches. This incomplete alignment discards vital cross-modal relationships and limits how well artificial intelligence models transfer to diverse clinical workflows.

The article introduces and evaluates the Multi-Granularity Cross-modal Alignment framework, an automated approach designed to learn generalizable medical visual representations from paired radiology reports. The system explicitly captures natural correspondences across three complementary semantic tiers: instance-level (matching full images to whole reports), token-level (aligning localized image regions to specific descriptive phrases via bidirectional cross-attention), and disease-level (grouping high-level clinical prototypes using cross-modal clustering).

The framework was pre-trained on approximately 217,000 frontal chest radiograph and report pairs from the MIMIC-CXR database. Evaluated across seven downstream medical benchmark datasets covering classification, localized object detection, and pixel-level semantic segmentation, the model demonstrated substantial data efficiency and diagnostic accuracy. Across three diagnostic classification benchmarks, the framework achieved top performance, showing an accuracy improvement of up to 8.3 percentage points on COVID-19 detection when using only 1% of labeled training data. In localized object detection tasks, it outperformed leading methods across all data sampling levels, delivering a 12.9% mean average precision on the RSNA Pneumonia benchmark with 1% training data. For semantic segmentation, the approach improved Dice overlap scores by up to 12.3 percentage points over previous multimodal models under constrained 1% annotation settings.

These findings indicate that integrating multi-level cross-modal alignment allows healthcare machine learning models to reach high diagnostic accuracy with minimal human annotation. By significantly reducing the volume of manually labeled cases required to train robust diagnostic tools, organizations can accelerate clinical deployment timelines, lower engineering costs, and reduce specialist annotation burdens. Direct comparisons also demonstrated that generic vision-language models pre-trained on natural images transfer poorly to medical tasks, underscoring the operational necessity of domain-specific medical pre-training.

Healthcare technology leaders and clinical deployment teams should prioritize multi-tier pre-training frameworks over single-level architectures when developing diagnostic vision tools. Organizations should pilot this framework on target clinical workflows where labeled data is scarce, while maintaining standard compliance reviews to audit source datasets for patient privacy and demographic biases before practical deployment. While the framework demonstrates robust statistical stability across repeated trials, its evaluation remains focused on chest X-rays without assessing retrieval tasks. Future research should expand the framework toward unified generative models and broader imaging modalities.

arXiv: 2210.06044
Cover for Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning

Abstract

Learning medical visual representations directly from paired radiology reports has become an emerging topic in representation learning. However, existing medical image-text joint learning methods are limited by instance or local supervision analysis, ignoring disease-level semantic correspondences. In this paper, we present a novel Multi-Granularity Cross-modal Alignment (MGCA) framework for generalized medical visual representation learning by harnessing the naturally exhibited semantic correspondences between medical image and radiology reports at three different levels, i.e., pathological region-level, instance-level, and disease-level. Specifically, we first incorporate the instance-wise alignment module by maximizing the agreement between image-report pairs. Further, for token-wise alignment, we introduce a bidirectional cross-attention strategy to explicitly learn the matching between fine-grained visual tokens and text tokens, followed by contrastive learning to align them. More important, to leverage the high-level inter-subject relationship semantic (e.g., disease) correspondences, we design a novel cross-modal disease-level alignment paradigm to enforce the cross-modal cluster assignment consistency. Extensive experimental results on seven downstream medical image datasets covering image classification, object detection, and semantic segmentation tasks demonstrate the stable and superior performance of our framework.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Overview
  • 3.2 Multi-granularity Cross-modal Alignment
  • 3.3 Overall Objective
  • 4 Experiments
  • 4.1 Pre-Training Setup
  • 4.2 Downstream Tasks and Experimental Setup
  • 4.3 Results
  • 4.4 Analysis of Our Framework
  • 5 Discussion and Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Multi-granularity cross-modal alignment framework

    model/method

    MGCA learns medical image representations from paired radiology images and reports by aligning cross-modal information at three granularities: global image–report instances, local pathological regions and report tokens, and high-level disease-related prototypes. For NN image–report pairs (xv,i,xt,i)(x_{v,i},x_{t,i}), an image encoder fvf_v produces a global image embedding viv_i and visual-token sequence Ri={rij}j=1SR_i=\{r_i^j\}_{j=1}^{S}, while a text encoder ftf_t produces a global report embedding tit_i and text-token sequence Zi={zik}k=1LZ_i=\{z_i^k\}_{k=1}^{L}. Two projection heads produce normalized global embeddings for instance and prototype alignment, and cross-attention modules transform local tokens for token-wise alignment. The three alignment losses are jointly optimized to obtain a visual encoder transferable to classification, detection, and segmentation.

  2. Knowl 2 — Instance-wise image–report alignment

    equation

    MGCA aligns each true image–report pair against mismatched pairs using symmetric temperature-scaled InfoNCE. For image–report pair ii, projection heads gvg_v and gtg_t map global embeddings vi,tiv_i,t_i to normalized vectors v~i,t~i∈Rd\tilde v_i,\tilde t_i\in\mathbb{R}^{d}, where dd is the projection dimension, and similarity is si,k=v~iTt~ks_{i,k}=\tilde v_i^{\mathsf T}\tilde t_k. For a minibatch B\mathcal{B} of size BB, the image-to-text and text-to-image losses are

    ℓiv→t=−log⁡exp⁡(si,i/τ1)∑k∈Bexp⁡(si,k/τ1),ℓit→v=−log⁡exp⁡(si,i/τ1)∑k∈Bexp⁡(t~iTv~k/τ1),\ell_i^{v\to t}=-\log\frac{\exp(s_{i,i}/\tau_1)}{\sum_{k\in\mathcal{B}}\exp(s_{i,k}/\tau_1)},\qquad \ell_i^{t\to v}=-\log\frac{\exp(s_{i,i}/\tau_1)}{\sum_{k\in\mathcal{B}}\exp(\tilde t_i^{\mathsf T}\tilde v_k/\tau_1)},

    where τ1\tau_1 is the instance-level temperature. The instance-wise alignment objective averages both directions:

    LITA=12B∑i∈B(ℓiv→t+ℓit→v).\mathcal{L}_{\mathrm{ITA}}=\frac{1}{2B}\sum_{i\in\mathcal{B}}\left(\ell_i^{v\to t}+\ell_i^{t\to v}\right).

    Thus, the matching image and report are pulled together while the other reports and images in the minibatch act as negatives.

  3. Knowl 3 — Bidirectional token-wise cross-attention alignment

    model/method

    MGCA uses bidirectional cross-attention to match local visual tokens with report tokens rather than relying only on global image–report similarity. The normalized visual and text tokens are r~ij,z~ik∈Rd\tilde r_i^j,\tilde z_i^k\in\mathbb{R}^{d}, with j∈{1,…,S}j\in\{1,\ldots,S\} and k∈{1,…,L}k\in\{1,\ldots,L\}. For a visual token, query, key, and value projections Q,K,V∈Rd×dQ,K,V\in\mathbb{R}^{d\times d} produce attention weights and a cross-modal text representation:

    αij→k=exp⁡((Qr~ij)T(Kz~ik)/d)∑u=1Lexp⁡((Qr~ij)T(Kz~iu)/d),oij=∑k=1Lαij→kVz~ik.\alpha_i^{j\rightarrow k}=\frac{\exp\left((Q\tilde r_i^j)^{\mathsf T}(K\tilde z_i^k)/\sqrt d\right)}{\sum_{u=1}^{L}\exp\left((Q\tilde r_i^j)^{\mathsf T}(K\tilde z_i^u)/\sqrt d\right)},\qquad o_i^j=\sum_{k=1}^{L}\alpha_i^{j\rightarrow k}V\tilde z_i^k.

    A symmetric text-to-image cross-attention operation produces an image representation for each text token. MGCA then applies per-token bidirectional InfoNCE: each visual token is contrasted with its matched cross-modal text representation against the other cross-modal text representations from the same pair, and each text token is treated analogously. The visual-token contribution is weighted by wijw_i^j, the final-layer attention from visual token jj to the image [CLS] token averaged across attention heads, so visually salient pathological regions contribute more strongly. The token-wise objective is LCTA=12(LLIA+LLTA)\mathcal{L}_{\mathrm{CTA}}=\tfrac12(\mathcal{L}_{\mathrm{LIA}}+\mathcal{L}_{\mathrm{LTA}}), where LIA and LTA are the local image-to-text and local text-to-image losses, respectively.

  4. Knowl 4 — Cross-modal disease-level prototype alignment

    model/method

    MGCA prevents image–report contrastive learning from treating every different patient as semantically unrelated by aligning disease-level cluster structure across modalities. For each normalized global pair (v~i,t~i)∈Rd×Rd(\tilde v_i,\tilde t_i)\in\mathbb{R}^{d}\times\mathbb{R}^{d}, an iterative Sinkhorn–Knopp clustering procedure produces soft visual and text assignment codes qv,i,qt,i∈RKq_{v,i},q_{t,i}\in\mathbb{R}^{K} over KK clusters. The model also learns KK shared cross-modal prototypes C={ck}k=1KC=\{c_k\}_{k=1}^{K}, with ck∈Rdc_k\in\mathbb{R}^{d}. Prototype probabilities are computed from cosine similarities:

    pv,i(k)=exp⁡(v~iTck/τ3)∑u=1Kexp⁡(v~iTcu/τ3),pt,i(k)=exp⁡(t~iTck/τ3)∑u=1Kexp⁡(t~iTcu/τ3), p_{v,i}^{(k)}=\frac{\exp(\tilde v_i^{\mathsf T}c_k/\tau_3)}{\sum_{u=1}^{K}\exp(\tilde v_i^{\mathsf T}c_u/\tau_3)},\qquad p_{t,i}^{(k)}=\frac{\exp(\tilde t_i^{\mathsf T}c_k/\tau_3)}{\sum_{u=1}^{K}\exp(\tilde t_i^{\mathsf T}c_u/\tau_3)},

    where τ3\tau_3 is the prototype-level temperature. Cross-modal prediction uses the text assignment qt,iq_{t,i} as a soft pseudo-label for the image probabilities pv,ip_{v,i} and the image assignment qv,iq_{v,i} as a soft pseudo-label for the text probabilities pt,ip_{t,i}. With CE(q,p)=−∑k=1Kq(k)log⁡p(k)\mathrm{CE}(q,p)=-\sum_{k=1}^{K}q^{(k)}\log p^{(k)}, the prototype objective is

    LCPA=12N∑i=1N[CE(qt,i,pv,i)+CE(qv,i,pt,i)].\mathcal{L}_{\mathrm{CPA}}=\frac{1}{2N}\sum_{i=1}^{N}\left[\mathrm{CE}(q_{t,i},p_{v,i})+\mathrm{CE}(q_{v,i},p_{t,i})\right].

    This encourages images and reports with related high-level disease semantics to share cross-modal cluster assignments instead of being pushed apart as ordinary negative instances.

  5. Knowl 5 — Joint MGCA training objective

    equation

    The MGCA pre-training loss is a weighted sum of the instance-, token-, and prototype-level alignment objectives:

    L=λ1LITA+λ2LCTA+λ3LCPA,\mathcal{L}=\lambda_1\mathcal{L}_{\mathrm{ITA}}+\lambda_2\mathcal{L}_{\mathrm{CTA}}+\lambda_3\mathcal{L}_{\mathrm{CPA}},

    where LITA\mathcal{L}_{\mathrm{ITA}} is the symmetric global image–report contrastive loss, LCTA\mathcal{L}_{\mathrm{CTA}} is the bidirectional local token alignment loss, LCPA\mathcal{L}_{\mathrm{CPA}} is the cross-modal prototype prediction loss, and λ1,λ2,λ3\lambda_1,\lambda_2,\lambda_3 balance the three granularities. The reported experiments use λ1=λ2=λ3=1\lambda_1=\lambda_2=\lambda_3=1.

  6. Knowl 6 — MIMIC-CXR pre-training configuration

    experimental setup

    MGCA is pre-trained on the JPG version of MIMIC-CXR 2.0.0. Lateral radiographs are removed because the target downstream datasets contain frontal chest images; the impression and findings portions of each report are retained, and empty reports or reports with fewer than three tokens are discarded, leaving approximately 217,000 image–report pairs. BioClinicalBERT is used as the text encoder and ViT-B/16 as the default image encoder; ResNet-50 is also evaluated for comparison with methods using that backbone. Training runs for 50 epochs on two RTX 3090 GPUs with batch size 144, AdamW, learning rate 2×10−52\times10^{-5}, weight decay 0.050.05, and a linear-warmup/cosine-annealing schedule initialized at 10−810^{-8} with 20 warmup epochs. The projection dimension is d=128d=128, temperatures are τ1=0.1\tau_1=0.1, τ2=0.07\tau_2=0.07, and τ3=0.2\tau_3=0.2, and the number of prototypes is K=500K=500.

  7. Knowl 7 — Transfer evaluation across seven medical datasets

    experimental setup

    The learned image encoder is evaluated on seven datasets spanning image classification, object detection, and semantic segmentation, using 1%, 10%, and 100% of downstream training data. Classification uses frozen-encoder linear evaluation on CheXpert, RSNA Pneumonia, and COVIDx-v6: CheXpert predicts five findings and reports AUROC, RSNA predicts normal versus pneumothorax and reports AUROC, and COVIDx predicts COVID-19, non-COVID pneumonia, or normal and reports accuracy. Detection uses a frozen ResNet-50 backbone in YOLOv3 on RSNA Pneumonia and Object CXR, with mAP computed over IoU thresholds from 0.4 through 0.75. Segmentation uses a frozen ResNet-50 encoder in U-Net on SIIM Pneumothorax and RSNA Pneumonia, with Dice score as the metric. For detection and segmentation, only the non-backbone layers are trained.

  8. Knowl 8 — Classification transfer performance

    empirical result

    MGCA with a ViT-B/16 image encoder achieves the best reported result in all nine linear-classification settings. With 1%, 10%, and 100% of downstream training data, its AUROC on CheXpert is 88.8, 89.1, and 89.7, respectively; its AUROC on RSNA Pneumonia is 89.1, 89.9, and 90.8; and its COVIDx accuracy is 74.8, 84.8, and 92.3. The ResNet-50 version obtains CheXpert AUROC 87.6, 88.0, and 88.2, RSNA AUROC 88.6, 89.1, and 89.9, and COVIDx accuracy 72.0, 83.5, and 90.5. Relative to GLoRIA-MIMIC at the 1% setting, the ViT-B/16 MGCA model improves CheXpert by 1.7 AUROC points, RSNA by 2.1 AUROC points, and COVIDx by 8.3 accuracy points, demonstrating strong transfer with very limited labels.

  9. Knowl 9 — Dense prediction transfer performance

    data/table

    MGCA improves localized medical prediction, especially in the low-label regime. Detection mAP for MGCA is 12.9, 16.8, and 24.9 on RSNA Pneumonia with 1%, 10%, and 100% training data, compared with 11.6, 16.1, and 24.8 for GLoRIA-MIMIC. On Object CXR, MGCA reports mAP below 1% at the 1% setting, 12.1 at 10%, and 19.2 at 100%; the corresponding GLoRIA-MIMIC values are below 1%, 8.9, and 16.6.

    Segmentation Dice scores for MGCA are 49.7, 59.3, and 64.2 on SIIM Pneumothorax and 63.0, 68.3, and 69.8 on RSNA Pneumonia for 1%, 10%, and 100% training data. GLoRIA-MIMIC obtains 37.4, 57.1, and 64.0 on SIIM and 60.3, 68.7, and 68.3 on RSNA. Thus, at 1% training data, MGCA improves over GLoRIA-MIMIC by 12.3 Dice points on SIIM and 2.7 Dice points on RSNA; the token-level alignment is particularly useful for learning localized representations.

  10. Knowl 10 — Complementarity of the three alignment modules

    data/table

    Ablations show that token- and prototype-level alignment each add value beyond instance-wise alignment, and their combination is complementary. The reported results use ViT-B/16 for CheXpert and RSNA classification and ResNet-50 for SIIM segmentation; classification metrics are AUROC and segmentation metrics are Dice, all in percent.

    ITA CTA CPA CheXpert AUROC RSNA AUROC
    used used used 1% 10% 100% 1% 10% 100%
    ✓ - - 87.6 88.2 88.5 88.4 89.5 90.5
    ✓ ✓ - 88.3 88.9 89.1 88.9 89.8 90.7
    ✓ - ✓ 88.5 88.9 89.0 88.6 89.2 90.4
    ✓ ✓ ✓ 88.8 89.1 89.7 89.1 89.9 90.8
    ITA CTA CPA SIIM Dice
    used used used 1% 10% 100%
    ✓ - - 25.0 43.2 59.9
    ✓ ✓ - 47.6 54.4 61.3
    ✓ - ✓ 37.4 46.7 55.0
    ✓ ✓ ✓ 49.7 59.3 64.2

    Adding CTA to ITA produces a larger SIIM improvement than adding CPA alone, consistent with CTA learning fine-grained pathological information. Jointly training ITA, CTA, and CPA gives the best result in every listed setting.

Coverage note — The qualitative token-correspondence and disease-cluster visualizations, the BLIP comparison, and repeat-run error-bar analysis were omitted as supporting analyses; the paper’s stated lack of retrieval experiments and future generation-based pre-training are also omitted because they are limitations and future work rather than additional contributions.

References

  1. 1.E. Alsentzer, J. R. Murphy, W. Boag, W.-H. Weng, D. Jin, T. Naumann, and M. McDermott. Publicly available clinical bert embeddings. arXiv preprint arXiv:1904.03323, 2019.
  2. 2.M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing Systems, 33:9912–9924, 2020.
  3. 3.K. Chaitanya, E. Erdil, N. Karani, and E. Konukoglu. Contrastive learning of global and local features for medical image segmentation with limited annotations. Advances in Neural Information Processing Systems, 33:12546–12558, 2020.
  4. 4.T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  5. 5.Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu. Uniter: Learning universal image-text representations. 2019.
  6. 6.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016.
  7. 7.M. Cornia, M. Stefanini, L. Baraldi, and R. Cucchiara. Meshed-memory transformer for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10578–10587, 2020.
  8. 8.M. Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. Advances in neural information processing systems, 26, 2013.
  9. 9.J. De Fauw, J. R. Ledsam, B. Romera-Paredes, S. Nikolov, N. Tomasev, S. Blackwell, H. Askham, X. Glorot, B. O’Donoghue, D. Visentin, et al. Clinically applicable deep learning for diagnosis and referral in retinal disease. Nature medicine, 24(9):1342–1350, 2018.
  10. 10.J. Devlin, M.-W. Chang, K. Lee, and Y. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  11. 11.H. Diao, Y. Zhang, L. Ma, and H. Lu. Similarity reasoning and filtration for image-text matching. arXiv preprint arXiv:2101.01368, 2021.
  12. 12.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, H. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  13. 13.M. Engilberge, L. Chevallier, P. Pérez, and M. Cord. Finding beans in burgers: Deep semantic-visual embedding with localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3984–3993, 2018.
  14. 14.A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun. Dermatologist-level classification of skin cancer with deep neural networks. nature, 542(7639):115–118, 2017.
  15. 15.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010.
  16. 16.F. Faghri, D. J. Fleet, J. R. Kiros, and S. Fidler. Vse++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612, 2017.
  17. 17.R. Feng, Z. Zhou, M. B. Gotway, and J. Liang. Parts2whole: Self-supervised contrastive learning via reconstruction. In Domain Adaptation and Representation Transfer, and Distributed and Collaborative Learning, pages 85–95. Springer, 2020.
  18. 18.S. for Imaging Informatics in Medicine. Siim-acr pneumothorax segmentation.
  19. 19.J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in Neural Information Processing Systems, 33:21271–21284, 2020.
  20. 20.V. Gulshan, L. Peng, M. Coram, M. C. Stumpe, D. Wu, A. Narayanaswamy, S. Venugopalan, K. Widner, T. Madams, J. Cuadros, et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. Jama, 316(22):2402–2410, 2016.
  21. 21.Y. Guo, M. Xu, J. Li, B. Ni, X. Zhu, Z. Sun, and Y. Xu. Hcsc: Hierarchical contrastive selective coding. arXiv preprint arXiv:2202.00455, 2022.
  22. 22.R. Hadsell, S. Chopra, and Y. LeCun. Dimensionality reduction by learning an invariant mapping. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), volume 2, pages 1735–1742. IEEE, 2006.
  23. 23.Y. Han, C. Chen, A. Tewfik, Y. Ding, and Y. Peng. Pneumonia detection on chest x-ray using radiomic features and contrastive learning. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 247–251. IEEE, 2021.
  24. 24.K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020.
  25. 25.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  26. 26.J. Healthcare. Object-cxr - automatic detection of foreign objects on chest x-rays.
  27. 27.S.-C. Huang, L. Shen, M. P. Lungren, and S. Yeung. Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3942–3951, 2021.
  28. 28.Y. Huang, Q. Wu, C. Song, and L. Wang. Learning semantic concepts and order for image and sentence matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6163–6171, 2018.
  29. 29.J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya, et al. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 590–597, 2019.
  30. 30.A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon. A survey on contrastive self-supervised learning. Technologies, 9(1):2, 2020.
  31. 31.A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C.-y. Deng, R. G. Mark, and S. Horng. Mimic-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, 6(1):1–8, 2019.
  32. 32.P. H. Le-Khac, G. Healy, and A. F. Smeaton. Contrastive representation learning: A framework and review. IEEE Access, 8:193907–193934, 2020.
  33. 33.K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He. Stacked cross attention for image-text matching. In Proceedings of the European Conference on Computer Vision (ECCV), pages 201–216, 2018.
  34. 34.J. Li, D. Li, C. Xiong, and S. Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. arXiv preprint arXiv:2201.12086, 2022.
  35. 35.J. Li, P. Zhou, C. Xiong, and S. C. Hoi. Prototypical contrastive learning of unsupervised representations. arXiv preprint arXiv:2005.04966, 2020.
  36. 36.K. Li, Y. Zhang, K. Li, Y. Li, and Y. Fu. Visual semantic reasoning for image-text matching. In Proceedings of the IEEE/CVF International conference on computer vision, pages 4654–4662, 2019.
  37. 37.Y. Li, P. Hu, Z. Liu, D. Peng, J. T. Zhou, and X. Peng. Contrastive clustering. In 2021 AAAI Conference on Artificial Intelligence (AAAI), 2021.
  38. 38.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  39. 39.I. Loshchilov and F. Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  40. 40.I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  41. 41.J. Lu, J. Yang, D. Batra, and D. Parikh. Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems, 29, 2016.
  42. 42.P. Müller, G. Kaissis, C. Zou, and D. Rückert. Joint learning of localized representations from medical images and reports. arXiv preprint arXiv:2112.02889, 2021.
  43. 43.A. v. d. Oord, Y. Li, and O. Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  44. 44.P. Rajpurkar, J. Irvin, R. L. Ball, K. Zhu, B. Yang, H. Mehta, T. Duan, D. Ding, A. Bagul, C. P. Langlotz, et al. Deep learning for chest radiograph diagnosis: A retrospective comparison of the chexnext algorithm to practicing radiologists. PLoS medicine, 15(11):e1002686, 2018.
  45. 45.J. Redmon and A. Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018.
  46. 46.O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  47. 47.G. Shih, C. C. Wu, S. S. Halabi, M. D. Kohli, L. M. Prevedello, T. S. Cook, A. Sharma, J. K. Amorosa, V. Arteaga, M. Galperin-Aizenberg, et al. Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia. Radiology: Artificial Intelligence, 1(1):e180041, 2019.
  48. 48.W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530, 2019.
  49. 49.A. Taleb, M. Kirchler, R. Monti, and C. Lippert. Contig: Self-supervised multimodal contrastive learning for medical imaging with genetics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20908–20921, 2022.
  50. 50.H. Tan and M. Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019.
  51. 51.L. Van der Maaten and G. Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  52. 52.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  53. 53.L. Wang, Z. Q. Lin, and A. Wong. Covid-net: A tailored deep convolutional neural network design for detection of covid-19 cases from chest x-ray images. Scientific Reports, 10(1):1–12, 2020.
  54. 54.X. Wang, Z. Liu, and S. X. Yu. Unsupervised feature learning by cross-level instance-group discrimination. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12586–12595, 2021.
  55. 55.X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2097–2106, 2017.
  56. 56.X. Wang, R. Zhang, C. Shen, T. Kong, and L. Li. Dense contrastive learning for self-supervised visual pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3024–3033, 2021.
  57. 57.E. Xie, J. Ding, W. Wang, X. Zhan, H. Xu, P. Sun, Z. Li, and P. Luo. Detco: Unsupervised contrastive learning for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8392–8401, 2021.
  58. 58.Z. Xie, Y. Lin, Z. Zhang, Y. Cao, S. Lin, and H. Hu. Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16684–16693, 2021.
  59. 59.J. Xu, S. De Mello, S. Liu, W. Byeon, T. Breuel, J. Kautz, and X. Wang. Groupvit: Semantic segmentation emerges from text supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18134–18144, 2022.
  60. 60.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning, pages 2048–2057. PMLR, 2015.
  61. 61.Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz. Contrastive learning of medical visual representations from paired images and text. arXiv preprint arXiv:2010.00747, 2020.
  62. 62.H.-Y. Zhou, X. Chen, Y. Zhang, R. Luo, L. Wang, and Y. Yu. Generalized radiograph representation learning via cross-supervision between images and free-text radiology reports. Nature Machine Intelligence, pages 1–9, 2022.

Citation

MLA
Wang, F., et al. “Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 33536–49, https://proceedings.neurips.cc/paper_files/paper/2022/file/d925bda407ada0df3190df323a212661-Paper-Conference.pdf.
APA
Wang, F., Zhou, Y., WANG, S., Vardhanabhuti, V., & Yu, L. (2022). Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning. Advances in Neural Information Processing Systems, 35, 33536–33549. https://proceedings.neurips.cc/paper_files/paper/2022/file/d925bda407ada0df3190df323a212661-Paper-Conference.pdf
Chicago
Wang, F., Y. Zhou, S. WANG, V. Vardhanabhuti, and L. Yu. 2022. “Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning”. Advances in Neural Information Processing Systems 35: 33536–49. https://proceedings.neurips.cc/paper_files/paper/2022/file/d925bda407ada0df3190df323a212661-Paper-Conference.pdf.
Harvard
Wang, F. et al. (2022) “Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 33536–33549. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/d925bda407ada0df3190df323a212661-Paper-Conference.pdf.
Vancouver
1. Wang F, Zhou Y, WANG S, Vardhanabhuti V, Yu L (2022) Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 33536–33549

BibTeX

@inproceedings{wang2022multi,
  title = {Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation Learning},
  author = {Wang, Fuying and Zhou, Yuyin and WANG, Shujun and Vardhanabhuti, Varut and Yu, Lequan},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {33536-33549},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/d925bda407ada0df3190df323a212661-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors