Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal contrastive learning

Multimodal contrastive learning is a machine learning framework that trains models to understand and align information across different data modalities, such as images, text, audio, or tabular records, within a shared representation space. In this approach, separate or unified neural network encoders process distinct data types to project them into a common embedding space. The training process uses a contrastive objective that maximizes the mathematical similarity between paired, semantically related instances across modalities while pushing apart representations of unrelated pairs. By capturing cross-modal associations without the need for extensive task-specific manual labels, multimodal contrastive learning produces versatile representations that support a wide range of downstream applications, including zero-shot classification, cross-modal retrieval, and multimodal feature transfer.

3 items

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

Siyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, Ee-Chien Chang

OrganizationsBeihang UniversityNational University of SingaporeSun Yat-sen UniversityThe Chinese University of Hong Kong

Why you should read this

Proposes BadCLIP, a dual-embedding guided backdoor attack on multimodal contrastive learning models that aligns visual trigger patterns with target text and image representations to evade state-of-the-art detection and persist through clean fine-tuning.

While existing backdoor attacks have successfully infected multimodal contrastive learning models such as CLIP, they can be easily countered by specialized backdoor defenses for MCL models. This paper reveals the threats in this practical scenario and introduces the BadCLIP attack, which is resistant to backdoor detection and model fine-tuning defenses. To achieve this, we draw motivations from the perspective of the Bayesian rule and propose a dual-embedding guided framework for backdoor attacks. Specifically, we ensure that visual trigger patterns approximate the textual target semantics in the embedding space, making it challenging to detect the subtle parameter variations induced by backdoor learning on such natural trigger patterns. Additionally, we optimize the visual trigger patterns to align the poisoned samples with target vision features in order to hinder backdoor unlearning through clean fine-tuning. Our experiments show a significant improvement in attack success rate (+45.3% ASR) over current leading methods, even against state-of-the-art backdoor defenses, highlighting our attack’s effectiveness in various scenarios, including downstream tasks. Our codes can be found at https://github.com/LiangSiyuan21/BadCLIP.

Added

2026-09-26

Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts

Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts

Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton, Neil Houlsby

OrganizationsGoogle

Why you should read this

Introduces LIMoE, the first large-scale multimodal Mixture of Experts model, demonstrating significant zero-shot ImageNet accuracy improvements over dense models of equivalent computational cost by effectively processing images and text simultaneously with robust, entropy-based regularization.

Large sparsely-activated models have obtained excellent performance in multiple domains. However, such models are typically trained on a single modality at a time. We present the Language-Image MoE, LIMoE, a sparse mixture of experts model capable of multimodal learning. LIMoE accepts both images and text simultaneously, while being trained using a contrastive loss. MoEs are a natural fit for a multimodal backbone, since expert layers can learn an appropriate partitioning of modalities. However, new challenges arise; in particular, training stability and balanced expert utilization, for which we propose an entropy-based regularization scheme. Across multiple scales, we demonstrate remarkable performance improvement over dense models of equivalent computational cost. LIMoE-L/16 trained comparably to CLIP-L/14 achieves 78.6% zero-shot ImageNet accuracy (vs. 76.2%), and when further scaled to H/14 (with additional data) it achieves 84.1%, comparable to state-of-the-art methods which use larger custom per-modality backbones and pre-training schemes. We analyse the quantitative and qualitative behavior of LIMoE, and demonstrate phenomena such as differing treatment of the modalities and the organic emergence of modality-specific experts.

Added

2026-02-23

Creative Commons License