Built independently by an author, for readers. Read the story and support ChapterPal

keyword

contrastive language-image pretraining

Contrastive language-image pretraining is a machine learning method used to train artificial intelligence models to understand both images and natural language within a shared representation space. The approach typically employs a dual-encoder architecture that processes paired images and textual captions concurrently through separate visual and language networks. During training on large-scale datasets, a contrastive loss function maximizes the mathematical similarity between corresponding image-text pairs while minimizing the similarity between unrelated pairs. By aligning visual content with descriptive language rather than relying on a fixed set of predefined classification labels, this training paradigm enables models to perform open-vocabulary visual recognition, zero-shot transfer, and cross-modal retrieval across a broad range of downstream tasks.

2 items

HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention

Shijie Geng, Jianbo Yuan, Yu Tian, Yuxiao Chen, Yongfeng Zhang

OrganizationsByteDanceRutgers University

Why you should read this

Introduces HiCLIP to integrate hierarchy-aware attention into contrastive vision-language pretraining, enabling the unsupervised discovery of multi-level semantics across images and text to improve cross-modal alignment on downstream tasks.

The success of large-scale contrastive vision-language pretraining (CLIP) has benefited both visual recognition and multimodal content understanding. The concise design brings CLIP the advantage in inference efficiency against other vision-language models with heavier cross-attention fusion layers, making it a popular choice for a wide spectrum of downstream tasks. However, CLIP does not explicitly capture the hierarchical nature of high-level and fine-grained semantics conveyed in images and texts, which is arguably critical to vision-language understanding and reasoning. To this end, we equip both the visual and language branches in CLIP with hierarchy-aware attentions, namely Hierarchy-aware CLIP (HiCLIP), to progressively discover semantic hierarchies layer-by-layer from both images and texts in an unsupervised manner. As a result, such hierarchical aggregation significantly improves the cross-modal alignment. To demonstrate the advantages of HiCLIP, we conduct qualitative analysis on its unsupervised hierarchy induction during inference, as well as extensive quantitative experiments on both visual recognition and vision-language downstream tasks.

Added

2026-09-26

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

Siyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, Ee-Chien Chang

OrganizationsBeihang UniversityNational University of SingaporeSun Yat-sen UniversityThe Chinese University of Hong Kong

Why you should read this

Proposes BadCLIP, a dual-embedding guided backdoor attack on multimodal contrastive learning models that aligns visual trigger patterns with target text and image representations to evade state-of-the-art detection and persist through clean fine-tuning.

While existing backdoor attacks have successfully infected multimodal contrastive learning models such as CLIP, they can be easily countered by specialized backdoor defenses for MCL models. This paper reveals the threats in this practical scenario and introduces the BadCLIP attack, which is resistant to backdoor detection and model fine-tuning defenses. To achieve this, we draw motivations from the perspective of the Bayesian rule and propose a dual-embedding guided framework for backdoor attacks. Specifically, we ensure that visual trigger patterns approximate the textual target semantics in the embedding space, making it challenging to detect the subtle parameter variations induced by backdoor learning on such natural trigger patterns. Additionally, we optimize the visual trigger patterns to align the poisoned samples with target vision features in order to hinder backdoor unlearning through clean fine-tuning. Our experiments show a significant improvement in attack success rate (+45.3% ASR) over current leading methods, even against state-of-the-art backdoor defenses, highlighting our attack’s effectiveness in various scenarios, including downstream tasks. Our codes can be found at https://github.com/LiangSiyuan21/BadCLIP.

Added

2026-09-26