Built independently by an author, for readers. Read the story and support ChapterPal

keyword

intra-modal self-similarity

Intra-modal self-similarity is a measure of the pairwise semantic, structural, or feature-level similarity among different data instances or elements within the same modality, such as between pairs of images or pairs of text passages. In multimodal machine learning and representation learning, it captures the internal relationships, shared semantic overlaps, and fine-grained correlations inherent to a single data type rather than comparing across different modalities. By modeling these within-modality relationships, systems can account for many-to-many correspondences and partial similarities across samples, providing richer supervisory signals to soften rigid one-to-one alignment targets during multimodal training.

1 item

SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger

SoftCLIP: Softer Cross-Modal Alignment Makes CLIP Stronger

Yuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu, Enwei Zhang, Ke Li, Jie Yang, Wei Liu, Xing Sun

OrganizationsShanghai Jiao Tong UniversityTencent

Why you should read this

Proposes a relaxed contrastive learning framework that uses fine-grained intra-modal self-similarity and negative-pair disentanglement as soft alignment targets, significantly improving CLIP's zero-shot classification performance on noisy web-scale image-text datasets.

During the preceding biennium, vision-language pre-training has achieved noteworthy success on several downstream tasks. Nevertheless, acquiring high-quality image-text pairs, where the pairs are entirely exclusive of each other, remains a challenging task, and noise exists in the commonly used datasets. To address this issue, we propose SoftCLIP, a novel approach that relaxes the strict one-to-one constraint and achieves a soft cross-modal alignment by introducing a softened target, which is generated from the fine-grained intra-modal self-similarity. The intra-modal guidance is indicative to enable two pairs have some local similarities and model many-to-many relationships between the two modalities. Besides, since the positive still dominates in the softened target distribution, we disentangle the negatives in the distribution to further boost the relation alignment with the negatives in the cross-modal learning. Extensive experiments demonstrate the effectiveness of SoftCLIP. In particular, on ImageNet zero-shot classification task, using CC3M/CC12M as pre-training dataset, SoftCLIP brings a top-1 accuracy improvement of 6.8%/7.2% over the CLIP baseline.

Added

2026-09-26