Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-head cross attention

Multi-head cross attention is a neural network mechanism in transformer architectures that allows a model to selectively focus on and integrate relevant information from one data sequence or modality into another. Unlike self-attention, in which queries, keys, and values are derived from the same input, cross attention generates its queries from a primary input stream while taking its keys and values from a separate context or external source. By projecting these inputs into multiple parallel attention heads, the mechanism enables the network to simultaneously learn diverse contextual relationships and correspondences across distinct representation subspaces before concatenating and projecting the results into a unified output.

1 item

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training

Chen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng, Deli Zhao, Jingren Zhou

OrganizationsAlibaba Group

Why you should read this

Proposes an open-book contrastive learning framework that uses online image-text retrieval from a reference set to augment visual embeddings, boosting zero-shot classification performance without requiring models to memorize vast training concepts.

Contrastive Language-Image Pre-training (CLIP) is attracting increasing attention for its impressive zero-shot recognition performance on different down-stream tasks. However, training CLIP is data-hungry and requires lots of image-text pairs to memorize various semantic concepts. In this paper, we propose a novel and efficient framework: Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) to augment embeddings by online retrieval. Specifically, we sample part of image-text data as a hold-out reference set. Given an input image, relevant image-text pairs are retrieved from the reference set to enrich the representation of input image. This process can be considered as an open-book exam: with the reference set as a cheat sheet, the proposed method doesn't need to memorize all visual concepts in the training data. It explores how to recognize visual concepts by exploiting correspondence between images and texts in the cheat sheet. The proposed RA-CLIP implements this idea and comprehensive experiments are conducted to show how RA-CLIP works. Performances on 10 image classification datasets and 2 object detection datasets show that RA-CLIP outperforms vanilla CLIP baseline by a large margin on zero-shot image classification task (+12.7%), linear probe image classification task (+6.9%) and zero-shot ROI classification task (+2.8%).

Added

2026-09-26