Stacked cross attention is a multimodal attention mechanism that aligns and measures semantic similarity between fine-grained components of two distinct data modalities, such as image regions and words in a sentence. Rather than comparing global feature vectors, the approach applies cross attention in two complementary directions, using elements of one modality as context to dynamically weight and aggregate relevant features from the other. This bidirectional, component-level alignment produces context-aware representations that capture latent semantic correspondences between visual elements and descriptive text, facilitating accurate cross-modal matching and retrieval.