Region-word similarity is a metric in multimodal artificial intelligence that measures the degree of semantic alignment and correspondence between a localized region of an image and an individual word in a text sequence. It is typically calculated by mapping visual feature representations of image regions, such as objects or background elements, alongside word embeddings into a shared semantic space, commonly using vector operations like cosine similarity. This fine-grained measurement allows computational models to capture local vision-language interactions, serving as a foundational element for cross-modal attention, visual grounding, and comprehensive image-text retrieval tasks.