A region-level caption is a targeted textual description that explains the visual content of a specific area, object, or segment within an image rather than describing the entire scene as a whole. In computer vision and multimodal artificial intelligence, this format grounds natural language to localized spatial regions, which are typically identified through bounding boxes, visual prompts, or pixel-level segmentation masks. Unlike global image captions that summarize an overall environment, region-level captions provide fine-grained descriptions of specific attributes, identities, states, and localized interactions, supporting detailed object-level and pixel-level reasoning across vision-language tasks.