Built independently by an author, for readers. Read the story and support ChapterPal

keyword

region-word alignment

Region-word alignment refers to the process of establishing fine-grained semantic correspondences between specific spatial regions within an image and individual words or phrases in accompanying text. In multimodal machine learning and computer vision, this technique maps local visual features, such as bounding boxes, segmentation masks, or object proposals, directly to their corresponding linguistic descriptions rather than relying solely on global, whole-image caption matching. By learning localized vision-language representations, region-word alignment enables models to accurately locate, identify, and ground distinct visual entities, making it a foundational component in tasks such as open-vocabulary object detection, phrase grounding, and detailed visual reasoning.

1 item

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Chuofan Ma, Yi Jiang, Xin Wen, Zehuan Yuan, Xiaojuan Qi

OrganizationsByteDanceUniversity of Hong Kong

Why you should read this

Proposes an open-vocabulary object detection framework that bypasses pre-aligned vision-language models by discovering co-occurring visual objects across captioned image groups to achieve state-of-the-art novel category detection on OV-LVIS.

Deriving reliable region-word alignment from image-text pairs is critical to learn object-level vision-language representations for open-vocabulary object detection. Existing methods typically rely on pre-trained or self-trained vision-language models for alignment, which are prone to limitations in localization accuracy or generalization capabilities. In this paper, we propose CoDet, a novel approach that overcomes the reliance on pre-aligned vision-language space by reformulating region-word alignment as a co-occurring object discovery problem. Intuitively, by grouping images that mention a shared concept in their captions, objects corresponding to the shared concept shall exhibit high co-occurrence among the group. CoDet then leverages visual similarities to discover the co-occurring objects and align them with the shared concept. Extensive experiments demonstrate that CoDet has superior performances and compelling scalability in open-vocabulary detection, e.g., by scaling up the visual backbone, CoDet achieves 37.0 AP_novel^m and 44.7 AP_all^m on OV-LVIS, surpassing the previous SoTA by 4.2 AP_novel^m and 9.8 AP_all^m. Code is available at https://github.com/CVMI-Lab/CoDet.

Added

2026-09-26