Built independently by an author, for readers. Read the story and support ChapterPal

keyword

CoDet

CoDet is a computer vision framework designed for open-vocabulary object detection that aligns image regions with textual words by discovering co-occurring objects across multiple image-text pairs. Rather than relying strictly on pre-aligned vision-language embedding spaces, the approach groups images that share common concepts in their captions and uses visual similarity to identify the recurring visual regions corresponding to those words. By framing region-word alignment as a co-occurring object discovery problem, CoDet enables detectors to learn fine-grained visual representations from large-scale image-text data without requiring extensive bounding-box annotations, improving both localization accuracy and the model ability to recognize novel object categories.

1 item

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Chuofan Ma, Yi Jiang, Xin Wen, Zehuan Yuan, Xiaojuan Qi

OrganizationsByteDanceUniversity of Hong Kong

Why you should read this

Proposes an open-vocabulary object detection framework that bypasses pre-aligned vision-language models by discovering co-occurring visual objects across captioned image groups to achieve state-of-the-art novel category detection on OV-LVIS.

Deriving reliable region-word alignment from image-text pairs is critical to learn object-level vision-language representations for open-vocabulary object detection. Existing methods typically rely on pre-trained or self-trained vision-language models for alignment, which are prone to limitations in localization accuracy or generalization capabilities. In this paper, we propose CoDet, a novel approach that overcomes the reliance on pre-aligned vision-language space by reformulating region-word alignment as a co-occurring object discovery problem. Intuitively, by grouping images that mention a shared concept in their captions, objects corresponding to the shared concept shall exhibit high co-occurrence among the group. CoDet then leverages visual similarities to discover the co-occurring objects and align them with the shared concept. Extensive experiments demonstrate that CoDet has superior performances and compelling scalability in open-vocabulary detection, e.g., by scaling up the visual backbone, CoDet achieves 37.0 AP_novel^m and 44.7 AP_all^m on OV-LVIS, surpassing the previous SoTA by 4.2 AP_novel^m and 9.8 AP_all^m. Code is available at https://github.com/CVMI-Lab/CoDet.

Added

2026-09-26