CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection
Chuofan MaYi JiangXin WenZehuan YuanXiaojuan Qi
Proposes an open-vocabulary object detection framework that bypasses pre-aligned vision-language models by discovering co-occurring visual objects across captioned image groups to achieve state-of-the-art novel category detection on OV-LVIS.
Traditional computer vision systems are typically restricted to identifying a fixed set of predefined object categories, limiting their effectiveness in open-world environments where novel and diverse objects frequently appear. While recent developments leverage web-scale image and text captions to train open-vocabulary detectors, these systems require precise alignments between specific image regions and text labels. Existing approaches rely heavily on pre-trained vision-language models to establish these connections, but such models often exhibit poor localization accuracy and struggle to generalize to unseen objects, creating a self-limiting cycle where training better detectors requires pre-existing high-quality detectors. The article introduces and evaluates CoDet, an open-vocabulary object detection framework that eliminates the need for pre-aligned vision-language models by reformulating region-word alignment as a visual co-occurrence discovery problem.
The framework operates on the principle that images sharing a descriptive concept in their captions will consistently contain visually similar representations of that object. Rather than attempting direct text-to-image region matching, CoDet groups images by shared caption terms and identifies regions that repeatedly appear across the group based on visual similarity. To prevent confusion when multiple objects appear together, the approach incorporates text guidance to weight semantic features, focusing visual comparisons on dimensions relevant to the target concept. The identified regions are combined into a prototypical visual representation of the concept and used to supervise the detector alongside standard detection data. The authors validated this methodology across standard open-vocabulary benchmarks, including OV-LVIS and OV-COCO, and tested its ability to generalize to new datasets without retraining.
The findings establish that CoDet consistently outperforms existing state-of-the-art methods in detecting novel categories. On the large-scale OV-LVIS benchmark, scaling the system with high-capacity visual backbones enabled CoDet to achieve a novel category mask average precision of 37.0 and an overall precision of 44.7, surpassing previous leading approaches by 4.2 and 9.8 points, respectively. Across standard model sizes, the framework consistently outperformed rival methods relying on vision-language models or size-based heuristics. In zero-shot transfer evaluations across different datasets, CoDet achieved approximately a 2 percentage point improvement over the best competing systems on both COCO and Objects365. Furthermore, comparative analyses revealed that visual co-occurrence discovery generates substantially more accurate and stable training signals than direct text-to-region matching, which frequently degrades due to incorrect initial label assignments.
These results demonstrate that visual correspondence across images provides a scalable and cost-effective alternative to manual region annotations and imperfect vision-language teacher models. Organizations building large-scale perception systems can train open-vocabulary models more effectively by pairing scalable visual backbones with uncurated web-crawled image-text data. For deployment, teams should scale concept group sizes on diverse web data while exercising caution on heavily curated datasets where artificial concept co-occurrences can introduce visual noise. Confidence in these results is supported by consistent benchmark improvements, though future work should evaluate hybrid approaches combining co-occurrence mechanisms with vision-language models to further enhance detection performance.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). ViLD establishes the foundational paradigm of distilling vision-language model embeddings into object detectors for open-vocabulary detection, which CoDet directly critiques and seeks to replace with co-occurrence discovery.
- Paper: RegionCLIP: Region-based Language-Image Pretraining, Yiwu Zhong et al. (2022). RegionCLIP investigates pretraining fine-grained region-text alignments from captioned images, introducing the central challenge of region-word alignment that CoDet solves via visual co-occurrence.
- Paper: ReCo: Retrieve and Co-segment for Zero-shot Transfer, Gyungin Shin et al. (2022). ReCo introduces the concept of discovering recurring visual concepts across retrieved images via co-segmentation and language gating, directly inspiring CoDet's visual co-occurrence framework.
- Paper: Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling, Dat Huynh et al. (2022). XPM explores cross-modal pseudo-labeling from image captions to supervise open-vocabulary object detectors, providing essential background on weakly supervised region alignment and noise management.
- Paper: Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks, Xiujun Li et al. (2020). Oscar demonstrates how using detected visual concepts as anchor points connects image regions to words, motivating CoDet's concept-grouped visual alignment.
- Paper: Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts, Soravit Changpinyo et al. (2021). Conceptual 12M provides the web-scale image-text pretraining dataset and conceptual vocabulary setup leveraged by modern open-vocabulary vision architectures.
- Paper: Object Recognition as Machine Translation: Learning a Lexicon for a Fixed Image Vocabulary, Pinar Duygulu et al. (2002). This seminal paper introduced formulating region-word alignment and object recognition as an unsupervised translation problem from co-occurring visual regions and text labels.
- Paper: Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection, Shilong Liu et al. (2023). Grounding DINO extends open-vocabulary object detection into a unified end-to-end grounded pre-training framework that can detect and localize open-set referring expressions and concepts.
- Paper: Few-Shot Object Detection with Foundation Models, Guangxing Han et al. (2024). FM-FSOD explores how pre-trained vision-language foundation models and visual prototypes can be leveraged for downstream few-shot novel object localization.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). PACL tackles the related open-vocabulary challenge at the pixel level by learning localized patch-text contrastive alignments from uncurated image-text pairs.
- Paper: PROB: Probabilistic Objectness for Open World Object Detection, Orr Zohar et al. (2023). PROB extends open-world object detection by using probabilistic modeling to identify unknown object candidates without requiring pre-aligned region-text teachers.
- Paper: CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data, Yihan Zeng et al. (2023). CLIP2 utilizes 2D open-vocabulary detectors to automatically harvest cross-modal triplets for open-vocabulary 3D point cloud perception.
