Built independently by an author, for readers. Read the story and support ChapterPal

keyword

co-occurrence discovery

Co-occurrence discovery is a visual learning process that identifies and localizes shared objects or visual patterns across a collection of related images. By analyzing groups of images that share a common semantic label, textual description, or contextual theme, the approach searches for visual similarities to isolate regions that repeatedly appear together across the group. This technique enables computer vision systems to distinguish foreground objects corresponding to shared concepts from varying backgrounds, facilitating object localization and establishing alignments between image regions and textual concepts without requiring fine-grained bounding-box annotations.

1 item

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Chuofan Ma, Yi Jiang, Xin Wen, Zehuan Yuan, Xiaojuan Qi

OrganizationsByteDanceUniversity of Hong Kong

Why you should read this

Proposes an open-vocabulary object detection framework that bypasses pre-aligned vision-language models by discovering co-occurring visual objects across captioned image groups to achieve state-of-the-art novel category detection on OV-LVIS.

Deriving reliable region-word alignment from image-text pairs is critical to learn object-level vision-language representations for open-vocabulary object detection. Existing methods typically rely on pre-trained or self-trained vision-language models for alignment, which are prone to limitations in localization accuracy or generalization capabilities. In this paper, we propose CoDet, a novel approach that overcomes the reliance on pre-aligned vision-language space by reformulating region-word alignment as a co-occurring object discovery problem. Intuitively, by grouping images that mention a shared concept in their captions, objects corresponding to the shared concept shall exhibit high co-occurrence among the group. CoDet then leverages visual similarities to discover the co-occurring objects and align them with the shared concept. Extensive experiments demonstrate that CoDet has superior performances and compelling scalability in open-vocabulary detection, e.g., by scaling up the visual backbone, CoDet achieves 37.0 AP_novel^m and 44.7 AP_all^m on OV-LVIS, surpassing the previous SoTA by 4.2 AP_novel^m and 9.8 AP_all^m. Code is available at https://github.com/CVMI-Lab/CoDet.

Added

2026-09-26