CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Chuofan MaYi JiangXin WenZehuan YuanXiaojuan Qi

article2023NeurIPS78 citations

Proposes an open-vocabulary object detection framework that bypasses pre-aligned vision-language models by discovering co-occurring visual objects across captioned image groups to achieve state-of-the-art novel category detection on OV-LVIS.

Listen

Traditional computer vision systems are typically restricted to identifying a fixed set of predefined object categories, limiting their effectiveness in open-world environments where novel and diverse objects frequently appear. While recent developments leverage web-scale image and text captions to train open-vocabulary detectors, these systems require precise alignments between specific image regions and text labels. Existing approaches rely heavily on pre-trained vision-language models to establish these connections, but such models often exhibit poor localization accuracy and struggle to generalize to unseen objects, creating a self-limiting cycle where training better detectors requires pre-existing high-quality detectors. The article introduces and evaluates CoDet, an open-vocabulary object detection framework that eliminates the need for pre-aligned vision-language models by reformulating region-word alignment as a visual co-occurrence discovery problem.

The framework operates on the principle that images sharing a descriptive concept in their captions will consistently contain visually similar representations of that object. Rather than attempting direct text-to-image region matching, CoDet groups images by shared caption terms and identifies regions that repeatedly appear across the group based on visual similarity. To prevent confusion when multiple objects appear together, the approach incorporates text guidance to weight semantic features, focusing visual comparisons on dimensions relevant to the target concept. The identified regions are combined into a prototypical visual representation of the concept and used to supervise the detector alongside standard detection data. The authors validated this methodology across standard open-vocabulary benchmarks, including OV-LVIS and OV-COCO, and tested its ability to generalize to new datasets without retraining.

The findings establish that CoDet consistently outperforms existing state-of-the-art methods in detecting novel categories. On the large-scale OV-LVIS benchmark, scaling the system with high-capacity visual backbones enabled CoDet to achieve a novel category mask average precision of 37.0 and an overall precision of 44.7, surpassing previous leading approaches by 4.2 and 9.8 points, respectively. Across standard model sizes, the framework consistently outperformed rival methods relying on vision-language models or size-based heuristics. In zero-shot transfer evaluations across different datasets, CoDet achieved approximately a 2 percentage point improvement over the best competing systems on both COCO and Objects365. Furthermore, comparative analyses revealed that visual co-occurrence discovery generates substantially more accurate and stable training signals than direct text-to-region matching, which frequently degrades due to incorrect initial label assignments.

These results demonstrate that visual correspondence across images provides a scalable and cost-effective alternative to manual region annotations and imperfect vision-language teacher models. Organizations building large-scale perception systems can train open-vocabulary models more effectively by pairing scalable visual backbones with uncurated web-crawled image-text data. For deployment, teams should scale concept group sizes on diverse web data while exercising caution on heavily curated datasets where artificial concept co-occurrences can introduce visual noise. Confidence in these results is supported by consistent benchmark improvements, though future work should evaluate hybrid approaches combining co-occurrence mechanisms with vision-language models to further enhance detection performance.

Cover for CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection

Abstract

Deriving reliable region-word alignment from image-text pairs is critical to learn object-level vision-language representations for open-vocabulary object detection. Existing methods typically rely on pre-trained or self-trained vision-language models for alignment, which are prone to limitations in localization accuracy or generalization capabilities. In this paper, we propose CoDet, a novel approach that overcomes the reliance on pre-aligned vision-language space by reformulating region-word alignment as a co-occurring object discovery problem. Intuitively, by grouping images that mention a shared concept in their captions, objects corresponding to the shared concept shall exhibit high co-occurrence among the group. CoDet then leverages visual similarities to discover the co-occurring objects and align them with the shared concept. Extensive experiments demonstrate that CoDet has superior performances and compelling scalability in open-vocabulary detection, e.g., by scaling up the visual backbone, CoDet achieves 37.0 AP_novel^m and 44.7 AP_all^m on OV-LVIS, surpassing the previous SoTA by 4.2 AP_novel^m and 9.8 AP_all^m. Code is available at https://github.com/CVMI-Lab/CoDet.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Preliminaries
  • 3.2 Aligning Regions and Words by Co-occurrence
  • 3.3 Discovering Co-occurring Objects across Images
  • 3.4 Training and Inference
  • 4 Experiments
  • 4.1 Benchmark Setup
  • 4.2 Implementation Details
  • 4.3 Benchmark Results
  • 4.4 Transfer to Other Datasets
  • 4.5 Visualization and Analysis
  • 4.6 Ablation Study
  • 5 Limitations and Conclusions
  • References
  • A A Heuristic Baseline for Co-occurrence Discovery
  • B Further Analysis on Different Alignment Strategies
  • C Visualization on OV-LVIS and OV-COCO
  • D Implementation Details

Knowls

  1. Knowl 1 — CoDet Open-Vocabulary Object Detection Framework

    model/method

    CoDet is an open-vocabulary object detection (OVD) framework that learns object-level vision-language representations from paired image-caption data without relying on pre-aligned vision-language models (VLMs) or hand-crafted bounding-box priors. Standard two-stage detectors (such as Mask R-CNN or CenterNet2) are adapted for open-vocabulary detection by making localization heads class-agnostic and replacing fixed classification weight matrices with text embeddings generated by a pre-trained language model (e.g., the text encoder of CLIP).

    To discover region-word alignments in weakly supervised caption data, CoDet reformulates cross-modal alignment as a visual co-occurrence discovery problem. Given an unbounded vocabulary Copen\mathcal{C}^{\text{open}}, text captions are parsed to extract concept words that correspond to noun synsets under the 'object' hierarchy in WordNet. Image-text pairs mentioning a shared concept cc are grouped into a concept group Gc\mathcal{G}_c. During training, a mini-group containing one query image and mm support images is sampled from Gc\mathcal{G}_c. CoDet estimates inter-image region correspondences between region proposals of the query image and support images, guided by the semantic text embedding of concept cc. The identified co-occurring proposals in the query image are aggregated into a prototypical visual feature vector fpf_p, which is directly aligned with the concept's text embedding through classification loss.

  2. Knowl 2 — Text-Guided Region-Region Similarity Estimation

    equation

    To determine visual correspondences for a target concept cc across different images while filtering out concurrent distracting concepts and handling intra-category variance, CoDet re-weights pairwise region similarities using the text embedding of concept cc.

    Let fi,fj∈Rdf_i, f_j \in \mathbb{R}^d denote the visual feature vectors of two region proposals extracted from the penultimate layer of the detector's classification head, where dd is the feature dimensionality. Let wc∈Rdw_c \in \mathbb{R}^d be the normalized language embedding of the shared concept cc generated by a CLIP text encoder. The text-guided similarity sijs_{ij} between proposal ii and proposal jj is defined as:

    sij=wˉc⊤(fi∥fi∥2∘fj∥fj∥2)s_{ij} = \bar{w}_c^\top \left( \frac{f_i}{\|f_i\|_2} \circ \frac{f_j}{\|f_j\|_2} \right)

    where ∘\circ represents the element-wise Hadamard product, ∥⋅∥2\|\cdot\|_2 denotes the ℓ2\ell_2-norm, and wˉc\bar{w}_c is a dimension-wise importance weight vector computed by:

    wˉc=d∣wc∣∥wc∥2\bar{w}_c = \sqrt{d} \frac{|w_c|}{\|w_c\|_2}

    with ∣⋅∣|\cdot| taking the element-wise absolute value. Weighting the dimension-wise visual product by ∣wc∣|w_c| amplifies feature dimensions that carry significant semantic activation for concept cc, making cross-image region matching concept-aware.

  3. Knowl 3 — Prototype-Based Co-Occurring Object Discovery and Region-Word Loss

    equation

    Given a query image with nn region proposals and mm support images each with nn proposals (total mnmn support proposals) sampled from a concept group associated with concept cc, pairwise text-guided similarities form a similarity matrix S∈Rn×mnS \in \mathbb{R}^{n \times mn}.

    A two-layer multi-layer perceptron (MLP) Φ:Rmn→R\Phi: \mathbb{R}^{mn} \to \mathbb{R} estimates the co-occurrence probability distribution p∈Rnp \in \mathbb{R}^n over the nn query region proposals:

    p=softmaxn(Φ(S))p = \text{softmax}_n(\Phi(S))

    The prototypical visual representation fp∈Rdf_p \in \mathbb{R}^d of the co-occurring object in the query image is computed as the soft weighted sum of its proposal features fi∈Rdf_i \in \mathbb{R}^d:

    fp=∑i=1npifif_p = \sum_{i=1}^n p_i f_i

    The prototype fpf_p is aligned with the concept label cc using a binary cross-entropy (BCE) classification loss over the open vocabulary Copen\mathcal{C}^{\text{open}}:

    Lregion-word=LBCE(Wfp,c)=−log⁡σ(sc)−∑k≠clog⁡(1−σ(sk))\mathcal{L}_{\text{region-word}} = \mathcal{L}_{\text{BCE}}(W f_p, c) = -\log \sigma(s_c) - \sum_{k \neq c} \log(1 - \sigma(s_k))

    where WW is the classifier weight matrix containing text embeddings of all concepts in Copen\mathcal{C}^{\text{open}}, sk=Wkfps_k = W_k f_p denotes the predicted classification score for concept kk, and σ(⋅)\sigma(\cdot) is the sigmoid function.

  4. Knowl 4 — Multi-Task Training Objective and Loss Formulation for CoDet

    equation

    CoDet is trained jointly on detection annotations Ddet\mathcal{D}^{\text{det}} and weakly annotated image-caption data Dcap\mathcal{D}^{\text{cap}}. For an input image II, the overall training objective L(I)\mathcal{L}(I) is defined as:

    L(I)={Lrpn+Lreg+Lclsif I∈DdetLregion-word+Limage-textif I∈Dcap\mathcal{L}(I) = \begin{cases} \mathcal{L}_{\text{rpn}} + \mathcal{L}_{\text{reg}} + \mathcal{L}_{\text{cls}} & \text{if } I \in \mathcal{D}^{\text{det}} \\ \mathcal{L}_{\text{region-word}} + \mathcal{L}_{\text{image-text}} & \text{if } I \in \mathcal{D}^{\text{cap}} \end{cases}

    where Lrpn\mathcal{L}_{\text{rpn}}, Lreg\mathcal{L}_{\text{reg}}, and Lcls\mathcal{L}_{\text{cls}} are standard region proposal network, bounding-box regression, and classification losses on annotated base categories Cbase\mathcal{C}^{\text{base}}.

    For caption images, Lregion-word\mathcal{L}_{\text{region-word}} aligns discovered object prototypes with extracted concept words, while Limage-text\mathcal{L}_{\text{image-text}} enforces global image-caption alignment using an image-wide proposal feature and caption language embedding optimized via binary cross-entropy treating other captions in the mini-batch as negatives. During inference, CoDet operates as a standard two-stage object detector using pre-computed text embeddings as classifier weights without requiring cross-image computation.

  5. Knowl 5 — Open-Vocabulary Object Detection Performance on OV-LVIS Benchmark

    data/table

    CoDet was evaluated on the OV-LVIS benchmark (866 base classes, 337 novel/rare classes) using CC3M (2.8M image-text pairs) as the caption source under a strict open-vocabulary setting (novel categories unknown during training). Metrics reported are mask average precision on novel categories (APnovelm\text{AP}^m_{\text{novel}}), common categories (APcm\text{AP}^m_c), frequent categories (APfm\text{AP}^m_f), and all categories (APallm\text{AP}^m_{\text{all}}).

    Method Backbone Supervision Strict APnovelm\text{AP}^m_{\text{novel}} APcm\text{AP}^m_c APallm\text{AP}^m_{\text{all}}
    ViLD RN50-FPN CLIP ✓ 16.6 24.6 25.5
    RegionCLIP RN50-C4 Caption ✓ 17.1 27.4 28.2
    DetPro RN50-FPN CLIP ✓ 19.8 25.6 25.9
    OV-DETR RN50-C4 Caption × 17.4 25.0 26.6
    PromptDet RN50-FPN Caption × 19.0 18.5 21.4
    Detic RN50 Caption × 19.5 - 30.9
    F-VLM RN50-FPN CLIP ✓ 18.6 - 24.2
    VLDet RN50 Caption ✓ 21.7 29.8 30.1
    BARON RN50-FPN CLIP ✓ 22.6 27.6 27.6
    CoDet (Ours) RN50 Caption ✓ 23.4 30.0 30.7
    RegionCLIP R50x4 (87M) Caption ✓ 22.0 32.1 32.3
    Detic SwinB (88M) Caption × 23.9 40.2 38.4
    F-VLM R50x4 (87M) CLIP ✓ 26.3 - 28.5
    VLDet SwinB (88M) Caption ✓ 26.3 39.4 38.1
    CoDet (Ours) SwinB (88M) Caption ✓ 29.4 39.5 39.2
    F-VLM R50x64 (420M) CLIP ✓ 32.8 - 34.9
    CoDet (Ours) EVA02-L (304M) Caption ✓ 37.0 46.3 44.7

    CoDet outperforms existing caption-supervised and distillation-based methods across backbones. When scaling the visual backbone from ResNet50 to Swin-B and EVA02-L, CoDet exhibits strong scaling behavior, achieving 37.0 APnovelm37.0\ \text{AP}^m_{\text{novel}} and 44.7 APallm44.7\ \text{AP}^m_{\text{all}}, outperforming the previous state of the art at comparable model size by 4.2 APnovelm4.2\ \text{AP}^m_{\text{novel}}.

  6. Knowl 6 — Cross-Dataset Transfer Detection on COCO and Objects365

    data/table

    To evaluate open-world generalization across domain and vocabulary shifts, the CoDet model trained on OV-LVIS (LVIS base + CC3M) with a ResNet50 backbone was evaluated directly on COCO (80 categories) and Objects365 v1 (365 categories) without any fine-tuning by plugging in the category text embeddings.

    Method COCO Objects365
    AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75}
    Supervised (Oracle) 46.5 67.6 50.9 25.6 38.6 28.0
    ViLD 36.6 55.6 39.8 11.8 18.2 12.6
    DetPro 34.9 53.8 37.4 12.1 18.8 12.9
    F-VLM 32.5 53.1 34.6 11.9 19.2 12.6
    BARON 36.2 55.7 39.1 13.6 21.0 14.5
    CoDet (Ours) 39.1 57.0 42.3 14.2 20.5 15.3

    CoDet achieves 39.1 AP39.1\ \text{AP} on COCO and 14.2 AP14.2\ \text{AP} on Objects365, exceeding the second-best transfer method (ViLD) by 2.5 AP2.5\ \text{AP} on COCO and 2.4 AP2.4\ \text{AP} on Objects365, despite ViLD utilizing a 32×32\times longer training schedule.

  7. Knowl 7 — Ablation of Text Guidance and Prototype-Based Discovery in CoDet

    data/table

    Ablation experiments conducted on the OV-COCO benchmark evaluate the individual contributions of text guidance in similarity estimation and prototype-based region aggregation against a heuristic co-occurrence baseline.

    Component Variation AP50novel\text{AP}^{\text{novel}}_{50} AP50base\text{AP}^{\text{base}}_{50} AP50all\text{AP}^{\text{all}}_{50}
    (a) Text Guidance Without text guidance 26.6 52.4 45.7
    With text guidance 30.6 52.3 46.6
    (b) Discovery Strategy Heuristic nearest-neighbor (ReCo-style) 26.9 52.4 45.7
    Prototype-based soft aggregation 30.6 52.3 46.6

    Introducing text guidance to re-weight feature dimensions improves novel class performance by 4.0 AP50novel4.0\ \text{AP}_{50}^{\text{novel}} by filtering out concurrent distracting objects. Replacing the heuristic hard selection strategy (which selects a single proposal via max-mean-argmax nearest neighbor rules) with soft MLP prototype aggregation yields a 3.7 AP50novel3.7\ \text{AP}_{50}^{\text{novel}} gain due to noise robustness and multi-instance aggregation capability.

  8. Knowl 8 — Impact of Concept Group Size on Open-Vocabulary Detection

    data/table

    The effect of varying the mini-group size m+1∈{2,4,8}m+1 \in \{2, 4, 8\} during training was ablated on both human-curated caption data (OV-COCO) and uncurated web image-text pairs (OV-LVIS with CC3M):

    Group Size OV-COCO OV-LVIS
    AP50novel\text{AP}^{\text{novel}}_{50} AP50base\text{AP}^{\text{base}}_{50} AP50all\text{AP}^{\text{all}}_{50} APnovelm\text{AP}^m_{\text{novel}} APcm\text{AP}^m_c APallm\text{AP}^m_{\text{all}}
    2 30.6 52.3 46.6 21.9 30.3 30.7
    4 29.9 51.2 45.6 21.8 30.2 30.6
    8 29.1 50.9 45.2 22.7 30.3 30.7

    On web-crawled CC3M data (OV-LVIS), increasing the concept group size improves novel detection performance (21.9→22.7 APnovelm21.9 \to 22.7\ \text{AP}^m_{\text{novel}}) because larger support groups reduce visual ambiguity among co-occurring concepts. In contrast, on human-curated COCO captions, increasing group size degrades novel performance (30.6→29.1 AP50novel30.6 \to 29.1\ \text{AP}^{\text{novel}}_{50}) because the fixed 80 COCO categories introduce strong category co-occurrence bias (e.g., ~50% of images contain 'people' and ~10% contain 'car'), creating persistent concurrent hard negatives.

  9. Knowl 9 — Pseudo-Label Quality and Self-Training Stability of Region Alignment Strategies

    empirical result

    Comparing pseudo-label quality across training iterations reveals significant differences between region-region visual similarity, region-word similarity, and hand-crafted max-size priors on the OV-COCO validation set:

    1. Cover Rate: Defined as the percentage of pseudo-bounding boxes whose mean Intersection over Union (mIoU) with the closest ground-truth box exceeds 0.5. At 90k iterations, region-region alignment (CoDet) achieves a cover rate of approximately 65%, whereas region-word similarity reaches roughly 47% and max-size prior reaches roughly 46%.
    2. Training Stability: Region-region similarity exhibits monotonic, steady improvements throughout self-training. Conversely, direct region-word similarity alignment exhibits unstable performance and initial degradation in novel category AP50\text{AP}_{50}. This instability stems from a negative feedback loop: early mismatched region-word pairs (e.g., associating 'seagull' with a 'dove') pull incorrect visual-text pairs closer in embedding space, compounding subsequent pseudo-labeling errors.
  10. Knowl 10 — Open-Vocabulary Object Detection Performance on OV-COCO Benchmark

    data/table

    Performance comparison on the OV-COCO dataset (48 base categories, 17 novel categories) trained using COCO Captions. The primary metric is bounding box AP50\text{AP}_{50} on novel categories (AP50novel\text{AP}^{\text{novel}}_{50}), base categories (AP50base\text{AP}^{\text{base}}_{50}), and all categories (AP50all\text{AP}^{\text{all}}_{50}).

    Method AP50novel\text{AP}^{\text{novel}}_{50} AP50base\text{AP}^{\text{base}}_{50} AP50all\text{AP}^{\text{all}}_{50}
    OVR-CNN 22.8 46.0 39.9
    ViLD 27.6 59.5 51.3
    RegionCLIP 26.8 54.8 47.5
    Detic 27.8 47.1 42.0
    OV-DETR 29.4 61.0 52.7
    PB-OVD 29.1 44.4 40.4
    VLDet 32.0 50.6 45.8
    CoDet (Ours) 30.6 52.3 46.6

    CoDet achieves 30.6 AP50novel30.6\ \text{AP}^{\text{novel}}_{50} and 52.3 AP50base52.3\ \text{AP}^{\text{base}}_{50}, outperforming most prior methods. The gap between CoDet and VLDet on OV-COCO is attributed to the high co-occurrence concentration of COCO's 80 base concepts across human-annotated captions, which introduces frequent co-occurring distracting instances during cross-image matching.

Coverage note — None was omitted; all key architectural components, equations, empirical benchmarks (OV-LVIS, OV-COCO, transfer detection), ablations, and analysis of training stability were captured.

References

  1. 1.Yutong Bai, Xinlei Chen, Alexander Kirillov, Alan Yuille, and Alexander C Berg. Point-level region contrast for object detection pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16061–16070, 2022. 3, 4
  2. 2.Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran. Zero-shot object detection. In Proceedings of the European Conference on Computer Vision (ECCV), pages 384–400, 2018. 2, 6
  3. 3.Hakan Bilen, Marco Pedersoli, and Tinne Tuytelaars. Weakly supervised object detection with convex clustering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1081–1089, 2015. 3
  4. 4.Hakan Bilen and Andrea Vedaldi. Weakly supervised deep detection networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2846–2854, 2016. 3
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, pages 213–229. Springer, 2020. 1
  6. 6.Peixian Chen, Kekai Sheng, Mengdan Zhang, Yunhang Shen, Ke Li, and Chunhua Shen. Open vocabulary object detection with proposal mining and prediction equalization. arXiv preprint arXiv:2206.11134, 2022. 3
  7. 7.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015. 6
  8. 8.Jang Hyun Cho, Utkarsh Mall, Kavita Bala, and Bharath Hariharan. Picie: Unsupervised semantic segmentation using invariance and equivariance in clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16794–16804, 2021. 3
  9. 9.Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid. Weakly supervised object localization with multi-fold multiple instance learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(1):189–203, 2016. 3
  10. 10.Berkan Demirel, Ramazan Gokberk Cinbis, and Nazli Ikizler-Cinbis. Zero-shot object detection by hybrid region embedding. In British Machine Vision Conference (BMVC), 2018. 2
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255. IEEE, 2009. 3
  12. 12.Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li. Learning to prompt for open-vocabulary object detection with vision-language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14084–14093, 2022. 6, 7
  13. 13.Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva-02: A visual representation for neon genesis. arXiv preprint arXiv:2303.11331, 2023. 7
  14. 14.Chengjian Feng, Yujie Zhong, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, and Lin Ma. Promptdet: Towards open-vocabulary detection using uncurated images. In Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX, page 701–717, 2022. 3, 6, 7
  15. 15.Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li, Ran Xu, Wenhao Liu, and Caiming Xiong. Open vocabulary object detection with pseudo bounding-box labels. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part X, pages 266–282. Springer, 2022. 1, 7
  16. 16.Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D Cubuk, Quoc V Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2918–2928, 2021. 17
  17. 17.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020. 3
  18. 18.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. In International Conference on Learning Representations, 2022. 3, 6, 7
  19. 19.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5356–5364, 2019. 6
  20. 20.Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T. Freeman. Unsupervised semantic segmentation by distilling feature correspondences. In International Conference on Learning Representations, 2022. 3, 4
  21. 21.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, pages 2961–2969, 2017. 1, 3
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 6
  23. 23.Olivier J Hénaff, Skanda Koppula, Evan Shelhamer, Daniel Zoran, Andrew Jaegle, Andrew Zisserman, João Carreira, and Relja Arandjelovic. Object discovery and representation networks. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVII, pages 123–143, 2022. 3
  24. 24.Hanzhe Hu, Jinshi Cui, and Liwei Wang. Region-aware contrastive learning for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16291–16301, 2021. 3, 4
  25. 25.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021. 1, 3
  26. 26.Weicheng Kuo, Yin Cui, Xiuye Gu, AJ Piergiovanni, and Anelia Angelova. Open-vocabulary object detection upon frozen vision and language models. In The Eleventh International Conference on Learning Representations, 2023. 3, 6, 7, 16
  27. 27.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10965–10975, 2022. 1, 3
  28. 28.Chuang Lin, Peize Sun, Yi Jiang, Ping Luo, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan, and Jianfei Cai. Learning object-language alignments for open-vocabulary object detection. In The Eleventh International Conference on Learning Representations, 2023. 1, 3, 5, 6, 7
  29. 29.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014. 1, 6
  30. 30.Weide Liu, Chi Zhang, Guosheng Lin, and Fayao Liu. Crnet: Cross-reference networks for few-shot segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4165–4173, 2020. 3
  31. 31.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 7
  32. 32.George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995. 6
  33. 33.Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, and Neil Houlsby. Simple open-vocabulary object detection. In Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part X, page 728–755, 2022. 3
  34. 34.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, 2014. 2
  35. 35.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 1, 2, 3, 4, 5
  36. 36.Shafin Rahman, Salman Khan, and Nick Barnes. Transductive learning for zero-shot object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6082–6091, 2019. 2
  37. 37.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 779–788, 2016. 1
  38. 38.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7263–7271, 2017. 7
  39. 39.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems, 28, 2015. 1
  40. 40.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018. 6
  41. 41.Gyungin Shin, Weidi Xie, and Samuel Albanie. Reco: Retrieve and co-segment for zero-shot transfer. In Advances in Neural Information Processing Systems, volume 35, pages 33754–33767, 2022. 3, 5, 9, 15
  42. 42.Zhao Shizhen, Gao Changxin, Shao Yuanjie, Li Lerenhan, Yu Changqian, Ji Zhong, and Sang Nong. Gtnet: Generative transfer network for zero-shot object detection. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2020. 3
  43. 43.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al. Sparse r-cnn: End-to-end object detection with learnable proposals. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14454–14463, 2021. 1
  44. 44.Sara Vicente, Vladimir Kolmogorov, and Carsten Rother. Cosegmentation revisited: Models and optimization. In Computer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part II 11, pages 465–479. Springer, 2010. 3
  45. 45.Sara Vicente, Carsten Rother, and Vladimir Kolmogorov. Object cosegmentation. In CVPR 2011, pages 2217–2224. IEEE, 2011. 3
  46. 46.Fang Wan, Chang Liu, Wei Ke, Xiangyang Ji, Jianbin Jiao, and Qixiang Ye. C-mil: Continuation multiple instance learning for weakly supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2199–2208, 2019. 3
  47. 47.Wenguan Wang, Tianfei Zhou, Fisher Yu, Jifeng Dai, Ender Konukoglu, and Luc Van Gool. Exploring cross-image pixel contrast for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7303–7313, 2021. 3, 4
  48. 48.Fangyun Wei, Yue Gao, Zhirong Wu, Han Hu, and Stephen Lin. Aligning pretraining for detection via object-level contrastive learning. Advances in Neural Information Processing Systems, 34:22682–22694, 2021. 7
  49. 49.Xin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang, and Xiaojuan Qi. Self-supervised visual representation learning with semantic grouping. In Advances in Neural Information Processing Systems, volume 35, pages 16423–16438, 2022. 3
  50. 50.Size Wu, Wenwei Zhang, Sheng Jin, Wentao Liu, and Chen Change Loy. Aligning bag of regions for open-vocabulary object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15254–15264, 2023. 3, 6, 7, 16
  51. 51.Bin Yan, Yi Jiang, Jiannan Wu, Dong Wang, Ping Luo, Zehuan Yuan, and Huchuan Lu. Universal instance perception as object discovery and retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. 1
  52. 52.Lewei Yao, Jianhua Han, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, and Hang Xu. Detclipv2: Scalable open-vocabulary object detection pre-training via word-region alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23497–23506, 2023. 3
  53. 53.Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, and Hang Xu. Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection. In Advances in Neural Information Processing Systems, volume 35, pages 9125–9138, 2022. 1
  54. 54.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. Transactions on Machine Learning Research, 2022. 3
  55. 55.Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021. 3
  56. 56.Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. Open-vocabulary detr with conditional matching. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX, pages 106–122. Springer, 2022. 3, 6, 7
  57. 57.Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. Open-vocabulary object detection using captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14393–14402, 2021. 1, 3, 6, 7
  58. 58.Feihu Zhang, Philip Torr, René Ranftl, and Stephan Richter. Looking beyond single images for contrastive semantic segmentation learning. In Advances in Neural Information Processing Systems, volume 34, pages 3285–3297, 2021. 3, 4
  59. 59.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. Regionclip: Region-based language-image pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16793–16803, 2022. 1, 3, 6, 7, 16
  60. 60.Tianfei Zhou, Wenguan Wang, Ender Konukoglu, and Luc Van Gool. Rethinking semantic segmentation: A prototype view. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2582–2593, 2022. 3, 4
  61. 61.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX, pages 350–368. Springer, 2022. 2, 3, 5, 6, 7, 8, 17
  62. 62.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Probabilistic two-stage detection. arXiv preprint arXiv:2103.07461, 2021. 6, 17
  63. 63.Pengkai Zhu, Hanxiao Wang, and Venkatesh Saligrama. Don’t even look once: Synthesizing features for zero-shot detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11693–11702, 2020. 3
  64. 64.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In International Conference on Learning Representations, 2021. 7

Citation

MLA
Ma, C., et al. “CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 71078–94, https://proceedings.neurips.cc/paper_files/paper/2023/file/e10a6a906ef323efaf708f76cf3c1d1e-Paper-Conference.pdf.
APA
Ma, C., Jiang, Y., Wen, X., Yuan, Z., & Qi, X. (2023). CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection. Advances in Neural Information Processing Systems, 36, 71078–71094. https://proceedings.neurips.cc/paper_files/paper/2023/file/e10a6a906ef323efaf708f76cf3c1d1e-Paper-Conference.pdf
Chicago
Ma, C., Y. Jiang, X. Wen, Z. Yuan, and X. Qi. 2023. “CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection”. Advances in Neural Information Processing Systems 36: 71078–94. https://proceedings.neurips.cc/paper_files/paper/2023/file/e10a6a906ef323efaf708f76cf3c1d1e-Paper-Conference.pdf.
Harvard
Ma, C. et al. (2023) “CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 71078–71094. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/e10a6a906ef323efaf708f76cf3c1d1e-Paper-Conference.pdf.
Vancouver
1. Ma C, Jiang Y, Wen X, Yuan Z, Qi X (2023) CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 71078–71094

BibTeX

@inproceedings{ma2023codet,
  title = {CoDet: Co-occurrence Guided Region-Word Alignment for Open-Vocabulary Object Detection},
  author = {Ma, Chuofan and Jiang, Yi and Wen, Xin and Yuan, Zehuan and Qi, Xiaojuan},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {71078-71094},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/e10a6a906ef323efaf708f76cf3c1d1e-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors