Aligning Bag of Regions for Open-Vocabulary Object Detection

Size WuWenwei ZhangSheng JinWentao LiuChen Change Loy

article2023CVPR170 citations

Proposes BARON, an open-vocabulary object detection framework that models contextual relations among multiple regions by encoding them as pseudo-word sentences for alignment with vision-language models, significantly boosting novel-category detection performance on COCO and LVIS benchmarks.

Listen

Traditional computer vision systems for object detection are limited to recognizing only the specific categories labeled during training, which restricts their deployment in complex, real-world environments. To overcome this limitation, open-vocabulary object detection leverages pre-trained vision-language models to recognize novel, unseen categories without requiring exhaustive manual annotations. However, existing approaches extract and align features from isolated image regions independently, failing to exploit the broader scene context and co-occurrence of multiple visual concepts that pre-trained vision-language models naturally learn from massive paired datasets.

The article evaluates and demonstrates a novel framework called BARON, which aligns groups of interrelated image regions rather than individual regions alone. The main objective is to establish whether grouping neighboring visual concepts and modeling them as contextual "bags" significantly enhances an open-vocabulary detector's ability to identify novel categories.

To achieve this, the authors implemented BARON on top of the standard Faster R-CNN detection architecture. The framework samples candidate bounding boxes around initial region proposals to form coherent groups of neighboring visual regions with balanced sizes. These regional features are projected into word embedding representations, combined with spatial positional data indicating relative size and position, and processed through a frozen text encoder from a pre-trained vision-language model. This composite embedding is then aligned with cropped visual representations from the vision-language model's image encoder using contrastive learning. The evaluation was conducted across major open-vocabulary benchmarks, specifically Common Objects in Context and Large Vocabulary Instance Segmentation, and tested for cross-dataset transfer.

The experimental findings show substantial performance improvements over existing state-of-the-art methods. On novel categories within the Common Objects in Context benchmark, BARON achieved a 34.0 average precision score, surpassing the previous leading approach by 4.6 points. On the Large Vocabulary Instance Segmentation benchmark, the framework improved novel mask average precision by 2.8 points over existing baselines. In addition, when trained with image caption supervision, the framework reached an average precision of 33.1 on novel categories, outperforming existing caption-supervised methods. Ablation analyses confirmed that incorporating spatial positional embeddings provides a critical boost of 7.1 points over unpositioned groupings, and transferring the trained model to external datasets consistently beat prior benchmarks.

These findings indicate that exploiting contextual relationships and visual co-occurrences significantly improves zero-shot recognition capabilities without requiring expensive new annotations or heavier base architectures. By effectively utilizing pre-trained foundation models, organizations can reduce the data labeling costs and operational overhead needed to deploy adaptable detection systems in dynamic environments.

Based on these results, engineering teams developing open-vocabulary visual detection systems should adopt contextual grouping strategies and preserve spatial relationships during distillation. Practitioners can also leverage caption-based supervision as a viable alternative when fine-grained visual teachers are unavailable. However, the study's scope is primarily focused on object co-occurrence structures, and the authors note that modeling more intricate compositional language structures remains an open challenge. The results provide high confidence within standard evaluation benchmarks, though deploying to domains with vastly different spatial distributions may require further validation.

No sufficiently relevant recommendations were found.

Cover for Aligning Bag of Regions for Open-Vocabulary Object Detection

Abstract

Pre-trained vision-language models (VLMs) learn to align vision and language representations on large-scale datasets, where each image-text pair usually contains a bag of semantic concepts. However, existing open-vocabulary object detectors only align region embeddings individually with the corresponding features extracted from the VLMs. Such a design leaves the compositional structure of semantic concepts in a scene under-exploited, although the structure may be implicitly learned by the VLMs. In this work, we propose to align the embedding of bag of regions beyond individual regions. The proposed method groups contextually interrelated regions as a bag. The embeddings of regions in a bag are treated as embeddings of words in a sentence, and they are sent to the text encoder of a VLM to obtain the bag-of-regions embedding, which is learned to be aligned to the corresponding features extracted by a frozen VLM. Applied to the commonly used Faster R-CNN, our approach surpasses the previous best results by 4.6 box AP50 and 2.8 mask AP on novel categories of open-vocabulary COCO and LVIS benchmarks, respectively. Code and models are available at https://github.com/wusize/ovdet.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Preliminaries
  • 3.2 Forming Bag of Regions
  • 3.3 Representing Bag of Regions
  • 3.4 Aligning Bag of Regions
  • 3.5 Caption Supervision
  • 4 Experiments
  • 4.1 Benchmark Results
  • 4.2 Ablation Study
  • 4.3 Further Analysis
  • 5 Discussion and Conclusion
  • A1 Implementation Details
  • A2 Sampling Strategy
  • A3 Pseudo Word Encoding
  • A4 Image-Guided Inference
  • A5 Detection Results
  • A6 Potential Negative Societal Impacts
  • References

Knowls

  1. Knowl 1 — BARON Framework for Open-Vocabulary Object Detection

    model/method

    Aligning Bag of Regions (BARON) is an open-vocabulary object detection framework that transfers compositional visual-semantic knowledge from pre-trained Vision-Language Models (VLMs) by aligning representations of groups of regions rather than individual regions in isolation.

    BARON is instantiated on a two-stage detector (e.g., Faster R-CNN). The standard fixed-class classifier head is replaced by a projection layer that maps region features into multiple pseudo-word embeddings in the text embedding space of a frozen VLM. To model context and co-occurring semantic concepts (such as objects and scene context), BARON samples contextually interrelated bounding boxes around region proposals to form a 'bag of regions'.

    The pseudo-words of all regions in a bag, combined with spatial positional embeddings, are concatenated as a sequence and fed into the frozen VLM text encoder T\mathcal{T} to generate a student bag-of-regions embedding ftf_t. Concurrently, the cropped image bounding box that encloses the entire bag of regions is fed into the frozen VLM image encoder V\mathcal{V} (with redundant contents outside the grouped regions masked in attention layers) to extract a teacher bag-of-regions embedding fvf_v. BARON aligns ftf_t and fvf_v via contrastive learning, enabling the detector to learn novel categories from multi-concept compositional contexts.

  2. Knowl 2 — Neighborhood Sampling Strategy for Forming Bags of Regions

    algorithm

    To construct a bag of regions that captures co-occurring objects while preventing the teacher image encoder from being distracted by redundant background or dominated by scale imbalances, BARON applies a neighborhood sampling procedure to region proposal network (RPN) candidates.

    Input: Region proposals {bi}i=1M\{b_i\}_{i=1}^M, where each bi=(xi,yi,Wi,Hi)b_i = (x_i, y_i, W_i, H_i), base sampling probability pb=0.3p_b = 0.3, aspect ratio scaling factor α=3.0\alpha = 3.0, Intersection over Foreground threshold IOF=0.1\mathrm{IOF} = 0.1, number of groups per proposal G=3G = 3
    Output: Sampled bags of regions {Bi,g}i=1,g=1M,G\{\mathcal{B}_{i, g}\}_{i=1, g=1}^{M, G}
    for each region proposal bib_i do
        Generate 8 neighboring candidate boxes surrounding bib_i with identical shape (Wi,Hi)(W_i, H_i) and overlapping with bib_i at IOF=0.1\mathrm{IOF} = 0.1
        Discard any candidate box with more than 2/32/3 of its area outside the image boundary
        Calculate sampling probability for left and right candidates: plr=pb⋅min⁡((Hi/Wi)α,1)p_{lr} = p_b \cdot \min\left( (H_i / W_i)^\alpha, 1 \right)
        Calculate sampling probability for top and bottom candidates: ptb=pb⋅min⁡((Wi/Hi)α,1)p_{tb} = p_b \cdot \min\left( (W_i / H_i)^\alpha, 1 \right)
        for g=1g = 1 to GG do
            Sample valid candidate boxes independently according to their adjusted probabilities
            Form bag Bi,g={bi}∪{sampled neighbor boxes}\mathcal{B}_{i, g} = \{b_i\} \cup \{\text{sampled neighbor boxes}\}
        end for
    end for
    return {Bi,g}\{\mathcal{B}_{i, g}\}
  3. Knowl 3 — Bag-of-Regions Representation with Positional Embeddings

    model/method

    Let a sampled bag of regions indexed by ii contain NiN_i regions {bji}j=0Ni−1\{b_j^i\}_{j=0}^{N_i-1}. For each region bjib_j^i, a linear projection layer extracts a sequence of pseudo-word vectors wji∈Rdw_j^i \in \mathbb{R}^d.

    To preserve the spatial relationship (relative box center coordinates and shapes) among regions in the bag—which provides sentence-like structure for the text encoder—BARON adds a spatial positional embedding pji∈Rdp_j^i \in \mathbb{R}^d to each region's pseudo-word vector wjiw_j^i. The student bag-of-regions embedding ftif_t^i is obtained by passing the concatenated sequence to the frozen VLM text encoder T\mathcal{T}:

    fti=T(w0i+p0i, w1i+p1i, …, wNi−1i+pNi−1i)f_t^i = \mathcal{T}\left(w_0^i + p_0^i, \, w_1^i + p_1^i, \, \dots, \, w_{N_i-1}^i + p_{N_i-1}^i\right)

    The teacher bag-of-regions embedding fvif_v^i is produced by feeding the minimal image crop bounding all regions {bji}j=0Ni−1\{b_j^i\}_{j=0}^{N_i-1} to the frozen VLM image encoder V\mathcal{V}:

    fvi=V(b0i,b1i,…,bNi−1i)f_v^i = \mathcal{V}\left(b_0^i, b_1^i, \dots, b_{N_i-1}^i\right)

    where image regions inside the crop that do not belong to any region bjib_j^i in the bag are masked out in the self-attention layers of V\mathcal{V}.

  4. Knowl 4 — Contrastive Bag Alignment Objective and Memory Queues

    equation

    Given GG sampled bags of regions in an image, BARON optimizes the mutual alignment between student bag embeddings {ftk}k=0G−1\{f_t^k\}_{k=0}^{G-1} and teacher bag embeddings {fvk}k=0G−1\{f_v^k\}_{k=0}^{G-1} using a symmetric InfoNCE loss:

    Lbag=−12∑k=0G−1(log⁡(pt,vk)+log⁡(pv,tk))\mathcal{L}_\mathrm{bag} = -\frac{1}{2} \sum_{k=0}^{G-1} \left( \log\left(p_{t, v}^k\right) + \log\left(p_{v, t}^k\right) \right)

    where the conditional alignment probabilities are defined as:

    pt,vk=exp⁡(τ′⋅⟨ftk,fvk⟩)∑l=0G−1exp⁡(τ′⋅⟨ftk,fvl⟩)p_{t, v}^k = \frac{\exp\left(\tau' \cdot \langle f_t^k, f_v^k \rangle\right)}{\sum_{l=0}^{G-1} \exp\left(\tau' \cdot \langle f_t^k, f_v^l \rangle\right)}

    pv,tk=exp⁡(τ′⋅⟨fvk,ftk⟩)∑l=0G−1exp⁡(τ′⋅⟨fvk,ftl⟩)p_{v, t}^k = \frac{\exp\left(\tau' \cdot \langle f_v^k, f_t^k \rangle\right)}{\sum_{l=0}^{G-1} \exp\left(\tau' \cdot \langle f_v^k, f_t^l \rangle\right)}

    Here, ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle denotes cosine similarity, and τ′\tau' is a learnable temperature parameter. To ensure a sufficient number of negative pairs when GG is small per image, the denominator sums over embeddings drawn from both the current batch and memory queues that store student and teacher bag embeddings from prior training iterations.

  5. Knowl 5 — Open-Vocabulary Region Classification via Pseudo-Words

    equation

    For an object proposal with projected pseudo-words ww, classification scores across CC candidate object categories are computed using category text embeddings generated by the frozen VLM text encoder T\mathcal{T}. For category c∈{0,…,C−1}c \in \{0, \dots, C-1\}, the category embedding fcf_c is obtained by passing the prompt string (e.g., 'a photo of {category} in the scene') to T\mathcal{T}.

    The predicted probability pcp_c of the region belonging to category cc is:

    pc=exp⁡(τ⋅⟨T(w),fc⟩)∑i=0C−1exp⁡(τ⋅⟨T(w),fi⟩)p_c = \frac{\exp\left(\tau \cdot \langle \mathcal{T}(w), f_c \rangle\right)}{\sum_{i=0}^{C-1} \exp\left(\tau \cdot \langle \mathcal{T}(w), f_i \rangle\right)}

    where ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle is cosine similarity and τ\tau is a temperature hyperparameter to scale logit magnitudes.

  6. Knowl 6 — Bag-of-Regions Distillation under Caption Supervision

    model/method

    BARON can be adapted to train directly from weak image-caption supervision instead of pre-trained VLM image encoder features.

    In this configuration, the teacher image crop feature fvf_v is replaced by the text embedding of the image caption fcap=T(caption)f_\mathrm{cap} = \mathcal{T}(\text{caption}). A bag of regions is formed by randomly sampling region proposals generated by the RPN. Because images may have multiple ground-truth captions and there is no fine-grained grounding between individual pseudo-words and caption words, the student bag-of-regions embedding is aligned to the set of caption embeddings using a soft cross-entropy loss (following UniCL). Individual region distillation Lindividual\mathcal{L}_\mathrm{individual} is omitted in this mode.

  7. Knowl 7 — Open-Vocabulary Object Detection Benchmark on OV-COCO

    data/table

    Evaluation on the open-vocabulary COCO benchmark (OV-COCO), split into 48 base categories and 17 novel categories. Metrics reported are box AP50\mathrm{AP}_{50} on novel categories (AP50novel\mathrm{AP}_{50}^\mathrm{novel}), base categories (AP50base\mathrm{AP}_{50}^\mathrm{base}), and all categories (AP50\mathrm{AP}_{50}).

    Method Supervision AP50novel\mathrm{AP}_{50}^\mathrm{novel} AP50base\mathrm{AP}_{50}^\mathrm{base} AP50\mathrm{AP}_{50}
    ViLD CLIP 27.6 59.5 51.2
    OV-DETR CLIP 29.4 61.0 52.7
    BARON (Ours) CLIP 34.0 60.4 53.5
    OVR-CNN Caption 22.8 46.0 39.9
    RegionCLIP Caption 26.8 54.8 47.5
    Detic Caption 27.8 51.1 45.0
    PB-OVD Caption 30.8 46.1 42.1
    VLDet Caption 32.0 50.6 45.8
    BARON (Ours) Caption 33.1 54.8 49.1
    Rasheed et al. CLIP + Caption 36.6 54.0 49.4
    BARON (Ours)†^\dagger CLIP + Caption 42.7 54.9 51.7

    †^{\dagger} denotes using proposals generated by MAVL. Under CLIP distillation, BARON outperforms OV-DETR by +4.6 novel AP50\mathrm{AP}_{50}. Under caption supervision, BARON exceeds VLDet and PB-OVD without requiring explicit pseudo-labeling.

  8. Knowl 8 — Open-Vocabulary Detection and Instance Segmentation on OV-LVIS

    data/table

    Performance comparison on the OV-LVIS benchmark (337 rare categories treated as novel, common and frequent categories treated as base). Evaluation includes bounding box and mask mean Average Precision across rare (APr\mathrm{AP}_r), common (APc\mathrm{AP}_c), frequent (APf\mathrm{AP}_f), and all categories (AP\mathrm{AP}).

    Object Detection Instance Segmentation
    Method Ensemble Learned Prompt APr\mathrm{AP}_r APc\mathrm{AP}_c APf\mathrm{AP}_f AP\mathrm{AP} APr\mathrm{AP}_r APc\mathrm{AP}_c APf\mathrm{AP}_f AP\mathrm{AP}
    ViLD - - 16.3 21.2 31.6 24.4 16.1 20.0 28.3 22.5
    OV-DETR - - - - - - 17.4 25.0 32.5 26.6
    BARON (Ours) - - 17.3 25.6 31.0 26.3 18.0 24.4 28.9 25.1
    ViLD ✓ - 16.7 26.5 34.2 27.8 16.6 24.6 30.3 25.5
    ViLD* ✓ - 17.4 27.5 31.9 27.5 16.8 25.6 28.5 25.2
    BARON (Ours) ✓ - 20.1 28.4 32.2 28.4 19.2 26.8 29.4 26.5
    DetPro ✓ ✓ 20.8 27.8 32.4 28.4 19.8 25.6 28.9 25.9
    BARON (Ours) ✓ ✓ 23.2 29.3 32.5 29.5 22.6 27.6 29.8 27.6

    ViLD* denotes the re-implemented ViLD in DetPro using SOCO pre-training on a 2×2\times schedule. When combined with score ensembling and learned prompt tuning, BARON reaches 22.622.6 novel mask APr\mathrm{AP}_r (+2.8 over DetPro).

  9. Knowl 9 — Cross-Dataset Zero-Shot Transfer of OV-LVIS Trained Models

    data/table

    Transfer performance of object detection models trained on OV-LVIS to Pascal VOC 2007 test, COCO validation, and Objects365 v2 validation sets without fine-tuning.

    Pascal VOC COCO Objects365
    Method AP50\mathrm{AP}_{50} AP75\mathrm{AP}_{75} AP\mathrm{AP} AP50\mathrm{AP}_{50} AP75\mathrm{AP}_{75} APs\mathrm{AP}_s APm\mathrm{AP}_m AP\mathrm{AP} AP50\mathrm{AP}_{50} AP75\mathrm{AP}_{75} APs\mathrm{AP}_s
    ViLD* 73.9 57.9 34.1 52.3 36.5 21.6 38.9 11.5 17.8 12.3 4.2
    BARON (Ours)‡^\ddagger 74.5 57.9 36.3 56.1 39.3 25.4 39.5 13.2 20.0 14.0 4.8
    DetPro 74.6 57.9 34.9 53.8 37.4 22.5 39.6 12.1 18.8 12.9 4.5
    BARON (Ours) 76.0 58.2 36.2 55.7 39.1 24.8 40.2 13.6 21.0 14.5 5.0

    ViLD* is the DetPro re-implementation. ‡^\ddagger denotes evaluation with hand-crafted prompts for a direct comparison with ViLD*. BARON demonstrates higher zero-shot generalization across all three datasets.

  10. Knowl 10 — Ablation of Bag Alignment Components and Positional Embeddings

    empirical result

    An ablation study on the OV-COCO benchmark isolates the contributions of individual region distillation (Lindividual\mathcal{L}_\mathrm{individual}), bag-of-regions distillation (Lbag\mathcal{L}_\mathrm{bag}), and region spatial Positional Embeddings (PE):

    # Lindividual\mathcal{L}_\mathrm{individual} Lbag\mathcal{L}_\mathrm{bag} PE AP50novel\mathrm{AP}_{50}^\mathrm{novel} AP50base\mathrm{AP}_{50}^\mathrm{base} AP50\mathrm{AP}_{50}
    1 ✓ - - 25.7 59.6 50.6
    2 - ✓ - 25.7 59.4 50.5
    3 - ✓ ✓ 32.8 60.1 53.0
    4 ✓ ✓ ✓ 34.0 60.4 53.5

    Aligning bag embeddings without spatial positional embeddings yields 25.7 AP50novel25.7\ \mathrm{AP}_{50}^\mathrm{novel}, identical to aligning individual regions alone. Incorporating positional embeddings causes novel AP to surge by +7.1+7.1 (from 25.725.7 to 32.832.8), showing that spatial organization is critical for the text encoder to model visual composition. Combining bag alignment with individual-level distillation yields an additional +1.2 AP50novel+1.2\ \mathrm{AP}_{50}^\mathrm{novel} gain (34.034.0 total), showing their complementarity.

  11. Knowl 11 — Impact of Sampling Strategies and Hyperparameters on Bag Alignment

    empirical result

    Ablations on OV-COCO evaluate region sampling strategies, box overlap, sampling probability pbp_b, bag counts, and pseudo-word length:

    1. Sampling Strategy: Neighborhood sampling achieves 34.0 AP50novel34.0\ \mathrm{AP}_{50}^\mathrm{novel} (and 32.232.2 in a reduced-budget version with 12 proposals per image), outperforming regular grid sampling (25.425.4) and random proposal sampling (27.327.3).
    2. Intersection over Foreground (IOF): Box overlap of IOF=0.1\mathrm{IOF} = 0.1 achieves the best performance (34.0 AP50novel34.0\ \mathrm{AP}_{50}^\mathrm{novel}), outperforming negative spacing IOF=−0.1\mathrm{IOF}=-0.1 (32.532.5), zero overlap IOF=0.0\mathrm{IOF}=0.0 (33.633.6), IOF=0.2\mathrm{IOF}=0.2 (33.833.8), and IOF=0.3\mathrm{IOF}=0.3 (33.733.7).
    3. Candidate Sampling Probability pbp_b: pb=0.3p_b = 0.3 gives 34.0 AP50novel34.0\ \mathrm{AP}_{50}^\mathrm{novel} compared to 33.733.7 (pb=0.1p_b=0.1) and 33.233.2 (pb=0.5p_b=0.5).
    4. Bags per Proposal: Sampling G=3G=3 bags per proposal produces 34.0 AP50novel34.0\ \mathrm{AP}_{50}^\mathrm{novel} compared to 32.632.6 (1 bag) and 33.233.2 (5 bags).
    5. Pseudo-Words per Region: Projecting each region to 6 pseudo-words achieves optimal performance (34.0 AP50novel34.0\ \mathrm{AP}_{50}^\mathrm{novel}), improving over 2 words (31.631.6) and 4 words (33.133.1), while 8 words degrades performance to 33.533.5.

Coverage note — None was omitted; all contributed methodology, formal equations, algorithmic steps, benchmark datasets, cross-dataset transfers, and ablation analyses are covered.

References

  1. 1.Edward H. Adelson. On seeing stuff: the perception of materials by humans and machines. In Bernice E. Rogowitz and Thrasyvoulos N. Pappas, editors, Human Vision and Electronic Imaging VI, SPIE Proceedings, pages 1–12, 2001.
  2. 2.Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran. Zero-shot object detection. In Eur. Conf. Comput. Vis., 2018.
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In Eur. Conf. Comput. Vis., 2020.
  4. 4.Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth N Balasubramanian. Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks. In IEEE Winter App. Comput. Vis., 2018.
  5. 5.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. Hybrid task cascade for instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  6. 6.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  7. 7.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  8. 8.Berkan Demirel, Ramazan Gokberk Cinbis, and Nazli Ikizler-Cinbis. Zero-shot object detection by hybrid region embedding. In Brit. Mach. Vis. Conf., 2018.
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In Int. Conf. Learn. Represent., 2021.
  10. 10.Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li. Learning to prompt for open-vocabulary object detection with vision-language model. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  11. 11.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. Int. J. Comput. Vis., 88(2):303–338, 2010.
  12. 12.Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov. Devise: A deep visual-semantic embedding model. Adv. Neural Inform. Process. Syst., 2013.
  13. 13.Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li, Ran Xu, Wenhao Liu, and Caiming Xiong. Open vocabulary object detection with pseudo bounding-box labels. Eur. Conf. Comput. Vis., 2022.
  14. 14.Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D. Cubuk, Quoc V. Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  15. 15.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. In Int. Conf. Learn. Represent., 2021.
  16. 16.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 5356–5364, 2019.
  17. 17.Nasir Hayat, Munawar Hayat, Shafin Rahman, Salman H. Khan, Syed Waqas Zamir, and Fahad Shahbaz Khan. Synthesizing the unseen for zero-shot object detection. In Asian Conf. Comput. Vis., 2020.
  18. 18.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In IEEE Conf. Comput. Vis. Pattern Recog., 2020.
  19. 19.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask R-CNN. In Int. Conf. Comput. Vis., 2017.
  20. 20.Dinesh Jayaraman and Kristen Grauman. Zero-shot recognition with unreliable attributes. In Adv. Neural Inform. Process. Syst., 2014.
  21. 21.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In Int. Conf. Mach. Learn., 2021.
  22. 22.Wonjae Kim, Bokyung Son, and Ildoo Kim. ViLT: Vision-and-language transformer without convolution or region supervision. In Int. Conf. Mach. Learn., 2021.
  23. 23.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollar. Panoptic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 9404–9413, 2019.
  24. 24.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation. Int. Conf. Learn. Represent., 2022.
  25. 25.Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Int. Conf. Mach. Learn., 2022.
  26. 26.Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. In Adv. Neural Inform. Process. Syst., 2021.
  27. 27.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jianfeng Gao. Grounded language-image pre-training. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  28. 28.Chuang Lin, Peize Sun, Yi Jiang, Ping Luo, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan, and Jianfei Cai. Learning object-language alignments for open-vocabulary object detection. Int. Conf. Learn. Represent., 2023.
  29. 29.Tsung-Yi Lin, Piotr Dollar, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. Feature pyramid networks for object detection. In Int. Conf. Comput. Vis., 2017.
  30. 30.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In Int. Conf. Comput. Vis., 2017.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Eur. Conf. Comput. Vis., 2014.
  32. 32.Li Liu, Wanli Ouyang, Xiaogang Wang, Paul W. Fieguth, Jie Chen, Xinwang Liu, and Matti Pietikainen. Deep learning for generic object detection: A survey. Int. J. Comput. Vis., 2020.
  33. 33.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Adv. Neural Inform. Process. Syst., 2019.
  34. 34.Muhammad Maaz, Hanoona Rasheed, Salman Khan, Fahad Shahbaz Khan, Rao Muhammad Anwer, and Ming-Hsuan Yang. Class-agnostic object detection with multi-modal transformer. In Eur. Conf. Comput. Vis., pages 512–531, 2022.
  35. 35.Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg Corrado, and Jeffrey Dean. Zero-shot learning by convex combination of semantic embeddings. In Int. Conf. Learn. Represent., 2014.
  36. 36.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  37. 37.Chao Peng, Tete Xiao, Zeming Li, Yuning Jiang, Xiangyu Zhang, Kai Jia, Gang Yu, and Jian Sun. MegDet: A large mini-batch object detector. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In Int. Conf. Mach. Learn., 2021.
  39. 39.Shafin Rahman, Salman H. Khan, and Nick Barnes. Transductive learning for zero-shot object detection. In Int. Conf. Comput. Vis., 2019.
  40. 40.Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. DenseCLIP: Language-guided dense prediction with context-aware prompting. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  41. 41.Hanoona Abdul Rasheed, Muhammad Maaz, Muhammd Uzair Khattak, Salman Khan, and Fahad Khan. Bridging the gap between object and image-level representations for open-vocabulary detection. In Adv. Neural Inform. Process. Syst., 2022.
  42. 42.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In Adv. Neural Inform. Process. Syst., 2015.
  43. 43.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In Int. Conf. Comput. Vis., 2019.
  44. 44.Hengcan Shi, Munawar Hayat, Yicheng Wu, and Jianfei Cai. ProposalCLIP: Unsupervised open-category object proposal generation via exploiting clip cues. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  45. 45.Ajinkya Tejankar, Maziar Sanjabi, Bichen Wu, Saining Xie, Madian Khabsa, Hamed Pirsiavash, and Hamed Firooz. A fistful of words: Learning transferable visual models from bag-of-words supervision. CoRR, abs/2112.13884, 2021.
  46. 46.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In Eur. Conf. Comput. Vis., 2020.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Adv. Neural Inform. Process. Syst., 2017.
  48. 48.Fangyun Wei, Yue Gao, Zhirong Wu, Han Hu, and Stephen Lin. Aligning pretraining for detection via object-level contrastive learning. Adv. Neural Inform. Process. Syst., 2021.
  49. 49.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019.
  50. 50.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao, Ce Liu, Lu Yuan, and Jianfeng Gao. Unified contrastive learning in image-text-label space. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  51. 51.Mert Yüksekgönül, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and why vision-language models behave like bags-of-words, and what to do about it? CoRR, abs/2210.01936, 2022.
  52. 52.Yuhang Zang, Wei Li, Kaiyang Zhou, Chen Huang, and Chen Change Loy. Open-vocabulary DETR with conditional matching. Eur. Conf. Comput. Vis., 2022.
  53. 53.Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. Open-vocabulary object detection using captions. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  54. 54.Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. LiT: Zero-shot transfer with locked-image text tuning. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  55. 55.Ye Zheng, Ruoran Huang, Chuanqi Han, Xi Huang, and Li Cui. Background learnable cascade for zero-shot object detection. In Asian Conf. Comput. Vis., 2020.
  56. 56.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. RegionCLIP: Region-based language-image pretraining. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  57. 57.Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In Eur. Conf. Comput. Vis., 2022.
  58. 58.Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krähenbühl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. Eur. Conf. Comput. Vis., 2022.
  59. 59.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: Deformable transformers for end-to-end object detection. Int. Conf. Learn. Represent., 2021.

Citation

MLA
Wu, S., et al. “Aligning Bag of Regions for Open-Vocabulary Object Detection”. arXiv, 2023, http://arxiv.org/abs/2302.13996v1.
APA
Wu, S., Zhang, W., Jin, S., Liu, W., & Loy, C. C. (2023). Aligning Bag of Regions for Open-Vocabulary Object Detection. arXiv. http://arxiv.org/abs/2302.13996v1
Chicago
Wu, S., W. Zhang, S. Jin, W. Liu, and C. C. Loy. 2023. “Aligning Bag of Regions for Open-Vocabulary Object Detection”. arXiv. http://arxiv.org/abs/2302.13996v1.
Harvard
Wu, S. et al. (2023) “Aligning Bag of Regions for Open-Vocabulary Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.13996v1.
Vancouver
1. Wu S, Zhang W, Jin S, Liu W, Loy CC (2023) Aligning Bag of Regions for Open-Vocabulary Object Detection. arXiv

BibTeX

@article{wu2023aligning,
  title = {Aligning Bag of Regions for Open-Vocabulary Object Detection},
  author = {Wu, Size and Zhang, Wenwei and Jin, Sheng and Liu, Wentao and Loy, Chen Change},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.13996v1},
  eprint = {2302.13996}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE