Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Xiuye GuTsung-Yi LinWeicheng KuoYin Cui

article2021ICLR1,398 citations

Introduces ViLD, a distillation framework that transfers multimodal knowledge from pretrained vision-language models into two-stage object detectors, enabling the detection of novel categories specified by arbitrary text without requiring additional bounding-box annotations.

Listen

Traditional computer vision systems identify only the specific object categories they were explicitly trained on, making it prohibitively expensive and time-consuming to scale detection to thousands of rare or real-world concepts. This article evaluates a new framework named ViLD (Vision and Language Knowledge Distillation), which aims to detect arbitrary object categories from text descriptions without requiring manual box-level annotations for every single target class.

The authors tackle this challenge by transferring knowledge from existing, highly capable vision-and-language models (such as CLIP and ALIGN) directly into standard two-stage object detectors. Rather than relying on rigid, hard-coded category classifiers, the proposed system uses class-agnostic proposal networks to locate potential objects and aligns their visual region representations with both the text and image embeddings produced by the teacher models. The approach was systematically evaluated across large-scale benchmarks, including LVIS (over 1,200 categories), COCO, PASCAL VOC, and Objects365.

The investigation produced several key findings. First, on the challenging LVIS benchmark, ViLD scored 16.1 average precision on novel, unseen categories using a standard ResNet-50 backbone, outperforming traditional fully supervised baselines by 3.8 points. Second, when equipped with a stronger teacher model (ALIGN) and an ensemble configuration, the system achieved 26.3 average precision on novel categories, coming within 3.7 points of heavily engineered, fully supervised challenge-winning models. Third, detectors trained with this method transferred smoothly across separate datasets without fine-tuning, reaching 72.2 average precision on PASCAL VOC and outperforming existing state-of-the-art open-vocabulary detectors on COCO by 4.8 points on novel classes and 11.4 points overall.

These results demonstrate that organizations can bypass the massive cost and operational bottlenecks of collecting niche bounding-box training data by leveraging pre-trained vision-language models. Additionally, because the distillation process runs offline during training, the resulting detector operates at standard inference speeds during deployment, avoiding the latency penalties of running massive vision-language models per proposal. The framework also enables interactive, on-the-fly detection where users can query fine-grained attributes and visual descriptors dynamically.

Decision-makers should consider adopting distillation-based open-vocabulary detection for applications with long-tailed or rapidly evolving visual classes, such as autonomous driving and catalog management. When deploying, teams should utilize dual-head ensemble architectures to resolve optimization trade-offs between base and novel categories, and select higher-capacity teacher models to maximize accuracy.

Users should remain mindful of certain limitations: the system can struggle with visually indistinguishable sub-species, highly distorted aspect ratios, and crowded scenes containing multiple overlapping items in a single region. However, given the strong quantitative validation across multiple public benchmarks, confidence in the framework's core open-vocabulary generalization capabilities remains high.

Cover for Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Abstract

We aim at advancing open-vocabulary object detection, which detects objects described by arbitrary text inputs. The fundamental challenge is the availability of training data. It is costly to further scale up the number of classes contained in existing object detection datasets. To overcome this challenge, we propose ViLD, a training method via Vision and Language knowledge Distillation. Our method distills the knowledge from a pretrained open-vocabulary image classification model (teacher) into a two-stage detector (student). Specifically, we use the teacher model to encode category texts and image regions of object proposals. Then we train a student detector, whose region embeddings of detected boxes are aligned with the text and image embeddings inferred by the teacher. We benchmark on LVIS by holding out all rare categories as novel categories that are not seen during training. ViLD obtains 16.1 mask APr_r with a ResNet-50 backbone, even outperforming the supervised counterpart by 3.8. When trained with a stronger teacher model ALIGN, ViLD achieves 26.3 APr_r. The model can directly transfer to other datasets without finetuning, achieving 72.2 AP50_{50} on PASCAL VOC, 36.6 AP on COCO and 11.8 AP on Objects365. On COCO, ViLD outperforms the previous state-of-the-art by 4.8 on novel AP and 11.4 on overall AP. Code and demo are open-sourced at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Localization for novel categories
  • 3.2 Open-vocabulary detection with cropped regions
  • 3.3 ViLD: Vision and Language knowledge Distillation.
  • 3.4 Model ensembling
  • 4 Experiments
  • 4.1 Benchmark settings
  • 4.2 Learning generalizable object proposals
  • 4.3 Open-vocabulary classifier on cropped regions
  • 4.4 Vision and language knowledge distillation
  • 4.5 Performance comparison on COCO dataset
  • 4.6 Transfer to other datasets
  • 4.7 Qualitative results
  • 5 Conclusion
  • References
  • A Additional qualitative results
  • B Analysis of CLIP on cropped regions
  • C Additional quantitative results
  • D More implementation details

Knowls

  1. Knowl 1 — ViLD: Open-Vocabulary Object Detection via Vision and Language Knowledge Distillation

    model/method

    Vision and Language Knowledge Distillation (ViLD) is a framework for training two-stage open-vocabulary object detectors (such as Mask R-CNN) to detect novel categories using only bounding box and mask annotations from base categories CB\mathcal{C}_B. ViLD distills knowledge from a pretrained open-vocabulary image classification model consisting of a text encoder T(⋅)\mathcal{T}(\cdot) and an image encoder V(⋅)\mathcal{V}(\cdot) (e.g., CLIP or ALIGN).

    The overall training loss combines a classification objective against fixed category text embeddings (ViLD-text) and a distillation objective aligning detector region embeddings with the teacher image encoder embeddings (ViLD-image):

    LViLD=LViLD-text+w⋅LViLD-image\mathcal{L}_{\text{ViLD}} = \mathcal{L}_{\text{ViLD-text}} + w \cdot \mathcal{L}_{\text{ViLD-image}}

    where ww is a distillation loss weighting hyperparameter.

    During inference, category text embeddings T(CB∪CN)\mathcal{T}(\mathcal{C}_B \cup \mathcal{C}_N) of both base categories CB\mathcal{C}_B and novel categories CN\mathcal{C}_N are passed to the classifier head, enabling open-vocabulary detection without altering the detector architecture or requiring proposal cropping through the heavy vision-language model.

  2. Knowl 2 — ViLD-Text Objective and Region Classification Formulation

    equation

    In ViLD-text, the standard learned classification matrix of a two-stage detector is replaced by category text embeddings generated offline by the pretrained text encoder T(⋅)\mathcal{T}(\cdot), along with a learnable background vector ebge_{\text{bg}}.

    Let ϕ(I)\phi(I) denote the backbone feature representation of image II, r∈Pr \in \mathcal{P} denote a region proposal generated online, and R(ϕ(I),r)\mathcal{R}(\phi(I), r) denote a lightweight head projecting RoI features into the vision-language embedding space to produce region embedding ere_r. Given base category text embeddings T(CB)={t1,t2,…,t∣CB∣}\mathcal{T}(\mathcal{C}_B) = \{t_1, t_2, \dots, t_{|\mathcal{C}_B|}\} and cosine similarity sim(a,b)=a⊤b∥a∥2∥b∥2\text{sim}(a, b) = \frac{a^\top b}{\|a\|_2 \|b\|_2}, the similarity logit vector is:

    z(r)=[sim(er,ebg),sim(er,t1),…,sim(er,t∣CB∣)]z(r) = \left[ \text{sim}(e_r, e_{\text{bg}}), \text{sim}(e_r, t_1), \dots, \text{sim}(e_r, t_{|\mathcal{C}_B|}) \right]

    The ViLD-text loss over N=∣P∣N = |\mathcal{P}| proposals per image is computed using cross-entropy with temperature parameter τ\tau:

    LViLD-text=1N∑r∈PLCE(softmax(z(r)τ),yr)\mathcal{L}_{\text{ViLD-text}} = \frac{1}{N} \sum_{r \in \mathcal{P}} \mathcal{L}_{\text{CE}}\left( \text{softmax}\left(\frac{z(r)}{\tau}\right), y_r \right)

    where yry_r is the ground-truth base class index or the background index if rr does not match any base category annotation.

  3. Knowl 3 — ViLD-Image Objective for Proposal Feature Distillation

    equation

    ViLD-image transfers visual representation knowledge for both base and novel object regions by distilling representations directly from the pretrained image encoder V(⋅)\mathcal{V}(\cdot) into the detector's region head R(⋅)\mathcal{R}(\cdot).

    For each training image, an offline proposal set r~∈P~\tilde{r} \in \tilde{\mathcal{P}} of size MM is extracted. To capture both local object appearance and contextual cues, proposals are cropped at 1.0×1.0\times and 1.5×1.5\times scales and passed through V(⋅)\mathcal{V}(\cdot) to form an ensembled, normalized target visual embedding:

    V(crop(I,r~{1×,1.5×}))=v∥v∥2,where v=V(crop(I,r~1×))+V(crop(I,r~1.5×))\mathcal{V}\left(\text{crop}(I, \tilde{r}_{\{1\times, 1.5\times\}})\right) = \frac{v}{\|v\|_2}, \quad \text{where } v = \mathcal{V}\left(\text{crop}(I, \tilde{r}_{1\times})\right) + \mathcal{V}\left(\text{crop}(I, \tilde{r}_{1.5\times})\right)

    The visual distillation loss is defined as the mean L1\mathcal{L}_1 distance between the student's region embedding and the teacher's ensembled image embedding:

    LViLD-image=1M∑r~∈P~∥V(crop(I,r~{1×,1.5×}))−R(ϕ(I),r~)∥1\mathcal{L}_{\text{ViLD-image}} = \frac{1}{M} \sum_{\tilde{r} \in \tilde{\mathcal{P}}} \left\| \mathcal{V}\left(\text{crop}(I, \tilde{r}_{\{1\times, 1.5\times\}})\right) - \mathcal{R}(\phi(I), \tilde{r}) \right\|_1

  4. Knowl 4 — ViLD-Ensemble Architecture and Score Fusion

    model/method

    Simultaneously optimizing the cross-entropy loss LViLD-text\mathcal{L}_{\text{ViLD-text}} and the L1\mathcal{L}_1 distillation loss LViLD-image\mathcal{L}_{\text{ViLD-image}} on a shared head introduces gradient contention. ViLD-ensemble resolves this by learning two separate heads with identical architectures on top of the backbone: one head trained via LViLD-text\mathcal{L}_{\text{ViLD-text}} yielding category confidence scores pi,ViLD-textp_{i, \text{ViLD-text}}, and a second head trained via LViLD-image\mathcal{L}_{\text{ViLD-image}} yielding confidence scores pi,ViLD-imagep_{i, \text{ViLD-image}} after applying text classifier embeddings T\mathcal{T}.

    During inference, predictions from both heads are merged using class-dependent weighted geometric averaging:

    pi,ensemble={pi,ViLD-textλ⋅pi,ViLD-image(1−λ),if i∈CBpi,ViLD-text(1−λ)⋅pi,ViLD-imageλ,if i∈CNp_{i, \text{ensemble}} = \begin{cases} p_{i, \text{ViLD-text}}^\lambda \cdot p_{i, \text{ViLD-image}}^{(1 - \lambda)}, & \text{if } i \in \mathcal{C}_B \\[6pt] p_{i, \text{ViLD-text}}^{(1 - \lambda)} \cdot p_{i, \text{ViLD-image}}^\lambda, & \text{if } i \in \mathcal{C}_N \end{cases}

    where λ∈[0,1]\lambda \in [0, 1] is a weighting hyperparameter (set to λ=2/3\lambda = 2/3). This biases the base category predictions towards the supervised classification head while relying more on the visually distilled head for novel categories.

  5. Knowl 5 — Class-Agnostic Localization for Novel Object Candidates

    model/method

    To localize novel categories without category-specific training annotations, the standard category-specific bounding box regression and mask prediction heads of two-stage detectors (e.g., Mask R-CNN) are replaced with class-agnostic localization modules.

    For every region of interest (RoI) extracted by the Region Proposal Network (RPN), the model predicts a single 4-coordinate bounding box delta and a single instance segmentation mask across all object categories. Training these shared heads exclusively on base annotations CB\mathcal{C}_B yields category-agnostic object localization representations that generalize directly to candidate objects belonging to novel classes CN\mathcal{C}_N.

  6. Knowl 6 — Open-Vocabulary Detection Performance on LVIS Benchmark

    data/table

    The open-vocabulary benchmark on LVIS v1 designates 866 frequent and common categories as base categories CB\mathcal{C}_B for training, holding out 337 rare categories as novel categories CN\mathcal{C}_N. The primary metric is mask average precision on rare classes (extAPr ext{AP}_r). ViLD trained with ResNet-50 and CLIP (ViT-B/32) significantly outperforms the supervised counterpart on rare categories. Scaling student backbones and using ALIGN (EfficientNet-l2 / BERT-large) as teacher closes the gap to fully supervised challenge-winning models.

    Backbone Method APr\text{AP}_r APc\text{AP}_c APf\text{AP}_f AP
    ResNet-50+ViT-B/32 CLIP on cropped regions 18.9 18.8 16.0 17.7
    ViLD-text+CLIP 22.6 24.8 29.2 26.1
    ResNet-50 Supervised-RFS (base+novel) 12.3 24.3 32.4 25.4
    GloVe baseline 3.0 20.1 30.4 21.2
    ViLD-text 10.1 23.9 32.5 24.9
    ViLD-image 11.2 11.3 11.1 11.2
    ViLD (w=0.5w=0.5) 16.1 20.0 28.3 22.5
    ViLD-ensemble (w=0.5w=0.5) 16.6 24.6 30.3 25.5
    EfficientNet-b7 ViLD-ensemble w/ ViT-L/14 (w=1.0w=1.0) 21.7 29.1 33.6 29.6
    ViLD-ensemble w/ ALIGN (w=1.0w=1.0) 26.3 27.2 32.9 29.3
    ResNeSt269+HTC 2020 Challenge winner (supervised) 30.0 41.9 46.0 41.5
  7. Knowl 7 — Open-Vocabulary Object Detection Results on COCO Benchmark

    data/table

    On the COCO open-vocabulary benchmark (48 base categories, 17 novel categories), bounding box AP at IoU=0.5\text{IoU}=0.5 (AP50\text{AP}_{50}) is reported for generalized open-vocabulary detection using a ResNet-50 backbone.

    Method Training source Novel AP Base AP Overall AP
    Bilen Vedaldi (2016) image-level labels in CB∪CN\mathcal{C}_B \cup \mathcal{C}_N 19.7 19.6 19.6
    Ye et al. (2019) image-level labels in CB∪CN\mathcal{C}_B \cup \mathcal{C}_N 20.3 20.1 20.1
    Bansal et al. (2018) instance-level labels in CB\mathcal{C}_B 0.31 29.2 24.9
    Zhu et al. (2020) instance-level labels in CB\mathcal{C}_B 3.41 13.8 13.0
    Rahman et al. (2020) instance-level labels in CB\mathcal{C}_B 4.12 35.9 27.9
    Zareian et al. (2021) captions in CB∪CN\mathcal{C}_B \cup \mathcal{C}_N + instances in CB\mathcal{C}_B 22.8 46.0 39.9
    CLIP on cropped regions CLIP image-text pairs + instances in CB\mathcal{C}_B 26.3 28.3 27.8
    ViLD-text CLIP image-text pairs + instances in CB\mathcal{C}_B 5.9 61.8 47.2
    ViLD-image CLIP image-text pairs + instances in CB\mathcal{C}_B 24.1 34.2 31.6
    ViLD (w=0.5w = 0.5) CLIP image-text pairs + instances in CB\mathcal{C}_B 27.6 59.5 51.3

    ViLD outperforms previous state-of-the-art methods on both novel categories (+4.8+4.8 AP over Zareian et al.) and overall categories (+11.4+11.4 AP), without requiring detection-specific caption pretraining.

  8. Knowl 8 — Direct Zero-Shot Transfer to External Detection Datasets

    data/table

    A ViLD model trained solely on LVIS with a ResNet-50 backbone can directly transfer to new datasets (PASCAL VOC 2007 test set, COCO validation set, and Objects365 v1 validation set) without weight finetuning by replacing the category text embeddings with the new label names while retaining the learned LVIS background embedding vector.

    Method PASCAL VOC COCO Objects365
    AP50\text{AP}_{50} AP75\text{AP}_{75} AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75}
    ViLD-text 40.5 31.6 28.8 43.4 31.4 10.4 15.8 11.1
    ViLD 72.2 56.7 36.6 55.6 39.8 11.8 18.2 12.6
    Finetuning 78.9 60.3 39.1 59.8 42.4 15.2 23.9 16.2
    Supervised 78.5 49.0 46.5 67.6 50.9 25.6 38.6 28.0

    Direct transfer via ViLD substantially improves upon ViLD-text (e.g., 72.272.2 vs 40.540.5 AP50\text{AP}_{50} on PASCAL VOC), demonstrating that visual distillation aligns detector region representations into the joint vision-language latent space.

  9. Knowl 9 — Recall Generalization of Region Proposal Networks to Novel Categories

    empirical result

    A Region Proposal Network (RPN) trained exclusively on base categories CB\mathcal{C}_B in LVIS generalizes effectively to novel categories CN\mathcal{C}_N, achieving bounding box average recall (extARr ext{AR}_r) that closely approaches an RPN supervised with both base and novel annotations:

    • At 100 proposals: base-only supervision achieves 39.3% ARr39.3\% \text{ AR}_r vs. 41.1% ARr41.1\% \text{ AR}_r for base+novel supervision (Δ=−1.8%\Delta = -1.8\%).
    • At 300 proposals: base-only supervision achieves 48.3% ARr48.3\% \text{ AR}_r vs. 50.9% ARr50.9\% \text{ AR}_r for base+novel supervision (Δ=−2.6%\Delta = -2.6\%).
    • At 1000 proposals: base-only supervision achieves 55.6% ARr55.6\% \text{ AR}_r vs. 57.0% ARr57.0\% \text{ AR}_r for base+novel supervision (Δ=−1.4%\Delta = -1.4\%).

    This demonstrates that localization modules do not overfit to known classes and provide candidate regions for unseen objects.

  10. Knowl 10 — Systematic Expansion of Detection Vocabulary with Attributes

    equation

    To scale an open-vocabulary detector's base vocabulary V={v1,…,vp}\mathcal{V} = \{v_1, \dots, v_p\} into a combined vocabulary of size p×qp \times q using an attribute set A={a1,…,aq}\mathcal{A} = \{a_1, \dots, a_q\}, the joint probability for region embedding ere_r is factorized under the conditional independence assumption vi⊥aj∣erv_i \perp a_j \mid e_r:

    Pr⁡(vi,aj∣er)=Pr⁡(vi∣er)⋅Pr⁡(aj∣er)\Pr(v_i, a_j \mid e_r) = \Pr(v_i \mid e_r) \cdot \Pr(a_j \mid e_r)

    where the individual category and attribute probabilities are computed using the pretrained text encoder T(⋅)\mathcal{T}(\cdot) and temperature τ\tau:

    Pr⁡(vi∣er)=softmaxi(sim(er,T(v))τ)\Pr(v_i \mid e_r) = \text{softmax}_i\left( \frac{\text{sim}(e_r, \mathcal{T}(v))}{\tau} \right)

    Pr⁡(aj∣er)=softmaxj(sim(er,T(a))τ)\Pr(a_j \mid e_r) = \text{softmax}_j\left( \frac{\text{sim}(e_r, \mathcal{T}(a))}{\tau} \right)

    This allows fine-grained object detection (e.g., color attributes or sub-species classification) by combining category and attribute text descriptions without joint retraining.

  11. Knowl 11 — Failure Modes of Direct Pretrained Classifier Inference on Object Proposals

    limitation

    Directly feeding cropped region proposals into a pretrained open-vocabulary image classifier (such as CLIP) exhibits distinct failure modes:

    1. Uncalibrated Localization Scores: Pretrained image classifiers evaluate object presence rather than box tightness, often assigning high classification confidence to partial or low-quality bounding boxes. (This is mitigated during crop inference by multiplying the classifier confidence score with the detector's proposal objectness score via geometric mean).
    2. Aspect Ratio Distortion: Standard image classifiers resize inputs to square dimensions (e.g., 224×224224 \times 224), causing extreme aspect ratio distortion for thin or elongated object proposals.
    3. Clutter and Co-occurring Objects: Contrastive image-level pretraining associates whole images with prominent caption entities, leading crop-based classifiers to be misled by secondary background objects inside the proposal box.
    4. Inference Latency: Evaluating thousands of cropped proposals individually through the image encoder results in an inference runtime roughly 630×630\times slower than standard two-stage detectors.

Coverage note — Omitted the full enumeration of all 63 prompt templates (listed in Appendix D) as the concept of prompt ensembling and its ablation performance are adequately covered in the main method and results.

References

  1. 1.Zeynep Akata, Mateusz Malinowski, Mario Fritz, and Bernt Schiele. Multi-cue zero-shot learning with strong supervision. In CVPR, 2016.
  2. 2.Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran. Zero-shot object detection. In ECCV, 2018.
  3. 3.Hakan Bilen and Andrea Vedaldi. Weakly supervised deep detection networks. In CVPR, 2016.
  4. 4.Yannick Le Cacheux, Herve Le Borgne, and Michel Crucianu. Modeling inter and intra-class relations in the triplet loss for zero-shot learning. In ICCV, 2019.
  5. 5.Berkan Demirel, Ramazan Gokberk Cinbis, and Nazli Ikizler-Cinbis. Zero-shot object detection by hybrid region embedding. In BMVC, 2018.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. NAACL, 2019.
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2020.
  8. 8.Mohamed Elhoseiny, Yizhe Zhu, Han Zhang, and Ahmed Elgammal. Link the head to the “beak”: Zero shot learning from noisy text description at part precision. In CVPR, 2017.
  9. 9.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2010.
  10. 10.Ali Farhadi, Ian Endres, Derek Hoiem, and David Forsyth. Describing objects by their attributes. In CVPR, 2009.
  11. 11.Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov. Devise: A deep visual-semantic embedding model. In NeurIPS, 2013.
  12. 12.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014.
  13. 13.Ross Girshick, Ilija Radosavovic, Georgia Gkioxari, Piotr Dollár, and Kaiming He. Detectron. https://github.com/facebookresearch/detectron, 2018.
  14. 14.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, 2019.
  15. 15.Nasir Hayat, Munawar Hayat, Shafin Rahman, Salman Khan, Syed Waqas Zamir, and Fahad Shahbaz Khan. Synthesizing the unseen for zero-shot object detection. In ACCV, 2020.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  17. 17.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  18. 18.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  19. 19.Dinesh Jayaraman and Kristen Grauman. Zero shot recognition with unreliable attributes. NeurIPS 2014, 2014.
  20. 20.Zhong Ji, Yanwei Fu, Jichang Guo, Yanwei Pang, Zhongfei Mark Zhang, et al. Stacked semanticsguided attention model for fine-grained zero-shot learning. In NeurIPS, 2018.
  21. 21.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. ICML, 2021.
  22. 22.KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Vineeth N Balasubramanian. Towards open world object detection. In CVPR, 2021.
  23. 23.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020.
  24. 24.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  25. 25.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  26. 26.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten. Exploring the limits of weakly supervised pretraining. In ECCV, 2018.
  27. 27.Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S Corrado, and Jeffrey Dean. Zero-shot learning by convex combination of semantic embeddings. ICLR, 2014.
  28. 28.Jeffrey Pennington, Richard Socher, and Christopher Manning. GloVe: Global vectors for word representation. In EMNLP, 2014.
  29. 29.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. ICML, 2021.
  30. 30.Shafin Rahman, Salman Khan, and Fatih Porikli. Zero-shot object detection: Learning to simultaneously recognize and localize novel concepts. In ACCV, 2018.
  31. 31.Shafin Rahman, Salman Khan, and Nick Barnes. Transductive learning for zero-shot object detection. In ICCV, 2019.
  32. 32.Shafin Rahman, Salman Khan, and Nick Barnes. Improved visual-semantic alignment for zero-shot object detection. In AAAI, 2020.
  33. 33.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In CVPR, 2017.
  34. 34.Shaoqing Ren, Kaiming He, Ross B Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS, 2015.
  35. 35.Marcus Rohrbach, Michael Stark, and Bernt Schiele. Evaluating knowledge transfer and zero-shot learning in a large-scale setting. In CVPR, 2011.
  36. 36.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019.
  37. 37.Jingru Tan, Gang Zhang, Hanming Deng, Changbao Wang, Lewei Lu, quanquan Li, and Jifeng Dai. Technical report: A good box is not a guarantee of a good mask. Joint COCO and LVIS workshop at ECCV 2020: LVIS Challenge Track, 2020.
  38. 38.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, 2019.
  39. 39.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  40. 40.Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011 dataset. 2011.
  41. 41.Xiaolong Wang, Yufei Ye, and Abhinav Gupta. Zero-shot recognition via semantic embeddings and knowledge graphs. In CVPR, 2018.
  42. 42.Guo-Sen Xie, Li Liu, Fan Zhu, Fang Zhao, Zheng Zhang, Yazhou Yao, Jie Qin, and Ling Shao. Region graph embedding network for zero-shot learning. In ECCV, 2020.
  43. 43.Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li, Danfeng Qin, and Jesse Berent. Cap2det: Learning to amplify weak caption supervision for object detection. In ICCV, 2019.
  44. 44.Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. Open-vocabulary object detection using captions. In CVPR, 2021.
  45. 45.Hang Zhao, Xavier Puig, Bolei Zhou, Sanja Fidler, and Antonio Torralba. Open vocabulary scene parsing. In ICCV, 2017.
  46. 46.Xiangyun Zhao, Samuel Schulter, Gaurav Sharma, Yi-Hsuan Tsai, Manmohan Chandraker, and Ying Wu. Object detection with a unified label space from multiple datasets. In ECCV, 2020.
  47. 47.Ye Zheng, Ruoran Huang, Chuanqi Han, Xi Huang, and Li Cui. Background learnable cascade for zero-shot object detection. In ACCV, 2020.
  48. 48.Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl. Simple multi-dataset detection. arXiv preprint arXiv:2102.13086, 2021.
  49. 49.Pengkai Zhu, Hanxiao Wang, and Venkatesh Saligrama. Don’t even look once: Synthesizing features for zero-shot detection. In CVPR, 2020.
  50. 50.Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai. Deformable convnets v2: More deformable, better results. In CVPR, 2019.

Citation

MLA
Gu, X., et al. “Open-vocabulary Object Detection via Vision and Language Knowledge Distillation”. ICLR 2022, 2021, http://arxiv.org/abs/2104.13921v3.
APA
Gu, X., Lin, T.-Y., Kuo, W., & Cui, Y. (2021). Open-vocabulary Object Detection via Vision and Language Knowledge Distillation. ICLR 2022. http://arxiv.org/abs/2104.13921v3
Chicago
Gu, X., T.-Y. Lin, W. Kuo, and Y. Cui. 2021. “Open-vocabulary Object Detection via Vision and Language Knowledge Distillation”. ICLR 2022. http://arxiv.org/abs/2104.13921v3.
Harvard
Gu, X. et al. (2021) “Open-vocabulary Object Detection via Vision and Language Knowledge Distillation”, ICLR 2022 [Preprint]. Available at: http://arxiv.org/abs/2104.13921v3.
Vancouver
1. Gu X, Lin T-Y, Kuo W, Cui Y (2021) Open-vocabulary Object Detection via Vision and Language Knowledge Distillation. ICLR 2022

BibTeX

@article{gu2021open,
  title = {Open-vocabulary Object Detection via Vision and Language Knowledge Distillation},
  author = {Gu, Xiuye and Lin, Tsung-Yi and Kuo, Weicheng and Cui, Yin},
  year = {2021},
  journal = {ICLR 2022},
  url = {http://arxiv.org/abs/2104.13921v3},
  eprint = {2104.13921}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission