Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Xiuye GuTsung-Yi LinWeicheng KuoYin Cui
Introduces ViLD, a distillation framework that transfers multimodal knowledge from pretrained vision-language models into two-stage object detectors, enabling the detection of novel categories specified by arbitrary text without requiring additional bounding-box annotations.
Traditional computer vision systems identify only the specific object categories they were explicitly trained on, making it prohibitively expensive and time-consuming to scale detection to thousands of rare or real-world concepts. This article evaluates a new framework named ViLD (Vision and Language Knowledge Distillation), which aims to detect arbitrary object categories from text descriptions without requiring manual box-level annotations for every single target class.
The authors tackle this challenge by transferring knowledge from existing, highly capable vision-and-language models (such as CLIP and ALIGN) directly into standard two-stage object detectors. Rather than relying on rigid, hard-coded category classifiers, the proposed system uses class-agnostic proposal networks to locate potential objects and aligns their visual region representations with both the text and image embeddings produced by the teacher models. The approach was systematically evaluated across large-scale benchmarks, including LVIS (over 1,200 categories), COCO, PASCAL VOC, and Objects365.
The investigation produced several key findings. First, on the challenging LVIS benchmark, ViLD scored 16.1 average precision on novel, unseen categories using a standard ResNet-50 backbone, outperforming traditional fully supervised baselines by 3.8 points. Second, when equipped with a stronger teacher model (ALIGN) and an ensemble configuration, the system achieved 26.3 average precision on novel categories, coming within 3.7 points of heavily engineered, fully supervised challenge-winning models. Third, detectors trained with this method transferred smoothly across separate datasets without fine-tuning, reaching 72.2 average precision on PASCAL VOC and outperforming existing state-of-the-art open-vocabulary detectors on COCO by 4.8 points on novel classes and 11.4 points overall.
These results demonstrate that organizations can bypass the massive cost and operational bottlenecks of collecting niche bounding-box training data by leveraging pre-trained vision-language models. Additionally, because the distillation process runs offline during training, the resulting detector operates at standard inference speeds during deployment, avoiding the latency penalties of running massive vision-language models per proposal. The framework also enables interactive, on-the-fly detection where users can query fine-grained attributes and visual descriptors dynamically.
Decision-makers should consider adopting distillation-based open-vocabulary detection for applications with long-tailed or rapidly evolving visual classes, such as autonomous driving and catalog management. When deploying, teams should utilize dual-head ensemble architectures to resolve optimization trade-offs between base and novel categories, and select higher-capacity teacher models to maximize accuracy.
Users should remain mindful of certain limitations: the system can struggle with visually indistinguishable sub-species, highly distorted aspect ratios, and crowded scenes containing multiple overlapping items in a single region. However, given the strong quantitative validation across multiple public benchmarks, confidence in the framework's core open-vocabulary generalization capabilities remains high.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. It introduces the CLIP contrastive vision-language pre-training paradigm that ViLD directly distills into two-stage object detectors for open-vocabulary recognition.
- Paper: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, Chao Jia et al. (2021). It establishes the ALIGN dual-encoder vision-language model, which serves as one of the key pre-trained teacher models distilled in ViLD's framework.
- Paper: LVIS: A Dataset for Large Vocabulary Instance Segmentation, Agrim Gupta et al. (2019). It provides the LVIS large-vocabulary benchmark and its rare-category distribution setup, which ViLD uses to evaluate open-vocabulary detection performance.
- Paper: DeViSE: A Deep Visual-Semantic Embedding Model, Andrea Frome et al. (2013). It presents the foundational approach of mapping visual features into text semantic embedding spaces to enable zero-shot object recognition.
- Paper: YOLO9000: Better, Faster, Stronger, Joseph Redmon et al. (2016). It pioneered expanding object detector vocabulary sizes to thousands of categories by leveraging joint training with classification data.
- Paper: Big Self-Supervised Models are Strong Semi-Supervised Learners, Ting Chen et al. (2020). It demonstrates how distillation from large pre-trained representation teachers effectively transfers general visual knowledge to downstream student architectures.
- Paper: Grounded Language-Image Pre-training, Liunian Harold Li et al. (2022). It extends open-vocabulary object detection beyond distillation by reformulating detection as unified phrase grounding with deep cross-modal fusion.
- Paper: Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection, Shilong Liu et al. (2023). It advances open-set object detection by integrating multi-stage vision-language cross-attention into transformer-based architectures like DINO.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). It surveys the broader landscape of vision-language models for visual tasks, providing a comprehensive analysis of distillation techniques including ViLD.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). It scales up vision-language integration to generalist multimodal LLMs capable of precise bounding-box localization and grounded dialogue.
