Object Detection in 20 Years: A Survey
Zhengxia ZouKeyan ChenZhenwei ShiYuhong GuoJieping Ye
Synthesizes twenty-five years of object detection research by analyzing milestone detector architectures, benchmark datasets, evaluation metrics, and speed-up techniques spanning classical feature-based methods to modern deep learning models.
Visual object detection addresses the core capability of identifying what objects are present in digital images and where they are located. This technology is critical across high-impact industries, powering real-world applications in autonomous driving, robotic perception, and automated surveillance. The article evaluates the historical progression, foundational breakthroughs, and algorithmic innovations that have shaped object detection over a quarter-century span, from the 1990s through 2022.
The review evaluates the transition across two primary development eras—early traditional frameworks utilizing handcrafted feature engineering and contemporary deep learning systems. It assesses architectural designs, core technical components such as multi-scale handling and loss functions, optimization and acceleration strategies, and modern transformer-based methods across standard industry benchmarks including PASCAL VOC and MS-COCO.
The analysis reveals several key findings across the technical landscape. First, deep learning triggered an immense accuracy leap; benchmark mean Average Precision rose from under 34% in early traditional models to over 70% in modern architectures on standard benchmarks. Second, the architectural divide between high-precision two-stage detectors and real-time one-stage detectors has substantially narrowed, particularly through improvements in balancing background samples and anchor-free keypoint formulations. Third, attention-based Vision Transformers have emerged as the leading architecture, occupying the top benchmark tiers and enabling fully end-to-end set prediction without relying on conventional bounding-box priors. Finally, practical deployment has been enabled by multifaceted speed-up techniques, including shared feature map calculations, lightweight separable convolutions, and mathematical frequency-domain accelerations.
These findings indicate that system accuracy is no longer the primary bottleneck for standard detection environments. Instead, development priorities are pivoting toward operational trade-offs involving computational complexity, hardware constraints, and inference latency. The growing reliance on massive supervised datasets presents data acquisition risks and costs, prompting the rise of weakly supervised learning and domain adaptation to handle unconstrained operating environments.
Organizations developing or deploying visual detection systems should pursue specialized architectural pathways based on operational requirements. Teams requiring high-throughput, edge-deployed intelligence should prioritize lightweight, one-stage, or anchor-free models optimized via network pruning and specialized convolutions. Conversely, applications demanding maximum localization precision should evaluate transformer-based architectures. Further investment is recommended in open-world detection, 3D multi-sensor fusion, and video temporal modeling before automated systems can operate reliably under complex, out-of-distribution real-world conditions.
The conclusions reflect established empirical results across standard public vision benchmarks. However, leaders should exercise caution when translating these findings directly to real-world edge environments, where factors such as small or heavily occluded objects, domain shift, sensor degradation, and restricted processing power present challenges not fully reflected in standard dataset evaluations.
- Paper: Object Detection with Discriminatively Trained Part-Based Models, Pedro F. Felzenszwalb et al. (2010). Felzenszwalb et al.'s deformable part-based models represent the definitive milestone of the pre-deep learning era that the survey analyzes to illustrate the foundations of classical object detection.
- Paper: Rich feature hierarchies for accurate object detection and semantic segmentation, Ross Girshick et al. (2014). R-CNN inaugurated the deep learning revolution in generic object detection, establishing the foundational two-stage paradigm thoroughly tracked throughout the survey.
- Paper: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks, Shaoqing Ren et al. (2015). Faster R-CNN introduced Region Proposal Networks to make two-stage detection end-to-end, serving as a core architectural baseline reviewed in the survey.
- Paper: You Only Look Once: Unified, Real-Time Object Detection, Joseph Redmon et al. (2016). YOLO pioneered the unified single-stage detection framework that the survey discusses extensively in its evolution of real-time detection systems.
- Paper: Feature Pyramid Networks for Object Detection, Tsung-Yi Lin et al. (2017). Feature Pyramid Networks introduced the essential top-down multi-scale feature representations that the survey highlights as a critical building block for modern detectors.
- Paper: Focal Loss for Dense Object Detection, Tsung-Yi Lin et al. (2017). RetinaNet solved the extreme foreground-background class imbalance with Focal Loss, a pivotal milestone analyzed in the survey's discussion of one-stage versus two-stage detectors.
- Paper: SSD: Single Shot MultiBox Detector, W. Liu et al. (2015). SSD established multi-scale feature map regression for single-shot detectors, forming another major architectural milestone reviewed in the survey.
- Paper: The Pascal Visual Object Classes Challenge: A Retrospective, M. Everingham et al. (2014). This retrospective on the PASCAL VOC challenge documents the standardized benchmarks and evaluation metrics that provided the empirical basis for early detection progress reviewed in the survey.
- Paper: ImageNet Large Scale Visual Recognition Challenge, Olga Russakovsky et al. (2014). ImageNet catalyzed the transition from hand-engineered features to deep convolutional detectors, serving as the core pre-training dataset surveyed in the paper.
- Paper: Detecting Faces in Images: A Survey, Ming-Hsuan Yang et al. (2002). This survey provides essential historical context on early appearance-based and hand-crafted feature detection techniques from the 1990s and 2000s that precede modern deep learning methods.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). DETR introduces end-to-end transformer-based object detection without anchor boxes or NMS, initiating the new architectural paradigm that followed the survey's primary CNN focus.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Deformable DETR resolves the slow convergence and multi-scale feature limitations of early detection transformers by integrating deformable attention.
- Paper: DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, Hao Zhang et al. (2023). DINO advances transformer-based object detection through denoising training and query formulation, establishing state-of-the-art end-to-end performance on benchmark datasets.
- Paper: Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection, Shilong Liu et al. (2023). Grounding DINO extends modern transformer detection into open-set and language-guided detection, moving beyond the fixed closed-set category paradigms covered in the survey.
- Paper: YOLOv4: Optimal Speed and Accuracy of Object Detection, Alexey Bochkovskiy et al. (2020). YOLOv4 optimizes single-stage real-time detection by consolidating modern architectural modifications and training strategies into a unified pipeline.
- Paper: YOLOX: Exceeding YOLO Series in 2021, Zheng Ge et al. (2021). YOLOX modernizes the real-time YOLO detector family by introducing an anchor-free design, decoupled heads, and dynamic label assignment.
- Paper: YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors, Chien-Yao Wang et al. (2023). YOLOv7 further advances real-time detector architectures using extended layer aggregation and trainable re-parameterization techniques.
- Paper: EfficientDet: Scalable and Efficient Object Detection, Mingxing Tan et al. (2020). EfficientDet systematically scales one-stage architectures across diverse computational budgets using weighted bidirectional feature pyramids and compound scaling.
- Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). RT-DETR develops the first real-time end-to-end transformer detector, successfully eliminating NMS while matching the latency and efficiency of modern YOLO models.
- Paper: MMDetection: Open MMLab Detection Toolbox and Benchmark, Kai Chen et al. (2019). MMDetection provides a unified, modular open-source codebase implementing the standard detection models, backbones, and loss functions discussed throughout the survey.
