Object Detection in 20 Years: A Survey

Zhengxia ZouKeyan ChenZhenwei ShiYuhong GuoJieping Ye

article2019Proceedings of the IEEE3,631 citations

Synthesizes twenty-five years of object detection research by analyzing milestone detector architectures, benchmark datasets, evaluation metrics, and speed-up techniques spanning classical feature-based methods to modern deep learning models.

Listen

Visual object detection addresses the core capability of identifying what objects are present in digital images and where they are located. This technology is critical across high-impact industries, powering real-world applications in autonomous driving, robotic perception, and automated surveillance. The article evaluates the historical progression, foundational breakthroughs, and algorithmic innovations that have shaped object detection over a quarter-century span, from the 1990s through 2022.

The review evaluates the transition across two primary development eras—early traditional frameworks utilizing handcrafted feature engineering and contemporary deep learning systems. It assesses architectural designs, core technical components such as multi-scale handling and loss functions, optimization and acceleration strategies, and modern transformer-based methods across standard industry benchmarks including PASCAL VOC and MS-COCO.

The analysis reveals several key findings across the technical landscape. First, deep learning triggered an immense accuracy leap; benchmark mean Average Precision rose from under 34% in early traditional models to over 70% in modern architectures on standard benchmarks. Second, the architectural divide between high-precision two-stage detectors and real-time one-stage detectors has substantially narrowed, particularly through improvements in balancing background samples and anchor-free keypoint formulations. Third, attention-based Vision Transformers have emerged as the leading architecture, occupying the top benchmark tiers and enabling fully end-to-end set prediction without relying on conventional bounding-box priors. Finally, practical deployment has been enabled by multifaceted speed-up techniques, including shared feature map calculations, lightweight separable convolutions, and mathematical frequency-domain accelerations.

These findings indicate that system accuracy is no longer the primary bottleneck for standard detection environments. Instead, development priorities are pivoting toward operational trade-offs involving computational complexity, hardware constraints, and inference latency. The growing reliance on massive supervised datasets presents data acquisition risks and costs, prompting the rise of weakly supervised learning and domain adaptation to handle unconstrained operating environments.

Organizations developing or deploying visual detection systems should pursue specialized architectural pathways based on operational requirements. Teams requiring high-throughput, edge-deployed intelligence should prioritize lightweight, one-stage, or anchor-free models optimized via network pruning and specialized convolutions. Conversely, applications demanding maximum localization precision should evaluate transformer-based architectures. Further investment is recommended in open-world detection, 3D multi-sensor fusion, and video temporal modeling before automated systems can operate reliably under complex, out-of-distribution real-world conditions.

The conclusions reflect established empirical results across standard public vision benchmarks. However, leaders should exercise caution when translating these findings directly to real-world edge environments, where factors such as small or heavily occluded objects, domain shift, sensor degradation, and restricted processing power present challenges not fully reflected in standard dataset evaluations.

arXiv: 1905.05055
Cover for Object Detection in 20 Years: A Survey

Abstract

Object detection, as of one the most fundamental and challenging problems in computer vision, has received great attention in recent years. Over the past two decades, we have seen a rapid technological evolution of object detection and its profound impact on the entire computer vision field. If we consider today's object detection technique as a revolution driven by deep learning, then back in the 1990s, we would see the ingenious thinking and long-term perspective design of early computer vision. This paper extensively reviews this fast-moving research field in the light of technical evolution, spanning over a quarter-century's time (from the 1990s to 2022). A number of topics have been covered in this paper, including the milestone detectors in history, detection datasets, metrics, fundamental building blocks of the detection system, speed-up techniques, and the recent state-of-the-art detection methods.

Table of Contents

  • I Introduction
  • II Object Detection in 20 Years
  • II-A A Road Map of Object Detection
  • II-A1 Milestones: Traditional Detectors
  • II-A2 Milestones: CNN based Two-stage Detectors
  • II-A3 Milestones: CNN based One-stage Detectors
  • II-B Object Detection Datasets and Metrics
  • II-B1 Datasets
  • II-B2 Metrics
  • II-C Technical Evolution in Object Detection
  • II-C1 Technical Evolution of Multi-Scale Detection
  • II-C2 Technical Evolution of Context Priming
  • II-C3 Technical Evolution of Hard Negative Mining
  • II-C4 Technical Evolution of Loss Function
  • II-C5 Technical Evolution of Non-Maximum Suppression
  • III Speed-Up of Detection
  • III-A Feature Map Shared Computation
  • III-B Cascaded Detection
  • III-C Network Pruning and Quantification
  • III-D Lightweight Network Design
  • III-D1 Factorizing Convolutions
  • III-D2 Group Convolution
  • III-D3 Depth-wise Separable Convolution
  • III-D4 Bottle-neck Design
  • III-D5 Detection with NAS
  • III-E Numerical Acceleration
  • III-E1 Speed Up with Integral Image
  • III-E2 Speed Up in Frequency Domain
  • III-E3 Vector Quantization
  • IV Recent Advances in Object Detection
  • IV-A Beyond Sliding Window Detection
  • IV-B Robust Detection of Rotation and Scale Changes
  • IV-B1 Rotation Robust Detection
  • IV-B2 Scale Robust Detection
  • IV-C Detection with Better Backbones
  • IV-D Improvements of Localization
  • IV-D1 Bounding Box Refinement
  • IV-D2 New Loss Functions for Accurate Localization
  • IV-E Learning with Segmentation Loss
  • IV-F Adversarial Training
  • IV-G Weakly Supervised Object Detection
  • IV-H Detection with Domain Adaptation
  • V Conclusion and Future Directions
  • References

Knowls

  1. Knowl 1 — Multi-Task Objective and Bounding Box Loss Formulations in Object Detection

    equation

    The standard optimization objective for generic object detection is formulated as a multi-task loss combining classification loss and bounding box localization loss:

    L(p,p∗,t,t∗)=Lcls(p,p∗)+βI(t)Lloc(t,t∗)L(p, p^*, t, t^*) = L_{\text{cls}}(p, p^*) + \beta I(t) L_{\text{loc}}(t, t^*)

    where:

    • p=(p0,p1,…,pK)p = (p_0, p_1, \dots, p_K) represents the predicted category probability distribution over KK object classes plus background.
    • p∗p^* is the ground-truth class label.
    • t=(tx,ty,tw,th)t = (t_x, t_y, t_w, t_h) is the predicted bounding box offsets or coordinates.
    • t∗t^* is the ground-truth bounding box coordinate representation.
    • β\beta is a balancing hyperparameter weighting localization relative to classification.
    • I(t)I(t) is an indicator function defined as:

    I(t)={1if IoU{a,a∗}>η0otherwiseI(t) = \begin{cases} 1 & \text{if } \text{IoU}\{a, a^*\} > \eta \\ 0 & \text{otherwise} \end{cases}

    where IoU{a,a∗}\text{IoU}\{a, a^*\} is the Intersection over Union between an anchor/reference element aa and the ground truth a∗a^*, and η\eta is a predefined IoU matching threshold (e.g., η=0.5\eta = 0.5). If a reference does not match any ground-truth object, its localization loss is excluded from training.

    For the localization loss term LlocL_{\text{loc}}, the Smooth L1L_1 loss provides robustness against outliers compared to L2L_2 loss:

    SmoothL1(x)={0.5x2if ∣x∣<1∣x∣−0.5otherwise\text{Smooth}_{L1}(x) = \begin{cases} 0.5 x^2 & \text{if } |x| < 1 \\ |x| - 0.5 & \text{otherwise} \end{cases}

    where xx denotes the coordinate residual between target and predicted values.

    To account for correlation among bounding box coordinates (x,y,w,h)(x, y, w, h) and align training directly with evaluation metrics, IoU-based loss functions are defined, starting with the base IoU loss:

    LIoU=−log⁡(IoU)L_{\text{IoU}} = -\log(\text{IoU})

    Subsequent extensions incorporate geometric metrics: Generalized IoU (GIoU) handles non-overlapping boxes (IoU=0\text{IoU} = 0), Distance-IoU (DIoU) optimizes the normalized Euclidean distance between box center points, and Complete IoU (CIoU) simultaneously penalizes overlap area, center distance, and aspect ratio discrepancies.

  2. Knowl 2 — Landmark Object Detection Benchmark Datasets and Scale Statistics

    data/table

    The development of generic object detection algorithms has been driven by benchmark datasets with increasing image counts, class numbers, and annotation densities.

    Dataset train validation trainval test
    images objects images objects images objects images objects
    VOC-2007 2,501 6,301 2,510 6,307 5,011 12,608 4,952 14,976
    VOC-2012 5,717 13,609 5,823 13,841 11,540 27,450 10,991 -
    ILSVRC-2014 456,567 478,807 20,121 55,502 476,688 534,309 40,152 -
    ILSVRC-2017 456,567 478,807 20,121 55,502 476,688 534,309 65,500 -
    MS-COCO-2015 82,783 604,907 40,504 291,875 123,287 896,782 81,434 -
    MS-COCO-2017 118,287 860,001 5,000 36,781 123,287 896,782 40,670 -
    Objects365-2019 600,000 9,623,000 38,000 479,000 638,000 10,102,000 100,000 1,700,000
    OID-2020 1,743,042 14,610,229 41,620 303,980 1,784,662 14,914,209 125,436 937,327

    Dataset progression illustrates three trends:

    1. Dataset scale grew from thousands of annotated instances (PASCAL VOC) to tens of millions (Open Images V4/2020, Objects365).
    2. Annotation depth advanced from basic 20-class bounding boxes in VOC to per-instance segmentation masks with 80 classes in MS-COCO, facilitating stricter evaluation protocols such as MS-COCO mAP@[.5,.95]\text{mAP}@[.5, .95] across multiple IoU thresholds.
    3. Object density and scale diversity shifted toward smaller objects (areas under 1% of the image) and dense, heavily occluded real-world scenes.
  3. Knowl 3 — Two-Stage versus One-Stage Deep Object Detection Paradigms

    model/method

    Deep learning-based object detectors fall into two distinct structural paradigms:

    1. Two-Stage Detectors (Coarse-to-Fine Pipeline):

      • First Stage: Generates a candidate pool of category-agnostic Region Proposals (e.g., via Selective Search in R-CNN/Fast R-CNN or via a Region Proposal Network (RPN) in Faster R-CNN).
      • Second Stage: Extracts fixed-length regional features (via RoI Pooling, RoI Align, or Spatial Pyramid Pooling), followed by class-specific classification and refined bounding box regression.
      • Characteristics: Achieves high precision and localization accuracy by filtering out easy background candidates early, but introduces higher computational latency and complex multi-step pipelines.
    2. One-Stage Detectors (Single-Step Dense Prediction):

      • Applies a single feed-forward neural network directly to the entire image, outputting class distributions and bounding box parameters across dense spatial positions in a single pass (e.g., YOLO, SSD, RetinaNet).
      • Characteristics: Highly optimized for real-time inference and deployment on edge hardware, but historically vulnerable to severe foreground-background class imbalance and lower recall on dense or small objects.

    Modern advancements bridge these paradigms using Feature Pyramid Networks (FPN) for cross-scale representation, Focal Loss for addressing dense candidate imbalance, and keypoint/Transformer-based architectures (CornerNet, CenterNet, DETR) that remove explicit anchor generation.

  4. Knowl 4 — Evolution of Multi-Scale Object Detection Architectures

    model/method

    Handling large scale variation and variable aspect ratios has evolved across distinct developmental phases:

    1. Feature Pyramids and Sliding Windows: Traditional approaches (e.g., Viola-Jones, HOG, DPM) repeatedly rescaled input images to build an image pyramid and slid fixed-dimension detector windows across each scale, using mixture models or part-based deformations to handle aspect ratios.
    2. Object Proposal Generation: Methods like Selective Search, EdgeBoxes, and BING extracted several thousand category-agnostic bounding box proposals per image, reducing exhaustive multi-scale sliding-window searches to region-specific evaluation.
    3. Multi-Reference and Multi-Resolution Detection: Deep networks integrated multi-scale handling directly into the feature hierarchy:
      • Multi-Reference (Anchors): Defining multiscale, multiratio predefined anchor boxes or reference points at each spatial grid location (e.g., Faster R-CNN, SSD, YOLOv3/v4).
      • Multi-Resolution (Feature Pyramids): Detecting objects of varying scales on different convolutional layers. Shallow layers with higher spatial resolution detect small objects, whereas deeper layers with high semantic context detect large objects (e.g., SSD, FPN).
    4. Anchor-Free and Keypoint-Based Representation: Eliminates hand-crafted anchor priors entirely:
      • Group-based keypoint detection: Predicts pairs/triplets of keypoints (e.g., corner pairs in CornerNet, extreme/center points in ExtremeNet) and groups them using associative embeddings.
      • Group-free point regression: Treats an object as a single center point or point set and regresses spatial boundaries directly (e.g., FCOS, CenterNet).
    5. Transformer-Based Direct Set Prediction: Architectures such as DETR and Deformable DETR model object detection as a direct bounding box set prediction task using global attention mechanisms and bipartite matching.
  5. Knowl 5 — Evolution of Non-Maximum Suppression (NMS) and NMS-Free Detection

    model/method

    Because dense sliding windows and overlapping anchor boxes generate duplicate detections for the same target, post-processing mechanisms are necessary to eliminate redundancy:

    1. Traditional Greedy NMS: Iteratively selects the detection candidate with the maximum confidence score, suppressing all surrounding candidate boxes that exhibit an IoU overlap with the selected box greater than a predefined threshold. Limitations include discarding true positive detections in dense crowds and lacking false-positive correction.
    2. Heuristic Improvements (Soft-NMS & Geometric NMS): Soft-NMS continuously decays the confidence scores of neighboring overlapping boxes via linear or Gaussian penalty functions rather than setting scores to zero immediately. DIoU-NMS incorporates center distance penalties to avoid suppressing occluded adjacent targets.
    3. Bounding Box Aggregation: Combines multiple overlapping candidate boxes into a unified final box through spatial averaging or clustering (e.g., Viola-Jones grouping, OverFeat, Weighted Boxes Fusion).
    4. Learning-Based NMS: Formulates candidate suppression as an end-to-end differentiable component, training neural networks or relation modules (e.g., RelationNet, LearnNMS) to re-score raw detections while explicitly modeling spatial and contextual interactions.
    5. NMS-Free End-to-End Detectors: Implements one-to-one label assignment during network training (assigning exactly one prediction box/point per ground-truth object via Hungarian bipartite matching or cost-based assignment), eliminating the need for heuristic or learned NMS post-processing entirely (e.g., CenterNet, DETR, POTO).
  6. Knowl 6 — Technical Evolution of Hard Negative Mining and Class Imbalance Management

    model/method

    In dense sliding-window and anchor-based detection, the ratio of background negative candidates to foreground objects can reach extremes of 107:110^7 : 1. Techniques for managing this class imbalance have undergone three main phases:

    1. Classical Bootstrapping: Used in early detectors (Viola-Jones, HOG, DPM), training begins on a manageable subset of negative samples. The model is iteratively evaluated on a large background dataset, collecting false positives ("hard negatives") and re-training the classifier until convergence, reducing computation over millions of uninformative negatives.
    2. Batch Weight Balancing and Heuristic Subsampling: Early deep detectors (e.g., R-CNN, Faster R-CNN, YOLOv1/v2) replaced iterative bootstrapping with uniform mini-batch sampling, fixing ratios between positive and negative proposals (e.g., 1:3 ratio in Faster R-CNN RPN sampling) or statically weighting foreground-background losses.
    3. Online Hard Example Mining (OHEM) and Loss Reshaping:
      • OHEM: Computes loss over all candidate regions in a forward pass, selects the top-K highest-loss examples, and only backpropagates gradients through these hard examples.
      • Focal Loss: Dynamically down-weights the loss assigned to easy, well-classified background examples during dense single-stage detector training:

    FL(pt)=−αt(1−pt)γlog⁡(pt)\text{FL}(p_t) = -\alpha_t (1 - p_t)^\gamma \log(p_t)

    where ptp_t is the model's estimated probability for the ground-truth class, αt\alpha_t addresses class prevalence imbalance, and γ≥0\gamma \ge 0 is a tunable focusing parameter that suppresses easy negative gradients.

  7. Knowl 7 — Computational Complexity of Convolutional Layer Acceleration Schemes

    theoretical result

    Replacing standard 2D convolutional layers with factorized and grouped alternatives reduces computational complexity (FLOPs):

    Let cc be the number of input channels, dd the number of output channels, and k×kk \times k the spatial kernel size of the filter.

    1. Standard 2D Convolution: O(d⋅k2⋅c)\mathcal{O}(d \cdot k^2 \cdot c)
    2. Filter Factorization (Spatial Decomposition): Decomposing a k×kk \times k 2D filter into two sequential 1D filters (k×1k \times 1 followed by 1×k1 \times k) or factorizing large kernels into a chain of k′×k′k' \times k' kernels (e.g., one 7×77 \times 7 into three 3×33 \times 3): O(d⋅k⋅c)orO(d⋅k′2⋅c)\mathcal{O}(d \cdot k \cdot c) \quad \text{or} \quad \mathcal{O}(d \cdot {k'}^2 \cdot c)
    3. Channel Factorization: Inserting low-rank linear bottleneck transformations between channel mappings (c→d′→dc \to d' \to d where d′<c,dd' < c, d): O(d′⋅k2⋅c)+O(d⋅k2⋅d′)\mathcal{O}(d' \cdot k^2 \cdot c) + \mathcal{O}(d \cdot k^2 \cdot d')
    4. Group Convolution: Dividing cc input channels and dd output filters into mm disjoint groups, performing independent convolutions within each group: O(d⋅k2⋅cm)\mathcal{O}\left(\frac{d \cdot k^2 \cdot c}{m}\right)
    5. Depthwise Separable Convolution: Setting the number of groups m=cm = c (one spatial k×kk \times k filter per input channel) followed by a 1×11 \times 1 pointwise convolution across all dd output channels: O(c⋅k2)+O(d⋅c)\mathcal{O}(c \cdot k^2) + \mathcal{O}(d \cdot c) This achieves an approximate 1d+1k2\frac{1}{d} + \frac{1}{k^2} computational cost reduction relative to standard convolutions.
  8. Knowl 8 — Multi-Scale Training via Scale Normalization for Image Pyramids (SNIP and SNIPER)

    model/method

    Standard multi-scale detector training applies gradient backpropagation across all object instances regardless of scale, causing extreme scale-mismatches against pre-trained backbone receptive fields.

    1. Scale Normalization for Image Pyramids (SNIP):

      • Constructs full image pyramids at both training and testing stages.
      • During training, for an image at resolution scale ii, bounding boxes with area aa are selectively filtered: if aa falls outside a designated valid scale range [smin⁡i,smax⁡i][s_{\min}^i, s_{\max}^i], its loss gradients are zeroed out.
      • Ensures that only objects with receptive-field-compatible resolutions contribute to network weight updates.
    2. SNIP with Efficient Resampling (SNIPER):

      • Alleviates the high computational cost of processing full high-resolution image pyramids in SNIP.
      • Generates fixed-size context chips (K×KK \times K pixel sub-regions) extracted from multiple image scales.
      • Positive chips are centered around ground-truth instances at scale-normalized sizes, while negative chips are sampled from unassigned proposal regions.
      • Discards irrelevant background regions, allowing high-resolution multi-scale training with significantly larger batch sizes and reduced GPU memory overhead.
  9. Knowl 9 — Context Priming Paradigms in Object Detection Systems

    model/method

    Contextual visual relationships improve detection precision and eliminate false positives across three primary paradigms:

    1. Local Context Integration: Enlarges the bounding box proposal boundaries or receptive fields to capture immediate surrounding visual cues (e.g., surrounding background pixels around a pedestrian or facial contours around face parts).
    2. Global Context Integration: Leverages global scene layout to guide object hypothesis validation:
      • Structural operations: Employs dilated/atrous convolutions, deformable convolutions, and global pooling layers to expand receptive fields over the entire image area.
      • Sequential & Attention modeling: Uses Recurrent Neural Networks (RNNs) across spatial features or self-attention Transformer blocks (Non-Local blocks, DETR) to capture long-range cross-image dependencies.
    3. Context Interactives: Directly models semantic dependencies among elements in the image:
      • Object-to-Object relations: Explicitly models co-occurrence, relative geometry, and spatial layout constraints between distinct bounding box proposals (e.g., Structure Inference Networks, Relation Networks for Object Detection).
      • Object-to-Scene relations: Conditions per-instance classification on the overall scene category (e.g., person-in-room vs. boat-on-water context priming).
  10. Knowl 10 — Numerical and Signal-Processing Acceleration in Object Detection

    model/method

    Low-level numerical acceleration techniques speed up feature extraction and linear scoring operations in object detection systems:

    1. Integral Image Computation: Exploits the integral-differential separability of convolution in signal processing:

    f(x)∗g(x)=(∫f(x) dx)∗(dg(x)dx)f(x) * g(x) = \left( \int f(x) \, dx \right) * \left( \frac{dg(x)}{dx} \right)

    When dg(x)dx\frac{dg(x)}{dx} is a sparse kernel (such as box filters or Haar wavelets), computing summations over arbitrary sub-regions reduces to O(1)\mathcal{O}(1) constant-time lookups using four array access operations on an integral image array. This principle extends to multi-channel histograms, forming "Integral HOG" maps that compute gradient orientation histograms over arbitrary rectangular cells in constant time.

    1. Frequency Domain Acceleration (Fourier Transform Convolution): Replaces window-wise spatial inner products between feature maps II and detection filter weights WW with pointwise multiplication in the Fourier domain via the Convolution Theorem:

    I∗W=F−1(F(I)⊙F(W))I * W = \mathcal{F}^{-1}\left( \mathcal{F}(I) \odot \mathcal{F}(W) \right)

    where F\mathcal{F} is the Discrete Fourier Transform, F−1\mathcal{F}^{-1} is the Inverse Fast Fourier Transform, and ⊙\odot denotes the Hadamard (pointwise) product.

    1. Vector Quantization (VQ): Compresses high-dimensional feature distributions and template parameters into compact codebooks of prototype vectors, converting costly floating-point matrix multiplications into discrete table lookups for linear detection scoring.
  11. Knowl 11 — Weakly Supervised Object Detection (WSOD) Frameworks

    model/method

    Weakly Supervised Object Detection (WSOD) trains localized bounding box detectors using only coarse image-level category labels, eliminating manual bounding box annotations:

    1. Multiple Instance Learning (MIL): Formulates WSOD by treating each image as a "bag" and its extracted candidate proposals (from unsupervised region generators) as "instances". A bag is positive if at least one instance contains the target class and negative if no instances contain it. Training alternates between selecting high-scoring proposals and updating class appearance models.
    2. Class Activation Mapping (CAM): Exploits the localization capability of standard CNN convolutional feature layers without explicit bounding box supervision. Projecting class-specific classification weights back onto final convolutional feature activation maps produces spatial heatmaps that identify discriminative object regions.
    3. Proposal Ranking and Erasing / Masking: Identifies complete object extents rather than merely the most discriminative parts by iteratively masking high-scoring regions. If masking a region causes a sharp drop in image classification confidence, the erased area is marked as containing the target object.
  12. Knowl 12 — Invariant Modeling for Rotation-Robust Object Detection

    model/method

    Standard horizontal bounding boxes and Cartesian convolutional operations struggle to capture objects exhibiting arbitrary orientations (e.g., in aerial remote sensing, scene text, and face detection). Rotation-robust approaches employ specific design strategies:

    1. Multi-Angle Model Collections & Data Augmentation: Directly augmenting training distributions with multi-angle rotational transformations, or training an ensemble of independent orientation-specific sub-detectors.
    2. Rotation-Invariant Loss Constraints: Introducing regularization constraints into the objective function (e.g., Fisher discriminative rotation-invariant losses) that enforce metric distance minimization between feature representations of the same object under varying rotation angles.
    3. Geometric Transformation Networks: Integrating Spatial Transformer Networks (STN) or RoI Transformers that learn affine and geometric parameters directly from proposals, rectifying oriented bounding boxes into axis-aligned canonical coordinates prior to classification.
    4. Polar Coordinate Feature Pooling: Replacing standard Cartesian grid RoI Pooling with polar coordinate pooling grids, ensuring extracted regional feature representations are invariant to 2D rotational shifts.

Coverage note — Omitted high-level survey introductions to unrelated vision tasks (image captioning, tracking, instance segmentation) and generic summaries of individual historical papers that are fully captured within the technical evolution and taxonomy knowls.

References

  1. 1.B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik, “Simultaneous detection and segmentation,” in ECCV. Springer, 2014, pp. 297–312.
  2. 2.——, “Hypercolumns for object segmentation and fine-grained localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 447–456.
  3. 3.J. Dai, K. He, and J. Sun, “Instance-aware semantic segmentation via multi-task network cascades,” in CVPR, 2016, pp. 3150–3158.
  4. 4.K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV. IEEE, 2017, pp. 2980–2988.
  5. 5.A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR, 2015, pp. 3128–3137.
  6. 6.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML, 2015, pp. 2048–2057.
  7. 7.Q. Wu, C. Shen, P. Wang, A. Dick, and A. van den Hengel, “Image captioning and visual question answering based on attributes and external knowledge,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 6, pp. 1367–1381, 2018.
  8. 8.K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao, C. Zhang, Z. Wang, R. Wang, X. Wang et al., “T-cnn: Tubelets with convolutional neural networks for object detection from videos,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 10, pp. 2896–2907, 2018.
  9. 9.Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature, vol. 521, no. 7553, p. 436, 2015.
  10. 10.P. Viola and M. Jones, “Rapid object detection using a boosted cascade of simple features,” in CVPR, vol. 1. IEEE, 2001, pp. I–I.
  11. 11.P. Viola and M. J. Jones, “Robust real-time face detection,” International journal of computer vision, vol. 57, no. 2, pp. 137–154, 2004.
  12. 12.N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in CVPR, vol. 1. IEEE, 2005, pp. 886–893.
  13. 13.P. Felzenszwalb, D. McAllester, and D. Ramanan, “A discriminatively trained, multiscale, deformable part model,” in CVPR. IEEE, 2008, pp. 1–8.
  14. 14.P. F. Felzenszwalb, R. B. Girshick, and D. McAllester, “Cascade object detection with deformable part models,” in CVPR. IEEE, 2010, pp. 2241–2248.
  15. 15.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, “Object detection with discriminatively trained part-based models,” IEEE transactions on pattern analysis and machine intelligence, vol. 32, no. 9, pp. 1627–1645, 2010.
  16. 16.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR, 2014, pp. 580–587.
  17. 17.K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in ECCV. Springer, 2014, pp. 346–361.
  18. 18.R. Girshick, “Fast r-cnn,” in ICCV, 2015, pp. 1440–1448.
  19. 19.S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems, 2015, pp. 91–99.
  20. 20.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in CVPR, 2016, pp. 779–788.
  21. 21.J. Redmon and A. Farhadi, “Yolov3: An incremental improvement,” arXiv preprint arXiv:1804.02767, 2018.
  22. 22.A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “Yolov4: Optimal speed and accuracy of object detection,” arXiv preprint arXiv:2004.10934, 2020.
  23. 23.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in ECCV. Springer, 2016, pp. 21–37.
  24. 24.T.-Y. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection.” in CVPR, vol. 1, no. 2, 2017, p. 4.
  25. 25.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” IEEE transactions on pattern analysis and machine intelligence, 2018.
  26. 26.H. Law and J. Deng, “Cornernet: Detecting objects as paired keypoints,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 734–750.
  27. 27.Z.-Q. Zhao, P. Zheng, S.-t. Xu, and X. Wu, “Object detection with deep learning: A review,” IEEE transactions on neural networks and learning systems, vol. 30, no. 11, pp. 3212–3232, 2019.
  28. 28.N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European Conference on Computer Vision. Springer, 2020, pp. 213–229.
  29. 29.D. G. Lowe, “Object recognition from local scale-invariant features,” in ICCV, vol. 2. Ieee, 1999, pp. 1150–1157.
  30. 30.——, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision, vol. 60, no. 2, pp. 91–110, 2004.
  31. 31.S. Belongie, J. Malik, and J. Puzicha, “Shape matching and object recognition using shape contexts,” CALIFORNIA UNIV SAN DIEGO LA JOLLA DEPT OF COMPUTER SCIENCE AND ENGINEERING, Tech. Rep., 2002.
  32. 32.T. Malisiewicz, A. Gupta, and A. A. Efros, “Ensemble of exemplar-svms for object detection and beyond,” in ICCV. IEEE, 2011, pp. 89–96.
  33. 33.R. B. Girshick, P. F. Felzenszwalb, and D. A. Mcallester, “Object detection with grammar models,” in Advances in Neural Information Processing Systems, 2011, pp. 442–450.
  34. 34.R. B. Girshick, From rigid templates to grammars: Object detection with structured models. Citeseer, 2012.
  35. 35.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems, 2012, pp. 1097–1105.
  36. 36.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Region-based convolutional networks for accurate object detection and segmentation,” IEEE transactions on pattern analysis and machine intelligence, vol. 38, no. 1, pp. 142–158, 2016.
  37. 37.M. A. Sadeghi and D. Forsyth, “30hz object detection with dpm v5,” in ECCV. Springer, 2014, pp. 65–79.
  38. 38.S. Zhang, L. Wen, X. Bian, Z. Lei, and S. Z. Li, “Single-shot refinement neural network for object detection,” in CVPR, 2018.
  39. 39.Y. Li, Y. Chen, N. Wang, and Z. Zhang, “Scale-aware trident networks for object detection,” arXiv preprint arXiv:1901.01892, 2019.
  40. 40.X. Zhou, D. Wang, and P. Krähenbühl, “Objects as points,” arXiv preprint arXiv:1904.07850, 2019.
  41. 41.Z. Tian, C. Shen, H. Chen, and T. He, “Fcos: Fully convolutional one-stage object detection,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9627–9636.
  42. 42.K. Chen, J. Pang, J. Wang, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Shi, W. Ouyang et al., “Hybrid task cascade for instance segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4974–4983.
  43. 43.X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” arXiv preprint arXiv:2010.04159, 2020.
  44. 44.Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” arXiv preprint arXiv:2103.14030, 2021.
  45. 45.J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, “Selective search for object recognition,” International journal of computer vision, vol. 104, no. 2, pp. 154–171, 2013.
  46. 46.R. B. Girshick, P. F. Felzenszwalb, and D. McAllester, “Discriminatively trained deformable part models, release 5,” http://people.cs.uchicago.edu/ rbg/latent-release5/.
  47. 47.S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis & Machine Intelligence, no. 6, pp. 1137–1149, 2017.
  48. 48.M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV. Springer, 2014, pp. 818–833.
  49. 49.J. Dai, Y. Li, K. He, and J. Sun, “R-fcn: Object detection via region-based fully convolutional networks,” in Advances in neural information processing systems, 2016, pp. 379–387.
  50. 50.Z. Li, C. Peng, G. Yu, X. Zhang, Y. Deng, and J. Sun, “Light-head r-cnn: In defense of two-stage object detector,” arXiv preprint arXiv:1711.07264, 2017.
  51. 51.J. Redmon and A. Farhadi, “Yolo9000: better, faster, stronger,” arXiv preprint, 2017.
  52. 52.C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,” arXiv preprint arXiv:2207.02696, 2022.
  53. 53.X. Zhou, J. Zhuo, and P. Krahenbuhl, “Bottom-up object detection by grouping extreme and center points,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 850–859.
  54. 54.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International journal of computer vision, vol. 88, no. 2, pp. 303–338, 2010.
  55. 55.M. Everingham, S. A. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes challenge: A retrospective,” International journal of computer vision, vol. 111, no. 1, pp. 98–136, 2015.
  56. 56.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” International Journal of Computer Vision, vol. 115, no. 3, pp. 211–252, 2015.
  57. 57.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in ECCV. Springer, 2014, pp. 740–755.
  58. 58.A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, T. Duerig, and V. Ferrari, “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,” IJCV, 2020.
  59. 59.R. Benenson, S. Popov, and V. Ferrari, “Large-scale interactive object segmentation with human annotators,” in CVPR, 2019.
  60. 60.S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 8430–8439.
  61. 61.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR. Ieee, 2009, pp. 248–255.
  62. 62.I. Krasin and T. e. a. Duerig, “Openimages: A public dataset for large-scale multi-label and multi-class image classification.” Dataset available from https://storage.googleapis.com/openimages/web/index.html, 2017.
  63. 63.P. Dollár, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: A benchmark,” in CVPR. IEEE, 2009, pp. 304–311.
  64. 64.P. Dollár, C. Wojek, B. Schiele, and P. Perona, “Pedestrian detection: An evaluation of the state of the art,” IEEE transactions on pattern analysis and machine intelligence, vol. 34, no. 4, pp. 743–761, 2012.
  65. 65.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “Overfeat: Integrated recognition, localization and detection using convolutional networks,” arXiv preprint arXiv:1312.6229, 2013.
  66. 66.C. Szegedy, A. Toshev, and D. Erhan, “Deep neural networks for object detection,” in Advances in neural information processing systems, 2013, pp. 2553–2561.
  67. 67.Z. Cai, Q. Fan, R. S. Feris, and N. Vasconcelos, “A unified multi-scale deep convolutional neural network for fast object detection,” in ECCV. Springer, 2016, pp. 354–370.
  68. 68.Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6154–6162.
  69. 69.Z. Yang, S. Liu, H. Hu, L. Wang, and S. Lin, “Reppoints: Point set representation for object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 9657–9666.
  70. 70.T. Malisiewicz, Exemplar-based representations for object detection, association and beyond. Carnegie Mellon University, 2011.
  71. 71.J. Hosang, R. Benenson, P. Dollár, and B. Schiele, “What makes for effective detection proposals?” IEEE transactions on pattern analysis and machine intelligence, vol. 38, no. 4, pp. 814–830, 2016.
  72. 72.J. Hosang, R. Benenson, and B. Schiele, “How good are detection proposals, really?” arXiv preprint arXiv:1406.6962, 2014.
  73. 73.B. Alexe, T. Deselaers, and V. Ferrari, “What is an object?” in CVPR. IEEE, 2010, pp. 73–80.
  74. 74.——, “Measuring the objectness of image windows,” IEEE transactions on pattern analysis and machine intelligence, vol. 34, no. 11, pp. 2189–2202, 2012.
  75. 75.M.-M. Cheng, Z. Zhang, W.-Y. Lin, and P. Torr, “Bing: Binarized normed gradients for objectness estimation at 300fps,” in CVPR, 2014, pp. 3286–3293.
  76. 76.D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov, “Scalable object detection using deep neural networks,” in CVPR, 2014, pp. 2147–2154.
  77. 77.K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, and Q. Tian, “Centernet: Keypoint triplets for object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6569–6578.
  78. 78.A. Torralba and P. Sinha, “Detecting faces in impoverished images,” MASSACHUSETTS INST OF TECH CAMBRIDGE ARTIFICIAL INTELLIGENCE LAB, Tech. Rep., 2001.
  79. 79.S. Zagoruyko, A. Lerer, T.-Y. Lin, P. O. Pinheiro, S. Gross, S. Chintala, and P. Dollár, “A multi-path network for object detection,” arXiv preprint arXiv:1604.02135, 2016.
  80. 80.X. Zeng, W. Ouyang, B. Yang, J. Yan, and X. Wang, “Gated bi-directional cnn for object detection,” in ECCV. Springer, 2016, pp. 354–369.
  81. 81.X. Zeng, W. Ouyang, J. Yan, H. Li, T. Xiao, K. Wang, Y. Liu, Y. Zhou, B. Yang, Z. Wang et al., “Crafting gbdnet for object detection,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 9, pp. 2109–2123, 2018.
  82. 82.W. Ouyang, K. Wang, X. Zhu, and X. Wang, “Learning chained deep features and classifiers for cascade in object detection,” arXiv preprint arXiv:1702.07054, 2017.
  83. 83.S. Gidaris and N. Komodakis, “Object detection via a multi-region and semantic segmentation-aware cnn model,” in ICCV, 2015, pp. 1134–1142.
  84. 84.Y. Zhu, C. Zhao, J. Wang, X. Zhao, Y. Wu, H. Lu et al., “Couplenet: Coupling global structure with local parts for object detection,” in ICCV, vol. 2, 2017.
  85. 85.C. Desai, D. Ramanan, and C. C. Fowlkes, “Discriminative models for multi-class object layout,” International journal of computer vision, vol. 95, no. 1, pp. 1–12, 2011.
  86. 86.S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick, “Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks,” in CVPR, 2016, pp. 2874–2883.
  87. 87.Z. Li, Y. Chen, G. Yu, and Y. Deng, “R-fcn++: Towards accurate region-based fully convolutional networks for object detection.” in AAAI, 2018.
  88. 88.S. Liu, D. Huang et al., “Receptive field block net for accurate and fast object detection,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 385–400.
  89. 89.X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803.
  90. 90.Q. Chen, Z. Song, J. Dong, Z. Huang, Y. Hua, and S. Yan, “Contextualizing object detection and classification,” IEEE transactions on pattern analysis and machine intelligence, vol. 37, no. 1, pp. 13–27, 2015.
  91. 91.S. Gupta, B. Hariharan, and J. Malik, “Exploring person context and local scene context for object detection,” arXiv preprint arXiv:1511.08177, 2015.
  92. 92.X. Chen and A. Gupta, “Spatial memory for context reasoning in object detection,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 4086–4096.
  93. 93.H. Hu, J. Gu, Z. Zhang, J. Dai, and Y. Wei, “Relation networks for object detection,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3588–3597.
  94. 94.Y. Liu, R. Wang, S. Shan, and X. Chen, “Structure inference net: Object detection using scene-level context and instance-level relationships,” in CVPR, 2018, pp. 6985–6994.
  95. 95.L. V. Pato, R. Negrinho, and P. M. Q. Aguiar, “Seeing without looking: Contextual rescoring of object detections for ap maximization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  96. 96.S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Hebert, “An empirical study of context in object detection,” in CVPR. IEEE, 2009, pp. 1271–1278.
  97. 97.C. Chen, M.-Y. Liu, O. Tuzel, and J. Xiao, “R-cnn for small object detection,” in Asian conference on computer vision. Springer, 2016, pp. 214–230.
  98. 98.J. Li, Y. Wei, X. Liang, J. Dong, T. Xu, J. Feng, and S. Yan, “Attentive contexts for object detection,” IEEE Transactions on Multimedia, vol. 19, no. 5, pp. 944–954, 2017.
  99. 99.H. A. Rowley, S. Baluja, and T. Kanade, “Human face detection in visual scenes,” in Advances in Neural Information Processing Systems, 1996, pp. 875–881.
  100. 100.C. P. Papageorgiou, M. Oren, and T. Poggio, “A general framework for object detection,” in ICCV. IEEE, 1998, pp. 555–562.
  101. 101.L. Zhang, L. Lin, X. Liang, and K. He, “Is faster r-cnn doing well for pedestrian detection?” in ECCV. Springer, 2016, pp. 443–457.
  102. 102.A. Shrivastava, A. Gupta, and R. Girshick, “Training region-based object detectors with online hard example mining,” in CVPR, 2016, pp. 761–769.
  103. 103.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 2818–2826.
  104. 104.R. Müller, S. Kornblith, and G. E. Hinton, “When does label smoothing help?” Advances in neural information processing systems, vol. 32, 2019.
  105. 105.J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang, “Unitbox: An advanced object detection network,” in Proceedings of the 2016 ACM on Multimedia Conference. ACM, 2016, pp. 516–520.
  106. 106.H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 658–666.
  107. 107.Z. Zheng, P. Wang, W. Liu, J. Li, R. Ye, and D. Ren, “Distance-iou loss: Faster and better learning for bounding box regression,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 12 993–13 000.
  108. 108.R. Vaillant, C. Monrocq, and Y. Le Cun, “Original approach for the localisation of objects in images,” IEE Proceedings-Vision, Image and Signal Processing, vol. 141, no. 4, pp. 245–250, 1994.
  109. 109.P. Henderson and V. Ferrari, “End-to-end training of object class detectors for mean average precision,” in Asian Conference on Computer Vision. Springer, 2016, pp. 198–213.
  110. 110.J. H. Hosang, R. Benenson, and B. Schiele, “Learning non-maximum suppression.” in CVPR, 2017, pp. 6469–6477.
  111. 111.Z. Tan, X. Nie, Q. Qian, N. Li, and H. Li, “Learning to rank proposals for object detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 8273–8281.
  112. 112.N. Bodla, B. Singh, R. Chellappa, and L. S. Davis, “Soft-nms—improving object detection with one line of code,” in ICCV. IEEE, 2017, pp. 5562–5570.
  113. 113.L. Tychsen-Smith and L. Petersson, “Improving object localization with fitness nms and bounded iou loss,” arXiv preprint arXiv:1711.00164, 2017.
  114. 114.Y. He, C. Zhu, J. Wang, M. Savvides, and X. Zhang, “Bounding box regression with uncertainty for accurate object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 2888–2897.
  115. 115.S. Liu, D. Huang, and Y. Wang, “Adaptive nms: Refining pedestrian detection in a crowd,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 6459–6468.
  116. 116.R. Rothe, M. Guillaumin, and L. Van Gool, “Non-maximum suppression for object detection by passing messages between windows,” in Asian Conference on Computer Vision. Springer, 2014, pp. 290–306.
  117. 117.D. Mrowca, M. Rohrbach, J. Hoffman, R. Hu, K. Saenko, and T. Darrell, “Spatial semantic regularisation for large scale object detection,” in ICCV, 2015, pp. 2003–2011.
  118. 118.R. Solovyev, W. Wang, and T. Gabruseva, “Weighted boxes fusion: Ensembling boxes from different object detection models,” Image and Vision Computing, vol. 107, p. 104117, 2021.
  119. 119.Z. Zheng, P. Wang, D. Ren, W. Liu, R. Ye, Q. Hu, and W. Zuo, “Enhancing geometric factors in model learning and inference for object detection and instance segmentation,” IEEE Transactions on Cybernetics, 2021.
  120. 120.J. Wang, L. Song, Z. Li, H. Sun, J. Sun, and N. Zheng, “End-to-end object detection with fully convolutional network,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 849–15 858.
  121. 121.C. Papageorgiou and T. Poggio, “A trainable system for object detection,” International journal of computer vision, vol. 38, no. 1, pp. 15–33, 2000.
  122. 122.L. Wan, D. Eigen, and R. Fergus, “End-to-end integration of a convolution network, deformable parts model and non-maximum suppression,” in CVPR, 2015, pp. 851–859.
  123. 123.Z. Zou, Z. Shi, Y. Guo, and J. Ye, “Object detection in 20 years: A survey,” arXiv preprint arXiv:1905.05055, 2019.
  124. 124.Q. Zhu, M.-C. Yeh, K.-T. Cheng, and S. Avidan, “Fast human detection using a cascade of histograms of oriented gradients,” in CVPR, vol. 2. IEEE, 2006, pp. 1491–1498.
  125. 125.F. Fleuret and D. Geman, “Coarse-to-fine face detection,” International Journal of computer vision, vol. 41, no. 1-2, pp. 85–107, 2001.
  126. 126.H. Li, Z. Lin, X. Shen, J. Brandt, and G. Hua, “A convolutional neural network cascade for face detection,” in CVPR, 2015, pp. 5325–5334.
  127. 127.K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, “Joint face detection and alignment using multitask cascaded convolutional networks,” IEEE Signal Processing Letters, vol. 23, no. 10, pp. 1499–1503, 2016.
  128. 128.Z. Cai, M. Saberian, and N. Vasconcelos, “Learning complexity-aware cascades for deep pedestrian detection,” in ICCV, 2015, pp. 3361–3369.
  129. 129.Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in Advances in neural information processing systems, 1990, pp. 598–605.
  130. 130.S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” arXiv preprint arXiv:1510.00149, 2015.
  131. 131.K. He and J. Sun, “Convolutional neural networks at constrained time cost,” in CVPR, 2015, pp. 5353–5360.
  132. 132.Z. Qin, Z. Li, Z. Zhang, Y. Bao, G. Yu, Y. Peng, and J. Sun, “Thundernet: Towards real-time generic object detection on mobile devices,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6718–6727.
  133. 133.R. J. Wang, X. Li, and C. X. Ling, “Pelee: A real-time object detection system on mobile devices,” in Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds. Curran Associates, Inc., 2018, pp. 1967–1976.
  134. 134.R. Huang, J. Pedoeem, and C. Chen, “Yolo-lite: a realtime object detection algorithm optimized for non-gpu computers,” in 2018 IEEE International Conference on Big Data (Big Data). IEEE, 2018, pp. 2503–2510.
  135. 135.H. Law, Y. Teng, O. Russakovsky, and J. Deng, “Cornernet-lite: Efficient keypoint based object detection,” arXiv preprint arXiv:1904.08900, 2019.
  136. 136.G. Yu, Q. Chang, W. Lv, C. Xu, C. Cui, W. Ji, Q. Dang, K. Deng, G. Wang, Y. Du et al., “Pp-picodet: A better real-time object detector on mobile devices,” arXiv preprint arXiv:2111.00902, 2021.
  137. 137.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in CVPR, 2016, pp. 2818–2826.
  138. 138.X. Zhang, J. Zou, X. Ming, K. He, and J. Sun, “Efficient and accurate approximations of nonlinear convolutional networks,” in CVPR, 2015, pp. 1984–1992.
  139. 139.X. Zhang, J. Zou, K. He, and J. Sun, “Accelerating very deep convolutional networks for classification and detection,” IEEE transactions on pattern analysis and machine intelligence, vol. 38, no. 10, pp. 1943–1955, 2016.
  140. 140.X. Zhang, X. Zhou, M. Lin, and J. Sun, “Shufflenet: An extremely efficient convolutional neural network for mobile devices,” 2017.
  141. 141.G. Huang, S. Liu, L. van der Maaten, and K. Q. Weinberger, “Condensenet: An efficient densenet using learned group convolutions,” group, vol. 3, no. 12, p. 11, 2017.
  142. 142.F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” arXiv preprint, pp. 1610–02 357, 2017.
  143. 143.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
  144. 144.M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR. IEEE, 2018, pp. 4510–4520.
  145. 145.Y. Li, J. Li, W. Lin, and J. Li, “Tiny-dsod: Lightweight object detection for resource-restricted usages,” arXiv preprint arXiv:1807.11013, 2018.
  146. 146.F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size,” arXiv preprint arXiv:1602.07360, 2016.
  147. 147.B. Wu, F. N. Iandola, P. H. Jin, and K. Keutzer, “Squeezedet: Unified, small, low power fully convolutional neural networks for real-time object detection for autonomous driving.” in CVPR Workshops, 2017, pp. 446–454.
  148. 148.T. Kong, A. Yao, Y. Chen, and F. Sun, “Hypernet: Towards accurate region proposal generation and joint object detection,” in CVPR, 2016, pp. 845–853.
  149. 149.Y. Chen, T. Yang, X. Zhang, G. Meng, C. Pan, and J. Sun, “Detnas: Neural architecture search on object detection,” arXiv preprint arXiv:1903.10979, 2019.
  150. 150.H. Xu, L. Yao, W. Zhang, X. Liang, and Z. Li, “Auto-fpn: Automatic network architecture adaptation for object detection beyond classification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6649–6658.
  151. 151.G. Ghiasi, T.-Y. Lin, and Q. V. Le, “Nas-fpn: Learning scalable feature pyramid architecture for object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 7036–7045.
  152. 152.J. Guo, K. Han, Y. Wang, C. Zhang, Z. Yang, H. Wu, X. Chen, and C. Xu, “Hit-detector: Hierarchical trinity architecture search for object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11 405–11 414.
  153. 153.N. Wang, Y. Gao, H. Chen, P. Wang, Z. Tian, C. Shen, and Y. Zhang, “Nas-fcos: Fast neural architecture search for object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11 943–11 951.
  154. 154.L. Yao, H. Xu, W. Zhang, X. Liang, and Z. Li, “Sm-nas: structural-to-modular neural architecture search for object detection,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 12 661–12 668.
  155. 155.C. Jiang, H. Xu, W. Zhang, X. Liang, and Z. Li, “Sp-nas: Serial-to-parallel backbone search for object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11 863–11 872.
  156. 156.P. Simard, L. Bottou, P. Haffner, and Y. LeCun, “Boxlets: a fast convolution algorithm for signal processing and neural networks,” in Advances in Neural Information Processing Systems, 1999, pp. 571–577.
  157. 157.X. Wang, T. X. Han, and S. Yan, “An hog-lbp human detector with partial occlusion handling,” in ICCV. IEEE, 2009, pp. 32–39.
  158. 158.F. Porikli, “Integral histogram: A fast way to extract histograms in cartesian spaces,” in CVPR, vol. 1. IEEE, 2005, pp. 829–836.
  159. 159.P. Dollár, Z. Tu, P. Perona, and S. Belongie, “Integral channel features,” 2009.
  160. 160.M. Mathieu, M. Henaff, and Y. LeCun, “Fast training of convolutional networks through ffts,” arXiv preprint arXiv:1312.5851, 2013.
  161. 161.H. Pratt, B. Williams, F. Coenen, and Y. Zheng, “Fcnn: Fourier convolutional neural networks,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 2017, pp. 786–798.
  162. 162.N. Vasilache, J. Johnson, M. Mathieu, S. Chintala, S. Piantino, and Y. LeCun, “Fast convolutional nets with fbfft: A gpu performance evaluation,” arXiv preprint arXiv:1412.7580, 2014.
  163. 163.O. Rippel, J. Snoek, and R. P. Adams, “Spectral representations for convolutional neural networks,” in Advances in neural information processing systems, 2015, pp. 2449–2457.
  164. 164.M. A. Sadeghi and D. Forsyth, “Fast template evaluation with vector quantization,” in Advances in neural information processing systems, 2013, pp. 2949–2957.
  165. 165.I. Kokkinos, “Bounding part scores for rapid detection with deformable part models,” in ECCV. Springer, 2012, pp. 41–50.
  166. 166.H. Zhu, X. Chen, W. Dai, K. Fu, Q. Ye, and J. Jiao, “Orientation robust object detection in aerial images using deep convolutional neural network,” in ICIP. IEEE, 2015, pp. 3735–3739.
  167. 167.B. Cai, Z. Jiang, H. Zhang, Y. Yao, and S. Nie, “Online exemplar-based fully convolutional network for aircraft detection in remote sensing images,” IEEE Geoscience and Remote Sensing Letters, no. 99, pp. 1–5, 2018.
  168. 168.G. Cheng, J. Han, P. Zhou, and L. Guo, “Multi-class geospatial object detection and geographic image classification based on collection of part detectors,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 98, pp. 119–132, 2014.
  169. 169.G. Cheng, P. Zhou, and J. Han, “Rifd-cnn: Rotation-invariant and fisher discriminative convolutional neural networks for object detection,” in CVPR, 2016, pp. 2884–2893.
  170. 170.——, “Learning rotation-invariant convolutional neural networks for object detection in vhr optical remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 54, no. 12, pp. 7405–7415, 2016.
  171. 171.G. Cheng, J. Han, P. Zhou, and D. Xu, “Learning rotation-invariant and fisher discriminative convolutional neural networks for object detection,” IEEE Transactions on Image Processing, vol. 28, no. 1, pp. 265–278, 2018.
  172. 172.X. Shi, S. Shan, M. Kan, S. Wu, and X. Chen, “Real-time rotation-invariant face detection with progressive calibration networks,” in CVPR, 2018, pp. 2295–2303.
  173. 173.M. Jaderberg, K. Simonyan, A. Zisserman et al., “Spatial transformer networks,” in Advances in neural information processing systems, 2015, pp. 2017–2025.
  174. 174.D. Chen, G. Hua, F. Wen, and J. Sun, “Supervised transformer network for efficient face detection,” in ECCV. Springer, 2016, pp. 122–138.
  175. 175.J. Ding, N. Xue, Y. Long, G.-S. Xia, and Q. Lu, “Learning roi transformer for oriented object detection in aerial images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 2849–2858.
  176. 176.B. Singh and L. S. Davis, “An analysis of scale invariance in object detection–snip,” in CVPR, 2018, pp. 3578–3587.
  177. 177.B. Singh, M. Najibi, and L. S. Davis, “Sniper: Efficient multi-scale training,” arXiv preprint arXiv:1805.09300, 2018.
  178. 178.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778.
  179. 179.M. Gao, R. Yu, A. Li, V. I. Morariu, and L. S. Davis, “Dynamic zoom-in network for fast object detection in large images,” in CVPR, 2018.
  180. 180.Y. Lu, T. Javidi, and S. Lazebnik, “Adaptive object detection using adjacency and zoom prediction,” in CVPR, 2016, pp. 2351–2359.
  181. 181.S. Qiao, W. Shen, W. Qiu, C. Liu, and A. L. Yuille, “Scalenet: Guiding object proposal generation in supermarkets and beyond.” in ICCV, 2017, pp. 1809–1818.
  182. 182.Z. Hao, Y. Liu, H. Qin, J. Yan, X. Li, and X. Hu, “Scale-aware face detection,” in CVPR, vol. 3, 2017.
  183. 183.C.-Y. Wang, H.-Y. M. Liao, Y.-H. Wu, P.-Y. Chen, J.-W. Hsieh, and I.-H. Yeh, “Cspnet: A new backbone that can enhance learning capability of cnn,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, 2020, pp. 390–391.
  184. 184.A. Newell, K. Yang, and J. Deng, “Stacked hourglass networks for human pose estimation,” in European conference on computer vision. Springer, 2016, pp. 483–499.
  185. 185.J. Gu, Z. Wang, J. Kuen, L. Ma, A. Shahroudy, B. Shuai, T. Liu, X. Wang, L. Wang, G. Wang et al., “Recent advances in convolutional neural networks,” arXiv preprint arXiv:1512.07108, 2015.
  186. 186.J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama et al., “Speed/accuracy trade-offs for modern convolutional object detectors,” in CVPR, vol. 4, 2017.
  187. 187.Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in CVPR, vol. 1, no. 2, 2018, p. 10.
  188. 188.R. N. Rajaram, E. Ohn-Bar, and M. M. Trivedi, “Refinenet: Iterative refinement for accurate object localization,” in ITSC. IEEE, 2016, pp. 1528–1533.
  189. 189.M.-C. Roh and J.-y. Lee, “Refining faster-rcnn for accurate object detection,” in Machine Vision Applications (MVA), 2017 Fifteenth IAPR International Conference on. IEEE, 2017, pp. 514–517.
  190. 190.B. Jiang, R. Luo, J. Mao, T. Xiao, and Y. Jiang, “Acquisition of localization confidence for accurate object detection,” in Proceedings of the ECCV, Munich, Germany, 2018, pp. 8–14.
  191. 191.S. Gidaris and N. Komodakis, “Locnet: Improving localization accuracy for object detection,” in CVPR, 2016, pp. 789–798.
  192. 192.S. Brahmbhatt, H. I. Christensen, and J. Hays, “Stuffnet: Using ‘stuff’to improve object detection,” in Applications of Computer Vision (WACV), 2017 IEEE Winter Conference on. IEEE, 2017, pp. 934–943.
  193. 193.A. Shrivastava and A. Gupta, “Contextual priming and feedback for faster r-cnn,” in ECCV. Springer, 2016, pp. 330–348.
  194. 194.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems, 2014, pp. 2672–2680.
  195. 195.A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” arXiv preprint arXiv:1511.06434, 2015.
  196. 196.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” arXiv preprint, 2017.
  197. 197.C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. P. Aitken, A. Tejani, J. Totz, Z. Wang et al., “Photo-realistic single image super-resolution using a generative adversarial network.” in CVPR, vol. 2, no. 3, 2017, p. 4.
  198. 198.J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan, “Perceptual generative adversarial networks for small object detection,” in CVPR, 2017.
  199. 199.Y. Bai, Y. Zhang, M. Ding, and B. Ghanem, “Sodmtgan: Small object detection via multi-task generative adversarial network,” Computer Vision–ECCV, pp. 8–14, 2018.
  200. 200.X. Wang, A. Shrivastava, and A. Gupta, “A-fast-rcnn: Hard positive generation via adversary for object detection,” in CVPR, 2017.
  201. 201.D. Zhang, J. Han, G. Cheng, and M.-H. Yang, “Weakly supervised object localization and detection: A survey,” IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 9, pp. 5866–5885, 2021.
  202. 202.T. G. Dietterich, R. H. Lathrop, and T. Lozano-Pérez, “Solving the multiple instance problem with axis-parallel rectangles,” Artificial intelligence, vol. 89, no. 1-2, pp. 31–71, 1997.
  203. 203.S. Andrews, I. Tsochantaridis, and T. Hofmann, “Support vector machines for multiple-instance learning,” in Advances in neural information processing systems, 2003, pp. 577–584.
  204. 204.R. G. Cinbis, J. Verbeek, and C. Schmid, “Weakly supervised object localization with multi-fold multiple instance learning,” IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 1, pp. 189–203, 2017.
  205. 205.D. P. Papadopoulos, J. R. Uijlings, F. Keller, and V. Ferrari, “We don’t need no bounding-boxes: Training object class detectors using only human verification,” in CVPR, 2016, pp. 854–863.
  206. 206.D. Zhang, W. Zeng, J. Yao, and J. Han, “Weakly supervised object detection using proposal-and semantic-level relationships,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  207. 207.P. Tang, X. Wang, S. Bai, W. Shen, X. Bai, W. Liu, and A. Yuille, “Pcl: Proposal cluster learning for weakly supervised object detection,” IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 1, pp. 176–191, 2018.
  208. 208.E. Sangineto, M. Nabi, D. Culibrk, and N. Sebe, “Self paced deep learning for weakly supervised object detection,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 3, pp. 712–725, 2018.
  209. 209.D. Zhang, J. Han, L. Zhao, and D. Meng, “Leveraging prior-knowledge for weakly supervised object detection under a collaborative self-paced curriculum learning framework,” International Journal of Computer Vision, vol. 127, no. 4, pp. 363–380, 2019.
  210. 210.Y. Zhu, Y. Zhou, Q. Ye, Q. Qiu, and J. Jiao, “Soft proposal networks for weakly supervised object localization,” in ICCV, 2017, pp. 1841–1850.
  211. 211.A. Diba, V. Sharma, A. M. Pazandeh, H. Pirsiavash, and L. Van Gool, “Weakly supervised cascaded convolutional networks.” in CVPR, vol. 1, no. 2, 2017, p. 8.
  212. 212.B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in CVPR, 2016, pp. 2921–2929.
  213. 213.H. Bilen and A. Vedaldi, “Weakly supervised deep detection networks,” in CVPR, 2016, pp. 2846–2854.
  214. 214.L. Bazzani, A. Bergamo, D. Anguelov, and L. Torresani, “Self-taught object localization with deep networks,” in Applications of Computer Vision (WACV), 2016 IEEE Winter Conference on. IEEE, 2016, pp. 1–9.
  215. 215.Y. Shen, R. Ji, S. Zhang, W. Zuo, and Y. Wang, “Generative adversarial learning towards fast weakly supervised detection,” in CVPR, 2018, pp. 5764–5773.
  216. 216.Y. Chen, W. Li, C. Sakaridis, D. Dai, and L. Van Gool, “Domain adaptive faster r-cnn for object detection in the wild,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3339–3348.
  217. 217.Y. Wang, R. Zhang, S. Zhang, M. Li, Y. Xia, X. Zhang, and S. Liu, “Domain-specific suppression for adaptive object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9603–9612.
  218. 218.L. Hou, Y. Zhang, K. Fu, and J. Li, “Informative and consistent correspondence mining for cross-domain weakly supervised object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9929–9938.
  219. 219.X. Zhu, J. Pang, C. Yang, J. Shi, and D. Lin, “Adapting object detectors via selective cross-domain alignment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 687–696.
  220. 220.K. Saito, Y. Ushiku, T. Harada, and K. Saenko, “Strong-weak distribution alignment for adaptive object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 6956–6965.
  221. 221.C.-D. Xu, X.-R. Zhao, X. Jin, and X.-S. Wei, “Exploring categorical regularization for domain adaptive object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 11 724–11 733.
  222. 222.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2223–2232.
  223. 223.T. Kim, M. Jeong, S. Kim, S. Choi, and C. Kim, “Diversify and match: A domain adaptive representation learning paradigm for object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 12 456–12 465.
  224. 224.N. Inoue, R. Furuta, T. Yamasaki, and K. Aizawa, “Cross-domain weakly-supervised object detection through progressive domain adaptation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 5001–5009.
  225. 225.H.-K. Hsu, C.-H. Yao, Y.-H. Tsai, W.-C. Hung, H.-Y. Tseng, M. Singh, and M.-H. Yang, “Progressive domain adaptation for object detection,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020, pp. 749–757.
  226. 226.B. Bosquet, M. Mucientes, and V. M. Brea, “Stdnet-st: Spatio-temporal convnet for small object detection,” Pattern Recognition, vol. 116, p. 107929, 2021.
  227. 227.C. Yang, Z. Huang, and N. Wang, “Querydet: Cascaded sparse query for accelerating high-resolution small object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 13 668–13 677.
  228. 228.P. Sun, Y. Jiang, E. Xie, W. Shao, Z. Yuan, C. Wang, and P. Luo, “What makes for end-to-end object detection?” in International Conference on Machine Learning. PMLR, 2021, pp. 9934–9944.
  229. 229.X. Zhou, X. Xu, W. Liang, Z. Zeng, S. Shimizu, L. T. Yang, and Q. Jin, “Intelligent small object detection for digital twin in smart manufacturing with industrial cyber-physical systems,” IEEE Transactions on Industrial Informatics, vol. 18, no. 2, pp. 1377–1386, 2021.
  230. 230.G. Cheng, X. Yuan, X. Yao, K. Yan, Q. Zeng, and J. Han, “Towards large-scale small object detection: Survey and benchmarks,” arXiv preprint arXiv:2207.14096, 2022.
  231. 231.Y. Wang, V. C. Guizilini, T. Zhang, Y. Wang, H. Zhao, and J. Solomon, “Detr3d: 3d object detection from multi-view images via 3d-to-2d queries,” in Conference on Robot Learning. PMLR, 2022, pp. 180–191.
  232. 232.Y. Wang, T. Ye, L. Cao, W. Huang, F. Sun, F. He, and D. Tao, “Bridged transformer for vision and point cloud 3d object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12 114–12 123.
  233. 233.X. Cheng, H. Xiong, D.-P. Fan, Y. Zhong, M. Harandi, T. Drummond, and Z. Ge, “Implicit motion handling for video camouflaged object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 13 864–13 873.
  234. 234.Q. Zhou, X. Li, L. He, Y. Yang, G. Cheng, Y. Tong, L. Ma, and D. Tao, “Transvod: End-to-end video object detection with spatial-temporal transformers,” arXiv preprint arXiv:2201.05047, 2022.
  235. 235.R. Cong, Q. Lin, C. Zhang, C. Li, X. Cao, Q. Huang, and Y. Zhao, “Cir-net: Cross-modality interaction and refinement for rgb-d salient object detection,” IEEE Transactions on Image Processing, 2022.
  236. 236.Y. Wang, L. Zhu, S. Huang, T. Hui, X. Li, F. Wang, and S. Liu, “Cross-modality domain adaptation for freespace detection: A simple yet effective baseline,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 4031–4042.
  237. 237.C. Feng, Y. Zhong, Z. Jie, X. Chu, H. Ren, X. Wei, W. Xie, and L. Ma, “Promptdet: Expand your detector vocabulary with uncurated images,” arXiv preprint arXiv:2203.16513, 2022.
  238. 238.Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li et al., “Regionclip: Region-based language-image pretraining,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 16 793–16 803.

Citation

MLA
Zou, Z., et al. “Object Detection in 20 Years: A Survey”. arXiv, 2019, http://arxiv.org/abs/1905.05055v3.
APA
Zou, Z., Chen, K., Shi, Z., Guo, Y., & Ye, J. (2019). Object Detection in 20 Years: A Survey. arXiv. http://arxiv.org/abs/1905.05055v3
Chicago
Zou, Z., K. Chen, Z. Shi, Y. Guo, and J. Ye. 2019. “Object Detection in 20 Years: A Survey”. arXiv. http://arxiv.org/abs/1905.05055v3.
Harvard
Zou, Z. et al. (2019) “Object Detection in 20 Years: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1905.05055v3.
Vancouver
1. Zou Z, Chen K, Shi Z, Guo Y, Ye J (2019) Object Detection in 20 Years: A Survey. arXiv

BibTeX

@article{zou2019object,
  title = {Object Detection in 20 Years: A Survey},
  author = {Zou, Zhengxia and Chen, Keyan and Shi, Zhenwei and Guo, Yuhong and Ye, Jieping},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1905.05055v3},
  eprint = {1905.05055}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF