Object Detection With Deep Learning: A Review

Zhong-Qiu ZhaoPeng ZhengShou-tao XuXindong Wu

article2018IEEE TNNLS4,860 citations

Systematizes the foundational architectures, optimization strategies, and practical training techniques of deep learning-based object detection while evaluating their performance across general benchmarks and specialized domains like face and pedestrian detection.

Listen

Deep learning has transformed object detection by replacing handcrafted features and shallow classifiers with convolutional networks that learn rich semantic representations directly from data. The review examines this shift across generic detection and three specialized taskssalient-object, face, and pedestrian detectionbecause accurate localization and classification underpin applications such as autonomous driving, surveillance, and image understanding. Traditional sliding-window pipelines stagnated after 2010 because exhaustive candidate generation was inefficient and low-level descriptors could not bridge the semantic gap; the arrival of large-scale labeled sets and GPU training removed those barriers and enabled end-to-end learning.

The authors set out to map the principal architectures, training strategies, and performance trade-offs that have emerged since the 2014 R-CNN breakthrough, while also supplying benchmark comparisons and forward-looking guidance. They synthesize the literature on region-proposal and single-shot regression families, trace the evolution of backbone networks, and evaluate representative methods on PASCAL VOC, Microsoft COCO, FDDB, and Caltech Pedestrian data sets.

Region-proposal pipelines (R-CNNFast/Faster R-CNNR-FCN, FPN, Mask R-CNN) deliver the highest accuracy by decoupling candidate generation from classification and bounding-box regression, yet they remain multi-stage and comparatively slow. Single-shot regressors (YOLO, SSD and variants) achieve real-time ratesoften above 30 fpsby casting detection as a unified grid or anchor-based prediction task, at a modest cost in localization precision. Across both families, multi-scale feature pyramids, hard-negative mining, and joint optimization of classification and regression consistently raise mean average precision by 515 points; the same ingredients also improve robustness on small or occluded instances that dominate challenging data sets such as COCO. Face and pedestrian detectors further benefit from part-based or scale-adaptive extensions, while salient-object methods gain from multi-context and boundary-aware supervision.

These advances translate directly into deployable systems: Faster R-CNN and its descendants now underpin production pipelines that must balance accuracy against latency, and the review shows that careful backbone selection plus batch normalization can cut inference time by an order of magnitude without sacrificing more than a few points of mAP. At the same time, the gap between laboratory benchmarks and real-world conditions remains large; even the best reported figures on COCO hover well below 40 % mAP when strict localization is required.

Future progress hinges on three practical directions. First, scale-adaptive and context-aware architectures must be made end-to-end trainable so that small-object performance improves without exhaustive image pyramids. Second, weakly supervised and self-supervised pre-training schemes are needed to reduce reliance on costly bounding-box annotations. Third, compact, hardware-aware modelsachieved through knowledge distillation, pruning, or training-from-scratch techniques such as DSODshould be pursued to meet the latency and memory constraints of embedded platforms. The review itself is limited to publications available through mid-2018; subsequent work on transformers and 3-D sensing will require periodic re-evaluation, yet the core architectural lessons remain a reliable foundation for those extensions.

Cover for Object Detection With Deep Learning: A Review

Abstract

Due to object detection's close relationship with video analysis and image understanding, it has attracted much research attention in recent years. Traditional object detection methods are built on handcrafted features and shallow trainable architectures. Their performance easily stagnates by constructing complex ensembles which combine multiple low-level image features with high-level context from object detectors and scene classifiers. With the rapid development in deep learning, more powerful tools, which are able to learn semantic, high-level, deeper features, are introduced to address the problems existing in traditional architectures. These models behave differently in network architecture, training strategy and optimization function, etc. In this paper, we provide a review on deep learning based object detection frameworks. Our review begins with a brief introduction on the history of deep learning and its representative tool, namely Convolutional Neural Network (CNN). Then we focus on typical generic object detection architectures along with some modifications and useful tricks to improve detection performance further. As distinct specific detection tasks exhibit different characteristics, we also briefly survey several specific tasks, including salient object detection, face detection and pedestrian detection. Experimental analyses are also provided to compare various methods and draw some meaningful conclusions. Finally, several promising directions and tasks are provided to serve as guidelines for future work in both object detection and relevant neural network based learning systems.

Table of Contents

  • I Introduction
  • II A Brief Overview of Deep Learning
  • II-A The History: Birth, Decline and Prosperity
  • II-B Architecture and Advantages of CNN
  • III Generic Object Detection
  • III-A Region Proposal Based Framework
  • III-A1 R-CNN
  • III-A2 SPP-net
  • III-A3 Fast R-CNN
  • III-A4 Faster R-CNN
  • III-A5 R-FCN
  • III-A6 FPN
  • III-A7 Mask R-CNN
  • III-A8 Multi-task Learning, Multi-scale Representation and Contextual Modelling
  • III-A9 Thinking in Deep Learning based Object Detection
  • III-B Regression//Classification Based Framework
  • III-B1 Pioneer Works
  • III-B2 YOLO
  • III-B3 SSD
  • III-C Experimental Evaluation
  • III-C1 PASCAL VOC 2007/2012
  • III-C2 Microsoft COCO
  • III-C3 Timing Analysis
  • IV Salient Object Detection
  • IV-A Deep learning in Salient Object Detection
  • IV-B Experimental Evaluation
  • V Face Detection
  • V-A Deep learning in Face Detection
  • V-B Experimental Evaluation
  • VI Pedestrian Detection
  • VI-A Deep learning in Pedestrian Detection
  • VI-B Experimental Evaluation
  • VII Promising Future Directions and Tasks
  • VIII Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy of Deep Learning Generic Object Detection Frameworks

    model/method

    Deep learning-based generic object detection frameworks are categorized into two primary paradigms based on their architectural pipeline:

    1. Region Proposal-Based (Two-Stage) Frameworks: This paradigm decomposes detection into candidate region extraction followed by classification and bounding-box refinement.

      • R-CNN: Generates approximately 2,000 candidate regions per image via bottom-up segmentation (such as Selective Search), warps each region to a fixed input size, extracts 4,096-dimensional representations using a Convolutional Neural Network (CNN), and classifies proposals using class-specific linear Support Vector Machines (SVMs) followed by bounding-box regression.
      • SPP-net: Incorporates a Spatial Pyramid Pooling layer after the final convolutional layer to accept arbitrary input image sizes and avoids redundant convolutional passes by extracting feature maps over the entire image once.
      • Fast R-CNN: Integrates feature extraction, classification, and bounding-box regression into a unified multi-task network via a Region of Interest (RoI) pooling layer.
      • Faster R-CNN: Introduces an in-network Region Proposal Network (RPN) that shares convolutional features with the downstream detection network, generating proposals via multi-scale and multi-aspect-ratio anchor boxes.
      • R-FCN: Utilizes fully convolutional networks with position-sensitive score maps to eliminate unshared per-RoI subnets.
      • Feature Pyramid Networks (FPN): Combines a bottom-up feature hierarchy with a top-down pathway and lateral connections to build semantically rich multi-scale feature pyramids.
      • Mask R-CNN: Extends Faster R-CNN by adding a parallel pixel-level instance mask branch and replacing RoI pooling with RoIAlign (bilinear interpolation) to preserve exact spatial alignment.
    2. Regression/Classification-Based (One-Stage) Frameworks: This paradigm eliminates the proposal generation step, mapping directly from input pixels to bounding-box coordinates and category probabilities in a single forward evaluation.

      • YOLO (You Only Look Once): Discretizes an image into an S×SS \times S spatial grid, where each grid cell simultaneously predicts BB bounding boxes, confidence scores, and CC conditional class probabilities.
      • SSD (Single Shot MultiBox Detector): Distributes default anchor boxes across multiple feature maps of varying spatial resolutions to detect objects of diverse scales in a single forward pass.
      • DSSD and DSOD: Enhance one-stage detection by integrating deconvolutional modules for context (DSSD) or learning object detectors from scratch without pre-trained ImageNet classifiers (DSOD).
  2. Knowl 2 — Fast R-CNN Multi-Task Loss Formulation

    equation

    In Fast R-CNN, network parameters for classification and bounding-box regression are jointly trained end-to-end using a multi-task loss defined for each candidate Region of Interest (RoI):

    L(p,u,tu,v)=Lcls(p,u)+λ[u1]Lloc(tu,v)L(p, u, t^u, v) = L_{\text{cls}}(p, u) + \lambda [u \ge 1] L_{\text{loc}}(t^u, v)

    where:

    • p=(p0,p1,,pC)p = (p_0, p_1, \dots, p_C) is the discrete probability distribution computed via softmax across C+1C+1 classes (CC object categories plus background class 00).
    • u{0,1,,C}u \in \{0, 1, \dots, C\} is the true class label of the RoI.
    • Lcls(p,u)=logpuL_{\text{cls}}(p, u) = -\log p_u is the classification log loss for the true class uu.
    • v=(vx,vy,vw,vh)v = (v_x, v_y, v_w, v_h) is the ground-truth bounding-box regression target tuple specifying scale-invariant translation and log-space height/width shifts.
    • tu=(txu,tyu,twu,thu)t^u = (t^u_x, t^u_y, t^u_w, t^u_h) is the predicted bounding-box parameter shift tuple for class uu.
    • [u1][u \ge 1] is the Iverson bracket indicator function, evaluating to 11 when u1u \ge 1 (foreground object RoI) and 00 when u=0u = 0 (background RoI, which ignores localization error).
    • λ\lambda is a hyperparameter weighting the relative contribution of the localization task.
    • Lloc(tu,v)L_{\text{loc}}(t^u, v) is the regression loss summing element-wise smoothL1\text{smooth}_{L1} errors over the four bounding-box coordinates:

    Lloc(tu,v)=i{x,y,w,h}smoothL1(tiuvi)L_{\text{loc}}(t^u, v) = \sum_{i \in \{x, y, w, h\}} \text{smooth}_{L1}(t^u_i - v_i)

  3. Knowl 3 — Smooth L1 Loss for Bounding Box Regression

    equation

    To prevent exploding gradients during training and increase robustness against localization outliers compared to standard L2L2 loss, deep learning object detectors utilize the smoothL1\text{smooth}_{L1} loss function:

    smoothL1(x)={0.5x2if x<1x0.5otherwise\text{smooth}_{L1}(x) = \begin{cases} 0.5 x^2 & \text{if } |x| < 1 \\ |x| - 0.5 & \text{otherwise} \end{cases}

    where xRx \in \mathbb{R} denotes the residual difference between a predicted bounding-box coordinate offset and its assigned ground-truth target. The loss is quadratic for small errors (x<1|x| < 1) and linear for larger errors (x1|x| \ge 1).

  4. Knowl 4 — Faster R-CNN Region Proposal Network Multi-Task Loss

    equation

    The Region Proposal Network (RPN) in Faster R-CNN is trained using a multi-task loss function defined over anchor locations in a mini-batch:

    L({pi},{ti})=1NclsiLcls(pi,pi)+λ1NregipiLreg(ti,ti)L(\{p_i\}, \{t_i\}) = \frac{1}{N_{\text{cls}}} \sum_i L_{\text{cls}}(p_i, p^*_i) + \lambda \frac{1}{N_{\text{reg}}} \sum_i p^*_i L_{\text{reg}}(t_i, t^*_i)

    where:

    • ii is the index of an anchor in a mini-batch.
    • pi[0,1]p_i \in [0, 1] is the predicted probability of anchor ii containing an object.
    • pi{0,1}p^*_i \in \{0, 1\} is the ground-truth binary label, set to 11 if the anchor has the highest Intersection-over-Union (IoU) overlap with a ground-truth box or has IoU>0.7\text{IoU} > 0.7 with any ground-truth box; it is set to 00 if IoU<0.3\text{IoU} < 0.3 with all ground-truth boxes. Anchors with intermediate overlap do not contribute to the loss.
    • ti=(ti,x,ti,y,ti,w,ti,h)t_i = (t_{i,x}, t_{i,y}, t_{i,w}, t_{i,h}) is the predicted 4-dimensional bounding-box coordinate parameterization.
    • ti=(ti,x,ti,y,ti,w,ti,h)t^*_i = (t^*_{i,x}, t^*_{i,y}, t^*_{i,w}, t^*_{i,h}) is the ground-truth box coordinate vector for positive anchor ii.
    • Lcls(pi,pi)L_{\text{cls}}(p_i, p^*_i) is binary log loss over the object versus non-object classes.
    • Lreg(ti,ti)=j{x,y,w,h}smoothL1(ti,jti,j)L_{\text{reg}}(t_i, t^*_i) = \sum_{j \in \{x, y, w, h\}} \text{smooth}_{L1}(t_{i,j} - t^*_{i,j}) is the smooth L1L1 regression loss activated only for positive anchors (pi=1p^*_i = 1).
    • NclsN_{\text{cls}} is the mini-batch size normalizing the classification term.
    • NregN_{\text{reg}} is the number of anchor locations normalizing the regression term.
    • λ\lambda is a balancing weighting coefficient.
  5. Knowl 5 — YOLO Single-Stage Detection Loss Function

    equation

    In the YOLO framework, an image is divided into an S×SS \times S grid. Each cell predicts BB bounding boxes, their confidence scores, and CC conditional class probabilities. The network optimizes the following multi-part sum-squared error loss:

    LYOLO=λcoordi=0S2j=0BIijobj[(xix^i)2+(yiy^i)2]L_{\text{YOLO}} = \lambda_{\text{coord}} \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{I}_{ij}^{\text{obj}} \left[ (x_i - \hat{x}_i)^2 + (y_i - \hat{y}_i)^2 \right] +λcoordi=0S2j=0BIijobj[(wiw^i)2+(hih^i)2]+ \lambda_{\text{coord}} \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{I}_{ij}^{\text{obj}} \left[ (\sqrt{w_i} - \sqrt{\hat{w}_i})^2 + (\sqrt{h_i} - \sqrt{\hat{h}_i})^2 \right] +i=0S2j=0BIijobj(CiC^i)2+λnoobji=0S2j=0BIijnoobj(CiC^i)2+ \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{I}_{ij}^{\text{obj}} (C_i - \hat{C}_i)^2 + \lambda_{\text{noobj}} \sum_{i=0}^{S^2} \sum_{j=0}^{B} \mathbb{I}_{ij}^{\text{noobj}} (C_i - \hat{C}_i)^2 +i=0S2Iiobjcclasses(pi(c)p^i(c))2+ \sum_{i=0}^{S^2} \mathbb{I}_i^{\text{obj}} \sum_{c \in \text{classes}} (p_i(c) - \hat{p}_i(c))^2

    where:

    • (xi,yi)(x_i, y_i) and (x^i,y^i)(\hat{x}_i, \hat{y}_i) denote the ground-truth and predicted center coordinates of the bounding box relative to grid cell ii.
    • (wi,hi)(w_i, h_i) and (w^i,h^i)(\hat{w}_i, \hat{h}_i) represent ground-truth and predicted width and height normalized relative to the full image dimensions (square roots are used to weight errors in small boxes more heavily than in large boxes).
    • CiC_i and C^i\hat{C}_i are ground-truth and predicted confidence scores for box predictor jj in cell ii, where ground-truth confidence is 11 if an object is present and 00 otherwise.
    • pi(c)p_i(c) and p^i(c)\hat{p}_i(c) denote the ground-truth and predicted conditional probabilities for class cc in cell ii.
    • Iiobj\mathbb{I}_i^{\text{obj}} is 11 if an object center falls within grid cell ii, and 00 otherwise.
    • Iijobj\mathbb{I}_{ij}^{\text{obj}} is 11 if the jj-th bounding box predictor in cell ii is responsible for predicting the ground-truth object (highest IoU with the object), and 00 otherwise.
    • Iijnoobj\mathbb{I}_{ij}^{\text{noobj}} is 11 if no object is present in that bounding box predictor.
    • λcoord\lambda_{\text{coord}} is a hyperparameter scaling the coordinate localization penalty.
    • λnoobj\lambda_{\text{noobj}} is a hyperparameter scaling down the confidence penalty from background cells containing no objects.
  6. Knowl 6 — Comparative Performance of Generic Object Detectors on PASCAL VOC Benchmarks

    data/table

    Object detection performance across 20 classes on the PASCAL VOC 2007 test set and PASCAL VOC 2012 test set, measured in Mean Average Precision (mAP in %):

    Method Backbone Training Data VOC 2007 mAP (%) VOC 2012 mAP (%)
    R-CNN AlexNet VOC 07 / 12 58.5 53.3
    R-CNN VGG16 VOC 07 / 12 66.0 62.4
    SPP-net ZF-Net VOC 07 60.9 -
    Fast R-CNN VGG16 VOC 07+12 70.0 68.4
    Faster R-CNN VGG16 VOC 07+12 73.2 70.4
    Faster R-CNN VGG16 VOC 07+12+COCO 78.8 75.9
    OHEM + Fast R-CNN VGG16 VOC 07+12(+COCO) 78.9 80.1
    ION VGG16 VOC 07+12+S 79.2 76.4
    MR-CNN S-CNN VGG16 VOC 07+12 78.2 73.9
    HyperNet VGG16 VOC 07+12 76.3 71.4
    YOLO Custom 24-conv VOC 07+12 63.4 57.9
    YOLO + Fast R-CNN - VOC 07+12 - 70.7
    YOLOv2 DarkNet-19 VOC 07+12+COCO - 78.2
    SSD300 VGG16 VOC 07+12+COCO 79.6 79.3
    SSD512 VGG16 VOC 07+12+COCO 81.6 82.2
    R-FCN ResNet101 VOC 07+12+COCO 83.6 85.0

    Key empirical findings:

    • Backbone capacity is critical to detection accuracy: replacing AlexNet with VGG16 in R-CNN raises VOC 2007 mAP from 58.5%58.5\% to 66.0%66.0\%, and ResNet101 in R-FCN achieves 85.0%85.0\% on VOC 2012.
    • Progression from multi-stage cached pipelines to end-to-end joint multi-task models (Fast R-CNN, 70.0%70.0\%) and internal region proposal networks (Faster R-CNN, 73.2%73.2\%) provides steady gains.
    • Augmenting training datasets with Microsoft COCO prior to VOC fine-tuning consistently yields substantial accuracy boosts across models.
  7. Knowl 7 — Object Detection Performance on Microsoft COCO Test-Dev Benchmark

    data/table

    Object detection results on the Microsoft COCO test-dev benchmark across 80 categories, evaluated by standard Average Precision across IoU thresholds 0.50:0.950.50:0.95 (AP\text{AP}), AP50\text{AP}^{50}, AP75\text{AP}^{75}, and scale-stratified subsets for small (APS\text{AP}^S), medium (APM\text{AP}^M), and large (APL\text{AP}^L) objects (all values in %):

    Method Training Set AP (0.5:0.95) AP50\text{AP}^{50} APS\text{AP}^S APM\text{AP}^M APL\text{AP}^L
    Fast R-CNN train 20.5 39.9 4.1 20.0 35.8
    ION train 23.6 43.2 6.4 24.1 38.3
    Faster R-CNN (VGG16) trainval 24.2 45.3 7.7 26.4 37.1
    OHEM + Fast R-CNN trainval 25.5 45.9 7.4 27.7 38.5
    YOLOv2 trainval35k 21.6 44.0 5.0 22.4 35.5
    SSD300 trainval35k 23.2 41.2 5.3 23.2 39.6
    SSD512 trainval35k 26.8 46.5 9.0 28.9 41.9
    R-FCN (ResNet101) trainval 29.2 51.5 10.8 32.8 45.0
    Multi-path trainval 33.2 51.9 13.6 37.2 47.8
    DSSD513 (ResNet101) trainval35k 33.2 53.3 13.0 35.4 51.1
    DSOD300 trainval 29.3 47.3 9.4 31.5 47.0
    FPN (ResNet101) trainval35k 36.2 59.1 18.2 39.0 48.2
    Mask R-CNN (ResNet101 + FPN) trainval35k 38.2 60.3 20.1 41.1 50.2
    Mask R-CNN (ResNeXt101 + FPN) trainval35k 39.8 62.3 22.1 43.2 51.2

    Key empirical findings:

    • Multi-scale feature construction with top-down pyramids and lateral connections (FPN) significantly improves small-object localization, increasing APS\text{AP}^S from 7.7%7.7\% (Faster R-CNN) to 18.2%18.2\%.
    • Combining instance segmentation multi-task supervision with high-capacity ResNeXt backbones and FPN (Mask R-CNN) achieves the highest detection performance (39.8%39.8\% AP).
    • One-stage detectors (YOLOv2, SSD300) produce lower detection and localization accuracy on small non-standard objects than two-stage pyramid architectures.
  8. Knowl 8 — Inference Latency and Throughput Benchmarks on PASCAL VOC 2007

    data/table

    Inference time (in seconds per image) and frame rate (in Frames Per Second, FPS) evaluated on an NVIDIA Titan X GPU (with CPU for Selective Search region proposal extraction where indicated) on the PASCAL VOC 2007 test set:

    Method Training Set mAP (%) Test Time (sec/img) Rate (FPS)
    Selective Search + R-CNN VOC 07 66.0 32.84 0.03
    Selective Search + SPP-net VOC 07 63.1 2.30 0.44
    Selective Search + Fast R-CNN VOC 07+12 66.9 1.72 0.60
    SDP + CRC VOC 07 68.9 0.47 2.10
    HyperNet (speed-up version) VOC 07+12 76.3 0.20 5.00
    MR-CNN S-CNN VOC 07+12 78.2 30.00 0.03
    ION VOC 07+12+S 79.2 1.92 0.50
    Faster R-CNN (VGG16) VOC 07+12 73.2 0.11 9.10
    Faster R-CNN (ResNet101) VOC 07+12 83.8 2.24 0.40
    R-FCN (ResNet101) VOC 07+12+COCO 83.6 0.17 5.90
    YOLO VOC 07+12 63.4 0.02 45.00
    SSD300 VOC 07+12 74.3 0.02 46.00
    SSD512 VOC 07+12 76.8 0.05 19.00
    YOLOv2 (544×544544 \times 544) VOC 07+12 78.6 0.03 40.00
    DSSD321 (ResNet101) VOC 07+12 78.6 0.07 13.60
    DSOD300 VOC 07+12+COCO 81.7 0.06 17.40
    PVANET+ VOC 07+12+COCO 83.8 0.05 21.70
    PVANET+ (compressed) VOC 07+12+COCO 82.9 0.03 31.30

    Key efficiency conclusions:

    • Eliminating external region proposal computation via in-network RPN accelerates processing from 0.60 FPS0.60\text{ FPS} (Fast R-CNN) to 9.10 FPS9.10\text{ FPS} (Faster R-CNN).
    • One-stage detectors (YOLO, SSD300, YOLOv2) achieve real-time rates of 4046 FPS40\text{--}46\text{ FPS} by framing detection as direct grid-based regression.
    • Model compression techniques (such as SVD layer compression in PVANET+) allow two-stage proposal networks to reach real-time rates (21.731.3 FPS21.7\text{--}31.3\text{ FPS}) while retaining high accuracy (82.983.8%82.9\text{--}83.8\% mAP).
  9. Knowl 9 — Salient Object Detection Evaluation Metrics: Weighted F-Measure and MAE

    equation

    Salient object detection models are evaluated using the weighted FF-measure (FβF_\beta) and the Mean Absolute Error (MAE):

    1. Weighted FF-Measure (FβF_\beta): Fβ=(1+β2)PrecisionRecallβ2Precision+RecallF_\beta = \frac{(1 + \beta^2) \cdot \text{Precision} \cdot \text{Recall}}{\beta^2 \cdot \text{Precision} + \text{Recall}} where β2\beta^2 is set to 0.30.3 to give more importance to precision over recall when assessing overlap between a binarized predicted saliency mask and the binary ground truth.

    2. Mean Absolute Error (MAE): MAE=1W×Hi=1Hj=1WS^(i,j)Z^(i,j)\text{MAE} = \frac{1}{W \times H} \sum_{i=1}^H \sum_{j=1}^W |\hat{S}(i, j) - \hat{Z}(i, j)| where S^(i,j)[0,1]\hat{S}(i, j) \in [0, 1] represents the continuous saliency value predicted at pixel (i,j)(i, j), Z^(i,j){0,1}\hat{Z}(i, j) \in \{0, 1\} is the ground-truth binary label at (i,j)(i, j), and WW and HH are image width and height. MAE directly measures average pixel-level numerical deviation.

  10. Knowl 10 — Salient Object Detection Benchmark Comparison

    data/table

    Quantitative performance of salient object detection models across four benchmark datasets (PASCAL-S, ECSSD, HKU-IS, SOD), evaluated by weighted FF-measure (wFβwF_\beta, higher is better) and Mean Absolute Error (MAE, lower is better):

    Method PASCAL-S ECSSD HKU-IS SOD
    wFβwF_\beta MAE wFβwF_\beta MAE wFβwF_\beta MAE wFβwF_\beta MAE
    CHM (Handcrafted) 0.631 0.222 0.722 0.195 0.728 0.158 0.655 0.249
    RC (Handcrafted) 0.640 0.225 0.741 0.187 0.726 0.165 0.657 0.242
    DRFI (Handcrafted) 0.679 0.221 0.787 0.166 0.783 0.143 0.712 0.215
    MC (CNN) 0.721 0.147 0.822 0.107 0.781 0.098 0.708 0.184
    MDF (CNN) 0.764 0.145 0.833 0.108 0.860 0.129 0.785 0.155
    LEGS (CNN) 0.756 0.157 0.827 0.118 0.770 0.118 0.707 0.205
    DSR (CNN) 0.697 0.128 0.872 0.037 0.833 0.040 - -
    MTDNN (CNN) 0.818 0.170 0.810 0.160 - - 0.781 0.150
    CRPSD (CNN) 0.776 0.063 0.849 0.046 0.821 0.043 - -
    DCL (CNN) 0.822 0.108 0.898 0.071 0.907 0.048 0.832 0.126
    ELD (CNN) 0.767 0.121 0.865 0.098 0.844 0.071 0.760 0.154
    NLDF (CNN) 0.831 0.099 0.905 0.063 0.902 0.048 0.810 0.143
    DSSC (CNN) 0.830 0.080 0.915 0.052 0.913 0.039 0.842 0.118

    Key empirical findings:

    • Deep CNN frameworks markedly outperform classical handcrafted contrast baselines across all datasets (e.g., DSSC achieves wFβ=0.915wF_\beta = 0.915 on ECSSD compared to 0.7870.787 for DRFI).
    • Architectures incorporating multi-scale contextual features, short skip-connections, and superpixel-guided segmentation (DCL, NLDF, DSSC) yield the lowest MAE and highest wFβwF_\beta.
  11. Knowl 11 — Pedestrian Detection Performance Breakdown on Caltech Dataset

    data/table

    Pedestrian detection evaluated on the Caltech Pedestrian dataset using Log-Average Miss Rate (L-AMR in %, where lower is better), averaged over false positives per image (FPPI) in [102,100][10^{-2}, 10^0] across standard occlusion and scale subsets:

    Method Reasonable All Far Medium Near None Partial Heavy
    Checkerboards+ (Handcrafted) 17.1 68.4 100.0 58.3 5.1 15.6 31.4 78.4
    LDCF++ (Handcrafted) 15.2 67.1 100.0 58.4 5.4 13.3 33.3 76.2
    SCF+AlexNet 23.3 70.3 100.0 62.3 10.2 20.0 48.5 74.7
    SA-FastRCNN 9.7 62.6 100.0 51.8 0.0 7.7 24.8 64.3
    MS-CNN 10.0 61.0 97.2 49.1 2.6 8.2 19.2 60.0
    DeepParts 11.9 64.8 100.0 56.4 4.8 10.6 19.9 60.4
    CompACT-Deep 11.8 64.4 100.0 53.2 4.0 9.6 25.1 65.8
    RPN+BF 9.6 64.7 100.0 53.9 2.3 7.7 24.2 74.2
    F-DNN+SS 8.2 50.3 77.5 33.2 2.8 6.7 15.1 53.4

    Key empirical findings:

    • Standard shallow CNN adaptations without explicit scale handling (SCF+AlexNet, 23.3%23.3\% L-AMR) performed worse than competitive handcrafted feature ensembles (LDCF++, 15.2%15.2\%).
    • Part-based representations (DeepParts, 19.9%19.9\% on partial occlusion) and scale-dependent layer pooling (MS-CNN, 19.2%19.2\%) provide significant robustness against severe body occlusion.
    • Multi-classifier fusion with soft rejection (F-DNN+SS) achieves top performance across all subsets (8.2%8.2\% on Reasonable, 50.3%50.3\% on All).
  12. Knowl 12 — Domain-Specific Adaptations across Object Detection Sub-Tasks

    model/method

    While generic object detection tackles diverse categories with high geometric variance, specific sub-tasks require dedicated network modifications:

    1. Face Detection:

      • Characteristics: High structural regularity but extreme scale ranges (101,00010\text{--}1,000 pixels), severe pose variations, and unconstrained occlusions.
      • Adaptations: Multi-task cascaded architectures (e.g., MTCNN, DenseBox) jointly optimize face bounding-box detection, facial landmark localization, and 3D face model fitting; scale-partitioned sub-networks (ScaleFace) and scale distribution histogram estimation adapt to wide scale variations.
    2. Pedestrian Detection:

      • Characteristics: Predominance of small object instances where standard RoI pooling can cause feature collapse, and error sources dominated by hard background confusion rather than inter-class competition.
      • Adaptations: Part-based CNN ensembles (e.g., DeepParts) handle partial occlusion; downstream boosted decision forests operate directly on high-resolution shared feature maps (RPN+BF); multi-spectral networks fuse RGB and thermal imaging channels.
    3. Salient Object Detection:

      • Characteristics: Identifying the most conspicuous foreground regions and producing continuous pixel-level probability masks rather than rectangular bounding boxes.
      • Adaptations: Multi-scale deep contrast feature extraction combines top-down task semantics with bottom-up local contrast; fully convolutional deconvolution streams are coupled with superpixel segmentation to produce crisp object boundaries.

Coverage note — Omitted individual descriptions of dozens of cited third-party classification backbones and minor engineering variants, as well as peripheral survey text on historical deep learning background, in favor of focusing on the paper's core contributed taxonomy, loss equations, comprehensive benchmark tables, and task-specific comparative analyses.

References

  1. 1.P. F. Felzenszwalb, R. B. Girshick, D. Mcallester, and D. Ramanan, ‘‘Object detection with discriminatively trained part-based models,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 9, p. 1627, 2010.
  2. 2.K. K. Sung and T. Poggio, ‘‘Example-based learning for view-based human face detection,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 20, no. 1, pp. 39–51, 2002.
  3. 3.C. Wojek, P. Dollar, B. Schiele, and P. Perona, ‘‘Pedestrian detection: An evaluation of the state of the art,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 34, no. 4, p. 743, 2012.
  4. 4.H. Kobatake and Y. Yoshinaga, ‘‘Detection of spicules on mammogram based on skeleton analysis.’’ IEEE Trans. Med. Imag., vol. 15, no. 3, pp. 235–245, 1996.
  5. 5.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, ‘‘Caffe: Convolutional architecture for fast feature embedding,’’ in ACM MM, 2014.
  6. 6.A. Krizhevsky, I. Sutskever, and G. E. Hinton, ‘‘Imagenet classification with deep convolutional neural networks,’’ in NIPS, 2012.
  7. 7.Z. Cao, T. Simon, S.-E. Wei, and Y. Sheikh, ‘‘Realtime multi-person 2d pose estimation using part affinity fields,’’ in CVPR, 2017.
  8. 8.Z. Yang and R. Nevatia, ‘‘A multi-scale cascade fully convolutional network face detector,’’ in ICPR, 2016.
  9. 9.C. Chen, A. Seff, A. L. Kornhauser, and J. Xiao, ‘‘Deepdriving: Learning affordance for direct perception in autonomous driving,’’ in ICCV, 2015.
  10. 10.X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, ‘‘Multi-view 3d object detection network for autonomous driving,’’ in CVPR, 2017.
  11. 11.A. Dundar, J. Jin, B. Martini, and E. Culurciello, ‘‘Embedded streaming deep neural networks accelerator with applications,’’ IEEE Trans. Neural Netw. & Learning Syst., vol. 28, no. 7, pp. 1572–1583, 2017.
  12. 12.R. J. Cintra, S. Duffner, C. Garcia, and A. Leite, ‘‘Low-complexity approximate convolutional neural networks,’’ IEEE Trans. Neural Netw. & Learning Syst., vol. PP, no. 99, pp. 1–12, 2018.
  13. 13.S. H. Khan, M. Hayat, M. Bennamoun, F. A. Sohel, and R. Togneri, ‘‘Cost-sensitive learning of deep feature representations from imbalanced data.’’ IEEE Trans. Neural Netw. & Learning Syst., vol. PP, no. 99, pp. 1–15, 2017.
  14. 14.A. Stuhlsatz, J. Lippel, and T. Zielke, ‘‘Feature extraction with deep neural networks by a generalized discriminant analysis.’’ IEEE Trans. Neural Netw. & Learning Syst., vol. 23, no. 4, pp. 596–608, 2012.
  15. 15.R. Girshick, J. Donahue, T. Darrell, and J. Malik, ‘‘Rich feature hierarchies for accurate object detection and semantic segmentation,’’ in CVPR, 2014.
  16. 16.R. Girshick, ‘‘Fast r-cnn,’’ in ICCV, 2015.
  17. 17.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, ‘‘You only look once: Unified, real-time object detection,’’ in CVPR, 2016.
  18. 18.S. Ren, K. He, R. Girshick, and J. Sun, ‘‘Faster r-cnn: Towards realtime object detection with region proposal networks,’’ in NIPS, 2015, pp. 91–99.
  19. 19.D. G. Lowe, ‘‘Distinctive image features from scale-invariant keypoints,’’ Int. J. of Comput. Vision, vol. 60, no. 2, pp. 91–110, 2004.
  20. 20.N. Dalal and B. Triggs, ‘‘Histograms of oriented gradients for human detection,’’ in CVPR, 2005.
  21. 21.R. Lienhart and J. Maydt, ‘‘An extended set of haar-like features for rapid object detection,’’ in ICIP, 2002.
  22. 22.C. Cortes and V. Vapnik, ‘‘Support vector machine,’’ Machine Learning, vol. 20, no. 3, pp. 273–297, 1995.
  23. 23.Y. Freund and R. E. Schapire, ‘‘A desicion-theoretic generalization of on-line learning and an application to boosting,’’ J. of Comput. & Sys. Sci., vol. 13, no. 5, pp. 663–671, 1997.
  24. 24.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, ‘‘Object detection with discriminatively trained part-based models,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, pp. 1627–1645, 2010.
  25. 25.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, ‘‘The pascal visual object classes challenge 2007 (voc 2007) results (2007),’’ 2008.
  26. 26.Y. LeCun, Y. Bengio, and G. Hinton, ‘‘Deep learning,’’ Nature, vol. 521, no. 7553, pp. 436–444, 2015.
  27. 27.N. Liu, J. Han, D. Zhang, S. Wen, and T. Liu, ‘‘Predicting eye fixations using convolutional neural networks,’’ in CVPR, 2015.
  28. 28.E. Vig, M. Dorr, and D. Cox, ‘‘Large-scale optimization of hierarchical features for saliency prediction in natural images,’’ in CVPR, 2014.
  29. 29.H. Jiang and E. Learned-Miller, ‘‘Face detection with the faster r-cnn,’’ in FG, 2017.
  30. 30.D. Chen, S. Ren, Y. Wei, X. Cao, and J. Sun, ‘‘Joint cascade face detection and alignment,’’ in ECCV, 2014.
  31. 31.D. Chen, G. Hua, F. Wen, and J. Sun, ‘‘Supervised transformer network for efficient face detection,’’ in ECCV, 2016.
  32. 32.D. Ribeiro, A. Mateus, J. C. Nascimento, and P. Miraldo, ‘‘A real-time pedestrian detector using deep learning for human-aware navigation,’’ arXiv:1607.04441, 2016.
  33. 33.F. Yang, W. Choi, and Y. Lin, ‘‘Exploit all the layers: Fast and accurate cnn object detector with scale dependent pooling and cascaded rejection classifiers,’’ in CVPR, 2016.
  34. 34.P. Druzhkov and V. Kustikova, ‘‘A survey of deep learning methods and software tools for image classification and object detection,’’ Pattern Recognition and Image Anal., vol. 26, no. 1, p. 9, 2016.
  35. 35.W. Pitts and W. S. McCulloch, ‘‘How we know universals the perception of auditory and visual forms,’’ The Bulletin of Mathematical Biophysics, vol. 9, no. 3, pp. 127–147, 1947.
  36. 36.D. E. Rumelhart, G. E. Hinton, and R. J. Williams, ‘‘Learning internal representation by back-propagation of errors,’’ Nature, vol. 323, no. 323, pp. 533–536, 1986.
  37. 37.G. E. Hinton and R. R. Salakhutdinov, ‘‘Reducing the dimensionality of data with neural networks,’’ Sci., vol. 313, pp. 504–507, 2006.
  38. 38.G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al., ‘‘Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,’’ IEEE Signal Process. Mag., vol. 29, no. 6, pp. 82–97, 2012.
  39. 39.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, ‘‘Imagenet: A large-scale hierarchical image database,’’ in CVPR, 2009.
  40. 40.L. Deng, M. L. Seltzer, D. Yu, A. Acero, A.-r. Mohamed, and G. Hinton, ‘‘Binary coding of speech spectrograms using a deep autoencoder,’’ in INTERSPEECH, 2010.
  41. 41.G. Dahl, A.-r. Mohamed, G. E. Hinton et al., ‘‘Phone recognition with the mean-covariance restricted boltzmann machine,’’ in NIPS, 2010.
  42. 42.G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, ‘‘Improving neural networks by preventing coadaptation of feature detectors,’’ arXiv:1207.0580, 2012.
  43. 43.S. Ioffe and C. Szegedy, ‘‘Batch normalization: Accelerating deep network training by reducing internal covariate shift,’’ in ICML, 2015.
  44. 44.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, ‘‘Overfeat: Integrated recognition, localization and detection using convolutional networks,’’ arXiv:1312.6229, 2013.
  45. 45.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, ‘‘Going deeper with convolutions,’’ in CVPR, 2015.
  46. 46.K. Simonyan and A. Zisserman, ‘‘Very deep convolutional networks for large-scale image recognition,’’ arXiv:1409.1556, 2014.
  47. 47.K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Deep residual learning for image recognition,’’ in CVPR, 2016.
  48. 48.V. Nair and G. E. Hinton, ‘‘Rectified linear units improve restricted boltzmann machines,’’ in ICML, 2010.
  49. 49.M. Oquab, L. Bottou, I. Laptev, J. Sivic et al., ‘‘Weakly supervised object recognition with convolutional neural networks,’’ in NIPS, 2014.
  50. 50.M. Oquab, L. Bottou, I. Laptev, and J. Sivic, ‘‘Learning and transferring mid-level image representations using convolutional neural networks,’’ in CVPR, 2014.
  51. 51.F. M. Wadley, ‘‘Probit analysis: a statistical treatment of the sigmoid response curve,’’ Annals of the Entomological Soc. of America, vol. 67, no. 4, pp. 549–553, 1947.
  52. 52.K. Kavukcuoglu, R. Fergus, Y. LeCun et al., ‘‘Learning invariant features through topographic filter maps,’’ in CVPR, 2009.
  53. 53.K. Kavukcuoglu, P. Sermanet, Y.-L. Boureau, K. Gregor, M. Mathieu, and Y. LeCun, ‘‘Learning convolutional feature hierarchies for visual recognition,’’ in NIPS, 2010.
  54. 54.M. D. Zeiler, D. Krishnan, G. W. Taylor, and R. Fergus, ‘‘Deconvolutional networks,’’ in CVPR, 2010.
  55. 55.H. Noh, S. Hong, and B. Han, ‘‘Learning deconvolution network for semantic segmentation,’’ in ICCV, 2015.
  56. 56.Z.-Q. Zhao, B.-J. Xie, Y.-m. Cheung, and X. Wu, ‘‘Plant leaf identification via a growing convolution neural network with progressive sample learning,’’ in ACCV, 2014.
  57. 57.A. Babenko, A. Slesarev, A. Chigorin, and V. Lempitsky, ‘‘Neural codes for image retrieval,’’ in ECCV, 2014.
  58. 58.J. Wan, D. Wang, S. C. H. Hoi, P. Wu, J. Zhu, Y. Zhang, and J. Li, ‘‘Deep learning for content-based image retrieval: A comprehensive study,’’ in ACM MM, 2014.
  59. 59.D. Tome, F. Monti, L. Baroffio, L. Bondi, M. Tagliasacchi, and S. Tubaro, ‘‘Deep convolutional neural networks for pedestrian detection,’’ Signal Process.: Image Commun., vol. 47, pp. 482–489, 2016.
  60. 60.Y. Xiang, W. Choi, Y. Lin, and S. Savarese, ‘‘Subcategory-aware convolutional neural networks for object proposals and detection,’’ in WACV, 2017.
  61. 61.Z.-Q. Zhao, H. Bian, D. Hu, W. Cheng, and H. Glotin, ‘‘Pedestrian detection based on fast r-cnn and batch normalization,’’ in ICIC, 2017.
  62. 62.J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, ‘‘Multimodal deep learning,’’ in ICML, 2011.
  63. 63.Z. Wu, X. Wang, Y.-G. Jiang, H. Ye, and X. Xue, ‘‘Modeling spatialtemporal clues in a hybrid deep learning framework for video classification,’’ in ACM MM, 2015.
  64. 64.K. He, X. Zhang, S. Ren, and J. Sun, ‘‘Spatial pyramid pooling in deep convolutional networks for visual recognition,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 9, pp. 1904–1916, 2015.
  65. 65.Y. Li, K. He, J. Sun et al., ‘‘R-fcn: Object detection via region-based fully convolutional networks,’’ in NIPS, 2016, pp. 379–387.
  66. 66.T.-Y. Lin, P. Dollar, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, ‘‘Feature pyramid networks for object detection,’’ in CVPR, 2017.
  67. 67.K. He, G. Gkioxari, P. Dollar, and R. B. Girshick, ‘‘Mask r-cnn,’’ in ICCV, 2017.
  68. 68.D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov, ‘‘Scalable object detection using deep neural networks,’’ in CVPR, 2014.
  69. 69.D. Yoo, S. Park, J.-Y. Lee, A. S. Paek, and I. So Kweon, ‘‘Attentionnet: Aggregating weak directions for accurate object detection,’’ in CVPR, 2015.
  70. 70.M. Najibi, M. Rastegari, and L. S. Davis, ‘‘G-cnn: an iterative grid based object detector,’’ in CVPR, 2016.
  71. 71.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, ‘‘Ssd: Single shot multibox detector,’’ in ECCV, 2016.
  72. 72.J. Redmon and A. Farhadi, ‘‘Yolo9000: better, faster, stronger,’’ arXiv:1612.08242, 2016.
  73. 73.C. Y. Fu, W. Liu, A. Ranga, A. Tyagi, and A. C. Berg, ‘‘Dssd: Deconvolutional single shot detector,’’ arXiv:1701.06659, 2017.
  74. 74.Z. Shen, Z. Liu, J. Li, Y. G. Jiang, Y. Chen, and X. Xue, ‘‘Dsod: Learning deeply supervised object detectors from scratch,’’ in ICCV, 2017.
  75. 75.G. E. Hinton, A. Krizhevsky, and S. D. Wang, ‘‘Transforming autoencoders,’’ in ICANN, 2011.
  76. 76.G. W. Taylor, I. Spiro, C. Bregler, and R. Fergus, ‘‘Learning invariance through imitation,’’ in CVPR, 2011.
  77. 77.X. Ren and D. Ramanan, ‘‘Histograms of sparse codes for object detection,’’ in CVPR, 2013.
  78. 78.J. R. Uijlings, K. E. Van De Sande, T. Gevers, and A. W. Smeulders, ‘‘Selective search for object recognition,’’ Int. J. of Comput. Vision, vol. 104, no. 2, pp. 154–171, 2013.
  79. 79.P. Sermanet, K. Kavukcuoglu, S. Chintala, and Y. LeCun, ‘‘Pedestrian detection with unsupervised multi-stage feature learning,’’ in CVPR, 2013.
  80. 80.P. Krahenbuhl and V. Koltun, ‘‘Geodesic object proposals,’’ in ECCV, 2014.
  81. 81.P. Arbelaez, J. Pont-Tuset, J. T. Barron, F. Marques, and J. Malik, ‘‘Multiscale combinatorial grouping,’’ in CVPR, 2014.
  82. 82.C. L. Zitnick and P. Dollar, ‘‘Edge boxes: Locating object proposals from edges,’’ in ECCV, 2014.
  83. 83.W. Kuo, B. Hariharan, and J. Malik, ‘‘Deepbox: Learning objectness with convolutional networks,’’ in ICCV, 2015.
  84. 84.P. O. Pinheiro, T.-Y. Lin, R. Collobert, and P. Dollar, ‘‘Learning to refine object segments,’’ in ECCV, 2016.
  85. 85.Y. Zhang, K. Sohn, R. Villegas, G. Pan, and H. Lee, ‘‘Improving object detection with deep convolutional networks via bayesian optimization and structured prediction,’’ in CVPR, 2015.
  86. 86.S. Gupta, R. Girshick, P. Arbelaez, and J. Malik, ‘‘Learning rich features from rgb-d images for object detection and segmentation,’’ in ECCV, 2014.
  87. 87.W. Ouyang, X. Wang, X. Zeng, S. Qiu, P. Luo, Y. Tian, H. Li, S. Yang, Z. Wang, C.-C. Loy et al., ‘‘Deepid-net: Deformable deep convolutional neural networks for object detection,’’ in CVPR, 2015.
  88. 88.K. Lenc and A. Vedaldi, ‘‘R-cnn minus r,’’ arXiv:1506.06981, 2015.
  89. 89.S. Lazebnik, C. Schmid, and J. Ponce, ‘‘Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories,’’ in CVPR, 2006.
  90. 90.F. Perronnin, J. Sanchez, and T. Mensink, ‘‘Improving the fisher kernel for large-scale image classification,’’ in ECCV, 2010.
  91. 91.J. Xue, J. Li, and Y. Gong, ‘‘Restructuring of deep neural network acoustic models with singular value decomposition.’’ in Interspeech, 2013.
  92. 92.S. Ren, K. He, R. Girshick, and J. Sun, ‘‘Faster r-cnn: Towards real-time object detection with region proposal networks,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 6, pp. 1137–1149, 2017.
  93. 93.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, ‘‘Rethinking the inception architecture for computer vision,’’ in CVPR, 2016.
  94. 94.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick, ‘‘Microsoft coco: Common objects in context,’’ in ECCV, 2014.
  95. 95.S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick, ‘‘Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks,’’ in CVPR, 2016.
  96. 96.A. Arnab and P. H. S. Torr, ‘‘Pixelwise instance segmentation with a dynamically instantiated network,’’ in CVPR, 2017.
  97. 97.J. Dai, K. He, and J. Sun, ‘‘Instance-aware semantic segmentation via multi-task network cascades,’’ in CVPR, 2016.
  98. 98.Y. Li, H. Qi, J. Dai, X. Ji, and Y. Wei, ‘‘Fully convolutional instanceaware semantic segmentation,’’ in CVPR, 2017.
  99. 99.M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, ‘‘Spatial transformer networks,’’ in CVPR, 2015.
  100. 100.S. Brahmbhatt, H. I. Christensen, and J. Hays, ‘‘Stuffnet: Using stuffto improve object detection,’’ in WACV, 2017.
  101. 101.T. Kong, A. Yao, Y. Chen, and F. Sun, ‘‘Hypernet: Towards accurate region proposal generation and joint object detection,’’ in CVPR, 2016.
  102. 102.A. Pentina, V. Sharmanska, and C. H. Lampert, ‘‘Curriculum learning of multiple tasks,’’ in CVPR, 2015.
  103. 103.J. Yim, H. Jung, B. Yoo, C. Choi, D. Park, and J. Kim, ‘‘Rotating your face using multi-task deep neural network,’’ in CVPR, 2015.
  104. 104.J. Li, X. Liang, J. Li, T. Xu, J. Feng, and S. Yan, ‘‘Multi-stage object detection with group recursive learning,’’ arXiv:1608.05159, 2016.
  105. 105.Z. Cai, Q. Fan, R. S. Feris, and N. Vasconcelos, ‘‘A unified multi-scale deep convolutional neural network for fast object detection,’’ in ECCV, 2016.
  106. 106.Y. Zhu, R. Urtasun, R. Salakhutdinov, and S. Fidler, ‘‘segdeepm: Exploiting segmentation and context in deep neural networks for object detection,’’ in CVPR, 2015.
  107. 107.W. Byeon, T. M. Breuel, F. Raue, and M. Liwicki, ‘‘Scene labeling with lstm recurrent neural networks,’’ in CVPR, 2015.
  108. 108.B. Moysset, C. Kermorvant, and C. Wolf, ‘‘Learning to detect and localize many objects from few examples,’’ arXiv:1611.05664, 2016.
  109. 109.X. Zeng, W. Ouyang, B. Yang, J. Yan, and X. Wang, ‘‘Gated bidirectional cnn for object detection,’’ in ECCV, 2016.
  110. 110.S. Gidaris and N. Komodakis, ‘‘Object detection via a multi-region and semantic segmentation-aware cnn model,’’ in CVPR, 2015.
  111. 111.M. Schuster and K. K. Paliwal, ‘‘Bidirectional recurrent neural networks,’’ IEEE Trans. Signal Process., vol. 45, pp. 2673–2681, 1997.
  112. 112.S. Zagoruyko, A. Lerer, T.-Y. Lin, P. O. Pinheiro, S. Gross, S. Chintala, and P. Dollar, ‘‘A multipath network for object detection,’’ arXiv:1604.02135, 2016.
  113. 113.A. Shrivastava, A. Gupta, and R. Girshick, ‘‘Training region-based object detectors with online hard example mining,’’ in CVPR, 2016.
  114. 114.S. Ren, K. He, R. Girshick, X. Zhang, and J. Sun, ‘‘Object detection networks on convolutional feature maps,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 7, pp. 1476–1481, 2017.
  115. 115.W. Ouyang, X. Wang, C. Zhang, and X. Yang, ‘‘Factors in finetuning deep model for object detection with long-tail distribution,’’ in CVPR, 2016.
  116. 116.S. Hong, B. Roh, K.-H. Kim, Y. Cheon, and M. Park, ‘‘Pvanet: Lightweight deep neural networks for real-time object detection,’’ arXiv:1611.08588, 2016.
  117. 117.W. Shang, K. Sohn, D. Almeida, and H. Lee, ‘‘Understanding and improving convolutional neural networks via concatenated rectified linear units,’’ in ICML, 2016.
  118. 118.C. Szegedy, A. Toshev, and D. Erhan, ‘‘Deep neural networks for object detection,’’ in NIPS, 2013.
  119. 119.P. O. Pinheiro, R. Collobert, and P. Dollar, ‘‘Learning to segment object candidates,’’ in NIPS, 2015.
  120. 120.C. Szegedy, S. Reed, D. Erhan, D. Anguelov, and S. Ioffe, ‘‘Scalable, high-quality object detection,’’ arXiv:1412.1441, 2014.
  121. 121.M. Everingham, L. Van Gool, C. Williams, J. Winn, and A. Zisserman, ‘‘The pascal visual object classes challenge 2012 (voc2012) results (2012),’’ in http://www.pascal-network.org/challenges/VOC/voc2011/workshop/index.html, 2011.
  122. 122.M. D. Zeiler and R. Fergus, ‘‘Visualizing and understanding convolutional networks,’’ in ECCV, 2014.
  123. 123.S. Xie, R. B. Girshick, P. Dollar, Z. Tu, and K. He, ‘‘Aggregated residual transformations for deep neural networks,’’ in CVPR, 2017.
  124. 124.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, ‘‘Deformable convolutional networks,’’ arXiv:1703.06211, 2017.
  125. 125.C. Rother, L. Bordeaux, Y. Hamadi, and A. Blake, ‘‘Autocollage,’’ ACM Trans. on Graphics, vol. 25, no. 3, pp. 847–852, 2006.
  126. 126.C. Jung and C. Kim, ‘‘A unified spectral-domain approach for saliency detection and its application to automatic object segmentation,’’ IEEE Trans. Image Process., vol. 21, no. 3, pp. 1272–1283, 2012.
  127. 127.W.-C. Tu, S. He, Q. Yang, and S.-Y. Chien, ‘‘Real-time salient object detection with a minimum spanning tree,’’ in CVPR, 2016.
  128. 128.J. Yang and M.-H. Yang, ‘‘Top-down visual saliency via joint crf and dictionary learning,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 3, pp. 576–588, 2017.
  129. 129.P. L. Rosin, ‘‘A simple method for detecting salient regions,’’ Pattern Recognition, vol. 42, no. 11, pp. 2363–2371, 2009.
  130. 130.T. Liu, Z. Yuan, J. Sun, J. Wang, N. Zheng, X. Tang, and H.-Y. Shum, ‘‘Learning to detect a salient object,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 33, no. 2, pp. 353–367, 2011.
  131. 131.J. Long, E. Shelhamer, and T. Darrell, ‘‘Fully convolutional networks for semantic segmentation,’’ in CVPR, 2015.
  132. 132.D. Gao, S. Han, and N. Vasconcelos, ‘‘Discriminant saliency, the detection of suspicious coincidences, and applications to visual recognition,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 31, pp. 989–1005, 2009.
  133. 133.S. Xie and Z. Tu, ‘‘Holistically-nested edge detection,’’ in ICCV, 2015.
  134. 134.M. Kummerer, L. Theis, and M. Bethge, ‘‘Deep gaze i: Boosting saliency prediction with feature maps trained on imagenet,’’ arXiv:1411.1045, 2014.
  135. 135.X. Huang, C. Shen, X. Boix, and Q. Zhao, ‘‘Salicon: Reducing the semantic gap in saliency prediction by adapting deep neural networks,’’ in ICCV, 2015.
  136. 136.L. Wang, H. Lu, X. Ruan, and M.-H. Yang, ‘‘Deep networks for saliency detection via local estimation and global search,’’ in CVPR, 2015.
  137. 137.H. Cholakkal, J. Johnson, and D. Rajan, ‘‘Weakly supervised top-down salient object detection,’’ arXiv:1611.05345, 2016.
  138. 138.R. Zhao, W. Ouyang, H. Li, and X. Wang, ‘‘Saliency detection by multi-context deep learning,’’ in CVPR, 2015.
  139. 139.C. Bak, A. Erdem, and E. Erdem, ‘‘Two-stream convolutional networks for dynamic saliency prediction,’’ arXiv:1607.04730, 2016.
  140. 140.S. He, R. W. Lau, W. Liu, Z. Huang, and Q. Yang, ‘‘Supercnn: A superpixelwise convolutional neural network for salient object detection,’’ Int. J. of Comput. Vision, vol. 115, no. 3, pp. 330–344, 2015.
  141. 141.X. Li, L. Zhao, L. Wei, M.-H. Yang, F. Wu, Y. Zhuang, H. Ling, and J. Wang, ‘‘Deepsaliency: Multi-task deep neural network model for salient object detection,’’ IEEE Trans. Image Process., vol. 25, no. 8, pp. 3919–3930, 2016.
  142. 142.Y. Tang and X. Wu, ‘‘Saliency detection via combining region-level and pixel-level predictions with cnns,’’ in ECCV, 2016.
  143. 143.G. Li and Y. Yu, ‘‘Deep contrast learning for salient object detection,’’ in CVPR, 2016.
  144. 144.X. Wang, H. Ma, S. You, and X. Chen, ‘‘Edge preserving and multi-scale contextual neural network for salient object detection,’’ arXiv:1608.08029, 2016.
  145. 145.M. Cornia, L. Baraldi, G. Serra, and R. Cucchiara, ‘‘A deep multi-level network for saliency prediction,’’ in ICPR, 2016.
  146. 146.G. Li and Y. Yu, ‘‘Visual saliency detection based on multiscale deep cnn features,’’ IEEE Trans. Image Process., vol. 25, no. 11, pp. 5012–5024, 2016.
  147. 147.J. Pan, E. Sayrol, X. Giro-i Nieto, K. McGuinness, and N. E. O’Connor, ‘‘Shallow and deep convolutional networks for saliency prediction,’’ in CVPR, 2016.
  148. 148.J. Kuen, Z. Wang, and G. Wang, ‘‘Recurrent attentional networks for saliency detection,’’ in CVPR, 2016.
  149. 149.Y. Tang, X. Wu, and W. Bu, ‘‘Deeply-supervised recurrent convolutional neural network for saliency detection,’’ in ACM MM, 2016.
  150. 150.X. Li, Y. Li, C. Shen, A. Dick, and A. Van Den Hengel, ‘‘Contextual hypergraph modeling for salient object detection,’’ in ICCV, 2013.
  151. 151.M.-M. Cheng, N. J. Mitra, X. Huang, P. H. Torr, and S.-M. Hu, ‘‘Global contrast based salient region detection,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 3, pp. 569–582, 2015.
  152. 152.H. Jiang, J. Wang, Z. Yuan, Y. Wu, N. Zheng, and S. Li, ‘‘Salient object detection: A discriminative regional feature integration approach,’’ in CVPR, 2013.
  153. 153.G. Lee, Y.-W. Tai, and J. Kim, ‘‘Deep saliency with encoded low level distance map and high level features,’’ in CVPR, 2016.
  154. 154.Z. Luo, A. Mishra, A. Achkar, J. Eichel, S. Li, and P.-M. Jodoin, ‘‘Non-local deep features for salient object detection,’’ in CVPR, 2017.
  155. 155.Q. Hou, M.-M. Cheng, X.-W. Hu, A. Borji, Z. Tu, and P. Torr, ‘‘Deeply supervised salient object detection with short connections,’’ arXiv:1611.04849, 2016.
  156. 156.Q. Yan, L. Xu, J. Shi, and J. Jia, ‘‘Hierarchical saliency detection,’’ in CVPR, 2013.
  157. 157.Y. Li, X. Hou, C. Koch, J. M. Rehg, and A. L. Yuille, ‘‘The secrets of salient object segmentation,’’ in CVPR, 2014.
  158. 158.V. Movahedi and J. H. Elder, ‘‘Design and perceptual validation of performance measures for salient object segmentation,’’ in CVPRW, 2010.
  159. 159.A. Borji, M.-M. Cheng, H. Jiang, and J. Li, ‘‘Salient object detection: A benchmark,’’ IEEE Trans. Image Process., vol. 24, no. 12, pp. 5706–5722, 2015.
  160. 160.C. Peng, X. Gao, N. Wang, and J. Li, ‘‘Graphical representation for heterogeneous face recognition,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 2, pp. 301–312, 2015.
  161. 161.C. Peng, N. Wang, X. Gao, and J. Li, ‘‘Face recognition from multiple stylistic sketches: Scenarios, datasets, and evaluation,’’ in ECCV, 2016.
  162. 162.X. Gao, N. Wang, D. Tao, and X. Li, ‘‘Face sketchcphoto synthesis and retrieval using sparse representation,’’ IEEE Trans. Circuits Syst. Video Technol., vol. 22, no. 8, pp. 1213–1226, 2012.
  163. 163.N. Wang, D. Tao, X. Gao, X. Li, and J. Li, ‘‘A comprehensive survey to face hallucination,’’ Int. J. of Comput. Vision, vol. 106, no. 1, pp. 9–30, 2014.
  164. 164.C. Peng, X. Gao, N. Wang, D. Tao, X. Li, and J. Li, ‘‘Multiple representations-based face sketch-photo synthesis.’’ IEEE Trans. Neural Netw. & Learning Syst., vol. 27, no. 11, pp. 2201–2215, 2016.
  165. 165.A. Majumder, L. Behera, and V. K. Subramanian, ‘‘Automatic facial expression recognition system using deep network-based data fusion,’’ IEEE Trans. Cybern., vol. 48, pp. 103–114, 2018.
  166. 166.P. Viola and M. Jones, ‘‘Robust real-time face detection,’’ Int. J. of Comput. Vision, vol. 57, no. 2, pp. 137–154, 2004.
  167. 167.J. Yu, Y. Jiang, Z. Wang, Z. Cao, and T. Huang, ‘‘Unitbox: An advanced object detection network,’’ in ACM MM, 2016.
  168. 168.S. S. Farfade, M. J. Saberian, and L.-J. Li, ‘‘Multi-view face detection using deep convolutional neural networks,’’ in ICMR, 2015.
  169. 169.S. Yang, P. Luo, C.-C. Loy, and X. Tang, ‘‘From facial parts responses to face detection: A deep learning approach,’’ in ICCV, 2015.
  170. 170.S. Yang, Y. Xiong, C. C. Loy, and X. Tang, ‘‘Face detection through scale-friendly deep convolutional networks,’’ in CVPR, 2017.
  171. 171.Z. Hao, Y. Liu, H. Qin, J. Yan, X. Li, and X. Hu, ‘‘Scale-aware face detection,’’ in CVPR, 2017.
  172. 172.H. Wang, Z. Li, X. Ji, and Y. Wang, ‘‘Face r-cnn,’’ arXiv:1706.01061, 2017.
  173. 173.X. Sun, P. Wu, and S. C. Hoi, ‘‘Face detection using deep learning: An improved faster rcnn approach,’’ arXiv:1701.08289, 2017.
  174. 174.L. Huang, Y. Yang, Y. Deng, and Y. Yu, ‘‘Densebox: Unifying landmark localization with end to end object detection,’’ arXiv:1509.04874, 2015.
  175. 175.Y. Li, B. Sun, T. Wu, and Y. Wang, ‘‘face detection with end-to-end integration of a convnet and a 3d model,’’ in ECCV, 2016.
  176. 176.K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, ‘‘Joint face detection and alignment using multitask cascaded convolutional networks,’’ IEEE Signal Process. Lett., vol. 23, no. 10, pp. 1499–1503, 2016.
  177. 177.I. A. Kalinovsky and V. G. Spitsyn, ‘‘Compact convolutional neural network cascadefor face detection,’’ in CEUR Workshop, 2016.
  178. 178.H. Qin, J. Yan, X. Li, and X. Hu, ‘‘Joint training of cascaded cnn for face detection,’’ in CVPR, 2016.
  179. 179.V. Jain and E. Learned-Miller, ‘‘Fddb: A benchmark for face detection in unconstrained settings,’’ Tech. Rep., 2010.
  180. 180.H. Li, Z. Lin, X. Shen, J. Brandt, and G. Hua, ‘‘A convolutional neural network cascade for face detection,’’ in CVPR, 2015.
  181. 181.B. Yang, J. Yan, Z. Lei, and S. Z. Li, ‘‘Aggregate channel features for multi-view face detection,’’ in IJCB, 2014.
  182. 182.N. Markus, M. Frljak, I. S. Pandɖziɖc, J. Ahlberg, and R. Forchheimer, ‘‘Object detection with pixel intensity comparisons organized in decision trees,’’ arXiv:1305.4537, 2013.
  183. 183.M. Mathias, R. Benenson, M. Pedersoli, and L. Van Gool, ‘‘Face detection without bells and whistles,’’ in ECCV, 2014.
  184. 184.J. Li and Y. Zhang, ‘‘Learning surf cascade for fast and accurate object detection,’’ in CVPR, 2013.
  185. 185.S. Liao, A. K. Jain, and S. Z. Li, ‘‘A fast and accurate unconstrained face detector,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 38, no. 2, pp. 211–223, 2016.
  186. 186.B. Yang, J. Yan, Z. Lei, and S. Z. Li, ‘‘Convolutional channel features,’’ in ICCV, 2015.
  187. 187.R. Ranjan, V. M. Patel, and R. Chellappa, ‘‘Hyperface: A deep multitask learning framework for face detection, landmark localization, pose estimation, and gender recognition,’’ arXiv:1603.01249, 2016.
  188. 188.P. Hu and D. Ramanan, ‘‘Finding tiny faces,’’ in CVPR, 2017.
  189. 189.Z. Jiang and D. Q. Huynh, ‘‘Multiple pedestrian tracking from monocular videos in an interacting multiple model framework,’’ IEEE Trans. Image Process., vol. 27, pp. 1361–1375, 2018.
  190. 190.D. Gavrila and S. Munder, ‘‘Multi-cue pedestrian detection and tracking from a moving vehicle,’’ Int. J. of Comput. Vision, vol. 73, pp. 41–59, 2006.
  191. 191.S. Xu, Y. Cheng, K. Gu, Y. Yang, S. Chang, and P. Zhou, ‘‘Jointly attentive spatial-temporal pooling networks for video-based person reidentification,’’ in ICCV, 2017.
  192. 192.Z. Liu, D. Wang, and H. Lu, ‘‘Stepwise metric promotion for unsupervised video person re-identification,’’ in ICCV, 2017.
  193. 193.A. Khan, B. Rinner, and A. Cavallaro, ‘‘Cooperative robots to observe moving targets: Review,’’ IEEE Trans. Cybern., vol. 48, pp. 187–198, 2018.
  194. 194.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, ‘‘Vision meets robotics: The kitti dataset,’’ Int. J. of Robotics Res., vol. 32, pp. 1231–1237, 2013.
  195. 195.Z. Cai, M. Saberian, and N. Vasconcelos, ‘‘Learning complexity-aware cascades for deep pedestrian detection,’’ in ICCV, 2015.
  196. 196.Y. Tian, P. Luo, X. Wang, and X. Tang, ‘‘Deep learning strong parts for pedestrian detection,’’ in CVPR, 2015.
  197. 197.P. Dollar, R. Appel, S. Belongie, and P. Perona, ‘‘Fast feature pyramids for object detection,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 36, no. 8, pp. 1532–1545, 2014.
  198. 198.S. Zhang, R. Benenson, and B. Schiele, ‘‘Filtered channel features for pedestrian detection,’’ in CVPR, 2015.
  199. 199.S. Paisitkriangkrai, C. Shen, and A. van den Hengel, ‘‘Pedestrian detection with spatially pooled features and structured ensemble learning,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 38, pp. 1243–1257, 2016.
  200. 200.L. Lin, X. Wang, W. Yang, and J.-H. Lai, ‘‘Discriminatively trained and-or graph models for object shape detection,’’ IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 5, pp. 959–972, 2015.
  201. 201.M. Mathias, R. Benenson, R. Timofte, and L. Van Gool, ‘‘Handling occlusions with franken-classifiers,’’ in ICCV, 2013.
  202. 202.S. Tang, M. Andriluka, and B. Schiele, ‘‘Detection and tracking of occluded people,’’ Int. J. of Comput. Vision, vol. 110, pp. 58–69, 2014.
  203. 203.L. Zhang, L. Lin, X. Liang, and K. He, ‘‘Is faster r-cnn doing well for pedestrian detection?’’ in ECCV, 2016.
  204. 204.Y. Tian, P. Luo, X. Wang, and X. Tang, ‘‘Deep learning strong parts for pedestrian detection,’’ in ICCV, 2015.
  205. 205.J. Liu, S. Zhang, S. Wang, and D. N. Metaxas, ‘‘Multispectral deep neural networks for pedestrian detection,’’ arXiv:1611.02644, 2016.
  206. 206.Y. Tian, P. Luo, X. Wang, and X. Tang, ‘‘Pedestrian detection aided by deep learning semantic tasks,’’ in CVPR, 2015.
  207. 207.X. Du, M. El-Khamy, J. Lee, and L. Davis, ‘‘Fused dnn: A deep neural network fusion approach to fast and robust pedestrian detection,’’ in WACV, 2017.
  208. 208.Q. Hu, P. Wang, C. Shen, A. van den Hengel, and F. Porikli, ‘‘Pushing the limits of deep cnns for pedestrian detection,’’ IEEE Trans. Circuits Syst. Video Technol., 2017.
  209. 209.D. Tome, L. Bondi, L. Baroffio, S. Tubaro, E. Plebani, and D. Pau, ‘‘Reduced memory region based deep convolutional neural network detection,’’ in ICCE-Berlin, 2016.
  210. 210.J. Hosang, M. Omran, R. Benenson, and B. Schiele, ‘‘Taking a deeper look at pedestrians,’’ in CVPR, 2015.
  211. 211.J. Li, X. Liang, S. Shen, T. Xu, J. Feng, and S. Yan, ‘‘Scale-aware fast r-cnn for pedestrian detection,’’ arXiv:1510.08160, 2015.
  212. 212.Y. Gao, M. Wang, Z.-J. Zha, J. Shen, X. Li, and X. Wu, ‘‘Visual-textual joint relevance learning for tag-based social image search,’’ IEEE Trans. Image Process., vol. 22, no. 1, pp. 363–376, 2013.
  213. 213.T. Kong, F. Sun, A. Yao, H. Liu, M. Lv, and Y. Chen, ‘‘Ron: Reverse connection with objectness prior networks for object detection,’’ in CVPR, 2017.
  214. 214.I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, ‘‘Generative adversarial nets,’’ in NIPS, 2014.
  215. 215.Y. Fang, K. Kuan, J. Lin, C. Tan, and V. Chandrasekhar, ‘‘Object detection meets knowledge graphs,’’ in IJCAI, 2017.
  216. 216.S. Welleck, J. Mao, K. Cho, and Z. Zhang, ‘‘Saliency-based sequential image attention with multiset prediction,’’ in NIPS, 2017.
  217. 217.S. Azadi, J. Feng, and T. Darrell, ‘‘Learning detection with diverse proposals,’’ in CVPR, 2017.
  218. 218.S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus, ‘‘End-to-end memory networks,’’ in NIPS, 2015.
  219. 219.P. Dabkowski and Y. Gal, ‘‘Real time image saliency for black box classifiers,’’ in NIPS, 2017.
  220. 220.B. Yang, J. Yan, Z. Lei, and S. Z. Li, ‘‘Craft objects from images,’’ in CVPR, 2016.
  221. 221.I. Croitoru, S.-V. Bogolin, and M. Leordeanu, ‘‘Unsupervised learning from video to detect foreground objects in single images,’’ in ICCV, 2017.
  222. 222.C. Wang, W. Ren, K. Huang, and T. Tan, ‘‘Weakly supervised object localization with latent category learning,’’ in ECCV, 2014.
  223. 223.D. P. Papadopoulos, J. R. R. Uijlings, F. Keller, and V. Ferrari, ‘‘Training object class detectors with click supervision,’’ in CVPR, 2017.
  224. 224.J. Huang, V. Rathod, C. Sun, M. Zhu, A. K. Balan, A. Fathi, I. Fischer, Z. Wojna, Y. S. Song, S. Guadarrama, and K. Murphy, ‘‘Speed/accuracy trade-offs for modern convolutional object detectors,’’ in CVPR, 2017.
  225. 225.Q. Li, S. Jin, and J. Yan, ‘‘Mimicking very efficient network for object detection,’’ in CVPR, 2017.
  226. 226.G. Hinton, O. Vinyals, and J. Dean, ‘‘Distilling the knowledge in a neural network,’’ Comput. Sci., vol. 14, no. 7, pp. 38–39, 2015.
  227. 227.A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, ‘‘Fitnets: Hints for thin deep nets,’’ Comput. Sci., 2014.
  228. 228.X. Chen, K. Kundu, Y. Zhu, A. G. Berneshawi, H. Ma, S. Fidler, and R. Urtasun, ‘‘3d object proposals for accurate object class detection,’’ in NIPS, 2015.
  229. 229.J. Dong, X. Fei, and S. Soatto, ‘‘Visual-inertial-semantic scene representation for 3d object detection,’’ in CVPR, 2017.
  230. 230.K. Kang, H. Li, T. Xiao, W. Ouyang, J. Yan, X. Liu, and X. Wang, ‘‘Object detection in videos with tubelet proposal networks,’’ in CVPR, 2017.

Citation

MLA
Zhao, Z.-Q., et al. “Object Detection With Deep Learning: A Review”. IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 11, 2019, pp. 3212–32, https://doi.org/10.1109/TNNLS.2018.2876865.
APA
Zhao, Z.-Q., Zheng, P., Xu, S.-T., & Wu, X. (2019). Object Detection With Deep Learning: A Review. IEEE Transactions on Neural Networks and Learning Systems, 30(11), 3212–3232. https://doi.org/10.1109/TNNLS.2018.2876865
Chicago
Zhao, Z.-Q., P. Zheng, S.-T. Xu, and X. Wu. 2019. “Object Detection With Deep Learning: A Review”. IEEE Transactions on Neural Networks and Learning Systems 30 (11): 3212–32. https://doi.org/10.1109/TNNLS.2018.2876865.
Harvard
Zhao, Z.-Q. et al. (2019) “Object Detection With Deep Learning: A Review”, IEEE Transactions on Neural Networks and Learning Systems, 30(11), pp. 3212–3232. Available at: https://doi.org/10.1109/TNNLS.2018.2876865.
Vancouver
1. Zhao Z-Q, Zheng P, Xu S-T, Wu X (2019) Object Detection With Deep Learning: A Review. IEEE Transactions on Neural Networks and Learning Systems 30:3212–3232

BibTeX

@article{Zhao_2019, title={Object Detection With Deep Learning: A Review}, volume={30}, ISSN={2162-2388}, url={http://dx.doi.org/10.1109/TNNLS.2018.2876865}, DOI={10.1109/tnnls.2018.2876865}, number={11}, journal={IEEE Transactions on Neural Networks and Learning Systems}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Zhao, Zhong-Qiu and Zheng, Peng and Xu, Shou-Tao and Wu, Xindong}, year={2019}, month=Nov, pages={3212–3232} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF