Dense Learning based Semi-Supervised Object Detection

Binghui ChenPengyu LiXiang ChenBiao WangLei ZhangXian-Sheng Hua

article2022CVPR84 citations

Proposes an anchor-free semi-supervised object detection framework that assigns dense pixel-level pseudo-labels via adaptive filtering and scale-consistent regularization to substantially outperform anchor-based methods on limited labeled data.

Listen

Deploying modern computer vision systems at scale is often constrained by the high cost and labor required to manually annotate large image datasets. Semi-supervised object detection addresses this bottleneck by using a small set of labeled images alongside abundant unlabeled data. However, existing semi-supervised methods almost exclusively rely on anchor-based architectures. While effective, these anchor-based systems require complex pre-processing and post-processing steps, making them difficult and inefficient to deploy on resource-constrained edge devices where one-stage, anchor-free detectors are preferred.

The article aims to design and evaluate the first anchor-free semi-supervised object detection framework, named Dense Learning. This approach bridges the practical deployment gap by demonstrating that anchor-free detectors can effectively utilize unlabeled data while outperforming traditional anchor-based models.

To overcome the challenge of noisy, dense pixel-level supervision inherent in anchor-free architectures, the authors developed a multi-part learning strategy. They introduced an adaptive filtering mechanism that categorizes predictions into foreground, background, and ignorable regions, supplemented by a lightweight network to remove high-confidence classification errors. In addition, an aggregated teacher model combines parameter updates over time with layer-to-layer connections to generate high-quality pseudo-labels. Finally, patch shuffling and multi-scale consistency regularization were applied to improve model generalization. The framework was evaluated across standard benchmarks, including the MS-COCO and PASCAL-VOC datasets, under both partially labeled and fully labeled settings.

The experimental findings show substantial performance improvements across all evaluation benchmarks. Under the MS-COCO benchmark with only 10% labeled data, the proposed method improved detection accuracy from a supervised baseline of 23.7% mean average precision to 36.2%, surpassing existing state-of-the-art anchor-based methods. On the PASCAL-VOC benchmark using unlabeled supplementary data, the method achieved up to 59.8% mean average precision, outperforming competing techniques by several percentage points. Ablation analyses confirmed that each component contributed measurably to accuracy, with adaptive filtering and aggregated teacher modeling providing the largest individual gains.

These results demonstrate that anchor-free architectures can match or exceed the accuracy of more cumbersome anchor-based models when training with limited labeled data. For technical organizations, this delivers two distinct operational advantages: significantly reduced data labeling costs and a streamlined model architecture that requires negligible pre- and post-processing, thereby lowering inference latency and hardware deployment costs on edge devices.

Organizations developing computer vision pipelines should consider adopting dense semi-supervised learning techniques when deploying models to resource-constrained environments. Engineering teams can leverage the publicly available codebase to pilot this anchor-free approach on internal datasets. While the reported results show high statistical confidence and consistent gains across multiple benchmark folds, practitioners should note that hyperparameter tuning, such as the weighting of unlabeled loss, remains sensitive to data scale and requires careful calibration during implementation.

arXiv: 2204.07300
Cover for Dense Learning based Semi-Supervised Object Detection

Abstract

Semi-supervised object detection (SSOD) aims to facilitate the training and deployment of object detectors with the help of a large amount of unlabeled data. Though various self-training based and consistency-regularization based SSOD methods have been proposed, most of them are anchor-based detectors, ignoring the fact that in many real-world applications anchor-free detectors are more demanded. In this paper, we intend to bridge this gap and propose a DenSe Learning (DSL) based anchor-free SSOD algorithm. Specifically, we achieve this goal by introducing several novel techniques, including an Adaptive Filtering strategy for assigning multi-level and accurate dense pixel-wise pseudo-labels, an Aggregated Teacher for producing stable and precise pseudo-labels, and an uncertainty-consistency-regularization term among scales and shuffled patches for improving the generalization capability of the detector. Extensive experiments are conducted on MS-COCO and PASCAL-VOC, and the results show that our proposed DSL method records new state-of-the-art SSOD performance, surpassing existing methods by a large margin. Codes can be found at https://github.com/chenbinghui1/DSL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methods
  • 3.1. Preliminary
  • 3.2. Adaptive Filtering Strategy
  • 3.3. MetaNet
  • 3.4. Aggregated Teacher
  • 3.5. Uncertainty Consistency
  • 4. Experiments
  • 4.1. Comparison with State-of-the-Arts
  • 4.2. Ablation Studies
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — DenSe Learning (DSL) Anchor-Free Semi-Supervised Object Detection Framework

    model/method

    DenSe Learning (DSL) is a semi-supervised object detection (SSOD) framework designed for anchor-free one-stage detectors (specifically evaluated using FCOS with a ResNet-50 backbone and Feature Pyramid Network). Unlike anchor-based detectors (such as Faster R-CNN) that perform region proposal matching and multi-stage refinement, anchor-free detectors directly assign supervision to dense pixel-wise spatial locations, making them particularly vulnerable to noisy pseudo-labels on unlabeled data.

    DSL addresses dense semi-supervised learning through four synergistic mechanisms:

    1. Adaptive Filtering (AF): Classifies pixel-level predictions on weakly augmented unlabeled images into foreground, background, and ignorable regions using dual class-adaptive confidence thresholds, eliminating ambiguous gradient updates.
    2. MetaNet: Uses class proxy representations derived from labeled data to filter out high-confidence classification false positives by reassigning them to ignorable regions.
    3. Aggregated Teacher (AT): Generates stable pseudo-labels by combining temporal Exponential Moving Average (EMA) parameter updates across training iterations with Recurrent Layer Aggregation (RLA) across depth in the backbone.
    4. Uncertainty-Consistency Regularization: Applies scale consistency constraints across feature pyramid levels between downsampled images and strongly augmented patch-shuffled images.

    The overall optimization objective for the student network is: L=Ls+αLu+Lscale\mathcal{L} = \mathcal{L}_s + \alpha \mathcal{L}_u + \mathcal{L}_{scale} where Ls\mathcal{L}_s is the supervised loss on labeled images, Lu\mathcal{L}_u is the unsupervised loss on unlabeled images with ignorable pixels masked out, Lscale\mathcal{L}_{scale} is the scale uncertainty consistency loss, and α\alpha is a weighting hyperparameter.

  2. Knowl 2 — Dense Pixel-Level Loss Formulation for Semi-Supervised Anchor-Free Detection

    equation

    In anchor-free semi-supervised object detection, dense spatial predictions at each location (h,w)(h, w) of an image feature map contain a classification score vector, bounding box regression offsets, and a centerness score. Let X={Xi}i=1Nl\mathcal{X} = \{X_i\}_{i=1}^{N_l} denote the labeled dataset with category annotations ph,w∗∈[0,C−1]p^*_{h,w} \in [0, C-1] (CC being the number of foreground classes) and bounding box annotations t∗t^*. Let U={Ui}i=1Nu\mathcal{U} = \{U_i\}_{i=1}^{N_u} denote the unlabeled dataset (Nu≫NlN_u \gg N_l) with estimated dense pseudo-labels pˉh,w∗∈{−1,0,…,C}\bar{p}^*_{h,w} \in \{-1, 0, \dots, C\}, where CC denotes the background class and −1-1 denotes an ignorable region.

    The supervised loss Ls\mathcal{L}_s and unsupervised loss Lu\mathcal{L}_u normalized by the number of positive pixels NposN_{pos} in the mini-batch are formulated as: Ls=1Npos∑i∑h,w(Lcls(Xi,h,w)+I{ph,w∗∈[0,C−1]}Lreg(Xi,h,w)+I{ph,w∗∈[0,C−1]}Lcenter(Xi,h,w))\mathcal{L}_s = \frac{1}{N_{pos}} \sum_{i} \sum_{h,w} \left( \mathcal{L}_{cls}(X_{i,h,w}) + \mathbb{I}_{\{p^*_{h,w} \in [0, C-1]\}} \mathcal{L}_{reg}(X_{i,h,w}) + \mathbb{I}_{\{p^*_{h,w} \in [0, C-1]\}} \mathcal{L}_{center}(X_{i,h,w}) \right) Lu=1Npos∑i∑h,w(I{pˉh,w∗≥0}Lcls(Ui,h,w)+I{pˉh,w∗∈[0,C−1]}Lreg(Ui,h,w)+I{pˉh,w∗∈[0,C−1]}Lcenter(Ui,h,w))\mathcal{L}_u = \frac{1}{N_{pos}} \sum_{i} \sum_{h,w} \left( \mathbb{I}_{\{\bar{p}^*_{h,w} \ge 0\}} \mathcal{L}_{cls}(U_{i,h,w}) + \mathbb{I}_{\{\bar{p}^*_{h,w} \in [0, C-1]\}} \mathcal{L}_{reg}(U_{i,h,w}) + \mathbb{I}_{\{\bar{p}^*_{h,w} \in [0, C-1]\}} \mathcal{L}_{center}(U_{i,h,w}) \right) where Lcls\mathcal{L}_{cls} is the focal classification loss, Lreg\mathcal{L}_{reg} is the IoU bounding box regression loss, Lcenter\mathcal{L}_{center} is the binary cross-entropy centerness loss from FCOS, and I{⋅}\mathbb{I}_{\{\cdot\}} is the indicator function. In the unsupervised loss, the indicator I{pˉh,w∗≥0}\mathbb{I}_{\{\bar{p}^*_{h,w} \ge 0\}} computes classification loss only for foreground and background pixels while ignoring pixels labeled −1-1.

  3. Knowl 3 — Adaptive Filtering Strategy for Dense Pseudo-Label Partitioning

    model/method

    In anchor-free detection, using a single score threshold to separate foreground from background causes severe label noise because true positive predictions with high and low bounding box IoUs overlap heavily with background predictions. The Adaptive Filtering (AF) strategy partitions dense predictions into three distinct categories using dual thresholds {τ1,τ2k}\{\tau_1, \tau_2^k\}: pˉh,w∗={Foreground: [0,…,C−1],ph,w≥τ2kIgnorable Region: [−1],τ1<ph,w<τ2kBackground: [C],ph,w≤τ1\bar{p}_{h,w}^* = \begin{cases} \text{Foreground: } [0, \dots, C-1], & p_{h,w} \ge \tau_2^k \\ \text{Ignorable Region: } [-1], & \tau_1 < p_{h,w} < \tau_2^k \\ \text{Background: } [C], & p_{h,w} \le \tau_1 \end{cases} where ph,wp_{h,w} is the combined prediction score (the product of classification score and centerness score) at location (h,w)(h, w).

    The background threshold is fixed to τ1=0.1\tau_1 = 0.1. The foreground threshold τ2k\tau_2^k is dynamically adapted for each class k∈[0,C−1]k \in [0, C-1] based on the model's prediction confidence on that class: τ2k=(∑h,wI{pˉh,w∗==k}ph,wNpos)βτ\tau_2^k = \left( \frac{\sum_{h,w} \mathbb{I}_{\{\bar{p}_{h,w}^* == k\}} p_{h,w}}{N_{pos}} \right)^\beta \tau where τ=0.35\tau = 0.35 is a fixed base reference threshold, β=0.7\beta = 0.7 adjusts focus on long-tail categories, and NposN_{pos} is the mini-batch positive pixel count. The adaptive threshold τ2k\tau_2^k is bounded within [0.25,0.35][0.25, 0.35]. Pixels in the ignorable region (pˉh,w∗=−1\bar{p}_{h,w}^* = -1) are excluded from both classification and regression gradient back-propagation.

  4. Knowl 4 — MetaNet Pseudo-Label Rectification

    model/method

    To filter out classification false-positive instances that achieve high prediction scores despite wrong category assignments, DSL employs a MetaNet implemented with a ResNet-50 backbone. MetaNet operates in a plug-and-play fashion without gradient back-propagation during student training.

    Prior to DSL training, all labeled instances from X\mathcal{X} are passed through MetaNet to compute a 1D feature representation for each instance. For each foreground class kk, a class proxy vector mkm_k is computed as the mean feature vector: mk=∑ifi,kNkm_k = \frac{\sum_i f_{i,k}}{N_k} where fi,kf_{i,k} is the 1D feature vector of the ii-th labeled instance belonging to class kk, and NkN_k is the total count of labeled instances of class kk.

    During training on unlabeled images, predicted foreground instances from the teacher model are passed through MetaNet. The cosine distance between the instance's feature vector and the corresponding class proxy mkm_k is computed. If the cosine distance is below a threshold d=0.6d = 0.6, the instance's label is reassigned from 'Foreground' to 'Ignorable Region' (−1-1), eliminating its influence on positive pseudo-supervision.

  5. Knowl 5 — Aggregated Teacher via Recurrent Layer and Temporal Aggregation

    model/method

    Standard Exponential Moving Average (EMA) teacher models update weights independently per layer across training iterations, ignoring inter-layer correlations. The Aggregated Teacher (AT) integrates both temporal parameter aggregation across iterations and Recurrent Layer Aggregation (RLA) across depth in the detector backbone.

    1. Temporal Parameter Aggregation: Teacher weights θ′\theta' are updated from student weights θ\theta at iteration tt using EMA with smoothing parameter ϵ=0.99\epsilon = 0.99: θ′t=ϵθ′t−1+(1−ϵ)θt\theta'^{t} = \epsilon \theta'^{t-1} + (1 - \epsilon) \theta^{t}

    2. Recurrent Layer Aggregation (RLA): In the backbone network, each layer ll propagates a hidden state tensor hlh_l to preserve inter-layer relationships: xl+1=θl+1[xl+hl]+xlx_{l+1} = \theta_{l+1}[x_l + h_l] + x_l hl+1=g2[g1[θl+1[xl+hl]]+hl]h_{l+1} = g_2[g_1[\theta_{l+1}[x_l + h_l]] + h_l] where xlx_l is the ll-th layer feature tensor, θl\theta_l denotes convolution parameters, h1h_1 is initialized to zero, and g1g_1 (1×11\times 1 convolution) and g2g_2 (3×33\times 3 convolution) are recurrent layers parameter-shared across adjacent layers within the same stage. When hl−1h_{l-1} is omitted, the formulation degenerates to a standard ResNet residual block. RLA is applied specifically across the backbone stages to improve pseudo-label stability and accuracy.

  6. Knowl 6 — Scale Uncertainty Consistency Regularization

    model/method

    In anchor-free architectures like FCOS, Feature Pyramid Networks (FPN) provide 5 resolution levels v∈{1,…,5}v \in \{1, \dots, 5\} to detect objects at different scales. To enforce scale and spatial invariance on unlabeled data, DSL minimizes the discrepancy between output dense score maps across adjacent pyramid levels under geometric transformations.

    For an unlabeled image UiU_i, two augmented inputs are generated:

    1. UspU_{sp}: A strongly augmented and patch-shuffled image.
    2. UdU_d: A downsampled image obtained by scaling UiU_i down by a factor of r=2r = 2.

    Because of the 2×2\times downsampling factor, the spatial resolution of pyramid level vv on UdU_d is identical to the spatial resolution of pyramid level v+1v+1 on UspU_{sp}. The scale uncertainty consistency loss enforces feature consistency across these aligned levels: Lscale=∑v=14∥pv[Ud]−pv+1[Usp]∥22\mathcal{L}_{scale} = \sum_{v=1}^4 \|p^v[U_d] - p^{v+1}[U_{sp}]\|_2^2 where pv[⋅]p^v[\cdot] denotes the dense score map at scale level vv. This constraint regularizes prediction uncertainty across both context shuffling and scaling variations.

  7. Knowl 7 — Patch Shuffle Augmentation Algorithm

    algorithm

    Patch Shuffle is a data augmentation procedure applied to unlabeled images to reduce the detector's dependency on global background context and improve robustness to spatial occlusion.

    Input: Unlabeled image UU, total iteration number JJ (J=2J=2)
    Output: Patch-shuffled image UpU_p
    U0←UU^0 \leftarrow U
    for j=0j = 0 to J−1J - 1 do
        Randomly select mode m∈{’horizontal’,’vertical’}m \in \{\text{'horizontal'}, \text{'vertical'}\}
        Randomly sample normalized split ratio s∼Uniform(0,1)s \sim \text{Uniform}(0, 1)
        Crop UjU^j into two sub-images along direction mm at split position ss
        Swap the order of the two sub-images and concatenate them into U^j\hat{U}^j
        Uj+1←U^jU^{j+1} \leftarrow \hat{U}^j
    end for
    return UJU^J

    The resulting image UpU_p is combined with standard strong augmentations (random horizontal flip, color jittering, and cutout) to form the transformed input UspU_{sp} used for uncertainty consistency regularization.

  8. Knowl 8 — Semi-Supervised Detection Performance on MS-COCO Partially Labeled Protocol

    empirical result

    On the MS-COCO benchmark under the Partially Labeled Data protocol (evaluating randomly sampled subsets of 1%, 2%, 5%, and 10% labeled data across 3 folds, using ResNet-50 FCOS detector), DSL sets new state-of-the-art performance, outperforming both its supervised baseline and prior anchor-based SSOD methods.

    Methods Deployment 1% 2% 5% 10%
    Supervised (Faster R-CNN) Hard 9.05±0.169.05 \pm 0.16 12.70±0.1512.70 \pm 0.15 18.47±0.2218.47 \pm 0.22 23.86±0.8123.86 \pm 0.81
    CSD Hard 11.12±0.1511.12 \pm 0.15 14.15±0.1314.15 \pm 0.13 18.79±0.1318.79 \pm 0.13 24.50±0.1524.50 \pm 0.15
    STAC Hard 13.97±0.3513.97 \pm 0.35 18.25±0.2518.25 \pm 0.25 24.38±0.1224.38 \pm 0.12 28.64±0.2128.64 \pm 0.21
    Instant-Teaching (IT) Hard 16.00±0.2016.00 \pm 0.20 20.70±0.3020.70 \pm 0.30 25.50±0.0525.50 \pm 0.05 29.45±0.1529.45 \pm 0.15
    ISMT Hard 18.88±0.7418.88 \pm 0.74 22.43±0.5622.43 \pm 0.56 26.37±0.2426.37 \pm 0.24 30.53±0.5230.53 \pm 0.52
    Humble Teacher Hard 16.96±0.3816.96 \pm 0.38 21.72±0.2421.72 \pm 0.24 27.70±0.1527.70 \pm 0.15 31.60±0.2831.60 \pm 0.28
    Unbiased Teacher (UB) Hard 20.75±0.1220.75 \pm 0.12 24.30±0.9724.30 \pm 0.97 28.27±0.1128.27 \pm 0.11 31.50±0.1031.50 \pm 0.10
    Soft Teacher (E2E) Hard 20.46±0.3920.46 \pm 0.39 – 30.74±0.0830.74 \pm 0.08 34.04±0.1434.04 \pm 0.14
    Supervised Baseline (FCOS) Easy 9.53±0.239.53 \pm 0.23 11.71±0.2611.71 \pm 0.26 18.74±0.1818.74 \pm 0.18 23.70±0.2223.70 \pm 0.22
    DSL (Ours, FCOS) Easy 22.03±0.28\mathbf{22.03 \pm 0.28} 25.19±0.37\mathbf{25.19 \pm 0.37} 30.87±0.24\mathbf{30.87 \pm 0.24} 36.22±0.18\mathbf{36.22 \pm 0.18}

    All values denote AP50:90AP_{50:90} (mAP, in %). While the anchor-free supervised FCOS baseline exhibits performance comparable to the anchor-based Faster R-CNN baseline (e.g., 23.70% vs. 23.86% at 10% labeled data), DSL achieves a +12.52% mAP improvement on the 10% setting, surpassing Soft Teacher by +2.18% mAP with simpler deployment requirements (no RPN/RoI processing).

  9. Knowl 9 — Semi-Supervised Detection Performance on Fully Labeled MS-COCO and PASCAL-VOC

    empirical result

    DSL improves object detection performance on full datasets by utilizing additional unlabeled images:

    1. MS-COCO Fully Labeled Protocol (100% labeled train set + 123k unlabeled images):

      • Supervised FCOS Baseline: 40.2%40.2\% mAP (AP50:90AP_{50:90}).
      • DSL (FCOS): 43.8%43.8\% mAP (+3.6%+3.6\% improvement).
      • Comparison: STAC improves 37.6→39.237.6 \to 39.2 (+1.6), ISMT improves 37.8→39.637.8 \to 39.6 (+1.8), Unbiased Teacher improves 40.2→41.340.2 \to 41.3 (+1.1), and Soft Teacher improves 40.9→44.540.9 \to 44.5 (+3.6).
    2. PASCAL-VOC Benchmark (trained on VOC07 labeled, evaluated on VOC07 test set):

    Unlabeled: VOC12 Unlabeled: VOC12 + COCO20
    Method Deployment AP50AP_{50} AP50:90AP_{50:90} AP50AP_{50} AP50:90AP_{50:90}
    Supervised (Faster R-CNN) Hard 72.75 42.04 72.75 42.04
    CSD Hard 74.70 – 75.10 –
    STAC Hard 77.45 44.64 79.08 46.01
    Instant-Teaching (IT) Hard 78.30 48.70 79.00 49.70
    ISMT Hard 77.23 46.23 77.75 49.59
    Unbiased Teacher (UB) Hard 77.37 48.69 78.82 50.34
    Supervised Baseline (FCOS) Easy 69.60 45.90 69.60 45.90
    DSL (Ours, FCOS) Easy 80.70 56.80 82.10 59.80

    On PASCAL-VOC with VOC12 + COCO20 unlabeled data, DSL achieves 82.10% AP50AP_{50} and 59.80% AP50:90AP_{50:90}, outperforming the best anchor-based method (Unbiased Teacher) by +3.28% AP50AP_{50} and +9.46% AP50:90AP_{50:90}.

  10. Knowl 10 — Ablation Studies on DSL Components and Hyperparameters

    empirical result

    Ablation experiments conducted on the MS-COCO 10% labeled data setting demonstrate the individual contribution of each proposed module and hyperparameter configuration:

    1. Component-wise Accumulation:

      • Supervised FCOS baseline: 23.7%23.7\% mAP
        • Adaptive Filtering (AF): 32.2%32.2\% mAP (+8.5%+8.5\% gain)
        • MetaNet: 32.5%32.5\% mAP (+0.3%+0.3\% gain)
        • Aggregated Teacher (AT): 34.5%34.5\% mAP (+2.0%+2.0\% gain)
        • Patch-Shuffle: 34.9%34.9\% mAP (+0.4%+0.4\% gain)
        • Lscale\mathcal{L}_{scale} (Full DSL): 36.2%36.2\% mAP (+1.3%+1.3\% gain)
    2. Filtering Strategies:

      • Single fixed threshold τ∈{0.05,0.1,0.2,0.3}\tau \in \{0.05, 0.1, 0.2, 0.3\} produces mAPs of 27.1%27.1\%, 28.8%28.8\%, 30.7%30.7\%, and 27.5%27.5\% respectively.
      • Dual threshold AF with fixed τ2∈{0.2,0.3,0.4}\tau_2 \in \{0.2, 0.3, 0.4\} achieves 34.3%34.3\%, 36.0%36.0\%, and 35.6%35.6\% mAP.
      • Class-adaptive threshold τ2k\tau_2^k achieves the best result (36.2%36.2\% mAP).
    3. Teacher Model Aggregation:

      • No teacher: 33.0%33.0\% mAP
      • Parameter EMA only: 34.1%34.1\% mAP
      • Recurrent Layer Aggregation (LA) only: 35.0%35.0\% mAP
      • Aggregated Teacher (EMA + LA): 36.2%36.2\% mAP
    4. Unsupervised Loss Weight α\alpha:

      • α=1\alpha=1: 33.9%33.9\% mAP; α=2\alpha=2: 35.4%35.4\% mAP; α=3\alpha=3: 36.2%36.2\% mAP; α=4\alpha=4: training fails (loss becomes NaN).

Coverage note — None was omitted; all key algorithmic components (Adaptive Filtering, MetaNet, Aggregated Teacher, Scale Uncertainty Consistency, Patch Shuffle) and all primary experimental results on MS-COCO and PASCAL-VOC are represented.

References

  1. 1.Sean Bell, C Lawrence Zitnick, Kavita Bala, and Ross Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2874–2883, 2016.
  2. 2.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi-supervised learning. arXiv preprint arXiv:1905.02249, 2019.
  3. 3.Zhaowei Cai, Quanfu Fan, Rogerio S Feris, and Nuno Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In European conference on computer vision, pages 354–370. Springer, 2016.
  4. 4.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154–6162, 2018.
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020.
  6. 6.Binghui Chen and Weihong Deng. Weakly-supervised deep self-learning for face recognition. In 2016 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6. IEEE, 2016.
  7. 7.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  8. 8.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  10. 10.Jay Heo, Hae Beom Lee, Saehoon Kim, Juho Lee, Kwang Joon Kim, Eunho Yang, and Sung Ju Hwang. Uncertainty-aware attention for reliable interpretation and prediction. arXiv preprint arXiv:1805.09653, 2018.
  11. 11.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  12. 12.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.
  13. 13.Lichao Huang, Yi Yang, Yafeng Deng, and Yinan Yu. Densebox: Unifying landmark localization with end to end object detection. arXiv preprint arXiv:1509.04874, 2015.
  14. 14.Jisoo Jeong, Seungeui Lee, Jeesoo Kim, and Nojun Kwak. Consistency-based semi-supervised learning for object detection. Advances in neural information processing systems, 32:10759–10768, 2019.
  15. 15.Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? arXiv preprint arXiv:1703.04977, 2017.
  16. 16.Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7482–7491, 2018.
  17. 17.Kang Kim and Hee Seok Lee. Probabilistic anchor assignment with iou prediction for object detection. In ECCV, 2020.
  18. 18.Tao Kong, Fuchun Sun, Huaping Liu, Yuning Jiang, Lei Li, and Jianbo Shi. Foveabox: Beyound anchor-based object detection. IEEE Transactions on Image Processing, 29:7389–7398, 2020.
  19. 19.Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016.
  20. 20.Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3, page 896, 2013.
  21. 21.Hyungtae Lee, Sungmin Eum, and Heesung Kwon. Me r-cnn: Multi-expert r-cnn for object detection. IEEE Transactions on Image Processing, 29:1030–1044, 2019.
  22. 22.Hengduo Li, Zuxuan Wu, Abhinav Shrivastava, and Larry S Davis. Rethinking pseudo labels for semi-supervised object detection. arXiv preprint arXiv:2106.00168, 2021.
  23. 23.Xinzhe Li, Qianru Sun, Yaoyao Liu, Qin Zhou, Shibao Zheng, Tat-Seng Chua, and Bernt Schiele. Learning to self-train for semi-supervised few-shot classification. Advances in Neural Information Processing Systems, 32:10276–10286, 2019.
  24. 24.Yanghao Li, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. Scale-aware trident networks for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6054–6063, 2019.
  25. 25.Tsungnan Lin, Bill G Horne, Peter Tino, and C Lee Giles. Learning long-term dependencies in narx recurrent neural networks. IEEE Transactions on Neural Networks, 7(6):1329–1338, 1996.
  26. 26.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017.
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  28. 28.Songtao Liu, Zeming Li, and Jian Sun. Self-emd: Self-supervised object detection without imagenet. arXiv preprint arXiv:2011.13677, 2020.
  29. 29.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In European conference on computer vision, pages 21–37. Springer, 2016.
  30. 30.Wei Liu, Shengcai Liao, Weiqiang Ren, Weidong Hu, and Yinan Yu. High-level semantic feature detection: A new perspective for pedestrian detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5187–5196, 2019.
  31. 31.Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda. Unbiased teacher for semi-supervised object detection. arXiv preprint arXiv:2102.09480, 2021.
  32. 32.Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence, 41(8):1979–1993, 2018.
  33. 33.Pytorch. https://pytorch.org/.
  34. 34.Ilija Radosavovic, Piotr Dollar, Ross Girshick, Georgia Gkioxari, and Kaiming He. Data distillation: Towards omni-supervised learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4119–4128, 2018.
  35. 35.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 779–788, 2016.
  36. 36.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28:91–99, 2015.
  37. 37.Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685, 2020.
  38. 38.Kihyuk Sohn, Zizhao Zhang, Chun-Liang Li, Han Zhang, Chen-Yu Lee, and Tomas Pfister. A simple semi-supervised learning framework for object detection. arXiv preprint arXiv:2005.04757, 2020.
  39. 39.Xiaolin Song, Binghui Chen, Pengyu Li, Biao Wang, and Honggang Zhang. Prnet++: Learning towards generalized occluded pedestrian detection via progressive refinement network. Neurocomputing, 2022.
  40. 40.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al. Sparse r-cnn: End-to-end object detection with learnable proposals. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14454–14463, 2021.
  41. 41.Yihe Tang, Weifeng Chen, Yijun Luo, and Yuting Zhang. Humble teachers teach better students for semi-supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3132–3141, 2021.
  42. 42.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in Neural Information Processing Systems, 30, 2017.
  43. 43.Wanxin Tian, Zixuan Wang, Haifeng Shen, Weihong Deng, Yiping Meng, Binghui Chen, Xiubao Zhang, Yuan Zhao, and Xiehe Huang. Learning better features for face detection with feature fusion and segmentation supervision. arXiv preprint arXiv:1811.08557, 2018.
  44. 44.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9627–9636, 2019.
  45. 45.Zhenyu Wang, Yali Li, Ye Guo, Lu Fang, and Shengjin Wang. Data-uncertainty guided multi-phase learning for semi-supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4568–4577, 2021.
  46. 46.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10687–10698, 2020.
  47. 47.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. arXiv preprint arXiv:2106.09018, 2021.
  48. 48.Qize Yang, Xihan Wei, Biao Wang, Xian-Sheng Hua, and Lei Zhang. Interactive self-training with mean teachers for semi-supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5941–5950, 2021.
  49. 49.Fisher Yu, Dequan Wang, Evan Shelhamer, and Trevor Darrell. Deep layer aggregation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2403–2412, 2018.
  50. 50.Jingyu Zhao, Yanwen Fang, and Guodong Li. Recurrence along depth: Deep convolutional neural networks with recurrent layer aggregation. Advances in Neural Information Processing Systems, 34, 2021.
  51. 51.Qiang Zhou, Chaohui Yu, Zhibin Wang, Qi Qian, and Hao Li. Instant-teaching: An end-to-end semi-supervised object detection framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4081–4090, 2021.
  52. 52.Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk, and Quoc Le. Rethinking pre-training and self-training. Advances in Neural Information Processing Systems, 33, 2020.

Citation

MLA
Chen, B., et al. “Dense Learning Based Semi-Supervised Object Detection”. arXiv, 2022, http://arxiv.org/abs/2204.07300v1.
APA
Chen, B., Li, P., Chen, X., Wang, B., Zhang, L., & Hua, X.-S. (2022). Dense Learning based Semi-Supervised Object Detection. arXiv. http://arxiv.org/abs/2204.07300v1
Chicago
Chen, B., P. Li, X. Chen, B. Wang, L. Zhang, and X.-S. Hua. 2022. “Dense Learning Based Semi-Supervised Object Detection”. arXiv. http://arxiv.org/abs/2204.07300v1.
Harvard
Chen, B. et al. (2022) “Dense Learning based Semi-Supervised Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.07300v1.
Vancouver
1. Chen B, Li P, Chen X, Wang B, Zhang L, Hua X-S (2022) Dense Learning based Semi-Supervised Object Detection. arXiv

BibTeX

@article{chen2022dense,
  title = {Dense Learning based Semi-Supervised Object Detection},
  author = {Chen, Binghui and Li, Pengyu and Chen, Xiang and Wang, Biao and Zhang, Lei and Hua, Xian-Sheng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.07300v1},
  eprint = {2204.07300}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE