Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors

Yen-Cheng LiuChih-Yao MaZsolt Kira

article2022CVPR136 citations

Proposes a semi-supervised object detection framework that extends pseudo-labeling to anchor-free detectors while using relative uncertainty between teacher and student predictions to filter out misleading bounding box regression targets.

Listen

Modern computer vision models require large amounts of human-labeled image data to achieve high performance, creating severe operational bottlenecks and high data annotation costs. While semi-supervised learning techniques help train models using small subsets of labeled images combined with abundant unlabeled images, existing semi-supervised frameworks focus almost entirely on traditional anchor-based detectors (which rely on predefined bounding boxes) and struggle to refine object boundaries accurately during training. The article demonstrates a general semi-supervised object detection framework called Unbiased Teacher v2 that extends successfully to computationally efficient anchor-free detectors while establishing an uncertainty-guided learning mechanism to generate precise bounding box predictions.

The authors evaluated their approach across standard computer vision benchmarks—including Common Objects in Context (COCO) and PASCAL VOC—under varying levels of supervision, ranging from fully labeled datasets down to as little as 0.5% labeled data. They diagnosed the failure modes of previous semi-supervised techniques when applied to anchor-free detectors, specifically identifying that standard techniques like centerness scores and center-based label assignment introduce heavy localization noise when trained with limited supervision. To address boundary prediction errors, the authors designed the Listen2Student mechanism, which predicts individual boundary uncertainties for both a primary "Student" model and a guiding "Teacher" model. The framework selects pseudo-labels only for specific box boundaries where the Teacher displays strictly lower uncertainty than the Student, eliminating misleading regression targets.

The findings show that the proposed framework consistently outperforms existing state-of-the-art semi-supervised methods across all tested benchmark settings. In anchor-free settings on COCO using only 0.5% labeled data, the model achieved an average precision of 16.25, outperforming the previous baseline of 10.27 by roughly 58%. In anchor-based settings with 10% labeled data, the method raised overall precision to 35.08, surpassing competing semi-supervised methods. Furthermore, an evaluation breakdown across strict boundary accuracy thresholds confirmed that comparing Student and Teacher relative uncertainties consistently improves fine-grained boundary precision, whereas conventional confidence thresholding degraded accuracy on the strictest evaluation metrics. Crucially, the method bridges the historical performance gap between anchor-free and anchor-based architectures in limited-label settings without adding computational overhead during final model deployment.

These results demonstrate that organizations can deploy lighter, anchor-free computer vision models in data-constrained environments without sacrificing localization accuracy, significantly reducing manual labeling costs and accelerating deployment timelines. The authors recommend adopting uncertainty-aware pseudo-labeling for regression tasks and reverting to standard, robust label assignments rather than complex center-sampling heuristics when training semi-supervised anchor-free models. However, the authors note several limitations: performance has not yet been verified on massive, uncurated unlabeled datasets (such as OpenImages), uncertainty estimation methods can be further optimized, and practical deployment must account for potential dataset domain shifts, unseen object categories, and demographic data biases arising from low-supervision training.

arXiv: 2206.09500
Cover for Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors

Abstract

With the recent development of Semi-Supervised Object Detection (SS-OD) techniques, object detectors can be improved by using a limited amount of labeled data and abundant unlabeled data. However, there are still two challenges that are not addressed: (1) there is no prior SS-OD work on anchor-free detectors, and (2) prior works are ineffective when pseudo-labeling bounding box regression. In this paper, we present Unbiased Teacher v2, which shows the generalization of SS-OD method to anchor-free detectors and also introduces Listen2Student mechanism for the unsupervised regression loss. Specifically, we first present a study examining the effectiveness of existing SS-OD methods on anchor-free detectors and find that they achieve much lower performance improvements under the semi-supervised setting. We also observe that box selection with centerness and the localization-based labeling used in anchor-free detectors cannot work well under the semi-supervised setting. On the other hand, our Listen2Student mechanism explicitly prevents misleading pseudo-labels in the training of bounding box regression; we specifically develop a novel pseudo-labeling selection mechanism based on the Teacher and Student's relative uncertainties. This idea contributes to favorable improvement in the regression branch in the semi-supervised setting. Our method, which works for both anchor-free and anchor-based methods, consistently performs favorably against the state-of-the-art methods in VOC, COCO-standard, and COCO-additional.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Background: Semi-supervised Object Detection and Pseudo-labeling
  • 3.2. Pseudo-labeling on Anchor-Free Detectors
  • 3.3. Listen2Student for Unsupervised Regression Loss
  • 3.3.1 Limitations of confidence thresholding for regression
  • 3.3.2 Listen2Student
  • 4. Experiments
  • 4.1. Settings and Implementation Details
  • 4.2. Results on Anchor-free Detector
  • 4.3. Results on Anchor-based Detector
  • 4.4. Effectiveness of Unsupervised Regression Loss
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Listen2Student Mechanism for Unsupervised Regression Loss

    model/method

    The Listen2Student mechanism selects pseudo-labels for bounding box regression in semi-supervised object detection (SS-OD) by comparing the relative localization uncertainties predicted by the Teacher model and the Student model on unlabeled images.

    In standard Teacher-Student SS-OD, selecting regression targets using classification confidence or a single box-level score (such as centerness or box IoU) introduces misleading instances where the Teacher's regression error exceeds the Student's regression error: ∥d~t−dg∥>∥d~s−dg∥\|\tilde{d}_t - d_g\| > \|\tilde{d}_s - d_g\|, where d~t\tilde{d}_t is the Teacher's regression prediction, d~s\tilde{d}_s is the Student's regression prediction, and dgd_g is the true ground-truth distance. Listen2Student addresses this via three components:

    1. Boundary-wise uncertainty prediction: An auxiliary branch parallel to the bounding box regression head predicts an uncertainty score δ\delta independently for each of the four box boundaries.
    2. Low-uncertainty student filtering: Boundaries for which the Student already has very low localization uncertainty (δs≤σs\delta_s \le \sigma_s, with σs=0.5\sigma_s = 0.5) are filtered out.
    3. Relative uncertainty selection: For boundary ii, the Teacher's prediction d~ti\tilde{d}_t^i is used as a supervisory regression pseudo-label for the Student's prediction d~si\tilde{d}_s^i only if the Teacher is strictly more confident than the Student by a margin σ\sigma: δti+σ≤δsi\delta_t^i + \sigma \le \delta_s^i where σ≥0\sigma \ge 0 (set to σ=0.1\sigma = 0.1). This evaluation is performed independently per boundary, allowing certain edges of a predicted box to provide supervision while others are ignored.
  2. Knowl 2 — Supervised Regression Loss with Boundary Uncertainty

    equation

    To learn boundary-level localization uncertainties without ground-truth uncertainty labels, the detector's regression head is jointly trained on labeled data using the negative power log-likelihood loss (NPLL):

    Lregsup=∑iηi(∑b((ds,b−dg,b)22δs,b2+12log⁡δs,b2)+2log⁡(2π))\mathcal{L}^{sup}_{reg} = \sum_i \eta_i \left( \sum_{b} \left( \frac{(d_{s,b} - d_{g,b})^2}{2 \delta_{s,b}^2} + \frac{1}{2} \log \delta_{s,b}^2 \right) + 2 \log(2\pi) \right)

    where:

    • ii indexes the foreground predicted object instances,
    • ηi\eta_i is the Intersection-over-Union (IoU) score between the predicted box and the corresponding ground-truth bounding box for instance ii,
    • b∈{left,right,top,bottom}b \in \{\text{left}, \text{right}, \text{top}, \text{bottom}\} indexes the four boundary directions,
    • ds,b∈R+d_{s,b} \in \mathbb{R}^+ is the Student model's predicted distance to boundary bb,
    • dg,b∈R+d_{g,b} \in \mathbb{R}^+ is the ground-truth distance to boundary bb,
    • δs,b>0\delta_{s,b} > 0 is the Student model's predicted localization uncertainty for boundary bb.
  3. Knowl 3 — Listen2Student Unsupervised Regression Loss Formulation

    equation

    For unlabeled images, the boundary-level unsupervised regression loss Lregunsup\mathcal{L}^{unsup}_{reg} is computed as:

    Lregunsup={∑i∥d~ti−d~si∥1,if δti+σ≤δsi0,otherwise\mathcal{L}^{unsup}_{reg} = \begin{cases} \sum_i \|\tilde{d}^i_t - \tilde{d}^i_s\|_1, & \text{if } \delta^i_t + \sigma \le \delta^i_s \\ 0, & \text{otherwise} \end{cases}

    where:

    • ii indexes each individual boundary of a detected object instance,
    • d~ti∈R+\tilde{d}^i_t \in \mathbb{R}^+ is the Teacher model's predicted distance for boundary ii,
    • d~si∈R+\tilde{d}^i_s \in \mathbb{R}^+ is the Student model's predicted distance for boundary ii,
    • δti>0\delta^i_t > 0 and δsi>0\delta^i_s > 0 are the localization uncertainties predicted by the Teacher and Student for boundary ii, respectively,
    • σ≥0\sigma \ge 0 is a fixed uncertainty margin hyperparameter (set to σ=0.1\sigma = 0.1).

    The loss enforces L1L_1 consistency between the Teacher and Student boundary predictions only when the Teacher's uncertainty is lower than the Student's uncertainty by at least σ\sigma.

  4. Knowl 4 — Adaptations of Anchor-Free Detectors for Semi-Supervised Learning

    model/method

    When applying pseudo-labeling semi-supervised object detection (SS-OD) frameworks to anchor-free detectors (specifically FCOS), three modifications to standard fully-supervised anchor-free training pipelines are required to prevent performance degradation:

    1. Pseudo-box selection via classification score only: Pseudo-boxes from the Teacher are filtered solely by thresholding the classification confidence score (scls>τs_{cls} > \tau, where τ=0.5\tau = 0.5) rather than the combined box score (sbox=scls×scenters_{box} = s_{cls} \times s_{center}). This prevents background regions with erroneously high centerness scores from being selected as pseudo-labels.
    2. Hard classification labels: The classification branch is trained on pseudo-labels using standard one-hot hard labels rather than localization-weighted soft classification labels.
    3. Standard label assignment: Every spatial location/pixel lying inside a pseudo-bounding box is assigned as foreground, discarding center-sampling (which restricts foreground assignment to a small region near the box center). This avoids severe precision/recall drops caused by center shifts and localization noise in pseudo-boxes.
  5. Knowl 5 — Centerness Bias in Semi-Supervised Anchor-Free Detectors

    empirical result

    In fully supervised anchor-free detectors such as FCOS, ranking and selecting bounding boxes by box score (sbox=scls×scenters_{box} = s_{cls} \times s_{center}) improves detection performance over using classification score sclss_{cls} alone (37.10 mAP vs. 33.50 mAP, a +3.60+3.60 mAP gain on COCO).

    However, under semi-supervised settings with limited labeled data, the centerness branch lacks sufficient negative supervision to suppress centerness scores on background instances. Consequently, the centerness scores dominate sboxs_{box}, leading to the selection of false-positive background pseudo-labels with high centerness but low classification confidence. In semi-supervised FCOS experiments, selecting pseudo-boxes using box score yields 15.12 mAP, whereas selecting pseudo-boxes using classification score alone yields 17.79 mAP (a difference of −2.67-2.67 mAP for box score selection).

  6. Knowl 6 — Detrimental Effect of Center-Sampling under Pseudo-Label Localization Noise

    empirical result

    In fully-supervised FCOS, center-sampling (assigning only pixels close to the center of ground-truth boxes as foreground) improves performance over standard label assignment (treating all pixels inside the box as foreground) from 37.10 mAP to 38.10 mAP (+1.00+1.00 mAP).

    Under semi-supervised learning, Teacher-generated pseudo-boxes inherently exhibit localization noise (such as center shifts and dimension errors). Applying center-sampling to noisy pseudo-boxes causes pixel-wise foreground precision to drop from 32% (standard) to 0% (center-sampling) and recall to drop from 55% to 0% on noisy samples. Across training on COCO, center-sampling degrades semi-supervised performance from 17.79 mAP to 14.96 mAP (a drop of −2.83-2.83 mAP). Standard label assignment is robust to localization noise and outperforms center-sampling in SS-OD.

  7. Knowl 7 — COCO-Standard Benchmark Results on Anchor-Free Detector (FCOS-ResNet50)

    data/table

    Evaluation of Unbiased Teacher v2 (Ours) against SS-OD baselines adapted to the anchor-free FCOS-ResNet50 architecture across various labeling ratios on the COCO-standard benchmark (evaluated on COCO2017-val with 5 runs; batch size 8 labeled / 8 unlabeled images):

    Methods 0.5% 1% 2% 5% 10%
    Supervised 5.42±0.015.42 \pm 0.01 8.43±0.038.43 \pm 0.03 11.97±0.0311.97 \pm 0.03 17.01±0.0117.01 \pm 0.01 20.98±0.0120.98 \pm 0.01
    CSD 5.76±0.555.76 \pm 0.55 9.23±0.089.23 \pm 0.08 12.53±0.0412.53 \pm 0.04 18.09±0.0818.09 \pm 0.08 22.06±0.0122.06 \pm 0.01
    STAC 8.79±0.128.79 \pm 0.12 11.97±0.1211.97 \pm 0.12 15.50±0.1615.50 \pm 0.16 20.36±0.0520.36 \pm 0.05 24.31±0.0224.31 \pm 0.02
    Unbiased Teacher 10.27±0.1310.27 \pm 0.13 14.61±0.1014.61 \pm 0.10 18.70±0.2118.70 \pm 0.21 23.99±0.1223.99 \pm 0.12 28.18±0.0128.18 \pm 0.01
    Ours 16.25±0.18\mathbf{16.25 \pm 0.18} 22.71±0.42\mathbf{22.71 \pm 0.42} 26.03±0.12\mathbf{26.03 \pm 0.12} 30.08±0.04\mathbf{30.08 \pm 0.04} 32.61±0.03\mathbf{32.61 \pm 0.03}

    Unbiased Teacher v2 outperforms adapted baseline methods across all supervision ratios, yielding improvements over the supervised baseline ranging from +10.83+10.83 mAP at 0.5% labels to +14.28+14.28 mAP at 1% labels.

  8. Knowl 8 — COCO-Standard Benchmark Results on Anchor-Based Detector (Faster-RCNN-ResNet50)

    data/table

    Performance comparison on the COCO-standard benchmark using Faster-RCNN-ResNet50 across labeled data percentages (0.5% to 10%), evaluated on COCO2017-val over 5 runs:

    Methods 0.5% 1% 2% 5% 10%
    Supervised 6.83±0.156.83 \pm 0.15 9.05±0.169.05 \pm 0.16 12.70±0.1512.70 \pm 0.15 18.47±0.2218.47 \pm 0.22 23.86±0.8123.86 \pm 0.81
    CSD 7.41±0.217.41 \pm 0.21 10.51±0.0610.51 \pm 0.06 13.93±0.1213.93 \pm 0.12 18.63±0.0718.63 \pm 0.07 22.46±0.0822.46 \pm 0.08
    STAC 9.78±0.539.78 \pm 0.53 13.97±0.3513.97 \pm 0.35 18.25±0.2518.25 \pm 0.25 24.38±0.1224.38 \pm 0.12 28.64±0.2128.64 \pm 0.21
    Humble Teacher - 16.96±0.3816.96 \pm 0.38 21.72±0.2421.72 \pm 0.24 27.70±0.1527.70 \pm 0.15 31.61±0.2831.61 \pm 0.28
    Instant Teaching - 18.05±0.1518.05 \pm 0.15 22.45±0.1522.45 \pm 0.15 26.75±0.0526.75 \pm 0.05 30.40±0.0530.40 \pm 0.05
    Unbiased Teacher 14.36±0.0914.36 \pm 0.09 18.33±0.1918.33 \pm 0.19 22.23±0.2122.23 \pm 0.21 26.65±0.3126.65 \pm 0.31 29.56±0.2429.56 \pm 0.24
    ISMT - 18.88±0.7418.88 \pm 0.74 22.43±0.5622.43 \pm 0.56 26.37±0.2426.37 \pm 0.24 30.53±0.5230.53 \pm 0.52
    Ours (batch 8/8) 17.51±0.24\mathbf{17.51 \pm 0.24} 21.84±0.13\mathbf{21.84 \pm 0.13} 26.14±0.01\mathbf{26.14 \pm 0.01} 30.06±0.14\mathbf{30.06 \pm 0.14} 33.50±0.03\mathbf{33.50 \pm 0.03}
    SoftTeacher* - 20.46±0.3920.46 \pm 0.39 - 30.74±0.0830.74 \pm 0.08 34.04±0.1434.04 \pm 0.14
    Ours* (batch 8/40) 21.02±0.49\mathbf{21.02 \pm 0.49} 24.79±0.30\mathbf{24.79 \pm 0.30} 28.23±0.05\mathbf{28.23 \pm 0.05} 32.05±0.04\mathbf{32.05 \pm 0.04} 35.02±0.02\mathbf{35.02 \pm 0.02}
    Unbiased Teacher†\dagger 16.94±0.2316.94 \pm 0.23 20.75±0.1220.75 \pm 0.12 24.30±0.0724.30 \pm 0.07 28.27±0.1128.27 \pm 0.11 31.50±0.1031.50 \pm 0.10
    Ours†\dagger (batch 32/32) 21.26±0.21\mathbf{21.26 \pm 0.21} 25.40±0.36\mathbf{25.40 \pm 0.36} 28.37±0.03\mathbf{28.37 \pm 0.03} 31.85±0.09\mathbf{31.85 \pm 0.09} 35.08±0.02\mathbf{35.08 \pm 0.02}

    Notation: †\dagger denotes labeled/unlabeled batch size 32/32, ∗* denotes batch size 8/40, and unannotated rows use batch size 8/8. Unbiased Teacher v2 consistently outperforms existing anchor-based SS-OD methods across all label fractions and batch settings.

  9. Knowl 9 — Strict IoU Evaluation and Fine-Grained AP Breakdown for Regression Methods

    data/table

    Average precision (AP) breakdown from AP55AP_{55} to AP95AP_{95} comparing unsupervised regression strategies on the anchor-based detector (Faster-RCNN-ResNet50). All models share the identical classification objectives and differ solely in the unsupervised regression loss:

    Method AP55AP_{55} AP60AP_{60} AP65AP_{65} AP70AP_{70} AP75AP_{75} AP80AP_{80} AP85AP_{85} AP90AP_{90} AP95AP_{95}
    No regression 29.71 27.34 24.64 21.38 17.55 13.27 8.33 3.45 0.35
    Confidence Thresholding 30.60 28.19 25.07 21.93 17.96 13.32 8.22 3.12 0.32
    (vs. No regression) +0.89 +0.85 +0.43 +0.55 +0.41 +0.05 -0.11 -0.33 -0.03
    Listen2Student (Ours) 30.78 28.59 26.19 23.05 19.64 15.61 10.47 5.06 0.58
    (vs. No regression) +1.07 +1.25 +1.56 +1.67 +2.09 +2.34 +2.14 +1.61 +0.23

    Standard confidence thresholding on box scores improves performance at loose IoU thresholds (AP55AP_{55} to AP75AP_{75}) but degrades performance at strict IoU thresholds (AP85AP_{85}, AP90AP_{90}, and AP95AP_{95}) due to misleading regression targets. In contrast, Listen2Student provides positive gains across all IoU thresholds, with the largest improvements occurring at high-precision thresholds (+2.34+2.34 at AP80AP_{80} and +2.14+2.14 at AP85AP_{85}).

  10. Knowl 10 — Semi-Supervised Detection Benchmarks on PASCAL VOC and COCO-Additional

    data/table

    Performance of Faster-RCNN-ResNet50 with Listen2Student on PASCAL VOC (VOC2007-test) and COCO-additional (COCO2017-val):

    PASCAL VOC Benchmark
    Methods Unlabeled Set AP50AP_{50} AP50:95AP_{50:95}
    Supervised (VOC07) None 76.70 43.60
    STAC VOC12 77.45 44.64
    ISMT VOC12 77.23 46.23
    Instant-Teaching VOC12 79.20 50.00
    Humble Teacher VOC12 80.94 53.04
    Unbiased Teacher VOC12 80.51 54.48
    Ours VOC12 81.29 56.87
    STAC VOC12 + COCO20cls 79.08 46.01
    ISMT VOC12 + COCO20cls 77.75 49.59
    Instant-Teaching VOC12 + COCO20cls 79.00 50.80
    Humble Teacher VOC12 + COCO20cls 81.29 54.41
    Unbiased Teacher VOC12 + COCO20cls 81.71 55.79
    Ours VOC12 + COCO20cls 82.04 58.08
    COCO-Additional Benchmark
    Methods (COCO2017-train labeled + COCO2017-unlabeled) mAP
    Supervised 40.90
    CSD 38.52
    STAC 39.21
    Humble Teacher 42.37
    Unbiased Teacher* (with scale jittering) 44.06
    SoftTeacher 44.50
    Ours 44.75

    When adding VOC12 and VOC12+COCO20cls as unlabeled data, the method reaches 56.87 mAP and 58.08 mAP (AP50:95AP_{50:95}), respectively. On COCO-additional, training for 720k iterations without inference threshold tuning reaches 44.75 mAP.

Coverage note — No substantial contributed material was omitted; the extraction covers the core anchor-free modifications, the Listen2Student relative uncertainty formulation and equations, the analysis of centerness bias and center-sampling failure modes, and all key benchmark results (COCO-standard, AP breakdown, VOC, and COCO-additional).

References

  1. 1.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), pages 5049–5059, 2019. 3
  2. 2.Zhaowei Cai, Quanfu Fan, Rogerio S Feris, and Nuno Vasconcelos. A unified multi-scale deep convolutional neural network for fast object detection. In Proceedings of the European Conference on Computer Vision (ECCV), 2016. 2
  3. 3.Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. Learning efficient object detection models with knowledge distillation. In Advances in Neural Information Processing Systems (NeurIPS), 2017. 5
  4. 4.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International Journal of Computer Vision (IJCV), 88(2):303–338, 2010. 6
  5. 5.Jiyang Gao, Jiang Wang, Shengyang Dai, Li-Jia Li, and Ram Nevatia. Note-rcnn: Noise tolerant ensemble rcnn for semi-supervised object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019. 1
  6. 6.Hongyu Guo, Yongyi Mao, and Richong Zhang. Mixup as locally linear out-of-manifold regularization. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 33, pages 3714–3722, 2019. 3
  7. 7.Yihui He, Chenchen Zhu, Jianren Wang, Marios Savvides, and Xiangyu Zhang. Bounding box regression with uncertainty for accurate object detection. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition, pages 2888–2897, 2019. 5
  8. 8.Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: A simple data processing method to improve robustness and uncertainty. Proceedings of the International Conference on Learning Representations (ICLR), 2020. 3
  9. 9.Jisoo Jeong, Seungeui Lee, Jeesoo Kim, and Nojun Kwak. Consistency-based semi-supervised learning for object detection. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 1, 3, 4, 6, 7, 8
  10. 10.Licheng Jiao, Fan Zhang, Fang Liu, Shuyuan Yang, Lingling Li, Zhixi Feng, and Rong Qu. A survey of deep learning-based object detection. IEEE Access, 7:128837–128868, 2019. 2
  11. 11.Tao Kong, Fuchun Sun, Huaping Liu, Yuning Jiang, Lei Li, and Jianbo Shi. Foveabox: Beyound anchor-based object detection. IEEE Transactions on Image Processing, 29:7389–7398, 2020. 1, 2
  12. 12.Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. In Proceedings of the International Conference on Learning Representations (ICLR), 2017. 3
  13. 13.Hei Law and Jia Deng. Cornernet: Detecting objects as paired keypoints. In Proceedings of the European Conference on Computer Vision (ECCV), 2018. 2
  14. 14.Youngwan Lee, Joong-won Hwang, Hyung-Il Kim, Kimin Yun, and Joungyoul Park. Localization uncertainty estimation for anchor-free object detection. arXiv preprint arXiv:2006.15607, 2020. 5
  15. 15.Xiang Li, Wenhai Wang, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss v2: Learning reliable localization quality estimation for dense object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 3, 6
  16. 16.Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 1, 3, 4, 6
  17. 17.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision (CVPR), pages 2980–2988, 2017. 2
  18. 18.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV), 2014. 6
  19. 19.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In European conference on computer vision (ECCV), pages 21–37. Springer, 2016. 1, 2
  20. 20.Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda. Unbiased teacher for semi-supervised object detection. In Proceedings of the International Conference on Learning Representations (ICLR), 2021. 1, 3, 4, 5, 6, 7, 8
  21. 21.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems (NeurIPS), pages 91–99, 2015. 1, 2
  22. 22.Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. Regularization with stochastic transformations and perturbations for deep semi-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), pages 1163–1171, 2016. 3
  23. 23.Muhamad Risqi U Saputra, Pedro PB de Gusmao, Yasin Almalioglu, Andrew Markham, and Niki Trigoni. Distilling knowledge from a deep pose regressor network. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019. 5
  24. 24.Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. Training region-based object detectors with online hard example mining. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2
  25. 25.Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 3, 5
  26. 26.Kihyuk Sohn, Zizhao Zhang, Chun-Liang Li, Han Zhang, Chen-Yu Lee, and Tomas Pfister. A simple semi-supervised learning framework for object detection. arXiv preprint arXiv:2005.04757, 2020. 1, 3, 4, 5, 6, 7, 8
  27. 27.Yihe Tang, Weifeng Chen, Yijun Luo, and Yuting Zhang. Humble teachers teach better students for semi-supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3132–3141, 2021. 3, 6, 7, 8
  28. 28.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in neural information processing systems (NeurIPS), pages 1195–1204, 2017. 3
  29. 29.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019. 1, 2, 3, 4, 6
  30. 30.Jiaqi Wang, Kai Chen, Shuo Yang, Chen Change Loy, and Dahua Lin. Region proposal by guided anchoring. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 1, 2
  31. 31.Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick. Detectron2. https://github.com/facebookresearch/detectron2, 2019. 6
  32. 32.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 2
  33. 33.Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. End-to-end semi-supervised object detection with soft teacher. arXiv preprint arXiv:2106.09018, 2021. 3, 6, 7, 8
  34. 34.Qize Yang, Xihan Wei, Biao Wang, Xian-Sheng Hua, and Lei Zhang. Interactive self-training with mean teachers for semi-supervised object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5941–5950, 2021. 7, 8
  35. 35.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 6023–6032, 2019. 3
  36. 36.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In Proc. International Conference on Learning Representations (ICLR), 2018. 3
  37. 37.Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z Li. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 1, 2, 3, 4, 6
  38. 38.Qiang Zhou, Chaohui Yu, Zhibin Wang, Qi Qian, and Hao Li. Instant-teaching: An end-to-end semi-supervised object detection framework. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 1, 3, 5, 7, 8
  39. 39.Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl. Objects as points. In arXiv preprint arXiv:1904.07850, 2019. 2, 4
  40. 40.Xingyi Zhou, Jiacheng Zhuo, and Philipp Krahenbuhl. Bottom-up object detection by grouping extreme and center points. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 2
  41. 41.Chenchen Zhu, Fangyi Chen, Zhiqiang Shen, and Marios Savvides. Soft anchor-point object detection. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 3, 6
  42. 42.Chenchen Zhu, Yihui He, and Marios Savvides. Feature selective anchor-free module for single-shot object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 1, 2

Citation

MLA
Liu, Y.-C., et al. “Unbiased Teacher V2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors”. arXiv, 2022, http://arxiv.org/abs/2206.09500v1.
APA
Liu, Y.-C., Ma, C.-Y., & Kira, Z. (2022). Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors. arXiv. http://arxiv.org/abs/2206.09500v1
Chicago
Liu, Y.-C., C.-Y. Ma, and Z. Kira. 2022. “Unbiased Teacher V2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors”. arXiv. http://arxiv.org/abs/2206.09500v1.
Harvard
Liu, Y.-C., Ma, C.-Y. and Kira, Z. (2022) “Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.09500v1.
Vancouver
1. Liu Y-C, Ma C-Y, Kira Z (2022) Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors. arXiv

BibTeX

@article{liu2022unbiased,
  title = {Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors},
  author = {Liu, Yen-Cheng and Ma, Chih-Yao and Kira, Zsolt},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.09500v1},
  eprint = {2206.09500}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE