Omni-DETR: Omni-Supervised Object Detection with Transformers

Pei WangZhaowei CaiHao YangGurumurthy SwaminathanNuno VasconcelosBernt SchieleStefano Soatto

article2022CVPR55 citations

Proposes an end-to-end transformer framework that unifies diverse weak annotations—including tags, points, and counts—through bipartite-matching pseudo-label filtering, demonstrating that mixed-supervision strategies can outperform fully annotated datasets under a fixed labeling budget.

Listen

Building modern computer vision systems for object detection typically requires extensive datasets where every object has a precise bounding box and a category tag. Generating these complete annotations is extremely slow and expensive—taking an estimated 346 seconds per image on standard benchmarks—which limits the ability of organizations to scale their datasets. While cheaper, weaker labeling formats exist (such as point clicks, object counts, or image-level tags), prior approaches struggled to extract meaningful performance gains from them, leading to the assumption that full annotations are always the most practical investment.

The article evaluates whether incorporating diverse weak annotations can improve object detection accuracy and deliver a better cost-accuracy trade-off than relying entirely on complete annotations. To demonstrate this, the authors introduce a unified system called Omni-DETR that trains detection models across any mixture of fully labeled, weakly labeled, and completely unlabeled data.

The evaluated approach combines a student-teacher training framework with a transformer-based detector (Deformable DETR). The teacher model generates initial candidate detections on weakly augmented images, which are then aligned with the available weak labels using a bipartite matching filter to create reliable synthetic training targets for the student model. The researchers tested this framework across five diverse benchmark datasets (MS-COCO, PASCAL VOC, CrowdHuman, Bees, and Objects365) and simulated various labeling cost budgets based on human annotation timings.

The findings demonstrate that weak annotations consistently provide meaningful improvements across all tested configurations. On standard benchmark subsets, adding weak labels to a semi-supervised baseline increased detection accuracy by 1.7 to 4.4 percentage points, with bounding boxes without tags and extreme-point clicks yielding the highest gains. Furthermore, the unified framework outperformed prior omni-supervised and weakly-supervised approaches by substantial margins, such as improving performance over the previous leading framework by roughly 5 to 10 percentage points. Most importantly, budget-aware experiments showed that mixing weak annotations with full annotations achieves a superior cost-to-accuracy trade-off compared to spending the entire budget on complete annotations alone. For instance, achieving a target accuracy on the dense Bees dataset required about 15 fewer annotation hours (a reduction from 40 hours to 25 hours), while on the CrowdHuman dataset, mixed annotations improved accuracy by approximately 4 percentage points under a fixed 330-hour budget.

These results indicate that organizations can significantly lower data collection costs and accelerate development timelines by adopting mixed annotation pipelines. Instead of uniform full labeling, data strategies can be customized to dataset characteristics: point clicks and counts are highly cost-effective for crowded or dense scenes, whereas bounding boxes without class labels are ideal for datasets with hundreds of difficult-to-differentiate categories. For fixed budgets, allocating a portion to complete labels and the remainder to cheaper weak labels consistently outperforms standard semi-supervised setups.

Decision-makers preparing new computer vision initiatives should evaluate their dataset characteristics and select tailored weak annotation strategies rather than defaulting to complete manual labeling. However, readers should note that the evaluation was limited to datasets containing up to roughly 120,000 images, and extreme-point annotations on certain datasets were simulated rather than collected live. Further validation on larger-scale corporate datasets is recommended before fully overhauling production annotation pipelines.

Cover for Omni-DETR: Omni-Supervised Object Detection with Transformers

Abstract

We consider the problem of omni-supervised object detection, which can use unlabeled, fully labeled and weakly labeled annotations, such as image tags, counts, points, etc., for object detection. This is enabled by a unified architecture, Omni-DETR, based on the recent progress on student-teacher framework and end-to-end transformer based object detection. Under this unified architecture, different types of weak labels can be leveraged to generate accurate pseudo labels, by a bipartite matching based filtering mechanism, for the model to learn. In the experiments, Omni-DETR has achieved state-of-the-art results on multiple datasets and settings. And we have found that weak annotations can help to improve detection performance and a mixture of them can achieve a better trade-off between annotation cost and accuracy than the standard complete annotation. These findings could encourage larger object detection datasets with mixture annotations. The code is available at https://github.com/amazon-research/omni-detr.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Omni-DETR
  • 3.1. Omni-labels
  • 3.2. Unified Framework
  • 3.3. Detection Architecture
  • 3.4. Training
  • 4. Pseudo-label Filtering
  • 4.1. Simple Pseudo-label Filtering
  • 4.2. Unified Pseudo-label Filtering
  • 4.2.1 No Annotation
  • 4.2.2 Weak Annotations
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Evaluation on Single Annotation
  • 5.3. Comparison with the State-of-the-art
  • 5.4. Ablation Study
  • 5.5. Budget-Aware Omni-Supervised Detection
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Omni-DETR Student-Teacher Framework for Omni-Supervised Object Detection

    model/method

    Omni-DETR is an omni-supervised object detection framework designed to jointly train on a fully labeled dataset Dl={(xil,yil)}i=1Nl\mathcal{D}^l = \{(x_i^l, y_i^l)\}_{i=1}^{N_l} and an omni-labeled (weakly labeled or unlabeled) dataset Do={(xio,yio)}i=1No\mathcal{D}^o = \{(x_i^o, y_i^o)\}_{i=1}^{N_o}. In the fully labeled set, each annotation yil={(bi,j,ci,j)}j=1Biy_i^l = \{(b_{i,j}, c_{i,j})\}_{j=1}^{B_i} comprises BiB_i bounding boxes bi,j∈R4b_{i,j} \in \mathbb{R}^4 paired with ground-truth class labels ci,j∈{1,…,C}c_{i,j} \in \{1, \dots, C\}, whereas each omni-label yioy_i^o can take any weaker annotation format (image tags, point coordinates, instance counts, or unlabeled data).

    The framework is built upon an end-to-end Deformable DETR detector organized into a student network Fs(x;θs)\mathcal{F}^s(x; \theta_s) and a teacher network Ft(x;θt)\mathcal{F}^t(x; \theta_t). Training consists of two stages:

    1. Burn-in Stage: The student network Fs\mathcal{F}^s is trained alone using standard supervised learning on the fully labeled set Dl\mathcal{D}^l. The teacher parameters θt\theta_t are then initialized by duplicating the trained student parameters θs\theta_s.

    2. Mutual Student-Teacher Learning Stage: For an omni-labeled image xox^o, weakly augmented (xo,wx^{o,w}) and strongly augmented (xo,sx^{o,s}) views are generated. The weakly augmented image is passed through the teacher to produce hypothesis predictions y^t=Ft(xo,w;θt)={y^cls,y^box}\hat{y}^t = \mathcal{F}^t(x^{o,w}; \theta_t) = \{\hat{y}_{\text{cls}}, \hat{y}_{\text{box}}\}. A unified pseudo-label filter T\mathcal{T} processes y^t\hat{y}^t using the available omni-labels yoy^o to generate filtered pseudo-labels y~t=T(y^t;yo)\tilde{y}^t = \mathcal{T}(\hat{y}^t; y^o), which supervise the student on the strongly augmented view xo,sx^{o,s}. For labeled instances xlx^l, both weakly augmented (xl,w,yl,w)(x^{l,w}, y^{l,w}) and strongly augmented (xl,s,yl,s)(x^{l,s}, y^{l,s}) versions are fed directly to the student.

    The student network parameters θs\theta_s are optimized using stochastic gradient descent on the total loss Ls\mathcal{L}^s:

    Ls=∑iL(xil,s,yil,s)+L(xil,w,yil,w)+∑iL(xio,s,y~it)\mathcal{L}^s = \sum_i \mathcal{L}(x_i^{l,s}, y_i^{l,s}) + \mathcal{L}(x_i^{l,w}, y_i^{l,w}) + \sum_i \mathcal{L}(x_i^{o,s}, \tilde{y}_i^t)

    where L=αLcls+βLbox\mathcal{L} = \alpha \mathcal{L}_{\text{cls}} + \beta \mathcal{L}_{\text{box}}, with Lcls\mathcal{L}_{\text{cls}} denoting classification loss, Lbox\mathcal{L}_{\text{box}} denoting generalized IoU plus L1L_1 bounding box regression loss, and α,β\alpha, \beta being loss weights. The teacher parameters θt\theta_t are updated via exponential moving average (EMA) with momentum k=0.9996k = 0.9996:

    θt←kθt+(1−k)θs\theta_t \leftarrow k \theta_t + (1 - k)\theta_s

  2. Knowl 2 — Omni-Label Annotation Taxonomy

    definition

    Omni-DETR formalizes six distinct weakly labeled or unlabeled annotation formats for an omni-labeled image xx, denoted yy, in addition to standard fully supervised labels:

    • None (Unlabeled): y=∅y = \emptyset, containing no annotation for the image.

    • Tags without Counts (TagsU): y={cj}j=1My = \{c_j\}_{j=1}^M, where cj∈{1,…,C}c_j \in \{1, \dots, C\} is an image-level category tag and MM is the number of distinct categories present in the image.

    • Tags with Counts (TagsK): y={(cj,nj)}j=1My = \{(c_j, n_j)\}_{j=1}^M, where cj∈{1,…,C}c_j \in \{1, \dots, C\} is the category tag and nj∈N≥1n_j \in \mathbb{N}_{\ge 1} is the exact instance count of class cjc_j in the image.

    • Points without Tags (PointsU): y={pj}j=1Py = \{p_j\}_{j=1}^P, where pj∈R2p_j \in \mathbb{R}^2 is a 2D point coordinate located on an object instance (such as its geometric center or a point within its segmentation mask) without category information, and PP is the total count of point annotations.

    • Points with Tags (PointsK): y={(pj,cj)}j=1Py = \{(p_j, c_j)\}_{j=1}^P, where pj∈R2p_j \in \mathbb{R}^2 is the point coordinate of the jj-th object and cj∈{1,…,C}c_j \in \{1, \dots, C\} is its associated category label.

    • Boxes without Tags (BoxesU): y={bj}j=1By = \{b_j\}_{j=1}^B, where bj∈R4b_j \in \mathbb{R}^4 is a high-quality 4-coordinate bounding box without category labels, and BB is the number of annotated boxes.

    • Extreme Clicking Boxes (BoxesEC): y={bj}j=1By = \{b_j\}_{j=1}^B, where each bj∈R4b_j \in \mathbb{R}^4 is a bounding box derived from four extreme clicked points (top, bottom, leftmost, rightmost pixels) of the object, reducing annotation cost by approximately 5×5\times compared to BoxesU with minimal quality degradation.

  3. Knowl 3 — Unified Hungarian Bipartite Matching for Pseudo-Label Filtering

    model/method

    Omni-DETR formulates pseudo-label filtering from weak annotations as a global set-to-set bipartite matching problem between teacher prediction hypotheses and ground-truth weak labels.

    Let the teacher network output KK predictions y^={y^cls,y^box}\hat{y} = \{\hat{y}_{\text{cls}}, \hat{y}_{\text{box}}\}, where logits y^cls=[z1,…,zK]T∈RK×C\hat{y}_{\text{cls}} = [z_1, \dots, z_K]^T \in \mathbb{R}^{K \times C} yield class probabilities pk∈[0,1]Cp_k \in [0, 1]^C across CC categories via softmax, and y^box=[b^1,…,b^K]T∈RK×4\hat{y}_{\text{box}} = [\hat{b}_1, \dots, \hat{b}_K]^T \in \mathbb{R}^{K \times 4} contains predicted bounding box coordinates. Given GG ground-truth weak omni-labels {gi}i=1G\{g_i\}_{i=1}^G with G≤KG \le K, the filtering finds an optimal permutation σ^∈℘K\hat{\sigma} \in \wp_K of the KK prediction indices:

    σ^=arg⁡min⁡σ∈℘K∑i=1GLmatch(gi,y^σ(i))\hat{\sigma} = \arg\min_{\sigma \in \wp_K} \sum_{i=1}^G \mathcal{L}_{\text{match}}(g_i, \hat{y}_{\sigma(i)})

    where Lmatch(gi,y^σ(i))\mathcal{L}_{\text{match}}(g_i, \hat{y}_{\sigma(i)}) is an annotation-specific pairwise matching cost between the ii-th ground-truth omni-label gig_i and the σ(i)\sigma(i)-th teacher prediction hypothesis y^σ(i)\hat{y}_{\sigma(i)}.

    The optimal one-to-one assignment σ^\hat{\sigma} is computed globally via the Hungarian algorithm. The resulting matched pairs serve as pseudo-labels {(bi∗,ci∗)}i=1G\{(b_i^*, c_i^*)\}_{i=1}^G to supervise the student detector, replacing heuristic single-label thresholding rules with a globally optimal assignment.

  4. Knowl 4 — Pseudo-Label Matching Costs for Unlabeled, Tag, and Point Supervisions

    equation

    In the Omni-DETR bipartite matching filter, matching costs and pseudo-label extraction vary according to the weak supervision modality:

    1. Unlabeled Images (None): Matching is bypassed. For query k∈[1,K]k \in [1, K], predicted class is c^k=arg⁡max⁡cpkc\hat{c}_k = \arg\max_c p_k^c and confidence score is sk=pkc^ks_k = p_k^{\hat{c}_k}. Pseudo-labels are selected via confidence thresholding with threshold τ=0.7\tau = 0.7:

    {(b^k,c^k)∣sk>τ,k∈[1,K]}\{(\hat{b}_k, \hat{c}_k) \mid s_k > \tau, k \in [1, K]\}

    1. Tags without Counts (TagsU): For tags yo={cj}j=1My^o = \{c_j\}_{j=1}^M, the instance count njn_j for each class cjc_j is predicted by:

    nj=max⁡(1,∣{k∈[1,K]∣pkcj>τ}∣)n_j = \max\left(1, \left|\left\{k \in [1, K] \mid p_k^{c_j} > \tau\right\}\right|\right)

    The ground-truth set is expanded to G=∑j=1MnjG = \sum_{j=1}^M n_j elements {gi}i=1G={ci}i=1G\{g_i\}_{i=1}^G = \{c_i\}_{i=1}^G with njn_j repetitions for class cjc_j. The matching cost is:

    Lmatcht(gi,y^σ(i))=1−pσ(i)ci\mathcal{L}_{\text{match}}^t(g_i, \hat{y}_{\sigma(i)}) = 1 - p_{\sigma(i)}^{c_i}

    yielding pseudo-labels {(b^σ^(i),ci)}i=1G\{(\hat{b}_{\hat{\sigma}(i)}, c_i)\}_{i=1}^G.

    1. Tags with Counts (TagsK): Given exact counts yo={(cj,nj)}j=1My^o = \{(c_j, n_j)\}_{j=1}^M, ground-truth tags are repeated njn_j times and matched using Lmatcht(gi,y^σ(i))=1−pσ(i)ci\mathcal{L}_{\text{match}}^t(g_i, \hat{y}_{\sigma(i)}) = 1 - p_{\sigma(i)}^{c_i}, yielding {(b^σ^(i),ci)}i=1G\{(\hat{b}_{\hat{\sigma}(i)}, c_i)\}_{i=1}^G.

    2. Points without Tags (PointsU): For point annotations yo={pi}i=1G⊂R2y^o = \{p_i\}_{i=1}^G \subset \mathbb{R}^2, the matching cost is:

    Lmatchp(gi,y^σ(i))=(di,σ(i)+ei,σ(i))⋅ηi,σ(i)\mathcal{L}_{\text{match}}^p(g_i, \hat{y}_{\sigma(i)}) = (d_{i, \sigma(i)} + e_{i, \sigma(i)}) \cdot \eta_{i, \sigma(i)}

    where di,σ(i)d_{i, \sigma(i)} is the L2L_2 distance between the center of predicted box b^σ(i)\hat{b}_{\sigma(i)} and point pip_i, normalized to [0,1][0, 1] across all K×GK \times G pairs via min-max normalization; ei,σ(i)=1−sσ(i)e_{i, \sigma(i)} = 1 - s_{\sigma(i)} with sσ(i)=max⁡cpσ(i)cs_{\sigma(i)} = \max_c p_{\sigma(i)}^c; and ηi,σ(i)=1\eta_{i, \sigma(i)} = 1 if point pip_i is inside predicted box b^σ(i)\hat{b}_{\sigma(i)} and +∞+\infty otherwise. Pseudo-labels are {(b^σ^(i),c^σ^(i))}i=1G\{(\hat{b}_{\hat{\sigma}(i)}, \hat{c}_{\hat{\sigma}(i)})\}_{i=1}^G.

    1. Points with Tags (PointsK): For yo={(pi,ci)}i=1Gy^o = \{(p_i, c_i)\}_{i=1}^G, the overall cost combines tag and point matching costs linearly with trade-off coefficient γ∈[0,1]\gamma \in [0, 1] (optimal at γ=0.5\gamma = 0.5):

    Lmatch(gi,y^σ(i))=γLmatcht(ci,y^σ(i))+(1−γ)Lmatchp(pi,y^σ(i))\mathcal{L}_{\text{match}}(g_i, \hat{y}_{\sigma(i)}) = \gamma \mathcal{L}_{\text{match}}^t(c_i, \hat{y}_{\sigma(i)}) + (1 - \gamma) \mathcal{L}_{\text{match}}^p(p_i, \hat{y}_{\sigma(i)})

    yielding pseudo-labels {(b^σ^(i),ci)}i=1G\{(\hat{b}_{\hat{\sigma}(i)}, c_i)\}_{i=1}^G.

  5. Knowl 5 — Pseudo-Label Matching Cost for Bounding Box Annotations

    equation

    When bounding boxes without category information are provided as omni-labels (yo={bi}i=1G⊂R4y^o = \{b_i\}_{i=1}^G \subset \mathbb{R}^4, covering both high-quality boxes BoxesU and extreme clicking boxes BoxesEC), Omni-DETR employs a spatial geometric matching cost between each ground-truth box gi=big_i = b_i and the σ(i)\sigma(i)-th predicted box b^σ(i)\hat{b}_{\sigma(i)}:

    Lmatchb(gi,y^σ(i))=λiouLiou(gi,b^σ(i))+λL1∥gi−b^σ(i)∥1\mathcal{L}_{\text{match}}^b(g_i, \hat{y}_{\sigma(i)}) = \lambda_{\text{iou}} \mathcal{L}_{\text{iou}}(g_i, \hat{b}_{\sigma(i)}) + \lambda_{L1} \|g_i - \hat{b}_{\sigma(i)}\|_1

    where Liou\mathcal{L}_{\text{iou}} is the Generalized Intersection over Union (GIoU) loss, ∥⋅∥1\|\cdot\|_1 is the L1L_1 coordinate distance loss, and λiou,λL1\lambda_{\text{iou}}, \lambda_{L1} are weighting coefficients.

    Following global assignment via Hungarian matching, the generated pseudo-labels are {(bi,c^σ^(i))}i=1G\{(b_i, \hat{c}_{\hat{\sigma}(i)})\}_{i=1}^G, where bib_i is the ground-truth box coordinate and c^σ^(i)=arg⁡max⁡cpσ^(i)c\hat{c}_{\hat{\sigma}(i)} = \arg\max_c p_{\hat{\sigma}(i)}^c is the predicted category from the matched teacher query σ^(i)\hat{\sigma}(i).

  6. Knowl 6 — Empirical Impact of Individual Weak Annotation Types on COCO

    data/table

    The individual impact of each weak annotation format in Omni-DETR was evaluated on the COCO-standard-10% split, using a ResNet-50 backbone initialized with ImageNet pretraining. The model was trained with 10% fully labeled data and 90% omni-labeled data under each annotation format. Performance is measured on COCO val2017 using mAP\text{mAP} (AP50:95\text{AP}_{50:95}), AP50\text{AP}_{50}, and AP75\text{AP}_{75}.

    Annotation Setting mAP AP50\text{AP}_{50} AP75\text{AP}_{75}
    10% supervision (baseline) 28.0 44.3 29.5
    + 90% None 32.4 49.3 34.5
    + 90% TagsU 34.7 52.4 37.2
    + 90% TagsK 35.2 53.5 37.7
    + 90% PointsU 34.1 51.9 36.2
    + 90% PointsK 35.7 54.2 38.6
    + 90% BoxesEC 36.4 54.6 39.3
    + 90% BoxesU 36.8 54.8 39.4

    The results indicate that adding unlabeled data (+90% None) improves the supervised baseline by +4.4% mAP. Introducing weak annotations provides an additional +1.7% to +4.4% mAP gain over the unlabeled semi-supervised baseline. Points without tags (PointsU) yield the smallest addition (+1.7% over None), while unclassified bounding boxes (BoxesU) yield the largest (+4.4% over None). Extreme Clicking boxes (BoxesEC) achieve 36.4% mAP, performing only 0.4% mAP below full-quality BoxesU while requiring five times less human annotation time. Providing count information with tags (+90% TagsK vs. +90% TagsU) yields a +0.5% mAP improvement, and adding class labels to points (+90% PointsK vs. +90% PointsU) yields a +1.6% mAP improvement.

  7. Knowl 7 — State-of-the-Art Comparisons in SSOD, WSSOD, and OSOD

    empirical result

    Omni-DETR establishes state-of-the-art performance across semi-supervised (SSOD), weakly semi-supervised (WSSOD), and omni-supervised (OSOD) object detection benchmarks:

    • Semi-Supervised Detection (SSOD): On COCO-standard splits, Omni-DETR achieves 18.6%, 23.2%, 30.2%, and 34.1% mAP on 1%, 2%, 5%, and 10% labeled data, outperforming Unbiased Teacher (28.3% and 31.5% at 5% and 10% labels). On VOC-07to12, Omni-DETR reaches 53.4% mAP, surpassing Humble Teacher (53.0%) and Unbiased Teacher (48.7%).

    • WSSOD with Tags: On COCO-35to80 (35K fully labeled, 80K omni-labeled images), Omni-DETR + TagsU achieves 39.4% mAP (+5.1% gain over its supervised Deformable DETR baseline of 34.3%), outperforming UFO2^2 (29.4% mAP, +0.3% gain over Faster R-CNN baseline). On COCO-standard with Tags, Omni-DETR achieves 20.1%, 31.7%, 35.9%, and 38.1% mAP on 1%, 5%, 10%, and 20% labeled splits, outperforming the method of Fang et al. (18.4%, 27.4%, 31.3%, 35.0% mAP). Omni-DETR trained on 5% labeled data (31.7% mAP) exceeds Fang et al. trained on 10% labeled data (31.3% mAP).

    • WSSOD with Points: On COCO-35to80 with PointsK, Omni-DETR achieves 40.2% mAP (+5.9% gain over baseline), compared to UFO2^2 at 30.1% mAP (+1.0% gain). On COCO-standard with PointsK, Omni-DETR reaches 32.5%, 37.1%, 39.0%, and 40.1% mAP across 5%, 10%, 20%, and 30% labeled splits, outperforming Point DETR (26.2%, 30.4%, 33.3%, and 34.8% mAP) by 5.3% to 6.7% mAP.

    • OSOD Fixed Budget Regimes: Under fixed annotation budgets on 10K COCO images with X%∈{80%,50%,20%}X\% \in \{80\%, 50\%, 20\%\} of budget spent on full annotations and the remainder on PointsK, Omni-DETR achieves 21.5%, 19.5%, and 9.1% mAP, outperforming UFO2^2 (14.1%, 11.1%, and 4.5% mAP).

  8. Knowl 8 — Ablations on Pseudo-Label Filtering, Matching Threshold, Weighting, and Pseudo Boxes

    empirical result

    Ablation experiments on the COCO-standard-10% setting demonstrate the role of key algorithmic components in Omni-DETR:

    • Unified Hungarian Matching vs. Heuristic Filtering: Global bipartite matching consistently outperforms local heuristic thresholding across all weak supervision types: TagsU achieves 34.7% mAP (unified) vs. 33.3% mAP (simple); TagsK achieves 35.2% vs. 33.8%; PointsU achieves 34.1% vs. 32.4%; PointsK achieves 35.7% vs. 34.6%.

    • Confidence Threshold τ\tau: In threshold sweeps over τ∈{0.5,0.6,0.7,0.8,0.9}\tau \in \{0.5, 0.6, 0.7, 0.8, 0.9\}, τ=0.7\tau = 0.7 achieves the highest performance for both unlabeled data (28.9%, 31.5%, 32.4%, 31.4%, 29.9% mAP) and TagsU (31.1%, 34.1%, 34.7%, 33.9%, 33.1% mAP), balancing pseudo-label recall and precision.

    • Point Matching Trade-off Weight γ\gamma: For PointsK matching cost Lmatch=γLmatcht+(1−γ)Lmatchp\mathcal{L}_{\text{match}} = \gamma \mathcal{L}_{\text{match}}^t + (1 - \gamma)\mathcal{L}_{\text{match}}^p, evaluating γ∈{0.00,0.25,0.50,0.75,1.00}\gamma \in \{0.00, 0.25, 0.50, 0.75, 1.00\} yields 34.1%, 35.3%, 35.7%, 35.5%, and 35.2% mAP, identifying γ=0.5\gamma = 0.5 as optimal.

    • Pseudo Bounding Box Supervision: Unlike prior semi-supervised methods that discard teacher bounding box predictions for unlabeled data due to uncertainty in classification confidence reflecting localization quality, using teacher pseudo bounding boxes to supervise student bounding box regression in Omni-DETR provides a consistent improvement of 0.5% to 1.0% mAP.

  9. Knowl 9 — Budget-Aware Annotation Costs and Accuracy Trade-Offs

    data/table

    Estimated human labeling costs (in seconds per image) across five benchmark datasets reflect substantial variance depending on category count, scene clutter, and object scale:

    Dataset TagsU TagsK PointsU PointsK BoxesEC BoxesU Fully
    Bees - 6.1 6.4 6.4 50.0 249.9 249.9
    CrowdHuman - 19.4 20.4 20.4 158.5 792.4 792.4
    VOC 20.0 21.0 2.2 22.9 16.8 84.0 102.6
    COCO 80.0 84.2 6.9 88.7 53.9 269.5 346.0
    Objects365 365.0 375.8 14.2 381.7 110.6 553.0 913.0

    Empirical evaluation of budget-constrained annotation policies on these datasets yields the following findings:

    • Mixture Superiority over Full Supervision: Mixed weak/full annotation policies (OSOD) consistently outperform the standard semi-supervised detection baseline (allocating the entire budget to full annotations and leaving remaining data unlabeled). On the Bees dataset at 40% mAP, mixing TagsK and PointsU reduces annotation cost from ~40 hours (fully labeled SSOD) to ~25 hours. On CrowdHuman at a budget of ~330 hours, OSOD improves mAP by approximately ~4% over the SSOD baseline.

    • Spatial vs. Category Annotation Efficiency: Annotating spatial locations (PointsU and BoxesEC) is universally more cost-effective than standard full annotation. Count annotations (TagsK) provide large gains in dense, crowded scenes with few classes (Bees, CrowdHuman).

    • Inefficiency of Tags for Large Vocabularies: In datasets with large category spaces (VOC with 20 classes, COCO with 80 classes, Objects365 with 365 classes), image tags (TagsU, TagsK, PointsK) are inefficient because verifying the presence or absence of hundreds of classes per image is prohibitively costly relative to the detection signal provided.

  10. Knowl 10 — Dataset Scale Scope and Dual-Use Considerations

    limitation

    The methodology and empirical conclusions of Omni-DETR have two primary limitations:

    1. Data Scale Generalization: All experimental evaluations were performed on datasets of up to approximately 120K120\text{K} images (the scale of MS-COCO). Whether the identified optimal annotation mixtures, bipartite matching behavior, and relative performance gains hold at ultra-large web scales (e.g., millions of images) has not been empirically verified.

    2. Societal and Dual-Use Impact: By substantially reducing the human annotation budget required to train competitive object detectors, the framework lowers barriers to creating customized detection models, which may facilitate unintended or unvetted deployment in surveillance and automated monitoring applications.

Coverage note — None was omitted; all core contributions, theoretical formulations, pseudo-label matching algorithms, experimental comparisons, ablations, budget analyses, and stated limitations are covered.

References

  1. 1.Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson. There are many consistent explanations of unlabeled data: Why you should average. In ICLR, 2019. 4
  2. 2.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What's the point: Semantic segmentation with point supervision. In ECCV, pages 549–565. Springer, 2016. 8
  3. 3.Bees. https://lila.science/datasets/boxes-on-bees-and-pollen. 2, 6
  4. 4.Hakan Bilen and Andrea Vedaldi. Weakly supervised deep detection networks. In CVPR, pages 2846–2854, 2016. 2, 3
  5. 5.Zhaowei Cai, Avinash Ravichandran, Subhransu Maji, Charless Fowlkes, Zhuowen Tu, and Stefano Soatto. Exponential moving average normalization for self-supervised and semi-supervised learning. In CVPR, pages 194–203, 2021. 4
  6. 6.Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In CVPR, pages 6154–6162, 2018. 2
  7. 7.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, pages 213–229. Springer, 2020. 2, 4, 5, 6
  8. 8.Akshay L Chandra, Sai Vikas Desai, Vineeth N Balasubramanian, Seishi Ninomiya, and Wei Guo. Active learning with point supervision for cost-effective panicle detection in cereal crops. Plant Methods, 16(1):1–16, 2020. 2
  9. 9.Liangyu Chen, Tong Yang, Xiangyu Zhang, Wei Zhang, and Jian Sun. Points as queries: Weakly semi-supervised object detection by points. In CVPR, pages 8823–8832, 2021. 2, 3, 6, 7
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. Ieee, 2009. 6
  11. 11.Piotr Dollár, Christian Wojek, Bernt Schiele, and Pietro Perona. Pedestrian detection: A benchmark. In CVPR, pages 304–311. IEEE, 2009. 1, 3
  12. 12.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88(2):303–338, 2010. 1, 3, 6
  13. 13.Shijie Fang, Yuhang Cao, Xinjiang Wang, Kai Chen, Dahua Lin, and Wayne Zhang. Wssod: A new pipeline for weakly-and semi-supervised object detection. arXiv preprint arXiv:2105.11293, 2021. 2, 3, 4, 6, 7
  14. 14.Jiyang Gao, Jiang Wang, Shengyang Dai, Li-Jia Li, and Ram Nevatia. Note-rcnn: Noise tolerant ensemble rcnn for semi-supervised object detection. In ICCV, pages 9508–9517, 2019. 2
  15. 15.Ross Girshick. Fast r-cnn. In ICCV, pages 1440–1448, 2015. 2
  16. 16.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, pages 580–587, 2014. 2
  17. 17.Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid. Multi-fold mil training for weakly supervised object localization. In CVPR, pages 2409–2416, 2014. 2
  18. 18.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In CVPR, pages 5356–5364, 2019. 3
  19. 19.Michael Gygli and Vittorio Ferrari. Efficient object annotation via speaking and pointing. IJCV, pages 1–15, 2019. 2, 3
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 6
  21. 21.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry P. Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. In UAI, pages 876–885. AUAI Press, 2018. 4
  22. 22.Jisoo Jeong, Seungeui Lee, Jeesoo Kim, and Nojun Kwak. Consistency-based semi-supervised learning for object detection. NeurIPS, 32:10759–10768, 2019. 2
  23. 23.Zequn Jie, Yunchao Wei, Xiaojie Jin, Jiashi Feng, and Wei Liu. Deep self-taught learning for weakly supervised object localization. In CVPR, pages 1377–1385, 2017. 1, 2, 3, 4
  24. 24.Harold W Kuhn. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955. 4, 5
  25. 25.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. arXiv preprint arXiv:1811.00982, 2018. 1, 3
  26. 26.Hengduo Li and et al. Rethinking pseudo labels for semi-supervised object detection. arXiv:2106.00168, 2021. 7
  27. 27.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, pages 2117–2125, 2017. 2, 4
  28. 28.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014. 1, 3, 6
  29. 29.Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. Ssd: Single shot multibox detector. In ECCV, pages 21–37. Springer, 2016. 2, 4
  30. 30.Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda. Unbiased teacher for semi-supervised object detection. arXiv preprint arXiv:2102.09480, 2021. 2, 3, 4, 5, 6, 7
  31. 31.Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari. Extreme clicking for efficient object annotation. In ICCV, pages 4930–4939, 2017. 3, 6, 8
  32. 32.Dim P Papadopoulos, Jasper RR Uijlings, Frank Keller, and Vittorio Ferrari. Training object class detectors with click supervision. In CVPR, pages 6374–6383, 2017. 1, 2, 3, 4
  33. 33.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In CVPR, pages 779–788, 2016. 2, 4
  34. 34.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 28:91–99, 2015. 2, 3, 4, 7
  35. 35.Zhongzheng Ren, Zhiding Yu, Xiaodong Yang, Ming-Yu Liu, Alexander G Schwing, and Jan Kautz. Ufo2: A unified framework towards omni-supervised object detection. In ECCV, pages 288–313. Springer, 2020. 1, 3, 5, 6, 7, 8
  36. 36.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR, pages 658–666, 2019. 6
  37. 37.Chuck Rosenberg, Martial Hebert, and Henry Schneiderman. Semi-supervised self-training of object detection models. 2005. 2
  38. 38.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In CVPR, pages 8430–8439, 2019. 1, 2, 3, 6
  39. 39.Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123, 2018. 2, 6
  40. 40.Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. In NeurIPS, 2020. 3
  41. 41.Kihyuk Sohn, Zizhao Zhang, Chun-Liang Li, Han Zhang, Chen-Yu Lee, and Tomas Pfister. A simple semi-supervised learning framework for object detection. arXiv preprint arXiv:2005.04757, 2020. 2, 7
  42. 42.Hyun Oh Song, Yong Jae Lee, Stefanie Jegelka, and Trevor Darrell. Weakly-supervised discovery of visual pattern configurations. arXiv preprint arXiv:1406.6507, 2014. 2
  43. 43.Hao Su, Jia Deng, and Li Fei-Fei. Crowdsourcing annotations for visual object detection. In AAAI workshop, 2012. 3, 8
  44. 44.Peng Tang, Xinggang Wang, Song Bai, Wei Shen, Xiang Bai, Wenyu Liu, and Alan Yuille. Pcl: Proposal cluster learning for weakly supervised object detection. TPAMI, 42(1):176–191, 2018. 1, 2, 3, 4
  45. 45.Yihe Tang, Weifeng Chen, Yijun Luo, and Yuting Zhang. Humble teachers teach better students for semi-supervised object detection. In CVPR, pages 3132–3141, June 2021. 2, 3, 7
  46. 46.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780, 2017. 2, 4
  47. 47.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In ICCV, pages 9627–9636, 2019. 2
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017. 2, 4
  49. 49.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In CVPR, pages 10687–10698, 2020. 2
  50. 50.Qize Yang, Xihan Wei, Biao Wang, Xian-Sheng Hua, and Lei Zhang. Interactive self-training with mean teachers for semi-supervised object detection. In CVPR, pages 5941–5950, 2021. 2
  51. 51.Qiang Zhou, Chaohui Yu, Zhibin Wang, Qi Qian, and Hao Li. Instant-teaching: An end-to-end semi-supervised object detection framework. In CVPR, pages 4081–4090, 2021. 2
  52. 52.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020. 2, 4, 6

Citation

MLA
Wang, P., et al. “Omni-DETR: Omni-Supervised Object Detection with Transformers”. arXiv, 2022, http://arxiv.org/abs/2203.16089v1.
APA
Wang, P., Cai, Z., Yang, H., Swaminathan, G., Vasconcelos, N., Schiele, B., & Soatto, S. (2022). Omni-DETR: Omni-Supervised Object Detection with Transformers. arXiv. http://arxiv.org/abs/2203.16089v1
Chicago
Wang, P., Z. Cai, H. Yang, et al. 2022. “Omni-DETR: Omni-Supervised Object Detection with Transformers”. arXiv. http://arxiv.org/abs/2203.16089v1.
Harvard
Wang, P. et al. (2022) “Omni-DETR: Omni-Supervised Object Detection with Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.16089v1.
Vancouver
1. Wang P, Cai Z, Yang H, Swaminathan G, Vasconcelos N, Schiele B, Soatto S (2022) Omni-DETR: Omni-Supervised Object Detection with Transformers. arXiv

BibTeX

@article{wang2022omni,
  title = {Omni-DETR: Omni-Supervised Object Detection with Transformers},
  author = {Wang, Pei and Cai, Zhaowei and Yang, Hao and Swaminathan, Gurumurthy and Vasconcelos, Nuno and Schiele, Bernt and Soatto, Stefano},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.16089v1},
  eprint = {2203.16089}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE