Omni-DETR: Omni-Supervised Object Detection with Transformers
Pei WangZhaowei CaiHao YangGurumurthy SwaminathanNuno VasconcelosBernt SchieleStefano Soatto
Proposes an end-to-end transformer framework that unifies diverse weak annotations—including tags, points, and counts—through bipartite-matching pseudo-label filtering, demonstrating that mixed-supervision strategies can outperform fully annotated datasets under a fixed labeling budget.
Building modern computer vision systems for object detection typically requires extensive datasets where every object has a precise bounding box and a category tag. Generating these complete annotations is extremely slow and expensive—taking an estimated 346 seconds per image on standard benchmarks—which limits the ability of organizations to scale their datasets. While cheaper, weaker labeling formats exist (such as point clicks, object counts, or image-level tags), prior approaches struggled to extract meaningful performance gains from them, leading to the assumption that full annotations are always the most practical investment.
The article evaluates whether incorporating diverse weak annotations can improve object detection accuracy and deliver a better cost-accuracy trade-off than relying entirely on complete annotations. To demonstrate this, the authors introduce a unified system called Omni-DETR that trains detection models across any mixture of fully labeled, weakly labeled, and completely unlabeled data.
The evaluated approach combines a student-teacher training framework with a transformer-based detector (Deformable DETR). The teacher model generates initial candidate detections on weakly augmented images, which are then aligned with the available weak labels using a bipartite matching filter to create reliable synthetic training targets for the student model. The researchers tested this framework across five diverse benchmark datasets (MS-COCO, PASCAL VOC, CrowdHuman, Bees, and Objects365) and simulated various labeling cost budgets based on human annotation timings.
The findings demonstrate that weak annotations consistently provide meaningful improvements across all tested configurations. On standard benchmark subsets, adding weak labels to a semi-supervised baseline increased detection accuracy by 1.7 to 4.4 percentage points, with bounding boxes without tags and extreme-point clicks yielding the highest gains. Furthermore, the unified framework outperformed prior omni-supervised and weakly-supervised approaches by substantial margins, such as improving performance over the previous leading framework by roughly 5 to 10 percentage points. Most importantly, budget-aware experiments showed that mixing weak annotations with full annotations achieves a superior cost-to-accuracy trade-off compared to spending the entire budget on complete annotations alone. For instance, achieving a target accuracy on the dense Bees dataset required about 15 fewer annotation hours (a reduction from 40 hours to 25 hours), while on the CrowdHuman dataset, mixed annotations improved accuracy by approximately 4 percentage points under a fixed 330-hour budget.
These results indicate that organizations can significantly lower data collection costs and accelerate development timelines by adopting mixed annotation pipelines. Instead of uniform full labeling, data strategies can be customized to dataset characteristics: point clicks and counts are highly cost-effective for crowded or dense scenes, whereas bounding boxes without class labels are ideal for datasets with hundreds of difficult-to-differentiate categories. For fixed budgets, allocating a portion to complete labels and the remainder to cheaper weak labels consistently outperforms standard semi-supervised setups.
Decision-makers preparing new computer vision initiatives should evaluate their dataset characteristics and select tailored weak annotation strategies rather than defaulting to complete manual labeling. However, readers should note that the evaluation was limited to datasets containing up to roughly 120,000 images, and extreme-point annotations on certain datasets were simulated rather than collected live. Further validation on larger-scale corporate datasets is recommended before fully overhauling production annotation pipelines.
- Paper: Deformable DETR: Deformable Transformers for End-to-End Object Detection, Xizhou Zhu et al. (2021). Omni-DETR directly adopts Deformable DETR as its core base detector architecture to achieve efficient multi-scale attention and end-to-end set prediction.
- Paper: End-to-End Object Detection with Transformers, Nicolas Carion et al. (2020). Reading the foundational DETR paper provides essential background on the set-based bipartite matching loss that Omni-DETR adapts to align weak annotations with teacher predictions.
- Paper: Sparse R-CNN: End-to-End Object Detection with Learnable Proposals, Peize Sun et al. (2020). This paper establishes iterative proposal refinement and set prediction mechanisms that inform transformer-based candidate generation and matching.
- Paper: Big Self-Supervised Models are Strong Semi-Supervised Learners, Ting Chen et al. (2020). This work introduces the self-training and teacher-student pseudo-label distillation framework upon which Omni-DETR builds its omni-supervised training regime.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). Understanding how semi-supervised learning techniques filter and refine noisy label estimates provides critical insight into Omni-DETR's weak-label matching filter.
- Paper: DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection, Hao Zhang et al. (2023). DINO advances DETR-based architectures using contrastive denoising and anchor query refinement, building upon the transformer detection principles explored in Omni-DETR.
- Paper: OW-DETR: Open-world Detection Transformer, Akshita Gupta et al. (2022). OW-DETR adapts Deformable DETR to open-world object detection and unknown pseudo-labeling, extending the transformer-based supervision paradigms used in Omni-DETR.
- Paper: Few-Shot Object Detection with Foundation Models, Guangxing Han et al. (2024). FM-FSOD pairs Deformable DETR proposals with foundation models to operate in low-data regimes, extending the weak- and limited-supervision focus of Omni-DETR.
- Paper: DETRs Beat YOLOs on Real-time Object Detection, Yian Zhao et al. (2024). RT-DETR optimizes DETR-based architectures for practical, real-time deployment, extending end-to-end transformer detection to computationally constrained scenarios.
- Paper: WildDet3D: Scaling Promptable 3D Detection in the Wild, Weikai Huang et al. (2026). WildDet3D scales promptable detection to 3D spaces using diverse input formats, generalizing Omni-DETR's motivation of learning from flexible and mixed supervision signals.
