SOOD: Towards Semi-Supervised Oriented Object Detection
Wei HuaDingkang LiangJingyu LiXiaolong LiuZhikang ZouXiaoqing YeXiang Bai
Presents the first semi-supervised oriented object detection framework by introducing rotation-aware adaptive weighting and global layout consistency losses to exploit unlabeled aerial imagery effectively.
Training computer vision models to identify objects in aerial imagery requires extensive bounding-box annotations. Labeling oriented bounding boxes for objects with arbitrary angles—such as vehicles, ships, and infrastructure—costs roughly 36.5% more than standard horizontal boxes. While semi-supervised learning methods use inexpensive unlabeled images to reduce labeling burdens, existing frameworks focus on horizontal objects and struggle with the unique challenges of aerial scenes, such as dense object clusters, small targets, and arbitrary rotations.
The article introduces and evaluates a semi-supervised oriented object detection framework called SOOD. Its primary objective is to demonstrate that incorporating rotation awareness and scene-level spatial layouts into a teacher-student pseudo-labeling framework enables high-accuracy oriented object detection with minimal labeled data.
The approach builds upon a dense pseudo-labeling architecture where a primary model (the student) learns from both labeled data and guidance generated by an ensemble model (the teacher). The researchers introduced two specialized loss mechanisms: an instance-level loss that dynamically weights training supervision based on the angular difference between predictions and pseudo-labels, and a global consistency loss that uses mathematical transport theory to match spatial distributions and confidence scores across the entire image. The framework was evaluated on the benchmark DOTA-v1.5 dataset, which contains over 400,000 annotated aerial instances across 16 categories, tested under 10%, 20%, 30%, and fully labeled data conditions against leading semi-supervised alternatives.
The experimental findings show that SOOD consistently outperforms existing semi-supervised methods across all evaluated data proportions. When trained on only 10% labeled data, SOOD achieved an accuracy score of 48.63 mAP, improving by +5.85 points over the supervised baseline and outperforming the prior leading method by +1.73 points. Under 20% and 30% labeled data settings, it reached 55.58 and 59.23 mAP, outperforming supervised models by +5.47 and +4.44 points, respectively. In fully labeled settings supplemented by unlabeled data, the framework reached 67.70 mAP, improving the baseline by +2.24 points. Furthermore, the framework generalized effectively to other detector architectures, delivering accuracy gains between +1.32 and +2.41 points.
These results indicate that organizations analyzing aerial and satellite imagery can achieve superior detection accuracy while cutting data labeling costs and operational timelines. By softly weighting rotation differences and enforcing overall layout consistency, the system prevents label noise from accumulating during training, overcoming the primary hurdle that caused some previous semi-supervised methods to degrade when given unlabeled data.
Organizations deploying aerial object detection should adopt rotation-aware semi-supervised frameworks to maximize model performance while controlling annotation budgets. Next operational steps should focus on piloting the framework across enterprise-specific aerial datasets and tuning pseudo-label sampling ratios. Further research is recommended to combine the separate rotation and layout modules into a unified architecture and expand the methodology to related domains such as 3D object detection and multi-oriented text recognition.
Confidence in these findings is strong given the rigorous multi-ratio testing on a standard benchmark and validation across multiple underlying detectors. However, stakeholders should note that the current implementation does not yet explicitly optimize for extreme scale variations or high aspect ratios, which represents the primary limitation when applying the framework to highly specialized imagery.
- Paper: DOTA: A Large-Scale Dataset for Object Detection in Aerial Images, Gui-Song Xia et al. (2017). Introduces the DOTA dataset and oriented bounding box evaluation benchmark upon which SOOD builds its aerial object detection framework.
- Paper: Unbiased Teacher v2: Semi-supervised Object Detection for Anchor-free and Anchor-based Detectors, Yen-Cheng Liu et al. (2022). Establishes modern teacher-student pseudo-labeling mechanisms for semi-supervised object detection that SOOD adapts for oriented aerial targets.
- Paper: Shape-Adaptive Selection and Measurement for Oriented Object Detection, Liping Hou et al. (2022). Explores the geometry and loss design challenges unique to arbitrary-oriented bounding box regression in aerial imagery.
- Paper: Spatial Transform Decoupling for Oriented Object Detection, Hongtian Yu et al. (2024). Extends oriented object detection architectures by decoupling position, scale, and angle predictions within Vision Transformers.
