Generalized UAV Object Detection via Frequency Domain Disentanglement
Kunyu WangXueyang FuYukun HuangChengzhi CaoGege ShiZheng-Jun Zha
Proposes a frequency-domain disentanglement framework using learnable spectral filters and instance-level contrastive learning to separate domain-invariant features from domain-specific variations, significantly improving drone-based object detection across unseen target environments.
Deploying unmanned aerial vehicles equipped with vision systems into unpredictable operational settings often leads to severe performance degradation. This decline occurs because models trained under standard conditions, such as clear daylight, struggle to generalize when exposed to unseen shifts in scene structure, lighting, or adverse weather like fog and nighttime darkness. Existing domain generalization methods typically rely on spatial convolutions that process local image regions, which fail to handle the drastic global appearance changes caused by aerial camera movement.
The article demonstrates that transforming aerial images into the frequency domain allows models to better separate universal object features from environmental distortions. By isolating amplitude spectrum signals via Fourier transform, the authors evaluate whether distinct frequency bands contribute differently to cross-domain detection performance.
To address this challenge, the authors design a framework utilizing two learnable filters to automatically extract domain-invariant components (which assist generalization) and domain-specific components (which hinder it). The model incorporates an instance-level contrastive loss to pull together representations of identical object classes across varied conditions while pushing apart background and domain-specific artifacts. Training is structured through an alternating optimization scheme between filter disentanglement and the primary object detection tasks. Credibility is established using benchmarks on two major aerial datasets (UAVDT and VisDrone2019-VID) evaluated across three unseen test environments: varying scene structures, diverse illumination conditions, and foggy weather.
Key experimental findings confirm the framework's effectiveness. First, preliminary tests show that eliminating specific frequency bands significantly impacts generalization, with statistical analysis revealing that middle- and high-frequency spectrums contain the vast majority of domain-invariant visual cues. Second, in single-dataset evaluations, the proposed method achieved average overall detection gains over baseline systems by 14.2% on AP50 and 9.0% on standard Average Precision, outperforming multiple state-of-the-art generalization models. Third, cross-dataset transfer tests verified robust generalizability, surpassing the baseline by 6.9% on AP50 and 5.8% on Average Precision across unseen scenarios. Finally, ablation studies confirm that frequency-domain disentanglement delivers substantially better detection accuracy than standard spatial-domain convolutions.
These findings suggest that frequency-based feature separation significantly lowers operational and safety risks when deploying aerial drones in dynamic or unmapped environments without requiring target-domain data during training. Organizations deploying vision-equipped aerial systems should consider integrating frequency-aware filtering into their detection pipelines to enhance model reliability under variable real-world conditions. Because this research represents an initial baseline for frequency-domain disentanglement in aerial detection, future initiatives should focus on refining adaptive filter architectures and validating performance across broader operational edge cases before wide-scale deployment.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). This comprehensive survey outlines foundational paradigms of domain generalization, providing the theoretical and conceptual background necessary for addressing out-of-distribution shifts without target domain access.
- Paper: Domain Separation Networks, Konstantinos Bousmalis et al. (2016). This work establishes the core concept of feature disentanglement into domain-invariant and domain-specific subspaces, which the source extends into the frequency domain.
- Paper: Domain Adaptive Faster R-CNN for Object Detection in the Wild, Yuhua Chen et al. (2018). It introduces multi-level (image and instance) domain adaptation for object detection architectures, informing the source's use of instance-level mechanisms for cross-domain detection.
- Paper: Deeper, Broader and Artier Domain Generalization, Da Li et al. (2017). This paper establishes standard formulation and multi-source visual benchmark protocols for domain generalization.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). It provides essential taxonomy and methodology on learning invariant visual representations to generalize across unseen domains.
- Paper: Decompose, Adjust, Compose: Effective Normalization by Playing with Frequency for Domain Generalization, Sangrok Lee et al. (2023). This work builds directly upon frequency-based domain generalization principles by analyzing phase and amplitude interactions to resolve content distortion during normalization.
- Paper: Frequency-Adaptive Dilated Convolution for Semantic Segmentation, Linwei Chen et al. (2024). It extends frequency-domain manipulation principles into architectural convolution operators by dynamically adapting receptive fields across distinct frequency bands.
- Paper: DI-V2X: Learning Domain-Invariant Representation for Vehicle-Infrastructure Collaborative 3D Object Detection, Xiang Li et al. (2024). It explores domain-invariant representation learning in vehicle object detection under cross-sensor and infrastructure domain shifts.
