Generalized UAV Object Detection via Frequency Domain Disentanglement

Kunyu WangXueyang FuYukun HuangChengzhi CaoGege ShiZheng-Jun Zha

article2023CVPR67 citations

Proposes a frequency-domain disentanglement framework using learnable spectral filters and instance-level contrastive learning to separate domain-invariant features from domain-specific variations, significantly improving drone-based object detection across unseen target environments.

Listen

Deploying unmanned aerial vehicles equipped with vision systems into unpredictable operational settings often leads to severe performance degradation. This decline occurs because models trained under standard conditions, such as clear daylight, struggle to generalize when exposed to unseen shifts in scene structure, lighting, or adverse weather like fog and nighttime darkness. Existing domain generalization methods typically rely on spatial convolutions that process local image regions, which fail to handle the drastic global appearance changes caused by aerial camera movement.

The article demonstrates that transforming aerial images into the frequency domain allows models to better separate universal object features from environmental distortions. By isolating amplitude spectrum signals via Fourier transform, the authors evaluate whether distinct frequency bands contribute differently to cross-domain detection performance.

To address this challenge, the authors design a framework utilizing two learnable filters to automatically extract domain-invariant components (which assist generalization) and domain-specific components (which hinder it). The model incorporates an instance-level contrastive loss to pull together representations of identical object classes across varied conditions while pushing apart background and domain-specific artifacts. Training is structured through an alternating optimization scheme between filter disentanglement and the primary object detection tasks. Credibility is established using benchmarks on two major aerial datasets (UAVDT and VisDrone2019-VID) evaluated across three unseen test environments: varying scene structures, diverse illumination conditions, and foggy weather.

Key experimental findings confirm the framework's effectiveness. First, preliminary tests show that eliminating specific frequency bands significantly impacts generalization, with statistical analysis revealing that middle- and high-frequency spectrums contain the vast majority of domain-invariant visual cues. Second, in single-dataset evaluations, the proposed method achieved average overall detection gains over baseline systems by 14.2% on AP50 and 9.0% on standard Average Precision, outperforming multiple state-of-the-art generalization models. Third, cross-dataset transfer tests verified robust generalizability, surpassing the baseline by 6.9% on AP50 and 5.8% on Average Precision across unseen scenarios. Finally, ablation studies confirm that frequency-domain disentanglement delivers substantially better detection accuracy than standard spatial-domain convolutions.

These findings suggest that frequency-based feature separation significantly lowers operational and safety risks when deploying aerial drones in dynamic or unmapped environments without requiring target-domain data during training. Organizations deploying vision-equipped aerial systems should consider integrating frequency-aware filtering into their detection pipelines to enhance model reliability under variable real-world conditions. Because this research represents an initial baseline for frequency-domain disentanglement in aerial detection, future initiatives should focus on refining adaptive filter architectures and validating performance across broader operational edge cases before wide-scale deployment.

  • Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). This comprehensive survey outlines foundational paradigms of domain generalization, providing the theoretical and conceptual background necessary for addressing out-of-distribution shifts without target domain access.
  • Paper: Domain Separation Networks, Konstantinos Bousmalis et al. (2016). This work establishes the core concept of feature disentanglement into domain-invariant and domain-specific subspaces, which the source extends into the frequency domain.
  • Paper: Domain Adaptive Faster R-CNN for Object Detection in the Wild, Yuhua Chen et al. (2018). It introduces multi-level (image and instance) domain adaptation for object detection architectures, informing the source's use of instance-level mechanisms for cross-domain detection.
  • Paper: Deeper, Broader and Artier Domain Generalization, Da Li et al. (2017). This paper establishes standard formulation and multi-source visual benchmark protocols for domain generalization.
  • Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). It provides essential taxonomy and methodology on learning invariant visual representations to generalize across unseen domains.
Cover for Generalized UAV Object Detection via Frequency Domain Disentanglement

Abstract

When deploying the Unmanned Aerial Vehicles object detection (UAV-OD) network to complex and unseen real-world scenarios, the generalization ability is usually reduced due to the domain shift. To address this issue, this paper proposes a novel frequency domain disentanglement method to improve the UAV-OD generalization. Specifically, we first verified that the spectrum of different bands in the image has different effects to the UAV-OD generalization. Based on this conclusion, we design two learnable filters to extract domain-invariant spectrum and domain-specific spectrum, respectively. The former can be used to train the UAV-OD network and improve its capacity for generalization. In addition, we design a new instance-level contrastive loss to guide the network training. This loss enables the network to concentrate on extracting domain-invariant spectrum and domain-specific spectrum, so as to achieve better disentangling results. Experimental results on three unseen target domains demonstrate that our method has better generalization ability than both the base-line method and state-of-the-art methods.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. Problem Definition
  • 3.2. Frequency-based Learnable Filtering
  • 3.3. Contrastive-based Frequency Disentanglement
  • 3.4. Training With the Alternating Optimization
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Domain Generalization Results
  • 4.4. Ablation Analysis
  • 4.5. Statistic Analysis on Learnable Filters
  • 4.6. Visualization Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Frequency-domain framework for single-source UAV object detection generalization

    model/method

    The paper addresses single-domain domain generalization for UAV object detection: a detector is trained only on a labeled source domain and must operate on multiple unseen target domains without access to target data. The proposed framework transforms each source image into amplitude and phase spectra, applies two learned amplitude masks, and reconstructs two image components: a domain-invariant component intended to support generalization and a domain-specific component intended to capture source-domain characteristics.

    The detector is based on YOLOv5. Its backbone is divided into an early part B1B_1 and a later part B2B_2. Both reconstructed components are passed through B1B_1 to obtain corresponding features. Object regions from the two feature maps are aligned with RoI-Align and projected by a two-hidden-layer MLP for an instance-level contrastive loss. The domain-invariant branch is passed through B2B_2 and the detection head HH for the bounding-box regression and classification losses, while the domain-specific branch is used to guide disentanglement.

  2. Knowl 2 — Learnable extraction of invariant and specific frequency components

    equation

    For a source image x∈RH×W×Cx\in\mathbb{R}^{H\times W\times C} with height HH, width WW, and CC channels, the two-dimensional Fourier transform is computed independently for each channel:

    F(x)(u,v)=∑h=0H−1∑w=0W−1x(h,w)exp⁡[−j2π(hHu+wWv)],\mathcal{F}(x)(u,v)=\sum_{h=0}^{H-1}\sum_{w=0}^{W-1}x(h,w)\exp\left[-j2\pi\left(\frac{h}{H}u+\frac{w}{W}v\right)\right],

    where h,wh,w are spatial indices, u,vu,v are frequency indices, and jj is the imaginary unit. If R(x)\mathcal{R}(x) and I(x)\mathcal{I}(x) are the real and imaginary parts of the Fourier signal, its amplitude and phase are

    A(x)(u,v)=[R2(x)(u,v)+I2(x)(u,v)]1/2,P(x)(u,v)=arctan⁡[I(x)(u,v)R(x)(u,v)].\mathcal{A}(x)(u,v)=\left[\mathcal{R}^{2}(x)(u,v)+\mathcal{I}^{2}(x)(u,v)\right]^{1/2},\qquad \mathcal{P}(x)(u,v)=\arctan\left[\frac{\mathcal{I}(x)(u,v)}{\mathcal{R}(x)(u,v)}\right].

    Two learnable filters ψsi,ψss∈RH×W×C\psi_{si},\psi_{ss}\in\mathbb{R}^{H\times W\times C}, whose elements are constrained to [0,1][0,1], extract the domain-invariant and domain-specific amplitudes by element-wise multiplication:

    xsiA=xsA⊗ψsi,xssA=xsA⊗ψss,x_{si}^{\mathcal{A}}=x_s^{\mathcal{A}}\otimes\psi_{si},\qquad x_{ss}^{\mathcal{A}}=x_s^{\mathcal{A}}\otimes\psi_{ss},

    where xsAx_s^{\mathcal{A}} is the source-image amplitude spectrum and ⊗\otimes denotes element-wise multiplication. Each filtered amplitude is combined with the original phase spectrum xsPx_s^{\mathcal{P}} and transformed by the inverse Fourier transform to produce the domain-invariant image component xsix_{si} and domain-specific image component xssx_{ss}. The invariant component is used for detection, whereas both components provide paired instance features for disentanglement.

  3. Knowl 3 — Instance-level contrastive frequency disentanglement

    equation

    For a source image containing nn labeled objects, the early backbone part B1B_1 produces feature maps from the invariant and specific components. RoI-Align converts the variable-sized regions into fixed-size tensors o^ij,o^sj∈Rs×s×c\hat{o}_{i_j},\hat{o}_{s_j}\in\mathbb{R}^{s\times s\times c}, where j∈{1,…,n}j\in\{1,\ldots,n\}, ss is the aligned spatial size, and cc is the feature-channel count. A projection head PP maps them to invariant embeddings Zi={zi1,…,zin}Z_i=\{z_{i_1},\ldots,z_{i_n}\} and specific embeddings Zs={zs1,…,zsn}Z_s=\{z_{s_1},\ldots,z_{s_n}\}. For embeddings u,vu,v, similarity is cosine similarity, sim⁡(u,v)=u⊤v/(∥u∥∥v∥)\operatorname{sim}(u,v)=u^{\top}v/(\|u\|\|v\|).

    Let Z^i=Zi∖{zij}\hat Z_i=Z_i\setminus\{z_{i_j}\} and Z^s=Zs∖{zsj}\hat Z_s=Z_s\setminus\{z_{s_j}\} for the current anchors, and define Za=Z^i∪ZsZ_a=\hat Z_i\cup Z_s and Zb=Zi∪Z^sZ_b=Z_i\cup\hat Z_s. With category labels yijy_{i_j} and ysjy_{s_j} inherited from the corresponding object regions, the contrastive loss is

    Lcon=∑zij∈Zi−1∣Z^i∣∑zik∈Z^iyij=yiklog⁡exp⁡(sim⁡(zij,zik)/τ)∑za∈Zaexp⁡(sim⁡(zij,za))+∑zsj∈Zs−1∣Z^s∣∑zsk∈Z^sysj=ysklog⁡exp⁡(sim⁡(zsj,zsk)/τ)∑zb∈Zbexp⁡(sim⁡(zsj,zb)).\begin{aligned} \mathcal{L}_{\mathrm{con}}={}&\sum_{z_{i_j}\in Z_i}\frac{-1}{|\hat Z_i|}\sum_{\substack{z_{i_k}\in\hat Z_i\\y_{i_j}=y_{i_k}}} \log\frac{\exp(\operatorname{sim}(z_{i_j},z_{i_k})/\tau)}{\sum_{z_a\in Z_a}\exp(\operatorname{sim}(z_{i_j},z_a))}\\ &+\sum_{z_{s_j}\in Z_s}\frac{-1}{|\hat Z_s|}\sum_{\substack{z_{s_k}\in\hat Z_s\\y_{s_j}=y_{s_k}}} \log\frac{\exp(\operatorname{sim}(z_{s_j},z_{s_k})/\tau)}{\sum_{z_b\in Z_b}\exp(\operatorname{sim}(z_{s_j},z_b))}. \end{aligned}

    Here τ\tau is the temperature. Same-category instances within the invariant group are positives for an invariant anchor, while all embeddings in the contrastive denominator that are not selected positives act as negatives; the same rule is applied to the specific group. The loss therefore pulls together same-category invariant instances and same-category specific instances while separating the corresponding other instances, providing gradients that train the two frequency filters.

  4. Knowl 4 — Alternating optimization of disentanglement and detection

    algorithm

    The trainable parameters are separated into a disentanglement group θ={ψsi,ψss,P}\theta=\{\psi_{si},\psi_{ss},P\} containing the two frequency filters and projection head, and a detection group η={B1,B2,H}\eta=\{B_1,B_2,H\} containing the detector backbone and head. The detection objective is

    Ldet=Lreg+Lcls,\mathcal{L}_{\mathrm{det}}=\mathcal{L}_{\mathrm{reg}}+\mathcal{L}_{\mathrm{cls}},

    where Lreg\mathcal{L}_{\mathrm{reg}} is the bounding-box regression loss and Lcls\mathcal{L}_{\mathrm{cls}} is the object-classification loss. At alternation tt, the method performs:

    Input: labeled source images and object annotations; parameters theta and eta; contrastive-loss weight lambda
    Output: trained frequency filters and UAV object detector
    Repeat for each alternation t during training:
        Fix eta at eta^(t-1)
        Update theta by minimizing lambda times L_con(theta, eta^(t-1))
        Fix theta at theta^t
        Update eta by minimizing L_det(theta^t, eta)
    Return theta and eta

    The alternating strategy prevents the detection objective from directly conflicting with the frequency-disentanglement objective. In the reported experiments, the balancing coefficient is λ=0.15\lambda=0.15.

  5. Knowl 5 — Different frequency bands have different effects on UAV detection generalization

    data/table

    A preliminary experiment trained the detector after rejecting selected amplitude-spectrum bands from source images and then evaluated it on three unseen UAVDT target conditions: different scene structures, nighttime illumination, and foggy weather. The row labeled Null retains the full spectrum. For the other rows, α\alpha and β\beta specify the normalized frequency-band thresholds used by the band-reject filter. The results show that spectral bands are not equally useful: removing bands can sharply reduce performance in one target condition while improving it in another, demonstrating the need for learned rather than fixed frequency selection.

    Could not parse LaTeX table

    AP50AP_{50}, AP75AP_{75}, and APAP are average precision at the indicated intersection-over-union thresholds and the aggregate average-precision metric, respectively. The Average columns average performance over the three unseen target domains.

  6. Knowl 6 — Datasets and training protocol for evaluating unseen-domain generalization

    experimental setup

    The principal experiment uses UAVDT, which contains approximately 41k frames and 840k bounding boxes originally divided into car, truck, and bus categories. Because truck and bus each account for fewer than 5% of the boxes, the experiments merge all three categories into one class. The daylight portion contains 23,741 images; 2,850 daylight images with different scene structures form one unseen target domain, and the remaining daylight images are used for source-domain training. The other unseen domains are 11,489 nighttime images for diverse illumination and 2,492 foggy images for adverse weather.

    For cross-dataset evaluation, the detector is trained on 16,238 daylight images from VisDrone2019-VID, which contains drone imagery from different places and heights and has ten predefined object classes, and is tested on the three UAVDT target conditions. The implementation uses PyTorch and eight NVIDIA 1080Ti GPUs, trains for 300 epochs with batch size 128, and uses Adam for the detector with learning rate 0.0010.001, momentum 0.90.9, five-epoch linear warmup, and lambda learning-rate decay. The projection head uses SGD with learning rate 0.050.05 and weight decay 10−410^{-4}; the frequency filters use SGD with learning rate 0.0010.001 and weight decay 10−410^{-4}. Both use five-epoch warmup and step decay. The contrastive temperature is τ=0.7\tau=0.7, and the loss-balancing coefficient is λ=0.15\lambda=0.15. Evaluation reports APAP, AP50AP_{50}, and AP75AP_{75}.

  7. Knowl 7 — Frequency disentanglement substantially improves UAVDT domain generalization

    data/table

    All methods in this comparison are trained on daylight UAVDT images and evaluated without target-domain training on different-scene daylight images, nighttime images, and foggy images. The proposed method is compared with a plain YOLOv5 baseline and several domain-generalization approaches adapted to object detection. It achieves the best value on 11 of the 12 reported metrics. Its average performance across the three unseen domains is 54.054.0 for AP50AP_{50}, 28.428.4 for AP75AP_{75}, and 29.429.4 for APAP, compared with 39.839.8, 18.618.6, and 20.420.4 for the baseline. The paper reports improvements over the baseline of 14.2%, 9.8%, and 9.0% on these three average metrics, respectively, and improvements over the runner-up of 4.5%, 1.8%, and 2.7%.

    Could not parse LaTeX table
  8. Knowl 8 — Cross-dataset generalization transfers the benefit beyond UAVDT training

    data/table

    To test whether the method generalizes across datasets rather than only across conditions within UAVDT, each detector is trained on VisDrone2019-VID daylight images and evaluated on the three unseen UAVDT target conditions. The proposed method obtains the best average result on all three aggregate metrics: 47.647.6 AP50AP_{50}, 21.621.6 AP75AP_{75}, and 24.424.4 APAP, compared with 40.740.7, 14.314.3, and 18.618.6 for the baseline. The paper reports average improvements over the baseline of 6.9%, 7.3%, and 5.8%, and improvements over the runner-up of 1.3%, 4.8%, and 2.6%, for AP50AP_{50}, AP75AP_{75}, and APAP, respectively.

    Could not parse LaTeX table
  9. Knowl 9 — Ablations establish the value of frequency filtering and alternating training

    data/table

    Replacing the two learned frequency filters with two spatial convolution blocks, each consisting of three convolution layers, while keeping the contrastive loss and training procedure fixed reduces average UAVDT performance. The frequency version reaches 54.054.0 AP50AP_{50}, 28.428.4 AP75AP_{75}, and 29.429.4 APAP, whereas the spatial version reaches 50.950.9, 24.624.6, and 26.326.3. This supports the paper's claim that frequency-domain filtering is more effective than a comparable spatial disentanglement mechanism for the tested UAV setting.

    Could not parse LaTeX table

    The paper also compares alternating optimization with joint optimization of all parameters using both detection and contrastive losses. Alternating optimization performs better on the average metrics: 54.0/28.4/29.454.0/28.4/29.4 versus 53.3/27.2/28.753.3/27.2/28.7 for joint training on AP50/AP75/APAP_{50}/AP_{75}/AP.

    Could not parse LaTeX table

    A sweep of the loss-balancing coefficient found that average generalization performance increases as λ\lambda rises from small values and peaks at λ=0.15\lambda=0.15. A sweep of the backbone split found the best performance when the split is made at block 4, determining the reported division between B1B_1 and B2B_2.

  10. Knowl 10 — Learned filters emphasize middle and high frequencies and remove domain-specific backgrounds

    empirical result

    The learned invariant filter ψsi\psi_{si} and specific filter ψss\psi_{ss} were divided into low-, middle-, and high-frequency regions, and the average filter value in each region was measured. The invariant-filter weights were 0.46450.4645, 0.58370.5837, and 0.58130.5813 for low-, middle-, and high-frequency bands, respectively; the corresponding specific-filter weights were 0.54950.5495, 0.37920.3792, and 0.38050.3805. Thus, the middle- and high-frequency regions receive relatively more weight in the invariant filter than in the specific filter, leading the paper to conclude that they contain more domain-invariant information than the low-frequency region.

    Qualitative reconstructions support this interpretation. Domain-invariant components from images with different scene structures and illumination conditions appear visually similar despite differences in the original images. Compared with Single-DGOD, the proposed invariant features more thoroughly suppress irrelevant background responses, including an advertising board in one scene and an area illuminated by a car lamp in another.

Coverage note — The paper's brief limitation statement—that this is an initial exploration and that more subtle frequency-disentanglement designs remain possible—was omitted as a separate knowl because it specifies no concrete failure case or tested limitation; the main qualitative visualization findings are included.

References

  1. 1.Qi Cai, Yingwei Pan, Chong-Wah Ngo, Xinmei Tian, Lingyu Duan, and Ting Yao. Exploring object relation in mean teacher for cross-domain detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11457–11466, 2019. 3
  2. 2.Fabio M Carlucci, Antonio D’Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi. Domain generalization by solving jigsaw puzzles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2229–2238, 2019. 6
  3. 3.Chaoqi Chen, Zebiao Zheng, Xinghao Ding, Yue Huang, and Qi Dou. Harmonizing transferability and discriminability for adapting object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8869–8878, 2020. 1, 3
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 5
  5. 5.Yuhua Chen, Wen Li, Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Domain adaptive faster r-cnn for object detection in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3339–3348, 2018. 1, 3
  6. 6.Dawei Du, Yuankai Qi, Hongyang Yu, Yifan Yang, Kaiwen Duan, Guorong Li, Weigang Zhang, Qingming Huang, and Qi Tian. The unmanned aerial vehicle benchmark: Object detection and tracking. In Proceedings of the European conference on computer vision (ECCV), pages 370–386, 2018. 5, 6
  7. 7.Milan Erdelj and Enrico Natalizio. Uav-assisted disaster management: Applications and open issues. In 2016 international conference on computing, networking and communications (ICNC), pages 1–5. IEEE, 2016. 1
  8. 8.Dayan Guan, Jiaxing Huang, Aoran Xiao, Shijian Lu, and Yanpeng Cao. Uncertainty-aware unsupervised domain adaptation in object detection. IEEE Transactions on Multimedia, 24:2502–2514, 2021. 1, 3
  9. 9.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020. 5
  10. 10.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 4
  11. 11.Eija Honkavaara, Heikki Saari, Jere Kaivosoja, Ilkka Polönen, Teemu Hakala, Paula Litkey, Jussi Mäkynen, and Liisa Pesonen. Processing and assessment of spectrometric, stereoscopic imagery collected using a lightweight uav spectral camera for precision agriculture. Remote Sensing, 5(10):5006–5039, 2013. 1
  12. 12.Jiaxing Huang, Dayan Guan, Aoran Xiao, and Shijian Lu. Fsdr: Frequency space domain randomization for domain generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6891–6902, 2021. 3
  13. 13.Zeyi Huang, Haohan Wang, Eric P Xing, and Dong Huang. Self-challenging improves cross-domain generalization. In European Conference on Computer Vision, pages 124–140. Springer, 2020. 6
  14. 14.G. Jocher, K. Nishimura, T. Mineeva, and R. Vilarino. Yolov5. https://github.com/ultralytics/yolov5, 2020. Accessed: 2020-07-10. 4
  15. 15.Benjamin Kiefer, Martin Messmer, and Andreas Zell. Diminishing domain bias by leveraging domain labels in object detection on uavs. In 2021 20th International Conference on Advanced Robotics (ICAR), pages 523–530. IEEE, 2021. 3
  16. 16.Seunghyeon Kim, Jaehoon Choi, Taekyung Kim, and Changick Kim. Self-training and adversarial background regularization for unsupervised domain adaptive one-stage object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6092–6101, 2019. 1
  17. 17.Taekyung Kim, Minki Jeong, Seunghyeon Kim, Seokeon Choi, and Changick Kim. Diversify and match: A domain adaptive representation learning paradigm for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12456–12465, 2019. 1
  18. 18.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  19. 19.Chuang Lin, Zehuan Yuan, Sicheng Zhao, Peize Sun, Changhu Wang, and Jianfei Cai. Domain-invariant disentangled network for generalizable object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8771–8780, 2021. 1, 2, 3
  20. 20.Hong Liu, Pinhao Song, and Runwei Ding. Towards domain generalization in underwater object detection. In 2020 IEEE International Conference on Image Processing (ICIP), pages 1971–1975. IEEE, 2020. 1, 3
  21. 21.Quande Liu, Cheng Chen, Jing Qin, Qi Dou, and Pheng-Ann Heng. Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous frequency space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1013–1023, 2021. 3
  22. 22.Sasanka Madawalagama, Niluka Munasinghe, SDPJ Dampegama, and L Samarakoon. Low cost aerial mapping with consumer-grade drones. In 37th Asian Conference on Remote Sensing, pages 1–8, 2016. 1
  23. 23.Payal Mittal, Raman Singh, and Akashdeep Sharma. Deep learning-based object detection in low-altitude uav datasets: A survey. Image and Vision Computing, 104:104046, 2020. 1
  24. 24.Henri J Nussbaumer. The fast fourier transform. In Fast Fourier Transform and Convolution Algorithms, pages 80–111. Springer, 1981. 2
  25. 25.Sebastian Ruder. An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747, 2016. 6
  26. 26.Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, and Kate Saenko. Strong-weak distribution alignment for adaptive object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6956–6965, 2019. 1, 3
  27. 27.Karthik Seemakurthy, Charles Fox, Erchan Aptoula, and Petra Bosilj. Domain generalisation for object detection. arXiv preprint arXiv:2203.05294, 2022. 1, 3
  28. 28.Eduard Semsch, Michal Jakob, Dusan Pavlicek, and Michal Pechoucek. Autonomous uav surveillance in complex urban environments. In 2009 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology, volume 2, pages 82–85. IEEE, 2009. 1
  29. 29.Jingye Wang, Ruoyi Du, Dongliang Chang, Kongming Liang, and Zhanyu Ma. Domain generalization via frequency-domain-based feature disentanglement and interaction. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4821–4829, 2022. 3
  30. 30.Aming Wu and Cheng Deng. Single-domain generalized object detection in urban scene via cyclic-disentangled self-distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 847–856, 2022. 1, 2, 3, 6, 8
  31. 31.Xin Wu, Wei Li, Danfeng Hong, Ran Tao, and Qian Du. Deep learning for unmanned aerial vehicle-based object detection and tracking: a survey. IEEE Geoscience and Remote Sensing Magazine, 10(1):91–124, 2021. 1
  32. 32.Zhenyu Wu, Karthik Suresh, Priya Narayanan, Hongyu Xu, Heesung Kwon, and Zhangyang Wang. Delving into robust object detection from unmanned aerial vehicles: A deep nuisance disentanglement approach. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1201–1210, 2019. 3
  33. 33.Chang-Dong Xu, Xing-Ran Zhao, Xin Jin, and Xiu-Shen Wei. Exploring categorical regularization for domain adaptive object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11724–11733, 2020. 3
  34. 34.Minghao Xu, Hang Wang, Bingbing Ni, Qi Tian, and Wenjun Zhang. Cross-domain detection via graph-induced prototype alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12355–12364, 2020. 3
  35. 35.Qinwei Xu, Ruipeng Zhang, Ya Zhang, Yanfeng Wang, and Qi Tian. A fourier-based framework for domain generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14383–14392, 2021. 3
  36. 36.Yanchao Yang and Stefano Soatto. Fda: Fourier domain adaptation for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4085–4095, 2020. 3
  37. 37.Xingxu Yao, Sicheng Zhao, Pengfei Xu, and Jufeng Yang. Multi-source domain adaptation for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3273–3282, 2021. 1, 3
  38. 38.Hongyang Yu, Guorong Li, Weigang Zhang, Qingming Huang, Dawei Du, Qi Tian, and Nicu Sebe. The unmanned aerial vehicle benchmark: Object detection, tracking and baseline. International Journal of Computer Vision, 128(5):1141–1159, 2020. 1
  39. 39.Xingxuan Zhang, Peng Cui, Renzhe Xu, Linjun Zhou, Yue He, and Zheyan Shen. Deep stable learning for out-of-distribution generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5372–5382, 2021. 6
  40. 40.Xingxuan Zhang, Zekai Xu, Renzhe Xu, Jiashuo Liu, Peng Cui, Weitao Wan, Chong Sun, and Chen Li. Towards domain generalization in object detection. arXiv preprint arXiv:2203.14387, 2022. 1, 2, 3
  41. 41.Ganlong Zhao, Guanbin Li, Ruijia Xu, and Liang Lin. Collaborative training between region proposal localization and classification for domain adaptive object detection. In European Conference on Computer Vision, pages 86–102. Springer, 2020. 3
  42. 42.Zhen Zhao, Yuhong Guo, Haifeng Shen, and Jieping Ye. Adaptive object detection with dual multi-label prediction. In European Conference on Computer Vision, pages 54–69. Springer, 2020. 3
  43. 43.Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy. Domain generalization: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. 2
  44. 44.Pengfei Zhu, Dawei Du, Longyin Wen, Xiao Bian, Haibin Ling, Qinghua Hu, Tao Peng, Jiayu Zheng, Xinyao Wang, Yue Zhang, et al. Visdrone-vid2019: The vision meets drone object detection in video challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 5, 6
  45. 45.Chenfan Zhuang, Xintong Han, Weilin Huang, and Matthew Scott. ifan: Image-instance full alignment networks for adaptive object detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 13122–13129, 2020. 3

Citation

MLA
Wang, K., et al. “Generalized UAV Object Detection via Frequency Domain Disentanglement”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 1064–73, https://doi.org/10.1109/CVPR52729.2023.00109.
APA
Wang, K., Fu, X., Huang, Y., Cao, C., Shi, G., & Zha, Z.-J. (2023). Generalized UAV Object Detection via Frequency Domain Disentanglement. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1064–1073. https://doi.org/10.1109/CVPR52729.2023.00109
Chicago
Wang, K., X. Fu, Y. Huang, C. Cao, G. Shi, and Z.-J. Zha. 2023. “Generalized UAV Object Detection via Frequency Domain Disentanglement”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1064–73. https://doi.org/10.1109/CVPR52729.2023.00109.
Harvard
Wang, K. et al. (2023) “Generalized UAV Object Detection via Frequency Domain Disentanglement”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 1064–1073. Available at: https://doi.org/10.1109/CVPR52729.2023.00109.
Vancouver
1. Wang K, Fu X, Huang Y, Cao C, Shi G, Zha Z-J (2023) Generalized UAV Object Detection via Frequency Domain Disentanglement. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 1064–1073

BibTeX

@inproceedings{Wang_2023, title={Generalized UAV Object Detection via Frequency Domain Disentanglement}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00109}, DOI={10.1109/cvpr52729.2023.00109}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wang, Kunyu and Fu, Xueyang and Huang, Yukun and Cao, Chengzhi and Shi, Gege and Zha, Zheng-Jun}, year={2023}, month=June, pages={1064–1073} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE