SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection

Wuyang LiXinyu LiuYixuan Yuan

article2022CVPR230 citationsBest Paper Finalist

Proposes a domain adaptive object detection framework that reformulates cross-domain alignment as a graph matching problem while generating hallucinated nodes to complete missing batch semantics, outperforming conventional prototype-based methods.

Listen

Deploying machine vision models in real-world applications often leads to severe performance degradation when the deployment conditions differ from the training environment. In critical tasks such as autonomous driving under adverse weather or robotic visual surveillance across different camera setups, models trained on clean data struggle with distribution shifts. While standard adaptation techniques attempt to align broad category averages across domains, they frequently overlook within-class diversity and suffer when certain object classes are absent from training mini-batches.

The main objective of the article is to demonstrate an unsupervised domain adaptation framework called SIGMA (Semantic-complete Graph Matching) that bridges the domain gap for object detection by completing missing category semantics and aligning structured feature graphs.

The evaluated approach introduces a graph-embedded semantic completion module that synthesizes missing categories using cross-domain statistics and a graph-guided memory bank. It then converts visual features into cross-image graphs and reformulates domain adaptation as a bipartite graph matching optimization problem. Rather than relying on simple category averages, the framework solves for fine-grained node-to-node correspondences using structural graph constraints. The authors validated this methodology across standard public benchmarks representing weather shifts (Cityscapes to Foggy Cityscapes), synthetic-to-real transitions (Sim10k to Cityscapes), and cross-camera sensor variations (KITTI to Cityscapes).

The key findings demonstrate that SIGMA consistently surpasses existing domain adaptation baselines. On the weather adaptation benchmark, SIGMA achieved detection accuracies of 43.5% and 44.2% mean Average Precision across two standard neural backbones, outperforming competing methods by up to 4.9 percentage points. In synthetic-to-real vehicle detection, it attained 53.7% accuracy, showing an adaptation gain of 13.9 percentage points over unadapted models. Similarly, in cross-camera transfer, the model led the benchmarks with 45.8% accuracy. Ablation analyses confirmed that both the generation of missing categories and structure-aware graph matching contributed significantly to these performance improvements.

These findings indicate that addressing within-class variability and batch-level semantic mismatches provides a more reliable pathway for deploying computer vision systems in unconstrained environments. For operational teams, this reduces the risk of missed detections and false alarms caused by environmental shifts, lowering safety risks in automated driving without requiring costly manual annotation in every target scenario.

Organizations developing autonomous systems should adopt structured graph matching and semantic completion strategies when adapting vision models to novel operating conditions. Technical teams should tune sampling capacities carefully, as findings indicate optimal adaptation occurs around moderate graph node sizes (approximately 100 to 200 nodes per feature map), whereas excessive sampling degrades optimization efficiency.

Confidence in the reported improvements is high across the evaluated driving benchmarks. However, the study focuses primarily on traffic scenes with predefined category overlaps. Stakeholders should conduct pilot evaluations when deploying the framework in non-vehicular domains or environments where target domain label distributions diverge substantially from the source domain.

  • Paper: Domain Adaptive Faster R-CNN for Object Detection in the Wild, Yuhua Chen et al. (2018). This foundational paper established the standard framework and benchmark protocols for unsupervised domain adaptive object detection via adversarial alignment on Faster R-CNN.
  • Paper: Conditional Adversarial Domain Adaptation, Mingsheng Long et al. (2017). It provides crucial theoretical and practical foundations for class-conditional domain alignment, directly motivating SIGMA's approach to resolving mismatched class semantics across domains.
  • Paper: Transfer Feature Learning with Joint Distribution Adaptation, Mingsheng Long et al. (2013). It formalizes the principles of joint distribution adaptation by matching marginal and class-conditional distributions, a core concept that SIGMA reformulates through graph matching.
  • Book: Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al. (2016). It introduces domain-adversarial training and gradient reversal layers, forming the essential baseline paradigm for feature alignment upon which modern DAOD methods build.
Cover for SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection

Abstract

Domain Adaptive Object Detection (DAOD) leverages a labeled domain to learn an object detector generalizing to a novel domain free of annotations. Recent advances align class-conditional distributions by narrowing down cross-domain prototypes (class centers). Though great success, they ignore the significant within-class variance and the domain-mismatched semantics within the training batch, leading to a sub-optimal adaptation. To overcome these challenges, we propose a novel Semantic-complete Graph Matching (SIGMA) framework for DAOD, which completes mismatched semantics and reformulates the adaptation with graph matching. Specifically, we design a Graph-embedded Semantic Completion module (GSC) that completes mismatched semantics through generating hallucination graph nodes in missing categories. Then, we establish cross-image graphs to model class-conditional distributions and learn a graph-guided memory bank for better semantic completion in turn. After representing the source and target data as graphs, we reformulate the adaptation as a graph matching problem, i.e., finding well-matched node pairs across graphs to reduce the domain gap, which is solved with a novel Bipartite Graph Matching adaptor (BGM). In a nutshell, we utilize graph nodes to establish semantic-aware node affinity and leverage graph edges as quadratic constraints in a structure-aware matching loss, achieving fine-grained adaptation with a node-to-node graph matching. Extensive experiments verify that SIGMA outperforms existing works significantly. Our code is available at https://github.com/CityU-AIM-Group/SIGMA.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Domain Adaptive Object Detection
  • 2.2. Graph Matching
  • 3. Motivation and Preliminaries
  • 4. Proposed Method
  • 4.1. Graph-embedded Semantic Completion
  • 4.2. Bipartite Graph Matching
  • 4.3. Model Optimization
  • 5. Experiments
  • 5.1. Datasets and Evaluation
  • 5.2. Implementation Details
  • 5.3. Comparison with State-of-the-arts
  • 5.4. Ablation Studies
  • 5.5. Sensitivity Analysis
  • 5.6. Qualitative Results
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — SIGMA semantic-complete graph-matching framework

    model/method

    SIGMA is an unsupervised domain-adaptive object detection framework that adapts a detector from an annotated source batch S={(xsi,ysi)}i=1BS=\{(x_s^i,y_s^i)\}_{i=1}^{B} to an unlabeled target batch T={xti}i=1BT=\{x_t^i\}_{i=1}^{B}, where BB is the batch size and ysiy_s^i contains source bounding-box labels. A shared feature extractor ϕ\phi converts both domains into visual feature maps. SIGMA then performs four main operations: (1) it samples fine-grained source and target feature nodes and projects them into a graphical space; (2) its Graph-embedded Semantic Completion module generates hallucinated nodes for object categories missing from either batch; (3) it builds cross-image graphs whose nodes represent class-conditional distributions and whose edges represent relationships among instances; and (4) its Bipartite Graph Matching adaptor aligns source and target graphs through semantic node affinities and graph-structure constraints. The full training objective combines detection, node classification, graph matching, node-level adversarial alignment, and global feature alignment losses.

  2. Knowl 2 — Domain-guided hallucination of missing categories

    model/method

    SIGMA completes domain-mismatched batch semantics by generating hallucinated graph nodes for categories that appear in only one domain. Let ΩsB\Omega_s^B and ΩtB\Omega_t^B be the object-category sets present in the source and target batches, respectively, and define Ωsmiss=ΩtB∖ΩsB\Omega_s^{\mathrm{miss}}=\Omega_t^B\setminus\Omega_s^B and Ωtmiss=ΩsB∖ΩtB\Omega_t^{\mathrm{miss}}=\Omega_s^B\setminus\Omega_t^B. Source feature maps use spatially uniform sampling inside ground-truth boxes for foreground nodes and sample a fraction 1/(C+1)1/(C+1) of outside-box pixels for background nodes, where CC is the number of object categories. Target feature maps use classification score maps Mt∈RC×W×HM_t\in\mathbb{R}^{C\times W\times H}: pixels with max⁡cMt(c,p)>τfg\max_c M_t(c,p)>\tau_{fg} are foreground candidates and a fraction 1/(C+1)1/(C+1) of pixels with max⁡cMt(c,p)<τbg\max_c M_t(c,p)<\tau_{bg} are background candidates. The implementation uses τfg=0.5\tau_{fg}=0.5 and τbg=0.05\tau_{bg}=0.05. A nonlinear projection maps sampled visual features into raw node embeddings.

    For a source-missing category ω∈Ωsmiss\omega\in\Omega_s^{\mathrm{miss}}, SIGMA computes the coordinate-wise scale vector σt(ω)∈RD\sigma_t^{(\omega)}\in\mathbb{R}^{D} from target nodes of category ω\omega and retrieves a source memory seed Ss(ω)∈RDS_s^{(\omega)}\in\mathbb{R}^{D} as the category-specific mean estimate. It samples a latent hallucination vector from a diagonal Gaussian and projects it into the graph space:

    xsh∼N ⁣(Ss(ω),diag⁡((σt(ω))2)),vsh=P(xsh),x_s^h\sim\mathcal{N}\!\left(S_s^{(\omega)},\operatorname{diag}\left((\sigma_t^{(\omega)})^2\right)\right),\qquad v_s^h=P(x_s^h),

    where DD is the node-embedding dimension and PP is a learnable linear projection. The target domain is completed symmetrically, using source statistics to generate target nodes for categories in Ωtmiss\Omega_t^{\mathrm{miss}}. Existing and hallucinated nodes together form the semantic-complete node sets used to construct the two graphs.

  3. Knowl 3 — Cross-image graph construction and graph-guided memory bank

    model/method

    For each domain, SIGMA constructs a cross-image graph G={V,E}G=\{V,E\} from the semantic-complete node matrix V∈RN×DV\in\mathbb{R}^{N\times D}, where NN is the number of sampled or hallucinated nodes and DD is the node dimension. A learnable projection WeW_e produces edge logits, which are row-normalized and randomly sparsified by edge dropping:

    A=EdgeDrop⁡ ⁣(softmax⁡[(VWe)(VWe)T]),A=\operatorname{EdgeDrop}\!\left(\operatorname{softmax}\left[(VW_e)(VW_e)^T\right]\right),

    where A∈RN×NA\in\mathbb{R}^{N\times N} is the adjacency matrix. A single graph-convolution layer propagates cross-image semantic information to each node:

    v~i=LN⁡ ⁣(vi+1∣Ni∣∑j∈NiAijvjWgcn),\widetilde v_i=\operatorname{LN}\!\left(v_i+\frac{1}{|\mathcal{N}_i|}\sum_{j\in\mathcal{N}_i}A_{ij}v_jW_{gcn}\right),

    where Ni\mathcal{N}_i is the neighborhood of node ii, WgcnW_{gcn} is learnable, and LN⁡\operatorname{LN} is layer normalization.

    SIGMA stores one graph-derived memory seed S(ω)∈RDS^{(\omega)}\in\mathbb{R}^{D} per category and domain. For every category present in a batch, spectral clustering is applied to the memory seed together with the enhanced nodes of that category. The seed-included cluster is retained, while the other cluster is discarded. If bb is the mean of the retained graph nodes excluding the seed, the memory update is

    S(ω)←sim⁡(b,S(ω))S(ω)+[1−sim⁡(b,S(ω))]b,S^{(\omega)}\leftarrow \operatorname{sim}(b,S^{(\omega)})S^{(\omega)}+\left[1-\operatorname{sim}(b,S^{(\omega)})\right]b,

    with adaptive momentum

    sim⁡(b,S(ω))=bTS(ω)∥b∥2 ∥S(ω)∥2.\operatorname{sim}(b,S^{(\omega)})=\frac{b^TS^{(\omega)}}{\|b\|_2\,\|S^{(\omega)}\|_2}.

    Only existing graph nodes, not hallucinated nodes, update the memory bank; this prevents noisy or prior-dependent hallucinations from corrupting category memory.

  4. Knowl 4 — Graph matching formulation of domain adaptation

    equation

    SIGMA represents the source and target batches as graphs GsG_s and GtG_t with adjacency matrices As∈RNs×NsA_s\in\mathbb{R}^{N_s\times N_s} and At∈RNt×NtA_t\in\mathbb{R}^{N_t\times N_t}, where NsN_s and NtN_t are the numbers of source and target nodes. Domain adaptation is formulated as a relaxed partial graph-matching problem. Let U∈RNs×NtU\in\mathbb{R}^{N_s\times N_t} be a unary node-affinity matrix and let Π∈[0,1]Ns×Nt\Pi\in[0,1]^{N_s\times N_t} be the continuous assignment matrix, with Πij\Pi_{ij} indicating the assignment strength between source node ii and target node jj. The optimization objective is

    min⁡Π  ∥As−ΠAtΠT∥F2−⟨U,Π⟩,\min_{\Pi}\;\|A_s-\Pi A_t\Pi^T\|_F^2-\langle U,\Pi\rangle,

    subject to

    Π1Nt≤1Ns,ΠT1Ns≤1Nt,\Pi\mathbf{1}_{N_t}\leq\mathbf{1}_{N_s},\qquad \Pi^T\mathbf{1}_{N_s}\leq\mathbf{1}_{N_t},

    where ∥⋅∥F\|\cdot\|_F is the Frobenius norm, ⟨U,Π⟩=∑i=1Ns∑j=1NtUijΠij\langle U,\Pi\rangle=\sum_{i=1}^{N_s}\sum_{j=1}^{N_t}U_{ij}\Pi_{ij}, and 1n\mathbf{1}_n is an all-ones vector of length nn. The first term matches graph structures, while the second favors semantically compatible node pairs. The inequalities allow the two graphs to have different numbers of nodes and relax one-hot permutation matching so that the objective can be optimized within a neural network.

  5. Knowl 5 — Cross-graph interaction and semantic-aware node affinity

    model/method

    The Bipartite Graph Matching adaptor first exchanges information between the source and target enhanced node matrices V~s\widetilde V_s and V~t\widetilde V_t using bidirectional cross-graph attention. With learnable query, key, value, and output projections Wq,Wk,Wv,WpW_q,W_k,W_v,W_p, the cross-aware node matrices are

    V^s=LN⁡ ⁣(softmax⁡ ⁣[(V~sWq)(V~tWk)T](V~tWv)Wp+V~s),\widehat V_s=\operatorname{LN}\!\left(\operatorname{softmax}\!\left[(\widetilde V_sW_q)(\widetilde V_tW_k)^T\right](\widetilde V_tW_v)W_p+\widetilde V_s\right),

    V^t=LN⁡ ⁣(softmax⁡ ⁣[(V~tWq)(V~sWk)T](V~sWv)Wp+V~t),\widehat V_t=\operatorname{LN}\!\left(\operatorname{softmax}\!\left[(\widetilde V_tW_q)(\widetilde V_sW_k)^T\right](\widetilde V_sW_v)W_p+\widetilde V_t\right),

    where softmax is applied row-wise and layer normalization is denoted by LN⁡\operatorname{LN}. An auxiliary classifier is trained on these nodes with cross-entropy: source nodes use ground-truth category labels, while target nodes use pseudo-labels from target score maps.

    For source node ii and target node jj, SIGMA computes a semantic-aware affinity rather than relying only on local visual similarity:

    Maffi,j=fmlp ⁣(fp(v^si)∥fp(v^tj)),M_{\mathrm{aff}}^{i,j}=f_{\mathrm{mlp}}\!\left(f_p(\widehat v_s^i)\mathbin{\|}f_p(\widehat v_t^j)\right),

    where fpf_p is a linear projection, ∥\mathbin{\|} denotes concatenation, and fmlpf_{\mathrm{mlp}} is an MLP with one output channel. Instance normalization followed by 20 iterations of a differentiable Sinkhorn operation converts MaffM_{\mathrm{aff}} into a doubly stochastic matrix M~aff\widetilde M_{\mathrm{aff}}. Positive entries of this matrix represent candidate cross-domain node correspondences.

  6. Knowl 6 — Structure-aware graph matching loss

    equation

    SIGMA trains the source-target affinity matrix with a matching loss that combines true-positive enhancement, false-positive suppression, and graph-structure consistency. Let YΠ∈{0,1}Ns×NtY_\Pi\in\{0,1\}^{N_s\times N_t} satisfy (YΠ)ij=1(Y_\Pi)_{ij}=1 exactly when source node ii and target node jj have the same category, and let 1\mathbf{1} be the all-ones matrix of the same size. For source and target adjacency matrices AsA_s and AtA_t, the loss is

    \mathcal{L}_{\mathrm{mat}}={}&\sum_i\frac{1}{N_s}\left[\max_j\left(\widetilde M_{\mathrm{aff}}\odot Y_\Pi\right)_{ij}-1\right]^2\\ &+\sum_{i,j}\frac{1}{\|\mathbf{1}-Y_\Pi\|_1}\left[\widetilde M_{\mathrm{aff}}\odot(\mathbf{1}-Y_\Pi)\right]_{ij}^{2}\\ &+\sum_{i,j}\frac{1}{N_sN_t}\left(A_s\widetilde M_{\mathrm{aff}}-\widetilde M_{\mathrm{aff}}A_t\right)_{ij}, \end{aligned}$$ where $\odot$ is element-wise multiplication and $\|\cdot\|_1$ is the entrywise $\ell_1$ norm. The first term raises the strongest same-category affinity for each source node, the second penalizes affinity assigned to different-category pairs, and the third imposes a quadratic structural constraint by making matched neighborhoods compatible across domains. The loss therefore aligns individual nodes while also preserving local graph relationships.
  7. Knowl 7 — Joint optimization with node and global adversarial alignment

    equation

    In addition to graph matching, SIGMA uses a node discriminator and a global discriminator. The node discriminator receives existing graph nodes through a gradient-reversal layer, three stacked fully connected–layer-normalization–ReLU blocks fbf_b, and a domain classifier fdcf_{dc}. With source domain label ds=1d_s=1 and target domain label dt=0d_t=0, its binary cross-entropy loss is

    LNA=−∑i=1Nslog⁡fdc(fb(vsi))−∑i=1Ntlog⁡[1−fdc(fb(vti))].\mathcal{L}_{\mathrm{NA}}=-\sum_{i=1}^{N_s}\log f_{dc}(f_b(v_s^i))-\sum_{i=1}^{N_t}\log\left[1-f_{dc}(f_b(v_t^i))\right].

    The global alignment loss LGA\mathcal{L}_{\mathrm{GA}} is an adversarial feature-alignment loss applied to the image-level features, and Ldet\mathcal{L}_{\mathrm{det}} is the standard supervised detection loss on source images. The total training objective is

    L=λ1Lnode+λ2Lmat+LNA+LGA+Ldet,\mathcal{L}=\lambda_1\mathcal{L}_{\mathrm{node}}+\lambda_2\mathcal{L}_{\mathrm{mat}}+\mathcal{L}_{\mathrm{NA}}+\mathcal{L}_{\mathrm{GA}}+\mathcal{L}_{\mathrm{det}},

    where Lnode\mathcal{L}_{\mathrm{node}} is auxiliary node-classification cross-entropy and the implementation uses λ1=λ2=0.1\lambda_1=\lambda_2=0.1. Thus, graph-specific semantic matching is trained jointly with both category supervision and domain-invariance objectives.

  8. Knowl 8 — Domain-adaptive detection evaluation protocol

    experimental setup

    SIGMA is evaluated under unsupervised domain adaptation, using labeled source images during training and unlabeled target images for adaptation and evaluation. The three transfer scenarios are Cityscapes →\rightarrow Foggy Cityscapes, Sim10k →\rightarrow Cityscapes, and KITTI →\rightarrow Cityscapes. Cityscapes contains 2,975 training and 500 validation images with eight annotated object categories; Foggy Cityscapes is a fog-corrupted version. Sim10k contains 10,000 synthetic images with annotated cars, and KITTI contains 7,481 real traffic-scene images with annotated cars.

    The detector uses either VGG-16 or ResNet-50 as feature extractor and is implemented in PyTorch. Training uses SGD with learning rate 0.00250.0025, batch size 44, momentum 0.90.9, and weight decay 5×10−45\times10^{-4}. At most 100 graph nodes are sampled from each feature map in each domain in the main experiments. Because graph matching can fail when a target batch has no usable nodes, SIGMA is warm-started before the BGM adaptor is introduced. Performance is reported with mean average precision over the detector's IoU thresholds, together with source-only performance and the adaptation gain over source-only training.

  9. Knowl 9 — SIGMA achieves state-of-the-art adaptation performance

    empirical result

    Across all three evaluated transfer scenarios, the proposed SIGMA framework substantially improves over source-only detection and prior domain-adaptive detectors. On Cityscapes →\rightarrow Foggy Cityscapes, SIGMA obtains 43.5%43.5\% mAP with VGG-16 and 44.2%44.2\% mAP with ResNet-50. Its VGG-16 per-category AP values are 46.946.9 for person, 48.448.4 for rider, 63.763.7 for car, 27.127.1 for truck, 50.750.7 for bus, 35.935.9 for train, 34.734.7 for motor, and 41.441.4 for bike; the source-only score is 18.4%18.4\% and the adaptation gain is 25.125.1 percentage points. With ResNet-50, the corresponding per-category AP values are 44.044.0, 43.943.9, 60.360.3, 31.631.6, 50.450.4, 51.551.5, 31.731.7, and 40.640.6, with source-only score 24.2%24.2\% and gain 20.020.0 points.

    On Sim10k →\rightarrow Cityscapes, SIGMA reaches 53.7%53.7\% mAP, with source-only score 39.8%39.8\% and adaptation gain 13.913.9 points. On KITTI →\rightarrow Cityscapes, it reaches 45.8%45.8\% mAP, with source-only score 34.4%34.4\% and gain 11.411.4 points. The method also exceeds same-detector comparison systems on Cityscapes →\rightarrow Foggy Cityscapes by 7.57.5, 2.62.6, and 3.93.9 mAP points over EPM, KTNet, and SSAL, respectively, and on Sim10k →\rightarrow Cityscapes by 4.74.7, 3.03.0, and 1.91.9 points over those methods.

  10. Knowl 10 — Component ablations and node-count sensitivity

    empirical result

    Ablations on Cityscapes →\rightarrow Foggy Cityscapes with VGG-16 show that every major SIGMA component contributes. The global-alignment baseline obtains 35.3%35.3\% mAP. Adding the complete Graph-embedded Semantic Completion module raises performance to 41.8%41.8\% mAP, while removing domain-guided node completion reduces it to 40.4%40.4\%, replacing the graph-guided memory bank with a common buffer gives 41.0%41.0\%, and removing the node discriminator gives 39.4%39.4\%. Adding the complete Bipartite Graph Matching adaptor reaches 43.5%43.5\% mAP; removing cross-graph interaction gives 42.8%42.8\%, replacing semantic-aware affinity with the simpler affinity gives 42.6%42.6\%, and removing the structure-aware matching loss gives 42.2%42.2\% in the tabulated ablation results.

    The maximum number of sampled source and target nodes per feature map also matters. Using only source nodes gives 36.8%36.8\% mAP, only target nodes gives 37.3%37.3\%, and using equal source-target counts of 2020, 5050, 100100, 200200, and 500500 gives 39.0%39.0\%, 41.0%41.0\%, 43.5%43.5\%, 43.9%43.9\%, and 42.6%42.6\%, respectively. Thus, increasing node coverage generally improves graph matching up to 200 nodes, while excessive node counts make matching optimization more difficult. For the matching loss, single-best matching with true-positive enhancement, false-positive suppression, and quadratic constraints obtains 43.5%43.5\% mAP at IoU 0.50.5, compared with 42.1%42.1\% using only true-positive enhancement and 43.2%43.2\% using true-positive enhancement plus false-positive suppression; multiple matching obtains 42.9%42.9\% with BCE and 43.1%43.1\% with MSE.

Coverage note — The qualitative detection examples and t-SNE visualizations, along with the full per-class scores of every competing method, are omitted because they corroborate the quantitative comparisons without adding an independent method or load-bearing result.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 4, 5
  2. 2.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv preprint arXiv:2105.04906, 2021. 2
  3. 3.Yuhua Chen, Wen Li, Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Domain adaptive faster r-cnn for object detection in the wild. In CVPR, pages 3339–3348, 2018. 1, 2, 4, 6
  4. 4.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, pages 3213–3223, 2016. 6
  5. 5.Jinhong Deng, Wen Li, Yuhua Chen, and Lixin Duan. Unbiased mean teacher for cross-domain object detection. In CVPR, pages 4091–4101, June 2021. 6
  6. 6.Kexue Fu, Shaolei Liu, Xiaoyuan Luo, and Manning Wang. Robust point cloud registration framework based on deep graph matching. In CVPR, pages 8893–8902, 2021. 3, 5
  7. 7.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In ICML, pages 1180–1189, 2015. 6
  8. 8.Quankai Gao, Fudong Wang, Nan Xue, Jin-Gang Yu, and Gui-Song Xia. Deep graph matching under quadratic constraint. In CVPR, pages 5069–5078, 2021. 3, 5, 7
  9. 9.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, pages 3354–3361, 2012. 7
  10. 10.Jiawei He, Zehao Huang, Naiyan Wang, and Zhaoxiang Zhang. Learnable graph matching: Incorporating graph partitioning with deep feature learning for multiple object tracking. In CVPR, pages 5299–5309, 2021. 3, 5
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 7
  12. 12.Cheng-Chun Hsu, Yi-Hsuan Tsai, Yen-Yu Lin, and Ming-Hsuan Yang. Every pixel matters: Center-aware feature alignment for domain adaptive object detector. In ECCV, pages 733–748, 2020. 1, 2, 6, 7, 8
  13. 13.Han-Kai Hsu, Chun-Han Yao, Yi-Hsuan Tsai, Wei-Chih Hung, Hung-Yu Tseng, Maneesh Singh, and Ming-Hsuan Yang. Progressive domain adaptation for object detection. In WACV, pages 749–757, 2020. 2
  14. 14.Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, and Kiyoharu Aizawa. Cross-domain weakly-supervised object detection through progressive domain adaptation. In CVPR, pages 5001–5009, 2018. 2
  15. 15.Matthew Johnson-Roberson, Charles Barto, Rounak Mehta, Sharath Nittur Sridhar, Karl Rosaen, and Ram Vasudevan. Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks? In ICRA, pages 746–753, 2017. 6
  16. 16.Taekyung Kim, Minki Jeong, Seunghyeon Kim, Seokeon Choi, and Changick Kim. Diversify and match: A domain adaptive representation learning paradigm for object detection. In CVPR, pages 12456–12465, 2019. 2
  17. 17.Congcong Li, Dawei Du, Libo Zhang, Longyin Wen, Tiejian Luo, Yanjun Wu, and Pengfei Zhu. Spatial attention pyramid network for unsupervised domain adaptation. In ECCV, pages 481–497. Springer, 2020. 2
  18. 18.Chuang Lin, Zehuan Yuan, Sicheng Zhao, Peize Sun, Changhu Wang, and Jianfei Cai. Domain-invariant disentangled network for generalizable object detection. In ICCV, pages 8771–8780, October 2021. 6
  19. 19.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, pages 2980–2988, 2017. 4
  20. 20.Eliane Maria Loiola, Nair Maria Maia de Abreu, Paulo Oswaldo Boaventura-Netto, Peter Hahn, and Tania Querido. A survey for the quadratic assignment problem. Eur. J. Oper. Res., 176(2):657–690, 2007. 3
  21. 21.Muhammad Akhtar Munir, Muhammad Haris Khan, M Saquib Sarfraz, and Mohsen Ali. Synergizing between self-training and adversarial learning for domain adaptive object detection. 2021. 2, 6, 7
  22. 22.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, pages 8024–8035, 2019. 7
  23. 23.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. 1, 4
  24. 24.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: towards real-time object detection with region proposal networks. In NeurIPS, pages 91–99, 2015. 1, 4
  25. 25.Farzaneh Rezaeianaran, Rakshith Shetty, Rahaf Aljundi, Daniel Olmeda Reino, Shanshan Zhang, and Bernt Schiele. Seeking similarities over differences: Similarity-based domain alignment for adaptive object detection. In ICCV, pages 9204–9213, 2021. 6
  26. 26.Yu Rong, Wenbing Huang, Tingyang Xu, and Junzhou Huang. Dropedge: Towards deep graph convolutional networks on node classification. ICLR, 2020. 4
  27. 27.Kuniaki Saito, Yoshitaka Ushiku, Tatsuya Harada, and Kate Saenko. Strong-weak distribution alignment for adaptive object detection. In CVPR, pages 6956–6965, 2019. 1, 2
  28. 28.Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Semantic foggy scene understanding with synthetic data. Int J Comput Vis, 126(9):973–992, 2018. 6
  29. 29.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 7
  30. 30.Richard Sinkhorn. A relationship between arbitrary positive matrices and doubly stochastic matrices. Ann. math. stat., 35:876–879, 1964. 5
  31. 31.X Yu Stella and Jianbo Shi. Multiclass spectral clustering. In ICCV, volume 2, pages 313–313. IEEE Computer Society, 2003. 5
  32. 32.Kun Tian, Chenghao Zhang, Ying Wang, Shiming Xiang, and Chunhong Pan. Knowledge mining and transferring for domain adaptive object detection. In ICCV, pages 9133–9142, October 2021. 1, 2, 3, 6, 7
  33. 33.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: Fully convolutional one-stage object detection. In ICCV, pages 9627–9636, 2019. 1, 4, 6, 7
  34. 34.Vibashan VS, Vikram Gupta, Poojan Oza, Vishwanath A. Sindagi, and Vishal M. Patel. Mega-cda: Memory guided attention for category-aware unsupervised domain adaptive object detection. In CVPR, pages 4516–4526, June 2021. 5, 6, 7
  35. 35.Yu Wang, Rui Zhang, Shuo Zhang, Miao Li, Yangyang Xia, Xishan Zhang, and Shaoli Liu. Domain-specific suppression for adaptive object detection. In CVPR, pages 9603–9612, June 2021. 2, 6
  36. 36.Aming Wu, Rui Liu, Yahong Han, Linchao Zhu, and Yi Yang. Vector-decomposed disentanglement for domain-invariant object detection. ICCV, 2021. 6
  37. 37.Minghao Xu, Hang Wang, Bingbing Ni, Qi Tian, and Wenjun Zhang. Cross-domain detection via graph-induced prototype alignment. In CVPR, pages 12355–12364, 2020. 1, 2, 3, 4, 6, 7
  38. 38.Junchi Yan, Xu-Cheng Yin, Weiyao Lin, Cheng Deng, Hongyuan Zha, and Xiaokang Yang. A short survey of recent advances in graph matching. In ACM ICMR, pages 167–174, 2016. 3
  39. 39.Xu Yang, Cheng Deng, Tongliang Liu, and Dacheng Tao. Heterogeneous graph attention network for unsupervised multiple-target domain adaptation. IEEE Trans. Pattern Anal. Mach. Intell., pages 1–1, 2020. 2, 6
  40. 40.Weilin Zhang and Yu-Xiong Wang. Hallucination improves few-shot object detection. In CVPR, pages 13008–13017, 2021. 2
  41. 41.Yixin Zhang, Zilei Wang, and Yushi Mao. Rpn prototype alignment for domain adaptive object detector. In CVPR, pages 12425–12434, June 2021. 1, 2, 3, 4, 6, 7
  42. 42.Yangtao Zheng, Di Huang, Songtao Liu, and Yunhong Wang. Cross-domain object detection through coarse-to-fine feature adaptation. In CVPR, pages 13766–13775, 2020. 1, 2, 3, 4, 5, 6, 7

Citation

MLA
Li, W., et al. “SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection”. arXiv, 2022, http://arxiv.org/abs/2203.06398v3.
APA
Li, W., Liu, X., & Yuan, Y. (2022). SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection. arXiv. http://arxiv.org/abs/2203.06398v3
Chicago
Li, W., X. Liu, and Y. Yuan. 2022. “SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection”. arXiv. http://arxiv.org/abs/2203.06398v3.
Harvard
Li, W., Liu, X. and Yuan, Y. (2022) “SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.06398v3.
Vancouver
1. Li W, Liu X, Yuan Y (2022) SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection. arXiv

BibTeX

@article{li2022sigma,
  title = {SIGMA: Semantic-complete Graph Matching for Domain Adaptive Object Detection},
  author = {Li, Wuyang and Liu, Xinyu and Yuan, Yixuan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.06398v3},
  eprint = {2203.06398}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE