Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection

Vibashan VSPoojan OzaVishal M. Patel

article2023CVPR120 citations

Proposes an Instance Relation Graph framework that models inter-proposal similarities to guide contrastive representation learning, enabling object detectors to adapt to new target domains without requiring access to source training data.

Listen

Deploying deep learning object detection models into new visual environments frequently leads to substantial performance drops caused by domain shift, such as changes in weather, camera hardware, or artistic styles. While conventional adaptation methods rely on simultaneous access to the original training data and new target data, real-world constraints such as strict privacy regulations, data proprietary concerns, and bandwidth limitations often make transmitting large source datasets impractical. Organizations therefore require techniques to adapt existing models to new operational domains using only the pre-trained model and unlabeled target data.

The article develops and evaluates a source-free domain adaptation framework that updates an object detector on unlabeled target imagery without accessing the original source dataset. The method introduces an Instance Relation Graph network paired with a graph-guided contrastive loss within a student-teacher knowledge distillation architecture to enhance target feature representations.

To establish credibility across varied deployment scenarios, the approach was tested on standard benchmarks covering four distinct domain shifts: adverse weather (adapting from clear to foggy driving scenes), cross-camera sensor variations, synthetic-to-real transfer, and realistic-to-artistic style transitions. Rather than relying on computationally heavy image-level contrastive learning, the method leverages class-agnostic proposals naturally generated by the detector's region proposal network as built-in augmentations. A graph neural network models the pairwise relationships between these proposals, generating positive and negative pairings to guide contrastive representation learning without requiring true target labels.

The experimental findings show significant performance gains across all evaluated scenarios. In adverse weather adaptation, the method achieved a 37.1 mean Average Precision, outperforming previous source-free baselines by 2.4 to 6.5 percentage points and surpassing many methods that require full source data access. In cross-camera and synthetic-to-real scenarios, the model delivered top-tier performance (reaching 46.9 and 45.2 mean Average Precision for car detection), improving upon competing source-free frameworks by roughly 2.0 to 3.3 percentage points. Ablation experiments confirmed that incorporating the relation graph network and its associated contrastive loss systematically improved target accuracy by 2.8 percentage points over basic student-teacher self-training.

These results demonstrate that organizations can successfully deploy and refine computer vision models on edge devices or decentralized environments without transmitting tens to hundreds of gigabytes of proprietary source data. This substantially reduces data transmission costs, mitigates privacy and regulatory risks, and shortens deployment timelines when expanding vision systems into unmodeled operating conditions.

Engineering and deployment teams should consider adopting graph-guided source-free adaptation pipelines when migrating detection systems to client-side hardware or privacy-constrained domains. Prior to full-scale rollout, teams should conduct pilot evaluations in the intended target environment to select appropriate prediction confidence thresholds (such as the 0.9 threshold utilized here) and graph relation cutoff parameters. While the methodology was validated across multiple standard benchmarks using two-stage detector architectures, performance depends on the initial quality of the source model and the presence of sufficient object proposals. Further testing is advised when applying this framework to extreme visual shifts or alternative single-stage detector backbones.

Cover for Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection

Abstract

Unsupervised Domain Adaptation (UDA) is an effective approach to tackle the issue of domain shift. Specifically, UDA methods try to align the source and target representations to improve generalization on the target domain. Further, UDA methods work under the assumption that the source data is accessible during the adaptation process. However, in real-world scenarios, the labelled source data is often restricted due to privacy regulations, data transmission constraints, or proprietary data concerns. The Source-Free Domain Adaptation (SFDA) setting aims to alleviate these concerns by adapting a source-trained model for the target domain without requiring access to the source data. In this paper, we explore the SFDA setting for the task of adaptive object detection. To this end, we propose a novel training strategy for adapting a source-trained object detector to the target domain without source data. More precisely, we design a novel contrastive loss to enhance the target representations by exploiting the objects relations for a given target domain input. These object instance relations are modelled using an Instance Relation Graph (IRG) network, which are then used to guide the contrastive representation learning. In addition, we utilize a student-teacher to effectively distill knowledge from source-trained model to target domain. Extensive experiments on multiple object detection benchmark datasets show that the proposed approach is able to efficiently adapt source-trained object detectors to the target domain, outperforming state-of-the-art domain adaptive detection methods. Code and models are provided in https://viudomain.github.io/irg-sfda-web/.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. Proposed method
  • 3.1. Preliminaries
  • 3.2. Graph-guided contrastive learning
  • 3.2.1 Instance Relation Graph (IRG):
  • 3.2.2 Graph Distillation Loss (GDL).
  • 3.2.3 Graph Contrastive Loss (GCL)
  • 3.3. Overall loss function
  • 4. Experiments and Results
  • 4.1. Quantitative comparison
  • 4.1.1 Adaptation to adverse weather:
  • 4.1.2 Realistic to artistic data adaptation:
  • 4.1.3 Synthetic to real-world adaptation
  • 4.2. Ablation analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Mean-Teacher Self-Training Framework for Source-Free Adaptive Object Detection

    model/method

    In source-free domain adaptive object detection (SFDA-OD), only a source-trained detector Θ\Theta and unlabeled target domain images Dt={xtn}n=1NtD_t = \{x_t^n\}_{n=1}^{N_t} are available during adaptation. To distill target domain knowledge while mitigating error propagation from noisy pseudo-labels, a mean-teacher framework is employed consisting of a student network parameterized by Θs\Theta_s and a teacher network parameterized by Θt\Theta_t, both initialized with Θ\Theta.

    The teacher network receives weakly augmented target images to produce pseudo-labels y~tn\tilde{y}_t^n by filtering out predictions with confidence lower than a threshold T=0.9T = 0.9. The student network processes strongly augmented versions of the same target images and is trained using the pseudo-label detection loss:

    LSLst=Lclsrpn(xtn,y~tn)+Lregrpn(xtn,y~tn)+Lclsroi(xtn,y~tn)+Lregroi(xtn,y~tn)L_{SL}^{st} = L_{cls}^{rpn}(x_t^n, \tilde{y}_t^n) + L_{reg}^{rpn}(x_t^n, \tilde{y}_t^n) + L_{cls}^{roi}(x_t^n, \tilde{y}_t^n) + L_{reg}^{roi}(x_t^n, \tilde{y}_t^n)

    where LclsrpnL_{cls}^{rpn} and LregrpnL_{reg}^{rpn} denote the classification and regression losses for the Region Proposal Network (RPN), and LclsroiL_{cls}^{roi} and LregroiL_{reg}^{roi} denote the classification and bounding-box regression losses for the RoI head.

    The network parameters are updated as follows:

    Θs←Θs+γ∂LSFDA∂Θs\Theta_s \leftarrow \Theta_s + \gamma \frac{\partial L_{SFDA}}{\partial \Theta_s}

    Θt←αΘt+(1−α)Θs\Theta_t \leftarrow \alpha \Theta_t + (1 - \alpha) \Theta_s

    where γ\gamma is the student learning rate, α∈[0,1]\alpha \in [0, 1] is the Exponential Moving Average (EMA) momentum coefficient (set to α=0.9\alpha = 0.9), and LSFDAL_{SFDA} is the total source-free domain adaptation loss.

  2. Knowl 2 — Instance Relation Graph Construction and Proposal Feature Aggregation

    model/method

    To model contextual and semantic relationships across class-agnostic object proposals without ground-truth labels on the target domain, the Instance Relation Graph (IRG) constructs a graph G=⟨V,E⟩G = \langle V, E \rangle over proposals extracted from an image.

    The node set V={v1,v2,…,vm}V = \{v_1, v_2, \dots, v_m\} contains mm RoI-pooled feature vectors vi∈Rdv_i \in \mathbb{R}^d corresponding to the mm object proposals generated by the teacher's Region Proposal Network (RPN), where m=300m = 300 and dd is the feature dimension. Teacher proposals are utilized for extracting both student and teacher RoI features because weak augmentations yield more consistent proposal locations.

    The edge matrix E=[eij]m×mE = [e_{ij}]_{m \times m} encodes pairwise relation affinities between proposal features viv_i and vjv_j:

    eij=exp⁡(Sij)∑k=1mexp⁡(Sik),Sij=f(vi)⋅g(vj)Te_{ij} = \frac{\exp(S_{ij})}{\sum_{k=1}^m \exp(S_{ik})}, \quad S_{ij} = f(v_i) \cdot g(v_j)^T

    where f(⋅)f(\cdot) and g(⋅)g(\cdot) are learnable linear projection functions.

    Given the RoI feature matrix F∈Rm×dF \in \mathbb{R}^{m \times d}, inter-instance feature aggregation is performed through a Graph Convolutional Network (GCN) layer parameterized by learnable weight matrix W∈Rd×dW \in \mathbb{R}^{d \times d}:

    F~=ReLU(EFW)\tilde{F} = \text{ReLU}(E F W)

    where F~∈Rm×d\tilde{F} \in \mathbb{R}^{m \times d} denotes the relation-aggregated proposal representations.

  3. Knowl 3 — Graph Distillation Loss for Instance Relation Regularization

    equation

    The Graph Distillation Loss (LGDLL_{GDL}) supervises the learnable parameters of the Instance Relation Graph (IRG) and maintains consistency between student and teacher representations.

    Both the original proposal features F∈Rm×dF \in \mathbb{R}^{m \times d} and the IRG-aggregated features F~∈Rm×d\tilde{F} \in \mathbb{R}^{m \times d} are fed into the classification head to compute class logits. Let Zst,Zte∈Rm×CZ_{st}, Z_{te} \in \mathbb{R}^{m \times C} denote student and teacher class logits obtained from FF, and let Z~st,Z~te∈Rm×C\tilde{Z}_{st}, \tilde{Z}_{te} \in \mathbb{R}^{m \times C} denote the logits obtained from F~\tilde{F}, where CC is the number of object classes. The Graph Distillation Loss is defined as:

    LGDL=KL(σ(Zst),σ(Z~st))+KL(σ(Zte),σ(Z~te))+KL(σ(Zst),σ(Zte))L_{GDL} = \text{KL}(\sigma(Z_{st}), \sigma(\tilde{Z}_{st})) + \text{KL}(\sigma(Z_{te}), \sigma(\tilde{Z}_{te})) + \text{KL}(\sigma(Z_{st}), \sigma(Z_{te}))

    where KL(p,q)=∑cpclog⁡(pc/qc)\text{KL}(p, q) = \sum_c p_c \log(p_c / q_c) represents the Kullback-Leibler divergence between predicted class probability distributions, and σ(⋅)\sigma(\cdot) denotes the softmax operator along the class dimension. Minimizing LGDLL_{GDL} trains the relation matrix EE end-to-end to capture reliable inter-instance correlations while aligning student and teacher prediction spaces.

  4. Knowl 4 — IRG-Guided Graph Contrastive Learning and Loss Formulation

    model/method

    To pull together feature embeddings of proposals corresponding to the same instance or category and repel disparate proposals without requiring target ground-truth annotations, the Graph Contrastive Loss (LGCLL_{GCL}) uses the relation matrix E=[eij]m×mE = [e_{ij}]_{m \times m} learned by the Instance Relation Graph (IRG).

    Pairwise binary labels Mij∈{0,1}M_{ij} \in \{0, 1\} between proposal ii and proposal jj are assigned by thresholding the normalized relation affinity eije_{ij} with hyperparameter ϵ\epsilon:

    Mij={1if eij≥ϵ0if eij<ϵM_{ij} = \begin{cases} 1 & \text{if } e_{ij} \ge \epsilon \\ 0 & \text{if } e_{ij} < \epsilon \end{cases}

    For proposal ii, RoI feature vector vi∈Rdv_i \in \mathbb{R}^d is linearly projected to key ki∈Rdkk_i \in \mathbb{R}^{d_k} and query qi∈Rdkq_i \in \mathbb{R}^{d_k} using weight matrices Wk,Wq∈Rdk×dW_k, W_q \in \mathbb{R}^{d_k \times d}:

    ki=Wkvi,qi=Wqvik_i = W_k v_i, \quad q_i = W_q v_i

    Let I={1,2,…,m}I = \{1, 2, \dots, m\} denote the set of proposal indices. For anchor proposal i∈Ii \in I, the set of all other candidate proposals is A(i)=I∖{i}A(i) = I \setminus \{i\}, and the set of positive proposal pairs is P(i)={p∈I:Mip=1}∖{i}P(i) = \{p \in I : M_{ip} = 1\} \setminus \{i\}. The Graph Contrastive Loss is formulated as:

    LGCL=∑i∈I−log⁡(1∣P(i)∣∑p∈P(i)exp⁡(qikpT)∑a∈A(i)exp⁡(qikaT))L_{GCL} = \sum_{i \in I} -\log \left( \frac{1}{|P(i)|} \sum_{p \in P(i)} \frac{\exp(q_i k_p^T)}{\sum_{a \in A(i)} \exp(q_i k_a^T)} \right)

    LGCLL_{GCL} is used exclusively to update the parameters of the student network, while the teacher network parameters are updated via Exponential Moving Average (EMA).

  5. Knowl 5 — Overall Optimization Objective for IRG-Guided SFDA Detection

    equation

    The complete objective function LSFDAL_{SFDA} for adapting a source-trained object detector to an unlabeled target domain without access to source data combines pseudo-label self-training, graph distillation, and graph-guided contrastive learning:

    LSFDA=LSLst+LGDL+LGCLL_{SFDA} = L_{SL}^{st} + L_{GDL} + L_{GCL}

    where:

    • LSLstL_{SL}^{st} is the student pseudo-label detection loss computed on teacher-filtered predictions (T=0.9T = 0.9): LSLst=Lclsrpn(xtn,y~tn)+Lregrpn(xtn,y~tn)+Lclsroi(xtn,y~tn)+Lregroi(xtn,y~tn)L_{SL}^{st} = L_{cls}^{rpn}(x_t^n, \tilde{y}_t^n) + L_{reg}^{rpn}(x_t^n, \tilde{y}_t^n) + L_{cls}^{roi}(x_t^n, \tilde{y}_t^n) + L_{reg}^{roi}(x_t^n, \tilde{y}_t^n)
    • LGDLL_{GDL} is the Graph Distillation Loss enforcing consistency across original RoI features and GCN-aggregated features: LGDL=KL(σ(Zst),σ(Z~st))+KL(σ(Zte),σ(Z~te))+KL(σ(Zst),σ(Zte))L_{GDL} = \text{KL}(\sigma(Z_{st}), \sigma(\tilde{Z}_{st})) + \text{KL}(\sigma(Z_{te}), \sigma(\tilde{Z}_{te})) + \text{KL}(\sigma(Z_{st}), \sigma(Z_{te}))
    • LGCLL_{GCL} is the Graph Contrastive Loss encouraging instance-discriminative representations on the target domain: LGCL=∑i∈I−log⁡(1∣P(i)∣∑p∈P(i)exp⁡(qikpT)∑a∈A(i)exp⁡(qikaT))L_{GCL} = \sum_{i \in I} -\log \left( \frac{1}{|P(i)|} \sum_{p \in P(i)} \frac{\exp(q_i k_p^T)}{\sum_{a \in A(i)} \exp(q_i k_a^T)} \right)
  6. Knowl 6 — Implementation and Optimization Details for IRG-SFDA

    experimental setup

    The IRG-guided source-free domain adaptation framework is implemented based on the Faster-RCNN detector with an ImageNet pre-trained ResNet-50 backbone.

    • Image Preprocessing: Target images are resized such that the shorter side is 600 pixels while maintaining aspect ratio. The adaptation batch size is set to 1.
    • Source Training: The source detection model is trained for 10 epochs using the SGD optimizer with a learning rate of 0.001 and momentum of 0.9.
    • Target Adaptation: Target adaptation is conducted for 10 epochs using SGD with a learning rate of 0.001 and momentum of 0.9.
    • Mean-Teacher and Graph Parameters: The teacher EMA momentum rate is α=0.9\alpha = 0.9. Pseudo-labels from the teacher with confidence exceeding threshold T=0.9T = 0.9 are selected for student supervision. The number of RPN proposals extracted per image is m=300m = 300.
    • Evaluation: Detection performance is measured by mean Average Precision (mAP) at an IoU threshold of 0.5 evaluated on the teacher network over the target domain evaluation set.
  7. Knowl 7 — Adverse Weather Adaptation Performance on Cityscapes to FoggyCityscapes

    data/table

    Performance is benchmarked for adapting from clear weather (Cityscapes, 2,975 training images) to adverse foggy weather (FoggyCityscapes, 500 validation images) across 8 categories: person (prsn), rider, car, truck, bus, train, motorcycle (mcycle), and bicycle (bcycle). Under SFDA, no source images are available during adaptation.

    Type Method prsn rider car truck bus train mcycle bcycle mAP
    Source Source Only 29.3 34.1 35.8 15.4 26.0 9.09 22.4 29.7 25.2
    UDA DA Faster 25.0 31.0 40.5 22.1 35.3 20.2 20.0 27.1 27.6
    UDA DMatch 30.8 40.5 44.3 27.2 38.4 34.5 28.4 32.2 34.6
    UDA MTOR 30.6 41.4 44.0 21.9 38.6 40.6 28.3 35.6 35.1
    UDA SWDA 29.9 42.3 43.5 24.5 36.2 32.6 30.0 35.3 34.3
    UDA CDN 35.8 45.7 50.9 30.1 42.5 29.8 30.8 36.5 36.6
    UDA Collaborative DA 32.7 44.4 50.1 21.7 45.6 25.4 30.1 36.8 35.9
    UDA iFAN DA 32.6 48.5 22.8 40.0 33.0 45.5 31.7 27.9 35.3
    UDA Instance DA 33.1 43.4 49.6 21.9 45.7 32.0 29.5 37.0 36.5
    UDA Progressive DA 36.0 45.5 54.4 24.3 44.1 25.8 29.1 35.9 36.9
    UDA Categorical DA 32.9 43.8 49.2 27.2 45.1 36.4 30.3 34.6 37.4
    UDA MeGA CDA 37.7 49.0 52.4 25.4 49.2 46.9 34.5 39.0 41.8
    UDA Unbiased DA 33.8 47.3 49.8 30.0 48.2 42.1 33.0 37.3 40.4
    SFDA SFOD 21.7 44.0 40.4 32.2 11.8 25.3 34.5 34.3 30.6
    SFDA SFOD-Mosaic 25.5 44.5 40.7 33.2 22.2 28.4 34.1 39.0 33.5
    SFDA HCL 26.9 46.0 41.3 33.0 25.0 28.1 35.9 40.7 34.6
    SFDA LODS 34.0 45.7 48.8 27.3 39.7 19.6 33.2 37.8 35.8
    SFDA Mean-Teacher 33.9 43.0 45.0 29.2 37.2 25.1 25.6 38.2 34.3
    SFDA IRG (Ours) 37.4 45.2 51.9 24.4 39.6 25.2 31.5 41.6 37.1
    Oracle Fully Supervised 38.7 46.9 56.7 35.5 49.4 44.7 35.9 38.8 43.1

    The IRG method reaches 37.1% mAP, outperforming previous SFDA approaches (SFOD at 30.6%, SFOD-Mosaic at 33.5%, HCL at 34.6%, LODS at 35.8%, and Mean-Teacher at 34.3%) and surpassing many full UDA methods that rely on labeled source domain data during adaptation.

  8. Knowl 8 — Synthetic-to-Real and Cross-Camera Adaptation Performance

    data/table

    Adaptation performance on the 'car' category is evaluated on two cross-domain benchmarks:

    1. Synthetic to Real (Sim10K→Cityscapes\text{Sim10K} \to \text{Cityscapes}): Adapting from 10,000 synthetic images rendered by the Grand Theft Auto engine to 2,975 real-world Cityscapes images.
    2. Cross-Camera (KITTI→Cityscapes\text{KITTI} \to \text{Cityscapes}): Adapting from 7,481 KITTI images to Cityscapes under different camera setups.
    Type Method Sim10k →\to Cityscapes (AP of Car) KITTI →\to Cityscapes (AP of Car)
    Source Source Only 32.0 33.9
    UDA DA Faster 38.9 38.5
    UDA Selective DA 43.0 42.5
    UDA MAF 41.1 41.0
    UDA Robust DA 42.5 42.9
    UDA Strong-Weak 40.1 37.9
    UDA ATF 42.8 42.1
    UDA Harmonizing 42.5 -
    UDA Cycle DA 41.5 41.7
    UDA MeGA CDA 44.8 43.0
    UDA Unbiased DA 43.1 -
    SFDA SFOD 42.3 43.6
    SFDA SFOD-Mosaic 42.9 44.6
    SFDA Mean-teacher 39.7 41.2
    SFDA IRG (Ours) 45.2 46.9

    IRG achieves 45.2% AP on Sim10K→Cityscapes\text{Sim10K} \to \text{Cityscapes} (+2.3% over SFOD-Mosaic, +0.4% over UDA method MeGA CDA) and 46.9% AP on KITTI→Cityscapes\text{KITTI} \to \text{Cityscapes} (+2.3% over SFOD-Mosaic, +3.9% over MeGA CDA), outperforming all prior SFDA and UDA baselines on these benchmarks.

  9. Knowl 9 — Real-to-Artistic Domain Adaptation Performance on Watercolor and Clipart

    data/table

    Cross-domain adaptation from real-world natural imagery (Pascal VOC 2007 + 2012) to artistic style domains is evaluated on two targets:

    1. Pascal VOC →\to Watercolor: 6 categories (bike, bird, car, cat, dog, person) evaluated on 1,000 testing images.
    2. Pascal VOC →\to Clipart: 20 categories evaluated on 1,000 unlabeled target images.
    Pascal-VOC →\to Watercolor
    Type Method bike bird car cat dog prsn mAP
    Source Source only 68.8 46.8 37.2 32.7 21.3 60.7 44.6
    UDA DA Faster 75.2 40.6 48.0 31.5 20.6 60.0 46.0
    UDA BDC Faster 68.6 48.3 47.2 26.5 21.7 60.5 45.5
    UDA BSR 82.8 43.2 49.8 29.6 27.6 58.4 48.6
    UDA WST 77.8 48.0 45.2 30.4 29.5 64.2 49.2
    UDA SWDA 71.3 52.0 46.6 36.2 29.2 67.3 50.4
    UDA HTCN 78.6 47.5 45.6 35.4 31.0 62.2 50.1
    UDA I3^3Net 81.1 49.3 46.2 35.0 31.9 65.7 51.5
    UDA Unbiased DA 88.2 55.3 51.7 39.8 43.6 69.9 55.6
    SFDA PL 74.6 46.5 45.1 27.3 25.9 54.4 46.1
    SFDA SFOD 76.2 44.9 49.3 31.6 30.6 55.2 47.9
    SFDA Mean-teacher 73.6 47.6 46.6 28.5 29.4 56.6 47.1
    SFDA IRG (Ours) 75.9 52.5 50.8 30.8 38.7 69.2 53.0

    On Pascal VOC →\to Watercolor, IRG attains 53.0% mAP, improving over SFOD (47.9%), Mean-Teacher (47.1%), and UDA methods SWDA (50.4%) and I3^3Net (51.5%). On Pascal VOC →\to Clipart (20 classes), IRG achieves 31.5% mAP, outperforming Source Only (27.8%), pseudo-labeling (PL, 28.2%), Mean-Teacher (29.1%), SFOD (29.5%), and UDA methods DAF (19.8%), ADDA (27.4%), and BDC Faster (25.6%).

  10. Knowl 10 — Ablation Analysis of Student-Teacher Augmentations, GDL, and GCL

    data/table

    An ablation study on Cityscapes→FoggyCityscapes\text{Cityscapes} \to \text{FoggyCityscapes} evaluates the individual and combined effects of augmentation pairings (Weak-Weak [WW], Strong-Strong [SS], Strong-Weak [SW]), Graph Distillation Loss (GDL), and Graph Contrastive Loss (GCL):

    Method PL GDL GCL prsn rider car truc bus train mcycle bcycle mAP
    Source Only 25.8 33.7 35.2 13.0 28.2 9.1 18.7 31.4 24.4
    MT + WW ✓ 35.8 42.6 43.9 23.1 32.7 11.0 29.9 38.7 32.2
    MT + SS ✓ 32.8 41.4 43.8 18.2 28.6 11.2 24.6 38.3 29.9
    MT + SW ✓ 33.9 43.0 45.0 29.1 37.2 25.1 25.5 38.2 34.3
    Ours (MT+SW+GDL) ✓ ✓ 37.2 43.1 51.0 28.6 40.1 21.2 28.2 37.1 35.9
    Ours (Full Model) ✓ ✓ ✓ 37.4 45.2 51.9 24.4 39.6 25.2 31.5 41.6 37.1
    1. Augmentation Strategy: Strong-Weak (SW) augmentation achieves 34.3% mAP, significantly outperforming Weak-Weak (32.2%) and Strong-Strong (29.9%) by enabling the student to learn robust representations under strong augmentation while receiving cleaner pseudo-labels from the weakly augmented teacher.
    2. Graph Distillation Loss (GDL): Adding GDL increases performance from 34.3% to 35.9% mAP (+1.6% mAP) by aligning original and graph-aggregated logits across the teacher and student networks.
    3. Graph Contrastive Loss (GCL): Adding IRG-guided GCL further improves performance from 35.9% to 37.1% mAP (+1.2% mAP) by reinforcing instance discrimination across proposals on target data.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Cai, H., Zheng, V.W., Chang, K.C.C.: A comprehensive survey of graph embedding: Problems, techniques, and applications. IEEE Transactions on Knowledge and Data Engineering 30(9), 1616–1637 (2018) 5
  2. 2.Cai, Q., Pan, Y., Ngo, C.W., Tian, X., Duan, L., Yao, T.: Exploring object relation in mean teacher for cross-domain detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 11457–11466 (2019) 3, 6
  3. 3.Chen, C., Zheng, Z., Ding, X., Huang, Y., Dou, Q.: Harmonizing transferability and discriminability for adapting object detectors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8869–8878 (2020) 7
  4. 4.Chen, C., Zheng, Z., Huang, Y., Ding, X., Yu, Y.: I3net: Implicit instance-invariant network for adapting one-stage object detectors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 12576–12585 (2021) 7
  5. 5.Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. In: International conference on machine learning. pp. 1597–1607. PMLR (2020) 2, 3, 4
  6. 6.Chen, T., Kornblith, S., Swersky, K., Norouzi, M., Hinton, G.E.: Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems 33, 22243–22255 (2020) 2, 4
  7. 7.Chen, Y.H., Chen, W.Y., Chen, Y.T., Tsai, B.C., Frank Wang, Y.C., Sun, M.: No more discrimination: Cross city adaptation of road scene segmenters. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1992–2001 (2017) 1
  8. 8.Chen, Y., Li, W., Sakaridis, C., Dai, D., Gool, L.V.: Domain adaptive faster r-cnn for object detection in the wild. 2018 IEEE Conference on Computer Vision and Pattern Recognition pp. 3339–3348 (2018) 3, 6, 7, 8
  9. 9.Chen, Y., Li, W., Sakaridis, C., Dai, D., Van Gool, L.: Domain adaptive faster r-cnn for object detection in the wild. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3339–3348 (2018) 1, 3
  10. 10.Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: The cityscapes dataset for semantic urban scene understanding. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3213–3223 (2016) 1, 6, 7
  11. 11.Deng, J., Li, W., Chen, Y., Duan, L.: Unbiased mean teacher for cross-domain object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4091–4101 (2021) 2, 4, 6, 7
  12. 12.Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., Tian, Q.: Centernet: Keypoint triplets for object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 6569–6578 (2019) 1
  13. 13.Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. International journal of computer vision 88(2), 303–338 (2010) 1, 7
  14. 14.Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., Lempitsky, V.: Domain-adversarial training of neural networks. The Journal of Machine Learning Research 17(1), 2096–2030 (2016) 1, 7, 8
  15. 15.Geiger, A., Lenz, P., Stiller, C., Urtasun, R.: Vision meets robotics: The kitti dataset. The International Journal of Robotics Research 32(11), 1231–1237 (2013) 1, 7
  16. 16.Gori, M., Monfardini, G., Scarselli, F.: A new model for learning in graph domains. In: Proceedings. 2005 IEEE International Joint Conference on Neural Networks, 2005. vol. 2, pp. 729–734. IEEE (2005) 3
  17. 17.He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.: Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9729–9738 (2020) 3
  18. 18.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016) 6
  19. 19.He, Z., Zhang, L.: Multi-adversarial faster-rcnn for unrestricted object detection. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 6668–6677 (2019) 1, 3, 7
  20. 20.He, Z., Zhang, L.: Domain adaptive object detection via asymmetric tri-way faster-rcnn. In: Proceedings of the European Conference on Computer Vision (2020) 7
  21. 21.Hegde, D., Patel, V.: Attentive prototypes for source-free unsupervised domain adaptive 3d object detection. arXiv preprint arXiv:2111.15656 (2021) 3
  22. 22.Hegde, D., Sindagi, V., Kilic, V., Cooper, A.B., Foster, M., Patel, V.: Uncertainty-aware mean teacher for source-free unsupervised domain adaptive 3d object detection. arXiv preprint arXiv:2109.14651 (2021) 3
  23. 23.Hoffman, J., Tzeng, E., Park, T., Zhu, J.Y., Isola, P., Saenko, K., Efros, A., Darrell, T.: Cycada: Cycle-consistent adversarial domain adaptation. In: International Conference on Machine Learning. pp. 1989–1998 (2018) 1
  24. 24.Hoffman, J., Wang, D., Yu, F., Darrell, T.: Fcns in the wild: Pixel-level adversarial and constraint-based adaptation. arXiv preprint arXiv:1612.02649 (2016) 1, 3
  25. 25.Hsu, C.C., Tsai, Y.H., Lin, Y.Y., Yang, M.H.: Every pixel matters: Center-aware feature alignment for domain adaptive object detector. In: European Conference on Computer Vision. pp. 733–748. Springer (2020) 3
  26. 26.Hsu, H.K., Yao, C.H., Tsai, Y.H., Hung, W.C., Tseng, H.Y., Singh, M., Yang, M.H.: Progressive domain adaptation for object detection. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 749–757 (2020) 6
  27. 27.Huang, J., Guan, D., Xiao, A., Lu, S.: Model adaptation: Historical contrastive learning for unsupervised domain adaptation without source data. Advances in Neural Information Processing Systems 34, 3635–3649 (2021) 1, 2, 3, 6
  28. 28.Inoue, N., Furuta, R., Yamasaki, T., Aizawa, K.: Cross-domain weakly-supervised object detection through progressive domain adaptation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 5001–5009 (2018) 1, 7, 8
  29. 29.Johnson-Roberson, M., Barto, C., Mehta, R., Sridhar, S.N., Rosaen, K., Vasudevan, R.: Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks? arXiv preprint arXiv:1610.01983 (2016) 7
  30. 30.Khodabandeh, M., Vahdat, A., Ranjbar, M., Macready, W.G.: A robust learning approach to domain adaptive object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 480–490 (2019) 3, 7
  31. 31.Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., Krishnan, D.: Supervised contrastive learning. Advances in neural information processing systems 33, 18661–18673 (2020) 3
  32. 32.Kim, S., Choi, J., Kim, T., Kim, C.: Self-training and adversarial background regularization for unsupervised domain adaptive one-stage object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 6092–6101 (2019) 7
  33. 33.Kim, T., Jeong, M., Kim, S., Choi, S., Kim, C.: Diversify and match: A domain adaptive representation learning paradigm for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 12456–12465 (2019) 6
  34. 34.Kim, Y., Cho, D., Han, K., Panda, P., Hong, S.: Domain adaptation without source data. IEEE Transactions on Artificial Intelligence (2021) 3, 6
  35. 35.Kipf, T.N., Welling, M.: Semi-supervised classification with graph convolutional networks. In: International Conference on Learning Representations (2017), https://openreview.net/forum?id=SJU4ayYgl 3
  36. 36.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25, 1097–1105 (2012) 6
  37. 37.Kundu, J.N., Venkat, N., Babu, R.V., et al.: Universal source-free domain adaptation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4544–4553 (2020) 1
  38. 38.Li, R., Jiao, Q., Cao, W., Wong, H.S., Wu, S.: Model adaptation: Unsupervised domain adaptation without source data. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 9641–9650 (2020) 3
  39. 39.Li, S., Ye, M., Zhu, X., Zhou, L., Xiong, L.: Source-free object detection by learning to overlook domain style. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8014–8023 (2022) 6
  40. 40.Li, X., Chen, W., Xie, D., Yang, S., Yuan, P., Pu, S., Zhuang, Y.: A free lunch for unsupervised domain adaptive object detection without source data. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 35, pp. 8474–8481 (2021) 2, 3, 6, 7, 8
  41. 41.Liang, J., Hu, D., Feng, J.: Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In: International Conference on Machine Learning. pp. 6028–6039. PMLR (2020) 1, 3
  42. 42.Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollar, P.: Focal loss for dense object detection. In: Proceedings of the IEEE international conference on computer vision. pp. 2980–2988 (2017) 1
  43. 43.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740–755. Springer (2014) 1
  44. 44.Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., Pietikainen, M.: Deep learning for generic object detection: A survey. International journal of computer vision 128(2), 261–318 (2020) 1
  45. 45.Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C.: Ssd: Single shot multibox detector. In: European conference on computer vision. pp. 21–37. Springer (2016) 1
  46. 46.Liu, Y.C., Ma, C.Y., He, Z., Kuo, C.W., Chen, K., Zhang, P., Wu, B., Kira, Z., Vajda, P.: Unbiased teacher for semi-supervised object detection. In: International Conference on Learning Representations (2021), https://openreview.net/forum?id=MJIve1zgR_ 2, 4
  47. 47.Liu, Y., Zhang, W., Wang, J.: Source-free domain adaptation for semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 1215–1224 (2021) 1, 3
  48. 48.Lo, S.Y., Oza, P., Chennupati, S., Galindo, A., Patel, V.M.: Spatio-temporal pixel-level contrastive learning-based source-free domain adaptation for video semantic segmentation. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023) 3
  49. 49.Morris, C., Ritzert, M., Fey, M., Hamilton, W.L., Lenssen, J.E., Rattan, G., Grohe, M.: Weisfeiler and leman go neural: Higher-order graph neural networks. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 33, pp. 4602–4609 (2019) 3
  50. 50.Nair, N.G., Patel, V.M.: Confidence guided network for atmospheric turbulence mitigation. In: 2021 IEEE International Conference on Image Processing (ICIP). pp. 1359–1363. IEEE (2021) 3
  51. 51.Nguyen, K., Tripathi, S., Du, B., Guha, T., Nguyen, T.Q.: In defense of scene graphs for image captioning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 1407–1416 (2021) 3
  52. 52.Oord, A.v.d., Li, Y., Vinyals, O.: Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2018) 3
  53. 53.Redmon, J., Divvala, S., Girshick, R., Farhadi, A.: You only look once: Unified, real-time object detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 779–788 (2016) 1
  54. 54.Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: Advances in neural information processing systems. pp. 91–99 (2015) 2, 6
  55. 55.RoyChowdhury, A., Chakrabarty, P., Singh, A., Jin, S., Jiang, H., Cao, L., Learned-Miller, E.: Automatic adaptation of object detectors to new domains using self-training. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 780–790 (2019) 3
  56. 56.Saito, K., Ushiku, Y., Harada, T., Saenko, K.: Strong-weak distribution alignment for adaptive object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6956–6965 (2019) 1, 3, 6, 7, 8
  57. 57.Saito, K., Watanabe, K., Ushiku, Y., Harada, T.: Maximum classifier discrepancy for unsupervised domain adaptation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3723–3732 (2018) 1
  58. 58.Sakaridis, C., Dai, D., Gool, L.V.: Semantic foggy scene understanding with synthetic data. International Journal of Computer Vision 126, 973–992 (2018) 6
  59. 59.Sindagi, V.A., nad R. Yasarla, P.O., Patel, V.M.: Prior-based domain adaptive object detection for hazy and rainy conditions. In: European Conference on Computer Vision (ECCV) (2020) 1, 3
  60. 60.Su, P., Wang, K., Zeng, X., Tang, S., Chen, D., Qiu, D., Wang, X.: Adapting object detectors with conditional domain normalization. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16. pp. 403–419. Springer (2020) 6
  61. 61.Tarvainen, A., Valpola, H.: Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780 (2017) 2, 4, 6, 7, 8
  62. 62.Tzeng, E., Hoffman, J., Saenko, K., Darrell, T.: Adversarial discriminative domain adaptation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 7167–7176 (2017) 1, 3
  63. 63.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30 (2017) 5
  64. 64.VS, V., Oza, P., Patel, V.M.: Towards online domain adaptive object detection. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 478–488 (2023) 3
  65. 65.VS, V., Oza, P., Sindagi, V.A., Gupta, V., Patel, V.M.: Mega-cda: Memory guided attention for category-aware unsupervised domain adaptive object detection (2021) 6, 7
  66. 66.Vs, V., Poster, D., You, S., Hu, S., Patel, V.M.: Meta-uda: Unsupervised domain adaptive thermal object detection using meta-learning. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 1412–1423 (2022) 3
  67. 67.Wu, A., Han, Y., Zhu, L., Yang, Y.: Instance-invariant domain adaptive object detection via progressive disentanglement. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021) 3, 6
  68. 68.Wu, F., Souza, A., Zhang, T., Fifty, C., Yu, T., Weinberger, K.: Simplifying graph convolutional networks. In: International conference on machine learning. pp. 6861–6871. PMLR (2019) 5
  69. 69.Wu, Z., Xiong, Y., Yu, S.X., Lin, D.: Unsupervised feature learning via non-parametric instance discrimination. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3733–3742 (2018) 2
  70. 70.Xia, H., Zhao, H., Ding, Z.: Adaptive adversarial network for source-free domain adaptation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 9010–9019 (2021) 1
  71. 71.Xu, C.D., Zhao, X.R., Jin, X., Wei, X.S.: Exploring categorical regularization for domain adaptive object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11724–11733 (2020) 6
  72. 72.Yang, J., Lu, J., Lee, S., Batra, D., Parikh, D.: Graph r-cnn for scene graph generation. In: Proceedings of the European conference on computer vision (ECCV). pp. 670–685 (2018) 3
  73. 73.Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., Darrell, T.: Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 2636–2645 (2020) 1
  74. 74.Zhang, Z., Cui, P., Zhu, W.: Deep learning on graphs: A survey. IEEE Transactions on Knowledge and Data Engineering (2020) 5
  75. 75.Zhao, G., Li, G., Xu, R., Lin, L.: Collaborative training between region proposal localization and classification for domain adaptive object detection. In: European Conference on Computer Vision. pp. 86–102. Springer (2020) 6, 7
  76. 76.Zhong, Y., Wang, L., Chen, J., Yu, D., Li, Y.: Comprehensive image captioning via scene graph decomposition. In: European Conference on Computer Vision. pp. 211–229. Springer (2020) 3
  77. 77.Zhu, X., Pang, J., Yang, C., Shi, J., Lin, D.: Adapting object detectors via selective cross-domain alignment. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 687–696 (2019) 7
  78. 78.Zhuang, C., Han, X., Huang, W., Scott, M.: ifan: Image-instance full alignment networks for adaptive object detection. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 13122–13129 (2020) 6
  79. 79.Zou, Z., Shi, Z., Guo, Y., Ye, J.: Object detection in 20 years: A survey. arXiv preprint arXiv:1905.05055 (2019) 1

Citation

MLA
VS, V., et al. “Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection”. arXiv, 2022, http://arxiv.org/abs/2203.15793v4.
APA
VS, V., Oza, P., & Patel, V. M. (2022). Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection. arXiv. http://arxiv.org/abs/2203.15793v4
Chicago
VS, V., P. Oza, and V. M. Patel. 2022. “Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection”. arXiv. http://arxiv.org/abs/2203.15793v4.
Harvard
VS, V., Oza, P. and Patel, V.M. (2022) “Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.15793v4.
Vancouver
1. VS V, Oza P, Patel VM (2022) Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection. arXiv

BibTeX

@article{vs2022instance,
  title = {Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection},
  author = {VS, Vibashan and Oza, Poojan and Patel, Vishal M.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.15793v4},
  eprint = {2203.15793}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE