Relation Networks for Object Detection

Han HuJiayuan GuZheng ZhangJifeng DaiYichen Wei

article2017CVPR1,362 citations

Proposes a lightweight relation module that jointly models geometric and appearance interactions between objects, establishing the first fully end-to-end object detector with integrated duplicate removal.

Listen

Modern computer vision systems rely heavily on deep convolutional neural networks to detect objects within images. However, conventional detection pipelines process candidate objects individually and rely on hand-crafted, sequential post-processing rules—such as non-maximum suppression—to eliminate duplicate detections. This individual evaluation ignores contextual visual and spatial relationships between co-occurring objects, creating a major performance bottleneck and preventing computer vision models from being trained as fully integrated, end-to-end systems.

The article demonstrates that explicitly modeling the relationships between candidate objects significantly enhances both instance recognition accuracy and duplicate removal. The primary objective is to evaluate whether an adapted attention-based module can jointly process multiple objects simultaneously and enable the first fully end-to-end object detector without requiring additional supervision.

To achieve this, the authors develop a lightweight object relation module inspired by attention mechanisms in natural language processing. The approach introduces a novel geometric weight that accounts for the relative spatial position and scale of bounding boxes alongside standard visual appearance features. The evaluation was conducted on the standard COCO benchmark dataset across 80 object categories, integrating the module into several top-performing detection architectures, including Faster R-CNN, Feature Pyramid Networks, and Deformable Convolutional Networks.

The experimental findings show clear performance gains across multiple benchmarks. Incorporating the relation module into the recognition head improved mean Average Precision by 2.3 to 3.2 points over standard baseline architectures, outperforming standard methods that merely increase network depth or width. When applied to duplicate removal, the module replaced traditional non-maximum suppression, achieving a 30.5 mean Average Precision compared to 29.6 for standard non-maximum suppression and 30.2 for Soft-NMS. Combining joint object recognition and learned duplicate removal in an end-to-end training setup increased overall accuracy to 31.0 on standard Faster R-CNN and achieved up to 39.0 on advanced architectures, adding less than 2% to 8% in computational overhead.

These findings prove that relationship modeling between objects is highly effective in modern deep learning architectures. By eliminating heuristic, hand-tuned post-processing steps, organizations can deploy unified, fully differentiable vision systems that deliver higher detection accuracy with negligible computational and operational overhead.

Engineering teams building or deploying region-based object detection systems should integrate relation modules directly into their recognition heads and replace hand-tuned post-processing steps with learned duplicate removal. Furthermore, the source suggests exploring the extension of this relation framework to adjacent computer vision tasks, including instance segmentation, action recognition, and visual question answering.

The results provide high confidence for standard region-based detection architectures on standard benchmarks. However, the approach has notable limitations when applied to dense sliding-window detection models, where processing extremely large numbers of candidate boxes becomes computationally expensive. Additionally, a detailed theoretical understanding of precisely what specific contextual relationships are learned across multi-layer relation networks remains preliminary and requires further investigation.

arXiv: 1711.11575
Cover for Relation Networks for Object Detection

Abstract

Although it is well believed for years that modeling relations between objects would help object recognition, there has not been evidence that the idea is working in the deep learning era. All state-of-the-art object detection systems still rely on recognizing object instances individually, without exploiting their relations during learning.

This work proposes an object relation module. It processes a set of objects simultaneously through interaction between their appearance feature and geometry, thus allowing modeling of their relations. It is lightweight and in-place. It does not require additional supervision and is easy to embed in existing networks. It is shown effective on improving object recognition and duplicate removal steps in the modern object detection pipeline. It verifies the efficacy of modeling object relations in CNN based detection. It gives rise to the first fully end-to-end object detector.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Object Relation Module
  • 4 Relation Networks For Object Detection
  • 4.1 Review of Object Detection Pipeline
  • 4.2 Relation for Instance Recognition
  • 4.3 Relation for Duplicate Removal
  • 4.4 End-to-End Object Detection
  • 5 Experiments
  • 5.1 Relation for Instance Recognition
  • 5.2 Relation for Duplicate Removal
  • 5.3 End-to-End Object Detection
  • 6 Conclusions
  • A1 Training Details
  • References

Knowls

  1. Knowl 1 — Object Relation Module Architecture

    model/method

    The object relation module adapts the self-attention mechanism to model spatial and appearance interactions among a set of NN object proposals simultaneously. Given an input set of NN objects, where the nn-th object has an appearance feature fAn∈Rdff_A^n \in \mathbb{R}^{d_f} and a 4D bounding box geometry fGn=(xn,yn,wn,hn)f_G^n = (x_n, y_n, w_n, h_n), the module computes an aggregated relation feature fR(n)f_R(n) for object nn over all other objects m∈{1,…,N}m \in \{1, \dots, N\}:

    fR(n)=∑m=1Nωmn⋅(WVfAm)f_R(n) = \sum_{m=1}^{N} \omega^{mn} \cdot (W_V f_A^m)

    where WV∈Rdv×dfW_V \in \mathbb{R}^{d_v \times d_f} is a learned linear projection matrix for values (dv=df/Nrd_v = d_f / N_r), and ωmn\omega^{mn} denotes the relation weight indicating the impact of object mm on object nn. The relation weight combines appearance similarity ωAmn\omega_A^{mn} with a geometric interaction weight ωGmn\omega_G^{mn}:

    ωmn=ωGmnexp⁡(ωAmn)∑k=1NωGknexp⁡(ωAkn)\omega^{mn} = \frac{\omega_G^{mn} \exp(\omega_A^{mn})}{\sum_{k=1}^{N} \omega_G^{kn} \exp(\omega_A^{kn})}

    The appearance weight ωAmn\omega_A^{mn} measures feature compatibility in a dkd_k-dimensional subspace via dot product:

    ωAmn=dot(WKfAm,WQfAn)dk\omega_A^{mn} = \frac{\text{dot}(W_K f_A^m, W_Q f_A^n)}{\sqrt{d_k}}

    where WK,WQ∈Rdk×dfW_K, W_Q \in \mathbb{R}^{d_k \times d_f} project input appearance features into key and query representations, and dk\sqrt{d_k} is a scaling factor.

    To capture diverse relationships across NrN_r distinct relation heads, the module computes NrN_r relation features {fRr(n)}r=1Nr\{f_R^r(n)\}_{r=1}^{N_r} in parallel, concatenates them, and adds the concatenated vector to the original input appearance feature:

    fAn←fAn+Concat[fR1(n),fR2(n),…,fRNr(n)]for all n∈{1,…,N}f_A^n \leftarrow f_A^n + \text{Concat}[f_R^1(n), f_R^2(n), \dots, f_R^{N_r}(n)] \quad \text{for all } n \in \{1, \dots, N\}

    Because the output feature has the identical channel dimension dfd_f as the input, the relation module operates in-place as a drop-in building block for neural detection architectures.

  2. Knowl 2 — Relative Geometry Embedding and Geometric Attention Weight Formulation

    equation

    The geometric interaction weight ωGmn\omega_G^{mn} between object mm with bounding box fGm=(xm,ym,wm,hm)f_G^m = (x_m, y_m, w_m, h_m) and object nn with bounding box fGn=(xn,yn,wn,hn)f_G^n = (x_n, y_n, w_n, h_n) is computed via a 4-dimensional relative spatial coordinate representation:

    E4D(fGm,fGn)=(log⁡(∣xm−xn∣wm),log⁡(∣ym−yn∣hm),log⁡(wnwm),log⁡(hnhm))T\mathcal{E}_{4D}(f_G^m, f_G^n) = \left( \log\left(\frac{|x_m - x_n|}{w_m}\right), \log\left(\frac{|y_m - y_n|}{h_m}\right), \log\left(\frac{w_n}{w_m}\right), \log\left(\frac{h_n}{h_m}\right) \right)^T

    Using logarithmic relative spatial offsets and scale ratios guarantees translation and scale invariance while ensuring non-zero gradients for distant objects.

    This 4D vector is mapped to a high-dimensional feature EG(fGm,fGn)∈Rdg\mathcal{E}_G(f_G^m, f_G^n) \in \mathbb{R}^{d_g} by applying sine and cosine functions across different predefined wavelengths. The resulting geometric embedding is projected to a scalar weight and clipped at zero:

    ωGmn=max⁡{0,WG⋅EG(fGm,fGn)}\omega_G^{mn} = \max\{0, W_G \cdot \mathcal{E}_G(f_G^m, f_G^n)\}

    where WG∈R1×dgW_G \in \mathbb{R}^{1 \times d_g} is a learned projection matrix. The zero-trimming operation acts as a ReLU non-linearity, enforcing geometric sparsity by eliminating relations between object pairs that do not satisfy required relative spatial configurations.

  3. Knowl 3 — Object Relation Module Algorithm and Computational Complexity

    algorithm

    The object relation module transforms NN input object appearance features {fAn}n=1N⊂Rdf\{f_A^n\}_{n=1}^N \subset \mathbb{R}^{d_f} and 4D geometry bounding boxes {fGn}n=1N\{f_G^n\}_{n=1}^N across NrN_r parallel relation heads.

    Input: N object appearance and geometry pairs {(f_A^n, f_G^n)}_{n=1}^N, input feature dimension d_f
    Hyperparameters: Number of relations N_r, key dimension d_k, geometry dimension d_g
    Learned weights: {W_K^r, W_Q^r, W_V^r, W_G^r}_{r=1}^{N_r}
    for each object index n in {1, ..., N} and relation head r in {1, ..., N_r} do
        for each object index m in {1, ..., N} do
            Compute 4D relative geometry E_4D(f_G^m, f_G^n)
            Embed E_4D to d_g-dimensional representation E_G(f_G^m, f_G^n)
            Compute geometric weight omega_G^{mn, r} = max(0, W_G^r * E_G(f_G^m, f_G^n))
            Compute appearance weight omega_A^{mn, r} = dot(W_K^r * f_A^m, W_Q^r * f_A^n) / sqrt(d_k)
        end for
        for each object index m in {1, ..., N} do
            Compute normalized relation weight omega^{mn, r} = (omega_G^{mn, r} * exp(omega_A^{mn, r})) / sum_{k=1}^N (omega_G^{kn, r} * exp(omega_A^{kn, r}))
        end for
        Compute relation feature f_R^r(n) = sum_{m=1}^N omega^{mn, r} * (W_V^r * f_A^m)
    end for
    for each object index n in {1, ..., N} do
        Compute output feature f_A^n = f_A^n + Concat[f_R^1(n), ..., f_R^{N_r}(n)]
    end for
    return Updated appearance features {f_A^n}_{n=1}^N

    The parameter complexity of one relation module is: O(Space)=Nr(2dfdk+dg)+df2\mathcal{O}(\text{Space}) = N_r(2 d_f d_k + d_g) + d_f^2

    The computational FLOP complexity per image is: O(Comp.)=Ndf(2Nrdk+df)+N2Nr(dg+dk+dfNr+1)\mathcal{O}(\text{Comp.}) = N d_f (2 N_r d_k + d_f) + N^2 N_r \left(d_g + d_k + \frac{d_f}{N_r} + 1\right)

    For default parameters Nr=16N_r = 16, dk=64d_k = 64, dg=64d_g = 64, and df=1024d_f = 1024, one relation module contains approximately 3 million parameters and executes 1.2 billion FLOPs for N=300N = 300 proposals.

  4. Knowl 4 — Relation-Enhanced Instance Recognition Head (2fc+RM)

    model/method

    In standard region-based object detectors (such as Faster R-CNN, FPN, and Deformable ConvNets), region-of-interest (RoI) pooled features undergo two fully connected layers (each of dimension 1024) before linear proposal classification and bounding box regression:

    {RoI_Featn}n=1N→FC1024⋅N→FC1024⋅N→Linear{(scoren,bboxn)}n=1N\{\text{RoI\_Feat}_n\}_{n=1}^N \xrightarrow{\text{FC}} 1024 \cdot N \xrightarrow{\text{FC}} 1024 \cdot N \xrightarrow{\text{Linear}} \{(\text{score}_n, \text{bbox}_n)\}_{n=1}^N

    The relation-enhanced head (2fc+RM) embeds r1r_1 relation modules after the first 1024-d FC layer and r2r_2 relation modules after the second 1024-d FC layer:

    {RoI_Featn}n=1N→FC1024⋅N→{RM}r11024⋅N→FC1024⋅N→{RM}r21024⋅N→Linear{(scoren,bboxn)}n=1N\{\text{RoI\_Feat}_n\}_{n=1}^N \xrightarrow{\text{FC}} 1024 \cdot N \xrightarrow{\{\text{RM}\}^{r_1}} 1024 \cdot N \xrightarrow{\text{FC}} 1024 \cdot N \xrightarrow{\{\text{RM}\}^{r_2}} 1024 \cdot N \xrightarrow{\text{Linear}} \{(\text{score}_n, \text{bbox}_n)\}_{n=1}^N

    Because the relation module is in-place and preserves the 1024 feature dimension, multiple modules can be stacked without altering the interface to subsequent classification and bounding box regression linear layers. Default parameter values are r1=1,r2=1r_1 = 1, r_2 = 1.

  5. Knowl 5 — Learnable Duplicate Removal Network via Object Relations

    model/method

    The duplicate removal network replaces heuristic non-maximum suppression (NMS) by framing duplicate removal as a learned binary classification task across a set of detected candidates.

    The network accepts NN detected objects output by the instance recognition head. Each object has an appearance feature vector fA∈R1024f_A \in \mathbb{R}^{1024}, a classification score s0∈[0,1]s_0 \in [0, 1], and a bounding box geometry fGf_G.

    1. Rank Feature Embedding: Input detections are sorted descending by s0s_0, assigning each candidate an integer rank k∈{1,…,N}k \in \{1, \dots, N\}. The scalar rank is embedded into a 128-dimensional continuous vector using sinusoidal wavelength encoding.
    2. Feature Fusion: The 1024-d appearance feature is projected to 128-d via learned linear weights WfW_f. The 128-d rank embedding is transformed by learned weights WfRW_{fR}. The two 128-d vectors are summed to construct the input appearance feature for the relation module (df=128d_f = 128).
    3. Relation Transformation: A relation module with Nr=16,dk=64,dg=64,df=128N_r = 16, d_k = 64, d_g = 64, d_f = 128 processes the N=100N = 100 highest-scoring candidate objects.
    4. Duplicate Probability Output: A linear classifier WsW_s followed by a sigmoid non-linearity predicts a binary correctness probability s1∈[0,1]s_1 \in [0, 1] for each candidate (s1=1s_1 = 1 for a true positive, s1=0s_1 = 0 for a duplicate detection).

    The final object detection score is defined multiplicatively as:

    sfinal=s0⋅s1s_{\text{final}} = s_0 \cdot s_1

    A valid detection must have both a high instance recognition confidence s0s_0 and a high duplicate removal confidence s1s_1.

  6. Knowl 6 — Multi-Threshold Loss and End-to-End Object Detection Training

    model/method

    In the duplicate removal step, candidate detections with Intersection over Union (IoU) ≥η\ge \eta relative to a ground-truth object are matched to it; the highest-scoring candidate is labeled positive (y=1y = 1), while all other matched candidates are labeled duplicates (y=0y = 0).

    To optimize performance across the COCO evaluation metric (mAP@[0.5:0.95]), the duplicate removal classifier WsW_s outputs multiple binary probabilities corresponding to five IoU thresholds simultaneously: η∈{0.5,0.6,0.7,0.8,0.9}\eta \in \{0.5, 0.6, 0.7, 0.8, 0.9\}. The binary cross-entropy loss is applied directly to the combined score s0s1s_0 s_1 across all classes and boxes B\mathcal{B}:

    Ldup=−1∣B∣∑i∈B[yilog⁡(s0(i)s1(i))+(1−yi)log⁡(1−s0(i)s1(i))]L_{\text{dup}} = -\frac{1}{|\mathcal{B}|} \sum_{i \in \mathcal{B}} \left[ y_i \log(s_0^{(i)} s_1^{(i)}) + (1 - y_i) \log(1 - s_0^{(i)} s_1^{(i)}) \right]

    Due to the multiplicative form s0s1s_0 s_1, gradients on non-duplicate background boxes with low s0s_0 (where s0<0.01s_0 < 0.01) remain near zero (∂L∂s1=s01−s0s1\frac{\partial L}{\partial s_1} = \frac{s_0}{1 - s_0 s_1}), naturally focusing learning on true duplicate candidates.

    In end-to-end training, the entire detector is jointly trained with combined loss:

    Ltotal=LRPN+Linst+LdupL_{\text{total}} = L_{\text{RPN}} + L_{\text{inst}} + L_{\text{dup}}

    where LRPNL_{\text{RPN}} is the region proposal loss, LinstL_{\text{inst}} comprises instance classification cross-entropy and bounding box regression smooth L1 losses, and LdupL_{\text{dup}} is the multi-threshold duplicate removal loss. During inference, the predicted duplicate probabilities across the five thresholds are averaged into a single score s1s_1.

  7. Knowl 7 — Ablation of Relation Module Hyperparameters and Geometric Representations

    empirical result

    Ablation experiments on MS COCO minival using a Faster R-CNN detector with a ResNet-50 backbone establish the contributions of the geometric attention formulation, number of relation heads NrN_r, and module repetitions {r1,r2}\{r_1, r_2\} (baseline 2fc head achieves 29.6 mAP):

    Setting Geometric Feature Usage Number of Relations NrN_r Number of Modules {r1,r2}\{r_1, r_2\}
    none unary ours 1 2 4 8 16 32 {1,0} {0,1} {1,1} {2,2} {4,4}
    mAP 30.3 31.1 31.9 30.5 30.6 31.3 31.7 31.9 31.7 31.7 31.4 31.9 32.5 32.8
    1. Geometric Attention: Incorporating the relative 2D geometric weight (ours, 31.9 mAP) outperforms omitting geometry entirely (none, 30.3 mAP) and unary embedding addition (unary, 31.1 mAP).
    2. Head Count NrN_r: Performance increases monotonically from Nr=1N_r=1 (30.5 mAP) to Nr=16N_r=16 (31.9 mAP), plateauing at Nr=32N_r=32 (31.7 mAP).
    3. Module Depth: Increasing module repetitions from {1,1}\{1, 1\} to {2,2}\{2, 2\} and {4,4}\{4, 4\} improves mAP to 32.5 and 32.8, respectively.
  8. Knowl 8 — Comparison of Duplicate Removal Network Against NMS and SoftNMS

    data/table

    Evaluated on MS COCO minival using Faster R-CNN with a ResNet-50 backbone (where the baseline instance recognition head achieves 29.6 mAP under standard NMS with IoU threshold Nt=0.6N_t = 0.6), the learnable duplicate removal network outperforms both greedy NMS and SoftNMS across multiple localization thresholds:

    Method Parameters / Settings mAP mAP50\text{mAP}_{50} mAP75\text{mAP}_{75}
    NMS Nt=0.3N_t = 0.3 29.0 51.4 29.4
    NMS Nt=0.4N_t = 0.4 29.4 52.1 29.5
    NMS Nt=0.5N_t = 0.5 29.6 51.9 29.7
    NMS Nt=0.6N_t = 0.6 29.6 50.9 30.1
    NMS Nt=0.7N_t = 0.7 28.4 46.6 30.7
    SoftNMS σ=0.2\sigma = 0.2 30.0 52.3 30.5
    SoftNMS σ=0.4\sigma = 0.4 30.2 51.7 31.3
    SoftNMS σ=0.6\sigma = 0.6 30.2 50.9 31.6
    SoftNMS σ=0.8\sigma = 0.8 29.9 49.9 31.6
    SoftNMS σ=1.0\sigma = 1.0 29.7 49.7 31.6
    Relation Network (standalone) η=0.5\eta = 0.5 30.3 51.9 31.5
    Relation Network (standalone) η=0.75\eta = 0.75 30.1 49.0 32.7
    Relation Network (standalone) η∈[0.5,0.9]\eta \in [0.5, 0.9] 30.5 50.2 32.4
    Relation Network (end-to-end) η∈[0.5,0.9]\eta \in [0.5, 0.9] 31.0 51.4 32.8

    Ablations on duplicate removal input features (evaluated at η=0.5\eta = 0.5, full model: 30.3 mAP) show that:

    • Removing the rank feature drops mAP to 26.6;
    • Using raw classification score s0s_0 instead of rank drops mAP to 28.3;
    • Removing the appearance feature drops mAP to 29.9;
    • Removing the relative geometry drops mAP to 28.1 (unary geometry yields 28.2).
  9. Knowl 9 — System-Level Detection Performance Across Architectures on MS COCO

    data/table

    System-level object detection results on MS COCO minival and test-dev across standard region-based detection frameworks (Faster R-CNN, Feature Pyramid Networks (FPN), and Deformable Convolutional Networks (DCN)) with ResNet-101 backbones and Online Hard Example Mining (OHEM). Transitions represent: baseline 2fc head + SoftNMS (σ=0.6\sigma = 0.6) →\rightarrow 2fc+RM head + SoftNMS →\rightarrow 2fc+RM head + end-to-end duplicate removal.

    Detector Test Set mAP mAP50\text{mAP}_{50} mAP75\text{mAP}_{75} # Params FLOPs
    Faster R-CNN minival 32.2→34.7→35.232.2 \rightarrow 34.7 \rightarrow 35.2 52.9→55.3→55.852.9 \rightarrow 55.3 \rightarrow 55.8 34.2→37.2→38.234.2 \rightarrow 37.2 \rightarrow 38.2 58.3M→64.3M→64.6M58.3\text{M} \rightarrow 64.3\text{M} \rightarrow 64.6\text{M} 122.2B→124.6B→124.9B122.2\text{B} \rightarrow 124.6\text{B} \rightarrow 124.9\text{B}
    Faster R-CNN test-dev 32.7→35.2→35.432.7 \rightarrow 35.2 \rightarrow 35.4 53.6→56.2→56.153.6 \rightarrow 56.2 \rightarrow 56.1 34.7→37.8→38.534.7 \rightarrow 37.8 \rightarrow 38.5 - -
    FPN minival 36.8→38.1→38.836.8 \rightarrow 38.1 \rightarrow 38.8 57.8→59.5→60.357.8 \rightarrow 59.5 \rightarrow 60.3 40.7→41.8→42.940.7 \rightarrow 41.8 \rightarrow 42.9 56.4M→62.4M→62.8M56.4\text{M} \rightarrow 62.4\text{M} \rightarrow 62.8\text{M} 145.8B→157.8B→158.2B145.8\text{B} \rightarrow 157.8\text{B} \rightarrow 158.2\text{B}
    FPN test-dev 37.2→38.3→38.937.2 \rightarrow 38.3 \rightarrow 38.9 58.2→59.9→60.558.2 \rightarrow 59.9 \rightarrow 60.5 41.4→42.3→43.341.4 \rightarrow 42.3 \rightarrow 43.3 - -
    DCN minival 37.5→38.1→38.537.5 \rightarrow 38.1 \rightarrow 38.5 57.3→57.8→57.857.3 \rightarrow 57.8 \rightarrow 57.8 41.0→41.3→42.041.0 \rightarrow 41.3 \rightarrow 42.0 60.5M→66.5M→66.8M60.5\text{M} \rightarrow 66.5\text{M} \rightarrow 66.8\text{M} 125.0B→127.4B→127.7B125.0\text{B} \rightarrow 127.4\text{B} \rightarrow 127.7\text{B}
    DCN test-dev 38.1→38.8→39.038.1 \rightarrow 38.8 \rightarrow 39.0 58.1→58.7→58.658.1 \rightarrow 58.7 \rightarrow 58.6 41.6→42.4→42.941.6 \rightarrow 42.4 \rightarrow 42.9 - -

    Adding the relation modules to the instance head yields consistent gains (+2.5 mAP on Faster R-CNN, +1.3 mAP on FPN, +0.6 mAP on DCN minival), and end-to-end duplicate removal training provides additional improvements of +0.4 to +0.7 mAP across all models with minimal computational overhead.

  10. Knowl 10 — Dense Detection Paradigm Limitation and Attention Interpretability

    limitation

    The object relation module is subject to two main limitations:

    1. Inapplicability to Dense Sliding-Window Detectors: The computation and memory complexity of the relation module scale quadratically with the number of proposals NN (specifically O(N2Nr(dg+dk+df/Nr+1))\mathcal{O}(N^2 N_r (d_g + d_k + d_f/N_r + 1))). While this overhead is low for sparse proposal sets (N≈100−300N \approx 100 - 300), it becomes computationally prohibitive for dense sliding-window detector architectures (such as SSD or YOLO) where NN reaches tens of thousands of candidate locations.
    2. Interpretability of Stacked Relation Modules: While individual relation weights in single-module heads indicate intuitive object-object interactions (e.g., surrounding context aiding a central object or a person feature supporting glove detection), interpreting the representations learned across multiple stacked relation modules remains an unsolved challenge.

Coverage note — No substantial contributed material was omitted. Qualitative visual examples from Figure 4 were summarized inside the limitation knowl as context for interpretability.

References

  1. 1.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh. Vqa: Visual question answering. In ICCV, 2015.
  2. 2.S. Bell, C. Lawrence Zitnick, K. Bala, and R. Girshick. Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. In CVPR, pages 2874–2883, 2016.
  3. 3.I. Biederman, R. J. Mezzanotte, and J. C. Rabinowitz. Scene perception: Detecting and judging objects undergoing relational violations. Cognitive psychology, 14(2):143–177, 1982.
  4. 4.N. Bodla, B. Singh, R. Chellappa, and L. S. Davis. Softnms–improving object detection with one line of code. In CVPR, 2017.
  5. 5.D. Britz, A. Goldie, T. Luong, and Q. Le. Massive exploration of neural machine translation architectures. arXiv preprint arXiv:1703.03906, 2017.
  6. 6.X. Chen and A. Gupta. Spatial memory for context reasoning in object detection. In ICCV, 2017.
  7. 7.M. J. Choi, A. Torralba, and A. S. Willsky. A tree-based context model for object recognition. TPAMI, 34(2):240–252, Feb 2012.
  8. 8.F. Chollet. Xception: Deep learning with depthwise separable convolutions. In CVPR, 2016.
  9. 9.J. Dai, Y. Li, K. He, and J. Sun. R-fcn: Object detection via region-based fully convolutional networks. In NIPS, 2016.
  10. 10.J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei. Deformable convolutional networks. In ICCV, 2017.
  11. 11.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  12. 12.S. K. Divvala, D. Hoiem, J. H. Hays, A. A. Efros, and M. Hebert. An empirical study of context in object detection. In CVPR, 2009.
  13. 13.Y. Duan, M. Andrychowicz, B. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba. One-shot imitation learning. arXiv preprint arXiv:1703.07326, 2017.
  14. 14.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes (VOC) Challenge. IJCV, 2010.
  15. 15.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. TPAMI, 2010.
  16. 16.C. Galleguillos and S. Belongie. Context based object categorization: A critical survey. In CVPR, 2010.
  17. 17.C. Galleguillos, A. Rabinovich, and S. Belongie. Object categorization using co-occurrence, location and appearance. In CVPR, 2008.
  18. 18.R. Girshick. Fast R-CNN. In ICCV, 2015.
  19. 19.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, 2014.
  20. 20.G. Gkioxari, R. Girshick, and J. Malik. Contextual action recognition with r* cnn. In ICCV, pages 1080–1088, 2015.
  21. 21.G. Gkioxari, R. B. Girshick, P. Doll'ar, and K. He. Detecting and recognizing human-object interactions. CoRR, abs/1704.07333, 2017.
  22. 22.S. Gupta, B. Hariharan, and J. Malik. Exploring person context and local scene context for object detection. CoRR, abs/1511.08177, 2015.
  23. 23.K. He, G. Gkioxari, P. Doll'ar, and R. Girshick. Mask r-cnn. arXiv preprint arXiv:1703.06870, 2017.
  24. 24.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. In ECCV, 2014.
  25. 25.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  26. 26.J. Hosang, R. Benenson, and B. Schiele. Learning non-maximum suppression. In ICCV, 2017.
  27. 27.J. Huang, V. Rathod, C. Sun, M. Zhu, A. Korattikara, A. Fathi, I. Fischer, Z. Wojna, Y. Song, S. Guadarrama, and K. Murphy. Speed/accuracy trade-offs for modern convolutional object detectors. arXiv preprint arXiv:1611.10012, 2016.
  28. 28.R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 123(1):32–73, 2017.
  29. 29.J. Li, Y. Wei, X. Liang, J. Dong, T. Xu, J. Feng, and S. Yan. Attentive contexts for object detection. arXiv preprint arXiv:1603.07415, 2016.
  30. 30.Y. Li, H. Qi, J. Dai, X. Ji, and Y. Wei. Fully convolutional instance-aware semantic segmentation. arXiv preprint arXiv:1611.07709, 2016.
  31. 31.T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Doll'ar. Focal loss for dense object detection. arXiv preprint arXiv:1708.02002, 2017.
  32. 32.T.-Y. Lin, P. Doll'ar, R. Girshick, K. He, B. Hariharan, and S. Belongie. Feature pyramid networks for object detection. In CVPR, 2017.
  33. 33.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Doll'ar. Focal loss for dense object detection. arXiv preprint arXiv:1708.02002, 2017.
  34. 34.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll'ar, and C. L. Zitnick. Microsoft COCO: Common objects in context. In ECCV. 2014.
  35. 35.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, and S. Reed. Ssd: Single shot multibox detector. In ECCV, 2016.
  36. 36.R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille. The role of context for object detection and semantic segmentation in the wild. In CVPR. 2014.
  37. 37.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. You only look once: Unified, real-time object detection. In CVPR, 2016.
  38. 38.S. Ren, K. He, R. Girshick, and J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In NIPS, 2015.
  39. 39.J. Shotton, J. Winn, C. Rother, and A. Criminisi. Textonboost: Joint appearance, shape and context modeling for multi-class object recognition and segmentation. In ECCV, 2006.
  40. 40.A. Shrivastava, A. Gupta, and R. Girshick. Training region-based object detectors with online hard example mining. In CVPR, 2016.
  41. 41.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In Advances in neural information processing systems, pages 568–576, 2014.
  42. 42.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  43. 43.R. K. Srivastava, K. Greff, and J. Schmidhuber. Highway networks. arXiv preprint arXiv:1505.00387, 2015.
  44. 44.R. Stewart, M. Andriluka, and A. Y. Ng. End-to-end people detection in crowded scenes. In ICCV, 2016.
  45. 45.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In CVPR, 2015.
  46. 46.A. Torralba, K. P. Murphy, W. T. Freeman, and M. A. Rubin. Context-based vision system for place and object recognition. In ICCV, 2003.
  47. 47.Z. Tu. Auto-context and its application to high-level vision tasks. In CVPR, 2008.
  48. 48.J. R. Uijlings, K. E. van de Sande, T. Gevers, and A. W. Smeulders. Selective search for object recognition. IJCV, 2013.
  49. 49.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017.
  50. 50.K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015.
  51. 51.B. Yao and L. Fei-Fei. Recognizing human-object interactions in still images by modeling the mutual context of objects and human poses. TPAMI, 34(9):1691–1703, Sept 2012.
  52. 52.C. L. Zitnick and P. Doll'ar. Edge boxes: Locating object proposals from edges. In ECCV, 2014.
  53. 53.B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le. Learning transferable architectures for scalable image recognition. arXiv preprint arXiv:1707.07012, 2017.

Citation

MLA
Hu, H., et al. “Relation Networks for Object Detection”. arXiv, 2017, http://arxiv.org/abs/1711.11575v2.
APA
Hu, H., Gu, J., Zhang, Z., Dai, J., & Wei, Y. (2017). Relation Networks for Object Detection. arXiv. http://arxiv.org/abs/1711.11575v2
Chicago
Hu, H., J. Gu, Z. Zhang, J. Dai, and Y. Wei. 2017. “Relation Networks for Object Detection”. arXiv. http://arxiv.org/abs/1711.11575v2.
Harvard
Hu, H. et al. (2017) “Relation Networks for Object Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.11575v2.
Vancouver
1. Hu H, Gu J, Zhang Z, Dai J, Wei Y (2017) Relation Networks for Object Detection. arXiv

BibTeX

@article{hu2017relation,
  title = {Relation Networks for Object Detection},
  author = {Hu, Han and Gu, Jiayuan and Zhang, Zheng and Dai, Jifeng and Wei, Yichen},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.11575v2},
  eprint = {1711.11575}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE