Structured Sparse R-CNN for Direct Scene Graph Generation

Yao TengLimin Wang

article2022CVPR79 citations

Proposes an end-to-end scene graph generation framework that replaces traditional multi-stage pipelines with learnable triplet queries and cascaded dynamic heads to predict entities and relationships directly.

Listen

Scene graph generation identifies objects within images alongside their visual relationships, providing a structured understanding crucial for advanced applications like image captioning and visual question answering. Existing methods typically rely on multi-stage pipelines that detect objects first and then construct dense, fully connected graphs to classify relationships. This standard workflow introduces significant computational redundancy and fails to effectively capture the natural sparsity of visual relationships. The article introduces Structured Sparse R-CNN, a unified, single-stage framework that models scene graph generation as a direct set prediction task, eliminating the need for explicit separate object detection and dense graph construction during inference.

To evaluate this framework, the authors conducted comprehensive experiments on standard visual benchmarks, including Visual Genome and Open Images versions 4 and 6. The architecture processes images through a standard convolutional neural network and progressively refines a sparse set of learnable triplet queries using a structured triplet detector. These queries encode initial spatial and semantic assumptions about object pairs and their relationships. To address sparse training labels and class imbalance, the approach incorporates knowledge distillation from an auxiliary object detector during training, an adaptive focusing loss parameter, and post-processing adjustments for long-tailed category distributions.

Across all evaluated benchmarks, the framework achieved state-of-the-art detection performance while delivering major speed improvements. On the Visual Genome dataset, the model processed images in 0.19 seconds per image—roughly two to three times faster than prominent existing approaches, which range from 0.38 to 0.67 seconds per image. The framework also set new performance highs on zero-shot recall and mean recall metrics, demonstrating a superior capability to recognize rare and unseen visual relationships. Furthermore, evaluations on Open Images confirmed top-tier performance, with the method outperforming existing benchmarks on weighted mean average precision metrics.

These findings demonstrate that direct, sparse set prediction can simplify visual relationship modeling while simultaneously boosting accuracy and processing speed. By removing multi-stage pipelines and explicit dense graph construction, the framework significantly cuts down computational latency and operational overhead. This efficiency makes deep visual reasoning far more viable for real-time systems, edge devices, and large-scale automated image analysis pipelines, without sacrificing the ability to recognize nuanced or uncommon interactions.

Organizations developing complex computer vision systems should consider adopting direct sparse prediction architectures over traditional multi-stage pipelines to streamline their infrastructure. When deploying in environments with severe category imbalances, teams can pair the model with logit adjustment techniques to prioritize rare relationships, though they should weigh slight trade-offs in overall weighted precision. Future work should focus on extending this direct prediction paradigm to video understanding and improving pseudo-labeling techniques to further reduce noise during training.

The reported conclusions are well-supported across multiple large-scale benchmarks and rigorous ablation studies. However, some practical caveats apply. Directly training sparse triplet detectors without auxiliary supervision remains difficult due to sparse dataset annotations, making the knowledge distillation strategy a necessary training component. Additionally, the presence of noisy pseudo-labels can slightly impact certain high-recall metrics when post-processing filters are applied.

Cover for Structured Sparse R-CNN for Direct Scene Graph Generation

Abstract

Scene graph generation (SGG) is to detect object pairs with their relations in an image. Existing SGG approaches often use multi-stage pipelines to decompose this task into object detection, relation graph construction, and dense or dense-to-sparse relation prediction. Instead, from a perspective on SGG as a direct set prediction, this paper presents a simple, sparse, and unified framework, termed as Structured Sparse R-CNN. The key to our method is a set of learnable triplet queries and a structured triplet detector which could be jointly optimized from the training set in an end-to-end manner. Specifically, the triplet queries encode the general prior for object pairs with their relations, and provide an initial guess of scene graphs for subsequent refinement. The triplet detector presents a cascaded architecture to progressively refine the detected scene graphs with the customized dynamic heads. In addition, to relieve the training difficulty of our method, we propose a relaxed and enhanced training strategy based on knowledge distillation from a Siamese Sparse R-CNN. We perform experiments on several datasets: Visual Genome and Open Images V4/V6, and the results demonstrate that our method achieves the state-of-the-art performance. In addition, we also perform in-depth ablation studies to provide insights on our structured modeling in triplet detector design and training strategies. The code and models are made available at https://github.com/MCG-NJU/Structured-Sparse-RCNN.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Approach
  • 3.1. Structured Sparse R-CNN
  • 3.2. Learning with Siamese Sparse R-CNN
  • 3.3. Imbalance Class Distribution
  • 4. Experiments
  • 4.1. Datasets and Evaluation Settings
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Comparisons with the State of the Art
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Structured Sparse R-CNN Architecture and Triplet Query Representation

    model/method

    Structured Sparse R-CNN formulates scene graph generation (SGG) as a direct sparse set prediction problem without requiring explicit object proposal generation or preceding graph construction during inference.

    The framework takes image features extracted via a convolutional neural network (CNN) with a Feature Pyramid Network (FPN) and processes a set of NN learnable triplet queries through a cascade of MM modular Triplet Detection Heads (e.g., M=6M=6, N=300N=300 or N=800N=800). Each learnable triplet query represents the data-driven spatial and semantic prior of a visual triplet and is parameterized by:

    • Two 4-dimensional normalized proposal bounding boxes (specifying the center coordinates, width, and height for the subject and object: bs,bo∈R4b_s, b_o \in \mathbb{R}^4).
    • Two 1024-dimensional object content vectors (encoding appearance/semantics of the subject and object: Xs,Xo∈R1024X_s, X_o \in \mathbb{R}^{1024}).
    • One 256-dimensional relation content vector (encoding relationship semantics and structure: Xr∈R256X_r \in \mathbb{R}^{256}).

    Each Triplet Detection Head progressively updates the bounding boxes, object classification scores, and relation classification scores across MM stages using dynamic convolution and structured feature fusion modules.

  2. Knowl 2 — Pair Fusion (PF) Module for Object Pair Context Modeling

    model/method

    In the object pair detection stage of each Triplet Detection Head, the Pair Fusion (PF) module establishes intra-pair interaction between the subject and object content vectors before dynamic feature interaction and bounding box refinement.

    Given the subject content vector Xs∈R1024X_s \in \mathbb{R}^{1024}, object content vector Xo∈R1024X_o \in \mathbb{R}^{1024}, and their respective learnable positional encodings Ps,Po∈R1024P_s, P_o \in \mathbb{R}^{1024}, the pair representation XpX_p and enhanced content vectors Xs′,Xo′X'_s, X'_o are computed via:

    Xp=ReLU(LN(W0sXs+W0oXo))X_p = \text{ReLU}(\text{LN}(W_0^s X_s + W_0^o X_o))

    Xs′=Xs+W1sXp+PsX'_s = X_s + W_1^s X_p + P_s

    Xo′=Xo+W1oXp+PoX'_o = X_o + W_1^o X_p + P_o

    where W0s,W0o,W1s,W1oW_0^s, W_0^o, W_1^s, W_1^o are learnable weight matrices, LN(⋅)\text{LN}(\cdot) denotes layer normalization, and ReLU(⋅)\text{ReLU}(\cdot) denotes the rectified linear activation function.

    The enhanced vectors Xs′X'_s and Xo′X'_o generate query and key vectors for multi-head self-attention across triplet queries, produce dynamic convolution filter kernels applied to RoI-aligned object image features, and feed into feed-forward networks (FFNs) for box regression and object category classification.

  3. Knowl 3 — Visual Entities to Relation (E2R) Fusion Module

    model/method

    The Visual Entities to Relation Fusion (E2R) module provides bottom-up structured connections within each Triplet Detection Head, passing localized visual and positional cues from the detected subject and object entities into the relation representation.

    Let Fs,Fo∈R1024F_s, F_o \in \mathbb{R}^{1024} denote the dynamic RoI-pooled visual features of the detected subject and object, and let Fr∈R256F_r \in \mathbb{R}^{256} denote the relation-level feature obtained from the union region RoI Align dynamically filtered by the relation content vector. With entity positional encodings Ps,PoP_s, P_o, the updated relation feature Fr′F'_r is computed as:

    Hr=WxReLU(LN(WrsFs))+WyReLU(LN(WroFo))H_r = W_x \text{ReLU}(\text{LN}(W_r^s F_s)) + W_y \text{ReLU}(\text{LN}(W_r^o F_o))

    Fr′=LN(Fr+Hr+WrpReLU(WpsPs+WpoPo))F'_r = \text{LN}\left(F_r + H_r + W_r^p \text{ReLU}(W_p^s P_s + W_p^o P_o)\right)

    where Wrs,Wro,Wx,Wy,Wps,Wpo,WrpW_r^s, W_r^o, W_x, W_y, W_p^s, W_p^o, W_r^p are learnable projection matrices, LN(⋅)\text{LN}(\cdot) denotes layer normalization, and ReLU(⋅)\text{ReLU}(\cdot) denotes the rectified linear unit.

    The enhanced relation vector Fr′F'_r is fed into a classification feed-forward network (FFN). The final relation prediction logits are computed as the element-wise sum of the master classification head on Fr′F'_r and an auxiliary classification branch predicting directly from concatenated entity features (Fs,Fo)(F_s, F_o).

  4. Knowl 4 — Two-Stage Triplet Label Assignment via Siamese Sparse R-CNN Distillation

    model/method

    To address the low recall and supervision sparsity of ground-truth relation annotations during end-to-end training, Structured Sparse R-CNN uses a two-stage bipartite label assignment strategy guided by a weight-sharing Siamese Sparse R-CNN branch.

    The Siamese Sparse R-CNN uses NauxN_{aux} auxiliary queries (e.g., Naux=100N_{aux}=100) to perform object detection independently. Objects detected by this auxiliary detector are paired to generate a pseudo-label candidate set UU, where predicted boxes not matching ground truth are retained and assigned hard pseudo-class labels.

    Training proceeds in two assignment stages:

    1. First Stage (Ground-Truth Matching): Bipartite Hungarian matching matches predictions to ground-truth triplets, minimized by loss LF\mathcal{L}_F:

    LF=λclsrLclsrg+∑i∈{s,o}(λclsiLclsig+λL1iLL1ig+λgiouiLgiouig)\mathcal{L}_F = \lambda_{cls_r} \mathcal{L}_{cls_r}^g + \sum_{i \in \{s, o\}} \left( \lambda_{cls_i} \mathcal{L}_{cls_i}^g + \lambda_{L1_i} \mathcal{L}_{L1_i}^g + \lambda_{giou_i} \mathcal{L}_{giou_i}^g \right)

    where Lclsg\mathcal{L}_{cls}^g denotes focal loss, LL1g\mathcal{L}_{L1}^g denotes L1 bounding-box regression loss, and Lgioug\mathcal{L}_{giou}^g denotes generalized IoU loss.

    1. Second Stage (Pseudo-Label Distillation): Unmatched predicted triplets are matched to pseudo-label pairs in UU via matching cost LBm\mathcal{L}_B^m:

    LBm=∑i∈{s,o}(ηL1iLL1iu+ηgiouiLgiouiu+IiuηclsiLclsiu)\mathcal{L}_B^m = \sum_{i \in \{s, o\}} \left( \eta_{L1_i} \mathcal{L}_{L1_i}^u + \eta_{giou_i} \mathcal{L}_{giou_i}^u + \mathbb{I}_i^u \eta_{cls_i} \mathcal{L}_{cls_i}^u \right)

    where Iiu=1\mathbb{I}_i^u = 1 if the entity in UU corresponds to a ground-truth object and 00 otherwise. Matched predictions are supervised via loss LB\mathcal{L}_B, assigning a background relation label Lclsr−\mathcal{L}_{cls_r}^-:

    LB=λclsrLclsr−+∑i∈{s,o}[ηclsiLclsiu+Iiu(ηL1iLL1iu+ηgiouiLgiouiu)]\mathcal{L}_B = \lambda_{cls_r} \mathcal{L}_{cls_r}^- + \sum_{i \in \{s, o\}} \left[ \eta_{cls_i} \mathcal{L}_{cls_i}^u + \mathbb{I}_i^u (\eta_{L1_i} \mathcal{L}_{L1_i}^u + \eta_{giou_i} \mathcal{L}_{giou_i}^u) \right]

    Hyperparameters are set to λclsr=λclsi=43\lambda_{cls_r} = \lambda_{cls_i} = \frac{4}{3}, λL1i=5\lambda_{L1_i} = 5, λgioui=2\lambda_{giou_i} = 2, ηclsi=13\eta_{cls_i} = \frac{1}{3}, ηL1i=54\eta_{L1_i} = \frac{5}{4}, and ηgioui=12\eta_{giou_i} = \frac{1}{2}.

  5. Knowl 5 — Adaptive Focusing Parameter for Object Entity Focal Loss

    equation

    To counteract entity frequency distortion caused by the fact that frequent individual objects appear with disproportionately higher multiplication when structured as subject-object pairs in triplets, the focusing parameter γ\gamma of the entity focal loss is dynamically calculated per object class cc:

    γ(c)=min⁡{2,3−(1−fc)μ(−log⁡(fc))1μ}\gamma(c) = \min\left\{2, 3 - (1 - f_c)^\mu (-\log(f_c))^{\frac{1}{\mu}}\right\}

    where cc is the object category index, fcf_c is the relative occurrence frequency of object category cc within annotated triplets, and μ\mu is a hyperparameter set to μ=4\mu = 4.

    This formulation adaptively down-weights the loss contributions of head entity classes that dominate the triplet pairing distribution.

  6. Knowl 6 — Scene Graph Detection Performance on Visual Genome Benchmark

    data/table

    Comparison of Structured Sparse R-CNN on the Visual Genome dataset under the Scene Graph Detection (SGDet) setting using a ResNeXt-101-FPN backbone. Evaluation metrics include Recall@K (R@K), zero-shot Recall@K (zR@K), mean Recall@K (mR@K) across thresholds K∈{20,50,100}K \in \{20, 50, 100\}, and per-image inference latency (in seconds). Model variants include standard 300 queries, 800 queries (∗*), Total Direct Effect debiasing (TDE), and Logit Adjustment (LA, τ=0.3\tau=0.3):

    Model R@20 R@50 R@100 zR@20 zR@50 zR@100 mR@20 mR@50 mR@100 Speed (s)
    IMP 18.1 25.9 31.2 0.2 0.4 0.8 2.8 4.2 5.3 0.43
    G-RCNN - 29.7 32.8 - - - - 5.8 6.6 -
    VTransE 24.5 31.3 35.5 - 1.9 2.6 5.1 6.8 8.0 0.40
    RelDN - 31.4 35.9 - - - - 6.0 7.3 -
    GPS-Net - 31.1 35.9 - - - - 7.0 8.6 -
    MOTIFS 25.1 32.1 36.9 - 0.1 0.2 4.1 5.5 6.8 0.45
    VCTree 24.5 31.9 36.2 0.1 0.3 0.7 5.4 7.4 8.7 0.67
    Transformer 25.6 33.0 37.4 0.0 0.1 0.3 6.0 8.1 9.6 0.38
    BGNN - 31.0 35.8 - - - - 10.7 12.6 -
    Ours 25.8 32.7 36.9 1.5 2.7 3.7 6.1 8.4 10.0 0.19
    Ours* 26.1 33.5 38.4 1.5 2.7 4.0 6.2 8.6 10.3 0.32
    MOTIFS_TDE 12.4 16.9 20.3 - 2.3 2.9 5.8 8.2 9.8 -
    VCTree_TDE 14.0 19.4 23.2 - 2.6 3.2 6.9 9.3 11.1 -
    Ours_TDE 14.5 18.3 21.0 1.8 2.7 3.6 10.8 15.0 18.5 0.29
    Ours*_TDE 15.0 19.7 22.9 1.6 2.7 3.8 9.8 14.6 18.0 0.54
    Ours_LA 18.4 23.3 26.5 1.9 2.9 4.0 13.5 17.9 21.4 0.19
    Ours*_LA 18.2 23.7 27.3 2.0 3.1 4.5 13.7 18.6 22.5 0.32

    Structured Sparse R-CNN achieves faster inference speed (0.19 s/image) and higher zero-shot Recall (3.7% zR@100 vs. ≤0.8%\le 0.8\% for non-debiased baselines). With logit adjustment (Ours*_{LA}), mR@100 reaches 22.5%.

  7. Knowl 7 — Scene Graph Generation Performance on Open Images V4 and V6 Benchmarks

    data/table

    Performance of Structured Sparse R-CNN compared to prior methods on Open Images (OI) V4 and Open Images V6 using identical backbones (ResNeXt-101-FPN). Metrics evaluated include mean Recall@50 (mR@50), micro-Recall@50 (R@50), weighted mean average precision of relations (wmAPrel\text{wmAP}_{rel}), weighted mean average precision of phrases (wmAPphr\text{wmAP}_{phr}), and the overall weighted score: scorewtd=0.2×R@50+0.4×wmAPrel+0.4×wmAPphr\text{score}_{wtd} = 0.2 \times \text{R@50} + 0.4 \times \text{wmAP}_{rel} + 0.4 \times \text{wmAP}_{phr}

    Open Images V4
    Model mR@50 R@50 wmAP_rel wmAP_phr score_wtd
    RelDN 70.40 75.66 36.13 39.91 45.21
    GPS-Net 69.50 74.65 35.02 39.40 44.70
    BGNN 72.11 75.46 37.76 41.70 46.87
    Ours 72.62 74.92 43.47 48.17 51.64
    Ours_LA 79.23 74.75 43.57 48.25 51.68
    Open Images V6
    MOTIFS 32.68 71.63 29.91 31.59 38.93
    RelDN 33.98 73.08 32.16 33.39 40.84
    VCTree 33.91 74.08 34.16 33.11 40.21
    G-RCNN 34.04 74.51 33.15 34.21 41.84
    GPS-Net 35.26 74.81 32.85 33.98 41.69
    BGNN 40.45 74.98 33.51 34.15 42.06
    Ours 42.84 76.66 41.47 43.64 49.38
    Ours_LA 50.73 75.70 41.14 43.24 48.89

    Structured Sparse R-CNN exceeds existing methods across both benchmarks, improving scorewtd\text{score}_{wtd} on OI V4 to 51.64 (vs. 46.87 for BGNN) and on OI V6 to 49.38 (vs. 42.06 for BGNN).

  8. Knowl 8 — Ablation Study on Structured Modules in Triplet Detection Head

    data/table

    Ablation study on the Visual Genome SGDet dataset analyzing the incremental impact of the relation feature vector (Rel), the Pair Fusion (PF) module, and the Visual Entities to Relation Fusion (E2R) module within Structured Sparse R-CNN:

    Rel E2R PF R@20 R@100 zR@20 zR@100 mR@20 mR@100 zR@100 (LA) mR@100 (LA)
    23.16 34.62 0.79 2.44 5.09 8.62 3.03 19.04
    ✓ 25.33 36.57 0.94 2.41 5.67 9.13 2.96 20.25
    ✓ ✓ 25.52 36.58 1.38 3.41 6.03 9.83 3.61 19.96
    ✓ ✓ 24.32 35.34 1.24 3.32 5.77 9.37 3.83 19.55
    ✓ ✓ ✓ 25.82 36.93 1.51 3.74 6.08 10.04 4.04 21.39

    The baseline treating relation detection purely as object pair detection without explicit relation vectors or structured fusion yields 23.16% R@20 and 8.62% mR@100. Combining Rel, PF, and E2R achieves 25.82% R@20 (+2.66%), 10.04% mR@100 (+1.42%), and 21.39% mR@100 under logit adjustment (+2.35%).

  9. Knowl 9 — Ablation Study on Triplet Label Assignment and Background Supervision

    data/table

    Ablation study on Visual Genome evaluating different triplet label assignment strategies and the effect of non-maximum suppression (NMS) post-processing on Structured Sparse R-CNN co-trained with Siamese Sparse R-CNN:

    TLA NMS R@20 R@100 zR@20 zR@100 mR@20 mR@100 zR@100 (LA) mR@100 (LA) Speed (s)
    full BG 24.62 35.00 1.21 3.11 5.75 9.20 3.52 18.57 0.19
    full BG ✓ 24.85 35.14 1.24 3.27 5.83 9.26 3.69 18.78 0.29
    no BG 23.21 35.99 1.09 2.82 5.25 9.63 3.35 20.35 0.19
    no BG ✓ 25.45 38.29 1.33 3.79 5.93 10.63 3.98 22.23 0.29
    p-label 25.82 36.93 1.51 3.74 6.08 10.04 4.04 21.39 0.19
    p-label ✓ 26.49 37.42 1.62 4.09 6.27 10.24 4.19 21.65 0.29

    Definitions of compared label assignment strategies:

    • full BG: Unmatched triplet object candidates are assigned the background class.
    • no BG: Background loss on unmatched triplet object candidates is omitted.
    • p-label: Proposed two-stage matching using pseudo-labels generated by Siamese Sparse R-CNN distillation.

    Without NMS, the pseudo-label strategy (p-label) achieves the highest performance across all recall metrics (25.82% R@20, 10.04% mR@100, 21.39% mR@100 with LA) while running at 0.19 s per image.

Coverage note — None was omitted; all contributed models, loss formulations, training distillation mechanisms, adaptive focal parameter calculations, and empirical evaluations on Visual Genome and Open Images V4/V6 are represented.

References

  1. 1.Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. CoRR, abs/1607.06450, 2016. 4
  2. 2.Hedi Ben-younes, Remi Cadene, Nicolas Thome, and Matthieu Cord. BLOCK: bilinear superdiagonal fusion for visual question answering and visual relationship detection. In AAAI, pages 8102–8109, 2019. 1
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, pages 213–229, 2020. 1, 3, 5
  4. 4.Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, and Shih-Fu Chang. Counterfactual critic multi-agent training for scene graph generation. In ICCV, pages 4612–4622, 2019. 6
  5. 5.Mingfei Chen, Yue Liao, Si Liu, Zhiyuan Chen, Fei Wang, and Chen Qian. Reformulating HOI detection as adaptive set prediction. In CVPR, pages 9004–9013, 2021. 2
  6. 6.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In ICCV, pages 764–773, 2017. 3
  7. 7.Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, and Petros Maragos. Attention-translation-relation network for scalable scene graph generation. In ICCV Workshops, pages 1754–1764, 2019. 8
  8. 8.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. Deep sparse rectifier neural networks. In AISTATS, volume 15 of JMLR Proceedings, pages 315–323, 2011. 4
  9. 9.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross B. Girshick. Mask R-CNN. In ICCV, pages 2980–2988, 2017. 4
  10. 10.Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. CoRR, abs/1503.02531, 2015. 2, 4
  11. 11.Drew A. Hudson and Christopher D. Manning. GQA: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, pages 6700–6709, 2019. 1
  12. 12.Xu Jia, Bert De Brabandere, Tinne Tuytelaars, and Luc Van Gool. Dynamic filter networks. In NIPS, pages 667–675, 2016. 4
  13. 13.Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J. Kim. HOTR: end-to-end human-object interaction detection with transformers. In CVPR, pages 74–83, 2021. 2
  14. 14.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In ICLR, 2017. 2
  15. 15.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis., 123(1):32–73, 2017. 2, 5
  16. 16.Harold W. Kuhn. The hungarian method for the assignment problem. In 50 Years of Integer Programming, pages 29–47. Springer, 2010. 5
  17. 17.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari. The open images dataset V4: unified image classification, object detection, and visual relationship detection at scale. CoRR, abs/1811.00982, 2018. 2, 5, 6
  18. 18.Rongjie Li, Songyang Zhang, Bo Wan, and Xuming He. Bipartite graph network with adaptive message passing for unbiased scene graph generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11109–11119, 2021. 2, 3, 6, 7, 8
  19. 19.Yikang Li, Wanli Ouyang, and Xiaogang Wang. Vip-cnn: A visual phrase reasoning convolutional neural network for visual relationship detection. CoRR, abs/1702.07191, 2017. 2
  20. 20.Yikang Li, Wanli Ouyang, Bolei Zhou, Jianping Shi, Chao Zhang, and Xiaogang Wang. Factorizable net: An efficient subgraph-based framework for scene graph generation. In ECCV, pages 346–363, 2018. 2
  21. 21.Tsung-Yi Lin, Piotr Dollar, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. Feature pyramid networks for object detection. In CVPR, pages 936–944, 2017. 3, 6
  22. 22.Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, pages 2999–3007, 2017. 5
  23. 23.Xin Lin, Changxing Ding, Jinquan Zeng, and Dacheng Tao. Gps-net: Graph property sensing network for scene graph generation. In CVPR, pages 3743–3752, 2020. 2, 5, 7, 8
  24. 24.Hengyue Liu, Ning Yan, Masood Mortazavi, and Bir Bhanu. Fully convolutional scene graph generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11546–11556, 2021. 2
  25. 25.Yen-Cheng Liu, Chih-Yao Ma, Zijian He, Chia-Wen Kuo, Kan Chen, Peizhao Zhang, Bichen Wu, Zsolt Kira, and Peter Vajda. Unbiased teacher for semi-supervised object detection. In ICLR, 2021. 2
  26. 26.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 6
  27. 27.Cewu Lu, Ranjay Krishna, Michael S. Bernstein, and Fei-Fei Li. Visual relationship detection with language priors. In ECCV, pages 852–869, 2016. 6
  28. 28.Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar. Long-tail learning via logit adjustment. In ICLR, 2021. 3, 5, 7
  29. 29.Alejandro Newell and Jia Deng. Pixels to graphs by associative embedding. In NIPS, pages 2171–2180, 2017. 2
  30. 30.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian D. Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR, pages 658–666, 2019. 5
  31. 31.Mamshad Nayeem Rizve, Kevin Duarte, Yogesh Singh Rawat, and Mubarak Shah. In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning. In ICLR, 2021. 2
  32. 32.Jiaxin Shi, Hanwang Zhang, and Juanzi Li. Explainable and explicit visual reasoning over scene graphs. In CVPR, pages 8376–8384, 2019. 1
  33. 33.Mohammed Suhail, Abhay Mittal, Behjat Siddiquie, Chris Broaddus, Jayan Eledath, Gerard G. Medioni, and Leonid Sigal. Energy-based learning for scene graph generation. In CVPR, pages 13936–13945, 2021. 7
  34. 34.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, and Ping Luo. Sparse R-CNN: end-to-end object detection with learnable proposals. In CVPR, pages 14454–14463, 2021. 1, 3, 4, 5, 6
  35. 35.Masato Tamura, Hiroki Ohashi, and Tomoaki Yoshinaga. QPIC: query-based pairwise human-object interaction detection with image-wide contextual information. In CVPR, pages 10410–10419, 2021. 2
  36. 36.Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. Unbiased scene graph generation from biased training. In CVPR, pages 3713–3722, 2020. 2, 3, 6, 8
  37. 37.Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu. Learning to compose dynamic tree structures for visual contexts. In CVPR, pages 6619–6628, 2019. 1, 2, 6, 7, 8
  38. 38.Yao Teng, Limin Wang, Zhifeng Li, and Gangshan Wu. Target adaptive context aggregation for video scene graph generation. In ICCV, pages 13668–13677, 2021. 2
  39. 39.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In ICML, pages 10347–10357, 2021. 5
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, pages 5998–6008, 2017. 3, 4, 7
  41. 41.Wenbin Wang, Ruiping Wang, Shiguang Shan, and Xilin Chen. Exploring context and visual pattern of relationship for scene graph generation. In CVPR, pages 8188–8197, 2019. 2
  42. 42.Wenbin Wang, Ruiping Wang, Shiguang Shan, and Xilin Chen. Sketching image gist: Human-mimetic hierarchical scene graph generation. In ECCV, pages 222–239, 2020. 2
  43. 43.Chen Wei, Kihyuk Sohn, Clayton Mellina, Alan L. Yuille, and Fan Yang. Crest: A class-rebalancing self-training framework for imbalanced semi-supervised learning. In CVPR, pages 10857–10866, 2021. 2
  44. 44.Saining Xie, Ross B. Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, pages 5987–5995, 2017. 3, 6
  45. 45.Danfei Xu, Yuke Zhu, Christopher B. Choy, and Li Fei-Fei. Scene graph generation by iterative message passing. In CVPR, pages 3097–3106, 2017. 1, 2, 6, 7
  46. 46.Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. Graph R-CNN for scene graph generation. In ECCV, pages 690–706, 2018. 1, 2, 7, 8
  47. 47.Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai. Auto-encoding scene graphs for image captioning. In CVPR, pages 10685–10694, 2019. 1
  48. 48.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Exploring visual relationship for image captioning. In ECCV, pages 711–727, 2018. 1
  49. 49.Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi. Neural motifs: Scene graph parsing with global context. In CVPR, pages 5831–5840, 2018. 1, 2, 6, 7, 8
  50. 50.Aixi Zhang, Yue Liao, Si Liu, Miao Lu, Yongliang Wang, Chen Gao, and Xiaobo Li. Mining the benefits of two-stage and one-stage HOI detection. Advances in Neural Information Processing Systems, 34, 2021. 2
  51. 51.Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua. Visual translation embedding network for visual relation detection. In CVPR, pages 3107–3115, 2017. 7
  52. 52.Ji Zhang, Kevin J. Shih, Ahmed Elgammal, Andrew Tao, and Bryan Catanzaro. Graphical contrastive losses for scene graph parsing. In CVPR, pages 11535–11543, 2019. 1, 2, 4, 7, 8
  53. 53.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable DETR: deformable transformers for end-to-end object detection. In ICLR, 2021. 3
  54. 54.Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, and Jian Sun. End-to-end human object interaction detection with HOI transformer. In CVPR, pages 11825–11834, 2021. 2

Citation

MLA
Teng, Y., and L. Wang. “Structured Sparse R-CNN for Direct Scene Graph Generation”. arXiv, 2021, http://arxiv.org/abs/2106.10815v2.
APA
Teng, Y., & Wang, L. (2021). Structured Sparse R-CNN for Direct Scene Graph Generation. arXiv. http://arxiv.org/abs/2106.10815v2
Chicago
Teng, Y., and L. Wang. 2021. “Structured Sparse R-CNN for Direct Scene Graph Generation”. arXiv. http://arxiv.org/abs/2106.10815v2.
Harvard
Teng, Y. and Wang, L. (2021) “Structured Sparse R-CNN for Direct Scene Graph Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2106.10815v2.
Vancouver
1. Teng Y, Wang L (2021) Structured Sparse R-CNN for Direct Scene Graph Generation. arXiv

BibTeX

@article{teng2021structured,
  title = {Structured Sparse R-CNN for Direct Scene Graph Generation},
  author = {Teng, Yao and Wang, Limin},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2106.10815v2},
  eprint = {2106.10815}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE