Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space

Yong ZhangYingwei PanTing YaoRui HuangTao MeiChang Wen Chen

article2023CVPR54 citations

Proposes a scene graph generation framework that uses pre-trained visual-semantic spaces to detect novel objects and relations from image captions without requiring expensive manual annotations.

Listen

Scene graph generation converts visual images into structured representations by identifying objects as nodes and their interactions as labeled edges. This capability is critical for fine-grained computer vision tasks, including automated image captioning, cross-modal retrieval, and autonomous robot decision-making. However, real-world deployment faces two significant hurdles: training models requires labor-intensive and costly manual bounding-box annotations, and conventional systems are restricted to a pre-defined, closed set of categories, causing them to fail when encountering unfamiliar objects in dynamic environments.

The article demonstrates an effective methodology called Visual-Semantic Space for Scene graph generation, or VS3, to overcome these bottlenecks. The primary objective is to evaluate how pre-trained vision-language models can be transferred to eliminate the need for expensive manual labels while enabling the system to recognize novel, unseen visual categories and their relationships in an open-vocabulary setting.

To achieve this, the authors leverage Grounded Language-Image Pre-training, an established foundational model that maps visual regions and descriptive text phrases into a shared visual-semantic space. The framework parses raw descriptive text captions into semantic relationships, automatically matches text entities to image regions using the pre-trained space, and uses these grounded graphs as low-cost supervision data. To infer relationships, a lightweight module combines visual and spatial coordinates between detected objects. The model freezes the underlying image and text encoders to prevent performance degradation and updates only the cross-modal and relationship-prediction components across benchmark evaluations on the Visual Genome dataset.

The findings show that VS3 sets new performance benchmarks across multiple configurations. In fully supervised settings, the larger variant achieved top-50 recall scores of 36.6 percent, outperforming prior state-of-the-art models by up to 3.4 percentage points while using a simpler architecture. Under language-only supervision, it scored 29.81 percent on top-50 recall, more than doubling earlier weakly supervised approaches and matching or exceeding several fully supervised systems. In the open-vocabulary setting involving 30 percent previously unseen object categories, the system demonstrated viable end-to-end detection and relationship mapping directly from raw images, a benchmark previously unsupported by standard two-stage object detectors.

These results demonstrate that leveraging pre-trained multi-modal foundations drastically cuts the financial and operational costs associated with manual data labeling while substantially lowering the risk of model failure in unconstrained environments. Systems capable of processing novel categories without retraining significantly improve adaptability and safety for automated perception pipelines, robotics, and content retrieval workflows.

Organizations developing vision-language solutions should consider adopting pre-trained grounding foundations and weakly supervised pipelines rather than relying exclusively on custom, manually annotated datasets. Future development should focus on testing these pipelines with richer, region-level descriptive text to minimize the performance drop caused by sparse captions and cross-domain data shifts.

The primary limitation highlighted in the analysis is that language-derived supervision tends to capture simpler, high-frequency relational predicates, such as basic spatial prepositions, compared to human-curated datasets. While confidence in the model's overall generalization and detection accuracy is high, practitioners should exercise caution when deploying systems in domains that require highly specialized or nuanced relationship classifications without fine-tuning.

Cover for Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space

Abstract

Scene graph generation (SGG) aims to abstract an image into a graph structure, by representing objects as graph nodes and their relations as labeled edges. However, two knotty obstacles limit the practicability of current SGG methods in real-world scenarios: 1) training SGG models requires time-consuming ground-truth annotations, and 2) the closed-set object categories make the SGG models limited in their ability to recognize novel objects outside of training corpora. To address these issues, we novelly exploit a powerful pre-trained visual-semantic space (VSS) to trigger language-supervised and open-vocabulary SGG in a simple yet effective manner. Specifically, cheap scene graph supervision data can be easily obtained by parsing image language descriptions into semantic graphs. Next, the noun phrases on such semantic graphs are directly grounded over image regions through region-word alignment in the pre-trained VSS. In this way, we enable open-vocabulary object detection by performing object category name grounding with a text prompt in this VSS. On the basis of visually-grounded objects, the relation representations are naturally built for relation recognition, pursuing open-vocabulary SGG. We validate our proposed approach with extensive experiments on the Visual Genome benchmark across various SGG scenarios (i.e., supervised / language-supervised, closed-set / open-vocabulary). Consistent superior performances are achieved compared with existing methods, demonstrating the potential of exploiting pre-trained VSS for SGG in more practical scenarios.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Notation & Overview
  • 3.2. The Proposed VS 3 Model
  • 3.3. Obtaining Language Scene Graph Supervision
  • 3.4. Transferring to Open-vocabulary SGG
  • 4. Experiments
  • 4.1. Datasets and Experimental Settings
  • 4.2. Fully Supervised SGG
  • 4.3. Language-supervised SGG
  • 4.4. Open-vocabulary SGG
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Visual-Semantic Space Model for Scene Graph Generation

    model/method

    The Visual-Semantic Space for Scene Graph Generation (VS3VS^3) framework extends grounded language-image pre-training (GLIP) to detect visual objects and predict pairwise relationships in a shared visual-semantic space (VSS).

    VS3VS^3 utilizes an image encoder EncI\text{Enc}_I (such as Swin Transformer) that extracts visual region embeddings O~∈RNˉ×d\tilde{\mathbf{O}} \in \mathbb{R}^{\bar{N} \times d} and a text encoder EncL\text{Enc}_L (such as BERT) that encodes a text prompt containing object category names into token embeddings P~∈RTˉ×d\tilde{\mathbf{P}} \in \mathbb{R}^{\bar{T} \times d}, where Nˉ\bar{N} is the number of candidate image regions, Tˉ\bar{T} is the sequence length of the prompt, and dd is the embedding dimension (d=256d = 256). A cross-modal fusion module provides bidirectional feature interaction between visual region features and text token features.

    Object detection is performed by computing alignment scores S^ground=O~P~⊤∈RNˉ×Tˉ\hat{S}_{ground} = \tilde{\mathbf{O}} \tilde{\mathbf{P}}^\top \in \mathbb{R}^{\bar{N} \times \bar{T}} and box coordinate predictions B^∈RNˉ×4\hat{B} \in \mathbb{R}^{\bar{N} \times 4}. The top-N′N' object detections (with N′=36N'=36) are retained via non-maximum suppression (NMS) to form candidate subject-object pairs (o~i,o~j)(\tilde{\mathbf{o}}_i, \tilde{\mathbf{o}}_j). A relation embedding module then constructs pairwise visual and spatial relation representations, which are passed to relation prediction heads for relateness estimation and predicate category classification.

    During training, the image encoder EncI\text{Enc}_I and text encoder EncL\text{Enc}_L remain frozen to preserve the open-vocabulary visual-semantic alignment, while only the cross-modal fusion module, relation embedding module, and relation prediction heads are optimized.

  2. Knowl 2 — Pairwise Relation Representation in VS3

    model/method

    In the VS3VS^3 scene graph generation model, candidate subject-object pairs (o~i,o~j)(\tilde{\mathbf{o}}_i, \tilde{\mathbf{o}}_j) with normalized bounding box coordinates (bi,bj)(\mathbf{b}_i, \mathbf{b}_j) are mapped into a pairwise relation representation pairi→j\mathbf{pair}_{i \rightarrow j} by concatenating visual and spatial features:

    pairi→j=cat[pairi→jvisual,pairi→jspatial]\mathbf{pair}_{i \rightarrow j} = \text{cat}[\mathbf{pair}_{i \rightarrow j}^{visual}, \mathbf{pair}_{i \rightarrow j}^{spatial}]

    The visual relation feature pairi→jvisual\mathbf{pair}_{i \rightarrow j}^{visual} is computed from region embeddings o~i,o~j∈Rd\tilde{\mathbf{o}}_i, \tilde{\mathbf{o}}_j \in \mathbb{R}^d through difference and sum mapping functions:

    pairi→jvisual=fdiff(o~i−o~j)+fsum(o~i+o~j)\mathbf{pair}_{i \rightarrow j}^{visual} = \mathbf{f}_{diff}(\tilde{\mathbf{o}}_i - \tilde{\mathbf{o}}_j) + \mathbf{f}_{sum}(\tilde{\mathbf{o}}_i + \tilde{\mathbf{o}}_j)

    where fdiff\mathbf{f}_{diff} and fsum\mathbf{f}_{sum} are two-layer multi-layer perceptrons (MLPs).

    The spatial relation feature pairi→jspatial\mathbf{pair}_{i \rightarrow j}^{spatial} captures geometric properties:

    pairi→jspatial=cat[bi,bj,dx,dy,dis,θ,Ai,Aj,I,U]\mathbf{pair}_{i \rightarrow j}^{spatial} = \text{cat}[\mathbf{b}_i, \mathbf{b}_j, dx, dy, dis, \theta, A_i, A_j, I, U]

    where:

    • (ctxi,ctyi)(ctx_i, cty_i) and (ctxj,ctyj)(ctx_j, cty_j) are normalized center coordinates of subject box bi\mathbf{b}_i and object box bj\mathbf{b}_j.
    • dx=ctxi−ctxjdx = ctx_i - ctx_j and dy=ctyi−ctyjdy = cty_i - cty_j are coordinate differences.
    • dis=dx2+dy2dis = \sqrt{dx^2 + dy^2} is the Euclidean center distance.
    • θ=arctan⁡(dy/dx)\theta = \arctan(dy / dx) is the relative spatial angle.
    • AiA_i and AjA_j denote the areas of the subject and object bounding boxes.
    • II and UU denote the areas of the bounding box intersection and union, respectively.
  3. Knowl 3 — Relation Prediction and Loss Formulation in VS3

    equation

    Given the pairwise relation representation pairi→j\mathbf{pair}_{i \rightarrow j}, VS3VS^3 predicts a scalar relateness score z^i→j∈[0,1]\hat{z}_{i \rightarrow j} \in [0, 1] indicating whether a relation exists between the object pair, and a semantic predicate probability distribution y^i→j∈[0,1]∣Cr∣\hat{\mathbf{y}}_{i \rightarrow j} \in [0, 1]^{|\mathcal{C}_r|} over the predicate set Cr\mathcal{C}_r:

    z^i→j=frelateness(pairi→j)=σ(MLPrel(pairi→j))\hat{z}_{i \rightarrow j} = f_{relateness}(\mathbf{pair}_{i \rightarrow j}) = \sigma(\text{MLP}_{rel}(\mathbf{pair}_{i \rightarrow j}))

    y^i→j=fsemantic(pairi→j)=Softmax(MLPsem(pairi→j))\hat{\mathbf{y}}_{i \rightarrow j} = f_{semantic}(\mathbf{pair}_{i \rightarrow j}) = \text{Softmax}(\text{MLP}_{sem}(\mathbf{pair}_{i \rightarrow j}))

    where σ\sigma is the sigmoid function, and MLPrel\text{MLP}_{rel} and MLPsem\text{MLP}_{sem} are multi-layer perceptrons.

    The relation recognition loss Lrel_rcg\mathcal{L}_{rel\_rcg} is the sum of a focal loss on relateness and a cross-entropy loss on predicate classification:

    Lrelateness=FL(z^i→j,zi→j)\mathcal{L}_{relateness} = \text{FL}(\hat{z}_{i \rightarrow j}, z_{i \rightarrow j})

    Lsemantic=CE(y^i→j,yi→j)\mathcal{L}_{semantic} = \text{CE}(\hat{\mathbf{y}}_{i \rightarrow j}, \mathbf{y}_{i \rightarrow j})

    Lrel_rcg=Lrelateness+Lsemantic\mathcal{L}_{rel\_rcg} = \mathcal{L}_{relateness} + \mathcal{L}_{semantic}

    where zi→j∈{0,1}z_{i \rightarrow j} \in \{0, 1\} is the ground-truth relateness label, yi→j\mathbf{y}_{i \rightarrow j} is the one-hot ground-truth predicate label vector, FL\text{FL} denotes focal loss, and CE\text{CE} denotes cross-entropy loss.

  4. Knowl 4 — Language Scene Graph Supervision Extraction and Grounding

    model/method

    To train scene graph generation models without manual bounding box annotations, weak supervision is extracted from image descriptions through a two-stage parsing and grounding pipeline:

    1. Semantic Graph Parsing: An image description is parsed into an unlocalized semantic graph SGtext={Otext,Rtext}SG_{text} = \{O_{text}, R_{text}\} using a scene graph parser, extracting entity noun phrases OtextO_{text} and relation triplets Rtext={⟨oi,pij,oj⟩}R_{text} = \{\langle o_i, p_{ij}, o_j \rangle\}. Parsed words are mapped to canonical dataset categories using direct string matching and WordNet synset matching.
    2. Semantic Graph Grounding: For each entity in OtextO_{text}, bounding box localization is performed using an off-the-shelf grounded language-image pre-training model (GLIP). A text prompt is formed by concatenating the parsed triplets (e.g., "woman playing piano. woman in room.") and fed with the image into GLIP. Each entity is assigned the image region exhibiting the highest visual-semantic alignment score with its label. Redundant candidate boxes referring to the same entity are merged via non-maximum suppression (NMS) on identical labels with intersection-over-union IoU≥0.9\text{IoU} \ge 0.9.

    The visually-grounded semantic graph is then used as weak supervision to train the scene graph generator.

  5. Knowl 5 — Transferring VS3 to Open-Vocabulary Scene Graph Generation

    model/method

    The VS3VS^3 model enables open-vocabulary scene graph generation (Ov-SGG), where unseen novel object categories Conovel=Cotarget∖Cobase≠∅\mathcal{C}_o^{novel} = \mathcal{C}_o^{target} \setminus \mathcal{C}_o^{base} \neq \emptyset must be detected and related to other objects at inference time.

    During training on base categories Cobase\mathcal{C}_o^{base}, the text prompt input to the text encoder is restricted to base category names:

    Prompttrain="name(c1).name(c2).…name(c∣Cobase∣)."\text{Prompt}_{train} = \text{"name}(c_1). \text{name}(c_2). \dots \text{name}(c_{|\mathcal{C}_o^{base}|}).\text{"}

    where ci∈Cobasec_i \in \mathcal{C}_o^{base}, and only relation triplets involving base classes are used for training.

    At inference time, the text prompt is switched to include all target categories:

    Prompttest="name(c1).name(c2).…name(c∣Cotarget∣)."\text{Prompt}_{test} = \text{"name}(c_1). \text{name}(c_2). \dots \text{name}(c_{|\mathcal{C}_o^{target}|}).\text{"}

    where ci∈Cotargetc_i \in \mathcal{C}_o^{target}.

    Because the underlying visual-semantic space (VSS) is preserved by keeping the encoders frozen, novel object category names embed closely to semantically related concepts and align directly with image regions without additional detector fine-tuning. Furthermore, because the visual and spatial relation embedding features are class-agnostic, pairwise relation recognition functions seamlessly for pairs involving novel object categories.

  6. Knowl 6 — Taxonomy of Scene Graph Generation Settings

    definition

    Scene graph generation (SGG) tasks are formalized along two orthogonal axes: the supervision source for scene graphs (SGSG) and the target inference object vocabulary (Cotarget\mathcal{C}_o^{target}).

    Supervision Source (SGSG) Closed-Set Inference (Conovel=∅\mathcal{C}_o^{novel} = \emptyset) Open-Vocabulary Inference (Conovel≠∅\mathcal{C}_o^{novel} \neq \emptyset)
    Manually annotated Fully supervised closed-set Fully supervised open-vocabulary
    Parsed from language descriptions Language-supervised closed-set Language-supervised open-vocabulary
    • Fully Supervised: Training uses manual ground-truth annotations of bounding boxes, object categories, and relation predicates.
    • Language-Supervised: Training uses weak supervision parsed and grounded from free-form natural language captions.
    • Closed-Set: The target inference vocabulary Cotarget\mathcal{C}_o^{target} contains only the base categories seen during training (Cobase\mathcal{C}_o^{base}).
    • Open-Vocabulary: The target inference vocabulary Cotarget\mathcal{C}_o^{target} includes unseen novel object categories Conovel=Cotarget∖Cobase\mathcal{C}_o^{novel} = \mathcal{C}_o^{target} \setminus \mathcal{C}_o^{base}.
  7. Knowl 7 — Fully Supervised Scene Graph Generation Performance on VG150

    empirical result

    Under the fully supervised SGDET setting on the Visual Genome benchmark (VG150, 150 object classes, 50 predicate classes), VS3VS^3 outperforms prior methods across Recall@20 (R@20), Recall@50 (R@50), and Recall@100 (R@100) under graph constraint.

    SGG Model Backbone R@20 R@50 R@100
    FCSGG HRNetW48 16.1 21.3 25.1
    SGTR R-101 - 24.6 28.4
    IMP VGG-16 14.6 20.7 24.5
    KERN VGG-16 - 27.1 29.8
    MOTIFS VGG-16 21.4 27.2 30.3
    RelDN VGG-16 21.1 28.3 32.7
    VTransE RX-101 23.0 29.7 34.3
    MOTIFS RX-101 25.1 32.1 36.9
    VCTREE RX-101 24.7 31.5 36.2
    SGNLS RX-101 24.6 31.8 36.3
    HL-Net RX-101 26.0 33.7 38.1
    VS3VS^3 Swin-T 26.1 34.5 39.2
    VS3VS^3 (w/o visual) Swin-T 23.1 31.6 36.7
    VS3VS^3 (w/o spatial) Swin-T 24.3 32.8 37.8
    VS3VS^3 Swin-L 27.8 36.6 41.5

    VS3VS^3 with Swin-L achieves an R@50 of 36.6 (a 2.9-point increase over HL-Net at 33.7) and R@100 of 41.5 (a 3.4-point increase over HL-Net at 38.1). Ablation of relation representation components shows that visual features contribute more significantly than spatial features: removing visual features reduces R@50 from 34.5 to 31.6 (-2.9 points), whereas removing spatial features reduces R@50 to 32.8 (-1.7 points).

  8. Knowl 8 — Language-Supervised Scene Graph Generation Performance on VG150

    empirical result

    When trained using weak scene graph supervision derived from language under the SGDET protocol on the VG150 test split, VS3VS^3 achieves superior performance compared to previous weakly supervised SGG methods across three supervision sources: unlocalized scene graphs, dense VG region captions, and image-level COCO captions.

    Text Supervision Source Model R@20 R@50 R@100
    Unlocalized graph VSPNet - 4.70 5.40
    Unlocalized graph LSWS - 7.30 8.73
    Unlocalized graph MOTIFS (WSGM) 4.12 5.59 6.45
    Unlocalized graph MOTIFS (SGNLS) 7.23 9.28 10.71
    Unlocalized graph MOTIFS (Li et al.) 9.09 11.39 12.89
    Unlocalized graph Uniter (SGNLS) 7.81 10.03 11.50
    Unlocalized graph Uniter (Li et al.) 9.57 11.80 13.15
    Unlocalized graph VS3VS^3 (Swin-T) 18.02 23.89 28.19
    Unlocalized graph VS3VS^3 (Swin-T + FreqBias) 20.06 26.72 31.75
    Unlocalized graph VS3VS^3 (Swin-L + FreqBias) 22.18 29.81 34.96
    VG caption LSWS - 3.85 4.04
    VG caption MOTIFS (SGNLS) 6.31 8.05 9.21
    VG caption MOTIFS (Li et al.) 8.25 10.50 11.98
    VG caption Uniter (SGNLS) - 9.20 10.30
    VG caption Uniter (Li et al.) 8.90 10.93 12.14
    VG caption VS3VS^3 (Swin-T) 11.78 16.25 19.70
    VG caption VS3VS^3 (Swin-L) 13.01 17.38 20.54
    COCO caption LSWS - 3.28 3.69
    COCO caption MOTIFS (Li et al.) 5.02 6.40 7.33
    COCO caption Uniter (SGNLS) - 5.80 6.70
    COCO caption Uniter (Li et al.) 5.42 6.74 7.62
    COCO caption VS3VS^3 (Swin-T) 5.59 7.30 8.62
    COCO caption VS3VS^3 (Swin-L) 6.04 8.15 9.90

    VS3VS^3 (Swin-L + FreqBias) trained on unlocalized scene graphs achieves R@50 of 29.81, outperforming several fully supervised baselines trained on human bounding box annotations. On the out-of-domain COCO caption setting, VS3VS^3 (Swin-L) reaches 8.15 R@50 and 9.90 R@100.

  9. Knowl 9 — Open-Vocabulary and Zero-Shot Object Scene Graph Generation Performance

    empirical result

    In open-vocabulary SGG, models are trained on 70% base object categories of VG150 and evaluated on two test splits: Open-Vocabulary SGG (Ov-SGG, evaluating on 70% base + 30% novel categories) and Zero-Shot Object SGG (ZsO-SGG, evaluating exclusively on the 30% novel categories).

    Under fully supervised training (reporting R@50 / R@100):

    Ov-SGG (70%+30%) ZsO-SGG (30%)
    Method PREDCLS SGDET PREDCLS SGDET
    IMP 40.02 / 43.40 - 37.01 / 39.46 -
    MOTIFS 41.14 / 44.70 - 39.53 / 41.14 -
    VCTREE 42.56 / 45.84 - 41.27 / 42.52 -
    TDE 38.29 / 40.38 - 34.15 / 36.37 -
    GCA 43.48 / 46.26 - 42.56 / 43.18 -
    EBM 44.09 / 46.95 - 43.27 / 44.03 -
    SVRP 47.62 / 49.94 - 45.75 / 48.39 -
    VS3VS^3 (Swin-T) 50.10 / 52.05 15.07 / 18.73 46.91 / 49.13 10.08 / 13.65
    VS3VS^3 (Swin-L) 55.88 / 58.18 23.13 / 28.49 54.44 / 57.35 21.51 / 27.62

    While previous methods could not evaluate open-vocabulary SGDET due to detector vocabulary bottlenecks, VS3VS^3 achieves 10.08 / 13.65 (Swin-T) and 21.51 / 27.62 (Swin-L) on ZsO-SGG SGDET.

    Under language-supervised open-vocabulary training (SGDET protocol, reporting R@50 / R@100):

    Model Supervision Source Ov-SGG (R@50 / R@100) ZsO-SGG (R@50 / R@100)
    VS3VS^3 (Swin-T) Manual annotation 15.07 / 18.73 10.08 / 13.65
    VS3VS^3 (Swin-T) VG caption 7.61 / 9.60 4.06 / 5.58
    VS3VS^3 (Swin-T) COCO caption 4.39 / 5.63 3.65 / 4.73
    VS3VS^3 (Swin-L) Manual annotation 23.13 / 28.49 21.51 / 27.62
    VS3VS^3 (Swin-L) VG caption 12.98 / 16.29 10.71 / 13.70
    VS3VS^3 (Swin-L) COCO caption 6.76 / 8.45 6.26 / 7.89

    With VS3VS^3 (Swin-L), the performance gap at R@50 between Ov-SGG and ZsO-SGG under VG caption supervision narrows to 2.27 points (12.98 vs 10.71), compared to a 3.55-point gap (7.61 vs 4.06) for Swin-T.

  10. Knowl 10 — Ablation on Caption Aggregation and Semantic Parser Complexity for SGG

    empirical result

    An ablation study evaluating different scene graph parsing setups on VS3VS^3 (Swin-T) trained on COCO caption supervision shows the impact of caption count and parser rule sophistication on downstream SGDET recall metrics on VG150:

    Captions Used Parser R@20 R@50 R@100
    Single caption Simple SG parser 5.07 6.25 7.36
    All captions (5/image) Simple SG parser 5.42 6.82 7.93
    All captions (5/image) Advanced SG parser 5.59 7.30 8.62

    Using all 5 human captions per image rather than a single caption yields a relative ~10% improvement in recall (R@50 increasing from 6.25 to 6.82), indicating that graph completeness derived from complementary text is important for supervision quality. Furthermore, using an advanced semantic parser capable of handling quantificational modifiers (e.g., "a lot of"), pronoun resolution (e.g., "it"), and plural nouns (e.g., "three men") boosts R@50 from 6.82 to 7.30 and R@100 from 7.93 to 8.62.

Coverage note — None. All primary contributions, including the model architecture, feature formulations, loss functions, grounding pipeline, open-vocabulary transfer mechanism, taxonomy, and comprehensive experimental results across supervised, language-supervised, and open-vocabulary SGG settings, have been fully captured.

References

  1. 1.Saeid Amiri, Kishan Chandan, and Shiqi Zhang. Reasoning with scene graphs for robot planning under partial observability. IEEE Robotics and Automation Letters, 7(2):5560–5567, 2022. 1, 2
  2. 2.Shizhe Chen, Qin Jin, Peng Wang, and Qi Wu. Say as you wish: Fine-grained control of image caption generation with abstract scene graphs. In CVPR, 2020. 1
  3. 3.Tianshui Chen, Weihao Yu, Riquan Chen, and Liang Lin. Knowledge-embedded routing network for scene graph generation. In CVPR, 2019. 6
  4. 4.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll'ar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015. 5
  5. 5.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In ECCV, 2020. 7
  6. 6.Jiuxiang Gu, Handong Zhao, Zhe Lin, Sheng Li, Jianfei Cai, and Mingyang Ling. Scene graph generation with external knowledge and image reconstruction. In CVPR, 2019. 1, 2
  7. 7.Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. Open-vocabulary object detection via vision and language knowledge distillation. In ICLR, 2021. 2, 3
  8. 8.Tao He, Lianli Gao, Jingkuan Song, and Yuan-Fang Li. Towards open-vocabulary scene graph generation with prompt-based finetuning. In ECCV, 2022. 3, 5, 7, 8
  9. 9.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICLR, 2021. 3
  10. 10.Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei. Image retrieval using scene graphs. In CVPR, 2015. 1, 2
  11. 11.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In ICCV, 2021. 3
  12. 12.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, 2019. 3
  13. 13.Boris Knyazev, Harm de Vries, Cătălina Cangea, Graham W Taylor, Aaron Courville, and Eugene Belilovsky. Generative compositional augmentations for scene graph prediction. In ICCV, 2021. 8
  14. 14.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. IJCV, 2017. 1, 2, 5
  15. 15.Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. Language-driven semantic segmentation. In ICLR, 2021. 2, 3
  16. 16.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. NeurIPS, 2021. 3
  17. 17.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, et al. Grounded language-image pre-training. In CVPR, 2022. 2, 3, 4, 6, 7
  18. 18.Manling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead, Di Lu, Heng Ji, and Shih-Fu Chang. Cross-media structured common space for multimedia event extraction. arXiv preprint arXiv:2005.02472, 2020. 1
  19. 19.Rongjie Li, Songyang Zhang, and Xuming He. Sgtr: End-to-end scene graph generation with transformer. In CVPR, 2022. 6
  20. 20.Rongjie Li, Songyang Zhang, Bo Wan, and Xuming He. Bipartite graph network with adaptive message passing for unbiased scene graph generation. In CVPR, 2021. 1, 2
  21. 21.Xingchen Li, Long Chen, Wenbo Ma, Yi Yang, and Jun Xiao. Integrating object-aware and interaction-aware knowledge for weakly supervised scene graph generation. In ACM MM, 2022. 2, 3, 5, 6, 7
  22. 22.Yehao Li, Yingwei Pan, Jingwen Chen, Ting Yao, and Tao Mei. X-modaler: A versatile and high-performance codebase for cross-modal analytics. In ACM MM, 2021. 1
  23. 23.Yehao Li, Yingwei Pan, Ting Yao, Jingwen Chen, and Tao Mei. Scheduled sampling in vision-language pretraining with decoupled encoder-decoder network. In AAAI, 2021. 3
  24. 24.Yongzhi Li, Duo Zhang, and Yadong Mu. Visual-semantic matching by exploring high-order attention and distraction. In CVPR, 2020. 1
  25. 25.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Doll'ar. Focal loss for dense object detection. In ICCV, 2017. 4
  26. 26.Xin Lin, Changxing Ding, Yibing Zhan, Zijian Li, and Dacheng Tao. Hl-net: Heterophily learning network for scene graph generation. In CVPR, 2022. 1, 2, 6
  27. 27.Hengyue Liu, Ning Yan, Masood Mortazavi, and Bir Bhanu. Fully convolutional scene graph generation. In CVPR, 2021. 6
  28. 28.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 3, 6
  29. 29.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. NeurIPS, 2019. 3
  30. 30.Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao, Hongyang Chao, and Tao Mei. Coco-bert: Improving video-language pre-training with contrastive cross-modal matching and denoising. In ACM MM, 2021. 3
  31. 31.Jiayuan Mao. Scenegraphparser, 2019. https : / / github . com / vacancy / SceneGraphParser (Access date: 2022-8-11). 3, 7
  32. 32.Victor Milewski, Marie-Francine Moens, and Iacer Calixto. Are scene graphs good enough to improve image captioning? arXiv preprint arXiv:2009.12313, 2020. 1
  33. 33.George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995. 3, 5, 6
  34. 34.Yingwei Pan, Yehao Li, Jianjie Luo, Jun Xu, Ting Yao, and Tao Mei. Auto-captions on gif: A large-scale video-sentence dataset for vision-language pre-training. In ACM Multimedia, 2022. 3
  35. 35.Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. X-linear attention networks for image captioning. In CVPR, 2020. 1
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICLR, 2021. 2, 3
  37. 37.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: towards real-time object detection with region proposal networks. IEEE TPAMI, 2016. 2
  38. 38.Brigit Schroeder and Subarna Tripathi. Structured query-based image retrieval using scene graphs. In CVPR Workshops, 2020. 1
  39. 39.Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning. Generating semantically precise scene graphs from textual descriptions for improved image retrieval. In Proceedings of the fourth workshop on vision and language, 2015. 3, 5, 7
  40. 40.Sahand Sharifzadeh, Sina Moayed Baharlou, and Volker Tresp. Classification by attention: Scene graph classification with prior knowledge. arXiv preprint arXiv:2011.10084, 2020. 1, 2
  41. 41.Jing Shi, Yiwu Zhong, Ning Xu, Yin Li, and Chenliang Xu. A simple baseline for weakly-supervised scene graph generation. In ICCV, 2021. 2, 3, 5, 6, 7
  42. 42.Motoharu Sonogashira, Masaaki Iiyama, and Yasutomo Kawanishi. Towards open-set scene graph generation with unknown objects. IEEE Access, 10:11574–11583, 2022. 2
  43. 43.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visual-linguistic representations. In ICLR, 2019. 3
  44. 44.Mohammed Suhail, Abhay Mittal, Behjat Siddiquie, Chris Broaddus, Jayan Eledath, Gerard Medioni, and Leonid Sigal. Energy-based learning for scene graph generation. In CVPR, 2021. 8
  45. 45.Rui Sun, Xuezhi Cao, Yan Zhao, Junchen Wan, Kun Zhou, Fuzheng Zhang, Zhongyuan Wang, and Kai Zheng. Multi-modal knowledge graphs for recommender systems. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, 2020. 1
  46. 46.Kaihua Tang. A scene graph generation codebase in pytorch, 2020. https://github.com/KaihuaTang/Scene-Graph-Benchmark.pytorch. 6
  47. 47.Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang. Unbiased scene graph generation from biased training. In CVPR, 2020. 1, 2, 5, 8
  48. 48.Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu. Learning to compose dynamic tree structures for visual contexts. In CVPR, 2019. 1, 2, 6, 8
  49. 49.Sijin Wang, Ruiping Wang, Ziwei Yao, Shiguang Shan, and Xilin Chen. Cross-modal scene graph matching for relationship-aware image-text retrieval. In WACV, 2020. 1
  50. 50.Danfei Xu, Yuke Zhu, Christopher B Choy, and Li Fei-Fei. Scene graph generation by iterative message passing. In CVPR, 2017. 1, 2, 5, 6, 7, 8
  51. 51.Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. Graph r-cnn for scene graph generation. In ECCV, 2018. 1, 2
  52. 52.Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai. Auto-encoding scene graphs for image captioning. In CVPR, 2019. 1
  53. 53.Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, and Hang Xu. Detclip: Dictionary-enriched visual-concept paralleled pre-training for open-world detection. arXiv preprint arXiv:2209.09407, 2022. 2, 3
  54. 54.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Exploring visual relationship for image captioning. In ECCV, 2018. 1
  55. 55.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Hierarchy parsing for image captioning. In ICCV, 2019. 1
  56. 56.Keren Ye and Adriana Kovashka. Linguistic structures as weak supervision for visual scene graph generation. In CVPR, 2021. 2, 5, 6, 7
  57. 57.Alireza Zareian, Svebor Karaman, and Shih-Fu Chang. Bridging knowledge graphs to generate scene graphs. In ECCV, 2020. 1, 2
  58. 58.Alireza Zareian, Svebor Karaman, and Shih-Fu Chang. Weakly supervised visual semantic parsing. In CVPR, 2020. 2, 7
  59. 59.Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. Open-vocabulary object detection using captions. In CVPR, 2021. 2, 3
  60. 60.Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi. Neural motifs: Scene graph parsing with global context. In CVPR, 2018. 1, 2, 6, 7, 8
  61. 61.Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, and Tat-Seng Chua. Visual translation embedding network for visual relation detection. In CVPR, 2017. 6
  62. 62.Ji Zhang, Kevin J Shih, Ahmed Elgammal, Andrew Tao, and Bryan Catanzaro. Graphical contrastive losses for scene graph parsing. In CVPR, 2019. 2, 6
  63. 63.Yong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, and Chang-Wen Chen. Exploring structure-aware transformer over interaction proposals for human-object interaction detection. In CVPR, 2022. 1
  64. 64.Yong Zhang, Yingwei Pan, Ting Yao, Rui Huang, Tao Mei, and Chang-Wen Chen. Boosting scene graph generation with visual relation saliency. ACM TOMM, 2023. 1
  65. 65.Yiwu Zhong, Jing Shi, Jianwei Yang, Chenliang Xu, and Yin Li. Learning to generate scene graph from natural language supervision. In ICCV, 2021. 2, 3, 5, 6, 7
  66. 66.Yiwu Zhong, Liwei Wang, Jianshu Chen, Dong Yu, and Yin Li. Comprehensive image captioning via scene graph decomposition. In ECCV, 2020. 1
  67. 67.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. Regionclip: Region-based language-image pretraining. In CVPR, 2022. 2, 3

Citation

MLA
Zhang, Y., et al. “Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 2915–24, https://doi.org/10.1109/CVPR52729.2023.00285.
APA
Zhang, Y., Pan, Y., Yao, T., Huang, R., Mei, T., & Chen, C.-W. (2023). Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2915–2924. https://doi.org/10.1109/CVPR52729.2023.00285
Chicago
Zhang, Y., Y. Pan, T. Yao, R. Huang, T. Mei, and C.-W. Chen. 2023. “Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2915–24. https://doi.org/10.1109/CVPR52729.2023.00285.
Harvard
Zhang, Y. et al. (2023) “Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 2915–2924. Available at: https://doi.org/10.1109/CVPR52729.2023.00285.
Vancouver
1. Zhang Y, Pan Y, Yao T, Huang R, Mei T, Chen C-W (2023) Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2915–2924

BibTeX

@inproceedings{Zhang_2023, title={Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00285}, DOI={10.1109/cvpr52729.2023.00285}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhang, Yong and Pan, Yingwei and Yao, Ting and Huang, Rui and Mei, Tao and Chen, Chang-Wen}, year={2023}, month=June, pages={2915–2924} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE