Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space
Yong ZhangYingwei PanTing YaoRui HuangTao MeiChang Wen Chen
Proposes a scene graph generation framework that uses pre-trained visual-semantic spaces to detect novel objects and relations from image captions without requiring expensive manual annotations.
Scene graph generation converts visual images into structured representations by identifying objects as nodes and their interactions as labeled edges. This capability is critical for fine-grained computer vision tasks, including automated image captioning, cross-modal retrieval, and autonomous robot decision-making. However, real-world deployment faces two significant hurdles: training models requires labor-intensive and costly manual bounding-box annotations, and conventional systems are restricted to a pre-defined, closed set of categories, causing them to fail when encountering unfamiliar objects in dynamic environments.
The article demonstrates an effective methodology called Visual-Semantic Space for Scene graph generation, or VS3, to overcome these bottlenecks. The primary objective is to evaluate how pre-trained vision-language models can be transferred to eliminate the need for expensive manual labels while enabling the system to recognize novel, unseen visual categories and their relationships in an open-vocabulary setting.
To achieve this, the authors leverage Grounded Language-Image Pre-training, an established foundational model that maps visual regions and descriptive text phrases into a shared visual-semantic space. The framework parses raw descriptive text captions into semantic relationships, automatically matches text entities to image regions using the pre-trained space, and uses these grounded graphs as low-cost supervision data. To infer relationships, a lightweight module combines visual and spatial coordinates between detected objects. The model freezes the underlying image and text encoders to prevent performance degradation and updates only the cross-modal and relationship-prediction components across benchmark evaluations on the Visual Genome dataset.
The findings show that VS3 sets new performance benchmarks across multiple configurations. In fully supervised settings, the larger variant achieved top-50 recall scores of 36.6 percent, outperforming prior state-of-the-art models by up to 3.4 percentage points while using a simpler architecture. Under language-only supervision, it scored 29.81 percent on top-50 recall, more than doubling earlier weakly supervised approaches and matching or exceeding several fully supervised systems. In the open-vocabulary setting involving 30 percent previously unseen object categories, the system demonstrated viable end-to-end detection and relationship mapping directly from raw images, a benchmark previously unsupported by standard two-stage object detectors.
These results demonstrate that leveraging pre-trained multi-modal foundations drastically cuts the financial and operational costs associated with manual data labeling while substantially lowering the risk of model failure in unconstrained environments. Systems capable of processing novel categories without retraining significantly improve adaptability and safety for automated perception pipelines, robotics, and content retrieval workflows.
Organizations developing vision-language solutions should consider adopting pre-trained grounding foundations and weakly supervised pipelines rather than relying exclusively on custom, manually annotated datasets. Future development should focus on testing these pipelines with richer, region-level descriptive text to minimize the performance drop caused by sparse captions and cross-domain data shifts.
The primary limitation highlighted in the analysis is that language-derived supervision tends to capture simpler, high-frequency relational predicates, such as basic spatial prepositions, compared to human-curated datasets. While confidence in the model's overall generalization and detection accuracy is high, practitioners should exercise caution when deploying systems in domains that require highly specialized or nuanced relationship classifications without fine-tuning.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). This seminal paper introduces iterative message passing for visually grounded scene graph generation, establishing the foundational problem setup and formulation that VS3 builds upon.
- Paper: Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations, Ranjay Krishna et al. (2016). It introduces the Visual Genome dataset and the formal ontology of objects, attributes, and relationships used to evaluate scene graph generation models, including VS3.
- Paper: Structured Sparse R-CNN for Direct Scene Graph Generation, Yao Teng et al. (2022). It presents a direct set-prediction framework for scene graph generation that helps frame the standard supervised baseline architectures and relational modules referenced by VS3.
- Paper: Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks, Xiujun Li et al. (2020). This work establishes the paradigm of using object tags and grounded image regions to align visual-semantic representations for vision-language tasks.
- Paper: Deep visual-semantic alignments for generating image descriptions, Andrej Karpathy et al. (2015). It introduces foundational deep visual-semantic alignments between image regions and text phrases, which underpins modern grounding and vision-language pre-training spaces.
- Paper: LXMERT: Learning Cross-Modality Encoder Representations from Transformers, Hao Tan et al. (2019). It provides key background on cross-modality transformer representations that capture both object relationships and visual-text alignments.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). It demonstrates a unified architecture aligning text tokens and image region features for multimodal reasoning and phrase grounding.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This survey provides a comprehensive synthesis of vision-language models and adaptation techniques across diverse downstream visual recognition tasks, situating open-vocabulary grounding in the broader literature.
- Paper: GSVA: Generalized Segmentation via Multimodal Large Language Models, Zhuofan Xia et al. (2024). It extends grounded multimodal reasoning by using multimodal large language models to resolve complex referring expressions, multi-target segmentation, and referent rejection.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). It builds on localized visual-language grounding by introducing visual chain-of-thought bounding box reasoning to systematically interpret fine-grained spatial and textual scene details.
- Paper: OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views, Francis Engelmann et al. (2024). It takes open-vocabulary visual-semantic feature alignment from 2D images and extends it to continuous 3D neural scene representations.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). It evaluates the complementary challenge of compositional spatial intelligence and relational consistency in text-to-image generative models.
