Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning
Juncheng LiJunlin XieLong QianLinchao ZhuSiliang TangFei WuYi YangYueting ZhuangXin Eric Wang
Introduces compositional temporal grounding benchmarks alongside a hierarchical variational cross-graph reasoning framework that aligns multi-level video and language structures to accurately localize video segments described by novel word combinations.
Video understanding systems are increasingly tasked with finding specific moments in long videos using natural language descriptions, a capability called temporal grounding. In practical applications, users describe events using varied combinations of words that an artificial intelligence model may not have encountered together during training. A reliable model must possess compositional generalization, which is the ability to understand novel combinations of familiar concepts, such as recognizing "throwing flowers" when it has only previously seen "throwing balls" and "smelling flowers." Existing benchmarks fail to properly evaluate this capability because their training and testing datasets contain virtually identical phrasing patterns.
The main objective of the article is to systematically evaluate how well leading temporal grounding systems generalize to novel word combinations and unseen words, and to introduce a novel modeling framework that improves this compositional reasoning.
To conduct this evaluation, the researchers developed a new benchmark by restructuring two standard video datasets into new test splits containing novel compositions and novel words, while eliminating video overlap between sets. They also developed a new framework called VISA (Variational Cross-Graph Reasoning). Instead of treating sentences and video clips as unstructured visual and textual blocks, VISA decomposes both videos and text into structured graphs across three levels: global events, local actions, and individual objects. It then applies a statistical alignment method called variational cross-graph learning to match elements across these two modalities.
The investigation produced four critical findings. First, existing leading models rely heavily on superficial correlations rather than actual sentence structure; when tested with randomly shuffled word orders, their performance barely degraded, showing they were essentially insensitive to syntax. Second, existing models suffered severe performance drops—falling by up to 20 percentage points—when tested on queries with novel word combinations. Third, the proposed VISA framework consistently outperformed all existing methods across every test split, achieving a 30.86% relative improvement in accuracy on Charades-CG and a 23.32% improvement on ActivityNet-CG for novel compositions. Fourth, error analysis revealed that verb-noun combinations were the most difficult for systems to resolve, requiring precise joint reasoning over both actions and objects.
These findings demonstrate that deploying standard video search models into operational settings poses operational risks, as systems may fail unexpectedly when users formulate real-world queries in unobserved ways. By explicitly modeling structured semantics, organizations can develop video search engines that are robust, predictable, and capable of understanding complex human instructions.
Organizations developing video intelligence solutions should adopt structured, multi-level semantic architectures rather than monolithic feature representations, and they should evaluate future systems against benchmarks specifically designed for compositional generalization. Future research should prioritize improving semantic sensitivity to subtle linguistic modifiers, such as adverbs, where current models still struggle.
- Paper: Scene Graph Generation by Iterative Message Passing, Danfei Xu et al. (2017). Introduces foundational scene graph generation via message passing, providing the structured relational representation that VISA extends to multi-level video-text grounding.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). Introduces the ActivityNet Captions dataset and the core formulation of temporally localizing natural language events in untrimmed video, which the source reorganizes into a compositional benchmark.
- Paper: The “Something Something” Video Database for Learning and Evaluating Visual Common Sense, Raghav Goyal et al. (2017). Pioneers the evaluation of fine-grained, compositional action-object interactions in video via structured caption templates.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Establishes baseline cross-modal transformer alignment and phrase grounding between text tokens and visual regions.
- Paper: VL-BERT: Pre-training of Generic Visual-Linguistic Representations, Weijie Su et al. (2019). Demonstrates generic visual-linguistic pre-training for aligning regions and phrases, foundational to cross-modal graph alignment.
- Paper: Sequence to Sequence -- Video to Text, Subhashini Venugopalan et al. (2015). Provides early sequence-to-sequence modeling for video-to-text mapping and highlights the necessity of temporal sequence modeling.
- Paper: TRACE: Temporal Grounding Video LLM via Causal Event Modeling, Yongxin Guo 0001 et al. (2025). Extends structured temporal grounding into multimodal large language models using causal event modeling and explicit timestamp-saliency triplets.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). Applies grounded temporal tuning and timestamp supervision to enhance long-form video reasoning in multimodal large language models.
- Paper: HierVL: Learning Hierarchical Video-Language Embeddings, Kumar Ashutosh et al. (2023). Advances multi-level video-language representations by hierarchically connecting fine-grained clip narrations with video-level summaries.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). Expands dynamic video understanding assessment across diverse multi-modal perception and cognition tasks, complementing compositional grounding benchmarks.
- Paper: Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space, Yong Zhang et al. (2023). Builds upon structured visual-semantic graphs to generate open-vocabulary scene representations using pre-trained vision-language spaces.
- Paper: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation, Kaiyue Sun 0001 et al. (2025). Investigates compositional generalization across action and attribute bindings in the complementary direction of text-to-video generation.
