Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng LiJunlin XieLong QianLinchao ZhuSiliang TangFei WuYi YangYueting ZhuangXin Eric Wang

article2022CVPR84 citations

Introduces compositional temporal grounding benchmarks alongside a hierarchical variational cross-graph reasoning framework that aligns multi-level video and language structures to accurately localize video segments described by novel word combinations.

Listen

Video understanding systems are increasingly tasked with finding specific moments in long videos using natural language descriptions, a capability called temporal grounding. In practical applications, users describe events using varied combinations of words that an artificial intelligence model may not have encountered together during training. A reliable model must possess compositional generalization, which is the ability to understand novel combinations of familiar concepts, such as recognizing "throwing flowers" when it has only previously seen "throwing balls" and "smelling flowers." Existing benchmarks fail to properly evaluate this capability because their training and testing datasets contain virtually identical phrasing patterns.

The main objective of the article is to systematically evaluate how well leading temporal grounding systems generalize to novel word combinations and unseen words, and to introduce a novel modeling framework that improves this compositional reasoning.

To conduct this evaluation, the researchers developed a new benchmark by restructuring two standard video datasets into new test splits containing novel compositions and novel words, while eliminating video overlap between sets. They also developed a new framework called VISA (Variational Cross-Graph Reasoning). Instead of treating sentences and video clips as unstructured visual and textual blocks, VISA decomposes both videos and text into structured graphs across three levels: global events, local actions, and individual objects. It then applies a statistical alignment method called variational cross-graph learning to match elements across these two modalities.

The investigation produced four critical findings. First, existing leading models rely heavily on superficial correlations rather than actual sentence structure; when tested with randomly shuffled word orders, their performance barely degraded, showing they were essentially insensitive to syntax. Second, existing models suffered severe performance drops—falling by up to 20 percentage points—when tested on queries with novel word combinations. Third, the proposed VISA framework consistently outperformed all existing methods across every test split, achieving a 30.86% relative improvement in accuracy on Charades-CG and a 23.32% improvement on ActivityNet-CG for novel compositions. Fourth, error analysis revealed that verb-noun combinations were the most difficult for systems to resolve, requiring precise joint reasoning over both actions and objects.

These findings demonstrate that deploying standard video search models into operational settings poses operational risks, as systems may fail unexpectedly when users formulate real-world queries in unobserved ways. By explicitly modeling structured semantics, organizations can develop video search engines that are robust, predictable, and capable of understanding complex human instructions.

Organizations developing video intelligence solutions should adopt structured, multi-level semantic architectures rather than monolithic feature representations, and they should evaluate future systems against benchmarks specifically designed for compositional generalization. Future research should prioritize improving semantic sensitivity to subtle linguistic modifiers, such as adverbs, where current models still struggle.

Cover for Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Abstract

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language descriptions, temporal grounding allows activity grounding beyond pre-defined classes and has received increasing attention in recent years. The semantic diversity is rooted in the principle of compositionality in linguistics, where novel semantics can be systematically described by combining known words in novel ways (compositional generalization). However, current temporal grounding datasets do not specifically test for the compositional generalizability. To systematically measure the compositional generalizability of temporal grounding models, we introduce a new Compositional Temporal Grounding task and construct two new dataset splits, i.e., Charades-CG and ActivityNet-CG. Evaluating the state-of-the-art methods on our new dataset splits, we empirically find that they fail to generalize to queries with novel combinations of seen words. To tackle this challenge, we propose a variational cross-graph reasoning framework that explicitly decomposes video and language into multiple structured hierarchies and learns fine-grained semantic correspondence among them. Experiments illustrate the superior compositional generalizability of our approach. The repository of this work is at https://github.com/YYJMJC/Compositional-Temporal-Grounding.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Compositional Temporal Grounding
  • 3.1. Problem Formulation
  • 3.2. Dataset Re-splitting
  • 4. Method
  • 4.1. Hierarchical Semantic Graph
  • 4.2. Variational Cross-Graph Correspondence
  • 4.3. Optimization
  • 5. Experiments
  • 5.1. Benchmarking the SOTA Methods
  • 5.2. Results on Compositional Temporal Grounding
  • 5.3. In-Depth Analysis
  • 5.4. Qualitative Analysis
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — Compositional Temporal Grounding Benchmark and Dataset Re-splitting Protocol

    experimental setup

    The Compositional Temporal Grounding (CTG) task evaluates a model's ability to temporally localize a target video segment given a query sentence containing unseen combinations of known words (novel compositions) or unseen words in context (novel words).

    To benchmark this capability, the Charades-STA and ActivityNet Captions datasets are re-organized into Charades-CG and ActivityNet-CG, respectively. The dataset construction protocol follows these steps:

    1. Instances predictable solely from video context are removed after merging original train and test sets.
    2. Query words are lemmatized and tagged for part-of-speech (nouns, adjectives, verbs, adverbs, prepositions) using dependency parsing, forming 5 composition types: verb-noun, adjective-noun, noun-noun, verb-adverb, and preposition-noun.
    3. For each composition type, a 2D co-occurrence matrix (first component ×\times second component) is constructed. Elements are sampled across rows and columns to ensure all individual constituent words appear in the training split, while held-out pairs form the novel-composition split.
    4. Held-out words form the novel-word split, and remaining seen compositions from the same test videos form the test-trivial split.
    5. Query assignments are partitioned such that there is zero video overlap between the training set and any testing split.
    Dataset Split Videos Queries
    Charades-CG Training 3555 8281
    Novel-Composition 2480 3442
    Novel-Word 588 703
    Test-Trivial 1689 3096
    ActivityNet-CG Training 9659 36724
    Novel-Composition 4202 12028
    Novel-Word 2011 3944
    Test-Trivial 4775 15712
  2. Knowl 2 — Hierarchical Semantic Graph Construction via Contextualized Learning and Aggregation

    model/method

    Given an untrimmed video VV and a query sentence QQ, both modalities are decomposed into unified three-level semantic hierarchies (global events, local actions, and atomic objects).

    Video Graph Initialization: Video VV is segmented into TT segments with C3D features Vt={fit}i=1KV_t = \{f_i^t\}_{i=1}^K. Off-the-shelf detectors extract N1N_1 object nodes {st,io}i=1N1∈Rd×N1\{s_{t,i}^o\}_{i=1}^{N_1} \in \mathbb{R}^{d \times N_1} and N2N_2 action nodes {st,ia}i=1N2∈Rd×N2\{s_{t,i}^a\}_{i=1}^{N_2} \in \mathbb{R}^{d \times N_2} per segment, initialized by GloVe embeddings of their category labels.

    Language Graph Initialization: Semantic Role Labeling (SRL) decomposes query QQ into predicates as action nodes {cia}i=1L2∈Rd×L2\{c_i^a\}_{i=1}^{L_2} \in \mathbb{R}^{d \times L_2} and arguments (noun phrases) as object nodes {ci,jo}j=1L1∈Rd×L1\{c_{i,j}^o\}_{j=1}^{L_1} \in \mathbb{R}^{d \times L_1}, initialized with GloVe embeddings.

    Semantic-Contextualized Learning (SCL): For three undirected edge types R={action-action,action-object,object-object}R = \{\text{action-action}, \text{action-object}, \text{object-object}\}, a relation-aware graph convolution updates node representations. For a video semantic node si∈{Sa,So}s_i \in \{S^a, S^o\}: α~ijr=(Wrsˉi)T(Wrsˉj),αijr=exp⁡(α~ijr)∑j∈Nirexp⁡(α~ijr)\tilde{\alpha}_{ij}^r = (W^r \bar{s}_i)^T (W^r \bar{s}_j), \quad \alpha_{ij}^r = \frac{\exp(\tilde{\alpha}_{ij}^r)}{\sum_{j \in \mathcal{N}_i^r} \exp(\tilde{\alpha}_{ij}^r)} s^i=∑r∈R∑j∈Nirαijr⋅(Ursˉj)\hat{s}_i = \sum_{r \in R} \sum_{j \in \mathcal{N}_i^r} \alpha_{ij}^r \cdot (U^r \bar{s}_j) where Nir\mathcal{N}_i^r is the neighborhood of sis_i under relation rr, and Wr,Ur∈Rd×dW^r, U^r \in \mathbb{R}^{d \times d} are relation-specific projection matrices. Stacking MM layers produces contextualized node representations S={si}i=1Nv∈RNv×dS = \{s_i\}_{i=1}^{N_v} \in \mathbb{R}^{N_v \times d} and language representations C={ci}i=1Ns∈RNs×dC = \{c_i\}_{i=1}^{N_s} \in \mathbb{R}^{N_s \times d}.

    Visual-Contextualized Learning (VCL): To infuse raw visual features into video nodes, a visual filter gjig_j^i is computed for each frame fjf_j in segment Vi={fj}j=1KV_i = \{f_j\}_{j=1}^K: gji=σ(Wg[si;fˉ;fj]+bg),fj′=fj⊙gjig_j^i = \sigma(W^g [s_i; \bar{f}; f_j] + b_g), \quad f_j' = f_j \odot g_j^i Fi=MaxPool⁡(f1′,…,fK′),si←Wv[si;Fi]F_i = \operatorname{MaxPool}(f_1', \dots, f_K'), \quad s_i \leftarrow W^v [s_i; F_i] where fˉ\bar{f} is the average-pooled segment feature, σ\sigma is the sigmoid function, and ⊙\odot is the Hadamard product.

    Hierarchical Semantics Aggregation (HSA): Global event nodes are initialized as NpN_p learnable queries {pi}i=1Np\{p_i\}_{i=1}^{N_p} and updated via multi-head graph self-attention over lower-level semantic nodes {sj}j=1Nv\{s_j\}_{j=1}^{N_v}: p~i=∑j=1Nvαijesj,αije=exp⁡((W1epi)T(W2esj))∑j=1Nvexp⁡((W1epi)T(W2esj))\tilde{p}_i = \sum_{j=1}^{N_v} \alpha_{ij}^e s_j, \quad \alpha_{ij}^e = \frac{\exp((W_1^e p_i)^T (W_2^e s_j))}{\sum_{j=1}^{N_v} \exp((W_1^e p_i)^T (W_2^e s_j))} Event nodes are combined with action and object nodes to produce unified hierarchical graphs M={mi}i=1Nm∈Rd×NmM = \{m_i\}_{i=1}^{N_m} \in \mathbb{R}^{d \times N_m} for video and H={hi}i=1Nh∈Rd×NhH = \{h_i\}_{i=1}^{N_h} \in \mathbb{R}^{d \times N_h} for language.

  3. Knowl 3 — Variational Cross-Graph Correspondence Learning Framework

    model/method

    Because fine-grained ground-truth correspondences between video graph M∈Rd×NmM \in \mathbb{R}^{d \times N_m} and language graph H∈Rd×NhH \in \mathbb{R}^{d \times N_h} are unannotated, the semantic alignment matrix is formulated as a latent variable z∈RNm×Nhz \in \mathbb{R}^{N_m \times N_h}. The objective maximizes the Evidence Lower Bound (ELBO) on the marginal log-likelihood log⁡p(Y∣M,H)\log p(Y|M, H) of target interval YY: LELBO(ϕ,θ)=Eqϕ(z∣M,H,Y)[log⁡pθ(Y∣M,H,z)]−KL⁡(qϕ(z∣M,H,Y) ∥ pθ(z∣M,H))\mathcal{L}^{\text{ELBO}}(\phi, \theta) = \mathbb{E}_{q_\phi(z|M, H, Y)} [\log p_\theta(Y|M, H, z)] - \operatorname{KL}(q_\phi(z|M, H, Y) \,\|\, p_\theta(z|M, H)) where:

    • pθ(z∣M,H)p_\theta(z|M, H) is the prior model parameterized by θ\theta, computed from graph-convolved nodes m~i,h~j\tilde{m}_i, \tilde{h}_j and global sentence feature q∈Rdq \in \mathbb{R}^d: z~ij=(W1sm~i)T(W2sq⊙h~j),zij=exp⁡(z~ij)∑j=1Nhexp⁡(z~ij)\tilde{z}_{ij} = (W_1^s \tilde{m}_i)^T (W_2^s q \odot \tilde{h}_j), \quad z_{ij} = \frac{\exp(\tilde{z}_{ij})}{\sum_{j=1}^{N_h} \exp(\tilde{z}_{ij})}
    • qϕ(z∣M,H,Y)q_\phi(z|M, H, Y) is the approximate posterior model parameterized by ϕ\phi, utilizing visual guidance m∗m^* (mean-pooled over video nodes falling inside ground-truth segment interval YY): z~ij=(W3sm∗⊙m~i)T(W4sq⊙h~j),zij=exp⁡(z~ij)∑j=1Nhexp⁡(z~ij)\tilde{z}_{ij} = (W_3^s m^* \odot \tilde{m}_i)^T (W_4^s q \odot \tilde{h}_j), \quad z_{ij} = \frac{\exp(\tilde{z}_{ij})}{\sum_{j=1}^{N_h} \exp(\tilde{z}_{ij})}
    • KL⁡(⋅∥⋅)\operatorname{KL}(\cdot\|\cdot) is computed row-wise over the correspondence probability distributions.
    • During training, zz is sampled/generated from the posterior qϕq_\phi; during inference, zz is generated using the learned prior pθp_\theta without ground-truth access.
  4. Knowl 4 — Cross-Graph Convolution and Likelihood Moment Prediction in VISA

    model/method

    Cross-modal interaction and moment boundary prediction in VISA proceed through level-wise cross-graph convolution and attentive segment fusion:

    Cross-Graph Convolution: Convolution occurs between corresponding hierarchical levels k∈{e,a,o}k \in \{e, a, o\} (events, actions, objects). For video node mikm_i^k and language nodes NHk\mathcal{N}_H^k: αijh2m=exp⁡((W1cmik)T(W2chjk))∑j∈NHkexp⁡((W1cmik)T(W2chjk))\alpha_{ij}^{h2m} = \frac{\exp((W_1^c m_i^k)^T (W_2^c h_j^k))}{\sum_{j \in \mathcal{N}_H^k} \exp((W_1^c m_i^k)^T (W_2^c h_j^k))} m~ik=(1−βik)⊙mik+βik⊙∑j∈NHkαijh2m⋅hjk\tilde{m}_i^k = (1 - \beta_i^k) \odot m_i^k + \beta_i^k \odot \sum_{j \in \mathcal{N}_H^k} \alpha_{ij}^{h2m} \cdot h_j^k where βik=σ(Ugmik+b)\beta_i^k = \sigma(U^g m_i^k + b) gates cross-modal information flow. Language nodes H~\tilde{H} are updated analogously.

    Likelihood Grounding Model pθ(Y∣M,H,z)p_\theta(Y|M, H, z):

    1. Latent correspondence z∈RNm×Nhz \in \mathbb{R}^{N_m \times N_h} aligns the language graph to the video graph: M′=zH~∈Rd×NmM' = z \tilde{H} \in \mathbb{R}^{d \times N_m}, followed by joint projection MJ=WJ[M~;M′]∈Rd×NmM^J = W^J [\tilde{M}; M'] \in \mathbb{R}^{d \times N_m}.
    2. Segment representations X={xt}t=1T∈Rd×TX = \{x_t\}_{t=1}^T \in \mathbb{R}^{d \times T} are refined via multi-head cross-attention using XX as queries and MJM^J as keys/values: X∗=MultiAttn⁡(X,MJ,MJ)X^* = \operatorname{MultiAttn}(X, M^J, M^J).
    3. Attentive pooling guided by sentence representation qq aggregates X∗X^* into video vector v∗v^*: v∗=∑i=1Tαiqxi∗,αiq=exp⁡((W1qq)T(W2qxi∗))∑i=1Texp⁡((W1qq)T(W2qxi∗))v^* = \sum_{i=1}^T \alpha_i^q x_i^*, \quad \alpha_i^q = \frac{\exp((W_1^q q)^T (W_2^q x_i^*))}{\sum_{i=1}^T \exp((W_1^q q)^T (W_2^q x_i^*))}
    4. An MLP regresses normalized start and end timestamps: (ts,te)=MLP⁡(v∗)(t^s, t^e) = \operatorname{MLP}(v^*). Optimization minimizes the smooth L1L_1 distance between (ts,te)(t^s, t^e) and normalized ground truth (t^s,t^e)∈[0,1](\hat{t}^s, \hat{t}^e) \in [0, 1].
  5. Knowl 5 — Comparative Benchmarking of SOTA Temporal Grounding Methods on Compositional Splits

    empirical result

    Existing temporal grounding models (proposal-based, proposal-free, RL-based, and weakly-supervised) experience severe performance degradation when moving from standard/trivial splits to novel-composition and novel-word splits. VISA substantially outperforms previous state-of-the-art methods across all splits on both Charades-CG and ActivityNet-CG.

    Charades-CG Test-Trivial Novel-Composition Novel-Word
    Method IoU=0.5 IoU=0.7 mIoU IoU=0.5 IoU=0.7 mIoU IoU=0.5 IoU=0.7 mIoU
    WSSL 15.33 5.46 18.31 3.61 1.21 8.26 2.79 0.73 7.92
    TSP-PRL 39.86 21.07 38.41 16.30 2.04 13.52 14.83 2.61 14.03
    TMN 18.75 8.16 19.82 8.68 4.07 10.14 9.43 4.96 11.23
    2D-TAN 48.58 26.49 44.27 30.91 12.23 29.75 29.36 13.21 28.47
    LGI 49.45 23.80 45.01 29.42 12.73 30.09 26.48 12.47 27.62
    VLSNet 45.91 19.80 41.63 24.25 11.54 31.43 25.60 10.07 30.21
    VISA (Ours) 53.20 26.52 47.11 45.41 22.71 42.03 42.35 20.88 40.18
    ActivityNet-CG Test-Trivial Novel-Composition Novel-Word
    Method IoU=0.5 IoU=0.7 mIoU IoU=0.5 IoU=0.7 mIoU IoU=0.5 IoU=0.7 mIoU
    WSSL 11.03 4.14 15.07 2.89 0.76 7.65 3.09 1.13 7.10
    TSP-PRL 34.27 18.80 37.05 14.74 1.43 12.61 18.05 3.15 14.34
    TMN 16.82 7.01 17.13 8.74 4.39 10.08 9.93 5.12 11.38
    2D-TAN 44.50 26.03 42.12 22.80 9.95 28.49 23.86 10.37 28.88
    LGI 43.56 23.29 41.37 23.21 9.02 27.86 23.10 9.03 26.95
    VLSNet 39.27 23.12 42.51 20.21 9.18 29.07 21.68 9.94 29.58
    VISA (Ours) 47.13 29.64 44.02 31.51 16.73 35.85 30.14 15.90 35.13

    On the Novel-Composition split, VISA achieves relative mIoU gains over the strongest baseline (VLSNet/LGI) of 30.86% on Charades-CG (42.03% vs. 31.43%) and 23.32% on ActivityNet-CG (35.85% vs. 29.07%).

  6. Knowl 6 — Ablation Analysis of VISA Architectural Modules

    empirical result

    An ablation study evaluating the individual components of the VISA framework on the Novel-Composition (Comp) and Novel-Word (Word) splits (measured by R@1, IoU=0.5) demonstrates the necessity of every sub-module:

    Method Charades-CG ActivityNet-CG
    Comp Word Comp Word
    1 w/o SCL (Semantic-Contextualized Learning) 43.75 40.16 29.03 29.41
    2 w/o VCL (Visual-Contextualized Learning) 42.26 38.62 29.34 28.09
    3 w/o HSA (Hierarchical Semantics Aggregation) 44.22 41.09 30.29 29.31
    4 w/o VCC (Variational Cross-Graph Correspondence) 41.08 37.54 27.32 26.37
    5 Detection-based (Raw detection/SRL features) 12.97 11.70 10.92 10.07
    6 Full VISA 45.41 42.35 31.51 30.14

    Key takeaways:

    • Removing Variational Cross-Graph Correspondence (VCC, replaced with direct cross-modal self-attention) leads to the largest drop (4.33% on Charades-CG Comp, 4.19% on ActivityNet-CG Comp), showing that unconstrained fusion disrupts fine-grained cross-modal semantic correspondence.
    • The detection-only baseline performs poorly (≈10–13% \approx 10\text{--}13\%), showing that raw symbolic detection labels alone are insufficient without contextual graph reasoning.
  7. Knowl 7 — Word Order Sensitivity Analysis of Temporal Grounding Models

    empirical result

    Word order sensitivity measures the relative performance degradation on R@1, IoU=0.5 when evaluating a model on queries with randomly permuted word order compared to original word order: Sensitivity=Perforig−PerfshuffledPerforig×100%\text{Sensitivity} = \frac{\text{Perf}_{\text{orig}} - \text{Perf}_{\text{shuffled}}}{\text{Perf}_{\text{orig}}} \times 100\%

    Method Charades-CG ActivityNet-CG
    Trivial Comp Word Trivial Comp Word
    2D-TAN 0.41 0.52 0.43 0.29 0.30 0.41
    LGI 0.28 0.23 0.16 0.31 0.22 0.19
    VLSNet 0.07 0.24 0.10 0.24 0.31 0.48
    VISA 24.14 29.80 33.97 22.09 27.60 31.89
    w/o SCL 19.64 24.31 29.72 18.07 24.64 28.73
    w/o VCC 21.32 26.73 30.88 20.15 25.46 29.79

    State-of-the-art temporal grounding baselines exhibit near-zero word order sensitivity (<0.55%<0.55\%), revealing that they rely primarily on superficial bag-of-words correlations. In contrast, VISA achieves 22.09%−33.97%22.09\% - 33.97\% sensitivity, confirming that its structured hierarchical graphs accurately capture linguistic syntax and compositionality.

  8. Knowl 8 — Compositional Grounding Generalization Across Linguistic Composition Types

    empirical result

    Model performance (R@1, IoU=0.5) varies across linguistic composition categories on both Charades-CG and ActivityNet-CG:

    Composition Type Charades-CG ActivityNet-CG
    w/o VCC w/o SCL VISA w/o VCC w/o SCL VISA
    Verb-Noun 36.56 38.82 41.37 24.41 26.32 28.89
    Adj-Noun 42.17 44.04 45.06 26.76 28.31 30.67
    Noun-Noun 40.38 42.56 43.41 29.51 30.20 33.93
    Verb-Adv 43.81 46.37 47.83 31.08 33.46 35.60
    Prep-Noun 44.12 47.86 48.61 34.78 36.03 37.35

    Across all models and both datasets, Verb-Noun compositions exhibit the lowest grounding accuracy (e.g., 41.37% on Charades-CG and 28.89% on ActivityNet-CG for VISA). This indicates that jointly identifying actions and objects in video and reasoning over their compositional interaction is the most difficult form of compositional generalization.

  9. Knowl 9 — Limitation in Discerning Subtle Adverbial Modifiers

    limitation

    The VISA framework exhibits failure cases when required to distinguish subtle semantic differences introduced by directional or manner adverbs modifying identical verbs (for example, differentiating "fly close" versus "fly away"). The current visual and linguistic graph parsing abstractions do not sufficiently separate fine-grained adverbial visual semantics.

Coverage note — None was omitted; all contributed aspects (task formulation, dataset splits, graph construction, variational cross-graph reasoning, comparative benchmarking, ablations, word-order sensitivity, and composition breakdown) are represented.

References

  1. 1.Jacob Andreas. Good-enough compositional data augmentation. arXiv preprint arXiv:1904.09545, 2019.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision, pages 213–229. Springer, 2020.
  3. 3.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  4. 4.Noam Chomsky. Syntactic structures. De Gruyter Mouton, 2009.
  5. 5.Yadong Ding, Yu Wu, Chengyue Huang, Siliang Tang, Fei Wu, Yi Yang, Wenwu Zhu, and Yueting Zhuang. Nap: Neural architecture search with pruning. Neurocomputing, 2022.
  6. 6.Yadong Ding, Yu Wu, Chengyue Huang, Siliang Tang, Yi Yang, Longhui Wei, Yueting Zhuang, and Qi Tian. Learning to learn by jointly optimizing neural architecture and weights. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  7. 7.Xuguang Duan, Wenbing Huang, Chuang Gan, Jingdong Wang, Wenwu Zhu, and Junzhou Huang. Weakly supervised dense event captioning in videos. arXiv preprint arXiv:1812.03849, 2018.
  8. 8.Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. Convolutional two-stream network fusion for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1933–1941, 2016.
  9. 9.Jerry A Fodor and Zenon W Pylyshyn. Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2):3–71, 1988.
  10. 10.Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on computer vision, pages 5267–5275, 2017.
  11. 11.Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. Allennlp: A deep semantic natural language processing platform. arXiv preprint arXiv:1803.07640, 2018.
  12. 12.Jonathan Gordon, David Lopez-Paz, Marco Baroni, and Diane Bouchacourt. Permutation equivariant models for compositional generalization in language. In International Conference on Learning Representations, 2019.
  13. 13.Madeleine Grunde-McLaughlin, Ranjay Krishna, and Maneesh Agrawala. Agqa: A benchmark for compositional spatio-temporal reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11287–11297, 2021.
  14. 14.Jiannan Guo, Haochen Shi, Yangyang Kang, Kun Kuang, Siliang Tang, Zhuoren Jiang, Changlong Sun, Fei Wu, and Yueting Zhuang. Semi-supervised active learning for semi-supervised models: exploit adversarial examples with graph-based virtual labels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2896–2905, 2021.
  15. 15.Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2901–2910, 2017.
  16. 16.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  17. 17.Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision, pages 706–715, 2017.
  18. 18.Brenden Lake and Marco Baroni. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In International conference on machine learning, pages 2873–2882. PMLR, 2018.
  19. 19.Brenden M Lake. Compositional generalization through meta sequence-to-sequence learning. arXiv preprint arXiv:1906.05381, 2019.
  20. 20.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436–444, 2015.
  21. 21.Juncheng Li, Siliang Tang, Fei Wu, and Yueting Zhuang. Walking with mind: Mental imagery enhanced embodied qa. In Proceedings of the 27th ACM International Conference on Multimedia, pages 1211–1219, 2019.
  22. 22.Juncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi, Xuanwen Huang, Fei Wu, Yi Yang, and Yueting Zhuang. Adaptive hierarchical graph reasoning with semantic coherence for video-and-language inference. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1867–1877, 2021.
  23. 23.Juncheng Li, Xin Wang, Siliang Tang, Haizhou Shi, Fei Wu, Yueting Zhuang, and William Yang Wang. Unsupervised reinforcement learning of transferable meta-skills for embodied navigation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12123–12132, 2020.
  24. 24.Mengze Li, Ming Kong, Kun Kuang, Qiang Zhu, and Fei Wu. Multi-task attribute-fusion model for fine-grained image recognition. In Optoelectronic Imaging and Multimedia Technology VII, 2020.
  25. 25.Mengze Li, Kun Kuang, Qiang Zhu, Xiaohong Chen, Qing Guo, and Fei Wu. Ib-m: A flexible framework to align an interpretable model and a black-box model. In BIBM, 2020.
  26. 26.Mengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang, Zhou Zhao, Jiaxu Miao, Wenqiao Zhang, Wenming Tan, Jin Wang, Peng Wang, Shiliang Pu, and Fei Wu. End-to-end modeling via information tree for one-shot natural language spatial video grounding.
  27. 27.Bingbin Liu, Serena Yeung, Edward Chou, De-An Huang, Li Fei-Fei, and Juan Carlos Niebles. Temporal modular networks for retrieving complex compositional activities in videos. In Proceedings of the European Conference on Computer Vision (ECCV), pages 552–568, 2018.
  28. 28.Shugao Ma, Leonid Sigal, and Stan Sclaroff. Learning activity progression in lstms for activity detection and early detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1942–1950, 2016.
  29. 29.Massimiliano Mancini, Muhammad Ferjad Naeem, Yongqin Xian, and Zeynep Akata. Open world compositional zero-shot learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5222–5230, 2021.
  30. 30.Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell. Something-else: Compositional action recognition with spatial-temporal interaction networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1049–1059, 2020.
  31. 31.Richard Montague et al. Universal grammar. 1974, pages 222–46, 1970.
  32. 32.Jonghwan Mun, Minsu Cho, and Bohyung Han. Local-global video-text interactions for temporal grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10810–10819, 2020.
  33. 33.Mitja Nikolaus, Mostafa Abdou, Matthew Lamm, Rahul Aralikatte, and Desmond Elliott. Compositional generalization in image captioning. arXiv preprint arXiv:1909.04402, 2019.
  34. 34.Maxwell I Nye, Armando Solar-Lezama, Joshua B Tenenbaum, and Brenden M Lake. Learning compositional rules via neural program synthesis. arXiv preprint arXiv:2003.05562, 2020.
  35. 35.Mayu Otani, Yuta Nakashima, Esa Rahtu, and Janne Heikkila. Uncovering hidden challenges in query-based video moment retrieval. arXiv preprint arXiv:2009.00325, 2020.
  36. 36.Jeffrey Pennington, Richard Socher, and Christopher D Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543, 2014.
  37. 37.Bharat Singh, Tim K Marks, Michael Jones, Oncel Tuzel, and Ming Shao. A multi-stream bi-directional recurrent neural network for fine-grained action detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1961–1970, 2016.
  38. 38.Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. Advances in neural information processing systems, 28:3483–3491, 2015.
  39. 39.Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen. Fine-grained action retrieval through multiple parts-of-speech embeddings. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 450–459, 2019.
  40. 40.Jie Wu, Guanbin Li, Si Liu, and Liang Lin. Tree-structured policy based progressive reinforcement learning for temporally language grounding in video. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 12386–12393, 2020.
  41. 41.Yitian Yuan, Xiaohan Lan, Long Chen, Wei Liu, Xin Wang, and Wenwu Zhu. A closer look at temporal sentence grounding in videos: Datasets and metrics. arXiv preprint arXiv:2101.09028, 2021.
  42. 42.Yitian Yuan, Lin Ma, Jingwen Wang, Wei Liu, and Wenwu Zhu. Semantic conditioned dynamic modulation for temporal sentence grounding in videos. arXiv preprint arXiv:1910.14303, 2019.
  43. 43.Yitian Yuan, Tao Mei, and Wenwu Zhu. To find where you talk: Temporal sentence localization in video with attention based location regression. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9159–9166, 2019.
  44. 44.Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, and Larry S Davis. Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1247–1257, 2019.
  45. 45.Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. Span-based localizing network for natural language video localization. arXiv preprint arXiv:2004.13931, 2020.
  46. 46.Shengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang, Zhou Zhao, Jianke Zhu, Jin Yu, Hongxia Yang, and Fei Wu. Devlbert: Learning deconfounded visio-linguistic representations. In MM 20: The 28th ACM International Conference on Multimedia, Virtual Event / Seattle, WA, USA, October 12-16, 2020, 2020.
  47. 47.Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. Learning 2d temporal adjacent networks for moment localization with natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 12870–12877, 2020.
  48. 48.Shengyu Zhang, Ziqi Tan, Zhou Zhao, Jin Yu, Kun Kuang, Tan Jiang, Jingren Zhou, Hongxia Yang, and Fei Wu. Comprehensive information integration modeling framework for video titling. In KDD 20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, CA, USA, August 23-27, 2020, 2020.
  49. 49.Shengyu Zhang, Dong Yao, Zhou Zhao, Tat-Seng Chua, and Fei Wu. Causerec: Counterfactual user sequence synthesis for sequential recommendation. In SIGIR 21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021, 2021.
  50. 50.Wenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang, Qingpeng Cai, Juncheng Li, Sihui Luo, and Yueting Zhuang. Magic: Multimodal relational graph adversarial inference for diverse and unpaired text-based image captioning. arXiv preprint arXiv:2112.06558, 2021.
  51. 51.Wenqiao Zhang, Haochen Shi, Siliang Tang, Jun Xiao, Qiang Yu, and Yueting Zhuang. Consensus graph representation learning for better grounded image captioning. In Proc 35 AAAI Conf on Artificial Intelligence, 2021.
  52. 52.Wenqiao Zhang, Xin Eric Wang, Siliang Tang, Haizhou Shi, Haochen Shi, Jun Xiao, Yueting Zhuang, and William Yang Wang. Relational graph learning for grounded video description generation. In Proceedings of the 28th ACM International Conference on Multimedia, pages 3807–3828, 2020.
  53. 53.Wenqiao Zhang, Lei Zhu, James Hallinan, Andrew Makmur, Shengyu Zhang, Qingpeng Cai, and Beng Chin Ooi. Boostmis: Boosting medical image semi-supervised learning with adaptive pseudo labeling and informative active annotation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  54. 54.Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Cross-task weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3537–3545, 2019.

Citation

MLA
Li, J., et al. “Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning”. arXiv, 2022, http://arxiv.org/abs/2203.13049v2.
APA
Li, J., Xie, J., Qian, L., Zhu, L., Tang, S., Wu, F., Yang, Y., Zhuang, Y., & Wang, X. E. (2022). Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning. arXiv. http://arxiv.org/abs/2203.13049v2
Chicago
Li, J., J. Xie, L. Qian, et al. 2022. “Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning”. arXiv. http://arxiv.org/abs/2203.13049v2.
Harvard
Li, J. et al. (2022) “Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.13049v2.
Vancouver
1. Li J, Xie J, Qian L, Zhu L, Tang S, Wu F, Yang Y, Zhuang Y, Wang XE (2022) Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning. arXiv

BibTeX

@article{li2022compositional,
  title = {Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning},
  author = {Li, Juncheng and Xie, Junlin and Qian, Long and Zhu, Linchao and Tang, Siliang and Wu, Fei and Yang, Yi and Zhuang, Yueting and Wang, Xin Eric},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.13049v2},
  eprint = {2203.13049}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE