TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Yongxin GuoJingyu LiuMingda LiQingbin LiuXi ChenXiaoying Tang

article2025ICLR112 citations

Proposes TRACE, a task-interleaved video language model that advances video temporal grounding by formulating predictions as causal event sequences combining timestamps, saliency scores, and captions rather than relying solely on text generation.

Listen

Video temporal grounding is essential for downstream applications such as automated video editing, content summarization, and moment retrieval. While modern video-based large language models have shown promise in handling multiple tasks without task-specific retraining, they traditionally rely purely on unstructured text generation. This reliance creates a mismatch with the natural structure of video data, which depends fundamentally on explicit timestamps and saliency scores alongside descriptions. As a result, existing models struggle to pinpoint precise event timings and assess moment importance accurately.

To address this limitation, the article presents a causal event modeling framework and develops a task-interleaved model named TRACE. Instead of treating video interpretation as pure narrative text, the model formalizes outputs as structured event triplets composed of timestamps, saliency scores, and textual descriptions. The primary objective is to evaluate whether structuring model outputs to mirror video temporality enables a unified model to achieve state-of-the-art accuracy across diverse temporal grounding tasks without losing broad language and reasoning capabilities.

The framework implements dedicated encoders and decoding heads to handle visual frames, timestamps, scores, and text separately, cycling between them using special synchronization tokens. This approach prevents numerical time and score data from corrupting the core text model's learned knowledge base. Training is executed in two stages across roughly two million samples spanning instruction tuning, video captioning, and question-answering datasets. Model evaluation focuses on standard benchmarks including YouCook2 for dense video captioning, Charades-STA for moment retrieval, and QVHighlights for highlight detection.

Empirical evaluations demonstrate substantial performance gains over existing temporal grounding and traditional video language models. In zero-shot testing, TRACE improves captioning quality and timing alignment on YouCook2, raising CIDEr scores by 3.1 points and F1 scores by 4.9%. On moment retrieval via Charades-STA, recall increases by 6.5% at an intersection-over-union threshold of 0.5. For highlight detection on QVHighlights, the model achieves a 10.3% boost in mean average precision and a 9.2% increase in top-ranked hit rate. Moreover, ablation tests show that omitting the independent task encoders and decoding heads causes instruction-following to collapse, proving the necessity of decoupled task processing. When fine-tuned, TRACE matches or exceeds non-generative, single-purpose baseline models while retaining generalist versatility.

These findings suggest that aligning model architecture with intrinsic video structure substantially improves temporal reasoning efficiency and precision. Organizations developing video analytics, content moderation, or media asset management systems can deploy unified foundational models rather than maintaining fragmented, task-specific pipelines. This consolidation reduces operational complexity, lower deployment costs, and improves zero-shot retrieval in large video repositories without requiring costly specialized fine-tuning for each new scenario.

Future development should focus on integrating explicit causality discovery mechanisms and causal graphs into the input prompts to help the model capture complex inter-event dependencies beyond standard left-to-right temporal sequences. Additionally, training data should be expanded to include timestamped question-and-answer pairs to bolster contextual reasoning. Practitioners adopting this approach should note that performance depends heavily on the accuracy of training annotations; incorporating broad datasets with noisy boundaries can slightly degrade precision on short-video tasks. Nonetheless, the architecture provides strong, highly reliable improvements across standard video understanding benchmarks.

Cover for TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Abstract

Video Temporal Grounding (VTG) is a crucial capability for video understanding models and plays a vital role in downstream tasks such as video browsing and editing. To effectively handle various tasks simultaneously and enable zero-shot prediction, there is a growing trend in employing video LLMs for VTG tasks. However, current video LLM-based methods rely exclusively on natural language generation, lacking the ability to model the clear structure inherent in videos, which restricts their effectiveness in tackling VTG tasks. To address this issue, this paper first formally introduces causal event modeling framework, which represents video LLM outputs as sequences of events, and predict the current event using previous events, video inputs, and textural instructions. Each event consists of three components: timestamps, salient scores, and textual captions. We then propose a novel task-interleaved video LLM called TRACE to effectively implement the causal event modeling framework in practice. The TRACE process visual frames, timestamps, salient scores, and text as distinct tasks, employing various encoders and decoding heads for each. Task tokens are arranged in an interleaved sequence according to the causal event modeling framework's formulation. Extensive experiments on various VTG tasks and datasets demonstrate the superior performance of TRACE compared to state-of-the-art video LLMs. Our model and code are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 TRACE
  • 3.1 Modeling the Inherent Structures of Videos
  • 3.2 TRACE: Task-Interleaved Temporal Grounding Video LLM
  • 3.2.1 Separated Multi-Task Processing
  • 3.2.2 Task-interleaved sequence modeling
  • 3.2.3 Adaptive Head-Switching Mechanism for Generation
  • 3.3 Training Strategy and Data Preparation
  • 4 Experiments
  • 4.1 Evaluation Datasets, Metrics, and Baseline Models.
  • 4.2 Performance of TRACE
  • 4.3 Ablation Studies of TRACE.
  • 4.4 Fine-tuned Performance of TRACE.
  • 5 Conclusion and Future Works
  • References
  • A Dataset Preparation
  • A.1 Details of data format
  • A.2 Processing InternVid
  • A.3 Processing VTG-IT
  • B Experiments
  • B.1 Detailed Experimental Settings
  • B.2 Additional Experiment Results
  • B.3 Case Studies
  • C Discussion on Causal Event Modeling
  • C.1 Causal Language Modeling VS. Causal Event Modeling.
  • C.2 Causal Event Modeling VS. Complete Causal Relationship Modeling/Discovey

Knowls

  1. Knowl 1 — Causal Event Modeling Framework for Video Temporal Grounding

    definition

    Causal Event Modeling is a framework for video large language models (video LLMs) that formulates model outputs as an ordered sequence of discrete event triplets rather than unstructured natural language text. Given textual instruction II and video frame inputs FF, the output sequence RR consists of KK events:

    R={e1,e2,…,eK}={(tk,sk,ck)∣1≤k≤K}R = \{e_1, e_2, \ldots, e_K\} = \{(t_k, s_k, c_k) \mid 1 \le k \le K\}

    where each event ek=(tk,sk,ck)e_k = (t_k, s_k, c_k) comprises:

    • tkt_k: the start and end timestamps of the event.
    • sks_k: the saliency or highlight score of the event.
    • ckc_k: the natural language caption or description of the event.

    The conditional joint probability distribution factorizes autoregressively across events and within each event component:

    P(R∣F,I)=∏k=1KP(ek∣e1:k−1,I,F)P(R \mid F, I) = \prod_{k=1}^K P(e_k \mid e_{1:k-1}, I, F)

    P(ek∣e1:k−1,I,F)=P(tk∣e1:k−1,I,F)⋅P(sk∣tk,e1:k−1,I,F)⋅P(ck∣sk,tk,e1:k−1,I,F)P(e_k \mid e_{1:k-1}, I, F) = P(t_k \mid e_{1:k-1}, I, F) \cdot P(s_k \mid t_k, e_{1:k-1}, I, F) \cdot P(c_k \mid s_k, t_k, e_{1:k-1}, I, F)

    Events are ordered chronologically by their occurrence timestamps (t1≤t2≤⋯≤tKt_1 \le t_2 \le \cdots \le t_K). Within each event triplet eke_k, timestamps tkt_k and saliency score sks_k are predicted prior to the text caption ckc_k, allowing temporal and saliency predictions to serve as conditional context for generating the descriptive caption.

  2. Knowl 2 — TRACE Task-Interleaved Video LLM Architecture

    model/method

    TRACE (TempoRAl grounding via Causal Event modeling) implements the causal event modeling framework by assigning distinct encoders and decoding heads to visual frames, timestamps, saliency scores, and text, centered around a shared pre-trained large language model backbone (Mistral-7B-v0.2):

    1. Visual Frame Encoding: For a video with TT frames, each frame is encoded into 576 visual tokens via a pre-trained CLIP ViT-L vision encoder. Slot-Based Compression reduces these to 8 visual tokens per frame. A dedicated time encoder converts the sampled frame's timestamp into 6 time tokens. Concatenating these yields 14 tokens per frame, sequenced chronologically across all TT frames to produce visual input FF.

    2. Text Processing: Textual instructions II and event descriptions ckc_k are tokenized and decoded using the standard tokenizer and language modeling head of Mistral-7B-v0.2. A delimiter token ⟨sync⟩ is introduced to mark the conclusion of text segments.

    3. Independent Timestamp and Saliency Score Modules: Separate encoder and decoding head pairs process timestamps tkt_k and scores sks_k. Each uses a 13-token vocabulary initialized from LLM token embeddings: digits ⟨0⟩ through ⟨9⟩, a decimal point ⟨.⟩, a delimiter ⟨sep⟩, and a task-switching token ⟨sync⟩.

      • Timestamps are formatted into fixed-width sequences with 4 integer digits, 1 decimal point, and 1 fractional digit per timestamp (e.g., [10.23, 125.37] tokenizes to ⟨0⟩⟨0⟩⟨1⟩⟨0⟩⟨.⟩⟨2⟩⟨sep⟩⟨0⟩⟨1⟩⟨2⟩⟨5⟩⟨.⟩⟨4⟩⟨sync⟩).
      • Saliency scores are formatted into 3 tokens (1 integer digit, 1 decimal point, 1 fractional digit) followed by ⟨sync⟩.
    4. Task-Interleaved Sequence Ordering: Input and output tokens are interleaved in the sequence: visual tokens FF, instruction tokens II, followed by event tokens t1,s1,c1,…,tK,sK,cKt_1, s_1, c_1, \ldots, t_K, s_K, c_K.

  3. Knowl 3 — Adaptive Head-Switching Generation Algorithm for TRACE

    algorithm

    TRACE generates structured event sequences autoregressively by switching decoding heads upon generating the special ⟨sync⟩ token.

    Input: Visual frame tokens FF, instruction tokens II, maximum generation steps LmaxL_{max}
    Output: Structured event sequence R={e1,e2,…,eK}R = \{e_1, e_2, \ldots, e_K\} where ek=(tk,sk,ck)e_k = (t_k, s_k, c_k)
    Initialize sequence S←[F,I]S \leftarrow [F, I]
    Initialize active_head ←\leftarrow TIME_HEAD
    Initialize current_event ←(t="",s="",c="")\leftarrow (t=\text{""}, s=\text{""}, c=\text{""})
    Initialize R←[]R \leftarrow []
    while length of S<LmaxS < L_{max} do
        if active_head == TIME_HEAD then
            token ←\leftarrow sample from TIME_HEAD(S)(S)
            Append token to SS
            if token == "⟨sync⟩" then
                active_head ←\leftarrow SCORE_HEAD
            else
                current_event.t←t \leftarrow current_event.t+t + token
            end if
        else if active_head == SCORE_HEAD then
            token ←\leftarrow sample from SCORE_HEAD(S)(S)
            Append token to SS
            if token == "⟨sync⟩" then
                active_head ←\leftarrow TEXT_HEAD
            else
                current_event.s←s \leftarrow current_event.s+s + token
            end if
        else if active_head == TEXT_HEAD then
            token ←\leftarrow sample from TEXT_HEAD(S)(S)
            Append token to SS
            if token == "⟨sync⟩" then
                Append current_event to RR
                current_event ←(t="",s="",c="")\leftarrow (t=\text{""}, s=\text{""}, c=\text{""})
                active_head ←\leftarrow TIME_HEAD
            else if token == "⟨eos⟩" then
                if current_event.c≠""c \ne \text{""} or current_event.t≠""t \ne \text{""} then
                    Append current_event to RR
                end if
                break
            else
                current_event.c←c \leftarrow current_event.c+c + token
            end if
        end if
    end while
    return RR

    The decoding process cyclically switches among the specialized heads in the order TIME_HEAD →\to SCORE_HEAD →\to TEXT_HEAD →\to TIME_HEAD until sequence completion.

  4. Knowl 4 — Two-Stage Training Pipeline for TRACE

    experimental setup

    TRACE is trained in two successive stages on 16 ATN 910B hardware accelerators using DeepSpeed ZeRO-3 Offload and a maximum sequence length of 4096:

    1. Stage 1: Initialization of Task Modules:

      • Trainable Modules: Vision compression layer, time encoder/head, score encoder/head, text tokenizer/head.
      • Frozen Modules: CLIP ViT-L vision encoder and Mistral-7B-v0.2 LLM backbone.
      • Training Data (1.9M samples): Video/image caption datasets (Valley, LLaVA Image, TextVR, ShareGPT4Video subset) and video temporal grounding datasets (VTG-IT).
      • Hyperparameters: 128 uniformly sampled video frames, batch size 128, learning rate 1×10−31\times 10^{-3} with a cosine decay schedule, trained for 1 epoch (~5 days).
    2. Stage 2: Instruction Tuning for Temporal Grounding Capacity:

      • Trainable Modules: Mistral-7B-v0.2 LLM backbone and all task modules.
      • Frozen Modules: CLIP ViT-L vision encoder.
      • Training Data (0.9M samples): VTG instruction tuning datasets (635K samples from VTG-IT, ActivityNet Captions, filtered InternVid subset), video caption datasets (284K samples compressed to 1 instruction per video from Valley, TextVR, ShareGPT4Video), and video question answering datasets (VideoChatGPT, Next-QA).
      • Hyperparameters: Video divided into 128 clips with 1 frame randomly sampled per clip, batch size 128, learning rate 5×10−65\times 10^{-6} with cosine schedule, trained for 2 epochs (~5 days).
  5. Knowl 5 — Zero-Shot Performance on Video Temporal Grounding Benchmarks

    empirical result

    In zero-shot evaluations across dense video captioning (Youcook2), moment retrieval (Charades-STA), and video highlight detection (QVHighlights), TRACE (7B) demonstrates state-of-the-art performance against traditional and temporal-grounding video LLMs.

    Model Youcook2 Charades-STA QVHighlights
    SODA_c CIDEr F1 R@1 (0.5) R@1 (0.7) mAP HIT@1
    Traditional Video LLMs
    Valley (7B) 0.1 0.0 1.5 4.7 1.6 10.9 15.2
    VideoChat (7B) 0.2 0.6 3.4 3.2 1.4 13.1 18.1
    Video-LLaMA (7B) 0.0 0.0 0.1 2.7 1.2 11.3 15.6
    Temporal Grounding Video LLMs
    TimeChat (7B) 1.2 3.4 12.6 32.2 13.4 14.5 23.9
    VTimeLLM (7B) - - - 27.5 11.4 - -
    VTimeLLM (13B) - - - 34.3 14.7 - -
    Momentor (7B) - - - 26.6 11.6 7.6 -
    HawkEye (7B) - - - 31.4 14.5 - -
    VTG-LLM (7B) 1.5 5.0 17.5 33.8 15.7 16.5 33.5
    TRACE (7B) 2.2 8.1 22.4 40.3 19.4 26.8 42.7

    Compared to VTG-LLM (7B), TRACE achieves improvements of +3.1 in CIDEr and +4.9 in F1 score on Youcook2; +6.5% Recall@1 (IoU=0.5) and +3.7% Recall@1 (IoU=0.7) on Charades-STA; and +10.3% mAP and +9.2% HIT@1 on QVHighlights. Additionally, the 7B TRACE model outperforms the 13B VTimeLLM model on Charades-STA.

  6. Knowl 6 — Fine-Tuned Performance on Video Temporal Grounding Benchmarks

    empirical result

    When fine-tuned for 3 epochs on task-specific benchmarks, TRACE achieves competitive performance with specialized non-generative models and surpasses generalist video LLM baselines.

    On Youcook2 (Dense Video Captioning):

    Model SODA_c CIDEr F1 Score
    Task-Specific Models
    PDVC 4.4 22.7 -
    Vid2Seq (no audio) 5.7 25.3 23.5
    CM2^2 5.3 31.7 28.4
    Generalist Video LLMs
    TimeChat 3.4 11.0 19.5
    VTG-LLM 3.6 13.4 20.6
    TRACE 6.7 35.5 31.8

    On Charades-STA (Moment Retrieval):

    Model R@1 (IoU=0.5) R@1 (IoU=0.7)
    Non-Generative Models
    InternVideo2-6B 70.0 49.0
    VDI 52.3 31.4
    Moment-DETR 55.7 34.2
    Generative / LLM Models
    HawkEye 58.3 28.8
    TimeChat 46.7 23.7
    VTG-LLM 57.2 33.4
    TRACE 61.7 41.4

    On QVHighlights (Video Highlight Detection), fine-tuned TRACE reaches 31.8 mAP and 51.5 HIT@1, outperforming TimeChat (21.7 mAP, 37.9 HIT@1) and VTG-LLM (24.1 mAP, 41.3 HIT@1).

  7. Knowl 7 — Ablation Analysis on TRACE Architecture and Frame Sampling

    empirical result

    Ablations on Youcook2 and Charades-STA (fine-tuned solely on VTG-IT) assess the contribution of the causal event framework, separate task modules, frame counts, and slot counts per frame.

    Variant Frames Youcook2 Charades-STA
    SODA_c CIDEr F1 R@1 (0.5) R@1 (0.7)
    Architecture Ablations
    w/o causal event modeling 96 1.4 4.3 17.2 29.7 14.0
    w/o independent encoder/heads 64 — Failed to Follow Instruction —
    TRACE (VTG-IT) 64 1.9 6.9 21.4 37.0 17.0
    Frame Number Scaling
    TRACE (VTG-IT) 8 1.4 5.0 18.6 28.8 13.6
    TRACE (VTG-IT) 64 1.9 6.9 21.4 37.0 17.0
    TRACE (VTG-IT) 128 2.1 7.5 21.4 41.2 20.0
    Slots per Frame (at 64 frames)
    8 slots / frame 64 1.9 6.9 21.4 37.0 17.0
    16 slots / frame 64 2.1 7.3 22.1 41.9 20.1

    Key takeaways:

    1. Formatting outputs as natural language (w/o causal event modeling) reduces Youcook2 CIDEr from 6.9 to 4.3 and Charades-STA Recall@1 (0.5) from 37.0 to 29.7, despite sampling more frames (96 vs 64).
    2. Adding time and score tokens directly into the LLM text tokenizer (w/o independent encoder/heads) disrupts pre-trained language knowledge, resulting in complete failure to follow instructions.
    3. Model performance scales monotonically with frame count (Recall@1 (0.5) rises from 28.8% at 8 frames to 41.2% at 128 frames) and with visual slots per frame (Recall@1 (0.5) rises from 37.0% at 8 slots to 41.9% at 16 slots).
  8. Knowl 8 — Saliency Score Annotation Synthesis via CLIP Frame-Caption Similarity

    model/method

    To generate non-uniform supervisory saliency scores for video highlight detection and summarization from dense video caption data (VTG-IT), an automated annotation synthesis procedure was used:

    1. Clip Splitting: Each annotated event is split into at most 20 sub-clips.
    2. Similarity Scoring: Visual frames in each sub-clip and the corresponding event caption are embedded using EVA-CLIP (ViT-G/14) to compute frame-text cosine similarities.
    3. Gaussian Quantization: Clip similarity scores are normalized to a Gaussian distribution and mapped to discrete saliency scores {1.0,1.2,1.4,…,5.0}\{1.0, 1.2, 1.4, \ldots, 5.0\} using 21 normal cumulative distribution percentile thresholds: [2.275%,3.593%,5.480%,8.076%,11.507%,15.866%,21.186%,27.425%,34.458%,42.074%,50.000%,57.926%,65.542%,72.575%,78.814%,84.134%,88.493%,91.924%,94.520%,96.407%,97.725%][2.275\%, 3.593\%, 5.480\%, 8.076\%, 11.507\%, 15.866\%, 21.186\%, 27.425\%, 34.458\%, 42.074\%, 50.000\%, 57.926\%, 65.542\%, 72.575\%, 78.814\%, 84.134\%, 88.493\%, 91.924\%, 94.520\%, 96.407\%, 97.725\%].
    4. Summarization Clip Selection: For each event, the sub-clip assigned the maximum saliency score is selected as the key summarization clip for the video summarization task.
  9. Knowl 9 — Performance on Event-Level Understanding and Causal Reasoning Benchmarks

    empirical result

    TRACE preserves fine-grained event understanding and temporal-causal reasoning capabilities, as evaluated on Event-Bench and E.T.Bench.

    On Event-Bench (evaluating atomic and composite reasoning across video events):

    Model Event Description Temporal Reasoning Causal Reasoning Episodic Reasoning Overall Avg.
    Video-LLaVA (7B) 12.82 5.50 0.00 7.20 5.87
    VideoChat2 (7B) 33.76 37.75 47.75 14.67 29.41
    ST-LLM (7B) 47.22 48.75 59.50 16.67 37.71
    VIM (7B) 48.08 51.25 61.25 18.67 41.64
    Gemini-1.5-Pro 48.50 47.50 41.75 38.67 43.24
    GPT-4o 54.27 56.75 58.25 37.33 53.33
    TRACE (7B) 55.56 49.25 54.50 43.00 50.46

    On E.T.Bench, TRACE (7B) achieves 46.8 F1 on Temporal Video Grounding (TVG), 45.2 F1 on Video Highlight Detection (VHD), 45.7 F1 on Dense Video Captioning (DVC), 27.3 F1 on Single-choice Action Localization (SLC), and 17.8 Recall on Temporal Event Moment retrieval (TEM), outperforming both VideoLLaMA2 (7B) and Qwen2-VL (7B).

  10. Knowl 10 — Limitations of Forward-Only Autoregressive Causal Event Modeling

    limitation

    Because TRACE is built on a decoder-only autoregressive LLM, its event prediction is strictly unidirectional: each event eke_k is generated conditioned solely on preceding events e1:k−1e_{1:k-1} and visual inputs FF. This imposes specific structural limitations:

    1. Inability to Capture Non-Sequential Event Graphs: Unidirectional autoregression cannot model complex non-chronological causal graphs, retroactive dependencies, or bidirectional event interactions in complex narratives.
    2. Absence of Explicit Symbolic Causality Structures: TRACE relies on implicit representation learning rather than explicit causality graphs CC. Incorporating external causality discovery models into the pipeline (via conditioning P(R∣F,I,C)P(R \mid F, I, C), explicit chain-of-thought graph generation P(C∣F,I)P(R∣C,F,I)P(C \mid F, I)P(R \mid C, F, I), or causality-guided visual attention masking) remains an unaddressed limitation.

Coverage note — No substantial contributed material from the paper was omitted; all primary architectural components, mathematical formulations, training stages, data preparation pipelines, zero-shot and fine-tuned results, ablation studies, and architectural limitations are included.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72, 2005.
  3. 3.Meinardus Boris, Batra Anil, Rohrbach Anna, and Rohrbach Marcus. The surprising effectiveness of multimodal large language models for video moment retrieval. arXiv preprint arXiv:2406.18113, 2024.
  4. 4.Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Bin Lin, Zhenyu Tang, et al. Sharegpt4video: Improving video understanding and generation with better captions. arXiv preprint arXiv:2406.04325, 2024a.
  5. 5.Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, and Jing Liu. Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset. Advances in Neural Information Processing Systems, 36, 2024b.
  6. 6.Tieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen, Chaofan Gan, Xiao Ma, Cheng Zhong, Yang Zhang, Yingxue Wang, Hui Lin, et al. Mecd: Unlocking multi-event causal discovery in video reasoning. arXiv preprint arXiv:2409.17647, 2024c.
  7. 7.Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476, 2024.
  8. 8.Yifan Du, Kun Zhou, Yuqi Huo, Yifan Li, Wayne Xin Zhao, Haoyu Lu, Zijia Zhao, Bingning Wang, Weipeng Chen, and Ji-Rong Wen. Towards event-oriented long video understanding. arXiv preprint arXiv:2406.14129, 2024.
  9. 9.Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 961–970, 2015.
  10. 10.Soichiro Fujita, Tsutomu Hirao, Hidetaka Kamigaito, Manabu Okumura, and Masaaki Nagata. Soda: Story oriented dense video captioning evaluation framework. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16, pp. 517–531. Springer, 2020.
  11. 11.Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. Tall: Temporal activity localization via language query. In Proceedings of the IEEE international conference on computer vision, pp. 5267–5275, 2017.
  12. 12.Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. Text-to-audio generation using instruction-tuned llm and latent diffusion model. arXiv preprint arXiv:2304.13731, 2023.
  13. 13.Rohit Girdhar and Deva Ramanan. Cater: A diagnostic dataset for compositional actions and temporal reasoning. arXiv preprint arXiv:1910.04744, 2019.
  14. 14.Yongxin Guo, Jingyu Liu, Mingda Li, Xiaoying Tang, Xi Chen, and Bo Zhao. Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding. arXiv preprint arXiv:2405.13382, 2024.
  15. 15.Michael Gygli, Helmut Grabner, Hayko Riemenschneider, and Luc Van Gool. Creating summaries from user videos. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, pp. 505–520. Springer, 2014.
  16. 16.Donghoon Han, Seunghyeon Seo, Eunhwan Park, Seong-Uk Nam, and Nojun Kwak. Unleash the potential of clip for video highlight detection. arXiv preprint arXiv:2404.01745, 2024.
  17. 17.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with temporal language. In Empirical Methods in Natural Language Processing (EMNLP), 2018a.
  18. 18.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with temporal language. In Empirical Methods in Natural Language Processing (EMNLP), 2018b.
  19. 19.Hang Hua, Yunlong Tang, Chenliang Xu, and Jiebo Luo. V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning. arXiv preprint arXiv:2404.12353, 2024.
  20. 20.Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. Vtimellm: Empower llm to grasp video moments. arXiv preprint arXiv:2311.18445, 2(3):9, 2023.
  21. 21.De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, and Jan Kautz. Lita: Language instructed temporal-localization assistant. arXiv preprint arXiv:2403.19046, 2024.
  22. 22.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  23. 23.Yang Jin, Linchao Zhu, and Yadong Mu. Complex video action reasoning via learnable markov logic network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3242–3251, 2022.
  24. 24.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  25. 25.Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi, and Seong Tae Kim. Do you remember? dense video captioning with cross-modal memory retrieval. arXiv preprint arXiv:2404.07610, 2024.
  26. 26.J Lei, TL Berg, and M Bansal. Qvhighlights: Detecting moments and highlights in videos via natural language queries.(2021). URL https://arxiv. org/abs/2107.09609.
  27. 27.Jie Lei, Tamara L Berg, and Mohit Bansal. Detecting moments and highlights in videos via natural language queries. Advances in Neural Information Processing Systems, 34:11846–11858, 2021.
  28. 28.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024.
  29. 29.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp. 19730–19742. PMLR, 2023a.
  30. 30.Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023b.
  31. 31.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao. Mvbench: A comprehensive multi-modal video understanding benchmark, 2023c.
  32. 32.Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao. Unmasked teacher: Towards training-efficient video foundation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 19948–19960, 2023d.
  33. 33.Yunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox, and Animesh Garg. Causal discovery in physical systems from videos. Advances in Neural Information Processing Systems, 33: 9180–9192, 2020.
  34. 34.Chen Liang, Wenguan Wang, Tianfei Zhou, and Yi Yang. Visual abductive reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 15565–15575, 2022.
  35. 35.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023a.
  36. 36.Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou. Univtg: Towards unified video-language temporal grounding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2794–2804, 2023b.
  37. 37.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36, 2024.
  38. 38.Ye Liu, Zongyang Ma, Zhongang Qi, Yang Wu, Ying Shan, and Chang Wen Chen. Et bench: Towards open-ended event-level video-language understanding. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  39. 39.Ye Liu, Siyuan Li, Yang Wu, Chang-Wen Chen, Ying Shan, and Xiaohu Qie. Umt: Unified multimodal transformers for joint video moment retrieval and highlight detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3042–3051, 2022.
  40. 40.Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, and Yang Liu. Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23045–23055, 2023a.
  41. 41.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Minghui Qiu, Pengcheng Lu, Tao Wang, and Zhongyu Wei. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207, 2023b.
  42. 42.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023.
  43. 43.Andreea-Maria Oncescu, Joao F Henriques, Yang Liu, Andrew Zisserman, and Samuel Albanie. Queryd: A video dataset with high-quality text and audio narrations. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2265–2269. IEEE, 2021.
  44. 44.Paritosh Parmar, Eric Peh, Ruirui Chen, Ting En Lam, Yuhan Chen, Elston Tan, and Basura Fernando. Causalchaos! dataset for comprehensive causal action question answering over longer causal chains grounded in dynamic visual scenes. arXiv preprint arXiv:2404.01299, 2024.
  45. 45.Long Qian, Juncheng Li, Yu Wu, Yaobo Ye, Hao Fei, Tat-Seng Chua, Yueting Zhuang, and Siliang Tang. Momentor: Advancing video large language model with fine-grained temporal reasoning, 2024.
  46. 46.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  47. 47.Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. Timechat: A time-sensitive multimodal large language model for long video understanding. arXiv preprint arXiv:2312.02051, 2023.
  48. 48.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18221–18232, 2024a.
  49. 49.Enxin Song, Wenhao Chai, Tian Ye, Jenq-Neng Hwang, Xi Li, and Gaoang Wang. Moviechat+: Question-aware sparse memory for long video question answering. arXiv preprint arXiv:2404.17176, 2024b.
  50. 50.Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. Tvsum: Summarizing web videos using titles. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5179–5187, 2015.
  51. 51.Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023.
  52. 52.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1207–1216, 2019.
  53. 53.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems, 35:10078–10093, 2022.
  54. 54.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  55. 55.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566–4575, 2015.
  56. 56.Teng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng, Ran Cheng, and Ping Luo. End-to-end dense video captioning with parallel decoding. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6847–6857, 2021.
  57. 57.Teng Wang, Jinrui Zhang, Feng Zheng, Wenhao Jiang, Ran Cheng, and Ping Luo. Learning grounded vision-language representation for versatile understanding in untrimmed videos. arXiv preprint arXiv:2303.06378, 2023a.
  58. 58.Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. Internvideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191, 2022.
  59. 59.Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023b.
  60. 60.Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Jilan Xu, Zun Wang, et al. Internvideo2: Scaling video foundation models for multimodal video understanding. arXiv preprint arXiv:2403.15377, 2024a.
  61. 61.Yueqian Wang, Xiaojun Meng, Jianxin Liang, Yuxuan Wang, Qun Liu, and Dongyan Zhao. Hawkeye: Training video-text llms for grounding text in videos. arXiv preprint arXiv:2403.10228, 2024b.
  62. 62.Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang, Dongyan Zhao, Yang Liu, and Zilong Zheng. Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge. arXiv preprint arXiv:2402.16050, 2024c.
  63. 63.Weijia Wu, Yuzhong Zhao, Zhuang Li, Jiahong Li, Hong Zhou, Mike Zheng Shou, and Xiang Bai. A large cross-modal video retrieval dataset with reading comprehension. Pattern Recognition, 157:110818, 2025.
  64. 64.Yongliang Wu, Xinting Hu, Yuyang Sun, Yizhou Zhou, Wenbo Zhu, Fengyun Rao, Bernt Schiele, and Xu Yang. Number it: Temporal grounding videos like flipping manga. arXiv preprint arXiv:2411.10332, 2024.
  65. 65.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9777–9786, 2021.
  66. 66.Yicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma, Hengwei Bian, Yatai Ji, Yujiu Yang, and Xiu Li. Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection. arXiv preprint arXiv:2311.16464, 2023.
  67. 67.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084, 2021.
  68. 68.Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Mi Zhang, Soham Ghosh, Yonghui Wu, and Jiahui Yu. Videococa: Video-text modeling with zero-shot transfer from contrastive captioners. arXiv preprint arXiv:2212.04979, 2022.
  69. 69.Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10714–10726, 2023.
  70. 70.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024.
  71. 71.Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum. Clevrer: Collision events for video representation and reasoning. arXiv preprint arXiv:1910.01442, 2019.
  72. 72.Abhay Zala, Jaemin Cho, Satwik Kottur, Xilun Chen, Barlas Oguz, Yashar Mehdad, and Mohit Bansal. Hierarchical video-moment retrieval and step-captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23056–23065, 2023a.
  73. 73.Abhay Zala, Jaemin Cho, Satwik Kottur, Xilun Chen, Barlas Oguz, Yashar Mehdad, and Mohit Bansal. Hierarchical video-moment retrieval and step-captioning. In CVPR, 2023b.
  74. 74.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. In Advances in Neural Information Processing Systems 34, 2021.
  75. 75.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023.
  76. 76.Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024.
  77. 77.Long Zhao, Nitesh B Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J Sun, Luke Friedman, Rui Qian, Tobias Weyand, Yue Zhao, et al. Videoprism: A foundational visual encoder for video understanding. arXiv preprint arXiv:2402.13217, 2024.
  78. 78.Luowei Zhou, Chenliang Xu, and Jason Corso. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  79. 79.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.

Citation

MLA
Guo, Y., et al. “TRACE: Temporal Grounding Video LLM via Causal Event Modeling”. arXiv, 2024, http://arxiv.org/abs/2410.05643v3.
APA
Guo, Y., Liu, J., Li, M., Liu, Q., Chen, X., & Tang, X. (2024). TRACE: Temporal Grounding Video LLM via Causal Event Modeling. arXiv. http://arxiv.org/abs/2410.05643v3
Chicago
Guo, Y., J. Liu, M. Li, Q. Liu, X. Chen, and X. Tang. 2024. “TRACE: Temporal Grounding Video LLM via Causal Event Modeling”. arXiv. http://arxiv.org/abs/2410.05643v3.
Harvard
Guo, Y. et al. (2024) “TRACE: Temporal Grounding Video LLM via Causal Event Modeling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.05643v3.
Vancouver
1. Guo Y, Liu J, Li M, Liu Q, Chen X, Tang X (2024) TRACE: Temporal Grounding Video LLM via Causal Event Modeling. arXiv

BibTeX

@article{guo2024trace,
  title = {TRACE: Temporal Grounding Video LLM via Causal Event Modeling},
  author = {Guo, Yongxin and Liu, Jingyu and Li, Mingda and Liu, Qingbin and Chen, Xi and Tang, Xiaoying},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.05643v3},
  eprint = {2410.05643}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors