TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Xiangyu ZengKunchang LiChenting WangXinhao LiTianxiang JiangZiang YanSongze LiYansong ShiZhengrong YueYi Wang

article2025ICLR111 citations

Presents TimeSuite, a grounded instruction-tuning framework that adapts multimodal language models for long-form video comprehension by pairing efficient token compression with explicit timestamp generation to curb hallucinations.

Listen

Multimodal artificial intelligence models have advanced significantly in understanding short videos, but they continue to struggle with long-form video content. Long video sequences contain complex temporal relationships, dynamic actions, and redundant visual information, which frequently causes models to lose track of key moments and produce hallucinated or inaccurate responses. Previous attempts to address this either created specialized architectures that damaged general question-answering capabilities or used broad compression methods that failed to pinpoint specific moments in time.

The article introduces TimeSuite, a comprehensive framework designed to adapt existing short-form video models for long video understanding by explicitly incorporating temporal grounding—the ability to identify exact start and end timestamps for specific events. The authors demonstrate that training a model to locate events temporally enables it to focus on relevant visual segments, thereby improving overall reasoning accuracy and reducing hallucinations.

The approach introduces three core innovations within a unified model named VideoChat-T. First, the architecture incorporates an efficient Token Shuffle compression scheme alongside a Temporal Adaptive Position Encoding module, which compresses visual tokens to reduce computational overhead while preserving temporal order. Second, the authors constructed TimePro, an instruction-tuning dataset comprising 349,000 high-quality annotations across nine temporal tasks. Third, they designed a new instruction-tuning task called Temporal Grounded Caption, which trains the model to simultaneously predict precise event timestamps and generate detailed segment descriptions.

The evaluation demonstrates several key findings:

  1. VideoChat-T achieved substantial gains on long-form video benchmarks, improving accuracy by 5.6% on the full Egoschema benchmark (reaching 60.0%) and by 6.8% on VideoMME (reaching 46.3% without subtitles), with an 8.7% gain on VideoMME's long-video subset.
  2. In zero-shot temporal grounding on the Charades-STA benchmark, the model scored 48.7% on top-1 recall at an intersection-over-union threshold of 0.5, outperforming the previous state-of-the-art model (TimeChat) by 16.5 percentage points; after targeted fine-tuning, it achieved 67.1%, matching or exceeding specialized expert models.
  3. Unlike prior specialized models that experienced severe capability drops, VideoChat-T preserved general short-video question-answering capabilities on MVBench, showing only a negligible 0.5% drop (59.9% vs. 60.4%) when trained with diverse datasets.
  4. The framework demonstrated extreme computational efficiency, utilizing only 3 tokens per frame and requiring 0.63 seconds of inference time per query on a single processor, consuming roughly 5.1% of the compute operations required by comparable leading models.

These findings indicate that temporal grounding serves as an effective mutual regularizer for video captioning and question answering. Rather than sacrificing general conversational capabilities to achieve temporal precision, models can attain both through diverse multi-task instruction data and parameter-efficient initialization. This architecture significantly lowers the compute cost and latency required for processing long video streams, making real-time long-form video analysis practical for operational deployment.

For future development, the article recommends incorporating chain-of-thought data structures to guide models through multi-step causal reasoning, as the current model still struggles to infer complex character intentions and deep logical context. Furthermore, research should focus on optimizing numerical output formats for dense tasks like highlight detection and exploring additional token compression techniques to eliminate redundancy in long video representations.

While the confidence in the empirical results is high across standardized benchmarks, the primary limitations involve tasks requiring dense numerical predictions—such as highlight detection with numerous discrete saliency scores—where language model text generation becomes a bottleneck. Readers should also note that while the model excels at temporal perception and direct question answering, human oversight remains necessary for tasks involving subtle, multi-step logical deductions.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for long video understanding, including a simple yet efficient framework to process long video sequence, a high-quality video dataset for grounded tuning of MLLMs, and a carefully-designed instruction tuning task to explicitly incorporate the grounding supervision in the traditional QA format. Specifically, based on VideoChat, we propose our long-video MLLM, coined as VideoChat-T, by implementing a token shuffling to compress long video tokens and introducing Temporal Adaptive Position Encoding (TAPE) to enhance the temporal awareness of visual representation. Meanwhile, we introduce the TimePro, a comprehensive grounding-centric instruction tuning dataset composed of 9 tasks and 349k high-quality grounded annotations. Notably, we design a new instruction tuning task type, called Temporal Grounded Caption, to peform detailed video descriptions with the corresponding time stamps prediction. This explicit temporal location prediction will guide MLLM to correctly attend on the visual content when generating description, and thus reduce the hallucination risk caused by the LLMs. Experimental results demonstrate that our TimeSuite provides a successful solution to enhance the long video understanding capability of short-form MLLM, achieving improvement of 5.6% and 6.8% on the benchmarks of Egoschema and VideoMME, respectively. In addition, VideoChat-T exhibits robust zero-shot temporal grounding capabilities, significantly outperforming the existing state-of-the-art MLLMs. After fine-tuning, it performs on par with the traditional supervised expert models.

Citation

MLA
Zeng, X., et al. “TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning”. arXiv, 2024, http://arxiv.org/abs/2410.19702v2.
APA
Zeng, X., Li, K., Wang, C., Li, X., Jiang, T., Yan, Z., Li, S., Shi, Y., Yue, Z., Wang, Y., Wang, Y., Qiao, Y., & Wang, L. (2024). TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning. arXiv. http://arxiv.org/abs/2410.19702v2
Chicago
Zeng, X., K. Li, C. Wang, et al. 2024. “TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning”. arXiv. http://arxiv.org/abs/2410.19702v2.
Harvard
Zeng, X. et al. (2024) “TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.19702v2.
Vancouver
1. Zeng X, Li K, Wang C, et al (2024) TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning. arXiv

BibTeX

@article{zeng2024timesuite,
  title = {TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning},
  author = {Zeng, Xiangyu and Li, Kunchang and Wang, Chenting and Li, Xinhao and Jiang, Tianxiang and Yan, Ziang and Li, Songze and Shi, Yansong and Yue, Zhengrong and Wang, Yi and Wang, Yali and Qiao, Yu and Wang, Limin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.19702v2},
  eprint = {2410.19702}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors