Built independently by an author, for readers. Read the story and support ChapterPal

keyword

temporal grounded caption task

A temporal grounded caption task is a multimodal machine learning task in which a system generates descriptive natural language text for events occurring in a video while simultaneously predicting the precise timestamps or time intervals corresponding to those described events. Unlike standard video captioning, which produces global narrative summaries without explicit time boundaries, this task directly pairs fine-grained textual statements with their exact temporal locations within the video timeline. By integrating temporal localization into the caption generation process, the task guides multimodal models to ground their textual descriptions directly in specific visual segments, thereby enhancing fine-grained temporal reasoning and mitigating the risk of descriptive hallucinations.

1 item

TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Xiangyu Zeng, Kunchang Li, Chenting Wang, Xinhao Li, Tianxiang Jiang, Ziang Yan, Songze Li, Yansong Shi, Zhengrong Yue, Yi Wang, Yali Wang, Yu Qiao, Limin Wang

OrganizationsFudan UniversityNanjing UniversityShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong UniversityShenzhen Institute of Advanced Technology, Chinese Academy of SciencesUniversity of Science and Technology of ChinaZhejiang University

Why you should read this

Presents TimeSuite, a grounded instruction-tuning framework that adapts multimodal language models for long-form video comprehension by pairing efficient token compression with explicit timestamp generation to curb hallucinations.

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for long video understanding, including a simple yet efficient framework to process long video sequence, a high-quality video dataset for grounded tuning of MLLMs, and a carefully-designed instruction tuning task to explicitly incorporate the grounding supervision in the traditional QA format. Specifically, based on VideoChat, we propose our long-video MLLM, coined as VideoChat-T, by implementing a token shuffling to compress long video tokens and introducing Temporal Adaptive Position Encoding (TAPE) to enhance the temporal awareness of visual representation. Meanwhile, we introduce the TimePro, a comprehensive grounding-centric instruction tuning dataset composed of 9 tasks and 349k high-quality grounded annotations. Notably, we design a new instruction tuning task type, called Temporal Grounded Caption, to peform detailed video descriptions with the corresponding time stamps prediction. This explicit temporal location prediction will guide MLLM to correctly attend on the visual content when generating description, and thus reduce the hallucination risk caused by the LLMs. Experimental results demonstrate that our TimeSuite provides a successful solution to enhance the long video understanding capability of short-form MLLM, achieving improvement of 5.6% and 6.8% on the benchmarks of Egoschema and VideoMME, respectively. In addition, VideoChat-T exhibits robust zero-shot temporal grounding capabilities, significantly outperforming the existing state-of-the-art MLLMs. After fine-tuning, it performs on par with the traditional supervised expert models.

Added

2026-09-26