Built independently by an author, for readers. Read the story and support ChapterPal

keyword

dense video captioning

Dense video captioning is an artificial intelligence and computer vision task that involves automatically identifying the temporal boundaries of multiple events within an untrimmed video and generating a natural language description for each localized event. Unlike standard video captioning, which produces a single broad summary for an entire video clip, dense video captioning requires a model to perform both temporal event localization, determining the start and end timestamps of distinct actions, and textual narration for every individual segment. By connecting temporal grounding with detailed language generation, dense video captioning enables a fine-grained, chronological understanding of complex, long-form video content for applications such as video search, automated surveillance, and multimodal dialogue systems.

4 items

StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

Xin Ding, Hao Wu, Yifan Yang, Shiqi Jiang, Qianxi Zhang, Donglin Bai, Zhibo Chen, Ting Cao

OrganizationsMicrosoftNanjing UniversityTsinghua UniversityUniversity of Science and Technology of China

Why you should read this

Proposes StreamMind, an event-gated video dialogue framework that pairs state-space feature extraction with selective language model invocation to enable proactive, real-time video understanding at 100 frames per second.

With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce StreamMind, a video LLM framework that achieves ultra-FPS streaming video processing (100 fps on a single A100) and enables proactive, always-on responses in real time, without explicit user intervention. To solve the key challenge of the contradiction between linear video streaming speed and quadratic transformer computation cost, we propose a novel perception-cognition interleaving paradigm named ''event-gated LLM invocation'', in contrast to the existing per-time-step LLM invocation. By introducing a Cognition Gate network between the video encoder and the LLM, LLM is only invoked when relevant events occur. To realize the event feature extraction with constant cost, we propose Event-Preserving Feature Extractor (EPFE) based on state-space method, generating a single perception token for spatiotemporal features. These techniques enable the video LLM with full-FPS perception and real-time cognition response. Experiments on Ego4D and SoccerNet streaming tasks, as well as standard offline benchmarks, demonstrate state-of-the-art performance in both model capability and real-time efficiency, paving the way for ultra-high-FPS applications, such as Game AI and interactive media. The code and data is available at this https URL.

Added

2026-09-29

TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Xiangyu Zeng, Kunchang Li, Chenting Wang, Xinhao Li, Tianxiang Jiang, Ziang Yan, Songze Li, Yansong Shi, Zhengrong Yue, Yi Wang, Yali Wang, Yu Qiao, Limin Wang

OrganizationsFudan UniversityNanjing UniversityShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong UniversityShenzhen Institute of Advanced Technology, Chinese Academy of SciencesUniversity of Science and Technology of ChinaZhejiang University

Why you should read this

Presents TimeSuite, a grounded instruction-tuning framework that adapts multimodal language models for long-form video comprehension by pairing efficient token compression with explicit timestamp generation to curb hallucinations.

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging for MLLMs. This paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for long video understanding, including a simple yet efficient framework to process long video sequence, a high-quality video dataset for grounded tuning of MLLMs, and a carefully-designed instruction tuning task to explicitly incorporate the grounding supervision in the traditional QA format. Specifically, based on VideoChat, we propose our long-video MLLM, coined as VideoChat-T, by implementing a token shuffling to compress long video tokens and introducing Temporal Adaptive Position Encoding (TAPE) to enhance the temporal awareness of visual representation. Meanwhile, we introduce the TimePro, a comprehensive grounding-centric instruction tuning dataset composed of 9 tasks and 349k high-quality grounded annotations. Notably, we design a new instruction tuning task type, called Temporal Grounded Caption, to peform detailed video descriptions with the corresponding time stamps prediction. This explicit temporal location prediction will guide MLLM to correctly attend on the visual content when generating description, and thus reduce the hallucination risk caused by the LLMs. Experimental results demonstrate that our TimeSuite provides a successful solution to enhance the long video understanding capability of short-form MLLM, achieving improvement of 5.6% and 6.8% on the benchmarks of Egoschema and VideoMME, respectively. In addition, VideoChat-T exhibits robust zero-shot temporal grounding capabilities, significantly outperforming the existing state-of-the-art MLLMs. After fine-tuning, it performs on par with the traditional supervised expert models.

Added

2026-09-26

TRACE: Temporal Grounding Video LLM via Causal Event Modeling

TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Yongxin Guo, Jingyu Liu, Mingda Li, Qingbin Liu, Xi Chen, Xiaoying Tang

OrganizationsGuangdong Provincial Key Laboratory of Future Networks of IntelligenceShenzhen Institute of Artificial Intelligence and Robotics for SocietyTencentThe Chinese University of Hong Kong

Why you should read this

Proposes TRACE, a task-interleaved video language model that advances video temporal grounding by formulating predictions as causal event sequences combining timestamps, saliency scores, and captions rather than relying solely on text generation.

Video Temporal Grounding (VTG) is a crucial capability for video understanding models and plays a vital role in downstream tasks such as video browsing and editing. To effectively handle various tasks simultaneously and enable zero-shot prediction, there is a growing trend in employing video LLMs for VTG tasks. However, current video LLM-based methods rely exclusively on natural language generation, lacking the ability to model the clear structure inherent in videos, which restricts their effectiveness in tackling VTG tasks. To address this issue, this paper first formally introduces causal event modeling framework, which represents video LLM outputs as sequences of events, and predict the current event using previous events, video inputs, and textural instructions. Each event consists of three components: timestamps, salient scores, and textual captions. We then propose a novel task-interleaved video LLM called TRACE to effectively implement the causal event modeling framework in practice. The TRACE process visual frames, timestamps, salient scores, and text as distinct tasks, employing various encoders and decoding heads for each. Task tokens are arranged in an interleaved sequence according to the causal event modeling framework's formulation. Extensive experiments on various VTG tasks and datasets demonstrate the superior performance of TRACE compared to state-of-the-art video LLMs. Our model and code are available at this https URL.

Added

2026-09-26

ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

Jr-Jen Chen, Yu-Chien Liao, Hsi-Che Lin, Yu-Chu Yu, Yen-Chun Chen, Yu-Chiang Frank Wang

OrganizationsMicrosoftNational Taiwan University

Why you should read this

Presents a benchmark and cost-effective data generation pipeline to evaluate and improve multimodal language models on complex cause-and-effect temporal reasoning where questions and answers span non-overlapping video segments.

We introduce REXTIME, a benchmark designed to rigorously test AI models’ ability to perform temporal reasoning within video events. Specifically, REXTIME focuses on reasoning across time, i.e. human-like understanding when the question and its corresponding answer occur in different video segments. This form of reasoning, requiring advanced understanding of cause-and-effect relationships across video segments, poses significant challenges to even the frontier multimodal large language models. To facilitate this evaluation, we develop an automated pipeline for generating temporal reasoning question-answer pairs, significantly reducing the need for labor-intensive manual annotations. Our benchmark includes 921 carefully vetted validation samples and 2,143 test samples, each manually curated for accuracy and relevance. Evaluation results show that while frontier large language models outperform academic models, they still lag behind human performance by a significant 14.3% accuracy gap. Additionally, our pipeline creates a training dataset of 9,695 machine generated samples without manual effort, which empirical studies suggest can enhance the across-time reasoning via fine-tuning.

Added

2026-09-26