Video ReCap: Recursive Captioning of Hour-Long Videos
Md Mohaiminul IslamNgan HoXitong YangTushar NagarajanLorenzo TorresaniGedas Bertasius
Presents a recursive video-language model and benchmark dataset that efficiently generate hierarchical captions across multiple temporal granularities for hour-long untrimmed videos.
Real-world video understanding poses a significant challenge because most natural videos last for minutes or hours and contain complex, nested human behaviors. However, current automated video captioning systems are primarily built for short clips spanning only 5 to 15 seconds. When applied to extended durations, existing models become computationally prohibitive, fail to compress repetitive visual details, and cannot connect immediate actions to overarching goals. As long-form video archives grow rapidly across industries, automated tools must be able to understand visual content across multiple time scales without requiring unsustainable computing resources.
The article introduces and evaluates Video ReCap, a recursive video-language model designed to generate natural language captions across multiple hierarchical levels for videos ranging from several seconds to two hours in duration. The article also introduces Ego4D-HCap, a newly curated benchmark designed to evaluate multi-level captioning on long, untrimmed egocentric video.
The researchers developed a recursive framework that processes long videos in three distinct stages: short clip-level captions describing atomic actions, medium-length segment descriptions covering intermediate steps, and long-range video summaries capturing high-level goals. The architecture uses an off-the-shelf visual encoder, a video-language alignment module built on frozen language models that compresses large volumes of features into compact representations, and a recursive text decoder. Captions generated at lower tiers serve as text inputs for higher tiers alongside sparsely sampled video features. To train the system effectively despite severe data scarcity at longer timescales, the authors implemented a psychologically inspired curriculum learning strategy—moving sequentially from short clips to full summaries—and used large language models to generate pseudo-annotations to expand training data.
The findings demonstrate substantial performance gains across all temporal granularities. Video ReCap outperformed leading baselines, including specialized models such as LaViLa and multimodal zero-shot approaches. For long-range video summaries, Video ReCap achieved a CIDEr metric score of 28.06 compared to 6.54 for standard LaViLa and 20.12 for heavily parameterized language model baselines. Incorporating large language model supervision further raised the segment description score to 46.88 and the full summary score to 29.34. The unified variant, Video ReCap-U, achieved highly competitive performance while utilizing only 113 million trainable parameters compared to 258 million to 586 million in comparative systems. Finally, applying the model’s generated hierarchical captions to long-form video question answering on the EgoSchema benchmark established a new state-of-the-art accuracy of 50.23%, outperforming the prior benchmark leader by 18.13 percentage points.
These results demonstrate that long-range video understanding does not require ingesting complete, uncompressed video streams at massive computational expense. Instead, passing concise text summaries and sparse visual features upward through a hierarchical pipeline enables scalable reasoning over hour-long footage. This recursive approach significantly lowers the operational and memory costs required for long-form video analysis while improving comprehension of intent, intermediate procedures, and complex context.
Organizations handling extensive video data should consider adopting hierarchical processing workflows rather than relying on flat, clip-by-clip analyses or brute-force vision models. Technical teams should explore the unified model architecture (Video ReCap-U) when computing budgets are constrained, as it delivers high accuracy with a substantially smaller parameter footprint. Future development efforts should focus on extending these recursive methods toward real-time stream captioning, conversational video interfaces, and interactive query systems.
The study's primary limitation lies in its primary evaluation on egocentric, activity-focused datasets (Ego4D), which may exhibit different narrative and temporal structures compared to third-person broadcast media, surveillance, or specialized industrial feeds. Additionally, generating higher-tier summaries relies partially on the quality of preceding lower-tier captions, creating potential risks of cascading errors if early action recognition fails. Nonetheless, the consistent performance gains across both captioning and complex question-answering benchmarks provide high confidence in the fundamental recursive methodology.
- Paper: Dense-Captioning Events in Videos, Ranjay Krishna et al. (2017). This paper establishes the foundational task of dense video event captioning and context modeling across extended timelines, which Video ReCap directly extends to recursive, multi-granularity hour-long summarization.
- Paper: Sequence to Sequence -- Video to Text, Subhashini Venugopalan et al. (2015). It introduces the classical sequence-to-sequence framework for translating video frame sequences into natural language descriptions, providing the core baseline paradigm that hierarchical video captioning models build upon.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). It demonstrates how to combine visual encoders with large language models for detailed video dialogue and question-answering, establishing the vision-language foundation used for multi-level video reasoning.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). It introduces scalable space-time attention and progressive multi-frame training for joint video-text representations, informing the video-language encoder design necessary for processing long video durations.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). It provides the foundational framework for learning joint video-text representations from instructional and long-form narrated videos at scale.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). It pioneered cross-modal transformer pre-training for video and language representations to capture high-level semantic event structures.
- Paper: Temporal Convolutional Networks for Action Segmentation and Detection, Colin Lea et al. (2016). It outlines hierarchical temporal modeling techniques for fine-grained action segmentation over long sequences that motivate curriculum learning across temporal granularities.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This technical report advances ultra-long video context understanding and fine-grained visual-language reasoning using explicit textual timestamps and deep cross-layer token injection.
