Built independently by an author, for readers. Read the story and support ChapterPal

keyword

recursive video-language model

A recursive video-language model is a multimodal artificial intelligence architecture designed to process video and natural language by hierarchically and repeatedly applying its representations across multiple temporal granularities. Rather than treating a video as a single uniform sequence of frames, the model operates recursively, first interpreting fine-grained, short-duration clips and then using those outputs to construct representations and text descriptions for progressively longer segments and full-length videos. This recursive processing allows the architecture to scale efficiently across videos of widely varying durations, maintaining fine-grained details of atomic actions while simultaneously capturing broad contextual narratives for tasks such as multi-level captioning, long-form summarization, and complex video question answering.

1 item

Video ReCap: Recursive Captioning of Hour-Long Videos

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, Gedas Bertasius

OrganizationsMetaUniversity of North Carolina at Chapel Hill

Why you should read this

Presents a recursive video-language model and benchmark dataset that efficiently generate hierarchical captions across multiple temporal granularities for hour-long untrimmed videos.

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for minutes or hours and have a complex hierarchical structure spanning different temporal granularities. We propose Video ReCap, a recursive video captioning model that can process video inputs of dramatically different lengths (from 1 second to 2 hours) and output video captions at multiple hierarchy levels. The recursive video-language architecture exploits the synergy between different video hierarchies and can process hour-long videos efficiently. We utilize a curriculum learning training scheme to learn the hierarchical structure of videos, starting from clip-level captions describing atomic actions, then focusing on segment-level descriptions, and concluding with generating summaries for hour-long videos. Furthermore, we introduce Ego4D-HCap dataset by augmenting Ego4D with 8,267 manually collected long-range video summaries. Our recursive model can flexibly generate captions at different hierarchy levels while also being useful for other complex video understanding tasks, such as VideoQA on EgoSchema. Data, code, and models are publicly available at https://sites.google.com/view/vidrecap.

Added

2026-09-26