keyword
recursive video-language model
A recursive video-language model is a multimodal artificial intelligence architecture designed to process video and natural language by hierarchically and repeatedly applying its representations across multiple temporal granularities. Rather than treating a video as a single uniform sequence of frames, the model operates recursively, first interpreting fine-grained, short-duration clips and then using those outputs to construct representations and text descriptions for progressively longer segments and full-length videos. This recursive processing allows the architecture to scale efficiently across videos of widely varying durations, maintaining fine-grained details of atomic actions while simultaneously capturing broad contextual narratives for tasks such as multi-level captioning, long-form summarization, and complex video question answering.
1 item

