Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video KV-cache retrieval

Video KV-cache retrieval is a mechanism used in video large language models and multimodal artificial intelligence systems to store and selectively fetch precomputed key-value states of processed video frames based on their relevance to a specific user query. Rather than repeatedly re-encoding video inputs or keeping massive sequences of visual tokens in active GPU memory, the system offloads intermediate attention states to system memory or storage and reloads only the query-pertinent key-value pairs during inference. This approach significantly reduces computational overhead and memory constraints during long-form or streaming video understanding, enabling efficient, low-latency reasoning and question answering over extended visual contexts while preserving essential temporal and visual information.

1 item

Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Shangzhe Di, Zhelun Yu, Guanghao Zhang, Haoyuan Li, Tao Zhong, Hao Cheng, Bolin Li, Wanggui He, Fangxun Shu, Hao Jiang

OrganizationsAlibaba GroupShanghai Jiao Tong University

Why you should read this

Develops ReKV, a training-free framework that enables low-latency streaming video question answering by offloading key-value caches to host memory and selectively retrieving query-relevant context into existing video large language models.

We propose ReKV, a novel training-free approach that enables efficient streaming video question-answering (StreamingVQA), by seamlessly integrating with existing Video Large Language Models (Video-LLMs). Traditional VideoQA systems struggle with long videos, as they must process entire videos before responding to queries, and repeat this process for each new question. In contrast, our approach analyzes long videos in a streaming manner, allowing for prompt responses as soon as user queries are received. Building on a common Video-LLM, we first incorporate a sliding-window attention mechanism, ensuring that input frames attend to a limited number of preceding frames, thereby reducing computational overhead. To prevent information loss, we store processed video key-value caches (KV-Caches) in RAM and disk, reloading them into GPU memory as needed. Additionally, we introduce a retrieval method that leverages an external retriever or the parameters within Video-LLMs to retrieve only query-relevant KV-Caches, ensuring both efficiency and accuracy in question answering. ReKV enables the separation of video encoding and question-answering across different processes and GPUs, significantly enhancing the efficiency of StreamingVQA. Through comprehensive experimentation, we validate the efficacy and practicality of our approach, which significantly boosts efficiency and enhances applicability over existing VideoQA models.

Added

2026-09-26