Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
Shangzhe DiZhelun YuGuanghao ZhangHaoyuan LiTao ZhongHao ChengBolin LiWanggui HeFangxun ShuHao Jiang
Develops ReKV, a training-free framework that enables low-latency streaming video question answering by offloading key-value caches to host memory and selectively retrieving query-relevant context into existing video large language models.
Real-world artificial intelligence applications in domains such as robotics, surveillance, and live broadcasting increasingly demand the ability to process continuous video streams and answer questions in real time. Conventional video question-answering systems generally operate in an offline manner, requiring an entire video to be processed beforehand and often reprocessing video frames for every new query. While existing video large language models can handle short clips, processing long streams often leads to severe memory overflow or visual information loss caused by aggressive frame subsampling and memory compression.
The article introduces and evaluates ReKV (Retrieve In-context Video Key-Value Cache), a training-free framework designed to enable efficient and accurate streaming video question-answering without altering or retraining underlying base models.
To overcome these computational bottlenecks, the authors developed an architecture that decouples video encoding from question-answering across separate processes and hardware. The framework encodes incoming continuous video chunk by chunk using a sliding-window attention mechanism that bounds short-term computational overhead. Rather than discarding past information, computed key-value caches are preserved and offloaded to system memory or disk. When a user asks a question, an in-context retrieval mechanism searches the stored cache to reload only the most query-relevant visual representations onto the graphics processing unit. The researchers evaluated ReKV across multiple base model sizes, from 0.5 billion to 72 billion parameters, using seven established offline and streaming video question-answering benchmarks.
The evaluation yielded several key findings. First, ReKV consistently improved question-answering accuracy across all tested models and benchmarks over baseline uniform sampling methods. When integrated with a 7-billion-parameter model, it raised accuracy on benchmarks such as MLVU from 64.7% to 68.5% and QAEGO4D from 52.8% to 56.0%. Second, ReKV maintained high, stable encoding throughput (11 frames per second for a 7-billion-parameter model and 17 frames per second for a 0.5-billion-parameter model) along with constant graphics memory usage, effectively eliminating out-of-memory errors on long videos. Third, internal cache retrieval—which reuses the model's existing attention representations—consistently outperformed external retrieval tools, reducing query latency from 5.8 seconds down to 3.3 seconds and lowering average computational operations by approximately 15%.
These findings demonstrate that high-performance streaming video analysis can be achieved cost-effectively by leveraging existing model architectures rather than training expensive specialized models from scratch. ReKV provides an operational blueprint for interactive video applications, amortizing encoding costs over multiple queries and scaling efficiently in high-concurrency environments. Organizations deploying real-time video intelligence should consider adopting decoupled architectures with internal key-value retrieval. For practical implementations, internal retrieval is strongly recommended over external models to minimize latency and hardware costs.
Nevertheless, some limitations remain. Offloading video caches requires significant auxiliary storage (approximately 18.8 gigabytes per hour of video for a 7-billion-parameter model), which may become unsustainable for continuous, days-long surveillance streams without future compression or quantization techniques. Additionally, the framework currently relies on fixed frame-retrieval counts and uniform frame grouping rather than dynamic, semantic video segmentation. While confidence in the reported experimental improvements is high across standard benchmarks, broader deployment in production streaming environments will require further testing on domain-specific workloads.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). Introduces the foundation vision-language architecture and spatiotemporal positional encoding principles evaluated as base models in ReKV's streaming framework.
- Paper: Efficient Streaming Language Models with Attention Sinks, Guangxuan Xiao et al. (2023). Establishes sliding-window attention and streaming key-value cache management principles that underlie ReKV's bounded online encoding pipeline.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). Provides the long-horizon video-language evaluation benchmark and defines the core bottlenecks of multi-hour video understanding that ReKV addresses.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). Presents key-value memory bandwidth analysis and foundational cache sharing mechanics essential for understanding key-value cache offloading and retrieval.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). Establishes the standard paradigm and conversational video question-answering formulation adapted by ReKV into an online streaming setting.
- Paper: ReToken: One Token to Improve Vision-Language Models for Visual Retrieval, Yao Xiao et al.. Introduces a single-token internal retrieval mechanism that learns compact visual query representations, complementing ReKV's training-free key-value cache retrieval.
- Paper: ThinK: Thinner Key Cache by Query-Driven Pruning, Yuhui Xu et al. (2025). Investigates channel-level key cache pruning to compress stored representations, addressing ReKV's limitation regarding large auxiliary storage footprints.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). Develops query-adaptive frame selection modules to dynamically prioritize decisive temporal moments instead of ReKV's uniform frame grouping.
- Paper: QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead, Amir Zandieh et al. (2025). Extends key-value cache efficiency through ultra-low-bit quantization, offering a concrete compression solution for ReKV's extensive offloaded video caches.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). Explores training-free visual token filtering and compression within multimodal models to further minimize prefill latency and memory demands.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). Proposes block-sparse attention pruning that can accelerate attention computations over the long-context visual key-value sequences managed by ReKV.
- Paper: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression, DeepSeek-AI (2026). Pushes system-level key-value compression limits with bounded replay schemes and persistent SSD offloading architectures designed for million-token multimodal contexts.
