Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Shangzhe DiZhelun YuGuanghao ZhangHaoyuan LiTao ZhongHao ChengBolin LiWanggui HeFangxun ShuHao Jiang

article2025ICLR120 citations

Develops ReKV, a training-free framework that enables low-latency streaming video question answering by offloading key-value caches to host memory and selectively retrieving query-relevant context into existing video large language models.

Listen

Real-world artificial intelligence applications in domains such as robotics, surveillance, and live broadcasting increasingly demand the ability to process continuous video streams and answer questions in real time. Conventional video question-answering systems generally operate in an offline manner, requiring an entire video to be processed beforehand and often reprocessing video frames for every new query. While existing video large language models can handle short clips, processing long streams often leads to severe memory overflow or visual information loss caused by aggressive frame subsampling and memory compression.

The article introduces and evaluates ReKV (Retrieve In-context Video Key-Value Cache), a training-free framework designed to enable efficient and accurate streaming video question-answering without altering or retraining underlying base models.

To overcome these computational bottlenecks, the authors developed an architecture that decouples video encoding from question-answering across separate processes and hardware. The framework encodes incoming continuous video chunk by chunk using a sliding-window attention mechanism that bounds short-term computational overhead. Rather than discarding past information, computed key-value caches are preserved and offloaded to system memory or disk. When a user asks a question, an in-context retrieval mechanism searches the stored cache to reload only the most query-relevant visual representations onto the graphics processing unit. The researchers evaluated ReKV across multiple base model sizes, from 0.5 billion to 72 billion parameters, using seven established offline and streaming video question-answering benchmarks.

The evaluation yielded several key findings. First, ReKV consistently improved question-answering accuracy across all tested models and benchmarks over baseline uniform sampling methods. When integrated with a 7-billion-parameter model, it raised accuracy on benchmarks such as MLVU from 64.7% to 68.5% and QAEGO4D from 52.8% to 56.0%. Second, ReKV maintained high, stable encoding throughput (11 frames per second for a 7-billion-parameter model and 17 frames per second for a 0.5-billion-parameter model) along with constant graphics memory usage, effectively eliminating out-of-memory errors on long videos. Third, internal cache retrieval—which reuses the model's existing attention representations—consistently outperformed external retrieval tools, reducing query latency from 5.8 seconds down to 3.3 seconds and lowering average computational operations by approximately 15%.

These findings demonstrate that high-performance streaming video analysis can be achieved cost-effectively by leveraging existing model architectures rather than training expensive specialized models from scratch. ReKV provides an operational blueprint for interactive video applications, amortizing encoding costs over multiple queries and scaling efficiently in high-concurrency environments. Organizations deploying real-time video intelligence should consider adopting decoupled architectures with internal key-value retrieval. For practical implementations, internal retrieval is strongly recommended over external models to minimize latency and hardware costs.

Nevertheless, some limitations remain. Offloading video caches requires significant auxiliary storage (approximately 18.8 gigabytes per hour of video for a 7-billion-parameter model), which may become unsustainable for continuous, days-long surveillance streams without future compression or quantization techniques. Additionally, the framework currently relies on fixed frame-retrieval counts and uniform frame grouping rather than dynamic, semantic video segmentation. While confidence in the reported experimental improvements is high across standard benchmarks, broader deployment in production streaming environments will require further testing on domain-specific workloads.

arXiv: 2503.00540
Cover for Streaming Video Question-Answering with In-context Video KV-Cache Retrieval

Abstract

We propose ReKV, a novel training-free approach that enables efficient streaming video question-answering (StreamingVQA), by seamlessly integrating with existing Video Large Language Models (Video-LLMs). Traditional VideoQA systems struggle with long videos, as they must process entire videos before responding to queries, and repeat this process for each new question. In contrast, our approach analyzes long videos in a streaming manner, allowing for prompt responses as soon as user queries are received. Building on a common Video-LLM, we first incorporate a sliding-window attention mechanism, ensuring that input frames attend to a limited number of preceding frames, thereby reducing computational overhead. To prevent information loss, we store processed video key-value caches (KV-Caches) in RAM and disk, reloading them into GPU memory as needed. Additionally, we introduce a retrieval method that leverages an external retriever or the parameters within Video-LLMs to retrieve only query-relevant KV-Caches, ensuring both efficiency and accuracy in question answering. ReKV enables the separation of video encoding and question-answering across different processes and GPUs, significantly enhancing the efficiency of StreamingVQA. Through comprehensive experimentation, we validate the efficacy and practicality of our approach, which significantly boosts efficiency and enhances applicability over existing VideoQA models.

Table of Contents

  • 1 Introduction
  • 2 StreamingVQA: Task Definition and Discussion
  • 3 ReKV: Retrieve In-context Video KV-Cache
  • 4 Experiments
  • 4.1 Benchmark and Metrics
  • 4.2 Implementation Details
  • 4.3 Ablations
  • 4.4 Offline Video Question-answering
  • 4.5 Streaming Video Question-answering
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Additional Implementation Details
  • A.1 Multi-processing Serving
  • A.2 Prompt Templates for VideoQA
  • A.3 KV-Cache Size Calculation
  • B Additional Experiments
  • B.1 Experiments with more Video-LLMs and Benchmark
  • B.2 Fair comparisons with Flash-VStream
  • B.3 Computational Complexity
  • C Limitations and Future Work

Knowls

  1. Knowl 1 — Task Definition of Streaming Video Question-Answering (StreamingVQA)

    definition

    Streaming Video Question-Answering (StreamingVQA\text{StreamingVQA}) is the task of continuously ingesting an incoming video stream while responding to natural language questions about past visual content at arbitrary timestamps.

    Formally, given a continuous video stream VT=[v1,v2,…,vT]V^T = [v_1, v_2, \dots, v_T] consisting of TT sequential frames and a set of NN questions Q={q1,q2,…,qN}Q = \{q_1, q_2, \dots, q_N\}, the goal is to predict the answer to any question qi∈Qq_i \in Q posed at time step tt (where 1≤t≤T1 \le t \le T) using strictly the sequence of frames observed up to that point: Vt=[v1,v2,…,vt]V^t = [v_1, v_2, \dots, v_t] without access to future frames vt+1,…,vTv_{t+1}, \dots, v_T.

    Unlike conventional offline video question-answering (OfflineVQA\text{OfflineVQA}), which assumes a pre-recorded video input of fixed duration and evaluates questions only after the entire sequence is ingested, extStreamingVQA ext{StreamingVQA} requires real-time incremental frame processing and immediate question answering at any point during ingestion. OfflineVQA\text{OfflineVQA} corresponds to the boundary condition of StreamingVQA\text{StreamingVQA} where t=Tt = T for all queries.

  2. Knowl 2 — Video Stream Encoding with Sliding-Window Attention in ReKV

    model/method

    In ReKV (Retrieve In-context Video KV-Cache), a decoder-based Video Large Language Model (Video-LLM) processes an incoming video stream VTV^T incrementally in chunks. To avoid computational overhead and memory overflow over long video streams, encoding uses a sliding-window causal attention mechanism.

    Let P={(kj,vj)}j=1lPP = \{(k_j, v_j)\}_{j=1}^{l_P} denote the accumulated past key and value vectors of length lPl_P, and let X={ti+lP}i=1lXX = \{t_{i+l_P}\}_{i=1}^{l_X} denote the newly arrived chunk of video tokens of length lXl_X. The active local key-value cache within a window of length lLl_L (with lL≤lPl_L \le l_P) is extracted as: L=P[lP−lL+1:lP]L = P_{[l_P - l_L + 1 : l_P]} where LkL_k and LvL_v represent the active key and value vectors in LL, respectively.

    The attention output OO for chunk XX is computed as: O=Attn(WQX,[Lk,WKX],[Lv,WVX])O = \text{Attn}\left(W_Q X, [L_k, W_K X], [L_v, W_V X]\right) where WQ,WK,WVW_Q, W_K, W_V are projection weight matrices of the attention layer, and [⋅,⋅][\cdot, \cdot] denotes concatenation along the sequence dimension.

    Key-value pairs for tokens that fall outside the active sliding window lLl_L are not discarded; instead, they are preserved by offloading to system RAM or secondary disk storage for subsequent query-driven retrieval.

  3. Knowl 3 — Internal and External In-Context Video KV-Cache Retrieval Mechanisms

    model/method

    When a query is received at time tt, ReKV selects a subset of rr relevant video frames (or ⌈r/b⌉\lceil r/b \rceil video blocks of block size bb) from the stored key-value cache to construct the prompt context. ReKV supports two retrieval formulations:

    1. External KV-Cache Retrieval: An external vision-language model (e.g., SigLIP) maps each video frame into an embedding v=fv(v)∈RDv = f_v(v) \in \mathbb{R}^D and the question into a text embedding q=ft(q)∈RDq = f_t(q) \in \mathbb{R}^D. The similarity score between frame vector vv and question vector qq is: Sim(v,q)=v⋅qτ∥v∥2∥q∥2\text{Sim}(v, q) = \frac{v \cdot q}{\tau \|v\|_2 \|q\|_2} where τ\tau is a temperature parameter. For block size b>1b > 1, frame embeddings within each block of bb consecutive frames are averaged before computing cosine similarity.

    2. Internal KV-Cache Retrieval: Internal retrieval extracts representations directly from the self-attention layers of the Video-LLM without requiring external encoders. For an attention layer, the representative vector of a frame containing NfN_f visual tokens is defined by averaging its key vectors across tokens: v=1Nf∑j=1Nfkj∈RD′v = \frac{1}{N_f} \sum_{j=1}^{N_f} k_j \in \mathbb{R}^{D'} where kjk_j is the jj-th key vector concatenated across all attention heads to form dimension D′D'. The question representation is similarly defined by averaging its query vectors across NqN_q tokens: q=1Nq∑k=1Nqqk∈RD′q = \frac{1}{N_q} \sum_{k=1}^{N_q} q_k \in \mathbb{R}^{D'} The similarity score is computed via cosine similarity with τ=1\tau = 1. Internal retrieval executes independently per self-attention layer, allowing distinct layers to retrieve different video blocks across the stream, while reusing already computed hidden representations.

  4. Knowl 4 — Question-Answering Context Formulation and Positional Encoding Strategy

    model/method

    During question answering in ReKV, the retrieved video key-value caches RkR_k and RvR_v are loaded onto the GPU to serve as context for answer generation. For query/decoding tokens XX, the attention operation is computed as: O=Attn(WQX,[Rk,WKX],[Rv,WVX])O = \text{Attn}\left(W_Q X, [R_k, W_K X], [R_v, W_V X]\right) where WQ,WK,WVW_Q, W_K, W_V are attention projection matrices, and Rk,RvR_k, R_v incorporate the retrieved video KV-cache, question tokens, and previously generated output tokens.

    Positional Encoding Implementation: Video-LLMs utilize Rotary Position Embeddings (RoPE). During continuous video stream encoding, RoPE operates within the local sliding window and applies a distance ceiling to distant tokens following the LM-Infinite strategy. During question answering, non-contiguous retrieved video tokens are assigned sequential relative positions (treating them as regular consecutive tokens with standard RoPE) rather than retaining their original temporal timestamps or assigning them static identical positions. Preserving sequential relative positions over retrieved tokens preserves temporal ordering information necessary for video reasoning.

  5. Knowl 5 — Multi-Processing Asynchronous Serving Architecture for StreamingVQA

    model/method

    ReKV decouples video stream encoding from question answering across distinct processes and GPUs:

    1. Video Ingestion Process: A dedicated primary process encodes the incoming video stream at a constant frame rate (e.g., 0.5 FPS) using sliding-window causal attention. It writes the computed key-value cache directly into host RAM, offloading to disk storage if memory capacity is exceeded. Ingestion runs continuously without stalling for query responses.
    2. Worker Process Pool: A pool of worker processes, each hosting an instance of the Video-LLM on separate GPUs, handles question answering.
    3. Query Handling: When a question arrives at timestamp tt, timestamp tt is logged to restrict context to VtV^t. An available worker process retrieves the top-rr relevant key-value vectors from the shared RAM/disk storage using in-context retrieval, reloads these vectors into its GPU memory, and generates the answer autoregressively without interrupting the primary encoding stream.
  6. Knowl 6 — Analytical Video KV-Cache Memory Calculation

    equation

    Assuming 16-bit floating-point precision (FP16, 2 bytes per element), the storage size in bytes of the complete key-value cache generated over TT frames is given by: Cache Size (bytes)=2×L×T×M×H×D×2\text{Cache Size (bytes)} = 2 \times L \times T \times M \times H \times D \times 2 where:

    • The leading factor 22 accounts for storing both Key and Value matrices,
    • LL is the number of transformer layers in the LLM backbone,
    • TT is the total number of encoded video frames (T=FPS×Duration in secondsT = \text{FPS} \times \text{Duration in seconds}),
    • MM is the number of visual tokens per frame,
    • HH is the number of key-value attention heads per layer,
    • DD is the hidden dimension per attention head,
    • The trailing factor 22 represents the byte width per parameter in FP16.

    For LLaVA-OV-7B (L=28L = 28, M=196M = 196, H=4H = 4, D=128D = 128), encoding a 1-hour video at 0.5 FPS0.5\text{ FPS} (T=1800T = 1800) yields: 2×28×1800×196×4×128×2 bytes≈18.8 GB2 \times 28 \times 1800 \times 196 \times 4 \times 128 \times 2 \text{ bytes} \approx 18.8\text{ GB} For LLaVA-OV-0.5B (L=24L = 24, M=196M = 196, H=2H = 2, D=64D = 64), encoding at 0.5 FPS0.5\text{ FPS} for 1 hour yields: 2×24×1800×196×2×64×2 bytes≈4.0 GB2 \times 24 \times 1800 \times 196 \times 2 \times 64 \times 2 \text{ bytes} \approx 4.0\text{ GB}

  7. Knowl 7 — Offline Long-Form Video Question-Answering Benchmark Performance

    data/table

    Evaluating ReKV (0.5 FPS stream encoding, retrieving 64 frames) across four long-form VideoQA benchmarks demonstrates consistent accuracy gains over baseline Video-LLMs without retraining.

    Method MLVU dev Acc. QAEGO4D test Acc. EgoSchema Acc. ActivityNet-QA Acc. / Score
    GPT-4V 49.2 - - 57.0 / -
    GPT-4o 64.6 - - - / -
    Gemini-1.5-Flash - - 65.7 55.3 / -
    Gemini-1.5-Pro - - 72.2 57.5 / -
    Video-ChatGPT-7B 31.3 - - - / -
    LLaMA-VID-7B 33.2 - - 47.4 / 3.30
    MiniGPT4-Video-7B 44.5 - - 44.3 / 3.35
    Video-LLaVA-7B 47.3 - - - / -
    LongVA-7B 56.3 - - 50.0 / -
    VideoStreaming - - 44.1 - / -
    Flash-VStream-7B 50.2 38.2 38.1 51.9 / 3.40
    LLaVA-OV-0.5B 53.2 42.6 29.6 50.5 / 3.02
    + ReKV (0.5 FPS →\to 64 Frames) 56.1 (+2.9) 50.0 (+7.4) 31.0 (+1.4) 52.1 (+1.6) / 3.15 (+0.13)
    LLaVA-OV-7B 64.7 52.8 59.8 56.6 / 3.29
    + ReKV (0.5 FPS →\to 64 Frames) 68.5 (+3.8) 56.0 (+3.2) 60.7 (+0.9) 60.4 (+3.8) / 3.52 (+0.23)

    Acc. denotes accuracy (%), and ActivityNet-QA Score is evaluated on a 1–5 scale using GPT-3.5-turbo.

  8. Knowl 8 — Streaming Video Question-Answering Performance, Latency, and Memory Footprint

    data/table

    Performance, speed, and memory usage on StreamingVQA benchmarks RVS-Ego (60-min videos) and RVS-Movie (30-min videos) tested on an NVIDIA A100 (80GB) GPU with 100 queries scattered across a 1-hour 1080P video at 0.5 FPS.

    Retrieval Method RVS-Ego RVS-Movie Running Speed Memory Usage
    Acc. Score Acc. Score Video Enc. Latency GPU KV-Cache
    Flash-VStream-7B 57.3 4.0 53.1 3.3 14 FPS 2.4s 20 GB -
    LLaVA-OV-7B
    Uniform Sampling 56.2 3.7 43.0 3.3 - 2.9s 21 GB -
    External Retrieval 62.4 3.9 53.6 3.5 11 FPS 5.8s 55 GB 18.8 GB/h
    Internal Retrieval 63.7 4.0 54.4 3.6 11 FPS 3.3s 38 GB 18.8 GB/h
    LLaVA-OV-0.5B
    Uniform Sampling 51.8 3.7 37.2 3.2 - 2.5s 7 GB -
    External Retrieval 54.1 3.8 44.7 3.4 17 FPS 4.1s 37 GB 4.0 GB/h
    Internal Retrieval 54.7 3.9 44.6 3.4 17 FPS 1.6s 19 GB 4.0 GB/h

    Internal retrieval achieves higher accuracy than external retrieval (63.7% vs 62.4% on RVS-Ego for 7B) while lowering QA response latency from 5.8s to 3.3s and peak GPU memory from 55 GB to 38 GB.

  9. Knowl 9 — Controlled Comparison and Query Complexity Scaling Against Memory Compression

    data/table

    Under identical architectures (CLIP-ViT-L/14, 2-layer MLP, Vicuna-7B-v1.5) and training data (InternVid-232K), ReKV is compared directly against the memory-compression baseline Flash-VStream.

    Accuracy Comparison Across Benchmarks:

    Method MLVU dev QAEGO4D test EgoSchema RVS-Movie RVS-Ego
    Base 49.8 39.0 42.6 47.2 54.1
    Base+Flash 51.0 37.4 41.2 50.1 55.4
    Base+ReKV 51.9 (+0.9) 40.5 (+3.1) 43.7 (+2.5) 51.9 (+1.8) 54.7 (-0.7)
    Original Flash 50.2 38.2 38.1 53.1 57.3

    Computational Complexity per QA Query vs. Query Frequency in a 1-Hour Video:

    #QAs TFLOPs / QA TMACs / QA
    Base Flash ReKV (Ext) ReKV (Int) Base Flash ReKV (Ext) ReKV (Int)
    100 22.4 15.5 21.7 18.5 11.2 7.8 10.8 9.2
    200 12.7 14.1 11.4 9.6 6.4 7.1 5.7 4.8
    360 8.5 13.8 6.8 5.6 4.3 6.8 3.3 2.8

    Because ReKV encodes the stream once and reuses the KV cache across queries, per-query TFLOPs and TMACs drop rapidly with increasing question frequency, outperforming Flash-VStream at high query rates (e.g., 5.6 vs 13.8 TFLOPs at 360 QAs).

  10. Knowl 10 — Ablation Analysis of Retrieval Recall, Frame Count, and Block Size

    empirical result

    Ablations on QAEGO4D-test-mc and MLVU-dev-mc using LLaVA-OV-7B establish the empirical characteristics of in-context retrieval:

    1. Recall vs. VideoQA Accuracy (QAEGO4D-test-mc):
    • Uniform Sampling: 6.1% recall →\to 53.0% Acc.
    • External Retrieval: 58.1% recall →\to 54.2% Acc.
    • Internal Retrieval: 70.5% recall →\to 56.0% Acc.
    • Oracle Retrieval (100% recall upper bound): 64.4% Acc. Frame recall exhibits a direct positive correlation with QA accuracy.
    1. Task Type Sensitivity (MLVU-dev-mc):
    • Internal retrieval improves Single Detail tasks (PlotQA from 69.8% to 76.3%) and Holistic tasks (Topic from 87.9% to 90.1%, Anomaly from 72.0% to 74.5%). External retrieval improves Needle (74.1% to 78.6%) and Ego (59.7% to 69.6%) but lags on Holistic tasks (Topic 84.5%, Anomaly 63.0%).
    1. Retrieved Frame Count (rr) and Block Size (bb):
    • Increasing retrieved frames r∈{8,16,32,48,64,80}r \in \{8, 16, 32, 48, 64, 80\} steadily increases accuracy until plateauing around r=64r=64 on MLVU, where additional frames introduce distracting context.
    • Increasing retrieval block size b∈{1,2,4,8,16}b \in \{1, 2, 4, 8, 16\} for fixed r=64r=64 decreases MLVU accuracy (from 68.5% at b=1b=1 to ~61% at b=16b=16) because multi-detail and holistic reasoning require temporally dispersed frames. In contrast, QAEGO4D accuracy remains flat (~56%) across block sizes due to its dependence on a single local event.
  11. Knowl 11 — Generalization Across Video-LLM Architectures and Scales

    empirical result

    Integrating ReKV without fine-tuning improves VideoQA accuracy across diverse Video-LLM architectures and scales on MLVU dev, QAEGO4D test, EgoSchema, and CGBench:

    • Video-LLaVA-7B (0.5 FPS →\to 8 Frames vs. 8 uniform frames): MLVU: 46.5%→47.2%46.5\% \to 47.2\% (+0.7); QAEGO4D: 37.0%→37.9%37.0\% \to 37.9\% (+0.9); EgoSchema: 41.4%→42.2%41.4\% \to 42.2\% (+0.8); CGBench: 18.7%→19.2%18.7\% \to 19.2\% (+0.5).
    • LongVA-7B (0.5 FPS →\to 32 Frames vs. 32 uniform frames): MLVU: 57.3%→58.6%57.3\% \to 58.6\% (+1.3); QAEGO4D: 42.8%→45.6%42.8\% \to 45.6\% (+2.8); EgoSchema: 42.5%→42.7%42.5\% \to 42.7\% (+0.2); CGBench: 26.1%→26.4%26.1\% \to 26.4\% (+0.3).
    • LLaVA-OV-0.5B (0.5 FPS →\to 64 Frames vs. 64 uniform frames): MLVU: 53.2%→56.1%53.2\% \to 56.1\% (+2.9); QAEGO4D: 42.6%→50.0%42.6\% \to 50.0\% (+7.4); EgoSchema: 29.6%→31.0%29.6\% \to 31.0\% (+1.4); CGBench: 21.4%→21.7%21.4\% \to 21.7\% (+0.3).
    • LLaVA-OV-7B (0.5 FPS →\to 64 Frames vs. 64 uniform frames): MLVU: 64.7%→68.5%64.7\% \to 68.5\% (+3.8); QAEGO4D: 52.8%→56.0%52.8\% \to 56.0\% (+3.2); EgoSchema: 59.8%→60.7%59.8\% \to 60.7\% (+0.9); CGBench: 31.1%→33.9%31.1\% \to 33.9\% (+2.8).
    • LLaVA-OV-72B (0.1 FPS →\to 32 Frames vs. 32 uniform frames with model sharding): MLVU: 71.9%→72.6%71.9\% \to 72.6\% (+0.7); QAEGO4D: 53.6%→57.0%53.6\% \to 57.0\% (+3.4); EgoSchema: 59.6%→62.0%59.6\% \to 62.0\% (+2.4); CGBench: 37.2%→40.5%37.2\% \to 40.5\% (+3.3).
  12. Knowl 12 — Limitations of the ReKV Streaming Framework

    limitation

    The ReKV framework possesses four primary limitations:

    1. Long-Stream Memory Growth: Although storing KV-caches in RAM/disk is feasible for hour-long streams (18.8 GB/h for 7B models), continuous execution over days or weeks (e.g., in 24/7 surveillance) leads to unbounded cache growth unless integrated with KV quantization, token pruning, or eviction.
    2. Fixed Temporal Block Grouping: Grouping consecutive frames using a static block size bb ignores semantic video boundaries and can segment continuous actions unnaturally.
    3. Static Retrieval Budget: ReKV retrieves a fixed number of frames rr across all queries, rather than dynamically scaling context size based on question complexity or temporal scope.
    4. Benchmark Limitations: There is a scarcity of dedicated StreamingVQA evaluation benchmarks with fine-grained temporal ground-truth annotations.

Coverage note — None. All contributed algorithms, mathematical formulations, systems designs, empirical benchmark evaluations, complexity analyses, ablations, and limitations have been extracted into self-contained knowls.

References

  1. 1.Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens. In CVPR Workshop, 2024.
  2. 2.Ivana Balazevic, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni, Skanda Koppula, and Olivier J Henaff. Memory consolidation enables long-context video understanding. In ICML, 2024.
  3. 3.Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. In ICML, 2022.
  4. 4.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR, 2015.
  5. 5.Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, and Limin Wang. Cg-bench: Clue-grounded question answering benchmark for long video understanding. In ICLR, 2025a.
  6. 6.Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. Videollm-online: Online video large language model for streaming video. In CVPR, 2024.
  7. 7.Qirui Chen, Shangzhe Di, and Weidi Xie. Grounded multi-hop videoqa in long-form egocentric videos. In AAAI, 2025b.
  8. 8.Shangzhe Di and Weidi Xie. Grounded question-answering in long egocentric videos. In CVPR, 2024.
  9. 9.Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. In CVPR, 2023.
  10. 10.Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, and Jun Wang. Human-like episodic memory for infinite context llms. In ICLR, 2025.
  11. 11.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The” something something” video database for learning and evaluating visual common sense. In ICCV, 2017.
  12. 12.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In CVPR, 2022.
  13. 13.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. Retrieval augmented language model pre-training. In ICML, 2020.
  14. 14.Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. Lm-infinite: Simple on-the-fly length generalization for large language models. arXiv preprint arXiv:2308.16137, 2023.
  15. 15.Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. In CVPR, 2024.
  16. 16.Lianghua Huang, Xin Zhao, and Kaiqi Huang. Got-10k: A large high-diversity benchmark for generic object tracking in the wild. TPAMI, 2019.
  17. 17.Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. Movienet: A holistic dataset for movie understanding. In ECCV, 2020.
  18. 18.Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius. Video recap: Recursive captioning of hour-long videos. In CVPR, 2024.
  19. 19.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatiotemporal reasoning in visual question answering. In CVPR, 2017.
  20. 20.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv:1705.06950, 2017.
  21. 21.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rockt ¨ aschel, et al. Retrieval-augmented gener- ¨ ation for knowledge-intensive nlp tasks. In NeurIPS, 2020.
  22. 22.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024a.
  23. 23.Jingyao Li, Han Shi, Xin Jiang, Zhenguo Li, Hong Xu, and Jiaya Jia. Quickllama: Query-aware inference acceleration for large language models. In COLING, 2025.
  24. 24.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023.
  25. 25.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In CVPR, 2024b.
  26. 26.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. In ECCV, 2024c.
  27. 27.Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. Snapkv: Llm knows what you are looking for before generation. In NeurIPS, 2024d.
  28. 28.Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-LLaVA: Learning united visual representation by alignment before projection. In EMNLP, 2024.
  29. 29.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. In ACL, 2024.
  30. 30.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. NeurIPS, 2023.
  31. 31.Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem. Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. In ECCV, 2018.
  32. 32.OpenAI. Gpt-4, March 2023a. URL https://cdn.openai.com/papers/gpt-4-system-card.pdf.
  33. 33.OpenAI. Gpt-4v, September 2023b. URL https://openai.com/index/gpt-4v-system-card/.
  34. 34.OpenAI. Gpt-4o, May 2024. URL https://openai.com/index/hello-gpt-4o/.
  35. 35.Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models. In NeurIPS, 2024.
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  37. 37.Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham. In-context retrieval-augmented language models. In ACL, 2023.
  38. 38.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 2024.
  39. 39.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  40. 40.Jiahao Wang, Guo Chen, Yifei Huang, Limin Wang, and Tong Lu. Memory-and-anticipation transformer for online action understanding. In ICCV, 2023.
  41. 41.Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. Long-term feature banks for detailed video understanding. In CVPR, 2019.
  42. 42.Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In CVPR, 2022.
  43. 43.Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao, Yankai Lin, Zhengyan Zhang, Zhiyuan Liu, and Maosong Sun. Infllm: Training-free long-context extrapolation for llms with an efficient context memory. In ICML Workshop, 2024a.
  44. 44.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. In ICLR, 2024b.
  45. 45.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa:next phase of question-answering to explaining temporal actions. In CVPR, 2021.
  46. 46.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In ACM Multimedia, 2017.
  47. 47.Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. Retrieval-augmented egocentric video captioning. In CVPR, 2024.
  48. 48.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. In NeurIPS, 2022.
  49. 49.X Ye. Calflops: A flops and params calculate tool for neural networks in pytorch framework, 2023.
  50. 50.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In AAAI, 2019.
  51. 51.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, 2023.
  52. 52.Ce Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang, Shoubin Yu, Mohit Bansal, and Gedas Bertasius. A simple llm framework for long-range video question-answering. arXiv preprint arXiv:2312.17235, 2023.
  53. 53.Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, and Xiaojie Jin. Flash-vstream: Memory-based real-time understanding for long video streams. arXiv preprint arXiv:2406.08085, 2024a.
  54. 54.Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024b.
  55. 55.Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava-next: A strong zero-shot video understanding model, April 2024c. URL https://llava-vl.github.io/blog/2024-04-30-llava-next-video/.
  56. 56.Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. Mlvu: A comprehensive benchmark for multi-task long video understanding. arXiv preprint arXiv:2406.04264, 2024a.
  57. 57.Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, and Cordelia Schmid. Streaming dense video captioning. In CVPR, 2024b.

Citation

MLA
Di, S., et al. “Streaming Video Question-Answering with In-context Video KV-Cache Retrieval”. arXiv, 2025, https://doi.org/10.48550/arxiv.2503.00540.
APA
Di, S., Yu, Z., Zhang, G., Li, H., Zhong, T., Cheng, H., Li, B., He, W., Shu, F., & Jiang, H. (2025). Streaming Video Question-Answering with In-context Video KV-Cache Retrieval. arXiv. https://doi.org/10.48550/arxiv.2503.00540
Chicago
Di, S., Z. Yu, G. Zhang, et al. 2025. “Streaming Video Question-Answering with In-context Video KV-Cache Retrieval”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2503.00540.
Harvard
Di, S. et al. (2025) “Streaming Video Question-Answering with In-context Video KV-Cache Retrieval”. arXiv. Available at: https://doi.org/10.48550/arxiv.2503.00540.
Vancouver
1. Di S, Yu Z, Zhang G, Li H, Zhong T, Cheng H, Li B, He W, Shu F, Jiang H (2025) Streaming Video Question-Answering with In-context Video KV-Cache Retrieval. https://doi.org/10.48550/arxiv.2503.00540

BibTeX

@misc{https://doi.org/10.48550/arxiv.2503.00540,
  doi = {10.48550/ARXIV.2503.00540},
  url = {https://arxiv.org/abs/2503.00540},
  author = {Di, Shangzhe and Yu, Zhelun and Zhang, Guanghao and Li, Haoyuan and Zhong, Tao and Cheng, Hao and Li, Bolin and He, Wanggui and Shu, Fangxun and Jiang, Hao},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Streaming Video Question-Answering with In-context Video KV-Cache Retrieval},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors