StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

Xin DingHao WuYifan YangShiqi JiangQianxi ZhangDonglin BaiZhibo ChenTing Cao

article2025ICCV48 citations

Proposes StreamMind, an event-gated video dialogue framework that pairs state-space feature extraction with selective language model invocation to enable proactive, real-time video understanding at 100 frames per second.

Listen

Real-time human-AI interaction systems, such as assistive AI companions and interactive media, require artificial intelligence models to continuously process live video streams and respond proactively without needing explicit user prompts at every step. However, current systems struggle with a fundamental computational bottleneck: processing incoming video frames sequentially using large language models creates severe latency, typically capping real-time performance at around 10 to 15 frames per second. This limitation prevents models from achieving true real-time alignment between ongoing real-world events and generated responses, especially in high-speed environments.

The article demonstrates and evaluates STREAMMIND, a novel video large language model framework designed to overcome this bottleneck. The objective is to achieve full-frame-rate streaming video dialogue by decoupling continuous visual perception from cognitive response generation, thereby enabling proactive, always-on AI interactions at speeds of up to 100 frames per second on a single graphics processing unit.

The approach introduces an event-gated architecture inspired by human cognition. Instead of invoking the resource-heavy large language model at every video frame, the system uses a compact, 56-million-parameter Event-Preserving Feature Extractor to process visual inputs with constant computational cost per frame. An intermediate Cognition Gate—constructed by fine-tuning the first four shallow layers of the language model—monitors the continuous perception stream against user queries and only triggers the deeper language model when a relevant event occurs. The authors evaluated the framework through extensive experiments across standard streaming and offline benchmarks, including the Ego4D and SoccerNet datasets, measuring response accuracy, timing validity, language quality, and processing throughput.

The experimental findings show substantial improvements over previous state-of-the-art systems. First, the framework processed video streams at up to 100 frames per second on modern hardware, representing roughly a tenfold throughput advantage over existing online baselines that falter past 10 frames per second. Second, the system achieved higher temporal responsiveness, improving trigger accuracy from roughly 32% to over 43% on Ego4D and from 31% to over 52% on SoccerNet. Third, the model enhanced language generation quality and correctness, boosting correctness on SoccerNet from 53.5% to 89.2% while maintaining lower perplexity and latency. Finally, the framework surpassed prior approaches across six traditional offline recognition and forecasting benchmarks.

These results imply that high-frame-rate, proactive video dialogue is computationally viable for deployment without requiring unsustainable hardware scaling. By reducing quadratic and cubic processing overheads to constant-cost perception loops, the framework substantially lowers inference latency and operational infrastructure costs. This makes real-time AI assistance practical for demanding domains such as high-refresh gaming AI, live sports commentary, and industrial human-robot collaboration.

Organizations developing interactive video and embodied AI systems should adopt decoupled, event-gated architectures rather than per-frame language model invocation. Teams implementing this architecture should utilize shallow language model layers for the gating mechanism and apply empirical loss-weighting formulas to manage the heavy data imbalance between event and non-event frames. The primary limitations include reliance on existing benchmark datasets with specific annotation densities and the necessity of tuning dataset-specific weighting parameters for optimal gating accuracy. Overall confidence in the performance gains is high across standard benchmarks, though organizations should validate the framework in domain-specific pilot environments before full production deployment.

arXiv: 2503.06220
Cover for StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition

Abstract

With the rise of real-world human-AI interaction applications, such as AI assistants, the need for Streaming Video Dialogue is critical. To address this need, we introduce StreamMind, a video LLM framework that achieves ultra-FPS streaming video processing (100 fps on a single A100) and enables proactive, always-on responses in real time, without explicit user intervention.

To solve the key challenge of the contradiction between linear video streaming speed and quadratic transformer computation cost, we propose a novel perception-cognition interleaving paradigm named ''event-gated LLM invocation'', in contrast to the existing per-time-step LLM invocation. By introducing a Cognition Gate network between the video encoder and the LLM, LLM is only invoked when relevant events occur. To realize the event feature extraction with constant cost, we propose Event-Preserving Feature Extractor (EPFE) based on state-space method, generating a single perception token for spatiotemporal features. These techniques enable the video LLM with full-FPS perception and real-time cognition response.

Experiments on Ego4D and SoccerNet streaming tasks, as well as standard offline benchmarks, demonstrate state-of-the-art performance in both model capability and real-time efficiency, paving the way for ultra-high-FPS applications, such as Game AI and interactive media. The code and data is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Task Discussion and Definition
  • 2.2 The Overview of StreamMind
  • 2.3 Perception Phase
  • 2.4 Cognition Gate
  • 2.5 The Training Strategy of StreamMind
  • 3 Experiment
  • 3.1 Experimental Settings
  • 3.2 Evaluate Metrics
  • 3.3 Comprehensive Evaluation on Streaming and Offline Video
  • 3.4 Real-time Inference Efficiency
  • 4 Perception Phase Visualization Experiment
  • 5 Ablation
  • 5.1 Impact of Silence-Response Sample Imbalance
  • 5.2 The Ablation of Cognition Gate
  • 5.3 The performance of Event-preserving Feature extractor
  • 6 Related Work
  • 6.1 Offline VideoLLMs.
  • 6.2 Online VideoLLMs.
  • 7 Visualization of demo
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Problem Formulation of Streaming Video Dialogue

    definition

    Streaming Video Dialogue (StreamingVD) extends conventional offline video-grounded question answering to real-time, continuous video streams where an agent must proactively decide when to respond and generate conversational outputs without requiring explicit, per-turn user triggers.

    Given an incoming video stream VT=[v1,v2,…,vT]V^T = [v_1, v_2, \dots, v_T] containing TT frames and a user query qq provided at start timestamp tst_s and active until termination time tet_e (1<ts<te≤T1 < t_s < t_e \le T), the model evaluates whether to output a text token sequence or remain silent at each discrete time step ti∈[ts,te]t_i \in [t_s, t_e]:

    max⁡P([Resti]∣[Ctx<ti],[Fti])\max P([\text{Res}_{t_i}] \mid [\text{Ctx}_{<t_i}], [F_{t_i}])

    where:

    • [Resti]∈{[Txtti],[EOSti]}[\text{Res}_{t_i}] \in \{[\text{Txt}_{t_i}], [\text{EOS}_{t_i}]\} represents generating a linguistic response [Txtti][\text{Txt}_{t_i}] or remaining silent using the end-of-sequence indicator [EOSti][\text{EOS}_{t_i}].
    • [Ctx<ti][\text{Ctx}_{<t_i}] denotes the historical context tokens accumulated from preceding video frames and user queries prior to timestamp tit_i.
    • [Fti][F_{t_i}] denotes visual features extracted from video frame vtiv_{t_i} at time step tit_i.
  2. Knowl 2 — Event-Gated LLM Invocation Paradigm

    model/method

    To overcome the cubic computational bottleneck O(n3)\mathcal{O}(n^3) of invoking a quadratic-complexity Large Language Model (LLM) at every incoming time step nn of an O(n)\mathcal{O}(n) video stream, STREAMMIND employs an event-gated perception-cognition interleaving framework that splits video understanding into three coordinated phases:

    1. Perception Phase: Visual frames are continuously ingested at full frame rate. A frozen spatial visual backbone (e.g., CLIP-ViT) extracts per-frame spatial embeddings, which are fed into a recurrent Event-Preserving Feature Extractor (EPFE) to produce a single compact perception token [Ftiper][F^{\text{per}}_{t_i}] and update internal hidden states. Tokens are stored in a Perception Memory buffer MtiperM^{\text{per}}_{t_i}.
    2. Cognition Gate Judgment: A lightweight Cognition Gate evaluates the current perception token [Ftiper][F^{\text{per}}_{t_i}] alongside the user query [Prompt][\text{Prompt}] at each step tit_i:

    max⁡P([Resti]∣G([Prompt],[Ftiper]))\max P([\text{Res}_{t_i}] \mid \mathcal{G}([\text{Prompt}], [F^{\text{per}}_{t_i}]))

    where [Resti]∈{⟨/response⟩,⟨/silence⟩}[\text{Res}_{t_i}] \in \{\langle /\text{response}\rangle, \langle /\text{silence}\rangle\}. The full LLM is bypassed if the gate outputs ⟨/silence⟩\langle /\text{silence}\rangle. 3. Cognition Phase: When the gate triggers a ⟨/response⟩\langle /\text{response}\rangle decision upon detecting a query-relevant event, historical perception tokens are sampled from MtiperM^{\text{per}}_{t_i} into a Cognition Pooling module and forwarded to the full LLM decoder to generate the final textual commentary or answer.

  3. Knowl 3 — Event-Preserving Feature Extractor

    model/method

    The Event-Preserving Feature Extractor (EPFE) replaces multi-frame cross-attention or 3D convolutional projection modules (e.g., Q-Former, STC) with a continuous state-space formulation based on the Selective State Space Model (Selective SSM / Mamba architecture). With approximately 56M56\text{M} parameters, EPFE compresses spatial embeddings into a single perception token per frame with constant O(1)\mathcal{O}(1) time complexity per step.

    For an input frame spatial feature xt=CLIP(vt)x_t = \text{CLIP}(v_t) at time step tt, the state update and token generation follow the discrete state-space formulation:

    ht+1=Aht+Bxth_{t+1} = A h_t + B x_t

    yt=Chty_t = C h_t

    where hth_t is the internal recurrent state vector, A,B,CA, B, C are learnable state-space matrices, and yt=[Ftper]y_t = [F^{\text{per}}_t] is the output perception token for time step tt. This recurrence enables the model to retain long-term event semantics across noise and distractors while computing tokens at a fixed per-frame cost.

  4. Knowl 4 — Cognition Gate via Shallow Layer Transfer

    model/method

    The Cognition Gate determines whether a query-relevant physical event has occurred based only on the prompt and the current frame's perception token. Because shallow feed-forward layers and self-attention blocks of an LLM possess foundational world knowledge and multimodal semantic parsing capabilities without the computational overhead of the full model, the gate is constructed via Shallow Layer Transfer.

    The gate copies and reuses the first 4 transformer layers of the LLM. It is fine-tuned autoregressively to output binary decision tokens ⟨/response⟩\langle /\text{response}\rangle or ⟨/silence⟩\langle /\text{silence}\rangle in a single forward step. This design leverages early-stage linguistic representations while maintaining low latency.

  5. Knowl 5 — Streaming Video Dataset Construction Algorithm

    algorithm

    To adapt offline narrated video datasets for training both the streaming LLM and the binary Cognition Gate, redundant narration timestamps are merged and negative silence indicators are inserted across non-annotated frames.

    Input: Offline video dataset C = {(c_i, t_i)}_{i=1}^n where c_i is the caption at timestamp t_i
    Output: Streaming dataset D with caption events and frame-level gate labels
    D = []
    cap = c_1
    time = t_1
    for i = 2 to n do
        if c_i != cap then
            Append (cap, time) to D
            cap = c_i
            time = t_i
        end if
    end for
    Append (cap, time) to D
    for each interval between annotated events d_i and d_{i+1} in D do
        for each frame f in the interval do
            if f is the initial frame of event d_i then
                Label f with </response>
            else
                Label f with </silence>
            end if
        end for
    end for
    return D
  6. Knowl 6 — Optimal Silence-Response Class Balancing Loss Weight

    equation

    During Cognition Gate training, the proportion of silence frames to response frames is highly skewed (for instance, 310:1310:1 in Ego4D and 71:171:1 in SoccerNet). Training with standard cross-entropy leads to decision collapse toward perpetual silence.

    A weighted cross-entropy loss is applied with class weight WsW_s for silence tokens and (1−Ws)(1 - W_s) for response tokens. Based on empirical evaluation across datasets with varying silence-to-response token ratios P=NsilenceNresponseP = \frac{N_{\text{silence}}}{N_{\text{response}}}, the optimal balancing weight WsoptW_s^{\text{opt}} follows the empirical scaling relation:

    Wsopt≈10PW_s^{\text{opt}} \approx 10 P

    For Ego4D (P≈310P \approx 310), setting Ws=0.15W_s = 0.15 achieves peak performance, whereas for SoccerNet (P≈71P \approx 71), Ws=0.03W_s = 0.03 is optimal.

  7. Knowl 7 — Two-Stage Training Strategy of STREAMMIND

    model/method

    STREAMMIND is trained in two sequential stages using supervised autoregressive language modeling objectives:

    1. Stage 1 (Representation Alignment): The spatial video encoder (CLIP) remains frozen, while the Event-Preserving Feature Extractor (EPFE) and the full Large Language Model (LLM) are jointly trained end-to-end on streaming video-caption sequences. This aligns EPFE spatiotemporal state features with the LLM embedding space using the standard causal language modeling cross-entropy loss.
    2. Stage 2 (Cognition Gate Training): The LLM backbone and EPFE are frozen. The Cognition Gate—initialized from the first 4 layers of the LLM—is trained to predict ⟨/response⟩\langle /\text{response}\rangle or ⟨/silence⟩\langle /\text{silence}\rangle given the current perception token and user prompt, optimized via weighted cross-entropy loss to handle class imbalance.

    Training is conducted at 2 fps on 8 NVIDIA A100 GPUs for one epoch per stage, using a Cosine Annealing learning rate schedule with peak learning rates of 2×10−52\times 10^{-5} in Stage 1 and 2×10−62\times 10^{-6} in Stage 2.

  8. Knowl 8 — Trigger Accuracy and Timing Validity Metrics

    definition

    To evaluate timing alignment across continuous streaming dialogues beyond turn-level latency, two global temporal metrics are defined:

    • Trigger Accuracy (TriggerAcc): Measures the proportion of query-relevant event timestamps at which the streaming video LLM correctly initiates a linguistic response (%\%).
    • Timing Validity (TimVal): Evaluates the joint binary decision sequence over the entire continuous video stream, quantifying the model's ability to speak at valid event intervals while remaining silent during non-event frames.
  9. Knowl 9 — Streaming Video Dialogue Online Benchmark Performance

    data/table

    STREAMMIND was evaluated against prior streaming video LLMs on the real-time Ego4D and SoccerNet-Caption datasets at 2 fps under identical backbone settings.

    Dataset Method TriggerAcc ↑\uparrow TimVal ↑\uparrow BLEU-1 BLEU-4 METEOR ROUGE-L Training Cost
    Ego4D VideoLLM-Online 32.34% 29.66% 66.01 35.25 31.12 63.06 24h
    Ego4D VideoLLM-MoD 32.36% 29.65% 65.34 35.21 30.65 63.02 14h
    Ego4D STREAMMIND (Ours) 43.34% 39.73% 67.12 39.26 31.60 65.71 (11+8)h
    SoccerNet VideoLLM-Online 31.25% 28.34% 75.36 64.23 50.92 81.57 12h
    SoccerNet VideoLLM-MoD 31.24% 28.12% 74.96 64.18 50.24 81.59 7h
    SoccerNet STREAMMIND (Ours) 52.18% 47.36% 82.78 66.70 51.43 82.04 (5+3)h

    Additional online metrics on Ego4D report TimeDiff ↓\downarrow (1.89 vs 2.04 for VideoLLM-Online and 2.15 for LION-FS), Fluency ↑\uparrow (60.2% vs 45.0% and 46.1%), and Correctness ↑\uparrow (77.3% vs 48.1% and 52.4%). On SoccerNet, STREAMMIND achieves TimeDiff of 14.02 (vs 15.62), Fluency of 70.35% (vs 46.35%), and Correctness of 89.2% (vs 53.5%).

  10. Knowl 10 — Streaming Video Inference Speed and High-FPS Processing

    empirical result

    Evaluating actual inference latency for processing 1 second of streaming video across input frame rates from 5 to 100 FPS on single NVIDIA A100 and H100 GPUs demonstrates that:

    • Baseline per-step invocation architectures (VideoLLM-Online and VideoLLM-MoD) exceed the 1-second real-time processing ceiling once video frame rate surpasses 10 FPS, rendering them incapable of full-frame-rate processing on standard 24–30 FPS film or 60–100 FPS gaming inputs.
    • STREAMMIND processes 1-second video chunks in well under 1 second across all tested rates up to 100 FPS on a single A100 GPU (and faster on H100), delivering full-frame-rate real-time perception and event-gated responsive dialogue.
  11. Knowl 11 — Ablations on Cognition Gate Architecture and Feature Extractors

    data/table

    Ablation studies on the Ego4D dataset evaluated the choice of Cognition Gate architecture, initialization scheme, gate depth, and perception feature extractor.

    Ablation Category Configuration TimeDiff ↓\downarrow TriggerAcc ↑\uparrow TimVal ↑\uparrow
    Gate Architecture Linear Layer 5.06 20.13% 17.65%
    MLP Projector + Linear 4.34 21.75% 18.33%
    Transformers + Linear 3.78 21.58% 18.97%
    Cross-Attention + Linear 3.64 24.34% 20.36%
    LLM Block (Ours) 1.89 43.34% 39.73%
    Gate Initialization Random Initialization 2.10 39.65% 37.35%
    SkipBlock Initialization 1.89 40.23% 37.67%
    EarlyBlock (First layers, Ours) 1.89 43.34% 39.73%
    Gate Depth (Blocks) 2 Blocks 1.92 41.53% 37.65%
    3 Blocks 1.90 42.35% 38.34%
    4 Blocks (Ours) 1.89 43.34% 39.73%
    5 Blocks 1.88 42.67% 38.56%
    6 Blocks 1.89 42.56% 38.34%
    Perception Extractor Q-Former 3.78 26.65% 25.31%
    STC Connector 3.56 27.54% 26.87%
    EPFE (Ours) 1.89 43.34% 39.73%

    The results confirm that extracting state-space features via EPFE and initializing a 4-block gate from the earliest layers of the LLM significantly outperforms alternative projection and attention architectures.

Coverage note — None was omitted; all key contributions—including the StreamingVD formulation, event-gated paradigm, EPFE module, Cognition Gate transfer method, balancing weight empirical formula, algorithm, online/offline experimental results, speed benchmarks, and ablation studies—are represented.

References

  1. 1.Roberto Amoroso, Gengyuan Zhang, Rajat Koner, Lorenzo Baraldi, Rita Cucchiara, and Volker Tresp. Perceive, query & reason: Enhancing video qa with question-guided temporal queries. arXiv preprint arXiv:2412.19304, 2024.
  2. 2.Sanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72, 2005.
  3. 3.Pietro Battistoni, Andrea Antonio Cantone, Mariarosaria Esposito, Rita Francese, Francesca Pia Perillo, Marco Romano, Monica Sebillo, and Giuliana Vitiello. Using artificial intelligence and companion robots to improve home healthcare for the elderly. In International Conference on Human-Computer Interaction, pages 3–17. Springer, 2023.
  4. 4.Jiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong, Wei Xia, Zhuowen Tu, and Stefano Soatto. Memot: Multi-object tracking with memory. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8090–8100, 2022.
  5. 5.Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh. Honeybee: Locality-enhanced projector for multimodal llm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13817–13827, 2024.
  6. 6.Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1–45, 2024.
  7. 7.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pretraining to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3558–3568, 2021.
  8. 8.Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. Videollm-online: Online video large language model for streaming video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18407–18418, 2024.
  9. 9.Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335, 2024.
  10. 10.Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24185–24198, 2024.
  11. 11.Ho Kei Cheng and Alexander G Schwing. Xmem: Longterm video object segmentation with an atkinson-shiffrin memory model. In European Conference on Computer Vision, pages 640–658. Springer, 2022.
  12. 12.Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. Videollama 2: Advancing spatialtemporal modeling and audio understanding in video-llms. arXiv preprint arXiv:2406.07476, 2024.
  13. 13.Shangzhe Di and Weidi Xie. Grounded question-answering in long egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12934–12943, 2024.
  14. 14.Shangzhe Di, Zhelun Yu, Guanghao Zhang, Haoyuan Li, Hao Cheng, Bolin Li, Wanggui He, Fangxun Shu, Hao Jiang, et al. Streaming video question-answering with in-context video kv-cache retrieval. In The Thirteenth International Conference on Learning Representations.
  15. 15.Yue Fan, Xiaojian Ma, Rujie Wu, Yuntao Du, Jiaqi Li, Zhi Gao, and Qing Li. Videoagent: A memory-augmented multimodal agent for video understanding. In European Conference on Computer Vision, pages 75–92. Springer, 2024.
  16. 16.Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training. arXiv preprint arXiv:2405.20978, 2024.
  17. 17.Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. Mmbench-video: A long-form multi-shot benchmark for holistic video understanding. Advances in Neural Information Processing Systems, 37:89098–89124, 2025.
  18. 18.Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19358–19369, 2023.
  19. 19.Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024.
  20. 20.Silvio Giancola, Mohieddine Amine, Tarek Dghaily, and Bernard Ghanem. Soccernet: A scalable dataset for action spotting in soccer videos. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 1711–1721, 2018.
  21. 21.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 18995–19012, 2022.
  22. 22.Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. CoRR, abs/2312.00752, 2023.
  23. 23.Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13504–13514, 2024.
  24. 24.Daoji Huang, Otmar Hilliges, Luc Van Gool, and Xi Wang. Palm: Predicting actions through language models@ ego4d long-term action anticipation challenge 2023. arXiv preprint arXiv:2306.16545, 2023.
  25. 25.De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, and Jan Kautz. Lita: Language instructed temporal-localization assistant. In European Conference on Computer Vision, pages 202–218. Springer, 2024.
  26. 26.Zhenpeng Huang, Xinhao Li, Jiaqi Li, Jing Wang, Xiangyu Zeng, Cheng Liang, Tao Wu, Xi Chen, Liang Li, and Limin Wang. Online video understanding: A comprehensive benchmark and memory-augmented method. arXiv preprint arXiv:2501.00584, 2024.
  27. 27.Yang Jin, Zhicheng Sun, Kun Xu, Liwei Chen, Hao Jiang, Quzhe Huang, Chengru Song, Yuliang Liu, Di Zhang, Yang Song, et al. Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization. arXiv preprint arXiv:2402.03161, 2024.
  28. 28.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023.
  29. 29.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023.
  30. 30.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024.
  31. 31.Wei Li, Bing Hu, Rui Shao, Leyang Shen, and Liqiang Nie. Lion-fs: Fast & slow video-language thinker as online video assistant. arXiv preprint arXiv:2503.03663, 2025.
  32. 32.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision, pages 323–340. Springer, 2024.
  33. 33.Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023.
  34. 34.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004.
  35. 35.Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26689–26699, 2024.
  36. 36.Kevin Lin, Faisal Ahmed, Linjie Li, Chung-Ching Lin, Ehsan Azarnasab, Zhengyuan Yang, Jianfeng Wang, Lin Liang, Zicheng Liu, Yumao Lu, et al. Mm-vid: Advancing video understanding with gpt-4v (ision). arXiv preprint arXiv:2310.19773, 2023.
  37. 37.Xingyu Liu, Joon-Young Lee, and Hailin Jin. Learning video representations from correspondence proposals. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, pages 4273–4281, 2019.
  38. 38.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634, 2023.
  39. 39.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Da Li, Pengcheng Lu, Tao Wang, Linmei Hu, Minghui Qiu, and Zhongyu Wei. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207, 2023.
  40. 40.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023.
  41. 41.Kelly Merrill Jr, Jihyun Kim, and Chad Collins. Ai companions for lonely individuals and the role of social presence. Communication Research Reports, 39(2):93–103, 2022.
  42. 42.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002.
  43. 43.Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, and Vittorio Ferrari. Connecting vision and language with localized narratives. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16, pages 647–664. Springer, 2020.
  44. 44.Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models. Advances in Neural Information Processing Systems, 37:119336–119360, 2025.
  45. 45.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021.
  46. 46.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Doll'ar. Designing network design spaces. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10428–10436, 2020.
  47. 47.Jiayuan Rao, Haoning Wu, Chang Liu, Yanfeng Wang, and Weidi Xie. Matchtime: Towards automatic soccer game commentary generation. arXiv preprint arXiv:2406.18530, 2024.
  48. 48.Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. Timechat: A time-sensitive multimodal large language model for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14313–14323, 2024.
  49. 49.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18221–18232, 2024.
  50. 50.Larry R Squire, Lisa Genzel, John T Wixted, and Richard G Morris. Memory consolidation. Cold Spring Harbor perspectives in biology, 7(8):a021766, 2015.
  51. 51.Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. Finegrained audio-visual joint representations for multimodal large language models. arXiv preprint arXiv:2310.05863, 2023.
  52. 52.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216, 2019.
  53. 53.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  54. 54.Ales Vysocky and Petr Novak. Human-robot collaboration in industry. MM Science Journal, 9(2):903–906, 2016.
  55. 55.Junke Wang, Dongdong Chen, Chong Luo, Xiyang Dai, Lu Yuan, Zuxuan Wu, and Yu-Gang Jiang. Chatvideo: A tracklet-centric multimodal and versatile video understanding system. arXiv preprint arXiv:2304.14407, 2023.
  56. 56.Qing Wang, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. Matchformer: Interleaving attention in transformers for feature matching. In Proceedings of the Asian Conference on Computer Vision, pages 2746–2762, 2022.
  57. 57.Xunguang Wang, Zhenlan Ji, Pingchuan Ma, Zongjie Li, and Shuai Wang. Instructta: Instruction-tuned targeted attack for large vision-language models. arXiv preprint arXiv:2312.01886, 2023.
  58. 58.Yuxuan Wang, Yueqian Wang, Pengfei Wu, Jianxin Liang, Dongyan Zhao, and Zilong Zheng. Lstp: Language-guided spatial-temporal prompt learning for long-form video-text understanding. arXiv e-prints, pages arXiv–2402, 2024.
  59. 59.Shiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang, Yan Gao, Qianli Xu, Tong Xu, Yao Hu, Enhong Chen, and Mike Zheng Shou. Videollm-mod: Efficient videolanguage streaming with mixture-of-depths vision computation. Advances in Neural Information Processing Systems, 37:109922–109947, 2025.
  60. 60.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786, 2021.
  61. 61.Haomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge, Lu Zhang, Jiawen Zhu, and Huchuan Lu. Streaming video understanding and multi-round interaction with memoryenhanced knowledge. arXiv preprint arXiv:2501.13468, 2025.
  62. 62.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653, 2017.
  63. 63.Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994, 2024.
  64. 64.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1686–1697, 2021.
  65. 65.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. Advances in Neural Information Processing Systems, 35:124–141, 2022.
  66. 66.Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10714–10726, 2023.
  67. 67.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9127–9134, 2019.
  68. 68.Jeffrey M Zacks, Nicole K Speer, Khena M Swallow, Todd S Braver, and Jeremy R Reynolds. Event perception: a mindbrain perspective. Psychological bulletin, 133(2):273, 2007.
  69. 69.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11975–11986, 2023.
  70. 70.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023.
  71. 71.Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, and Xiaojie Jin. Flash-vstream: Memorybased real-time understanding for long video streams. arXiv preprint arXiv:2406.08085, 2024.
  72. 72.Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024.
  73. 73.Ruohong Zhang, Liangke Gui, Zhiqing Sun, Yihao Feng, Keyang Xu, Yuanhan Zhang, Di Fu, Chunyuan Li, Alexander Hauptmann, Yonatan Bisk, et al. Direct preference optimization of video large multimodal models from language model reward. arXiv preprint arXiv:2404.01258, 2024.
  74. 74.Ruoyu Zhang, Lulu Wang, Yi He, Tongling Pan, Zhengtao Yu, and Yingna Li. Tpcap: Unlocking zero-shot image captioning with trigger-augmented and multi-modal purification modules. arXiv preprint arXiv:2502.11024, 2025.
  75. 75.Yue Zhao, Ishan Misra, Philipp Kr¨ahenb¨uhl, and Rohit Girdhar. Learning video representations from large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6586–6597, 2023.
  76. 76.Xingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan, Austin Myers, Xuehan Xiong, Arsha Nagrani, and Cordelia Schmid. Streaming dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18243–18252, 2024.

Citation

MLA
Ding, X., et al. “StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue Through Event-Gated Cognition”. arXiv, 2025, http://arxiv.org/abs/2503.06220v3.
APA
Ding, X., Wu, H., Yang, Y., Jiang, S., Bai, D., Chen, Z., & Cao, T. (2025). StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition. arXiv. http://arxiv.org/abs/2503.06220v3
Chicago
Ding, X., H. Wu, Y. Yang, et al. 2025. “StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue Through Event-Gated Cognition”. arXiv. http://arxiv.org/abs/2503.06220v3.
Harvard
Ding, X. et al. (2025) “StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.06220v3.
Vancouver
1. Ding X, Wu H, Yang Y, Jiang S, Bai D, Chen Z, Cao T (2025) StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition. arXiv

BibTeX

@article{ding2025streammind,
  title = {StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition},
  author = {Ding, Xin and Wu, Hao and Yang, Yifan and Jiang, Shiqi and Bai, Donglin and Chen, Zhibo and Cao, Ting},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.06220v3},
  eprint = {2503.06220}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/