OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Junbo NiuYifei LiZiyang MiaoChunjiang GeYuanhang ZhouQihao HeXiaoyi DongHaodong DuanShuangrui DingRui Qian

article2025CVPR151 citations

Introduces OVO-Bench, a fine-grained evaluation benchmark that measures how effectively video large language models process dynamic video streams across backward tracing, real-time perception, and forward active responding scenarios.

Listen

Artificial intelligence systems designed for real-world interactive applications, such as autonomous driving and robotic assistants, must continuously process live video streams and respond to queries at specific points in time. Most current video models and benchmarks, however, operate in an offline setting where complete videos are available in advance for static analysis. This creates a critical disconnect between standard evaluation metrics and the actual requirements of dynamic, real-time video understanding.

The article introduces OVO-Bench, a comprehensive benchmark designed to evaluate how effectively video language models handle time-sensitive reasoning across streaming video. It establishes an evaluation taxonomy based on three core operational modes: tracing back to past events, perceiving ongoing real-time activities, and actively delaying responses until sufficient future information becomes available.

To construct the benchmark, the researchers combined existing video datasets and web-collected sources into a diverse pool of 644 videos across seven domains, spanning lengths from several minutes to half an hour. They developed 2,814 fine-grained meta-annotations with precise event timestamps using a hybrid semi-automated generation and human-curation process. These annotations support 12 distinct evaluation tasks. The benchmark was used to assess 11 leading video language models, spanning proprietary systems, open-source offline architectures, and specialized streaming models, alongside human baselines.

The evaluation revealed several critical findings. First, all evaluated artificial intelligence systems fall substantially short of human capability: human agents achieved an overall average score of 92.81%, whereas the top-performing model, Gemini 1.5 Pro, reached only 63.00%. Second, offline proprietary models transferred better to real-time perception than dedicated online streaming models, which achieved overall scores between 33.61% and 41.78%. Third, current models suffer heavily from hallucinations and lack temporal prioritization, meaning they struggle to distinguish current events from similar past or future occurrences. Finally, inference latency remains a severe bottleneck, with standard models requiring approximately 4 seconds to process 64-frame inputs, rendering live interaction impractical under existing hardware and software configurations.

These findings indicate that current video language models are not yet dependable for high-stakes, time-critical deployments. Deploying existing models in live-assistant or autonomous settings poses notable operational and safety risks due to frequent hallucinations, poor temporal localization, and processing latency. While offline models demonstrate stronger visual reasoning, their architectural demands make real-time streaming interaction difficult without major latency trade-offs.

Organizations developing video assistants and interactive vision systems should avoid relying purely on offline benchmark scores to predict real-world performance. Technical teams should focus development efforts on creating efficient online architectures that improve streaming inference speeds while maintaining the deep reasoning capabilities of offline models. Additionally, training pipelines must explicitly incorporate temporal prioritization and active responding strategies so models learn when to withhold answers until adequate visual evidence appears.

The benchmark relies on simulated streaming conditions for offline models via video segmenting and uses multiple-choice formats for several core tasks. Although these methodologies provide a standardized and rigorous testbed, future evaluations should expand to fully continuous, open-ended live dialogues as real-time streaming architectures mature.

Cover for OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Abstract

Temporal Awareness—the ability to reason dynamically based on the timestamp when a question is raised—is the key distinction between offline and online video LLMs. Unlike offline models, which rely on complete videos for static, post hoc analysis, online models process video streams incrementally and dynamically adapt their responses based on the timestamp at which the question is posed. Despite its significance, temporal awareness has not been adequately evaluated in existing benchmarks. To fill this gap, we present OVO-Bench (Online-VideO-Benchmark), a novel video benchmark that emphasizes the importance of timestamps for advanced online video understanding capability benchmarking. OVO-Bench evaluates the ability of video LLMs to reason and respond to events occurring at specific timestamps under three distinct scenarios: (1) Backward tracing: trace back to past events to answer the question. (2) Real-time understanding: understand and respond to events as they unfold at the current timestamp.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 OVO-Bench
  • 3.1 Online Video Understanding Mode Taxonomy
  • 3.1.1 Backward Tracing
  • 3.1.2 Real-Time Visual Perception
  • 3.1.3 Forward Active Responding
  • 3.2 Benchmark Construction
  • 3.2.1 Video and Annotation Collection
  • 3.2.2 Prompt Generation
  • 3.3 Datasets Statistics
  • 4 Experiments
  • 4.1 Models and Evaluation Strategies
  • 4.2 Main Results
  • 4.3 Comparison between online Video-LLMs and offline Video-LLMs
  • 4.4 Forward Active Responding
  • 5 Conclusion and Future Work
  • 6 More Details of Evaluation
  • 6.1 Evaluation for Online Models on Forward Active Responding
  • 6.2 Prompt Design for Offline Models on Forward Active Responding
  • 6.3 Prompt Design for Models on Backward Tracing and Real-Time Visual Perception
  • 7 More Details of Benchmark Construction
  • 7.1 Human-annotated QA Generation
  • 8 Additional Dataset Analysis
  • 8.1 Task and Sample Distribution
  • 8.2 Query Timestamps and Video Duration
  • 9 Limitations
  • 10 Licenses
  • 11 Data Examples
  • References

Knowls

  1. Knowl 1 — Online Video Understanding Formulation Across Three Temporal Modes

    definition

    Online video understanding models real-world, always-on agents that continuously process streaming video inputs and dynamically adapt their responses depending on the exact timestamp at which a query is raised. Given a user text query Qt0Q_{t_0} posed at timestamp t0t_0 and a continuous streaming video input X(−∞,+∞)X_{(-\infty, +\infty)}, the reasoning process is formulated across three distinct temporal modes:

    1. Backward Tracing: Answering queries about historical events that occurred prior to the current window: Rt0=P(Qt0,X(−∞,−T])R_{t_0} = P(Q_{t_0}, X_{(-\infty, -T]}) where T>0T > 0 denotes a temporal threshold defining the boundary of recent/ongoing events, and Rt0R_{t_0} denotes the model response at t0t_0.

    2. Real-Time Visual Perception: Answering queries about ongoing events unfolding at the immediate current timestamp: Rt0=P(Qt0,X(−T,t0])R_{t_0} = P(Q_{t_0}, X_{(-T, t_0]})

    3. Forward Active Responding: Processing queries posed at t0t_0 where the necessary information has not yet appeared, requiring the model to withhold answering until sufficient visual evidence accumulates in subsequent stream frames: R(t0,+∞)=P(Qt0,X(t0,+∞))R_{(t_0, +\infty)} = P(Q_{t_0}, X_{(t_0, +\infty)}).

  2. Knowl 2 — OVO-Bench Task Taxonomy Across Online Perception Modes

    definition

    OVO-Bench structures online video understanding into 12 core tasks organized across three problem-solving modes:

    • Backward Tracing:

      • Episodic Memory (EPM): Backtracking and retrieving specific key moments from past video frames.
      • Action Sequence Identification (ASI): Identifying the correct chronological order of preceding human actions.
      • Hallucination Detection (HLD): Detecting and rejecting queries regarding events or entities that never occurred in previous video frames.
    • Real-Time Visual Perception:

      • Spatial Understanding (STU): Reasoning about spatial relationships between objects present in nearby frames.
      • Object Recognition (OJR): Identifying objects visible in current frames.
      • Attribute Recognition (ATR): Recognizing entity properties (such as color, texture, and size) in current frames.
      • Action Recognition (ACR): Identifying ongoing actions performed by agents in the current frame.
      • Optical Character Recognition (OCR): Reading text and characters displayed within current frames.
      • Future Prediction (FPD): Forecasting immediate subsequent scene transitions, object state modifications, or upcoming actions.
    • Forward Active Responding:

      • Repetition Event Count (REC): Incrementally tracking and counting repeated short-term or long-term action cycles over time.
      • Sequential Steps Recognition (SSR): Detecting when an ongoing multi-step procedure transitions from one phase to the next.
      • Clues Reveal Responding (CRR): Delaying response generation until sufficient informative clues emerge in future video inputs.
  3. Knowl 3 — OVO-Bench Dataset Composition and Annotation Protocol

    experimental setup

    OVO-Bench contains 644 unique videos spanning 7 domains (including Sports, Video Games, Egocentric, and Tutorials) with video durations ranging from a few minutes to 30 minutes, with a mean query timepoint of 428.89 seconds. The benchmark comprises 2,814 fine-grained question-answer (QA) pairs equipped with precise event timestamps.

    The benchmark construction follows a three-stage pipeline:

    1. Video Curation: Clips are sourced from human-annotated video benchmarks—QA-Ego4D and OpenEQA for EPM; STAR, YouCook2, CrossTask, HiREST, and COIN for ASI; Perception-Test and Thumos for REC; COIN for SSR; MovieNet for CRR; and Ego4D for Real-Time Visual Perception—supplemented by diverse YouTube web-crawled videos. All source clips are drawn from validation or test splits.
    2. Meta-Annotation: Existing datasets with precise event-level timestamps are repurposed directly. For datasets with video-level QA lacking precise temporal localization, Gemini-1.5-Pro is prompted to produce coarse candidate timestamps, followed by manual human verification and boundary adjustment.
    3. Visually Grounded Distractor Generation: For multiple-choice questions (ranging from 2 to 5 options), plausible distractors are generated by prompting Video-LLMs with original QA pairs and video clips to inject grounded visual misinformation from the footage, followed by human curation and random option shuffling to prevent positional and linguistic bias.
  4. Knowl 4 — Multiple-Triggering Evaluation Pipeline for Forward Active Responding

    model/method

    To evaluate offline Video-LLMs in forward active responding settings requiring continuous adaptation, OVO-Bench employs a multiple-triggering dense query protocol:

    1. Temporal Query Slicing: Given a query Qt0Q_{t_0} posed at time t0t_0, the model is queried at discrete timestamps {t1,t2,…,tN}\{t_1, t_2, \dots, t_N\} (t0<t1<⋯<tNt_0 < t_1 < \dots < t_N) along the video timeline surrounding target clue events.
    2. Visual Truncation: At each step tkt_k, the visual stream is truncated to X[0,tk]X_{[0, t_k]} to simulate an online feed up to that point.
    3. Independent Triggering Decision: The model independently evaluates whether the visual evidence accumulated up to tkt_k provides sufficient information to answer Qt0Q_{t_0} or if it must continue waiting.
    4. Evaluation Scoring:
      • In Clues Reveal Responding (CRR) and Sequential Steps Recognition (SSR), the scoring metric rewards early accurate identification once clue frames appear while penalizing premature responses (hallucinations) triggered before clue arrival.
      • In Repetition Event Count (REC), the scoring metric evaluates incremental count accuracy as successive action iterations occur across the temporal sequence.
  5. Knowl 5 — Comparative Evaluation of Multimodal Models on OVO-Bench

    data/table

    The performance of eleven Multimodal Large Language Models (MLLMs), a blind LLM baseline, and human evaluators on OVO-Bench across Real-Time Visual Perception (RT), Backward Tracing (BT), and Forward Active Responding (FAR) tasks is summarized below (all values are accuracy percentages):

    Model # Frames Real-Time Visual Perception Backward Tracing Forward Active Responding Overall
    OCR ACR ATR STU FPD OJR Avg EPM ASI HLD Avg REC SSR CRR Avg Avg
    Human Agents - 93.96 92.57 94.83 92.70 91.09 94.02 93.20 92.59 93.02 91.37 92.33 95.48 89.67 93.56 92.90 92.81
    Blind LLMs
    GPT-4-turbo - 28.86 24.77 25.67 33.76 27.72 26.63 27.90 42.76 48.65 70.05 53.82 - - 52.92 - -
    Proprietary Offline Multimodal Models
    Gemini 1.5 Pro 1 fps 85.91 66.97 79.31 58.43 63.37 61.96 69.32 58.59 76.35 52.64 62.54 35.53 74.24 61.67 57.15 63.00
    GPT-4o 64 69.80 64.22 71.55 51.12 70.30 59.78 64.46 57.91 75.68 48.66 60.75 27.58 73.21 59.40 53.40 59.54
    Open-source Offline Multimodal Models
    Qwen2-VL-72B 64 65.77 60.55 69.83 51.69 69.31 54.35 61.92 52.53 60.81 57.53 56.95 38.83 64.07 45.00 49.30 56.27
    LLaVA-Video-7B 64 69.13 58.72 68.83 49.44 74.26 59.78 63.52 56.23 57.43 7.53 40.40 34.10 69.95 60.42 54.82 52.91
    LLaVA-OneVision-7B 64 66.44 57.80 73.28 53.37 71.29 61.96 64.02 54.21 55.41 21.51 43.71 25.64 67.09 58.75 50.50 52.74
    Qwen2-VL-7B 64 60.40 50.46 56.03 47.19 66.34 55.43 55.98 47.81 35.48 56.08 46.46 31.66 65.82 48.75 48.74 50.39
    InternVL-V2-8B 64 67.11 60.55 63.79 46.07 68.32 56.52 60.39 48.15 57.43 24.73 43.44 26.50 59.14 54.14 46.60 50.15
    LongVU-7B 1 fps 53.69 53.21 62.93 47.75 68.32 59.78 57.61 40.74 59.46 4.84 35.01 12.18 69.48 60.83 47.50 46.71
    Open-source Online Multimodal Models
    Flash-VStream-7B 1 fps 24.16 29.36 28.45 33.71 25.74 28.80 28.37 39.06 37.16 5.91 27.38 8.02 67.25 60.00 45.09 33.61
    VideoLLM-online-8B 2 fps 8.05 23.85 12.07 14.04 45.54 21.20 20.79 22.22 18.80 12.18 17.73 - - - - -
    Dispider 1 fps 57.72 49.54 62.07 44.94 61.39 51.63 54.55 48.48 55.41 4.30 36.06 18.05 37.36 48.75 34.72 41.78

    Proprietary offline models achieve the highest scores, with Gemini 1.5 Pro attaining 63.00% and GPT-4o achieving 59.54% overall average accuracy. Open-source offline models range between 46.71% and 56.27%, while dedicated online streaming architectures show lower overall performance (33.61% for Flash-VStream-7B and 41.78% for Dispider). All evaluated models exhibit a substantial performance gap compared to human accuracy (92.81%).

  6. Knowl 6 — Performance Gap Between Offline and Online Streaming Video-LLMs

    empirical result

    When evaluated in simulated streaming setups via temporal clip truncation Video[0:ti]Video[0:t_i], state-of-the-art offline Video-LLMs substantially outperform specialized online streaming models across all evaluation dimensions:

    • On Real-Time Visual Perception, Gemini 1.5 Pro reaches 69.32% average accuracy and GPT-4o reaches 64.46%, compared to 28.37% for Flash-VStream-7B, 20.79% for VideoLLM-online-8B, and 54.55% for Dispider.
    • On Backward Tracing, Gemini 1.5 Pro achieves 62.54% and Qwen2-VL-72B reaches 56.95%, compared to 27.38% for Flash-VStream-7B and 36.06% for Dispider.
    • Overall benchmark averages for offline models (Gemini 1.5 Pro at 63.00%, GPT-4o at 59.54%, Qwen2-VL-72B at 56.27%) consistently exceed streaming architectures (Dispider at 41.78%, Flash-VStream-7B at 33.61%).

    This indicates that the global context processing and representation capacity of offline models transfer effectively to simulated streaming inputs, whereas existing streaming memory compression and frame-selection methods suffer from loss of fine-grained visual information.

  7. Knowl 7 — Temporal Prioritization Failure and Hallucination Susceptibility in Video-LLMs

    empirical result

    Evaluation across the tasks of OVO-Bench demonstrates two critical limitations in current Video-LLMs:

    1. Lack of Temporal Prioritization: Models struggle to isolate the target query timeframe when multiple distractor scenes matching query entities appear earlier in the stream. On Spatial Understanding (STU) and Action Recognition (ACR), the top model (Gemini 1.5 Pro) achieves only 58.43% and 66.97% accuracy, respectively, compared to 92.70% and 92.57% for human annotators.
    2. Severe Hallucinations (HLD): Open-source and streaming models frequently fail negative rejection when asked about absent events, scoring very low on Hallucination Detection (e.g., Dispider at 4.30%, LongVU-7B at 4.84%, Flash-VStream-7B at 5.91%, and LLaVA-Video-7B at 7.53%). Although Gemini 1.5 Pro achieves the highest model score at 52.64%, it still lags human performance (91.37%) by 38.73 percentage points.
  8. Knowl 8 — Inference Latency Scaling Bottleneck in Streaming Video-LLMs

    limitation

    Inference latencies of evaluated Video-LLMs exhibit exponential growth with increasing frame counts. When processing 64 visual frames, efficient models—such as Qwen2-VL-7B (tested on four NVIDIA A100 GPUs) and Flash-VStream (tested on a single NVIDIA A100 GPU)—require approximately 4 seconds on average to return a single response. This latency creates a major barrier for real-time conversational streaming dialogue, demonstrating that current architectures cannot support simultaneous high frame capacity and real-time responsiveness without further inference efficiency innovations.

Coverage note — No substantial contributed material was omitted; all primary taxonomy definitions, benchmark construction steps, evaluation protocols, experimental data, and analytical findings are covered.

References

  1. 1.Leonard Barmann and Alex Waibel. Where did i leave my keys?-episodic-memory-based question answering on ego-centric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1560–1568, 2022.
  2. 2.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015.
  3. 3.Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao, Yong Jae Lee, and Jianwei Yang. Temporalbench: Bench-marking fine-grained temporal understanding for multimodal video models, 2024.
  4. 4.Keshigeyan Chandrasegaran, Agrim Gupta, Lea M Hadzic, Taran Kota, Jimming He, Cristobal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. Hourvideo: 1-hour video-language understanding. arXiv preprint arXiv:2411.04998, 2024.
  5. 5.Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. Videollm-online: Online video large language model for streaming video. In CVPR, 2024.
  6. 6.Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al. Are we on the right way for evaluating large vision-language models? arXiv preprint arXiv:2403.20330, 2024.
  7. 7.Xiuyuan Chen, Yuan Lin, Yuchen Zhang, and Weiran Huang. Autoeval-video: An automatic benchmark for assessing large vision language models in open-ended video question answering, 2024.
  8. 8.Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv preprint arXiv:2312.14238, 2023.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  10. 10.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500, 2023.
  11. 11.Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. Mmbench-video: A long-form multi-shot benchmark for holistic video under-standing. arXiv preprint arXiv:2406.14515, 2024.
  12. 12.Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever compre-hensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024.
  13. 13.Alex Gorban, Haroon Idrees, Yu-Gang Jiang, A Roshan Za-mir, Ivan Laptev, Mubarak Shah, and Rahul Sukthankar. Thumos challenge: Action recognition with a large number of classes, 2015.
  14. 14.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18995–19012, 2022.
  15. 15.Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding supplementary material.
  16. 16.Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. Movienet: A holistic dataset for movie under-standing. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16, pages 709–727. Springer, 2020.
  17. 17.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel-man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli-hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  18. 18.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE con-ference on computer vision and pattern recognition, pages 2758–2766, 2017.
  19. 19.Yu-Gang Jiang, Jingen Liu, A Roshan Zamir, George Toderici, Ivan Laptev, Mubarak Shah, and Rahul Sukthankar. Thumos challenge: Action recognition with a large number of classes, 2014.
  20. 20.Peng Jin, Ryuichi Takanobu, Caiwan Zhang, Xiaochun Cao, and Li Yuan. Chat-univi: Unified visual representation em-powers large language models with image and video under-standing. arXiv preprint arXiv:2311.08046, 2023.
  21. 21.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024.
  22. 22.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023.
  23. 23.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understand-ing benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024.
  24. 24.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. 2024.
  25. 25.Junming Lin, Zheng Fang, Chi Chen, Zihao Wan, Fuwen Luo, Peng Li, Yang Liu, and Maosong Sun. Streamingbench: Assessing the gap for mllms to achieve streaming video under-standing. arXiv preprint arXiv:2411.03628, 2024.
  26. 26.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023.
  27. 27.Ruyang Liu, Chen Li, Haoran Tang, Yixiao Ge, Ying Shan, and Ge Li. St-llm: Large language models are effective temporal learners. In European Conference on Computer Vision, pages 1–18. Springer, 2025.
  28. 28.Ye Liu, Zongyang Ma, Zhongang Qi, Yang Wu, Chang Wen Chen, and Ying Shan. E.t. bench: Towards open-ended event-level video-language understanding. In Neural Information Processing Systems (NeurIPS), 2024.
  29. 29.Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, et al. Openeqa: Embodied question answering in the era of foundation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16488–16498, 2024.
  30. 30.Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, et al. Openeqa: Embodied question answering in the era of foundation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16488–16498, 2024.
  31. 31.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding, 2023.
  32. 32.Salman Khan Muhammad Maaz, Hanoona Rasheed and Fahad Khan. Video-chatgpt: Towards detailed video un-derstanding via large vision and language models. ArXiv 2306.05424, 2023.
  33. 33.OpenAI. Gpt-4 technical report, 2023. Technical report.
  34. 34.OpenAI. Hello gpt-4o. https://openai.com/index/hello-gpt-4o, 2024.
  35. 35.Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Mateusz Malinowski, Yi Yang, Carl Doersch, et al. Perception test: A diagnostic benchmark for multimodal video models. Advances in Neural Information Processing Systems, 36, 2024.
  36. 36.Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models. Advances in Neural Information Processing Systems, 37:119336–119360, 2024.
  37. 37.Rui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Dahua Lin, and Jiaqi Wang. Dispider: Enabling video llms with active real-time interaction via disentangled perception, decision, and reaction. arXiv preprint arXiv:2501.03218, 2025.
  38. 38.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  39. 39.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021.
  40. 40.Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, Zhuang Liu, Hu Xu, Hyunwoo J. Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra. Longvu: Spatiotemporal adaptive compression for long video-language understanding. arXiv preprint arXiv:2410.17434, 2024.
  41. 41.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Xun Guo, Tian Ye, Yan Lu, Jenq-Neng Hwang, and Gaoang Wang. Moviechat: From dense token to sparse memory for long video understanding, 2023.
  42. 42.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216, 2019.
  43. 43.Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  44. 44.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  45. 45.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
  46. 46.Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, et al. Lvbench: An extreme long video understanding benchmark. arXiv preprint arXiv:2406.08035, 2024.
  47. 47.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023.
  48. 48.Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan. Star: A benchmark for situated reasoning in real-world videos. arXiv preprint arXiv:2405.09711, 2024.
  49. 49.Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long-context interleaved video-language understanding, 2024.
  50. 50.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786, 2021.
  51. 51.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653, 2017.
  52. 52.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016.
  53. 53.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9127–9134, 2019.
  54. 54.Abhay Zala, Jaemin Cho, Satwik Kottur, Xilun Chen, Barlas Oguz, Yashar Mehdad, and Mohit Bansal. Hierarchical video-moment retrieval and step-captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 23056–23065, 2023.
  55. 55.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023.
  56. 56.Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, and Xiaojie Jin. Flash-vstream: Memory-based real-time understanding for long video streams, 2024.
  57. 57.Jiacheng Zhang, Yang Jiao, Shaoxiang Chen, Jingjing Chen, and Yu-Gang Jiang. Eventhallusion: Diagnosing event hallucinations in video llms. arXiv preprint arXiv:2409.16597, 2024.
  58. 58.Pan Zhang, Xiaoyi Dong, Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Wenwei Zhang, Hang Yan, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition, 2023.
  59. 59.Pan Zhang, Xiaoyi Dong, Yuhang Cao, Yuhang Zang, Rui Qian, Xilin Wei, Lin Chen, Yifei Li, Junbo Niu, Shuangrui Ding, et al. Internlm-xcomposer2. 5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions. arXiv preprint arXiv:2412.09596, 2024.
  60. 60.Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data, 2024.
  61. 61.Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. Mlvu: A comprehensive benchmark for multi-task long video understanding. arXiv preprint arXiv:2406.04264, 2024.
  62. 62.Luowei Zhou, Chenliang Xu, and Jason Corso. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018.
  63. 63.Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Cross-task weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3537–3545, 2019.

Citation

MLA
Niu, J., et al. “OVO-Bench: How Far Is Your Video-LLMs from Real-World Online Video Understanding?”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 18902–13, https://doi.org/10.1109/CVPR52734.2025.01761.
APA
Niu, J., Li, Y., Miao, Z., Ge, C., Zhou, Y., He, Q., Dong, X., Duan, H., Ding, S., Qian, R., Zhang, P., Zang, Y., Cao, Y., He, C., & Wang, J. (2025). OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18902–18913. https://doi.org/10.1109/CVPR52734.2025.01761
Chicago
Niu, J., Y. Li, Z. Miao, et al. 2025. “OVO-Bench: How Far Is Your Video-LLMs from Real-World Online Video Understanding?”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18902–13. https://doi.org/10.1109/CVPR52734.2025.01761.
Harvard
Niu, J. et al. (2025) “OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 18902–18913. Available at: https://doi.org/10.1109/CVPR52734.2025.01761.
Vancouver
1. Niu J, Li Y, Miao Z, et al (2025) OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 18902–18913

BibTeX

@inproceedings{Niu_2025, title={OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?}, url={http://dx.doi.org/10.1109/CVPR52734.2025.01761}, DOI={10.1109/cvpr52734.2025.01761}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Niu, Junbo and Li, Yifei and Miao, Ziyang and Ge, Chunjiang and Zhou, Yuanhang and He, Qihao and Dong, Xiaoyi and Duan, Haodong and Ding, Shuangrui and Qian, Rui and Zhang, Pan and Zang, Yuhang and Cao, Yuhang and He, Conghui and Wang, Jiaqi}, year={2025}, month=June, pages={18902–18913} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE