OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
Junbo NiuYifei LiZiyang MiaoChunjiang GeYuanhang ZhouQihao HeXiaoyi DongHaodong DuanShuangrui DingRui Qian
Introduces OVO-Bench, a fine-grained evaluation benchmark that measures how effectively video large language models process dynamic video streams across backward tracing, real-time perception, and forward active responding scenarios.
Artificial intelligence systems designed for real-world interactive applications, such as autonomous driving and robotic assistants, must continuously process live video streams and respond to queries at specific points in time. Most current video models and benchmarks, however, operate in an offline setting where complete videos are available in advance for static analysis. This creates a critical disconnect between standard evaluation metrics and the actual requirements of dynamic, real-time video understanding.
The article introduces OVO-Bench, a comprehensive benchmark designed to evaluate how effectively video language models handle time-sensitive reasoning across streaming video. It establishes an evaluation taxonomy based on three core operational modes: tracing back to past events, perceiving ongoing real-time activities, and actively delaying responses until sufficient future information becomes available.
To construct the benchmark, the researchers combined existing video datasets and web-collected sources into a diverse pool of 644 videos across seven domains, spanning lengths from several minutes to half an hour. They developed 2,814 fine-grained meta-annotations with precise event timestamps using a hybrid semi-automated generation and human-curation process. These annotations support 12 distinct evaluation tasks. The benchmark was used to assess 11 leading video language models, spanning proprietary systems, open-source offline architectures, and specialized streaming models, alongside human baselines.
The evaluation revealed several critical findings. First, all evaluated artificial intelligence systems fall substantially short of human capability: human agents achieved an overall average score of 92.81%, whereas the top-performing model, Gemini 1.5 Pro, reached only 63.00%. Second, offline proprietary models transferred better to real-time perception than dedicated online streaming models, which achieved overall scores between 33.61% and 41.78%. Third, current models suffer heavily from hallucinations and lack temporal prioritization, meaning they struggle to distinguish current events from similar past or future occurrences. Finally, inference latency remains a severe bottleneck, with standard models requiring approximately 4 seconds to process 64-frame inputs, rendering live interaction impractical under existing hardware and software configurations.
These findings indicate that current video language models are not yet dependable for high-stakes, time-critical deployments. Deploying existing models in live-assistant or autonomous settings poses notable operational and safety risks due to frequent hallucinations, poor temporal localization, and processing latency. While offline models demonstrate stronger visual reasoning, their architectural demands make real-time streaming interaction difficult without major latency trade-offs.
Organizations developing video assistants and interactive vision systems should avoid relying purely on offline benchmark scores to predict real-world performance. Technical teams should focus development efforts on creating efficient online architectures that improve streaming inference speeds while maintaining the deep reasoning capabilities of offline models. Additionally, training pipelines must explicitly incorporate temporal prioritization and active responding strategies so models learn when to withhold answers until adequate visual evidence appears.
The benchmark relies on simulated streaming conditions for offline models via video segmenting and uses multiple-choice formats for several core tasks. Although these methodologies provide a standardized and rigorous testbed, future evaluations should expand to fully continuous, open-ended live dialogues as real-time streaming architectures mature.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Video-MME establishes foundational protocols and benchmark standards for evaluating multimodal LLMs across diverse video durations, providing the offline baseline from which OVO-Bench shifts toward online streaming evaluation.
- Paper: LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding, Haoning Wu et al. (2024). LongVideoBench formalizes fine-grained temporal retrieval and reasoning over extended multimodal video contexts, defining the long-form evaluation paradigms adapted by OVO-Bench.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). HourVideo establishes benchmarks for long-horizon egocentric video-language comprehension that directly motivate OVO-Bench's operational testing of temporal dependencies.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). MVBench introduces multi-task temporal evaluation frameworks for dynamic video reasoning, serving as a core predecessor to time-sensitive video understanding benchmarks.
- Paper: ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos, Jr-Jen Chen et al. (2024). ReXTime provides essential methodology for benchmarking reasoning across non-simultaneous, temporally separated video events.
- Paper: Grounded Question-Answering in Long Egocentric Videos, Shangzhe Di et al. (2024). GroundVQA demonstrates the integration of temporal moment localization with question answering in continuous first-person video, informing online video query design.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench establishes standardized, robust multi-choice evaluation techniques and circular evaluation strategies for multimodal foundation models.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Video-LLaVA provides the foundational architecture for aligning video representations within large language models evaluated in streaming benchmarks.
- Paper: Streaming Video Question-Answering with In-context Video KV-Cache Retrieval, Shangzhe Di et al. (2025). ReKV develops an in-context key-value cache retrieval architecture that directly tackles the latency and streaming memory bottlenecks highlighted by OVO-Bench.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). TimeSuite enhances video LLMs via temporal grounded tuning to solve the exact temporal localization failures and hallucinations identified in OVO-Bench's evaluation.
- Paper: Adaptive Keyframe Sampling for Long Video Understanding, Xi Tang et al. (2025). This work introduces adaptive keyframe sampling to filter informative visual evidence dynamically, addressing the streaming inference latency challenges noted in OVO-Bench.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). This paper presents lightweight, question-aware frame selection modules to optimize downstream multimodal model efficiency and reduce processing latency.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Qwen3-VL incorporates explicit textual timestamps and interleaved spatial-temporal modeling to address long-video context limits and fine-grained temporal understanding.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). Qwen3.5-Omni applies streaming-oriented Thinker-Talker architectures and speech-video alignment for real-time interactive omnimodal understanding.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). VLWM extends real-time temporal and visual comprehension into dynamic world modeling and long-horizon action planning for interactive agents.
