AVA: Towards Agentic Video Analytics with Vision Language Models
Yuxuan YanShiqi JiangTing CaoYifan YangQianqian YangYuanchao ShuYuqing YangLili Qiu
Introduces AVA, an agentic video analytics system that builds real-time Event Knowledge Graphs to enable vision language models to accurately reason over ultra-long video streams exceeding ten hours.
Modern video analytics systems are critical for domains such as traffic management, industrial monitoring, and public safety. However, existing commercial systems are largely restricted to rigid, predefined tasks like simple object detection or short-term event tracking. While cutting-edge vision-language models offer the reasoning required for open-ended, natural language inquiries, their limited context windows make analyzing continuous video streams or ultra-long recordings spanning hundreds of hours computationally prohibitive. Most existing retrieval techniques struggle with complex multi-hop reasoning, temporal continuity, and query-focused summarization over extensive video archives.
The article demonstrates the design, implementation, and evaluation of AVA, an advanced system powered by vision-language models that achieves open-ended reasoning over ultra-long and continuous video streams. The objective is to evaluate how combining near-real-time event-centric indexing with proactive, agentic search over knowledge graphs can overcome the context length and cost limitations inherent in current video analytics.
The approach introduces a two-stage architecture evaluated through empirical experiments across standard benchmarks and real-world long-duration video collections. During the indexing phase, AVA breaks video feeds into short segments, generates descriptions using a compact vision-language model, and merges semantically related segments into an Event Knowledge Graph that maps temporally ordered events and associated entities. During the query phase, the system conducts a three-dimensional retrieval across events, entities, and visual frames, followed by an agentic tree search that navigates temporal relationships to gather relevant context before synthesizing answers via a thought-consistency framework. Performance was benchmarked on standard datasets like LVBench and VideoMME-Long, alongside AVA-100, a newly introduced benchmark containing 100 hours of continuous footage across traffic, wildlife, urban, and daily activity scenarios.
The key findings show that AVA significantly outperforms both baseline vision-language models and specialized retrieval systems. On the ultra-long AVA-100 benchmark, AVA achieved 75.8% accuracy, surpassing existing baselines by approximately 20.8%. On public benchmarks, AVA attained 62.3% accuracy on LVBench (a 16.9% lead over previous approaches) and 64.1% on VideoMME-Long (a 5.2% improvement). For complex temporal reasoning tasks on LVBench, AVA exceeded baseline accuracy by 35.6%. In terms of operational efficiency, the system constructs its knowledge graphs in near real time, processing between 2.5 and 6.7 frames per second on standard edge and server hardware, and maintains consistent analytical accuracy even when video length expands to 10 hours.
These results demonstrate that organizations can deploy scalable, open-ended video intelligence without suffering from context-window degradation or unsustainable cloud inference costs. By converting video streams into structured event graphs, organizations can enable natural language questioning over continuous operations, reducing query latency and network bandwidth while expanding analytics capabilities from simple detection to causal, multi-step problem solving.
For practical implementation, engineering and operations teams should pilot event knowledge graph indexing on edge-deployed hardware for critical monitoring streams, maintaining an agentic search depth of three to balance retrieval accuracy and latency. Organizations must also tailor model selection based on cost and responsiveness trade-offs, pairing compact models for continuous graph construction with higher-capacity reasoning models only during final answer generation.
Key limitations include computational latency during multi-path agentic searching, which requires up to a few minutes per complex query, and lower accuracy on fine-grained visual tasks such as precise object counting. Readers can place high confidence in AVA's general temporal reasoning and summarization accuracy, but should exercise caution and consider integrating specialized vision tools when high-precision numerical counting or real-time query responses are mandatory.
No sufficiently relevant recommendations were found.
- Paper: From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations, Yuchen Guan et al. (2026). It carries long-video analytics beyond AVA’s event-graph retrieval by distilling video knowledge into compact model adapters for much faster query-time responses.
