AVA: Towards Agentic Video Analytics with Vision Language Models

Yuxuan YanShiqi JiangTing CaoYifan YangQianqian YangYuanchao ShuYuqing YangLili Qiu

article2025NSDI4 citations

Introduces AVA, an agentic video analytics system that builds real-time Event Knowledge Graphs to enable vision language models to accurately reason over ultra-long video streams exceeding ten hours.

Listen

Modern video analytics systems are critical for domains such as traffic management, industrial monitoring, and public safety. However, existing commercial systems are largely restricted to rigid, predefined tasks like simple object detection or short-term event tracking. While cutting-edge vision-language models offer the reasoning required for open-ended, natural language inquiries, their limited context windows make analyzing continuous video streams or ultra-long recordings spanning hundreds of hours computationally prohibitive. Most existing retrieval techniques struggle with complex multi-hop reasoning, temporal continuity, and query-focused summarization over extensive video archives.

The article demonstrates the design, implementation, and evaluation of AVA, an advanced system powered by vision-language models that achieves open-ended reasoning over ultra-long and continuous video streams. The objective is to evaluate how combining near-real-time event-centric indexing with proactive, agentic search over knowledge graphs can overcome the context length and cost limitations inherent in current video analytics.

The approach introduces a two-stage architecture evaluated through empirical experiments across standard benchmarks and real-world long-duration video collections. During the indexing phase, AVA breaks video feeds into short segments, generates descriptions using a compact vision-language model, and merges semantically related segments into an Event Knowledge Graph that maps temporally ordered events and associated entities. During the query phase, the system conducts a three-dimensional retrieval across events, entities, and visual frames, followed by an agentic tree search that navigates temporal relationships to gather relevant context before synthesizing answers via a thought-consistency framework. Performance was benchmarked on standard datasets like LVBench and VideoMME-Long, alongside AVA-100, a newly introduced benchmark containing 100 hours of continuous footage across traffic, wildlife, urban, and daily activity scenarios.

The key findings show that AVA significantly outperforms both baseline vision-language models and specialized retrieval systems. On the ultra-long AVA-100 benchmark, AVA achieved 75.8% accuracy, surpassing existing baselines by approximately 20.8%. On public benchmarks, AVA attained 62.3% accuracy on LVBench (a 16.9% lead over previous approaches) and 64.1% on VideoMME-Long (a 5.2% improvement). For complex temporal reasoning tasks on LVBench, AVA exceeded baseline accuracy by 35.6%. In terms of operational efficiency, the system constructs its knowledge graphs in near real time, processing between 2.5 and 6.7 frames per second on standard edge and server hardware, and maintains consistent analytical accuracy even when video length expands to 10 hours.

These results demonstrate that organizations can deploy scalable, open-ended video intelligence without suffering from context-window degradation or unsustainable cloud inference costs. By converting video streams into structured event graphs, organizations can enable natural language questioning over continuous operations, reducing query latency and network bandwidth while expanding analytics capabilities from simple detection to causal, multi-step problem solving.

For practical implementation, engineering and operations teams should pilot event knowledge graph indexing on edge-deployed hardware for critical monitoring streams, maintaining an agentic search depth of three to balance retrieval accuracy and latency. Organizations must also tailor model selection based on cost and responsiveness trade-offs, pairing compact models for continuous graph construction with higher-capacity reasoning models only during final answer generation.

Key limitations include computational latency during multi-path agentic searching, which requires up to a few minutes per complex query, and lower accuracy on fine-grained visual tasks such as precise object counting. Readers can place high confidence in AVA's general temporal reasoning and summarization accuracy, but should exercise caution and consider integrating specialized vision tools when high-precision numerical counting or real-time query responses are mandatory.

arXiv: 2505.00254I-ESC/Project-Ava

No sufficiently relevant recommendations were found.

Cover for AVA: Towards Agentic Video Analytics with Vision Language Models

Abstract

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent emergence of Vision Language Models (VLMs) as transformative technologies offers significant potential for enabling open-ended video understanding, reasoning, and analytics. Nevertheless, their limited context windows present challenges when processing ultra-long video content, which is prevalent in real-world applications. To address this, we introduce AVA, a VLM-powered system designed for open-ended, advanced video analytics. AVA incorporates two key innovations: (1) the near real-time construction of Event Knowledge Graphs (EKGs) for efficient indexing of long or continuous video streams, and (2) an agentic retrieval-generation mechanism that leverages EKGs to handle complex and diverse queries. Comprehensive evaluations on public benchmarks, LVBench and VideoMME-Long, demonstrate that AVA achieves state-of-the-art performance, attaining 62.3% and 64.1% accuracy, respectively-significantly surpassing existing VLM and video Retrieval-Augmented Generation (RAG) systems. Furthermore, to evaluate video analytics in ultra-long and open-world video scenarios, we introduce a new benchmark, AVA-100. This benchmark comprises 8 videos, each exceeding 10 hours in duration, along with 120 manually annotated, diverse, and complex question-answer pairs. On AVA-100, AVA achieves top-tier performance with an accuracy of 75.8%. The source code of AVA is available at this https URL. The AVA-100 benchmark can be accessed at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work and Motivation
  • 2.1 Video Analytics System and VLMs
  • 2.2 Long Video Understandings
  • 2.3 Retrieval Augmented Generation
  • 3 Ava System Overview
  • 4 Near-Real-Time Index Construction
  • 4.1 Event Knowledge Graph
  • 4.2 Semantic Chunking
  • 4.3 Entity Extraction and Linking
  • 5 Agentic Retrieval and Generation
  • 5.1 Tri-View Retrieval
  • 5.2 Agentic Searching on Graph
  • 5.3 Consistency Enhanced Generation
  • 6 Implementation
  • 7 Evaluation
  • 7.1 Evaluation Settings
  • 7.1.1 Benchmarks
  • 7.2 Baselines
  • 7.3 Overall Evaluation
  • 7.3.1 Overall Performance
  • 7.3.2 Performance on Different Query Categories
  • 7.3.3 Performance under Different Configurations
  • 7.3.4 Performance on Different Video Lengths
  • 7.3.5 System Overhead
  • 7.4 Ablation Evaluation
  • 7.4.1 Different Index Construction Methods
  • 7.4.2 Different Tree Search Depths
  • 7.4.3 Different Consistency Evaluation Settings
  • 8 Limitations and Future Work
  • 9 Conclusion
  • References
  • A Ava-100 Benchmark
  • A.1 Benchmark overview
  • A.2 Selected Scenarios and Data Sources
  • A.2.1 Human Daily Activities
  • A.2.2 City Walking
  • A.2.3 Traffic Monitoring
  • A.2.4 Wildlife Monitoring
  • A.3 Prompts

Knowls

  1. Knowl 1 — AVA’s architecture for open-ended analytics on long video

    model/method

    AVA is a vision-language-model (VLM) system designed to answer open-ended questions about long or continuous video, including temporal-reasoning and query-focused summarization questions. Its pipeline has two stages: a small VLM converts incoming video into an Event Knowledge Graph (EKG), and a retrieval-and-generation agent uses that index to gather relevant event descriptions and, when needed, their associated raw frames before producing an answer. The index is built continuously so its processing can keep up with a live stream; this design avoids sending an entire long video to a VLM at query time. AVA’s stated design goals are to support video collections of hundreds of hours or more, near-real-time index construction, and both fact-retrieval and complex multi-hop queries. These are design goals; measured throughput is reported separately.

  2. Knowl 2 — Event Knowledge Graph representation

    definition

    An Event Knowledge Graph (EKG) represents a video as temporally ordered events, entities identified in those events, and relations among them. Formally, G=(E,U,R)G=(E,U,R), where E={ei}E=\{e_i\} is the ordered set of video events, U={uj}U=\{u_j\} is the set of extracted entities, and R=Ree∪Ruu∪RueR=R_{ee}\cup R_{uu}\cup R_{ue}. Here, ReeR_{ee} contains temporal relations between events, such as before and after; RuuR_{uu} contains semantic relations between entities; and RueR_{ue} records an entity’s participation and contextual role in an event. Unlike a graph that represents entities and their attributes over an entire video without event structure, the EKG retains event-level context and temporal evolution, supporting event summaries and multi-hop temporal queries.

  3. Knowl 3 — Semantic chunking of video into events

    model/method

    AVA uses semantic chunking to form event-sized units despite events having different durations. It first divides the video into fixed-duration uniform chunks (3 seconds in the described implementation) and uses a small VLM to produce a textual description for each chunk. It compares descriptions of neighboring chunks with BERTScore and merges chunks judged semantically similar. A merge is permitted when every pair of uniform chunks within the proposed semantic chunk exceeds the merging threshold (0.65 in the implementation); the similarity at the boundary between adjacent semantic chunks must be below a separate, sufficiently low threshold, whose value is not specified. After merging, the VLM summarizes each semantic chunk. This process provides event descriptions and temporal extents for EKG construction; the paper reports scheduling the pairwise similarity computations in parallel.

  4. Knowl 4 — Entity extraction, deduplication, and event linking

    model/method

    For each semantic event, AVA uses a small VLM to extract entities and their relationships. Because the same entity may be described differently in different events, AVA encodes entity descriptions with JinaCLIP and clusters the resulting vectors using standard K-means rather than relying on exact string matching. Descriptions assigned to a cluster are treated as one linked entity, represented by the centroid of their embedding vectors. The resulting EKG is stored in five logical tables: events, entities, event-to-event relations, entity-to-entity relations, and entity-to-event relations. AVA also embeds raw video frames with JinaCLIP and associates each frame with its corresponding event, so retrieved graph information can be connected to visual evidence.

  5. Knowl 5 — Tri-view retrieval and event ranking

    equation

    For a query, AVA retrieves candidates in three views: event descriptions, entity representations, and raw-frame visual embeddings. Entities and frames are linked back to their associated events through the EKG. Since the views have different similarity scores, AVA combines their rankings with normalized Borda-style scores. For view mm, let EmE_m be its set of retrieved top-KK events and let sim⁡m(e)\operatorname{sim}_m(e) be the query similarity score for event ee in that view. The normalized score and aggregate event score are pm(e)=sim⁡m(e)∑e′∈Emsim⁡m(e′)p_m(e)=\frac{\operatorname{sim}_m(e)}{\sum_{e'\in E_m}\operatorname{sim}_m(e')} and s(e)=∑mpm(e)s(e)=\sum_m p_m(e), with an event absent from a view contributing no score from that view. AVA ranks the retrieved events by s(e)s(e); this ranking is also used to limit less relevant events during agentic search.

  6. Knowl 6 — Agentic graph search for temporal and multi-hop queries

    algorithm

    AVA begins with tri-view retrieval on the user’s query; the retrieved event list forms the root of a search tree. At each search node it can apply four actions: Forward retrieves temporally subsequent events; Backward retrieves preceding events; Re-query asks an LLM to generate alternative keywords and retrieves events for that new query; and Summary and Answer asks an LLM to answer from the current event descriptions and terminates that path. The search rolls out these actions until a configured maximum depth. In the implementation, maximum depth is 3 and the maintained event list is capped at 16; if it grows beyond the cap, lower-ranked events are dropped using the tri-view ranking. A depth-three example yields 13 answer-producing pathways. This exploration is intended to add neighboring temporal context and alternative query perspectives for summaries and multi-hop questions, rather than relying only on events matching the original query.

  7. Knowl 7 — Thought-consistency answer selection and visual verification

    model/method

    At each Summary and Answer node, AVA samples multiple chain-of-thought responses, then scores each distinct answer using both answer agreement and consistency among the reasoning traces that produced it. If nn samples produce answers aia_i with traces rir_i, the agreement score for answer aa is the fraction of samples with that answer. Its thought-consistency score is the mean BERTScore across pairs of traces associated with aa. AVA combines the scores as Sfinal(a)=λSa(a)+(1−λ)Sr(a)S_{\mathrm{final}}(a)=\lambda S_a(a)+(1-\lambda)S_r(a), where λ\lambda weights answer agreement and the implementation uses λ=0.3\lambda=0.3. The candidate with the highest combined score is selected for each answer node. AVA then selects the top two nodes with different answers, retrieves the raw frames linked to their events, and uses a VLM to generate visually grounded Check Frames and Answer responses; consistency scoring is applied to these responses as well. The implementation uses eight self-consistency samples, with sampling temperature between 0.5 and 0.7.

  8. Knowl 8 — AVA-100 ultra-long video analytics benchmark

    data/table

    AVA-100 was introduced to evaluate long-horizon analytics and reasoning on ultra-long video. It contains eight videos totaling 99.2 hours and 120 manually annotated question-answer pairs across human daily activities, city walking, traffic monitoring, and wildlife monitoring. The first four videos are moving, first-person recordings; the last four are fixed-camera, third-person recordings. Human annotators supplied reference answers, while GPT-4o generated multiple-choice distractors that were manually checked. The video-level statistics are:

    Video ID Duration (hours) QA pairs View
    ego-1 12.7 22 First-person (moving)
    ego-2 11.7 19 First-person (moving)
    citytour-1 12.0 19 First-person (moving)
    citytour-2 10.5 20 First-person (moving)
    traffic-1 14.9 12 Third-person (fixed)
    traffic-2 13.9 13 Third-person (fixed)
    wildlife-1 12.0 8 Third-person (fixed)
    wildlife-2 11.5 7 Third-person (fixed)
    Total 99.2 120 –

    The benchmark’s questions target tasks such as activity-sequence reasoning, landmark and route recall, time-specific traffic counts, and identifying animals across footage. Evaluation uses accuracy on multiple-choice questions.

  9. Knowl 9 — Accuracy on long-video and ultra-long analytics benchmarks

    empirical result

    AVA attained 62.3% accuracy on LVBench and 64.1% on VideoMME-Long, which the paper reports as state-of-the-art results and gains of 16.9% and 5.2%, respectively, over the strongest compared systems. LVBench contains 103 videos and 1,549 questions, with an average video duration of about 4,100 seconds; VideoMME-Long contains 300 videos and 900 questions, with an average duration of about 2,400 seconds. On AVA-100’s eight videos exceeding 10 hours each, AVA reached 75.8% accuracy. The paper reports an approximately 20.8% gain over vectorized-retrieval baselines and 26.9% over uniform-sampling baselines on AVA-100. On LVBench, AVA also improved over the Gemini-1.5-Pro uniform-sampling and vectorized-retrieval baselines across temporal grounding, summarization, reasoning, entity recognition, event understanding, and key-information retrieval; the reported gains for these categories, in that order, are 16%, 5.3%, 35.6%, 21.2%, 17.5%, and 18.9%.

  10. Knowl 10 — Measured indexing throughput and index-construction comparison

    empirical result

    AVA’s index constructor processed LVBench videos at an input rate of 2 FPS at an average of 6.7 FPS on two A100 GPUs, 4.4 FPS on one RTX 4090, and 2.5 FPS on one RTX 3090. Thus, the reported single-4090 rate exceeded the tested input rate, while the single-3090 rate did not. In a separate ablation on a 1.2-hour LVBench subset (20 videos and 305 questions), using Qwen2.5-7B to build the index and Qwen2.5-14B for answer generation, AVA’s EKG achieved 39.7 accuracy with 0.31 hours of construction overhead. Text-graph baselines MiniRAG and LightRAG achieved 28.1 accuracy with 3.49 hours and 30.6 accuracy with 3.52 hours, respectively, under the reported comparison setup. For query-time generation measured on one A100, tri-view retrieval with JinaCLIP took 0.44 seconds and used 0.8 GB of GPU memory; agentic search took 101.5 seconds with Qwen2.5-14B (30 GB) or 174.2 seconds with Qwen2.5-32B (40 GB); consistency-enhanced generation took 45.8 seconds with Qwen2.5-VL-7B (31 GB) or 14.2 seconds with API-based Gemini-1.5-Pro (GPU memory not reported). These measurements identify LLM agentic search as the largest listed query-time latency.

Coverage note — The paper’s stated limitations and proposed future directions—fixed, costly tree search and weaker performance on specialized tasks such as precise counting—are omitted because they are prospective qualifications rather than core method or evaluation results; the detailed depth and consistency-parameter ablations are likewise not included separately.

References

  1. 1.4KWorldWandering. 4kworldwandering youtube channel. https://www.youtube.com/@4KWorldWandering, 2025.
  2. 2.Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  3. 3.Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. Phi-4 technical report. arXiv preprint arXiv:2412.08905, 2024.
  4. 4.Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, et al. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras. arXiv preprint arXiv:2503.01743, 2025.
  5. 5.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  6. 6.Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923, 2025.
  7. 7.Romil Bhardwaj, Zhengxu Xia, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Nikolaos Karianakis, Kevin Hsieh, Paramvir Bahl, and Ion Stoica. Ekya: Continuous learning of video analytics models on edge compute servers. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2022.
  8. 8.Arkansas Critter Cam. Arkansas critter cam youtube channel. https://www.youtube.com/@arkansascrittercam611, 2025.
  9. 9.Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024.
  10. 10.LMDeploy Contributors. Lmdeploy: A toolkit for compressing, deploying, and serving llm. https://github.com/InternLM/lmdeploy, 2023.
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of conference of the North American chapter of the association for computational linguistics (NAACL), 2019.
  12. 12.Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130, 2024.
  13. 13.Tianyu Fan, Jingyuan Wang, Xubin Ren, and Chao Huang. Minirag: Towards extremely simple retrieval-augmented generation. arXiv preprint arXiv:2501.06713, 2025.
  14. 14.Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. In Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  15. 15.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  16. 16.Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. Lightrag: Simple and fast retrieval-augmented generation. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024.
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  18. 18.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654, 2020.
  19. 19.Samvit Jain, Xun Zhang, Yuhao Zhou, Ganesh Ananthanarayanan, Junchen Jiang, Yuanchao Shu, Paramvir Bahl, and Joseph Gonzalez. Spatula: Efficient Cross-camera Video Analytics on Large Camera Networks. In ACM/IEEE Symposium on Edge Computing (SEC), 2020.
  20. 20.Shiqi Jiang, Zhiqi Lin, Yuanchun Li, Yuanchao Shu, and Yunxin Liu. Flexible high-resolution object detection on edge devices with tunable latency. In ACM International Conference on Mobile Computing and Networking (Mobicom), 2021.
  21. 21.Mehrdad Khani, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang, Ravi Netravali, Yuanchao Shu, Mohammad Alizadeh, and Victor Bahl. Recl: Responsive resource-efficient continuous learning for video analytics. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2023.
  22. 22.Andreas Koukounas, Georgios Mastrapas, Michael Günther, Bo Wang, Scott Martens, Isabelle Mohr, Saba Sturua, Mohammad Kalim Akram, Joan Fontanals Martínez, Saahil Ognawala, et al. Jina clip: Your clip model is also your text retriever. arXiv preprint arXiv:2405.20204, 2024.
  23. 23.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  24. 24.Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. Awq: Activation-aware weight quantization for on-device llm compression and acceleration. In International Conference on Machine Learning Systems (MLSys), 2024.
  25. 25.Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, et al. Nvila: Efficient frontier visual language models. In Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  26. 26.Nature Live. Nature live youtube channel. https://www.youtube.com/@nature-live, 2025.
  27. 27.Yan Lu, Shiqi Jiang, Ting Cao, and Yuanchao Shu. Turbo: Opportunistic enhancement for edge video analytics. In Proceedings of ACM Conference on Embedded Networked Sensor Systems (Sensys), 2022.
  28. 28.Yongdong Luo, Xiawu Zheng, Xiao Yang, Guilin Li, Haojia Lin, Jinfa Huang, Jiayi Ji, Fei Chao, Jiebo Luo, and Rongrong Ji. Video-rag: Visually-aligned retrieval-augmented long video comprehension. In Advances in Neural Information Processing Systems (NeurIPS), 2025.
  29. 29.Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun, Shutao Li, Hamid Rezatofighi, and Jianfei Cai. Drvideo: Document retrieval based long video understanding. In Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  30. 30.Ya Nan, Shiqi Jiang, and Mo Li. Large-scale video analytics with cloud–edge collaborative continuous learning. ACM Trans. Sen. Netw., 2023.
  31. 31.Arthi Padmanabhan, Neil Agarwal, Anand Iyer, Ganesh Ananthanarayanan, Yuanchao Shu, Nikolaos Karianakis, Guoqing Harry Xu, and Ravi Netravali. Gemel: Model merging for memory-efficient,real-time video analytics at the edge. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2023.
  32. 32.Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2024.
  33. 33.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021.
  34. 34.Microsoft Research. Lazygraphrag: Setting a new standard for quality and cost. lazygraphrag-setting-a-new-standard-for-quality-and-cost, 2024.
  35. 35.Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, et al. Longvu: Spatiotemporal adaptive compression for long video-language understanding. In International Conference on Machine Learning (ICML), 2025.
  36. 36.Yunhang Shen, Chaoyou Fu, Shaoqi Dong, Xiong Wang, Peixian Chen, Mengdan Zhang, Haoyu Cao, Ke Li, Xiawu Zheng, Yan Zhang, et al. Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray. arXiv preprint arXiv:2502.05177, 2025.
  37. 37.Vibhaalakshmi Sivaraman, Pantea Karimi, Vedantha Venkatapathy, Mehrdad Khani, Sadjad Fouladi, Mohammad Alizadeh, Frédo Durand, and Vivienne Sze. Gemino: Practical and robust neural compression for video conferencing. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2024.
  38. 38.Mingxing Tan, Ruoming Pang, and Quoc V Le. Efficientdet: Scalable and efficient object detection. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  39. 39.Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  40. 40.Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  41. 41.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In International Conference on Computer Vision (ICCV), 2015.
  42. 42.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024.
  43. 43.Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Xiaotao Gu, Shiyu Huang, Bin Xu, Yuxiao Dong, et al. Lvbench: An extreme long video understanding benchmark. In International Conference on Computer Vision (ICCV), 2025.
  44. 44.Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu, Li Cao, and Liqiang Nie. Adaretake: Adaptive redundancy reduction to perceive longer for video-language understanding. arXiv preprint arXiv:2503.12559, 2025.
  45. 45.Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy. Videoagent: Long-form video understanding with large language model as agent. In European Conference on Computer Vision (ECCV), 2024.
  46. 46.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR), 2023.
  47. 47.Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon, Feng Cheng, Gedas Bertasius, and Mohit Bansal. Videotree: Adaptive tree-based video representation for llm reasoning on long videos. In Conference on Computer Vision and Pattern Recognition (CVPR), 2025.
  48. 48.Zeyuan Yang, Delin Chen, Xueyang Yu, Maohao Shen, and Chuang Gan. Vca: Video curious agent for long video understanding. In International Conference on Computer Vision (ICCV), 2025.
  49. 49.Chen-Lin Zhang, Jianxin Wu, and Yin Li. Actionformer: Localizing moments of actions with transformers. In European Conference on Computer Vision (ECCV), 2022.
  50. 50.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations (ICLR), 2020.
  51. 51.Xu Zhang, Yiyang Ou, Siddhartha Sen, and Junchen Jiang. Sensei: Aligning video streaming quality with dynamic user sensitivity. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2021.
  52. 52.Yiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand Iyer, Yuanchao Shu, Victor Bahl, Z Morley Mao, and Mosharaf Chowdhury. Vulcan: Automatic query planning for live ml analytics. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2024.
  53. 53.Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024.
  54. 54.Siyun Zhao, Yuqing Yang, Zilong Wang, Zhiyuan He, Luna K Qiu, and Lili Qiu. Retrieval augmented generation (rag) and beyond: A comprehensive survey on how to make your llms use external data more wisely. arXiv preprint arXiv:2409.14924, 2024.
  55. 55.Yue Zheng, Yuhao Chen, Bin Qian, Xiufang Shi, Yuanchao Shu, and Jiming Chen. A review on edge large language models: Design, execution, and applications. ACM Comput. Surv., 2025.

Citation

MLA
Yan, Y., et al. “AVA: Towards Agentic Video Analytics with Vision Language Models”. arXiv, 2025, https://doi.org/10.48550/arxiv.2505.00254.
APA
Yan, Y., Jiang, S., Cao, T., Yang, Y., Yang, Q., Shu, Y., Yang, Y., & Qiu, L. (2025). AVA: Towards Agentic Video Analytics with Vision Language Models. arXiv. https://doi.org/10.48550/arxiv.2505.00254
Chicago
Yan, Y., S. Jiang, T. Cao, et al. 2025. “AVA: Towards Agentic Video Analytics with Vision Language Models”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2505.00254.
Harvard
Yan, Y. et al. (2025) “AVA: Towards Agentic Video Analytics with Vision Language Models”. arXiv. Available at: https://doi.org/10.48550/arxiv.2505.00254.
Vancouver
1. Yan Y, Jiang S, Cao T, Yang Y, Yang Q, Shu Y, Yang Y, Qiu L (2025) AVA: Towards Agentic Video Analytics with Vision Language Models. https://doi.org/10.48550/arxiv.2505.00254

BibTeX

@misc{https://doi.org/10.48550/arxiv.2505.00254,
  doi = {10.48550/ARXIV.2505.00254},
  url = {https://arxiv.org/abs/2505.00254},
  author = {Yan, Yuxuan and Jiang, Shiqi and Cao, Ting and Yang, Yifan and Yang, Qianqian and Shu, Yuanchao and Yang, Yuqing and Qiu, Lili},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {AVA: Towards Agentic Video Analytics with Vision Language Models},
  publisher = {arXiv},
  year = {2025},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/