Adaptive Keyframe Sampling for Long Video Understanding

Xi TangJihao QiuLingxi XieYunjie TianJianbin JiaoQixiang Ye

article2025CVPR118 citations

Proposes a plug-and-play keyframe selection algorithm that balances prompt relevance with temporal coverage to improve long video question-answering accuracy in multimodal large language models without exceeding token limits.

Listen

Artificial intelligence systems designed for visual and language understanding struggle to process long video files because the volume of visual data quickly exceeds their computing and context capacity. Standard industry solutions typically sample a fixed, small subset of frames at evenly spaced intervals, a practice known as uniform sampling. However, this arbitrary selection frequently skips pivotal moments, causing artificial intelligence models to miss essential visual evidence and return inaccurate answers.

The article demonstrates an optimization method called Adaptive Keyframe Sampling to overcome these processing constraints. The primary objective is to evaluate whether intelligently pre-filtering and selecting the most informative video frames before feeding them to an AI model improves overall video comprehension accuracy without requiring changes to the underlying model architecture.

To evaluate this approach, the authors integrated the method as a plug-and-play module across three standard multimodal language models and tested performance on two long-video benchmarks containing footage up to an hour long. The method uses a smaller, secondary vision-language model to score candidate frames based on two balanced criteria: prompt relevance, which assesses how well a frame relates to the user's specific query, and temporal coverage, which ensures selected frames are distributed across the video timeline to prevent clustering and redundant data capture. Candidate frames were extracted at varying sampling frequencies to evaluate computational trade-offs, and video subtitles were intentionally excluded to evaluate purely visual reasoning.

The findings show consistent performance gains across all evaluated systems. Incorporating the sampling technique improved video question-answering accuracy across the board, lifting a standard 7-billion-parameter model's baseline scores by up to 5.0 percentage points on LongVideoBench and 2.3 percentage points on VideoMME. Notably, a 7-billion-parameter open-source model enhanced with this sampling technique achieved 62.7% accuracy on LongVideoBench, outperforming larger proprietary systems such as GPT-4V and Gemini-1.5-Flash that evaluated four times as many frames. The analysis also confirmed that prompt-guided frame selection successfully adapts across diverse tasks, including video description and specific moment retrieval, while maintaining strong accuracy even when candidate frames are sampled at low frequencies down to one frame every four seconds.

These results indicate that pre-filtering visual inputs is a highly effective, cost-efficient strategy for deploying artificial intelligence on complex, high-dimensional media. Rather than expending substantial compute budget to expand model context windows or process massive video files in their entirety, organizations can achieve superior performance with smaller, faster models by improving the quality of the visual data supplied to them. This provides an immediate operational pathway to reduce computing overhead and cloud infrastructure costs.

For near-term deployment, technical teams should consider integrating lightweight pre-filtering modules before the primary vision processing pipeline in long-video applications. System architects can customize the balance between prompt relevance and temporal coverage depending on the use case, prioritizing relevance for pinpoint retrieval tasks and coverage for global video summarization. Further work should explore refining pre-filtering efficiency to lower processing latency even further when analyzing ultra-long video streams.

Confidence in these findings is supported by consistent gains across multiple benchmark datasets and model families. However, decision-makers should account for several limitations: the method introduces minor computational overhead during the initial frame scoring phase, and its effectiveness depends in part on the capability of the smaller scoring model to correctly interpret the prompt. In addition, benchmark tests relied on multiple-choice formats without audio or subtitle inputs, meaning performance under complex, multi-modal production conditions warrants targeted pilot testing.

Cover for Adaptive Keyframe Sampling for Long Video Understanding

Abstract

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and (2) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Principles of Keyframe Selection
  • 3.3. Adaptive Keyframe Sampling
  • 4. Experiments
  • 4.1. Experimental Setup and Details
  • 4.2. Comparison to the State-of-the-Art
  • 4.3. Diagnostic on Keyframe Selection
  • 4.4. Ablative Studies
  • 4.5. Generalization to Other Tasks
  • 5. Conclusions
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Adaptive Keyframe Sampling as a plug-and-play video pre-filter

    model/method

    Adaptive Keyframe Sampling (AKS) is a prompt-conditioned module placed before a video-based multimodal large language model (MLLM)'s visual encoder. For a video V∈RT×W×H×CV \in \mathbb{R}^{T \times W \times H \times C} with TT frames, frame tt is encoded into visual tokens FtF_t, and a text instruction is denoted by QQ. Given a fixed visual-token budget corresponding to MM frames, AKS returns an index set I⊆{1,…,T}I \subseteq \{1,\ldots,T\} with ∣I∣=M|I|=M. Only the tokens {Ft:t∈I}\{F_t:t\in I\} are passed to the frozen MLLM.

    AKS uses a vision-language model to score the relevance of every candidate frame to QQ, then chooses frames by balancing high prompt relevance against temporal coverage. This design lets the same MLLM select different visual contexts for different questions without retraining or changing the MLLM parameters.

  2. Knowl 2 — Ideal keyframe selection objective

    equation

    The ideal keyframe-selection function is defined as

    KSM⁡(Q,F)=arg⁡max⁡I⊆{1,…,T}, ∣I∣=M  Gˉ({Ft:t∈I}),\operatorname{KSM}(Q,F)=\underset{I\subseteq\{1,\ldots,T\},\ |I|=M}{\arg\max}\;\bar{G}\left(\{F_t:t\in I\}\right),

    where QQ is the text prompt, F={F1,…,FT}F=\{F_1,\ldots,F_T\} is the sequence of frame-level visual-token sets, MM is the permitted number of frames, and Gˉ\bar{G} is a complementary confidence measure for the MLLM output produced from the selected visual tokens. The objective is not directly usable in practice because it contains exponentially many candidate subsets and because there is no direct supervision identifying the ideal keyframes: a correct answer does not imply that the selected frames are optimal, and an incorrect answer does not identify which frame-selection error caused it.

  3. Knowl 3 — Relevance–coverage formulation of keyframe selection

    equation

    AKS approximates the ideal objective by maximizing prompt relevance and temporal coverage:

    KSM⁡(Q,F)=arg⁡max⁡I⊆{1,…,T}, ∣I∣=M[∑t∈Ir(Q,Ft)+λ c(I)].\operatorname{KSM}(Q,F)=\underset{I\subseteq\{1,\ldots,T\},\ |I|=M}{\arg\max}\left[\sum_{t\in I}r(Q,F_t)+\lambda\,c(I)\right].

    Here r(Q,Ft)r(Q,F_t) is the relevance score between prompt QQ and frame representation FtF_t, c(I)c(I) is a coverage score for the selected timestamps over the video duration, and λ≥0\lambda\geq 0 controls the trade-off. Relevance favors frames containing visual information useful for answering the prompt; coverage discourages selecting redundant neighboring frames and rewards distributing the fixed number of frames across distinct temporal regions. AKS estimates rr with a cheaper vision-language model rather than repeatedly evaluating the target MLLM.

  4. Knowl 4 — Recursive temporal coverage estimator

    model/method

    AKS estimates temporal coverage by recursively partitioning the video time axis. The normalized timestamp interval [0,T)[0,T) is first divided into two equal bins, [0,T/2)[0,T/2) and [T/2,T)[T/2,T). If m1m_1 and m2m_2 selected frames fall in the two bins, an imbalance penalty proportional to ∣m1−m2∣|m_1-m_2| is applied; balanced occupancy indicates stronger coverage. Each bin is then split into two equal sub-bins and the same occupancy-balance criterion is applied recursively.

    The recursion continues through a maximum depth LL satisfying L≤⌈log⁡2M⌉L\leq\lceil\log_2 M\rceil, where MM is the number of selected frames. This binned approximation is motivated by the temporal-spacing idea behind Ripley’s KK-function, but avoids pairwise computation by treating two timestamps as locally related when they fall in the same bin. The resulting coverage term is therefore a multilevel penalty against uneven allocation of keyframes across time.

  5. Knowl 5 — Adaptive judge-and-split selection algorithm

    algorithm

    The AKS adaptive sampler, called ADA in the paper, takes a video, prompt QQ, frame budget MM, maximum recursion depth LL, and relevance threshold sthrs_{\mathrm{thr}} as input. The implementation samples candidate frames from the raw video at 1 frame per second and computes each candidate's prompt-frame score with BLIP image-text matching.

    Input: Prompt Q, candidate frames with timestamps, frame budget M, maximum depth L, threshold s_thr
    Output: M selected keyframes
    Compute relevance score r(Q, F_t) for every candidate frame
    Start recursive selection on the full time interval with allocation M and depth 0
    RecursiveSelect(current interval, allocated count k, depth):
        If k is 0, return no frames
        If k is 1, return the highest-scoring frame in the interval
        Compute s_all as the mean score of all candidate frames in the interval
        Compute s_top as the mean score of the k highest-scoring frames in the interval
        If depth has reached L or s_top - s_all exceeds s_thr:
            Return the k highest-scoring frames in the interval
        Split the interval into two equal temporal sub-intervals
        Allocate k frames as evenly as possible between the two sub-intervals
        Recursively select frames in each sub-interval
        Return the union of the recursively selected frames

    The threshold condition preserves highly relevant frames when the best scores are clearly separated from the average; otherwise, the interval is split so that the selected set covers multiple temporal regions. The paper replaces explicit tuning of the objective weight λ\lambda with the score-gap threshold sthrs_{\mathrm{thr}}. At the maximum depth, the remaining allocation is filled with the highest-scoring frames subject to the recursive bin allocation.

  6. Knowl 6 — TOP, BIN, and UNI are limiting sampling strategies

    theoretical result

    The relevance–coverage objective yields three interpretable limiting strategies. When λ=0\lambda=0, coverage is ignored and the solution is TOP sampling: select the MM frames with the largest relevance scores, regardless of their temporal locations. When λ→+∞\lambda\rightarrow+\infty, coverage dominates and BIN sampling selects the highest-scoring frame in each temporal bin; if there are more bins than the available MM frames, it retains the highest-scoring bin champions. If the relevance scorer is constant across time, BIN degenerates to UNI, the uniform-sampling baseline.

    ADA is the finite-trade-off strategy between these limits. It concentrates frames in high-relevance regions when the score evidence is strong, but recursively enforces temporal spread when relevance peaks do not justify concentrating the entire frame budget.

  7. Knowl 7 — Experimental protocol for long-video MLLM evaluation

    experimental setup

    AKS was evaluated with the LMMs-Eval framework on the LongVideoBench validation set and VideoMME, both of which contain long videos that can exceed one hour. The evaluation used multiple-choice questions without video subtitles, so the comparison focused on visual keyframe selection rather than auxiliary textual information.

    The target MLLMs were Qwen2-VL-7B, LLaVA-OV-7B, and LLaVA-Video-7B, with their original parameters unchanged. Their input frames were replaced by AKS-selected frames while preserving each model's frame budget: 32 frames for Qwen2-VL and LLaVA-OV, and 64 frames for LLaVA-Video. Candidate frames were sampled at 1 frame per second. BLIP encoded the prompt and each candidate image, and its image-text matching score served as r(Q,Ft)r(Q,F_t). ADA was used as the default AKS strategy.

  8. Knowl 8 — AKS improves long-video question answering

    data/table

    The central quantitative comparison measures video question-answering accuracy in percent on LongVideoBench validation and VideoMME. AKS changes only the selected frames, while the underlying MLLM remains the same. It consistently improves all three 7B baselines and allows LLaVA-Video-7B to exceed several much larger or proprietary systems.

    Could not parse LaTeX table

    The gains from AKS are +5.0+5.0 and +2.3+2.3 percentage points for Qwen2-VL on LongVideoBench and VideoMME, +4.5+4.5 and +1.9+1.9 for LLaVA-OV, and +3.8+3.8 and +0.9+0.9 for LLaVA-Video. LLaVA-Video-7B with AKS reaches 62.7%62.7\% on LongVideoBench, exceeding the reported 61.9%61.9\% of LLaVA-Video-72B without AKS, and also exceeds GPT-4V and Gemini-1.5-Flash on that benchmark.

  9. Knowl 9 — ADA outperforms uniform, relevance-only, and coverage-only sampling

    data/table

    Using the unchanged LLaVA-Video-7B MLLM, the sampling strategies differ only in how they select the same type of frame input. UNI is uniform sampling, TOP selects the highest-scoring frames, BIN emphasizes temporal bin coverage, and ADA adaptively combines relevance and coverage.

    Could not parse LaTeX table

    ADA gives the best result on both benchmarks. TOP is stronger on LongVideoBench, whose questions often focus on one moment and therefore benefit from concentrating frames around a strong relevance peak. BIN is stronger than TOP on VideoMME, whose questions more often require evidence from multiple moments. ADA retains the advantages of both behaviors by allocating frames adaptively.

  10. Knowl 10 — AKS remains effective with sparse candidate sampling and different scorers

    data/table

    The paper evaluates the computational trade-off in AKS by reducing the candidate-frame sampling frequency and by replacing BLIP as the relevance scorer. Accuracy is reported for LLaVA-Video-7B; the frame count is the number ultimately fed to the MLLM.

    Could not parse LaTeX table

    Even at 0.10.1 fps, AKS remains above uniform sampling for the tested LongVideoBench frame budgets; on VideoMME, 0.250.25 fps is generally sufficient to surpass the uniform baseline. The scorer comparison was:

    Could not parse LaTeX table

    BLIP performs best among the tested scorers on LongVideoBench, whereas CLIP performs best on VideoMME at every listed frame budget. Thus AKS does not depend on one specific relevance model, although scorer choice affects the benchmark-specific accuracy.

  11. Knowl 11 — AKS transfers to referring and captioning

    empirical result

    AKS was also applied without retraining to video referring and video captioning using LLaVA-Video-7B. For referring, the prompt was changed to questions such as “What is [target] doing in the video?”; for captioning, the prompt requested a video description and answer options were removed.

    In the reported qualitative comparisons, uniform sampling produced incorrect descriptions such as claiming that a woman wearing sunglasses was walking through a garden, or that a shirtless man was leading donkeys. AKS selected frames showing the relevant actions and produced descriptions that the woman was sitting at a table with a drink and that the shirtless man was standing in a small rectangular pool. For captioning, uniform sampling described a tropical landscape, whereas AKS selected temple views and generated a description of the temple's carvings, statues, stone structure, and surrounding environment. These examples support the claim that AKS improves visual context selection beyond multiple-choice video question answering.

Coverage note — The full $L$–$s_{\mathrm{thr}}$ hyperparameter grid and the paper's additional frame-by-frame qualitative examples were omitted because they mainly refine the reported strategy trade-offs rather than add a separate load-bearing contribution.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022. 2
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 2
  3. 3.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022. 2
  6. 6.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500, 2023. 2
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 2
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2
  9. 9.Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024. 1, 2, 5
  10. 10.Wei Han, Hui Chen, Min-Yen Kan, and Soujanya Poria. Self-adaptive sampling for accurate video question answering on image text models. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2522–2534, 2024. 2
  11. 11.Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim. Ma-lmm: Memory-augmented large multimodal model for long-term video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13504–13514, 2024. 2
  12. 12.Mojan Javaheripi, Sebastien Bubeck, Marah Abdin, Jyoti Aneja, Sebastien Bubeck, Caio Cesar Teodoro Mendes, Weizhu Chen, Allie Del Giorno, Ronen Eldan, Sivakanth Gopi, et al. Phi-2: The surprising power of small language models. Microsoft Research Blog, 2023. 2
  13. 13.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment Anything. arXiv preprint arXiv:2304.02643, 2023. 2
  14. 14.Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmentation via Large Language Model. arXiv preprint arXiv:2308.00692, 2023. 1, 2
  15. 15.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 2, 5
  16. 16.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR, 2022. 4, 5, 7
  17. 17.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. arXiv preprint arXiv:2301.12597, 2023. 2
  18. 18.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding, 2024. 2
  19. 19.Zhaowei Li, Qi Xu, Dong Zhang, Hang Song, Yiqing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Van Tu Vu, Zhida Huang, and Tao Wang. Groundinggpt:language enhanced multi-modal grounding model, 2024. 2
  20. 20.Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection, 2023. 1, 2, 5
  21. 21.Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26689–26699, 2024. 5
  22. 22.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. arXiv preprint arXiv:2304.08485, 2023. 1, 2
  23. 23.Jiajun Liu, Yibing Wang, Hanghang Ma, Xiaoping Wu, Xiaoqi Ma, Xiaoming Wei, Jianbin Jiao, Enhua Wu, and Jie Hu. Kangaroo: A powerful video-language model supporting long-context video input. arXiv preprint arXiv:2408.15542, 2024. 2
  24. 24.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10012–10022, 2021. 2
  25. 25.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Da Li, Pengcheng Lu, Tao Wang, Linmei Hu, Minghui Qiu, and Zhongyu Wei. Valley: Video assistant with large language model enhanced ability, 2023. 2
  26. 26.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models, 2023. 2
  27. 27.Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Mubarak Shah, and Fahad Khan. Pg-video-llava: Pixel grounding large video-language models, 2023. 2
  28. 28.OpenAI. Gpt-4v. https://openai.com/index/gpt-4v-system-card/, 2023. 5
  29. 29.OpenAI. Hello gpt-4o. https://openai . com /index/hello-gpt-4o/, 2024. 5
  30. 30.Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models. arXiv preprint arXiv:2405.16009, 2024. 2
  31. 31.Jihao Qiu, Yuan Zhang, Xi Tang, Lingxi Xie, Tianren Ma, Pengyu Yan, David Doermann, Qixiang Ye, and Yunjie Tian. Artemis: Towards referential understanding in complex videos. arXiv preprint arXiv:2406.00258, 2024. 2
  32. 32.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 2, 3, 4, 5, 7
  33. 33.Brian D Ripley. The second-order analysis of stationary point processes. Journal of applied probability, 13(2):255–266, 1976. 4
  34. 34.Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, et al. Longvu: Spatiotemporal adaptive compression for long video-language understanding. arXiv preprint arXiv:2410.17434, 2024. 2
  35. 35.Dachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang, Chun Yuan, and Jiaqi Wang. Crossget: Cross-guided ensemble of tokens for accelerating vision-language transformers. arXiv preprint arXiv:2305.17455, 2023. 2
  36. 36.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18221–18232, 2024. 2
  37. 37.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. 5
  38. 38.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022. 2
  39. 39.Yunjie Tian, Lingxi Xie, Xiaopeng Zhang, Jiemin Fang, Haohang Xu, Wei Huang, Jianbin Jiao, Qi Tian, and Qixiang Ye. Semantic-aware generation for self-supervised visual representation learning. arXiv preprint arXiv:2111.13163, 2021. 2
  40. 40.Yunjie Tian, Lingxi Xie, Zhaozhi Wang, Longhui Wei, Xiaopeng Zhang, Jianbin Jiao, Yaowei Wang, Qi Tian, and Qixiang Ye. Integrally Pre-Trained Transformer Pyramid Networks. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18610–18620. IEEE, 2023. 2
  41. 41.Yunjie Tian, Tianren Ma, Lingxi Xie, Jihao Qiu, Xi Tang, Yuan Zhang, Jianbin Jiao, Qi Tian, and Qixiang Ye. Chatterbox: Multi-round multimodal referring and grounding. arXiv preprint arXiv:2401.13307, 2024. 1, 2
  42. 42.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 2
  43. 43.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 2, 5
  44. 44.Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy. Videoagent: Long-form video understanding with large language model as agent. In European Conference on Computer Vision, pages 58–76. Springer, 2025. 2
  45. 45.Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon, Feng Cheng, Gedas Bertasius, and Mohit Bansal. Videotree: Adaptive tree-based video representation for llm reasoning on long videos. arXiv preprint arXiv:2405.19209, 2024. 2
  46. 46.Hongchen Wei and Zhenzhong Chen. Visual context window extension: A new perspective for long video understanding. arXiv preprint arXiv:2409.20018, 2024. 2
  47. 47.Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, and Bohan Zhuang. Longvlm: Efficient long video understanding via large language models. arXiv preprint arXiv:2404.03384, 2024. 2
  48. 48.Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long-context interleaved video-language understanding. arXiv preprint arXiv:2407.15754, 2024. 1, 2, 5
  49. 49.Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994, 2024. 1, 5
  50. 50.Fuzhao Xue, Yukang Chen, Dacheng Li, Qinghao Hu, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Zhijian Liu, et al. Longvila: Scaling long-context visual language models for long videos. arXiv preprint arXiv:2408.10188, 2024. 2
  51. 51.An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024. 5
  52. 52.Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800, 2024. 5
  53. 53.Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. Self-chained image-language model for video localization and question answering. Advances in Neural Information Processing Systems, 36, 2024. 7
  54. 54.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414, 2022. 2
  55. 55.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11975–11986, 2023. 5
  56. 56.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding, 2023. 2
  57. 57.Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, et al. Lmms-eval: Reality check on the evaluation of large multimodal models. arXiv preprint arXiv:2407.12772, 2024. 5
  58. 58.Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024. 2
  59. 59.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022. 2
  60. 60.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. arXiv preprint arXiv:2307.03601, 2023. 1, 2
  61. 61.Xiaosong Zhang, Yunjie Tian, Lingxi Xie, Wei Huang, Qi Dai, Qixiang Ye, and Qi Tian. Hivit: A simpler and more efficient design of hierarchical vision transformer. In The Eleventh International Conference on Learning Representations, 2022. 2
  62. 62.Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. 1, 2, 3, 5
  63. 63.Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Wancai Zhang, Zhifeng Li, Wei Liu, and Li Yuan. Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment, 2024. 2

Citation

MLA
Tang, X., et al. “Adaptive Keyframe Sampling for Long Video Understanding”. arXiv, 2025, http://arxiv.org/abs/2502.21271v1.
APA
Tang, X., Qiu, J., Xie, L., Tian, Y., Jiao, J., & Ye, Q. (2025). Adaptive Keyframe Sampling for Long Video Understanding. arXiv. http://arxiv.org/abs/2502.21271v1
Chicago
Tang, X., J. Qiu, L. Xie, Y. Tian, J. Jiao, and Q. Ye. 2025. “Adaptive Keyframe Sampling for Long Video Understanding”. arXiv. http://arxiv.org/abs/2502.21271v1.
Harvard
Tang, X. et al. (2025) “Adaptive Keyframe Sampling for Long Video Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.21271v1.
Vancouver
1. Tang X, Qiu J, Xie L, Tian Y, Jiao J, Ye Q (2025) Adaptive Keyframe Sampling for Long Video Understanding. arXiv

BibTeX

@article{tang2025adaptive,
  title = {Adaptive Keyframe Sampling for Long Video Understanding},
  author = {Tang, Xi and Qiu, Jihao and Xie, Lingxi and Tian, Yunjie and Jiao, Jianbin and Ye, Qixiang},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.21271v1},
  eprint = {2502.21271}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE