M-LLM Based Video Frame Selection for Efficient Video Understanding

Kai HuFeng GaoXiaohan NiePeng ZhouSon TranTal NeimanLingyun WangMubarak ShahRaffay HamidBing Yin

article2025CVPR113 citations

Develops a lightweight, plug-and-play video frame selector trained with spatial and temporal pseudo-labels to replace uniform sampling, boosting question-answering accuracy and efficiency for frozen multimodal LLMs on long- and medium-context video benchmarks.

Listen

Multi-modal artificial intelligence systems increasingly process video to answer complex user queries. However, standard architectures struggle with the trade-off between analyzing dense visual data and managing limited computational context windows. Most current approaches rely on uniform frame sampling, extracting images at fixed time intervals across a video. This conventional strategy frequently misses brief, decisive actions while ingesting redundant, uninformative visual data, which leads to degraded visual reasoning accuracy and inflated computational costs.

The main objective of the article is to develop and evaluate a lightweight, adaptive frame selection module that identifies the most query-relevant video frames before passing them to downstream language and vision models. By doing so, the article aims to demonstrate that targeted, question-aware frame selection significantly improves video question-answering accuracy and processing efficiency compared to standard uniform sampling.

To accomplish this, the authors engineered a plug-and-play frame selector built upon a compact language model paired with aggressive visual compression, reducing each candidate frame from standard high-token representations to just nine visual tokens. Because human-labeled frame importance data is scarce, the selector is trained using automated pseudo-labels that combine spatial relevance scores from a multimodal model with temporal reasoning from a text-based language model analyzing frame captions. The system processes an initial pool of 128 uniformly sampled frames, calculates an importance score for each frame relative to the question, and selects the top candidates using a greedy algorithm with non-maximum suppression to prevent the selection of redundant neighboring frames. The framework was evaluated across multiple standard medium- and long-context benchmarks, including ActivityNet-QA, NExT-QA, EgoSchema, VideoMME, and LongVideoBench.

The key findings show consistent performance and efficiency gains across diverse benchmarks. First, integrating the adaptive selector improved the question-answering accuracy of various leading video models, yielding gains across every tested benchmark without requiring any modifications or fine-tuning of the downstream models themselves. Second, the selector enables models to achieve equal or superior accuracy using half as many input frames; for example, downstream models processing four selectively chosen frames matched or outperformed systems using eight uniformly sampled frames. Third, on long-form video benchmarks with runtimes averaging nearly eight minutes, models utilizing selected frames consistently outperformed uniform sampling baselines across every tested frame budget. Finally, ablation studies showed that combining spatial and temporal reasoning during supervision produced markedly higher accuracy than relying on simple image-text similarity metrics.

These results demonstrate that question-aware frame selection directly addresses the computational bottleneck of long video analysis. In enterprise settings, processing fewer frames per query lowers hardware memory requirements, reduces inference latency, and decreases operational cloud computing costs. Furthermore, the plug-and-play design ensures compatibility with existing foundation models without expensive retraining.

Organizations deploying video analysis and reasoning systems should consider adopting front-end adaptive frame selectors rather than relying solely on uniform frame sampling or expanding context window sizes. For immediate implementation, teams can integrate compact, fine-tuned selectors into existing inference pipelines to cut computational loads while boosting accuracy. Future efforts should explore end-to-end joint training architectures, expand training across broader synthetic and real-world datasets, and optimize pseudo-labeling pipelines to minimize upstream training overhead.

Confidence in these findings is high, supported by consistent empirical improvements across diverse model families and standardized evaluation suites. However, decision-makers should note certain limitations: the training pipeline depends heavily on synthetic supervision generated by external large models, which may inherit hallucinations or labeling noise. Additionally, while the selector is computationally lightweight, it introduces a preliminary inference step whose latency and overhead must be factored into real-time or ultra-low-latency production environments.

Cover for M-LLM Based Video Frame Selection for Efficient Video Understanding

Abstract

Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context videos. However, it could lose crucial context in certain periods of a video, so that the downstream M-LLM may not have sufficient visual information to answer a question. To attack this pain point, we propose a light-weight M-LLM-based frame selection method that adaptively select frames that are more relevant to users' queries. In order to train the proposed frame selector, we introduce two supervision signals (i) Spatial signal, where single frame importance score by prompting an M-LLM; (ii) Temporal signal, in which multiple frames selection by prompting Large Language Model (LLM) using the captions of all frame candidates. The selected frames are then digested by a frozen downstream video M-LLM for visual reasoning and question answering. Empirical results show that the proposed M-LLM video frame selector improves the performances various downstream video Large Language Model (video-LLM) across medium (ActivityNet, NExT-QA) and long (EgoSchema, LongVideoBench) context video question answering benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Rethinking Uniform Sampling in Video LLMs
  • 3.2. Design of the Frame Selector
  • 3.3. Pseudo Labels for the Frame Selector
  • 3.4. Training of the Frame Selector
  • 4. Experiments
  • 4.1. Experiment Setup
  • 4.2. Comparison with SOTA Video-LLMs
  • 4.3. Ablation Studies
  • 4.4. Visualization of selected frames
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Architecture of M-LLM-Based Video Frame Selector

    model/method

    To select the most informative video frames for a question without running a heavy downstream video multi-modal large language model (video M-LLM) over all frames, a lightweight M-LLM-based frame selector processes a dense set of nn video frames along with the text query QQ in a single forward pass.

    Let the dense input video frames be [x1,x2,…,xn][x_1, x_2, \dots, x_n], where each frame xi∈RH×W×3x_i \in \mathbb{R}^{H \times W \times 3} has resolution H×WH \times W. A pre-trained visual encoder fvf_v extracts features from each frame, which are mapped to the hidden dimension dd of a decoder-only large language model (LLM) via an alignment projector gag_a and spatially pooled to a small number of visual tokens mm per frame:

    hi=AvgPooling(ga(fv(xi))),hi∈Rm×dh_i = \text{AvgPooling}(g_a(f_v(x_i))), \quad h_i \in \mathbb{R}^{m \times d}

    Aggressive spatial pooling (e.g., m=3×3=9m = 3 \times 3 = 9 tokens per frame, compared to 12×12=14412 \times 12 = 144 tokens in typical downstream models) preserves video context while drastically reducing sequence length.

    Let Q∈Rl×dQ \in \mathbb{R}^{l \times d} be the text embedding of the input question with length ll. A learnable score query token qscore∈R1×dq_{\text{score}} \in \mathbb{R}^{1 \times d} is appended to the sequence of visual and question tokens. The concatenated sequence is processed by the LLM backbone:

    e1,…,en,eQ,eq=LLM(h1,…,hn,Q,qscore)e_1, \dots, e_n, e_Q, e_q = \text{LLM}(h_1, \dots, h_n, Q, q_{\text{score}})

    Because of causal self-attention, the hidden state eq∈Rde_q \in \mathbb{R}^d corresponding to qscoreq_{\text{score}} at the penultimate transformer block aggregates information from all visual tokens and question tokens across both spatial and temporal dimensions. A multi-layer perceptron (MLP) score projector gsg_s maps eqe_q to an nn-dimensional importance score vector s∈Rns \in \mathbb{R}^n:

    s=gs(eq)=MLP(eq),s∈Rns = g_s(e_q) = \text{MLP}(e_q), \quad s \in \mathbb{R}^n

    where the ii-th scalar entry sis_i indicates the relevance of frame xix_i to question QQ.

  2. Knowl 2 — Greedy Non-Maximum Suppression Frame Sampling

    algorithm

    Given an nn-dimensional frame importance score vector s∈Rns \in \mathbb{R}^n predicted by the frame selector for nn densely sampled video frames, selecting the top-kk scores naively often picks redundant neighboring frames from identical temporal windows. To ensure visual diversity and coverage, a greedy selection algorithm with non-maximum suppression (NMS) suppresses neighboring indices within a window ω=⌊n/(4k)⌋\omega = \lfloor n / (4k) \rfloor each time a peak frame is chosen.

    Input: Importance score vector s∈Rns \in \mathbb{R}^n, number of frames to select kk
    Initialize neighbor suppression gap ω←⌊n/(4k)⌋\omega \leftarrow \lfloor n / (4k) \rfloor
    Initialize selected index list Is←[]I_s \leftarrow []
    for step in 1,…,k1, \dots, k do
        i←arg⁡max⁡(s)i \leftarrow \arg\max(s)
        Append index ii to IsI_s
        for each index j∈{1,…,n}j \in \{1, \dots, n\} do
            if ∣i−j∣≤ω|i - j| \le \omega then
                s[j]←−1s[j] \leftarrow -1
            end if
        end for
    end for
    Is←sort(Is)I_s \leftarrow \text{sort}(I_s)
    Return: IsI_s

    The selected indices IsI_s are then used to extract the kk key frames fed into a downstream video M-LLM for answer generation.

  3. Knowl 3 — Spatial and Temporal Pseudo-Labeling for Video Frame Importance

    model/method

    Because ground-truth frame-level importance annotations for video question answering are unavailable, pseudo-labels are generated by combining independent single-frame spatial reasoning with multi-frame temporal reasoning:

    1. Spatial Pseudo-Labels: Each of the nn uniformly sampled video frames xix_i is evaluated independently by prompting a multi-modal LLM (M-LLM, such as Qwen2-VL-7B) with the question using chain-of-thought (CoT) reasoning to produce an explanation followed by a boolean evaluation ('True' or 'False'). The normalized probability of the 'True' token is used as the raw score:

    sispatial, raw=pTruepTrue+pFalses_i^{\text{spatial, raw}} = \frac{p_{\text{True}}}{p_{\text{True}} + p_{\text{False}}}

    The spatial pseudo-label vector is then normalized by its maximum element: sispatial=sispatial, raw/max⁡jsjspatial, raws_i^{\text{spatial}} = s_i^{\text{spatial, raw}} / \max_j s_j^{\text{spatial, raw}}.

    1. Temporal Pseudo-Labels: To overcome the single-frame model's inability to observe global temporal dynamics (e.g., action sequencing), an M-LLM first generates concise captions for all nn frames. A text-only LLM (such as GPT-4o mini) then receives all nn frame captions simultaneously along with the question and outputs a discrete list of top-kk most helpful frame indices. Frames included in this list receive a score of 11, and excluded frames receive 00, forming a binary temporal pseudo-label vector stemporal∈{0,1}ns^{\text{temporal}} \in \{0, 1\}^n.

    2. Combined Supervision: To balance visual grounding with temporal context while mitigating text-caption information loss, the final supervision target for training the frame selector is the element-wise average:

    sipseudo=sispatial+sitemporal2,∀i∈{1,…,n}s_i^{\text{pseudo}} = \frac{s_i^{\text{spatial}} + s_i^{\text{temporal}}}{2}, \quad \forall i \in \{1, \dots, n\}

  4. Knowl 4 — Two-Stage Training Scheme for the Video Frame Selector

    model/method

    The lightweight frame selector (comprising a pre-trained visual encoder fvf_v, alignment projector gag_a, base LLM backbone, learnable score query qscoreq_{\text{score}}, and score projection MLP gsg_s) is trained in two stages:

    • Stage 1 (Alignment and Initialization): The vision encoder fvf_v and the base LLM backbone are kept frozen. The parameters of gag_a, qscoreq_{\text{score}}, and gsg_s are trained by alternating between two loss functions:

      1. Visual Instruction Following: Standard autoregressive cross-entropy loss on the downstream text answer given ground-truth QA pairs, aligning visual embeddings into the LLM embedding space.
      2. Importance Score Prediction: Binary cross-entropy (BCE) loss between the predicted frame score vector s∈Rns \in \mathbb{R}^n and the combined pseudo-label vector spseudo∈[0,1]ns^{\text{pseudo}} \in [0, 1]^n, providing an initial score projection capability.
    • Stage 2 (Task-Specific LoRA Adaptation): The selector is trained exclusively on the Importance Score Prediction task with BCE loss against spseudos^{\text{pseudo}}. In addition to updating gag_a, qscoreq_{\text{score}}, and gsg_s, Low-Rank Adaptation (LoRA) parameters are inserted into the LLM backbone and optimized to specialize its multi-modal reasoning for frame selection.

  5. Knowl 5 — Video Question Answering Performance Across Benchmarks

    data/table

    Equipping frozen downstream video and multi-image M-LLMs with frames chosen by the lightweight frame selector (1.5B1.5\text{B} parameters) consistently improves question-answering accuracy across short-, medium-, and long-context benchmarks compared to standard uniform frame sampling with the same number of downstream frames.

    Model Model Size ActivityNet-QA NExT-QA EgoSchema VideoMME
    (Acc / Corr) (Acc %) (Acc %) (Avg Acc %)
    PLLaVA 7B 56.3 / 3.5 - - -
    PLLaVA + Selector 7B + 1.5B 57.6 / 3.5 - - -
    PLLaVA 34B 60.9 / 3.7 - - -
    PLLaVA + Selector 34B + 1.5B 62.3 / 3.6 - - -
    LLaVA-NeXT-Video 7B 53.5 / 3.2 62.4 45.8 -
    LLaVA-NeXT-Video + Selector 7B + 1.5B 55.1 / 3.4 63.4 47.2 -
    LLaVA-NeXT-Video 34B 58.8 / 3.4 68.1 48.6 -
    LLaVA-NeXT-Video + Selector 34B + 1.5B 60.2 / 3.5 69.3 50.6 -
    Idefics2 8B - 68.0 56.6 -
    Idefics2 + Selector 8B + 1.5B - 69.1 57.9 -
    Qwen2-VL 7B - 77.6 64.6 58.1
    Qwen2-VL + Selector 7B + 1.5B - 78.4 65.9 58.7

    On ActivityNet-QA, GPT-3.5 evaluates accuracy (percentage) and correctness score (scale 1–5). On NExT-QA, EgoSchema (5-choice), and VideoMME, metrics reflect question-answering accuracy. In all cases, the selector improves zero-shot downstream reasoning without updating the downstream M-LLM weights.

  6. Knowl 6 — Inference Latency and Frame Efficiency Comparison

    data/table

    Evaluating LLaVA-NeXT-Video 34B on NExT-QA and LongVideoBench demonstrates that using the frame selector allows the downstream model to achieve equal or superior accuracy with half the number of frames (nn selected frames vs. 2n2n uniform frames), leading to lower overall inference latency on an NVIDIA A100 GPU (measured in float16 precision with batch size 1).

    Sampling Configuration NExT-QA NExT-QA NExT-QA NExT-QA NExT-QA LongVideoBench
    Acc@C Acc@T Acc@D Acc (%) Speed (s) Acc (%)
    Uniform 4 frames 67.2 61.2 73.9 66.4 0.56 45.3
    Uniform 8 frames 68.7 62.5 76.9 68.1 0.92 46.9
    Uniform 16 frames 69.1 63.6 76.8 68.7 1.71 48.1
    Uniform 32 frames 69.5 64.3 78.4 69.3 3.40 49.7
    Selector 128→4128 \to 4 frames 68.5 64.5 75.7 68.5 0.76 49.5
    Selector 128→8128 \to 8 frames 69.3 64.9 77.5 69.3 1.12 49.9
    Selector 128→16128 \to 16 frames 69.4 64.8 78.5 69.5 1.91 49.8
    Selector 128→32128 \to 32 frames 69.2 65.6 78.7 69.6 3.50 50.0

    On NExT-QA, selecting 4 frames from 128 dense candidates (128→4128 \to 4) achieves 68.5%68.5\% total accuracy in 0.76 s0.76\text{ s}, outperforming uniform 8-frame sampling (68.1%68.1\% at 0.92 s0.92\text{ s}). On LongVideoBench (average video duration 473 s473\text{ s}), selector 128→4128 \to 4 (49.5%49.5\%) outperforms uniform 16 frames (48.1%48.1\%).

  7. Knowl 7 — Ablation of Pseudo-Labeling and Selection Methods

    data/table

    An ablation study evaluated on LLaVA-NeXT-Video 7B compares downstream video QA accuracy on ActivityNet-QA and NExT-QA when using different methods to assign frame importance scores prior to greedy NMS sampling.

    Selection Method ActivityNet-QA (Acc %) NExT-QA (Acc %)
    Uniform sampling 53.5 62.4
    CLIP image-text similarity 53.7 62.2
    SeViLA pseudo-labels 54.0 63.2
    Spatial pseudo-labels only 54.2 63.6
    Spatial temporal pseudo-labels 55.5 63.9
    Trained lightweight selector 55.1 63.4

    Simple similarity (CLIP) offers minimal gain over uniform sampling. Spatial pseudo-labels with chain-of-thought outperform single-frame score approaches like SeViLA. Combining spatial and temporal pseudo-labels gives the highest score, and the trained lightweight selector distills this supervision to achieve near-oracle pseudo-label performance (55.1%55.1\% vs 55.5%55.5\%, 63.4%63.4\% vs 63.9%63.9\%) at a fraction of the computational inference cost.

  8. Knowl 8 — Ablation of Frame Selector Visual Tokens and LLM Backbone Size

    data/table

    Ablations on LLaVA-NeXT-Video 7B demonstrate that the frame selector can operate with minimal visual tokens per frame and a small LLM backbone without compromising downstream video question answering accuracy.

    Ablated Factor Value ActivityNet-QA NExT-QA EgoSchema
    Tokens per Frame No selector (Uniform) 53.5 62.4 45.8
    1 token 53.2 62.7 46.6
    9 tokens (3×33 \times 3) 55.1 63.4 47.2
    25 tokens (5×55 \times 5) 55.3 63.6 47.3
    Selector LLM Size No selector (Uniform) 53.5 62.4 45.8
    0.5B 53.8 62.8 46.4
    1.5B 55.1 63.4 47.2
    7B 55.5 64.0 47.9

    Reducing visual representation to 9 tokens per frame (3×33 \times 3 pooling) provides nearly the same accuracy as 25 tokens while reducing token length. Similarly, scaling the selector LLM backbone from 0.5B to 1.5B brings significant improvements, while moving to 7B provides diminishing returns at higher computational cost, establishing 9 tokens and a 1.5B backbone as the optimal efficiency-accuracy trade-off.

Coverage note — None was omitted; all key architectural components, pseudo-label generation algorithms, training stages, benchmark comparisons, efficiency analyses, and ablation studies are fully represented.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 technical report. arXiv:2303.08774, 2023. 2
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022. 1, 2
  3. 3.Anthropic. The claude 3 model family: Opus, sonnet, haiku. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf, 2024. Accessed: 2024-09-18. 1, 2
  4. 4.Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens. arXiv preprint arXiv:2404.03413, 2024. 1, 6
  5. 5.Shyamal Buch, Cristobal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles. Revisiting the” video” in video-language understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2917–2927, 2022. 2
  6. 6.Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024. 2
  7. 7.Ran Cui, Tianwen Qian, Pai Peng, Elena Daskalaki, Jingjing Chen, Xiaowei Guo, Huyang Sun, and Yu-Gang Jiang. Video moment retrieval from text queries via single frame annotation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 1033–1043, 2022. 2
  8. 8.Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. Mmbench-video: A long-form multi-shot benchmark for holistic video understanding. Advances in Neural Information Processing Systems, 37:89098–89124, 2024. 7
  9. 9.Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075, 2024. 6
  10. 10.Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, and Mike Zheng Shou. Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14773–14783, 2023. 2
  11. 11.Songhao Han, Wei Huang, Hairong Shi, Le Zhuo, Xiu Su, Shifeng Zhang, Xu Zhou, Xiaojuan Qi, Yue Liao, and Si Liu. Videoespresso: A large-scale chain-of-thought dataset for fine-grained video reasoning via core frame selection. arXiv preprint arXiv:2411.14794, 2024. 3
  12. 12.Wei Han, Hui Chen, Min-Yen Kan, and Soujanya Poria. Self-adaptive sampling for accurate video question answering on image text models. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2522–2534, 2024. 3
  13. 13.Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan. Chat-univi: Unified visual representation empowers large language models with image and video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13700–13710, 2024. 6
  14. 14.Sungdong Kim, Jin-Hwa Kim, Jiyoung Lee, and Minjoon Seo. Semi-parametric video-grounded text generation. arXiv preprint arXiv:2301.11507, 2023. 2
  15. 15.Hugo Laurençon, Leo Tronchon, Matthieu Cord, and Victor Sanh. What matters when building vision-language models? arXiv preprint arXiv:2405.02246, 2024. 2, 6, 7
  16. 16.Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 1
  17. 17.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023. 2
  18. 18.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023. 3
  19. 19.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, Limin Wang, and Yu Qiao. MVBench: A comprehensive multi-modal video understanding benchmark. arXiv:2311.17005, 2023. 2
  20. 20.Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024. 6
  21. 21.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision, pages 323–340. Springer, 2025. 2, 3, 6
  22. 22.Jianxin Liang, Xiaojun Meng, Yueqian Wang, Chang Liu, Qun Liu, and Dongyan Zhao. End-to-end video question answering with frame scoring mechanisms and adaptive sampling. arXiv preprint arXiv:2407.15047, 2024. 3
  23. 23.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023. 2, 6
  24. 24.Haogeng Liu, Qihang Fan, Tingkai Liu, Linjie Yang, Yunzhe Tao, Huaibo Huang, Ran He, and Hongxia Yang. Video-teller: Enhancing cross-modal generation with fusion and decoupling. arXiv preprint arXiv:2310.04991, 2023. 2
  25. 25.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 2
  26. 26.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. LLaVA-NeXT: Improved reasoning, ocr, and world knowledge, 2024. 1, 2
  27. 27.Yuqi Liu, Pengfei Xiong, Luhui Xu, Shengming Cao, and Qin Jin. Ts2-net: Token shift and selection transformer for text-video retrieval. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XIV, pages 319–335. Springer, 2022. 2
  28. 28.Yuan Liu, Zhongyin Zhao, Ziyuan Zhuang, Le Tian, Xiao Zhou, and Jie Zhou. Points: Improving your vision-language model with affordable strategies. arXiv preprint arXiv:2409.04828, 2024. 2
  29. 29.Haoyu Lu, Mingyu Ding, Nanyi Fei, Yuqi Huo, and Zhiwu Lu. LGDN: Language-guided denoising network for video-language modeling. In Advances in Neural Information Processing Systems, 2022. 2
  30. 30.Ruipu Luo, Ziwang Zhao, Min Yang, Junwei Dong, Da Li, Pengcheng Lu, Tao Wang, Linmei Hu, Minghui Qiu, and Zhongyu Wei. Valley: Video assistant with large language model enhanced ability. arXiv preprint arXiv:2306.07207, 2023. 2
  31. 31.Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 2024. 2, 6
  32. 32.Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. Egoschema: A diagnostic benchmark for very long-form video language understanding. Advances in Neural Information Processing Systems, 36:46212–46244, 2023. 6
  33. 33.Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al. MM1: Methods, analysis & insights from multimodal llm pre-training. arXiv:2403.09611, 2024. 2
  34. 34.TB OpenAI. Chatgpt: Optimizing language models for dialogue. OpenAI, 2022. 1
  35. 35.Jongwoo Park, Kanchana Ranasinghe, Kumara Kahatapitiya, Wonjeong Ryoo, Donghyun Kim, and Michael S Ryoo. Too many frames, not all useful: Efficient strategies for long-form video qa. arXiv preprint arXiv:2406.09396, 2024. 2
  36. 36.Tianwen Qian, Ran Cui, Jingjing Chen, Pai Peng, Xiaowei Guo, and Yu-Gang Jiang. Locate before answering: Answer guided question localization for video question answering. In IEEE transactions on multimedia, 2022. 2
  37. 37.Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya, and Michael S Ryoo. Understanding long videos in one multimodal language model pass. arXiv preprint arXiv:2403.16998, 2024. 3, 4, 12
  38. 38.Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. Timechat: A time-sensitive multimodal large language model for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14313–14323, 2024. 6
  39. 39.Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Xun Guo, Tian Ye, Yan Lu, Jenq-Neng Hwang, et al. MovieChat: From dense token to sparse memory for long video understanding. arXiv:2307.16449, 2023. 2
  40. 40.Reuben Tan, Ximeng Sun, Ping Hu, Jui-hsien Wang, Hanieh Deilamsalehy, Bryan A Plummer, Bryan Russell, and Kate Saenko. Koala: Key frame-conditioned long video-llm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13581–13591, 2024. 3
  41. 41.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv:2312.11805, 2023. 1, 2
  42. 42.Qwen Team. Qwen2.5: A party of foundation models, 2024. 6
  43. 43.Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. arXiv preprint arXiv:2406.16860, 2024. 2
  44. 44.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023. 1, 2
  45. 45.Jiawei Wang, Liping Yuan, and Yuchen Zhang. Tarsier: Recipes for training and evaluating large video description models. arXiv preprint arXiv:2407.00634, 2024. 6, 7
  46. 46.Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 1, 6, 7
  47. 47.Xijun Wang, Junbang Liang, Chun-Kai Wang, Kenan Deng, Yu Lou, Ming C Lin, and Shan Yang. Vila: Efficient video-language alignment for video question answering. In European Conference on Computer Vision, pages 186–204. Springer, 2024. 3
  48. 48.Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy. Videoagent: Long-form video understanding with large language model as agent, 2024. 3
  49. 49.Zixu Wang, Yujie Zhong, Yishu Miao, Lin Ma, and Lucia Specia. Contrastive video-language learning with fine-grained frame sampling. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 694–705. Association for Computational Linguistics, 2022. 2
  50. 50.Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon, Feng Cheng, Gedas Bertasius, and Mohit Bansal. VideoTree: Adaptive tree-based video representation for llm reasoning on long videos. arXiv:2405.19209, 2024. 3
  51. 51.Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. Vary: Scaling up the vision vocabulary for large vision-language model. In European Conference on Computer Vision, pages 408–424. Springer, 2025. 2
  52. 52.Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, and Bohan Zhuang. Longvlm: Efficient long video understanding via large language models. In European Conference on Computer Vision, pages 453–470. Springer, 2025. 2, 8
  53. 53.Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long-context interleaved video-language understanding. arXiv preprint arXiv:2407.15754, 2024. 6
  54. 54.Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786, 2021. 6
  55. 55.Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. PLLaVA: Parameter-free llava extension from images to videos for video dense captioning. arXiv:2404.16994, 2024. 1, 2, 3, 6
  56. 56.Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen, Zhengfeng Lai, Haiming Gang, Kai Kang, and Afshin Dehghan. Slowfast-llava: A strong training-free baseline for video large language models. arXiv preprint arXiv:2407.15841, 2024. 6, 7
  57. 57.Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal. Self-chained image-language model for video localization and question answering. Advances in Neural Information Processing Systems, 36, 2024. 2, 3, 4, 7, 12
  58. 58.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. Activitynet-qa: A dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 9127–9134, 2019. 6
  59. 59.Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao. ActivityNet-QA: A dataset for understanding complex web videos via question answering. In AAAI, 2019. 1
  60. 60.Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. arXiv preprint arXiv:2303.15343, 2023. 6
  61. 61.Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava-next: A strong zero-shot video understanding model, 2024. 1, 3
  62. 62.Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. Llava-next: A strong zero-shot video understanding model, 2024. 6, 7
  63. 63.Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Ziwei Liu, and Chunyuan Li. Video instruction tuning with synthetic data, 2024. 6

Citation

MLA
Hu, K., et al. “M-LLM Based Video Frame Selection for Efficient Video Understanding”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 13702–12, https://doi.org/10.1109/CVPR52734.2025.01279.
APA
Hu, K., Gao, F., Nie, X., Zhou, P., Tran, S., Neiman, T., Wang, L., Shah, M., Hamid, R., Yin, B., & Chilimbi, T. (2025). M-LLM Based Video Frame Selection for Efficient Video Understanding. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13702–13712. https://doi.org/10.1109/CVPR52734.2025.01279
Chicago
Hu, K., F. Gao, X. Nie, et al. 2025. “M-LLM Based Video Frame Selection for Efficient Video Understanding”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13702–12. https://doi.org/10.1109/CVPR52734.2025.01279.
Harvard
Hu, K. et al. (2025) “M-LLM Based Video Frame Selection for Efficient Video Understanding”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 13702–13712. Available at: https://doi.org/10.1109/CVPR52734.2025.01279.
Vancouver
1. Hu K, Gao F, Nie X, et al (2025) M-LLM Based Video Frame Selection for Efficient Video Understanding. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 13702–13712

BibTeX

@inproceedings{Hu_2025, title={M-LLM Based Video Frame Selection for Efficient Video Understanding}, url={http://dx.doi.org/10.1109/CVPR52734.2025.01279}, DOI={10.1109/cvpr52734.2025.01279}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Hu, Kai and Gao, Feng and Nie, Xiaohan and Zhou, Peng and Tran, Son and Neiman, Tal and Wang, Lingyun and Shah, Mubarak and Hamid, Raffay and Yin, Bing and Chilimbi, Trishul}, year={2025}, month=June, pages={13702–13712} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE