Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
Zhenhailong WangManling LiRuochen XuLuowei ZhouJie LeiXudong LinShuohang WangZiyi YangChenguang ZhuDerek Hoiem
Proposes VidIL, a framework that decomposes video content into multi-level textual descriptions via frozen image-language models, enabling large language models to perform generative video tasks with few-shot prompting without requiring any video pretraining or finetuning.
Artificial intelligence systems often struggle to generalize to new video understanding tasks when provided with only a few annotated examples. Current video-language models typically focus solely on encoding visual data without generating text, or they rely heavily on computationally expensive pretraining and finetuning over millions of video-text pairs. Moreover, conventional video models struggle to bridge the gap between noisy speech transcripts and rich visual scenes across time.
The article evaluates a modular framework, named VidIL, to demonstrate that pretrained large language models can perform diverse video-to-text generative tasks using only a few examples, completely eliminating the need for video-specific pretraining or finetuning.
The approach converts video content into a structured, unified textual format across three hierarchical tiers: visual tokens (identifying objects, events, and attributes via image encoders), frame-level captions, and video-level summaries. These elements are arranged chronologically using temporal transition markers (such as "First," "Then," and "Finally") and combined with optional speech transcripts. A frozen language model receives this structured text alongside a small set of dynamically selected in-context examples to generate outputs across multiple benchmarks, including open-domain and instructional video datasets.
Key findings demonstrate the effectiveness and efficiency of this approach:
- In video future event prediction, the framework achieved 72.0% accuracy with only 10 labeled examples, outperforming fully supervised models trained on over 20,000 video instances (which achieved 68.4%).
- In few-shot video question answering, the 5-shot model achieved 21.2% accuracy on MSR-VTT and 39.1% on MSVD, surpassing zero-shot baselines by large margins and outperforming heavy models like Flamingo-3B across multiple shot settings.
- In video captioning, incorporating speech transcripts lifted instructional captioning quality from a 27.0 baseline score to 111.6, demonstrating robust cross-domain flexibility.
- In semi-supervised text-video retrieval, using the model to generate synthetic labels on unlabeled videos improved retrieval recall across open-domain benchmarks, achieving results comparable to models trained on fully human-annotated data.
These results show that organizations can bypass expensive, specialized video pretraining pipelines by chaining off-the-shelf image models with general-purpose language models. This substantially cuts training compute costs, accelerates deployment timelines, and simplifies the integration of multi-modal streams like audio and text transcripts into existing workflows.
Organizations developing video analytics should adopt modular prompting architectures when tackling low-resource or rapid-adaptation video tasks. Teams should also implement dynamic in-context example selection rather than static prompting to maximize output accuracy. Where large pools of unlabeled video exist, teams can employ this framework as an automated pseudo-labeling tool to bootstrap downstream retrieval models.
Confidence in these findings is strong across generative, question answering, and retrieval benchmarks. However, leaders should note that converting visual data into pure text inevitably discards fine-grained spatial and low-level visual details, making this approach less suitable for specialized spatial localization tasks. Furthermore, because the framework relies on large language models trained on massive internet data, organizations must implement safeguards to monitor and mitigate potential social biases in generated outputs.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). Provides the foundational paradigm of pairing narrated instructional video with textual transcripts for scalable multimodal learning that VidIL adapts.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Introduces joint image-video representations and retrieval benchmarks that serve as the baseline foundations for VidIL's video-to-text generative framework.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). Establishes the methodology of converting continuous video into discrete visual tokens aligned with language models.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). Supplies foundational cross-modal region-text alignment principles used by image descriptor models to extract discrete visual tokens.
- Paper: LXMERT: Learning Cross-Modality Encoder Representations from Transformers, Hao Tan et al. (2019). Pioneers the extraction of visual entities and cross-modality language reasoning that underlies modular descriptor-to-text prompting.
- Paper: Sequence to Sequence -- Video to Text, Subhashini Venugopalan et al. (2015). Introduces sequence-to-sequence video captioning on the MSVD benchmark, which VidIL aims to solve in a few-shot, tuning-free manner.
- Paper: VL-BERT: Pre-training of Generic Visual-Linguistic Representations, Weijie Su et al. (2019). Establishes generic visual-linguistic pretraining architectures for interpreting grounded visual elements with Transformer backbones.
- Paper: Unsupervised Learning of Video Representations using LSTMs, Nitish Srivastava et al. (2015). Presents foundational temporal modeling and future frame prediction concepts evaluated under VidIL's few-shot event prediction setup.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). Extends zero-shot and few-shot video understanding from descriptor-prompted LLMs to end-to-end instruction-tuned conversational video-language models.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). Builds on multimodal video-language prompting by integrating dedicated audio and visual adapters directly into frozen LLMs for conversational interaction.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Advances the concept of unifying visual domains before LLM projection by aligning image and video representations in a shared feature space.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). Continues the idea of discrete visual tokenization for LLMs by decoupling static keyframes and temporal motion tokens for unified video-language pre-training.
- Paper: Video ReCap: Recursive Captioning of Hour-Long Videos, Md Mohaiminul Islam et al. (2024). Generalizes VidIL's multi-tiered video captioning hierarchy to recursively generate clip, segment, and summary captions for long-form video.
- Paper: VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset, Sihan Chen et al. (2023). Expands VidIL's integration of speech transcripts and visual descriptions into a scaled omni-modality foundation model encompassing vision, audio, and subtitles.
- Paper: Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models, Wenhao Wu et al. (2023). Explores bidirectional video-text alignment and text phrase association to enhance few-shot and zero-shot video recognition.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). Refines the visual descriptor pipeline by using language models to adaptively select the most relevant video frames before multimodal processing.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). Establishes a comprehensive multi-task benchmark specifically addressing the dynamic temporal evaluation challenges highlighted in modular video-LLM systems.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Provides a comprehensive multimodal LLM video evaluation benchmark across diverse domains, durations, audio, and subtitle streams.
