Revealing Single Frame Bias for Video-and-Language Learning
Jie LeiTamara L. BergMohit Bansal
Reveals a pervasive static appearance bias in standard video-and-language benchmarks by showing that single-frame training paired with inference-time frame ensembling outperforms multi-frame methods, while proposing two new action-focused retrieval tasks to properly evaluate temporal reasoning.
Standard artificial intelligence systems for video and language understanding typically process multiple video frames during training to capture changes over time. However, training multi-frame models demands substantial computational power, memory, and time, which significantly increases development and infrastructure costs. This raises an important question for technology leaders: does processing full multi-frame sequences during training provide enough value to justify these high computational expenses, or can simpler, more efficient methods achieve comparable performance?
The article evaluates whether a model trained using only a single randomly selected frame per video can match or exceed the performance of complex multi-frame systems across standard video-language tasks, such as video retrieval and question answering. In doing so, it investigates whether widely used benchmark datasets rely on genuine temporal understanding or are instead dominated by static visual cues like background scenes and still objects.
To conduct this evaluation, the researchers developed an architecture called SINGULARITY. The model pairs standard vision and language encoders with a cross-modal fusion encoder. During training, the system samples only a single frame per video, drastically cutting computational overhead. During testing and deployment, the model samples multiple frames and applies an early fusion strategy, combining all visual features before making a video-level prediction. The approach was evaluated across six established video retrieval and question answering benchmarks using pre-training datasets ranging from 5.4 million to 17.3 million image-text and video-text pairs, and compared against baseline systems trained on up to 400 million examples.
The findings show that single-frame training achieves state-of-the-art results across standard benchmarks while dramatically lowering resource requirements. Pre-trained on just 5.4 million examples, the single-frame model matched or outperformed existing multi-frame models trained on dozens to hundreds of millions of samples. Expanding pre-training to 17.3 million examples pushed performance even higher, achieving top retrieval scores on benchmarks such as DiDeMo and ActivityNet Captions. In efficiency terms, the single-frame approach required up to 16 times less pre-training computation than comparable models and trained between 2.8 and 8.5 times faster during task adaptation, allowing substantially larger batch sizes on standard hardware. Furthermore, early fusion consistently outperformed traditional late fusion methods (such as score averaging), which frequently suffer from noisy individual frame predictions.
These results demonstrate a substantial static appearance bias in popular video-language benchmarks: standard tests largely evaluate whether a model recognizes objects and environments rather than actions occurring over time. For enterprise applications where static visual matching is sufficient (such as tagging or general video search), organizations can deploy lightweight single-frame architectures to achieve major cost savings and faster deployment cycles without sacrificing accuracy. However, because standard benchmarks mask shortcomings in temporal reasoning, the researchers adapted the action-focused Something-Something v2 dataset into two new retrieval benchmarks. On these motion-critical tasks, the single-frame model underperformed multi-frame baselines by significant margins (such as a 10.9-point deficit in template retrieval), confirming that single-frame training cannot replace temporal modeling when tracking true physical actions.
Organizations should adopt a split strategy based on use-case requirements. Teams focused on general video search, cataloging, and high-volume question answering should leverage single-frame training with early fusion to maximize throughput and minimize cloud compute costs. Conversely, teams building applications that depend strictly on sequence and motion (such as safety monitoring, gesture control, or fine-grained action classification) must incorporate temporal modeling modules, such as the multi-frame temporal variant introduced in the article. Researchers and evaluation teams should immediately incorporate fine-grained action benchmarks to ensure systems are tested on genuine temporal understanding rather than static visual shortcuts.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Its joint image-and-video retrieval model establishes the single-frame training and frame-count trade-offs that this paper probes in video-language learning.
- Paper: All in One: Exploring Unified Video-Language Pre-Training, Jinpeng Wang et al. (2023). Its unified video-language pretraining approach provides context for the source paper’s comparison of sparse-frame and multi-frame training.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). Its video-transformer designs and frame-count ablations clarify the computational and modeling choices behind multi-frame video inputs.
- Paper: LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding, Haoning Wu et al. (2024). Building on the source’s finding that common benchmarks permit single-frame shortcuts, LongVideoBench tests whether long-context models improve when questions require evidence across video moments.
- Paper: Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos, Chiara Plizzari et al. (2025). EgoTempo carries the source’s concern about static-frame shortcuts into egocentric video, testing questions that demand genuine temporal understanding.
