Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Max BainArsha NagraniGül VarolAndrew Zisserman
Proposes an end-to-end spatiotemporal visual architecture trained jointly on static images and video captions alongside the WebVid-2M dataset, achieving state-of-the-art text-to-video retrieval performance with an order of magnitude less training data.
Video search systems are increasingly critical across commercial and enterprise platforms, yet training machine learning models to accurately match text queries to relevant videos remains computationally expensive and data-inefficient. Existing techniques rely heavily on complex combinations of pre-extracted features or require pretraining on massive, noisy instructional video datasets that consume immense computing power. Furthermore, video and image retrieval models are traditionally developed on separate tracks despite substantial overlaps in the visual information they convey.
The main objective of the article is to design, train, and evaluate a unified, end-to-end dual-encoder retrieval architecture that flexibly learns from both image and video caption datasets. By treating static images as single-frame video snapshots frozen in time, the article demonstrates how joint training can dramatically improve retrieval accuracy while lowering computational requirements.
The authors implemented a visual transformer architecture with divided space-time attention alongside a lightweight text encoder, mapping visual and text inputs directly into a shared embedding space. To support this approach, they curated a new pretraining dataset called WebVid-2M containing 2.5 million video-text pairs with well-formed, visually aligned captions. They evaluated their approach by pretraining the model on combinations of WebVid-2M and image caption datasets using a progressive training schedule that increases frame counts over time, followed by finetuning across standard industry video retrieval benchmarks, including MSR-VTT, MSVD, DiDeMo, and LSMDC.
The key findings show that the proposed unified model consistently outperforms established benchmarks. First, the model achieves state-of-the-art text-to-video retrieval accuracy across multiple benchmarks while relying exclusively on visual data, outperforming complex systems that use multiple pre-extracted expert features and audio signals. Second, high-quality, visually aligned pretraining data delivers superior results compared to larger datasets; pretraining on 5.5 million combined image-video pairs outperformed systems trained on uncurated datasets twenty times larger. Third, the progressive curriculum training schedule cut computing time by roughly one-half to two-thirds while matching or exceeding the accuracy of models trained on full-frame sequences from the beginning. Finally, the model demonstrated strong zero-shot retrieval capabilities out of the box without target-dataset finetuning.
These findings have direct practical implications for operational cost, computational resource management, and system architecture. Because the model maps video and text into independent embeddings, search indexing scales linearly rather than quadratically at runtime, enabling fast approximate nearest-neighbor search for large-scale video catalogs. Organizations can reduce pretraining infrastructure expenses by prioritizing smaller, higher-quality datasets and adopting joint image-video training workflows instead of building separate pipelines.
Based on these results, decision-makers deploying enterprise video search should transition toward joint vision-language encoders and adopt progressive frame training to optimize GPU utilization. When curating training data, engineering teams should prioritize diverse, well-aligned caption pairs from multiple sources over scaling single-source datasets, as results indicate diminishing returns when scaling a single domain. Further investigation should explore combining the model with multi-dataset training mixtures to capture even broader visual distributions.
Confidence in these findings is high across standard academic video retrieval benchmarks, supported by extensive ablation studies. However, decision-makers should note that evaluations were conducted primarily on short video clips under research benchmark conditions. Deployment to real-world industrial environments with long-form video or highly domain-specific video content may require localized pilot testing and further domain adaptation.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). This paper introduces TimeSformer and divided space-time attention, which directly provides the foundational visual transformer architecture that Frozen in Time adapts for joint video and image encoding.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). This work establishes the MSR-VTT dataset, which serves as one of the primary downstream benchmarks for evaluating text-to-video retrieval in Frozen in Time.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). This seminal paper introduces joint self-supervised pre-training across video and language using Transformer architectures, establishing the conceptual paradigm of multimodal video-text representations.
- Paper: UNITER: UNiversal Image-TExt Representation Learning, Yen-Chun Chen et al. (2020). This work presents foundational multimodal transformer pre-training techniques for joint vision-and-language representations that inspire subsequent end-to-end retrieval frameworks.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). This work directly utilizes the WebVid-2M dataset curated by Frozen in Time to train its video transformer adapters for conversational video understanding.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). This paper advances joint image and video representation learning by unifying multimodal alignment before projection into large language models.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). This study extends the paradigm of unified image-video training to large multimodal models through transfer learning and joint representation across visual modalities.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This paper builds upon the principle of joint image and video pre-training to establish generative video diffusion architectures.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). This work develops a comprehensive multi-dimensional benchmark suite to evaluate video-text alignment and generation quality in modern video models.
