Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Max BainArsha NagraniGül VarolAndrew Zisserman

article2021ICCV1,664 citations

Proposes an end-to-end spatiotemporal visual architecture trained jointly on static images and video captions alongside the WebVid-2M dataset, achieving state-of-the-art text-to-video retrieval performance with an order of magnitude less training data.

Listen

Video search systems are increasingly critical across commercial and enterprise platforms, yet training machine learning models to accurately match text queries to relevant videos remains computationally expensive and data-inefficient. Existing techniques rely heavily on complex combinations of pre-extracted features or require pretraining on massive, noisy instructional video datasets that consume immense computing power. Furthermore, video and image retrieval models are traditionally developed on separate tracks despite substantial overlaps in the visual information they convey.

The main objective of the article is to design, train, and evaluate a unified, end-to-end dual-encoder retrieval architecture that flexibly learns from both image and video caption datasets. By treating static images as single-frame video snapshots frozen in time, the article demonstrates how joint training can dramatically improve retrieval accuracy while lowering computational requirements.

The authors implemented a visual transformer architecture with divided space-time attention alongside a lightweight text encoder, mapping visual and text inputs directly into a shared embedding space. To support this approach, they curated a new pretraining dataset called WebVid-2M containing 2.5 million video-text pairs with well-formed, visually aligned captions. They evaluated their approach by pretraining the model on combinations of WebVid-2M and image caption datasets using a progressive training schedule that increases frame counts over time, followed by finetuning across standard industry video retrieval benchmarks, including MSR-VTT, MSVD, DiDeMo, and LSMDC.

The key findings show that the proposed unified model consistently outperforms established benchmarks. First, the model achieves state-of-the-art text-to-video retrieval accuracy across multiple benchmarks while relying exclusively on visual data, outperforming complex systems that use multiple pre-extracted expert features and audio signals. Second, high-quality, visually aligned pretraining data delivers superior results compared to larger datasets; pretraining on 5.5 million combined image-video pairs outperformed systems trained on uncurated datasets twenty times larger. Third, the progressive curriculum training schedule cut computing time by roughly one-half to two-thirds while matching or exceeding the accuracy of models trained on full-frame sequences from the beginning. Finally, the model demonstrated strong zero-shot retrieval capabilities out of the box without target-dataset finetuning.

These findings have direct practical implications for operational cost, computational resource management, and system architecture. Because the model maps video and text into independent embeddings, search indexing scales linearly rather than quadratically at runtime, enabling fast approximate nearest-neighbor search for large-scale video catalogs. Organizations can reduce pretraining infrastructure expenses by prioritizing smaller, higher-quality datasets and adopting joint image-video training workflows instead of building separate pipelines.

Based on these results, decision-makers deploying enterprise video search should transition toward joint vision-language encoders and adopt progressive frame training to optimize GPU utilization. When curating training data, engineering teams should prioritize diverse, well-aligned caption pairs from multiple sources over scaling single-source datasets, as results indicate diminishing returns when scaling a single domain. Further investigation should explore combining the model with multi-dataset training mixtures to capture even broader visual distributions.

Confidence in these findings is high across standard academic video retrieval benchmarks, supported by extensive ablation studies. However, decision-makers should note that evaluations were conducted primarily on short video clips under research benchmark conditions. Deployment to real-world industrial environments with long-form video or highly domain-specific video content may require localized pilot testing and further domain adaptation.

Cover for Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

Abstract

Our objective in this work is video-text retrieval - in particular a joint embedding that enables efficient text-to-video retrieval. The challenges in this area include the design of the visual architecture and the nature of the training data, in that the available large scale video-text training datasets, such as HowTo100M, are noisy and hence competitive performance is achieved only at scale through large amounts of compute. We address both these challenges in this paper. We propose an end-to-end trainable model that is designed to take advantage of both large-scale image and video captioning datasets. Our model is an adaptation and extension of the recent ViT and Timesformer architectures, and consists of attention in both space and time. The model is flexible and can be trained on both image and video text datasets, either independently or in conjunction. It is trained with a curriculum learning schedule that begins by treating images as 'frozen' snapshots of video, and then gradually learns to attend to increasing temporal context when trained on video datasets. We also provide a new video-text pretraining dataset WebVid-2M, comprised of over two million videos with weak captions scraped from the internet. Despite training on datasets that are an order of magnitude smaller, we show that this approach yields state-of-the-art results on standard downstream video-retrieval benchmarks including MSR-VTT, MSVD, DiDeMo and LSMDC.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Model Architecture
  • 3.2 Training Strategy
  • 4 Experiments
  • 4.1 Pretraining Datasets
  • 4.2 Downstream Datasets
  • 4.3 Implementation Details
  • 4.4 Ablation Study
  • 4.5 Curriculum strategy
  • 4.6 Comparison to the State of the Art
  • 5 Extension: Scaling up Further
  • 6 Conclusion
  • A Additional Benchmark Results
  • B Architectural Details
  • B.1 Video Encoder
  • B.2 Text Encoder
  • C Architectural Ablations
  • C.1 Video Backbone
  • C.2 Text Backbone
  • C.3 Space-Time Attention
  • C.4 Temporal Expansion
  • D WebVid-2M Dataset Details
  • E WebVid-10M Extension
  • References

Knowls

  1. Knowl 1 — Space-Time Transformer Dual Encoder Architecture for Video-Text Retrieval

    model/method

    The Frozen in Time dual encoder maps video clips (or static images) and text descriptions into a shared semantic embedding space via separate encoders.

    Visual Space-Time Transformer Encoder:

    1. Patch Extraction: An input video clip X∈RM×3×H×W\boldsymbol{X} \in \mathbb{R}^{M \times 3 \times H \times W} consisting of MM frames with spatial dimensions H×WH \times W (where M=1M = 1 for static images) is partitioned into M×NM \times N non-overlapping spatial patches of size P×PP \times P, where N=HW/P2N = HW / P^2. For standard resolution 224×224224 \times 224 with P=16P = 16, N=196N = 196.
    2. Linear Projection: Patches x∈RM×N×3×P×P\boldsymbol{x} \in \mathbb{R}^{M \times N \times 3 \times P \times P} are projected to dimension D=768D = 768 via a 2D convolutional layer with kernel size PP and stride PP, producing tokens z∈RMN×D\boldsymbol{z} \in \mathbb{R}^{MN \times D}.
    3. Positional Encoding: Spatial positional embeddings Es∈RN×D\boldsymbol{E}^s \in \mathbb{R}^{N \times D} and temporal positional embeddings Et∈RM×D\boldsymbol{E}^t \in \mathbb{R}^{M \times D} are added to patch embeddings at spatial index p∈{1,…,N}p \in \{1, \dots, N\} and frame index m∈{1,…,M}m \in \{1, \dots, M\}: zp,m(0)=zp,m+Eps+Emt\boldsymbol{z}^{(0)}_{p,m} = \boldsymbol{z}_{p,m} + \boldsymbol{E}^s_p + \boldsymbol{E}^t_m A learned classification token [CLS]∈R1×D[\text{CLS}] \in \mathbb{R}^{1 \times D} is prepended to the sequence.
    4. Divided Space-Time Attention Stack: The sequence is processed through a stack of ∣ℓ∣=12|\ell| = 12 space-time attention blocks (with 12 heads). In each block ℓ\ell, temporal self-attention is applied across tokens at the same spatial location across all frames, followed by spatial self-attention across tokens within the same frame. A modified residual connection connects the input of the block directly to the output of the spatial self-attention layer (before the feed-forward MLP).
    5. Video Embedding: The [CLS][\text{CLS}] token from the final block is projected through a linear layer to produce the visual embedding x∈Rdproj\boldsymbol{x} \in \mathbb{R}^{d_{\text{proj}}} with dproj=256d_{\text{proj}} = 256, normalized to unit length.

    Text Transformer Encoder: Text captions are processed using a multi-layer bidirectional transformer (DistilBERT base-uncased). The [CLS][\text{CLS}] output of the final layer is projected via a linear layer to y∈Rdproj\boldsymbol{y} \in \mathbb{R}^{d_{\text{proj}}} (dproj=256d_{\text{proj}} = 256) and normalized to unit length.

    Similarity and Retrieval Complexity: The similarity score between a video and text is computed via the dot product x⊤y\boldsymbol{x}^\top \boldsymbol{y}. Because the two pathways are completely decoupled, gallery video representations can be pre-indexed, enabling text-to-video retrieval over tt text queries and vv gallery videos with computational complexity O(t+v)O(t + v), supporting fast approximate nearest neighbor search.

  2. Knowl 2 — Bidirectional Video-Text Symmetric Contrastive Objective

    equation

    The dual encoder is trained end-to-end on paired video-text and image-text batches using a bidirectional InfoNCE-style contrastive loss. For a mini-batch of BB paired visual and textual instances, where normalized visual embeddings are xi∈Rd\boldsymbol{x}_i \in \mathbb{R}^d and normalized text embeddings are yj∈Rd\boldsymbol{y}_j \in \mathbb{R}^d with ∥xi∥2=∥yj∥2=1\|\boldsymbol{x}_i\|_2 = \|\boldsymbol{y}_j\|_2 = 1, matching pairs (i=j)(i = j) are treated as positive instances and all mismatched pairs (i≠j)(i \neq j) are treated as negative instances.

    The total training loss is the sum of the video-to-text loss Lv2tL_{\text{v2t}} and the text-to-video loss Lt2vL_{\text{t2v}}:

    L=Lv2t+Lt2vL = L_{\text{v2t}} + L_{\text{t2v}}

    where:

    Lv2t=−1B∑i=1Blog⁡(exp⁡(xi⊤yi/σ)∑j=1Bexp⁡(xi⊤yj/σ))L_{\text{v2t}} = -\frac{1}{B} \sum_{i=1}^B \log \left( \frac{\exp(\boldsymbol{x}_i^\top \boldsymbol{y}_i / \sigma)}{\sum_{j=1}^B \exp(\boldsymbol{x}_i^\top \boldsymbol{y}_j / \sigma)} \right)

    Lt2v=−1B∑i=1Blog⁡(exp⁡(yi⊤xi/σ)∑j=1Bexp⁡(yi⊤xj/σ))L_{\text{t2v}} = -\frac{1}{B} \sum_{i=1}^B \log \left( \frac{\exp(\boldsymbol{y}_i^\top \boldsymbol{x}_i / \sigma)}{\sum_{j=1}^B \exp(\boldsymbol{y}_i^\top \boldsymbol{x}_j / \sigma)} \right)

    Here, σ>0\sigma > 0 is a scalar temperature hyperparameter, fixed to σ=0.05\sigma = 0.05 during pretraining and finetuning.

  3. Knowl 3 — Joint Video and Image Training via Alternating Batches and Zero-Temporal Initialization

    model/method

    The Frozen in Time architecture unifies image-text and video-text pretraining by treating an image as a single-frame video clip (M=1M = 1).

    Initialization:

    • The spatial self-attention parameters of the space-time transformer are initialized with weights from a Vision Transformer (ViT) pre-trained on ImageNet-21k.
    • The temporal self-attention parameters are initialized to zero.
    • Due to the residual connection around the temporal attention block, the model initially acts as an exact per-frame ViT feature extractor and gradually learns temporal dependencies across frames as training proceeds.

    Alternating Batch Strategy: Training alternates between batches of image-text pairs (such as Conceptual Captions CC3M or COCO) and video-text pairs (such as WebVid-2M). Because the self-attention computational and memory cost scales with O(M2)O(M^2) in the number of frames MM, image batches (M=1M = 1) can use much larger batch sizes (e.g., B=96B = 96 for M=1M = 1, compared to B=24B = 24 for M=4M = 4 and B=16B = 16 for M=8M = 8), maximizing GPU throughput and contrastive negative diversity.

  4. Knowl 4 — Temporal Curriculum Learning via Positional Embedding Expansion

    algorithm

    To reduce the computational burden of training multi-frame video transformers, training follows a temporal curriculum schedule that progressively increases the number of input frames MM over time. When transitioning from a model trained on mm frames to MM frames (M>mM > m), the learned temporal positional embedding matrix Et∈Rm×D\boldsymbol{E}^t \in \mathbb{R}^{m \times D} must be expanded to Et′∈RM×D\boldsymbol{E}^{t\prime} \in \mathbb{R}^{M \times D}.

    Input: Pretrained temporal embedding Et∈Rm×DE^t \in \mathbb{R}^{m \times D}, source length mm, target length MM, expansion method method ∈{"zero","bilinear","nearest"}\in \{\text{"zero"}, \text{"bilinear"}, \text{"nearest"}\}
    Output: Expanded temporal embedding Et′∈RM×DE^{t\prime} \in \mathbb{R}^{M \times D}
    if method == "zero" then
        Initialize Et′∈RM×DE^{t\prime} \in \mathbb{R}^{M \times D} to zeros
        E1:m,:t′←E1:m,:tE^{t\prime}_{1:m, :} \leftarrow E^t_{1:m, :}
    else if method == "bilinear" then
        for k←1k \leftarrow 1 to MM do
            u←1+(k−1)×m−1M−1u \leftarrow 1 + (k - 1) \times \frac{m - 1}{M - 1}
            i←⌊u⌋i \leftarrow \lfloor u \rfloor
            α←u−i\alpha \leftarrow u - i
            if i<mi < m then
                Ek,:t′←(1−α)Ei,:t+αEi+1,:tE^{t\prime}_{k, :} \leftarrow (1 - \alpha) E^t_{i, :} + \alpha E^t_{i+1, :}
            else
                Ek,:t′←Em,:tE^{t\prime}_{k, :} \leftarrow E^t_{m, :}
            end if
        end for
    else if method == "nearest" then
        for k←1k \leftarrow 1 to MM do
            u←round(1+(k−1)×m−1M−1)u \leftarrow \text{round}\left(1 + (k - 1) \times \frac{m - 1}{M - 1}\right)
            Ek,:t′←Eu,:tE^{t\prime}_{k, :} \leftarrow E^t_{u, :}
        end for
    end if
    return Et′E^{t\prime}

    This frame curriculum (e.g., 1⇒4⇒81 \Rightarrow 4 \Rightarrow 8 frames) achieves equal or superior retrieval accuracy compared to direct 8-frame training while reducing total pretraining time from 98.0 GPU hours to 36.0 GPU hours.

  5. Knowl 5 — WebVid-2M and WebVid-10M Video-Text Pretraining Datasets

    definition

    WebVid-2M and its extension WebVid-10M are large-scale web-scraped video-text datasets designed for end-to-end multi-modal pretraining.

    Dataset Characteristics:

    • WebVid-2M: Comprises 2.5 million video clip-text pairs totaling approximately 13,000 hours of video, with an average clip duration of 18 seconds. Over 275,000 clips exceed 30 seconds in length.
    • WebVid-10M: Extends WebVid-2M to 10.0 million video clip-text pairs totaling approximately 52,000 hours of video with an average clip duration of 18 seconds.

    Comparison to ASR-derived Video Datasets (e.g., HowTo100M):

    1. Text Quality and Alignment: Captions in WebVid are written descriptions scraped from web video alt-text, which are visually aligned with the video content. In contrast, HowTo100M relies on automatic speech recognition (ASR) narration, which frequently lacks temporal alignment, punctuation, and direct visual relevance.
    2. Lexical Diversity: WebVid-2M captions average 12 words per sentence with a Measure of Textual Lexical Diversity (MTLD) of 203, whereas HowTo100M clip narrations average 4 words with an MTLD of 13.5.
    3. Pretraining Efficiency: Pretraining on 2.5M WebVid pairs achieves superior downstream retrieval performance on MSR-VTT (26.0% R@1) compared to pretraining on a 17M subset of HowTo100M (24.1% R@1), despite WebVid-2M being nearly an order of magnitude smaller.
  6. Knowl 6 — Video Frame Sampling and Test-Time Multi-Clip Aggregation

    model/method

    To represent variable-length videos with a fixed number of frames MM during training and testing:

    1. Training Frame Sampling: Given an input video of length LL frames, the video is subdivided into MM equal non-overlapping temporal segments. A single frame is sampled uniformly at random from each of the MM segments (following Temporal Segment Networks).
    2. Test-Time Evaluation and Aggregation: At test time, multi-frame video clips are sampled across the full video duration using a fixed temporal stride S=2 secondsS = 2\text{ seconds}. For the kk-th evaluated clip, an embedding vector vk∈Rdproj\boldsymbol{v}_k \in \mathbb{R}^{d_{\text{proj}}} is produced by the visual encoder. The final video representation vˉ\bar{\boldsymbol{v}} is computed by averaging the clip embeddings: vˉ=1K∑k=1Kvk\bar{\boldsymbol{v}} = \frac{1}{K} \sum_{k=1}^K \boldsymbol{v}_k where KK is the number of evaluated clips spanning the video.
  7. Knowl 7 — Comparative Evaluation of Video and Text Backbones for Video-Text Retrieval

    empirical result

    Ablations on the MSR-VTT 1K-A text-to-video retrieval benchmark demonstrate the superiority of the space-time transformer over 2D and 3D convolutional networks, as well as the parameter efficiency of DistilBERT.

    All models were pretrained on WebVid-2M and finetuned on the MSR-VTT training set (using 4 input frames for video models, and 1 frame for ResNet-101):

    Video Backbone #params R@1 R@10 MedR
    ResNet-101 45M 11.5 44.1 14.5
    S3D-G 76M 3.6 20.4 59.5
    R(3D)-101 85M 9.3 38.3 20.0
    Space-Time Transformer (224/16 B) 114M 26.8 68.2 4.0
    Text Backbone #params R@1 R@10 MedR
    t5-small 60.5M 15.1 51.4 10.0
    t5-base 222.9M 24.0 62.8 6.0
    distilbert-base-uncased 66.4M 26.8 68.2 4.0
    bert-base-uncased 109.5M 27.5 67.3 4.0

    Key takeaways:

    • The Space-Time Transformer outperforms 3D ResNet-101 by +17.5% R@1 and S3D-G by +23.2% R@1.
    • DistilBERT matches the retrieval performance of full BERT-base (26.8% vs 27.5% R@1) while requiring approximately 40% fewer parameters (66.4M vs 109.5M), making it the preferred lightweight text backbone.
  8. Knowl 8 — State-of-the-Art Video Retrieval Performance Across Downstream Benchmarks

    empirical result

    The Frozen in Time model achieves state-of-the-art results on multiple standard text-to-video retrieval benchmarks, using only raw pixels without pre-extracted multi-modal expert features.

    MSR-VTT (1k-A Test Split, Text-to-Video Retrieval):

    Method Pretraining Dataset #pairs R@1 R@5 MedR
    ClipBERT COCO, VisGenome 5.6M 22.0 46.8 6.0
    MMT (7 experts) HowTo100M 136M 26.6 57.1 4.0
    Support Set HowTo100M 136M 30.1 58.5 3.0
    Ours (CC3M) CC3M 3.0M 25.5 54.5 4.0
    Ours (CC3M+WV-2M) CC3M + WebVid-2M 5.5M 31.0 59.5 3.0
    Ours (CC3M+WV-2M+COCO) CC3M + WV-2M + COCO 6.1M 32.5 61.5 3.0
    Zero-Shot Retrieval
    HT MIL-NCE HowTo100M 136M 7.5 21.2 38.0
    Support Set HowTo100M 136M 8.7 23.0 31.0
    Ours (CC3M+WV-2M) CC3M + WebVid-2M 5.5M 23.2 44.6 7.0
    Ours (CC3M+WV-2M+COCO) CC3M + WV-2M + COCO 6.1M 24.7 46.9 7.0

    Other Video Retrieval Benchmarks:

    • MSVD: Achieves 33.7% R@1, 64.7% R@5, MedR 3.0, outperforming Support Set with HowTo100M pretraining (28.4% R@1).
    • DiDeMo: Achieves 31.0% R@1 (without ground truth proposals) and 34.6% R@1 (with ground truth proposals), outperforming ClipBERT (20.4% R@1). In a zero-shot setting, it attains 21.1% R@1 without finetuning.
    • LSMDC: Achieves 15.0% R@1 and 30.8% R@5, outperforming MMT (12.9% R@1) and CE (11.2% R@1).
    • ActivityNet Captions (val1k): Achieves 28.8% R@1 and 60.9% R@5, comparable to Support Set (29.2% R@1) despite using 20x fewer training pairs.
  9. Knowl 9 — Scaling and Source Diversity Effects in Multi-Modal Pretraining

    empirical result

    Evaluating pretraining across various image and video data source combinations on MSR-VTT 1k-A text-to-video retrieval demonstrates that adding smaller datasets from diverse sources yields greater performance gains than scaling up data from a single source.

    Pretraining Data #pairs R@1 R@5 R@10 MedR
    COCO 0.6M 27.2 56.1 67.5 4.0
    WebVid-2M 2.5M 27.5 56.6 67.6 4.0
    WebVid-10M 10.0M 28.9 57.2 68.6 4.0
    CC3M + WebVid-2M 5.0M 31.0 59.5 70.5 3.0
    CC3M + WebVid-2M + COCO 5.6M 32.5 61.5 71.2 3.0
    CC3M + WebVid-10M 13.0M 33.4 59.2 70.7 3.0
    CC3M + CC12M + WebVid-10M 25.0M 34.0 61.4 73.1 3.0

    Key observations:

    • Adding COCO captions (0.57M pairs) to CC3M + WebVid-2M improves R@1 by +1.5% (from 31.0% to 32.5%), whereas adding 7.5M additional WebVid videos (moving from WebVid-2M to WebVid-10M with CC3M) improves R@1 by +2.4% (to 33.4%), demonstrating high data efficiency from source diversity.
    • Increasing pretraining data volume from 0.6M to 25.0M pairs continuously improves downstream retrieval without saturating.
  10. Knowl 10 — Cross-Modal Text-to-Image Retrieval on Flickr30K

    empirical result

    The Frozen in Time visual-text encoder transfers effectively to static image retrieval tasks without requiring task-specific architectural modifications or external object detectors (such as Faster R-CNN).

    Evaluated on the standard Flickr30K 1k test set for text-to-image retrieval:

    Method Visual PT Data / Size R@1 R@5 R@10
    SCANM Visual Genome Objects (3.8M) 48.6 77.7 85.2
    IMRAM Visual Genome Objects (3.8M) 53.9 79.4 87.2
    SGRAF Visual Genome Objects (3.8M) 58.5 83.0 88.8
    Ours (CC3M) CC3M (3.0M) 54.2 83.2 89.8
    Ours (CC3M + WV-2M) CC3M + WebVid-2M (5.5M) 61.0 87.5 92.7

    Key takeaways:

    • When pretrained on CC3M and WebVid-2M, the model attains 61.0% R@1, outperforming region-feature-based models such as SGRAF (58.5% R@1).
    • Joint pretraining on video data (WebVid-2M) provides a +6.8% improvement in R@1 on static image retrieval compared to pretraining on CC3M alone, demonstrating reciprocal cross-modal transfer between video and image representations.

Coverage note — None was omitted; all primary architectural innovations, equations, curriculum training strategies, dataset definitions, and empirical benchmark evaluations on MSR-VTT, MSVD, DiDeMo, LSMDC, ActivityNet, and Flickr30K are fully covered.

References

  1. 1.Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. Selfsupervised multimodal versatile networks. In NeurIPS, 2020. 2, 3
  2. 2.Elad Amrani, Rami Ben Ari, Daniel Rotman, and Alex Bronstein. Noise estimation using density estimation for self-supervised multimodal learning. arXiv preprint arXiv:2003.03186, 2020. 8
  3. 3.Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell. Localizing moments in video with natural language. In ICCV, 2017. 2, 5
  4. 4.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In ICCV, 2015. 1
  5. 5.Max Bain, Arsha Nagrani, Andrew Brown, and Andrew Zisserman. Condensed movies: Story based retrieval with contextual embeddings. In ACCV, 2020. 2, 5
  6. 6.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv:2102.05095, 2021. 2, 3, 4, 11
  7. 7.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-toend object detection with transformers. In ECCV, 2020. 3
  8. 8.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? A new model and the Kinetics dataset. In CVPR, 2017. 2, 3
  9. 9.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, 2021. 9
  10. 10.David Chen and William B Dolan. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 190–200, 2011. 2, 5, 7, 8
  11. 11.Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. IMRAM: Iterative matching with recurrent attention memory for cross-modal image-text retrieval, 2020. 8, 9
  12. 12.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C. Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server, 2015. 9
  13. 13.Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng. A2-nets: Double attention networks. arXiv preprint arXiv:1810.11579, 2018. 3
  14. 14.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. UNITER: Universal image-text representation learning, 2020. 9
  15. 15.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019. 3
  16. 16.Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. On the relationship between self-attention and convolutional layers. arXiv preprint arXiv:1911.03584, 2019. 3
  17. 17.J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, 2019. 3, 10
  18. 18.Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. Similarity reasoning and filtration for image-text matching, 2021. 8, 9
  19. 19.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 3, 4
  20. 20.Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. Vse++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612, 2017. 8
  21. 21.Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. Multi-modal transformer for video retrieval. In ECCV, 2020. 2, 3, 8, 9, 10
  22. 22.Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell. Actionvlad: Learning spatio-temporal aggregation for action classification. In CVPR, 2017. 3
  23. 23.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? In CVPR, 2018. 2
  24. 24.Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei. Relation networks for object detection. In CVPR, 2018. 3
  25. 25.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. arXiv preprint arXiv:2102.05918, 2021. 9
  26. 26.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. The Kinetics human action video dataset. CoRR, abs/1705.06950, 2017. 2
  27. 27.Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel. Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539, 2014. 8
  28. 28.Bruno Korbar, Fabio Petroni, Rohit Girdhar, and Lorenzo Torresani. Video understanding as machine translation. arXiv preprint arXiv:2006.07203, 2020. 8
  29. 29.Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In ICCV, 2017. 1, 2, 5, 10
  30. 30.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision, 123(1):32–73, 2017. 2
  31. 31.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching, 2018. 8, 9
  32. 32.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. arXiv preprint arXiv:2102.06183, 2021. 2, 3, 5, 8, 11
  33. 33.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. Tvqa: Localized, compositional video question answering. arXiv preprint arXiv:1809.01696, 2018. 1
  34. 34.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. Hero: Hierarchical encoder for video+ language omni-representation pre-training. EMNLP, 2020. 8
  35. 35.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV, 2020. 9
  36. 36.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014. 1, 2
  37. 37.Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, and Zhongyuan Wang. Hit: Hierarchical transformer with momentum contrast for video-text retrieval, 2021. 2
  38. 38.Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. Use what you have: Video retrieval using representations from collaborative experts. In Proc. BMVC, 2019. 1, 2, 3, 5, 8, 9, 10
  39. 39.Chenxu Luo and Alan Yuille. Grouped spatial-temporal aggretation for efficient action recognition. In ICCV, 2019. 4
  40. 40.Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Xilin Chen, and Ming Zhou. UniVL: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020. 8
  41. 41.Philip M McCarthy and Scott Jarvis. Mtld, vocd-d, and hd-d: A validation study of sophisticated approaches to lexical diversity assessment. Behavior research methods, 42(2):381–392, 2010. 5
  42. 42.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In CVPR, 2020. 2, 3
  43. 43.Antoine Miech, Ivan Laptev, and Josef Sivic. Learning a text-video embedding from incomplete and heterogeneous data. arXiv, 2018. 1, 2, 3, 9
  44. 44.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In ICCV, 2019. 2, 5, 8
  45. 45.Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K Roy-Chowdhury. Learning joint embedding with multimodal cues for cross-modal video-text retrieval. In Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval, pages 19–27, 2018. 8
  46. 46.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. Image transformer. In ICML, 2018. 3
  47. 47.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019. 6
  48. 48.Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, Joao Henriques, and Andrea Vedaldi. Support-set bottlenecks for video-text representation learning. arXiv preprint arXiv:2010.02824, 2020. 2, 5, 6, 7, 8, 10
  49. 49.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. 2
  50. 50.Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. Standalone self-attention in vision models. arXiv preprint arXiv:1906.05909, 2019. 3
  51. 51.Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics, 1:25–36, 2013. 5
  52. 52.Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. A dataset for movie description. In CVPR, 2015. 5
  53. 53.Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele. Movie description. International Journal of Computer Vision, 123(1):94–120, 2017. 2, 5
  54. 54.Marcus Rohrbach, Sikandar Amin, Mykhaylo Andriluka, and Bernt Schiele. A database for fine grained activity detection of cooking activities. In CVPR, 2012. 5
  55. 55.Andrew Rouditchenko, Angie Boggust, David Harwath, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, et al. AVLnet: Learning audio-visual language representations from instructional videos. arXiv preprint arXiv:2006.09199, 2020. 2, 8
  56. 56.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108, 2019. 6, 10
  57. 57.Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid. Look before you speak: Visually contextualized utterances. CVPR, 2021. 2
  58. 58.Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan, Vedanuj Goswami, Matt Feiszli, and Lorenzo Torresani. Only time can tell: Discovering temporal data for temporal modeling. In WACV, 2021. 2
  59. 59.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018. 2, 4, 5
  60. 60.Gunnar A Sigurdsson, Gul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In ECCV, 2016. 5
  61. 61.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020. 3
  62. 62.Du Tran, Heng Wang, L. Torresani, Jamie Ray, Y. LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In CVPR, 2018. 2
  63. 63.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 1, 3
  64. 64.Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko. Translating videos to natural language using deep recurrent neural networks. arXiv preprint arXiv:1412.4729, 2014. 8
  65. 65.Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and tell: Lessons learned from the 2015 MSCOCO image captioning challenge. IEEE transactions on pattern analysis and machine intelligence, 39(4):652–663, 2016. 1
  66. 66.Liwei Wang, Yin Li, and Svetlana Lazebnik. Learning deep structure-preserving image-text embeddings. In CVPR, 2016. 1
  67. 67.L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool. Temporal segment networks for action recognition in videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(11):2740–2755, 2019. 4
  68. 68.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. 3
  69. 69.Xiaohan Wang, Linchao Zhu, and Yi Yang. T2vlad: Globallocal sequence alignment for text-video retrieval, 2021. 8
  70. 70.Chao-Yuan Wu, Ross B. Girshick, Kaiming He, Christoph Feichtenhofer, and Philipp Krahenbuhl. A multigrid method for efficiently training video models. In CVPR, 2020. 3
  71. 71.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In ECCV, 2018. 2
  72. 72.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016. 2, 5
  73. 73.Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. Image captioning with semantic attention. In CVPR, 2016. 1
  74. 74.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014. 6
  75. 75.Youngjae Yu, Jongseok Kim, and Gunhee Kim. A joint sequence fusion model for video question answering and retrieval. In ECCV, 2018. 8, 9
  76. 76.Andrew Zhai and Hao-Yu Wu. Classification is a strong baseline for deep metric learning. In BMVC, 2019. 3
  77. 77.Bowen Zhang, Hexiang Hu, and Fei Sha. Cross-modal and hierarchical modeling of video and text. In ECCV, 2018. 8
  78. 78.Luowei Zhou, Chenliang Xu, and Jason Corso. Towards automatic learning of procedures from web instructional videos. In AAAI, 2018. 2, 5
  79. 79.Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong. End-to-end dense video captioning with masked transformer. In CVPR, 2018. 2
  80. 80.Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In CVPR, 2020. 8

Citation

MLA
Bain, M., et al. “Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval”. arXiv, 2021, http://arxiv.org/abs/2104.00650v2.
APA
Bain, M., Nagrani, A., Varol, G., & Zisserman, A. (2021). Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. arXiv. http://arxiv.org/abs/2104.00650v2
Chicago
Bain, M., A. Nagrani, G. Varol, and A. Zisserman. 2021. “Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval”. arXiv. http://arxiv.org/abs/2104.00650v2.
Harvard
Bain, M. et al. (2021) “Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2104.00650v2.
Vancouver
1. Bain M, Nagrani A, Varol G, Zisserman A (2021) Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. arXiv

BibTeX

@article{bain2021frozen,
  title = {Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval},
  author = {Bain, Max and Nagrani, Arsha and Varol, Gül and Zisserman, Andrew},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2104.00650v2},
  eprint = {2104.00650}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/