VideoPrism: A Foundational Visual Encoder for Video Understanding

Long ZhaoNitesh Bharadwaj GundavarapuLiangzhe YuanHao ZhouShen YanJennifer J. SunLuke FriedmanRui QianTobias WeyandYue Zhao

article2024ICML122 citations

Presents VideoPrism, a foundational visual encoder pretrained via a two-stage contrastive and masked modeling pipeline that achieves state-of-the-art results across 31 video understanding benchmarks using a single frozen model.

Listen

Building a universal artificial intelligence model for video understanding remains a critical challenge because existing systems typically struggle to balance static visual appearance with complex motion over time. Most prior approaches either adapt image-based models that miss dynamic motion or rely on specialized architectures tailored to narrow benchmarks. The article introduces VideoPrism, a general-purpose visual encoder designed to handle a wide spectrum of video understanding tasks—such as classification, localization, retrieval, captioning, and question answering—using a single frozen model.

To build this foundational encoder, the researchers assembled a massive pretraining dataset comprising 36 million high-quality video-caption pairs and 582 million video clips with noisy text transcripts. They developed a unique two-stage pretraining strategy: the first stage uses contrastive learning on video-text pairs to align visual content with linguistic concepts, and the second stage continues training on video-only data using masked video modeling. This second stage incorporates global-local knowledge distillation to prevent forgetting visual concepts and applies a token shuffling mechanism that forces the model to learn complex motion and contextual structures rather than relying on superficial shortcut patterns. VideoPrism was evaluated across 33 diverse benchmarks spanning general web video, complex human actions, and specialized scientific domains such as neuroscience and ecology.

VideoPrism achieved state-of-the-art performance on 31 of the 33 evaluated benchmarks when evaluated as a single frozen encoder. On the VideoGLUE benchmark, the giant model configuration surpassed prior best-performing foundation models across all eight core tasks, achieving substantial gains such as a 22.4-point increase in mean average precision on the Charades dataset and an 11.8-point increase on spatiotemporal action localization. In zero-shot text-video retrieval, the model showed gains of up to 9.9% on ActivityNet and 9.3% on VATEX, while its base configuration consistently outperformed several larger competing models. Furthermore, VideoPrism matched or outperformed specialized domain-expert models across all scientific benchmarks, including behavioral analysis of fruit flies, mice, chimpanzees, and wild animals.

These findings demonstrate that organizations can deploy a single, frozen foundational video encoder to serve diverse downstream applications without incurring the prohibitive computational, financial, and memory expenses of fine-tuning large models for every specific task. By simultaneously capturing rich semantics and temporal dynamics, VideoPrism eliminates the historical trade-off between motion-centric reasoning and appearance-heavy understanding. The authors recommend adopting frozen-encoder backbones as standard building blocks for multimodal language systems and video analysis pipelines across both enterprise and scientific research environments.

Confidence in these findings is high given the broad empirical validation across standard benchmarks and scientific use cases, supported by strict data de-duplication to prevent evaluation leakage. However, stakeholders should note specific limitations: pretraining relied partly on noisy text annotations, and the model was evaluated primarily on short clips sampling 16 frames. Future work should focus on integrating this encoder into long-form video understanding systems and exploring more extensive conversational evaluation protocols.

arXiv: 2402.13217

No sufficiently relevant recommendations were found.

Cover for VideoPrism: A Foundational Visual Encoder for Video Understanding

Table of Contents

  • 1. Introduction
  • 2. Approach
  • 2.1. Pretraining data
  • 2.2. Model architecture
  • 2.3. Training algorithm
  • 2.3.1. STAGE 1: VIDEO-TEXT CONTRASTIVE TRAINING
  • 2.3.2. STAGE 2: MASKED VIDEO MODELING
  • 3. Experiments
  • 3.1. Classification and spatiotemporal localization
  • 3.2. Zero-shot video-text retrieval and classification
  • 3.3. Zero-shot video captioning and QA
  • 3.4. CV for science tasks
  • 3.5. Ablation study
  • 3.6. Limitations
  • 4. Related work
  • 5. Conclusion
  • Acknowledgements
  • Impact statement
  • References
  • A. Pretraining data
  • A.1. Data curation
  • A.2. Corpus analysis
  • B. Model architecture
  • C. Implementation details
  • C.1. Stage 1
  • C.2. Stage 2
  • D. Evaluation data
  • E. VideoGLUE
  • E.1. Tasks and task heads for VideoPrism
  • E.2. Adaptations
  • E.3. Results
  • F. Zero-shot video-text retrieval
  • F.1. Implementation details
  • F.2. Zero-shot classification on Charades-STA
  • F.3. Additional results on MSRVTT
  • F.4. Additional results on Kinetics-600
  • G. Gluing VideoPrism with PaLM-2
  • H. CV for Science
  • I. Ablation studies
  • I.1. Data
  • I.2. Model design
  • I.3. Scaling properties

Knowls

  1. Knowl 1 — Two-stage video-text and video-only pretraining

    model/method

    VideoPrism is trained in two stages to combine language-derived semantics with video-only contextual learning. In Stage 1, a video encoder and a text encoder are trained on video-text pairs with a symmetric cross-entropy contrastive objective over pairwise similarities within each minibatch. A multi-head attention pooler (MAP) converts the video encoder’s token features into a global video embedding. Training alternates minibatches from different datasets using alternating gradient descent (AGD); the spatial video encoder is initialized from CoCa, and the stage also includes WebLI image–alt-text data. The resulting Stage 1 encoder supplies semantic global and token-level targets for Stage 2.

    In Stage 2, the video encoder is initialized from Stage 1 and trained on video clips without text. A frozen Stage 1 encoder processes each intact clip as the teacher. The student receives masked video patches and learns both to match the teacher’s token-level embeddings and to reproduce its global video embedding. For token-level prediction, visible student embeddings are combined with mask tokens, randomly shuffled, and only then given positional embeddings before entering a four-layer Transformer decoder. This prevents the decoder from simply copying visible tokens in their original positions. A separate four-layer Transformer decoder and MAP layer predict the global embedding from visible student features; this path uses neither token shuffling nor positional embeddings. The two cosine-distance distillation losses have equal weight.

    The video encoder uses 8 sampled frames during pretraining. Both stages use Adafactor and batch size 4096; Stage 1 runs for 200,000 iterations with 20,000 warmup iterations and linear learning-rate decay, while Stage 2 runs for 300,000 iterations with 25,000 warmup iterations and cosine decay. Stage 1 drops 50% of tokens using tube masking; Stage 2 masks 65% using BEVT masking. Stage 2 excludes WebLI because it is image-based.

  2. Knowl 2 — Hybrid pretraining corpus

    data/table

    VideoPrism’s pretraining corpus combines 36.1 million high-quality manually captioned video clips with roughly 582 million clips paired with noisier parallel text, for approximately 618 million clips from about 275 million videos. The mixture is intended to provide both reliable semantic supervision and scale across diverse video sources. The corpus is deduplicated against all 33 evaluation benchmarks, and benchmark training sets such as Kinetics are not added during pretraining or post-pretraining.

    Source Video source Text source Quality Videos (M) Clips (M)
    Anonymous-Corpus #1 Web Manual captions High 36.1 36.1
    WTS-70M YouTube Metadata Low 55.1 55.1
    YT-Temporal-180M YouTube ASR Low 2.3 87.8
    VideoCC YouTube Retrieved image captions Low 133.5 191.1
    InternVid YouTube VLM/LLM-generated captions Medium 2.8 7.0
    Anonymous-Corpus #2 YouTube ASR Low 44.6 170.3
    Anonymous-Corpus #3 YouTube VLM/LLM-generated captions Medium 36.7 71.5

    The in-house ASR corpus uses ASR sentence boundaries for clips and filters pairs using metadata and a text-video groundedness score. The in-house generated-caption corpus uses vision-language models and an LLM, and filters out primarily talking-head videos and static clips. The manually captioned corpus consists of commercially licensed stock videos uploaded with contributor-written captions and is not filtered.

  3. Knowl 3 — Factorized video encoder architecture

    model/method

    VideoPrism is based on a Vision Transformer with factorized spatial and temporal processing. An input video is partitioned into non-overlapping patches. A spatial Transformer first models interactions among patches at each temporal index; a temporal Transformer then models interactions across temporal indices. The temporal module has four layers, and the model retains the spatiotemporal token sequence rather than applying global average pooling after spatial encoding, preserving fine-grained features for localization and other dense tasks. Spatial and temporal learnable positional embeddings are separate.

    The giant configuration uses a ViT-Giant spatial encoder with 1 billion parameters; VideoPrism-B uses a ViT-Base spatial encoder. The reported pretraining configuration uses 288 × 288 resolution, 18 × 18 patches, and 8 uniformly sampled frames. Evaluation uses 16 frames, with temporal positional embeddings interpolated for the changed frame count. The factorized design was selected to balance compute and memory costs, particularly for the large contrastive-training batches.

  4. Knowl 4 — Ablations identify the effects of scale, distillation, and shuffling

    empirical result

    Frozen-feature MAP probing on Kinetics-400 (K400) and Something-Something v2 (SSv2) shows that the successive additions in VideoPrism’s training design improve the joint appearance- and motion-focused results. The progression begins with a video-text contrastive model trained on 150 million clips; each subsequent row modifies the preceding configuration.

    Configuration SSv2 top-1 accuracy (%) K400 top-1 accuracy (%)
    Contrastive baseline (150M clips) 50.0 81.7
    Use full data corpus (600M clips) 55.4 83.8
    Contrastive and MAE in one stage 55.9 82.7
    Use two-stage training 60.9 81.9
    Add global distillation 61.8 83.3
    Add token shuffling (full model) 63.6 84.2

    A separate component ablation on VideoPrism-B reports K400 and SSv2 top-1 accuracy and AVA mean average precision (mAP). The full model scores 84.2, 63.6, and 30.6, respectively. Removing token shuffling changes these scores to 83.6, 61.8, and 29.4; removing global distillation changes them to 83.4, 64.2, and 29.0. Thus, shuffling particularly benefits SSv2, while global distillation improves the appearance-focused K400 and AVA results; the reported SSv2 score is slightly higher without global distillation.

  5. Knowl 5 — Frozen-backbone performance on video-only understanding

    empirical result

    On the eight-dataset VideoGLUE evaluation, VideoPrism is used as a frozen video encoder and only task heads are trained on downstream training sets. Metrics are top-1 accuracy for K400, Moments-in-Time (MiT), SSv2, and Diving48 (D48), and mAP for Charades, ActivityNet temporal localization, AVA spatiotemporal localization, and AVA-Kinetics (AVA-K). All scores below are percentages.

    Model K400 MiT SSv2 D48 Charades ActivityNet AVA AVA-K
    VideoPrism-B 84.2 40.8 63.6 67.4 40.4 36.6 30.6 31.8
    VideoPrism-g 87.2 45.5 68.5 71.3 62.3 37.8 36.2 37.3

    The paper reports that VideoPrism outperforms the compared frozen-backbone models on every listed dataset, and that scaling from B to giant improves every score. Across the broader VideoGLUE adaptation evaluation—which also includes multi-layer pooling, low-rank adapters, and end-to-end tuning—the reported VideoGLUE score for VideoPrism-B is 51.25, 13.6% above the second-ranked model.

  6. Knowl 6 — Zero-shot video-text retrieval

    empirical result

    For zero-shot retrieval, VideoPrism is paired with a text encoder tuned to match the frozen video encoder’s embeddings using the first-stage pretraining data. On the reported MSRVTT 1K-A split (1,000 test videos), VATEX, and ActivityNet benchmarks, VideoPrism-g obtains the highest reported scores among the compared methods in both retrieval directions. Recall@1 and Recall@5 are reported below.

    MSRVTT (1K-A) VATEX ActivityNet
    Model Direction R@1 R@5 R@1 R@5 R@1 R@5 R@1 R@5 R@1 R@5 R@1 R@5
    Text to video Video to text Text to video Video to text Text to video Video to text
    VideoPrism-B 51.4 74.4 50.2 73.2 57.7 88.5 76.2 93.7 49.6 76.7 47.9 75.0
    VideoPrism-g 52.7 77.2 51.7 75.2 62.5 91.0 77.1 95.6 52.7 79.4 50.3 77.1

    The scores show gains for both model sizes across datasets and directions. The strongest results are on VATEX video-to-text retrieval, where VideoPrism-g reaches R@1 77.1 and R@5 95.6, and on ActivityNet, where it reaches R@1 52.7 for text-to-video and 50.3 for video-to-text.

  7. Knowl 7 — Zero-shot video classification and temporal retrieval

    empirical result

    VideoPrism’s zero-shot text alignment is evaluated by matching videos to text descriptions of class labels or candidate answers. The benchmarks include appearance-focused K400, motion-focused SSv2-Temporal and SSv2-Events, NExT-QA ATP-Hard, multi-label Charades, and the temporal-reasoning Charades-STA task. Scores are percentages; K400 reports top-1/top-5 accuracy, SSv2 reports temporal/events accuracy, NExT-QA ATP-Hard and Charades-STA report multi-choice (MC) accuracy, and Charades reports mAP.

    Benchmark Metric VideoPrism-B VideoPrism-g Best prior score
    K400 Top-1 / Top-5 71.3 / 91.7 76.4 / 94.3 72.0 / 90.5
    SSv2-Temporal Accuracy 16.1 18.6 15.2
    SSv2-Events Accuracy 11.9 15.7 11.4
    NExT-QA ATP-Hard MC accuracy 31.3 32.7 27.6
    Charades mAP 29.2 32.4 25.8
    Charades-STA MC accuracy 50.0 50.4 47.2

    The K400 best-prior values are VideoCoCa-g scores; the prior scores for the other rows are respectively VideoCon-L, VideoCon-L, TACT-B, VideoCoCa-g, and VideoCoCa-g. VideoPrism-g exceeds the prior score on every listed metric except K400 top-1, while VideoPrism-B exceeds the prior score on every listed metric except K400 top-1. Charades-STA is evaluated by trimming videos to annotated temporal segments and retrieving the correct description from the descriptions associated with that video.

  8. Knowl 8 — Video captioning and question answering with frozen language models

    empirical result

    For generative tasks, VideoPrism-B is connected to a frozen PaLM-2 language decoder through a trainable one-layer Perceiver Resampler that produces 256 video tokens. The video encoder and language model remain frozen; the models are not tuned separately for captioning and QA. Captioning is evaluated with CIDEr, while QA reports top-1 accuracy on MSRVTT-QA and MSVD-QA and WUPS on NExT-QA. QA uses two-shot text-only prompts.

    Captioning model MSRVTT CIDEr VATEX CIDEr YouCook2 CIDEr
    VideoPrism-B + PaLM-2-1B 40.3 24.2 52.3
    VideoPrism-B + PaLM-2-8B 38.5 31.7 63.6
    QA model MSRVTT-QA top-1 (%) MSVD-QA top-1 (%) NExT-QA WUPS
    VideoPrism-B + PaLM-2-1B 28.5 39.5 23.8
    VideoPrism-B + PaLM-2-8B 32.0 47.1 27.4

    The systems are competitive with methods that freeze both vision and language components, and the paper reports that they lead that comparison except on VATEX captioning. On QA, the larger PaLM-2 configuration scores higher than the smaller one on all three benchmarks.

  9. Knowl 9 — Transfer to scientific behavior videos

    empirical result

    The scientific-video evaluation uses a shared frozen VideoPrism encoder with task-specific heads trained on expert-annotated data. The tasks cover fruit-fly behavior (Fly vs. Fly), mouse behavior (CalMS21 and CRIM13), Kenyan animal behavior (KABR), and chimpanzee spatiotemporal action localization (ChimpACT). Scores are mAP except for KABR, which reports macro-accuracy; CRIM13 gives side-view (S) and top-view (T) results.

    Model Fly vs. Fly CalMS21 CRIM13 (S / T) KABR ChimpACT
    Domain expert 88.6 88.9 – 61.9 24.4
    VideoPrism-B 89.1 91.1 64.5 / 64.9 61.6 28.8
    VideoPrism-g 92.0 91.5 65.9 / 66.8 63.3 31.5

    VideoPrism-B exceeds the listed domain-expert score on Fly vs. Fly, CalMS21, and ChimpACT, and is 0.3 points below the KABR expert score. VideoPrism-g exceeds the listed expert scores on all datasets with an expert value; increasing model size also improves all five reported VideoPrism task results.

  10. Knowl 10 — Stated limitations

    limitation

    VideoPrism’s use of noisy parallel text creates a risk that incomplete or biased captions will affect its representations and performance. Long-video understanding remains difficult because the current model is designed for short clips and uses 16 sampled frames as input at evaluation. The frozen-backbone evaluation setting is practical for reducing adaptation cost, but the authors note that some applications may benefit more from end-to-end fine-tuning or parameter-efficient adaptation.

Coverage note — Detailed corpus-distribution analyses and exhaustive downstream task-head and optimizer recipes are omitted because they provide supporting characterization or implementation detail rather than separate main contributions.

References

  1. 1.Akbari, H., Yuan, L., Qian, R., Chuang, W.-H., Chang, S.-F., Cui, Y., and Gong, B. VATT: Transformers for multimodal self-supervised learning from raw video, audio and text. In NeurIPS, 2021.
  2. 2.Akbari, H., Kondratyuk, D., Cui, Y., Hornung, R., Wang, H., and Adam, H. Alternating gradient descent and mixture-of-experts for integrated multimodal perception. In NeurIPS, 2023.
  3. 3.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: A visual language model for few-shot learning. In NeurIPS, 2022.
  4. 4.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. PaLM 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  5. 5.Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučič, M., and Schmid, C. ViViT: A video vision transformer. In ICCV, 2021.
  6. 6.Bagad, P., Tapaswi, M., and Snoek, C. G. Test of time: Instilling video-language models with a sense of time. In CVPR, 2023.
  7. 7.Bain, M., Nagrani, A., Varol, G., and Zisserman, A. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, 2021.
  8. 8.Bain, M., Nagrani, A., Varol, G., and Zisserman, A. A CLIP-Hitchhiker’s guide to long video retrieval. arXiv preprint arXiv:2205.08508, 2022.
  9. 9.Bansal, H., Bitton, Y., Szpektor, I., Chang, K.-W., and Grover, A. VideoCon: Robust video-language alignment via contrast captions. arXiv preprint arXiv:2311.10111, 2023.
  10. 10.Bao, H., Dong, L., Piao, S., and Wei, F. BEiT: BERT pre-training of image transformers. In ICLR, 2022.
  11. 11.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  12. 12.Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., et al. AudioLM: A language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2023.
  13. 13.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In NeurIPS, 2020.
  14. 14.Buch, S., Eyzaguirre, C., Gaidon, A., Wu, J., Fei-Fei, L., and Niebles, J. C. Revisiting the “video” in video-language understanding. In CVPR, 2022.
  15. 15.Burgos-Artizzu, X. P., Dollar, P., Lin, D., Anderson, D. J., and Perona, P. Social behavior recognition in continuous video. In CVPR, 2012.
  16. 16.Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J. ActivityNet: A large-scale video benchmark for human activity understanding. In CVPR, 2015.
  17. 17.Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A. A short note about Kinetics-600. arXiv preprint arXiv:1808.01340, 2018.
  18. 18.Chen, G., Zheng, Y.-D., Wang, J., Xu, J., Huang, Y., Pan, J., Wang, Y., Wang, Y., Qiao, Y., Lu, T., et al. VideoLLM: Modeling video sequence with large language models. arXiv preprint arXiv:2305.13292, 2023a.
  19. 19.Chen, S. and Huang, D. Elaborative rehearsal for zero-shot action recognition. In ICCV, 2021.
  20. 20.Chen, S., Li, H., Wang, Q., Zhao, Z., Sun, M., Zhu, X., and Liu, J. VAST: A vision-audio-subtitle-text omni-modality foundation model and dataset. In NeurIPS, 2023b.
  21. 21.Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., Beyer, L., et al. PaLI: A jointly-scaled multilingual language-image model. In ICLR, 2023c.
  22. 22.Cheng, F., Wang, X., Lei, J., Crandall, D., Bansal, M., and Bertasius, G. VindLU: A recipe for effective video-and-language pretraining. In CVPR, 2023.
  23. 23.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, 2019.
  24. 24.Dima, D., Doughty, H., Farinella, G. M., Antonino, F., Evangelos, K., Ma, J., Davide, M., Munro, J., Toby, P., Price, W., et al. Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100. IJCV, 130(1):33–55, 2022.
  25. 25.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  26. 26.Eyjolfsdottir, E., Branson, S., Burgos-Artizzu, X. P., Hoopfer, E. D., Schor, J., Anderson, D. J., and Perona, P. Detecting social actions of fruit flies. In ECCV, 2014.
  27. 27.Fang, H., Xiong, P., Xu, L., and Chen, Y. CLIP2Video: Mastering video-text retrieval via image CLIP. arXiv preprint arXiv:2106.11097, 2021.
  28. 28.Fang, Y., Wang, W., Xie, B., Sun, Q.-S., Wu, L. Y., Wang, X., Huang, T., Wang, X., and Cao, Y. EVA: Exploring the limits of masked visual representation learning at scale. In CVPR, 2022.
  29. 29.Fang, Y., Sun, Q., Wang, X., Huang, T., Wang, X., and Cao, Y. EVA-02: A visual representation for neon genesis. arXiv preprint arXiv:2303.11331, 2023.
  30. 30.Feichtenhofer, C., Fan, H., Malik, J., and He, K. SlowFast networks for video recognition. In ICCV, 2019.
  31. 31.Feichtenhofer, C., Fan, H., Xiong, B., Girshick, R., and He, K. A large-scale study on unsupervised spatiotemporal representation learning. In CVPR, 2021.
  32. 32.Feichtenhofer, C., Fan, H., Li, Y., and He, K. Masked autoencoders as spatiotemporal learners. In NeurIPS, 2022.
  33. 33.Fellbaum, C. WordNet and wordnets. In Barber, A. (ed.), Encyclopedia of Language and Linguistics, pp. 2–665. Elsevier, 2005.
  34. 34.Fu, T.-J., Li, L., Gan, Z., Lin, K., Wang, W. Y., Wang, L., and Liu, Z. VIOLET: End-to-end video-language transformers with masked visual-token modeling. arXiv preprint arXiv:2111.12681, 2021.
  35. 35.Gao, J., Sun, C., Yang, Z., and Nevatia, R. TALL: Temporal activity localization via language query. In ICCV, 2017.
  36. 36.Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I. ImageBind: One embedding space to bind them all. In CVPR, 2023.
  37. 37.Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al. The “something something” video database for learning and evaluating visual common sense. In ICCV, 2017a.
  38. 38.Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017b.
  39. 39.Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al. Ego4D: Around the world in 3,000 hours of egocentric video. In CVPR, 2022.
  40. 40.Gu, C., Sun, C., Ross, D. A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., et al. AVA: A video dataset of spatiotemporally localized atomic visual actions. In CVPR, 2018.
  41. 41.Gutmann, M. and Hyvarinen, A. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In AISTATS, 2010.
  42. 42.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. In CVPR, 2022.
  43. 43.He, X., Chen, S., Ma, F., Huang, Z., Jin, X., Liu, Z., Fu, D., Yang, Y., Liu, J., and Feng, J. VLAB: Enhancing video language pre-training by feature adapting and blending. arXiv preprint arXiv:2305.13167, 2023.
  44. 44.Hendricks, L. A. and Nematzadeh, A. Probing image-language transformers for verb understanding. In ACL, 2021.
  45. 45.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In ICLR, 2022.
  46. 46.Huang, G., Pang, B., Zhu, Z., Rivera, C., and Soricut, R. Multimodal pretraining for dense video captioning. arXiv preprint arXiv:2011.11760, 2020.
  47. 47.Huang, T.-H., Ferraro, F., Mostafazadeh, N., Misra, I., Agrawal, A., Devlin, J., Girshick, R., He, X., Kohli, P., Batra, D., et al. Visual storytelling. In NAACL-HLT, 2016.
  48. 48.Jain, P., Kar, P., et al. Non-convex optimization for machine learning. Foundations and Trends® in Machine Learning, 10(3-4):142–363, 2017.
  49. 49.Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 2021.
  50. 50.Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al. The Kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  51. 51.Kholiavchenko, M., Kline, J., Ramirez, M., Stevens, S., Sheets, A., Babu, R., Banerji, N., Campolongo, E., Thompson, M., Van Tiel, N., et al. KABR: In-situ dataset for kenyan animal behavior recognition from drone videos. In WACV, 2024.
  52. 52.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In ICLR, 2015.
  53. 53.Krishna, R., Hata, K., Ren, F., Fei-Fei, L., and Carlos Niebles, J. Dense-captioning events in videos. In ICCV, 2017.
  54. 54.Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W. Set Transformer: A framework for attention-based permutation-invariant neural networks. In ICML, 2019.
  55. 55.Lei, J., Berg, T. L., and Bansal, M. Revealing single frame bias for video-and-language learning. In ACL, 2023.
  56. 56.Li, A., Thotakuri, M., Ross, D. A., Carreira, J., Vostrikov, A., and Zisserman, A. The AVA-Kinetics localized human actions video dataset. arXiv preprint arXiv:2005.00214, 2020.
  57. 57.Li, J., Li, D., Xiong, C., and Hoi, S. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022.
  58. 58.Li, J., Li, D., Savarese, S., and Hoi, S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023a.
  59. 59.Li, K., Wang, Y., Li, Y., Wang, Y., He, Y., Wang, L., and Qiao, Y. Unmasked teacher: Towards training-efficient video foundation models. In ICCV, 2023b.
  60. 60.Li, L., Gan, Z., Lin, K., Lin, C.-C., Liu, Z., Liu, C., and Wang, L. LAVENDER: Unifying video-language understanding as masked language modeling. In CVPR, 2023c.
  61. 61.Li, W., Zhu, L., Wen, L., and Yang, Y. DeCap: Decoding CLIP latents for zero-shot captioning via text-only training. In ICLR, 2023d.
  62. 62.Li, Y., Li, Y., and Vasconcelos, N. RESOUND: Towards action recognition without representation bias. In ECCV, 2018.
  63. 63.Li, Y., Fan, H., Hu, R., Feichtenhofer, C., and He, K. Scaling language-image pre-training via masking. In CVPR, 2023e.
  64. 64.Li, Y., Wang, C., and Jia, J. LLaMA-VID: An image is worth 2 tokens in large language models. arXiv preprint arXiv:2311.17043, 2023f.
  65. 65.Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, P., and Yuan, L. Video-LLaVA: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023a.
  66. 66.Lin, K. Q., Wang, J., Soldan, M., Wray, M., Yan, R., XU, E. Z., Gao, D., Tu, R.-C., Zhao, W., Kong, W., et al. Egocentric video-language pretraining. In NeurIPS, 2022.
  67. 67.Lin, W., Karlinsky, L., Shvetsova, N., Possegger, H., Kozinski, M., Panda, R., Feris, R., Kuehne, H., and Bischof, H. Match, expand and improve: Unsupervised finetuning for zero-shot action recognition with language knowledge. In ICCV, 2023b.
  68. 68.Liu, R., Huang, J., Li, G., Feng, J., Wu, X., and Li, T. H. Revisiting temporal modeling for clip-based image-to-video knowledge transferring. In CVPR, 2023.
  69. 69.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In ICLR, 2019.
  70. 70.Luo, H., Ji, L., Zhong, M., Chen, Y., Lei, W., Duan, N., and Li, T. CLIP4Clip: An empirical study of clip for end to end video clip retrieval and captioning. Neurocomputing, 508:293–304, 2022.
  71. 71.Ma, X., Kaufhold, S. P., Su, J., Zhu, W., Terwilliger, J., Meza, A., Zhu, Y., Rossano, F., and Wang, Y. ChimpACT: A longitudinal dataset for understanding chimpanzee behaviors. arXiv preprint arXiv:2310.16447, 2023.
  72. 72.Maaz, M., Rasheed, H., Khan, S., and Khan, F. S. VideoChatGPT: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023.
  73. 73.McCloskey, M. and Cohen, N. J. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pp. 109–165. Elsevier, 1989.
  74. 74.Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., and Sivic, J. HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips. In ICCV, 2019.
  75. 75.Momeni, L., Caron, M., Nagrani, A., Zisserman, A., and Schmid, C. Verbs in action: Improving verb understanding in video-language models. In ICCV, 2023.
  76. 76.Monfort, M., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S. A., Yan, T., Brown, L., Fan, Q., Gutfreund, D., Vondrick, C., et al. Moments in Time dataset: one million videos for event understanding. IEEE TPAMI, 42 (2):502–508, 2019.
  77. 77.Monfort, M., Jin, S., Liu, A., Harwath, D., Feris, R., Glass, J., and Oliva, A. Spoken Moments: Learning joint audio-visual representations from video descriptions. In CVPR, 2021.
  78. 78.Nagrani, A., Seo, P. H., Seybold, B., Hauth, A., Manen, S., Sun, C., and Schmid, C. Learning audio-video modalities from image captions. In ECCV, 2022.
  79. 79.Ni, B., Peng, H., Chen, M., Zhang, S., Meng, G., Fu, J., Xiang, S., and Ling, H. Expanding language-image pre-trained models for general video recognition. In ECCV, 2022.
  80. 80.Noroozi, M. and Favaro, P. Unsupervised learning of visual representations by solving Jigsaw puzzles. In ECCV, 2016.
  81. 81.Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. DINOv2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
  82. 82.Peng, Z., Dong, L., Bao, H., Ye, Q., and Wei, F. BEiT v2: Masked image modeling with vector-quantized visual tokenizers. arXiv preprint arXiv:2208.06366, 2022.
  83. 83.Piergiovanni, A., Nobel, I., Kim, D., Ryoo, M. S., Gomes, V., and Angelova, A. Mirasol3B: A multimodal autoregressive model for time-aligned and contextual modalities. arXiv preprint arXiv:2311.05698, 2023.
  84. 84.Pitcher-Cooper, C., Seth, M., Kao, B., Coughlan, J. M., and Yoon, I. You Described, We Archived: A rich audio description dataset. Journal on Technology and Persons with Disabilities, 2023.
  85. 85.Qian, R., Meng, T., Gong, B., Yang, M.-H., Wang, H., Belongie, S., and Cui, Y. Spatiotemporal contrastive video representation learning. In CVPR, 2021.
  86. 86.Qian, R., Li, Y., Yuan, L., Gong, B., Liu, T., Brown, M., Yang, M.-H., Adam, H., and Cui, Y. On temporal granularity in self-supervised video representation learning. In BMVC, 2022.
  87. 87.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  88. 88.Recasens, A., Luc, P., Alayrac, J.-B., Wang, L., Strub, F., Tallec, C., Malinowski, M., Patřaucěan, V., Altché, F., Valko, M., et al. Broaden your views for self-supervised video learning. In ICCV, 2021.
  89. 89.Regneri, M., Rohrbach, M., Wetzel, D., Thater, S., Schiele, B., and Pinkal, M. Grounding action descriptions in videos. Transactions of the Association for Computational Linguistics, 1:25–36, 2013.
  90. 90.Ren, S., He, K., Girshick, R., and Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. In NeurIPS, 2015.
  91. 91.Sevilla-Lara, L., Zha, S., Yan, Z., Goswami, V., Feiszli, M., and Torresani, L. Only time can tell: Discovering temporal data for temporal modeling. In WACV, 2021.
  92. 92.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. In ICML, 2018.
  93. 93.Sigurdsson, G. A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A. Hollywood in Homes: Crowdsourcing data collection for activity understanding. In ECCV, 2016.
  94. 94.Singh, A., Chakraborty, O., Varshney, A., Panda, R., Feris, R., Saenko, K., and Das, A. Semi-supervised action recognition with temporal contrastive learning. In CVPR, 2021.
  95. 95.Singh, A., Hu, R., Goswami, V., Couairon, G., Galuba, W., Rohrbach, M., and Kiela, D. FLAVA: A foundational language and vision alignment model. In CVPR, 2022.
  96. 96.Stroud, J. C., Lu, Z., Sun, C., Deng, J., Sukthankar, R., Schmid, C., and Ross, D. A. Learning video representations from textual web supervision. arXiv preprint arXiv:2007.14937, 2020.
  97. 97.Sun, J. J., Karigo, T., Chakraborty, D., Mohanty, S. P., Wild, B., Sun, Q., Chen, C., Anderson, D. J., Perona, P., Yue, Y., et al. The multi-agent behavior dataset: Mouse dyadic social interactions. In NeurIPS D&B, 2021a.
  98. 98.Sun, J. J., Kennedy, A., Zhan, E., Anderson, D. J., Yue, Y., and Perona, P. Task programming: Learning data efficient behavior representations. In CVPR, 2021b.
  99. 99.Tan, J., Wang, C., Li, B., Li, Q., Ouyang, W., Yin, C., and Yan, J. Equalization loss for long-tailed object recognition. In CVPR, 2020.
  100. 100.Tang, Y., Ding, D., Rao, Y., Zheng, Y., Zhang, D., Zhao, L., Lu, J., and Zhou, J. COIN: A large-scale dataset for comprehensive instructional video analysis. In CVPR, 2019.
  101. 101.Tang, Y., Bi, J., Xu, S., Song, L., Liang, S., Wang, T., Zhang, D., An, J., Lin, J., Zhu, R., et al. Video understanding with large language models: A survey. arXiv preprint arXiv:2312.17432, 2023.
  102. 102.Tong, Z., Song, Y., Wang, J., and Wang, L. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In NeurIPS, 2022.
  103. 103.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In NeurIPS, 2017.
  104. 104.Voigtlaender, P., Changpinyo, S., Pont-Tuset, J., Soricut, R., and Ferrari, V. Connecting vision and language with video localized narratives. In CVPR, 2023.
  105. 105.Wang, J., Chen, D., Wu, Z., Luo, C., Zhou, L., Zhao, Y., Xie, Y., Liu, C., Jiang, Y.-G., and Yuan, L. OmniVL: One foundation model for image-language and video-language tasks. In NeurIPS, 2022a.
  106. 106.Wang, J., Ge, Y., Yan, R., Ge, Y., Lin, K. Q., Tsutsui, S., Lin, X., Cai, G., Wu, J., Shan, Y., et al. All in one: Exploring unified video-language pre-training. In CVPR, 2023a.
  107. 107.Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., and Qiao, Y. VideoMAE v2: Scaling video masked autoencoders with dual masking. In CVPR, 2023b.
  108. 108.Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Jiang, Y.-G., Zhou, L., and Yuan, L. BEVT: BERT pre-training of video transformers. In CVPR, 2022b.
  109. 109.Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Yuan, L., and Jiang, Y.-G. Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning. In CVPR, 2023c.
  110. 110.Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al. Image as a foreign language: BEiT pretraining for vision and vision-language tasks. In CVPR, 2023d.
  111. 111.Wang, X., Wu, J., Chen, J., Li, L., Wang, Y.-F., and Wang, W. Y. VATEX: A large-scale, high-quality multilingual dataset for video-and-language research. In ICCV, 2019.
  112. 112.Wang, Y., Li, K., Li, Y., He, Y., Huang, B., Zhao, Z., Zhang, H., Xu, J., Liu, Y., Wang, Z., et al. InternVideo: General video foundation models via generative and discriminative learning. arXiv preprint arXiv:2212.03191, 2022c.
  113. 113.Wang, Y., He, Y., Li, Y., Li, K., Yu, J., Ma, X., Chen, X., Wang, Y., Luo, P., Liu, Z., et al. InternVid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023e.
  114. 114.Wang, Z., Li, M., Xu, R., Zhou, L., Lei, J., Lin, X., Wang, S., Yang, Z., Zhu, C., Hoiem, D., et al. Language models with image descriptors are strong few-shot video-language learners. In NeurIPS, 2022d.
  115. 115.Wang, Z., Blume, A., Li, S., Liu, G., Cho, J., Tang, Z., Bansal, M., and Ji, H. Paxion: Patching action knowledge in video-language foundation models. In NeurIPS, 2023f.
  116. 116.Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., and Feichtenhofer, C. Masked feature prediction for self-supervised visual pre-training. In CVPR, 2022.
  117. 117.Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021.
  118. 118.Wu, W., Sun, Z., and Ouyang, W. Revisiting classifier: Transferring vision-language models for video recognition. In AAAI, 2023.
  119. 119.Wu, Z. and Palmer, M. Verb semantics and lexical selection. In ACL, 1994.
  120. 120.Wu, Z., Weng, Z., Peng, W., Yang, X., Li, A., Davis, L. S., and Jiang, Y.-G. Building an open-vocabulary video CLIP model with better architectures, optimization and data. IEEE TPAMI, 2024.
  121. 121.Xiao, J., Shang, X., Yao, A., and Chua, T.-S. NExT-QA: Next phase of question-answering to explaining temporal actions. In CVPR, 2021.
  122. 122.Xiong, Y., Zhao, L., Gong, B., Yang, M.-H., Schroff, F., Liu, T., Hsieh, C.-J., and Yuan, L. Spatiotemporally discriminative video-language pre-training with text grounding. arXiv preprint arXiv:2303.16341, 2023.
  123. 123.Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y. Video question answering via gradually refined attention over appearance and motion. In ACM MM, 2017.
  124. 124.Xu, H., Ghosh, G., Huang, P.-Y., Okhonko, D., Aghajanyan, A., Metze, F., Zettlemoyer, L., and Feichtenhofer, C. VideoCLIP: Contrastive pre-training for zero-shot video-text understanding. In EMNLP, 2021.
  125. 125.Xu, H., Ye, Q., Yan, M., Shi, Y., Ye, J., Xu, Y., Li, C., Bi, B., Qian, Q., Wang, W., et al. mPLUG-2: A modularized multi-modal foundation model across text, image and video. In ICML, 2023.
  126. 126.Xu, J., Mei, T., Yao, T., and Rui, Y. MSR-VTT: A large video description dataset for bridging video and language. In CVPR, 2016.
  127. 127.Xu, M., Zhao, C., Rojas, D. S., Thabet, A., and Ghanem, B. G-TAD: Sub-graph localization for temporal action detection. In CVPR, 2020.
  128. 128.Xue, H., Hang, T., Zeng, Y., Sun, Y., Liu, B., Yang, H., Fu, J., and Guo, B. Advancing high-resolution video-language representation with large-scale video transcriptions. In CVPR, 2022.
  129. 129.Xue, H., Sun, Y., Liu, B., Fu, J., Song, R., Li, H., and Luo, J. CLIP-ViP: Adapting pre-trained image-text model to video-language representation alignment. In ICLR, 2023.
  130. 130.Yan, S., Zhu, T., Wang, Z., Cao, Y., Zhang, M., Ghosh, S., Wu, Y., and Yu, J. VideoCoCa: Video-text modeling with zero-shot transfer from contrastive captioners. arXiv preprint arXiv:2212.04979, 2022.
  131. 131.Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C. Zero-shot video question answering via frozen bidirectional language models. In NeurIPS, 2022.
  132. 132.Yarom, M., Bitton, Y., Changpinyo, S., Aharoni, R., Herzig, J., Lang, O., Ofek, E., and Szpektor, I. What you see is what you read? improving text-image alignment evaluation. In NeurIPS, 2023.
  133. 133.Ye, Q., Xu, G., Yan, M., Xu, H., Qian, Q., Zhang, J., and Huang, F. HiTeA: Hierarchical temporal-aware video-language pre-training. In ICCV, 2023.
  134. 134.Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., and Wu, Y. CoCa: Contrastive captioners are image-text foundation models. TMLR, 2022. ISSN 2835-8856. URL https://openreview.net/forum?id=Ee277P3AYC.
  135. 135.Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., Li, C., et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021.
  136. 136.Yuan, L., Qian, R., Cui, Y., Gong, B., Schroff, F., Yang, M.-H., Adam, H., and Liu, T. Contextualized spatio-temporal contrastive learning with self-supervision. In CVPR, 2022.
  137. 137.Yuan, L., Gundavarapu, N. B., Zhao, L., Zhou, H., Cui, Y., Jiang, L., Yang, X., Jia, M., Weyand, T., Friedman, L., et al. VideoGLUE: Video general understanding evaluation of foundation models. arXiv preprint arXiv:2307.03166, 2023.
  138. 138.Zellers, R., Lu, X., Hessel, J., Yu, Y., Park, J. S., Cao, J., Farhadi, A., and Choi, Y. MERLOT: Multimodal neural script knowledge models. In NeurIPS, 2021.
  139. 139.Zellers, R., Lu, J., Lu, X., Yu, Y., Zhao, Y., Salehi, M., Kusupati, A., Hessel, J., Farhadi, A., and Choi, Y. MERLOT Reserve: Neural script knowledge through vision and language and sound. In CVPR, 2022.
  140. 140.Zeng, A., Attarian, M., Ichter, B., Choromanski, K., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M., Sindhwani, V., et al. Socratic models: Composing zero-shot multimodal reasoning with language. arXiv preprint arXiv:2204.00598, 2022.
  141. 141.Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. Scaling vision transformers. In CVPR, 2022a.
  142. 142.Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L. LiT: Zero-shot transfer with locked-image text tuning. In CVPR, 2022b.
  143. 143.Zhang, H., Li, X., and Bing, L. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023a.
  144. 144.Zhang, X., Chen, J., Yuan, J., Chen, Q., Wang, J., Wang, X., Han, S., Chen, X., Pi, J., Yao, K., Han, J., Ding, E., and Wang, J. CAE v2: Context autoencoder with CLIP latent alignment. TMLR, 2023b. ISSN 2835-8856. URL https://openreview.net/forum?id=f36LaK7M0F.
  145. 145.Zhao, Y., Zhao, L., Zhou, X., Wu, J., Chu, C.-T., Miao, H., Schroff, F., Adam, H., Liu, T., Gong, B., et al. Distilling vision-language models on millions of videos. arXiv preprint arXiv:2401.06129, 2024.
  146. 146.Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. iBOT: Image BERT pre-training with online tokenizer. In ICLR, 2022.
  147. 147.Zhou, L., Xu, C., and Corso, J. Towards automatic learning of procedures from web instructional videos. In AAAI, 2018.
  148. 148.Zhu, B., Lin, B., Ning, M., Yan, Y., Cui, J., Wang, H., Pang, Y., Jiang, W., Zhang, J., Li, Z., et al. LanguageBind: Extending video-language pretraining to N-modality by language-based semantic alignment. In ICLR, 2024.

Citation

MLA
Zhao, L., et al. “VideoPrism: A Foundational Visual Encoder for Video Understanding”. arXiv, 2024, https://doi.org/10.48550/arxiv.2402.13217.
APA
Zhao, L., Gundavarapu, N. B., Yuan, L., Zhou, H., Yan, S., Sun, J. J., Friedman, L., Qian, R., Weyand, T., Zhao, Y., Hornung, R., Schroff, F., Yang, M.-H., Ross, D. A., Wang, H., Adam, H., Sirotenko, M., Liu, T., & Gong, B. (2024). VideoPrism: A Foundational Visual Encoder for Video Understanding. arXiv. https://doi.org/10.48550/arxiv.2402.13217
Chicago
Zhao, L., N. B. Gundavarapu, L. Yuan, et al. 2024. “VideoPrism: A Foundational Visual Encoder for Video Understanding”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.2402.13217.
Harvard
Zhao, L. et al. (2024) “VideoPrism: A Foundational Visual Encoder for Video Understanding”. arXiv. Available at: https://doi.org/10.48550/arxiv.2402.13217.
Vancouver
1. Zhao L, Gundavarapu NB, Yuan L, et al (2024) VideoPrism: A Foundational Visual Encoder for Video Understanding. https://doi.org/10.48550/arxiv.2402.13217

BibTeX

@misc{https://doi.org/10.48550/arxiv.2402.13217,
  doi = {10.48550/ARXIV.2402.13217},
  url = {https://arxiv.org/abs/2402.13217},
  author = {Zhao, Long and Gundavarapu, Nitesh B. and Yuan, Liangzhe and Zhou, Hao and Yan, Shen and Sun, Jennifer J. and Friedman, Luke and Qian, Rui and Weyand, Tobias and Zhao, Yue and Hornung, Rachel and Schroff, Florian and Yang, Ming-Hsuan and Ross, David A. and Wang, Huisheng and Adam, Hartwig and Sirotenko, Mikhail and Liu, Ting and Gong, Boqing},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {VideoPrism: A Foundational Visual Encoder for Video Understanding},
  publisher = {arXiv},
  year = {2024},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/