All in One: Exploring Unified Video-Language Pre-Training

Jinpeng WangYixiao GeRui YanYuying GeKevin Qinghong LinSatoshi TsutsuiXudong LinGuanyu CaiJianping WuYing Shan

article2023CVPR257 citations

Introduces an end-to-end unified video-language framework that processes raw video and text inputs within a single shared transformer backbone via a parameter-free temporal token rolling mechanism, significantly reducing computational cost while matching competitive multi-network models.

Listen

Mainstream artificial intelligence systems that connect video and text have grown increasingly complex and resource-intensive. Most standard architectures rely on separate, heavy feature extractors for video and text before combining them in a dedicated fusion network. While effective, this multi-stage approach demands massive computational power, large memory footprints, and slow processing speeds, which creates significant bottlenecks for real-world deployment and large-scale applications.

The article introduces and evaluates the All-in-one Transformer, a unified, lightweight framework designed to process raw video pixels and text tokens end-to-end within a single shared neural network. The central objective is to demonstrate that a single, modality-agnostic architecture can achieve state-of-the-art multimodal performance while eliminating the need for dedicated unimodal encoders and complex fusion layers.

To achieve this, the authors developed a parameter-free temporal token rolling technique that exchanges visual information across video frames without adding parameters or computational bloat. Rather than analyzing dozens of frames per video, the approach samples only three frames per clip. The model was pre-trained using standard video-text matching and masked language modeling objectives on large public video and image datasets (including WebVid, HowTo100M, and CC3M) via a balanced co-training strategy. The researchers then evaluated the system across ten benchmark datasets covering video question answering, text-to-video retrieval, multiple-choice reasoning, captioning, and action recognition.

The evaluation produced four primary findings. First, the All-in-one Transformer achieved competitive or state-of-the-art results across downstream tasks while requiring substantially fewer parameters and computational operations—often using less than half the parameters and a fraction of the floating-point operations of leading alternatives. Second, the temporal token rolling module reduced self-attention computational complexity by roughly two-thirds compared to standard flattened token processing while effectively capturing motion cues. Third, on text-to-video retrieval benchmarks such as MSR-VTT, the model achieved a top-1 recall of 41.8%, outperforming specialized multi-component systems. Finally, the unified design successfully supported fast unimodal extraction for retrieval via contrastive learning, reducing matching complexity from multiplicative to additive.

These findings demonstrate that heavyweight unimodal encoders are not essential for high-performance video-language understanding. By unifying modalities in a single backbone, organizations can significantly reduce hardware infrastructure costs, simplify software pipelines, and lower inference latency. This streamlined approach makes multimodal video search and analysis feasible in high-throughput environments where multi-model pipelines are commercially impractical.

Organizations developing or deploying multimodal AI should consider transitioning from fragmented multi-encoder architectures toward unified transformer backbones to capture efficiency and maintenance gains. For future development, engineering teams should evaluate the model on specific domain data and explore extending the architecture to broader single-modality tasks. Further research should focus on refining fine-grained word-region alignment and testing whether performance holds across tasks requiring fine-grained temporal sequence understanding.

Cover for All in One: Exploring Unified Video-Language Pre-Training

Abstract

Mainstream Video-Language Pre-training (VLP) models [10, 26, 64] consist of three parts, a video encoder, a text encoder, and a video-text fusion Transformer. They pursue better performance via utilizing heavier unimodal encoders or multimodal fusion Transformers, resulting in increased parameters with lower efficiency in downstream tasks. In this work, we for the first time introduce an end-to-end VLP model, namely all-in-one Transformer, that embeds raw video and textual signals into joint representations using a unified backbone architecture. We argue that the unique temporal information of video data turns out to be a key barrier hindering the design of a modality-agnostic Transformer. To overcome the challenge, we introduce a novel and effective token rolling operation to encode temporal representations from video clips in a non-parametric manner. The careful design enables the representation learning of both video-text multimodal inputs and unimodal inputs using a unified model. Our pre-trained all-in-one Transformer is transferred to various downstream video-text tasks after fine-tuning, including text-video retrieval, video-question answering, multiple choice and video captioning. State-of-the-art performances with the minimal model FLOPs on ten datasets demonstrate the superiority of our method compared to the competitive counterparts. The code and pretrained models are available at https://github.com/showlab/all-in-one.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Unified Video-language Transformer
  • 3.2. Temporal Token Rolling
  • 3.3. Training Objectives
  • 3.3.1 Pre-training
  • 3.3.2 BVTC for fast downstream retrieval
  • 3.4. Image Video Co-Training
  • 3.5. Setup
  • 3.5.1 Model Variants.
  • 3.5.2 Pre-training & Fine-tuning.
  • 3.6. Downstream Tasks Settings
  • 4. Main Results
  • 4.1. Multimodal Understanding Tasks
  • 4.1.1 Video-question Answering.
  • 4.1.2 Multiple-choice.
  • 4.2. Video-text Alignment Task
  • 4.2.1 Text-to-video Retrieval.
  • 4.3. Video Captioning and Action Recognition
  • 4.3.1 Video Captioning.
  • 4.3.2 Action Recognition via Linear Probe.
  • 5. Visualization
  • 6. Conclusions & Future Work
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Unified raw-input All-in-one Transformer

    model/method

    The All-in-one Transformer is an end-to-end video-language pre-training architecture that processes raw video pixels and raw text tokens with one shared Transformer backbone, without a separate visual encoder, language encoder, or modality-specific fusion Transformer. For each video, SS frames are sparsely sampled, each frame is divided into visual patches, and the patches are linearly projected into visual tokens. Learnable spatial-temporal position embeddings and modality-type embeddings are added to the visual tokens. Text words are mapped to tokens with a word-embedding layer and receive modality-type embeddings as well.

    The text tokens and visual tokens from the sampled frames are concatenated and passed through NN Transformer blocks. In block dd, where d∈{1,…,N}d\in\{1,\ldots,N\}, zd−1z^{d-1} is the incoming token sequence, zrd−1z_r^{d-1} is its temporally rolled version, MSA is multi-head self-attention, and MLP is a multilayer perceptron:

    zrd−1=TTR⁡(zd−1),zd=MLP⁡(MSA⁡(zrd−1)).z_r^{d-1}=\operatorname{TTR}(z^{d-1}),\qquad z^d=\operatorname{MLP}(\operatorname{MSA}(z_r^{d-1})).

    Here TTR denotes the parameter-free Temporal Token Rolling module. The Transformer weights are initialized from a pre-trained ViT or DeiT, while the model adds only a lightweight text tokenizer and task-specific heads beyond the shared Transformer. Because the backbone is modality-agnostic, it can process video-text pairs as well as video-only or text-only inputs.

  2. Knowl 2 — Parameter-free temporal token rolling

    model/method

    Temporal Token Rolling enables the shared Transformer to model video time without adding temporal-attention parameters or a modality-specific video encoder. Let mm be the number of text tokens, nn the number of visual patch tokens per frame, and SS the number of sampled frames. At the input to each Transformer block, a selected portion of visual tokens is cyclically shifted by one position along the temporal dimension, while the remaining visual tokens stay in their original frame. Self-attention is then applied to each frame-level group containing m+nm+n tokens, so text tokens can attend to visual tokens originating from different frames.

    A naive flattened design would apply self-attention to all m+Snm+Sn tokens simultaneously, with computational complexity O((m+Sn)2)O((m+Sn)^2). Temporal Token Rolling instead uses complexity O(S(m+n)2)O(S(m+n)^2), which is approximately a factor of SS cheaper when the visual-token count dominates the text-token count. The operation is non-parametric and does not increase the Transformer block's parameter count; it exposes progressively longer temporal dependencies through repeated rolling and self-attention blocks.

  3. Knowl 3 — Video-text pre-training objectives

    model/method

    All-in-one pre-training uses video-text matching and masked language modeling on paired video-text examples. For video-text matching, a paired video is replaced by a video from another example with probability 0.50.5, creating positive and negative pairs. A linear head applied to the final [CLS] representation predicts one of two matching labels, and the negative log-likelihood is used as the matching loss.

    For masked language modeling, each text token is independently selected for masking with probability 0.150.15. The model predicts the original token using the remaining text tokens together with the visual tokens from the video. Thus, the same shared Transformer is trained both to determine whether a video and sentence belong together and to use video context for recovering masked words.

  4. Knowl 4 — Backbone-shared video-text contrastive retrieval

    model/method

    The paper introduces Backbone-shared Video-text Contrastive learning (BVTC) to make retrieval practical with a one-stream video-language model. During BVTC, a video and a text are fed independently through the same All-in-one Transformer backbone rather than jointly. The resulting video and text features are passed through separate modality-specific projection heads into a common embedding space, where a symmetric contrastive loss is applied in both text-to-video and video-to-text directions.

    At retrieval time, each candidate video and text is encoded once and their projected features are compared with cosine similarity. For kk videos and ll texts, this avoids evaluating a separate joint Transformer for every video-text pair: the paper reports a reduction in retrieval computation from O(kl)O(k l) to O(k+l)O(k+l) for the feature-extraction pipeline. This preserves the shared-backbone representation while providing the fast candidate scoring characteristic of dual-stream retrieval systems.

  5. Knowl 5 — Balanced image-video co-training

    model/method

    All-in-one can be trained jointly on image-text and video-text data, but treating every image as a one-frame video was found to damage temporal learning and destabilize training. The proposed alternative samples half of each training batch from image-text examples and half from video-text examples. Image-text inputs bypass Temporal Token Rolling and enter the self-attention blocks directly, whereas video-text inputs pass through Temporal Token Rolling before self-attention. A single shared pretext head is used for both modalities.

    The weighted co-training loss is

    Lct=∑iwvideoiL(yvideoi)+∑jwimagejL(yimagej)\mathcal{L}_{ct}=\sum_i w_{\mathrm{video}}^i\mathcal{L}(y_{\mathrm{video}}^i)+\sum_j w_{\mathrm{image}}^j\mathcal{L}(y_{\mathrm{image}}^j),

    where ii indexes video-text examples, jj indexes image-text examples, yvideoiy_{\mathrm{video}}^i and yimagejy_{\mathrm{image}}^j are their pretext-task labels, L\mathcal{L} is the per-example pretext loss, and wvideoiw_{\mathrm{video}}^i and wimagejw_{\mathrm{image}}^j are the corresponding sample weights.

  6. Knowl 6 — Training configuration and model scaling

    experimental setup

    The default experiments pre-train All-in-one on WebVid2.5M and HowTo100M; an All-in-one-B+ variant additionally uses CC3M, and an All-in-one-B* variant uses CC3M, WebVid2.5M, and YT-Temporal180M. Text is tokenized with the bert-base-uncased tokenizer. The default video input consists of three randomly sampled frames, each resized to 224×224224\times224 pixels. Image-video co-training uses CC3M as the additional image-text source.

    The three reported model configurations are All-in-one-Ti with embedding dimension 192, 3 attention heads, 12M parameters, and throughput 745; All-in-one-S with embedding dimension 384, 6 heads, 33M parameters, and throughput 285; and All-in-one-B with embedding dimension 768, 12 heads, 110M parameters, and throughput 89. Throughput is reported for videos at 224×224224\times224 resolution. All-in-one-B is the default model unless otherwise specified. The model is transferred to text-to-video retrieval, video question answering, multiple-choice understanding, video captioning, and action recognition.

  7. Knowl 7 — Video question answering and multiple-choice performance

    empirical result

    All-in-one achieves strong video question answering performance while using only three sampled frames in most experiments and one frame in one TGIF-QA setting. On the TGIF-QA Action, Transition, and FrameQA sub-tasks, All-in-one-B using one frame obtains accuracies of 92.992.9, 94.294.2, and 62.562.5, respectively. All-in-one-B+ trained on CC3M, WebVid2.5M, and HowTo100M with three frames obtains 96.396.3, 95.595.5, and 67.367.3; the corresponding All-in-one-B* trained with YT-Temporal180M instead of HowTo100M obtains 95.595.5, 94.794.7, and 66.366.3. For comparison, VIOLET obtains 87.187.1 and 93.693.6 on the Action and Transition sub-tasks while using 16 frames.

    On open-ended question answering, All-in-one-B+ obtains 44.644.6 accuracy on MSRVTT-QA, 48.248.2 on MSVD-QA, and 71.571.5 on TVQA; All-in-one-B* obtains 46.846.8, 48.348.3, and 72.072.0, respectively. On the MSRVTT multiple-choice task, All-in-one-B and All-in-one-B+ obtain 91.491.4 and 91.991.9 accuracy, while the All-in-one-B+ zero-shot result is 82.282.2. On first-view Ego4D multiple-choice evaluation, All-in-one-B obtains 36.5236.52 zero-shot accuracy and 65.8965.89 after fine-tuning, compared with 32.4732.47 and 60.3260.32 for Frozen and 27.3427.34 and 59.4459.44 for VATT.

  8. Knowl 8 — Text-to-video retrieval with low computational cost

    empirical result

    After fine-tuning, All-in-one-B+ produces strong text-to-video retrieval results using 110M parameters, three frames, and 58.7G FLOPs. On the MSRVTT 9K/7K training splits, the reported (R@1,R@5,R@10)(R@1,R@5,R@10) values are (39.7,67.8,76.1)(39.7,67.8,76.1) and (35.9,66.1,75.1)(35.9,66.1,75.1) when pre-trained on CC3M and WebVid2.5M, and (41.8,68.5,76.7)(41.8,68.5,76.7) and (37.3,66.4,75.6)(37.3,66.4,75.6) when HowTo100M is added. All-in-one-B trained on HowTo100M alone obtains (29.5,63.3,71.9)(29.5,63.3,71.9) on the 9K split and (26.5,59.4,69.8)(26.5,59.4,69.8) on the 7K split.

    On ActivityNet Caption, All-in-one-B obtains (R@1,R@5,R@10,MdR)=(21.5,50.3,65.5,6.0)(R@1,R@5,R@10,\mathrm{MdR})=(21.5,50.3,65.5,6.0) and All-in-one-B+ obtains (22.4,53.7,67.7,5.0)(22.4,53.7,67.7,5.0). On DiDeMo, the corresponding results are (31.2,60.5,72.1,3.0)(31.2,60.5,72.1,3.0) and (32.7,61.4,73.5,3.0)(32.7,61.4,73.5,3.0). The computational advantage is substantial relative to the compared one-stream systems: VIOLET uses 198M parameters, 16 frames, and 351.4G FLOPs; Frozen uses 232M parameters, 8 frames, and 217.3G FLOPs; and ClipBERT uses 137M parameters, 16 effective frames, and 183.2G FLOPs.

  9. Knowl 9 — Single-modality transfer and video captioning

    empirical result

    The unified representation transfers to action recognition with a frozen All-in-one backbone and learned linear classifiers on the final [CLS] representation. With three frames, All-in-one-B obtains Top-1/Top-5/Top-10 accuracies of 49.8/79.8/90.749.8/79.8/90.7 on Kinetics-400, 51.9/84.1/93.451.9/84.1/93.4 on HMDB51, and 81.1/93.8/95.581.1/93.8/95.5 on UCF101. With eight frames, the corresponding results are 52.4/83.2/92.952.4/83.2/92.9, 54.7/88.2/95.254.7/88.2/95.2, and 82.8/95.1/96.982.8/95.1/96.9. Training separate pretext heads for image-text and video-text inputs further gives 53.2/83.5/92.753.2/83.5/92.7, 55.2/89.1/95.855.2/89.1/95.8, and 84.1/95.7/97.884.1/95.7/97.8 on the three datasets.

    For video captioning, a lightweight language-modeling head is added to All-in-one. All-in-one-B+ obtains BLEU-4/METEOR/CIDEr scores of 12.5/20.4/56.312.5/20.4/56.3 on TVC and 11.2/13.9/114.511.2/13.9/114.5 on YouCook2. These results demonstrate that the same unified backbone supports both discriminative single-modality transfer and generative video captioning.

  10. Knowl 10 — Observed alignment behavior and limitations

    limitation

    The paper's qualitative analysis masks verbs and nouns in video descriptions and examines the visual regions associated with the predictions. All-in-one usually recovers the correct masked words; when it is wrong, the prediction can be semantically close to the target, such as predicting a synonym-like word. Temporal token rolling also causes verb predictions such as walking or waving to attend to motion-relevant regions rather than only static central patches.

    The authors identify fine-grained word-to-region matching as an unresolved difficulty. They also state that temporal modeling in the unified architecture has not been fully investigated and that applying All-in-one more broadly to other single-modality tasks remains future work.

Coverage note — Detailed supplementary ablations, complete baseline rows, and extended qualitative examples were omitted because they are not load-bearing for reconstructing the proposed architecture, training procedures, and principal findings.

References

  1. 1.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems, 34, 2021.
  2. 2.Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein. Noise estimation using density estimation for self-supervised multimodal learning. In AAAI, 2021.
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1728–1738, 2021.
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding. arXiv preprint arXiv:2102.05095, page 4, 2021.
  5. 5.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  6. 6.David Chen and William B Dolan. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, pages 190–200, 2011.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  9. 9.Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang. Heterogeneous memory enhanced multimodal attention model for video question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1999–2007, 2019.
  10. 10.Tsu-Jui Fu, Linjie Li, Zhe Gan, Kevin Lin, William Yang Wang, Lijuan Wang, and Zicheng Liu. Violet: End-to-end video-language transformers with masked visual-token modeling. arXiv preprint arXiv:2111.12681, 2021.
  11. 11.Yuying Ge, Yixiao Ge, Xihui Liu, Jinpeng Wang, Jianping Wu, Ying Shan, Xiaohu Qie, and Ping Luo. Miles: visual bert pre-training with injected language semantics for video-text retrieval. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXV, pages 691–708. Springer, 2022.
  12. 12.Shijie Geng, Ji Zhang, Zuohui Fu, Peng Gao, Hang Zhang, and Gerard de Melo. Character matters: Video story understanding with character-aware relations. arXiv preprint arXiv:2005.08646, 2020.
  13. 13.Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra. Omnivore: A single model for many visual modalities. arXiv preprint arXiv:2201.08377, 2022.
  14. 14.Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. arXiv preprint arXiv:2110.07058, 2021.
  15. 15.Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2758–2766, 2017.
  16. 16.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021.
  17. 17.Jianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao, and Yue Gao. Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 11101–11108, 2020.
  18. 18.Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D Yoo. Gaining extra supervision via multi-task learning for multi-modal video question answering. In 2019 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2019.
  19. 19.Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D Yoo. Progressive attention memory network for movie story question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8337–8346, 2019.
  20. 20.Junyeong Kim, Minuk Ma, Trung Pham, Kyungsu Kim, and Chang D Yoo. Modality shifting attention network for multi-modal video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10106–10115, 2020.
  21. 21.Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583–5594. PMLR, 2021.
  22. 22.Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision, pages 706–715, 2017.
  23. 23.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, pages 32–73, 2017.
  24. 24.Thomas S. Kuhn. The structure of scientific revolutions. University of Chicago Press, 1962.
  25. 25.Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. Hierarchical conditional relation networks for video question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9972–9981, 2020.
  26. 26.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7331–7341, 2021.
  27. 27.Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. Tvqa+: Spatio-temporal grounding for video question answering. arXiv preprint arXiv:1904.11574, 2019.
  28. 28.Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, et al. Value: A multi-task benchmark for video-and-language understanding evaluation. arXiv preprint arXiv:2106.04632, 2021.
  29. 29.Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. arXiv preprint arXiv:2012.15409, 2020.
  30. 30.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121–137. Springer, 2020.
  31. 31.Ji Lin, Chuang Gan, and Song Han. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7083–7093, 2019.
  32. 32.Kevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed, Zhe Gan, Zicheng Liu, Yumao Lu, and Lijuan Wang. Swinbert: End-to-end transformers with sparse attention for video captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17949–17958, 2022.
  33. 33.Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. 2022.
  34. 34.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  35. 35.Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. Use what you have: Video retrieval using representations from collaborative experts. arXiv preprint arXiv:1907.13487, 2019.
  36. 36.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  37. 37.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019.
  38. 38.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9879–9889, 2020.
  39. 39.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2630–2640, 2019.
  40. 40.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  41. 41.Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid. Look before you speak: Visually contextualized utterances. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16877–16887, 2021.
  42. 42.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018.
  43. 43.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7464–7473, 2019.
  44. 44.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7464–7473, 2019.
  45. 45.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  46. 46.Alex Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan, Xudong Lin, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Object-aware video-language pre-training for retrieval. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022.
  47. 47.Jianfeng Wang, Xiaowei Hu, Zhe Gan, Zhengyuan Yang, Xiyang Dai, Zicheng Liu, Yumao Lu, and Lijuan Wang. Ufo: A unified transformer for vision-language representation learning. arXiv preprint arXiv:2111.10023, 2021.
  48. 48.Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022.
  49. 49.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks for action recognition in videos. IEEE transactions on pattern analysis and machine intelligence, pages 2740–2755, 2018.
  50. 50.Mengmeng Wang, Jiazheng Xing, and Yong Liu. Actionclip: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472, 2021.
  51. 51.Yujia Xie, Xiangfeng Wang, Ruijia Wang, and Hongyuan Zha. A fast proximal point method for computing exact wasserstein distance. In Uncertainty in artificial intelligence, pages 433–453. PMLR, 2020.
  52. 52.Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia, pages 1645–1653, 2017.
  53. 53.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016.
  54. 54.Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo. Clip-vip: Adapting pretrained image-text model to video-language representation alignment. arXiv preprint arXiv:2209.06430, 2022.
  55. 55.Rui Yan, Mike Zheng Shou, Yixiao Ge, Alex Jinpeng Wang, Xudong Lin, Guanyu Cai, and Jinhui Tang. Video-text pre-training with learned regions. arXiv preprint arXiv:2112.01194, 2021.
  56. 56.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Just ask: Learning to answer questions from millions of narrated videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1686–1697, 2021.
  57. 57.Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Zero-shot video question answering via frozen bidirectional language models. In Advances in Neural Information Processing Systems, 2022.
  58. 58.Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. Taco: Token-aware cascade contrastive learning for video-text alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11562–11572, 2021.
  59. 59.Youngjae Yu, Jongseok Kim, and Gunhee Kim. A joint sequence fusion model for video question answering and retrieval. In Proceedings of the European Conference on Computer Vision (ECCV), pages 471–487, 2018.
  60. 60.Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. Merlot reserve: Neural script knowledge through vision and language and sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16375–16387, 2022.
  61. 61.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. Advances in Neural Information Processing Systems, 34, 2021.
  62. 62.Bowen Zhang, Hexiang Hu, and Fei Sha. Cross-modal and hierarchical modeling of video and text. In Proceedings of the European Conference on Computer Vision (ECCV), pages 374–390, 2018.
  63. 63.Bowen Zhang, Jiahui Yu, Christopher Fifty, Wei Han, Andrew M Dai, Ruoming Pang, and Fei Sha. Co-training transformer with videos and images improves action recognition. arXiv preprint arXiv:2112.07175, 2021.
  64. 64.Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8746–8755, 2020.

Citation

MLA
Wang, A. J., et al. “All in One: Exploring Unified Video-Language Pre-training”. arXiv, 2022, http://arxiv.org/abs/2203.07303v1.
APA
Wang, A. J., Ge, Y., Yan, R., Ge, Y., Lin, X., Cai, G., Wu, J., Shan, Y., Qie, X., & Shou, M. Z. (2022). All in One: Exploring Unified Video-Language Pre-training. arXiv. http://arxiv.org/abs/2203.07303v1
Chicago
Wang, A. J., Y. Ge, R. Yan, et al. 2022. “All in One: Exploring Unified Video-Language Pre-training”. arXiv. http://arxiv.org/abs/2203.07303v1.
Harvard
Wang, A.J. et al. (2022) “All in One: Exploring Unified Video-Language Pre-training”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.07303v1.
Vancouver
1. Wang AJ, Ge Y, Yan R, Ge Y, Lin X, Cai G, Wu J, Shan Y, Qie X, Shou MZ (2022) All in One: Exploring Unified Video-Language Pre-training. arXiv

BibTeX

@article{wang2022all,
  title = {All in One: Exploring Unified Video-Language Pre-training},
  author = {Wang, Alex Jinpeng and Ge, Yixiao and Yan, Rui and Ge, Yuying and Lin, Xudong and Cai, Guanyu and Wu, Jianping and Shan, Ying and Qie, Xiaohu and Shou, Mike Zheng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.07303v1},
  eprint = {2203.07303}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE