VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

Sihan ChenHandong LiQunbo WangZijia ZhaoMingzhen SunXinxin ZhuJing Liu

article2023NeurIPS202 citations

Presents VAST-27M, an automatically generated 27-million-clip omni-modality video caption dataset, along with a unified foundation model capable of processing vision, audio, subtitles, and text across diverse retrieval, captioning, and question-answering tasks.

Listen

Modern artificial intelligence systems increasingly rely on automated video understanding to power content search, description, and interactive question answering. However, conventional video-language models predominantly focus on pairing visual frames with text descriptions, often neglecting audio signals and spoken subtitles. Environmental sounds and speech provide essential context that reduces ambiguity and deepens comprehension. Existing training datasets either rely on raw transcriptions that fail to describe visual events or contain expensive, manually generated captions that cannot easily scale to meet real-world demands.

The article aims to introduce and evaluate VAST-27M, an automatically generated large-scale video dataset pairing omni-modality captions with video tracks, and VAST, a foundational model designed to perceive and integrate vision, audio, and subtitle modalities for cross-modal tasks.

To construct the dataset without expensive human annotation, the authors built an automated pipeline using open-domain video clips. Separate vision and audio captioning models were trained on established public corpora to produce single-modality descriptions. An off-the-shelf large language model then synthesized these descriptions alongside raw subtitles and instructional prompts into comprehensive omni-modality captions, yielding 27 million video clips with 297 million total captions. The VAST model was constructed using dedicated vision, audio, and text encoders, trained across combined contrastive matching and text generation objectives, and evaluated across numerous public benchmarks for retrieval, captioning, and question answering.

The experimental findings show that the omni-modality approach significantly outperforms prior specialized models across multiple domains. VAST achieved 22 new state-of-the-art results across various multimodal benchmarks. In text-to-video retrieval, the model improved retrieval accuracy over prior leading systems by roughly 5 to 17 percentage points on major benchmarks like MSRVTT and YouCook2. In audio-text retrieval, the model gained approximately 5 to 10 percentage points over previous baselines. Furthermore, on video captioning benchmarks, VAST outperformed larger models like GIT2 while using only about 22% of the parameter count and roughly 3% of the training data scale.

These results demonstrate that synthesizing vision, environmental sound, and spoken dialogue into a single pretraining framework dramatically improves model performance and efficiency without requiring massive increases in model size or costly human labeling. Integrating multiple sensory tracks mitigates modality gaps and enhances cross-domain transfer, offering organizations a more cost-effective blueprint for building high-performing multimodal AI systems.

Organizations developing video and audio intelligence should adopt automated caption synthesis pipelines and multimodal pretraining rather than relying strictly on visual-text alignments. Further engineering should explore integrating generative large language models directly into the architecture to expand contextual reasoning and deploying pilot evaluations in production settings.

The authors note that because the dataset generation relies on automated captioners and an off-the-shelf language model, the corpus may inherit the underlying biases of those base tools. While confidence in the benchmark improvements is high, decision-makers should account for potential dataset-specific noise and validate the pipeline across specialized or higher-stakes domains before wide operational deployment.

arXiv: 2305.18500
Cover for VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

Abstract

Vision and text have been fully explored in contemporary video-text foudational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audio, and Subtitle, and Text by exploring an automatically generated large-scale omni-modality video caption dataset called VAST-27M. Specifically, we first collect 27 million open-domain video clips and separately train a vision and an audio captioner to generate vision and audio captions. Then, we employ an off-the-shelf Large Language Model (LLM) to integrate the generated captions, together with subtitles and instructional prompts into omni-modality captions. Based on the proposed VAST-27M dataset, we train an omni-modality video-text foundational model named VAST, which can perceive and process vision, audio, and subtitle modalities from video, and better support various tasks including vision-text, audio-text, and multi-modal video-text tasks (retrieval, captioning and QA). Extensive experiments have been conducted to demonstrate the effectiveness of our proposed VAST-27M corpus and VAST foundation model. VAST achieves 22 new state-of-the-art results on various cross-modality benchmarks. Code, model and dataset will be released at https://github.com/TXH-mercury/VAST.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Cross-Modality Pretraining Corpus
  • 2.2 Multi-Modality Learning
  • 3 Dataset
  • 3.1 Data Collection of VAST-27M
  • 3.2 Statistics of VAST-27M
  • 4 Approach
  • 4.1 Basic Framework
  • 4.2 Pretraining Objectives
  • 4.3 Modality Grouping
  • 5 Experiments
  • 5.1 Implementation Details
  • 5.2 Comparison to State-of-the-Art Models
  • 5.3 Comparison to Open-Source Cross-Modality Training Corpus
  • 5.4 Ablation Study
  • 6 Conclusion, Broader Impact and Limitation
  • Acknowledgments
  • References
  • Appendix
  • A More about VAST Foundation Model
  • A.1 Pretraining Settings
  • A.2 Downstream Datasets Descriptions
  • A.3 Finetuning Settings
  • A.4 Detailed Comparisons to State-of-the-Art Methods
  • B.2 Prompts for Omni-Modality Caption Generation
  • B.3 More Examples

Knowls

  1. Knowl 1 — VAST-27M omni-modality video-caption corpus

    data/table

    VAST-27M contains 27 million open-domain video clips sampled from HD_VILA_100M. Each clip is paired with 11 automatically generated captions: five vision captions, five audio captions, and one omni-modality caption (OMC) that integrates vision, audio, and subtitle information, yielding 297 million captions in total. The clips cover more than 15 categories, including music, gaming, education, entertainment, and animals. The reported average caption lengths are 12.5 for vision captions, 7.2 for audio captions, and 32.4 for omni-modality captions. Unlike subtitle-only or vision-only corpora, VAST-27M provides separate supervision for vision-text, audio-text, and jointly modeled vision-audio-subtitle-text learning.

  2. Knowl 2 — Automatic generation of omni-modality captions

    algorithm

    The VAST-27M construction pipeline has two stages. First, a vision captioner and an audio captioner independently generate modality-specific descriptions for each video clip. The vision captioner is pretrained on CC4M, CC12M, and 100 million randomly selected LAION-400M image-text pairs, then fine-tuned on MSCOCO, VATEX, MSRVTT, and MSVD so that it captures both objects and actions. The audio captioner is trained on VALOR-1M and WavCaps without a second fine-tuning stage, because the available downstream audio-captioning data are relatively small and could cause overfitting.

    For every selected clip, each captioner produces five captions using Top-KK sampling with K=10K=10. The pipeline randomly selects three vision captions and three audio captions, combines them with the clip's raw subtitle and an instructional prompt, and supplies them to Vicuna-13B. The prompt asks the language model to summarize the modalities into one natural sentence, avoid simply concatenating the inputs, and give equal weight to vision, audio, and speech. Vicuna-13B produces the final OMC.

    The source clips are filtered to have durations from 5 to 30 seconds, to contain vision, audio, and subtitle tracks, and to be evenly sampled from the original 3.3 million long-form videos. This filtering produces the 27 million clips used in VAST-27M.

  3. Knowl 3 — VAST multimodal encoder architecture

    model/method

    VAST is an end-to-end Transformer foundation model with three modality encoders: a ViT vision encoder, a BEATs audio encoder, and a BERT text encoder. The text encoder tokenizes subtitles and omni-modality captions with WordPiece and uses cross-attention layers both to fuse modalities and to decode captions.

    The vision encoder processes raw images or sparsely sampled video frames and produces a vision representation fvf_v. Audio is divided into 10-second segments, zero-padded when necessary, converted into 64-dimensional log-Mel filterbank spectrograms using a 25 ms Hamming window, and encoded as faf_a. Subtitle and OMC sequences produce text representations fsf_s and fomcf_{omc}. The global representation of each modality is its [CLS][\mathrm{CLS}] feature, denoted fvgf_{vg}, fagf_{ag}, fsgf_{sg}, and fomcgf_{omcg}.

    For an omni-modality video, VAST concatenates the global features from vision, audio, and subtitle streams to obtain a global representation, while its local representation concatenates the corresponding unpooled sequential features after independent linear projections align their hidden dimensions. This allows the same architecture to process images, videos, audio, subtitles, and captions.

  4. Knowl 4 — Omni-modality pretraining objectives

    equation

    VAST trains on paired omni-modality videos and omni-modality captions using three objectives. Let a minibatch contain BB matched pairs (vi,ci)(v_i,c_i), where viv_i is an omni-modality video representation and cic_i is its OMC representation. Let v^i\hat v_i and c^i\hat c_i be their projected, ℓ2\ell_2-normalized global features, let sim⁡(v^,c^)=v^⊤c^\operatorname{sim}(\hat v,\hat c)=\hat v^\top\hat c, and let τ\tau be a learned temperature scale. The omni-modality video-caption contrastive loss is

    LOM-VCC=−12B∑i=1B[log⁡exp⁡(τsim⁡(v^i,c^i))∑j=1Bexp⁡(τsim⁡(v^i,c^j))+log⁡exp⁡(τsim⁡(v^i,c^i))∑j=1Bexp⁡(τsim⁡(v^j,c^i))].\mathcal{L}_{\mathrm{OM\text{-}VCC}}=-\frac{1}{2B}\sum_{i=1}^{B}\left[\log\frac{\exp(\tau\operatorname{sim}(\hat v_i,\hat c_i))}{\sum_{j=1}^{B}\exp(\tau\operatorname{sim}(\hat v_i,\hat c_j))}+\log\frac{\exp(\tau\operatorname{sim}(\hat v_i,\hat c_i))}{\sum_{j=1}^{B}\exp(\tau\operatorname{sim}(\hat v_j,\hat c_i))}\right].

    The omni-modality video-caption matching objective, OM-VCM, uses the local video representation as cross-attention context for the caption tokens. A two-layer MLP predicts pvcm∈[0,1]p_{vcm}\in[0,1], the probability that the video-caption pair is matched. For label y∈{0,1}y\in\{0,1\}, the objective is binary cross-entropy, −E[ylog⁡pvcm+(1−y)log⁡(1−pvcm)]-\mathbb{E}[y\log p_{vcm}+(1-y)\log(1-p_{vcm})], with hard negative pairs mined during training.

    The omni-modality video-caption generation objective, OM-VCG, masks 60% of the OMC tokens, activates cross-attention to the local video representation, applies causal self-attention, and predicts the masked tokens autoregressively. If cmc_m is a masked token, c<mc_{<m} is its preceding token prefix, and vv is the omni-modality video representation, then

    LOM-VCG=−E(v,c)[log⁡P(cm∣c<m,v)].\mathcal{L}_{\mathrm{OM\text{-}VCG}}=-\mathbb{E}_{(v,c)}\left[\log P(c_m\mid c_{<m},v)\right].

    The total omni-modality loss gives the three objectives equal weight: LOM=LOM-VCC+LOM-VCM+LOM-VCG\mathcal{L}_{\mathrm{OM}}=\mathcal{L}_{\mathrm{OM\text{-}VCC}}+\mathcal{L}_{\mathrm{OM\text{-}VCM}}+\mathcal{L}_{\mathrm{OM\text{-}VCG}}.

  5. Knowl 5 — Modality grouping for missing downstream tracks

    model/method

    To reduce the mismatch between pretraining inputs and downstream datasets that may lack audio or subtitles, VAST explicitly trains modality-specific groups in addition to full omni-modality training. The groups are vision-text (V-T), audio-text (A-T), vision-audio-text (VA-T), vision-subtitle-text (VS-T), and vision-audio-subtitle-text (VAS-T). Vision captions supervise V-T, audio captions supervise A-T, and OMCs supervise VA-T, VS-T, and VAS-T.

    The final training objective is the sum of the omni-modality objective and the corresponding group objectives:

    L=LOM+LV-T+LA-T+LVA-T+LVS-T.\mathcal{L}=\mathcal{L}_{\mathrm{OM}}+\mathcal{L}_{\mathrm{V\text{-}T}}+\mathcal{L}_{\mathrm{A\text{-}T}}+\mathcal{L}_{\mathrm{VA\text{-}T}}+\mathcal{L}_{\mathrm{VS\text{-}T}}.

    This grouping exposes the text encoder to each modality combination that can occur at adaptation time rather than requiring every downstream example to contain all three video tracks.

  6. Knowl 6 — Pretraining and downstream adaptation configuration

    experimental setup

    The full VAST model has 1.3 billion parameters and is initialized with EVA-CLIP ViT-G, BEATs, and BERT-B encoders. Pretraining combines VAST-27M, VALOR-1M, WavCaps, CC14M, and 110 million randomly sampled LAION-400M image-text pairs. Captions in CC14M and LAION are replaced by captions generated by the trained vision captioner. The main training description uses 64 Tesla V100 GPUs, batch size 1024, an initial learning rate of 10−410^{-4} with linear decay, and approximately 200,000 training steps; the detailed appendix schedule totals 205,000 steps across the corpora. Each video example uses one randomly sampled frame and two randomly sampled 10-second audio clips.

    For retrieval, VAST first ranks candidates with the contrastive score and reranks the top 50 with the matching score. Captioning uses beam search with beam size 3. Question answering is formulated as open-ended generation in which the question is supplied as a prefix and the answer is generated without answer constraints. The reported primary metrics are Recall@1 for retrieval, CIDEr for captioning, and accuracy for question answering.

  7. Knowl 7 — Broad cross-modality benchmark performance

    empirical result

    VAST is reported to obtain 22 new state-of-the-art results across vision-text, audio-text, and multimodal video-text benchmarks. On multimodal text-to-video retrieval, its Recall@1 scores are 63.9 on MSRVTT, 50.4 on YouCook2, 80.0 on VALOR-32K, 83.0 on VATEX, 72.0 on DiDeMo, and 70.5 on ActivityNet. On multimodal video captioning, its CIDEr scores are 78.0 on MSRVTT, 198.8 on YouCook2, 62.0 on VALOR-32K, 99.5 on VATEX with SCST fine-tuning, and 74.1 on TVC. Its multimodal QA accuracies are 50.1 on MSRVTT-QA, 80.7 on MUSIC-AVQA, and 50.4 on ActivityNet-QA.

    VAST also improves audio-text retrieval, reaching Recall@1 values of 25.1 on ClothoV1, 26.9 on ClothoV2, and 52.0 on AudioCaps, corresponding to reported gains of 7.6, 5.4, and 9.8 points over the prior state of the art. On image-text tasks, it obtains Recall@1 values of 68.0 on MSCOCO and 91.0 on Flickr30K, an MSCOCO caption CIDEr of 149.0 with SCST, and an MSCOCO SPICE score of 27.0. Zero-shot text-to-video Recall@1 reaches 49.3 on MSRVTT and 55.5 on DiDeMo.

  8. Knowl 8 — Caption-source quality demonstrated by controlled corpus comparisons

    empirical result

    Controlled pretraining experiments show that the caption types in VAST-27M provide useful and complementary supervision. When only visual video content is used for both pretraining and fine-tuning, the model trained on VAST-27M vision captions obtains MSVD scores of 47.7 retrieval, 149.6 CIDEr, and 55.3 QA accuracy, and MSRVTT scores of 49.1 retrieval, 68.9 CIDEr, and 46.8 QA accuracy. These are the best results among the compared vision-text corpora. Treating raw subtitles as captions performs substantially worse, with MSVD scores of 37.3, 124.5, and 40.3 and MSRVTT scores of 40.3, 61.9, and 45.1. OMC pretraining is stronger than subtitle supervision but is slightly below dedicated vision captions when only vision is available.

    For audio-text pretraining, VAST-27M audio captions produce ClothoV2 retrieval and captioning scores of 23.1 and 48.9, and AudioCaps scores of 47.4 retrieval and 76.9 captioning, outperforming the compared VALOR-1M and WavCaps settings on all four measurements.

    For full omni-modality video-caption learning, VAST-27M OMC pretraining obtains MSRVTT retrieval, captioning, and QA scores of 53.0, 71.7, and 47.9; YouCook2 retrieval and captioning scores of 44.0 and 187.5; and VALOR-32K retrieval and captioning scores of 58.8 and 48.5. Replacing LLM-generated OMCs with a simple concatenation of vision captions, audio captions, and subtitles reduces these results to 44.4, 71.4, 47.7; 36.9, 159.0; and 49.9, 48.1, respectively, supporting the value of LLM-based modality integration.

  9. Knowl 9 — Effect of multimodal pretraining and modality grouping

    empirical result

    Ablations on MSRVTT, YouCook2, and VALOR-32K show that adding audio and subtitles improves a vision-only model, and the gains are larger after omni-modality pretraining. With omni-modality pretraining and vision-only fine-tuning, the model scores 47.5 retrieval, 67.4 CIDEr, and 46.7 QA accuracy on MSRVTT; 14.0 retrieval and 93.6 CIDEr on YouCook2; and 56.5 retrieval and 45.0 CIDEr on VALOR-32K. Using vision, audio, and subtitles during both fine-tuning and evaluation raises these values to 53.0, 71.7, 47.9; 44.0, 187.5; and 58.8, 48.5, respectively.

    Using all modality groups during pretraining and vision-audio-subtitle inputs during fine-tuning further produces 55.3 retrieval, 71.7 CIDEr, and 48.2 QA accuracy on MSRVTT; 45.9 retrieval and 188.0 CIDEr on YouCook2; and 62.2 retrieval and 48.4 CIDEr on VALOR-32K. Models that rely on subtitle supervision can degrade on VALOR-32K when subtitles are absent in most evaluation videos. The modality-grouping objective mitigates this pretraining–fine-tuning inconsistency and gives more uniform performance across vision-, subtitle-, and audio-oriented benchmarks.

  10. Knowl 10 — Limitations of the corpus and foundation model

    limitation

    The paper identifies three limitations. VAST-27M is large but does not exhaust the need for more diverse and larger-scale omni-modality corpora. Although VAST supports many retrieval, captioning, and QA tasks, integrating a large language model is still considered necessary to further improve its generalization capabilities. Finally, the dataset-generation pipeline depends on Vicuna-13B and the captioner training uses existing open-source cross-modality corpora, so VAST-27M and the VAST model may inherit biases from those language models, captioners, and source datasets.

Coverage note — No substantial contributed material was deliberately omitted; detailed benchmark split descriptions and auxiliary examples were excluded because they support evaluation rather than adding independent methodological content.

References

  1. 1.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020.
  2. 2.J. Devlin, M.-W. Chang, K. Lee, and Y. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018.
  3. 3.M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 1728–1738.
  4. 4.A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2630–2640.
  5. 5.H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo, “Advancing high-resolution video-language representation with large-scale video transcriptions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5036–5045.
  6. 6.R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi, “Merlot: Multimodal neural script knowledge models,” Advances in Neural Information Processing Systems, vol. 34, pp. 23 634–23 651, 2021.
  7. 7.S. Chen, X. He, L. Guo, X. Zhu, W. Wang, J. Tang, and J. Liu, “Valor: Vision-audio-language omni-perception pretraining model and dataset,” arXiv preprint arXiv:2304.08345, 2023.
  8. 8.W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
  9. 9.S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3558–3568.
  10. 10.P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 2556–2565.
  11. 11.J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, pp. 776–780.
  12. 12.K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 736–740.
  13. 13.C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 119–132.
  14. 14.I. Martin Morato and A. Mesaros, “Diversity and bias in audio captioning datasets,” 2021.
  15. 15.M. Wu, H. Dinkel, and K. Yu, “Audio caption: Listen and tell,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 830–834.
  16. 16.Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5.
  17. 17.F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proceedings of the 21st ACM international conference on Multimedia, 2013, pp. 411–412.
  18. 18.X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” arXiv preprint arXiv:2303.17395, 2023.
  19. 19.J. Wang, Z. Yang, X. Hu, L. Li, K. Lin, Z. Gan, Z. Liu, C. Liu, and L. Wang, “Git: A generative image-to-text transformer for vision and language,” arXiv preprint arXiv:2205.14100, 2022.
  20. 20.L. Li, J. Lei, Z. Gan, L. Yu, Y.-C. Chen, R. Pillai, Y. Cheng, L. Zhou, X. E. Wang, W. Y. Wang et al., “Value: A multi-task benchmark for video-and-language understanding evaluation,” arXiv preprint arXiv:2106.04632, 2021.
  21. 21.V. Gabeur, C. Sun, K. Alahari, and C. Schmid, “Multi-modal transformer for video retrieval,” in European Conference on Computer Vision. Springer, 2020, pp. 214–229.
  22. 22.S. Chen, X. Zhu, D. Hao, W. Liu, J. Liu, Z. Zhao, L. Guo, and J. Liu, “Mm21 pre-training for video understanding challenge: Video captioning with pretraining techniques,” in Proceedings of the 29th ACM International Conference on Multimedia, 2021, pp. 4853–4857.
  23. 23.H. Luo, L. Ji, B. Shi, H. Huang, N. Duan, T. Li, J. Li, T. Bharti, and M. Zhou, “Univl: A unified video and language pre-training model for multimodal understanding and generation,” arXiv preprint arXiv:2002.06353, 2020.
  24. 24.P. H. Seo, A. Nagrani, and C. Schmid, “Look before you speak: Visually contextualized utterances,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 16 877–16 887.
  25. 25.A. Rouditchenko, A. Boggust, D. Harwath, B. Chen, D. Joshi, S. Thomas, K. Audhkhasi, H. Kuehne, R. Panda, R. Feris et al., “Avlnet: Learning audio-visual language representations from instructional videos,” arXiv preprint arXiv:2006.09199, 2020.
  26. 26.R. Zellers, J. Lu, X. Lu, Y. Yu, Y. Zhao, M. Salehi, A. Kusupati, J. Hessel, A. Farhadi, and Y. Choi, “Merlot reserve: Neural script knowledge through vision and language and sound,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 16 375–16 387.
  27. 27.Z. Yang, Y. Fang, C. Zhu, R. Pryzant, D. Chen, Y. Shi, Y. Xu, Y. Qian, M. Gao, Y.-L. Chen et al., “i-code: An integrative and composable multimodal learning framework,” arXiv preprint arXiv:2205.01818, 2022.
  28. 28.J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” arXiv preprint arXiv:2201.12086, 2022.
  29. 29.C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki, “Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,” arXiv preprint arXiv:2111.02114, 2021.
  30. 30.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision. Springer, 2014, pp. 740–755.
  31. 31.X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 4581–4591.
  32. 32.J. Xu, T. Mei, T. Yao, and Y. Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5288–5296.
  33. 33.D. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” in Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, 2011, pp. 190–200.
  34. 34.H. Touvron, H. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, L. Rozière, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
  35. 35.A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, D. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.
  36. 36.S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, and F. Wei, “Beats: Audio pre-training with acoustic tokenizers,” arXiv preprint arXiv:2212.09058, 2022.
  37. 37.Y. Wu, M. Schuster, Z. Chen, Q. V. Le, Q. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey et al., “Google’s neural machine translation system: Bridging the gap between human and machine translation,” arXiv preprint arXiv:1609.08144, 2016.
  38. 38.J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in neural information processing systems, vol. 34, pp. 9694–9705, 2021.
  39. 39.Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao, “Eva-clip: Improved training techniques for clip at scale,” arXiv preprint arXiv:2303.15389, 2023.
  40. 40.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning. PMLR, 2021, pp. 8748–8763.
  41. 41.S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7008–7024.
  42. 42.B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in Proceedings of the IEEE international conference on computer vision, 2015, pp. 2641–2649.
  43. 43.Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo, “Tgif: A new dataset and benchmark on animated gif description,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4641–4650.
  44. 44.D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang, “Video question answering via gradually refined attention over appearance and motion,” in Proceedings of the 25th ACM international conference on Multimedia, 2017, pp. 1645–1653.
  45. 45.Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the v in vqa matter: Elevating the role of image understanding in visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 6904–6913.
  46. 46.J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” arXiv preprint arXiv:2301.12597, 2023.
  47. 47.W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som et al., “Image as a foreign language: Beit pretraining for all vision and vision-language tasks,” arXiv preprint arXiv:2208.10442, 2022.
  48. 48.P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in International Conference on Machine Learning. PMLR, 2022, pp. 23 318–23 340.
  49. 49.W. Kuo, A. Piergiovanni, D. Kim, X. Luo, B. Caine, W. Li, W. Ogale, L. Zhou, A. Dai, A. Chen et al., “Mammut: A simple architecture for joint learning for multimodal tasks,” arXiv preprint arXiv:2303.16839, 2023.
  50. 50.X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer et al., “Pali: A jointly-scaled multilingual language-image model,” arXiv preprint arXiv:2209.06794, 2022.
  51. 51.L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,” in Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  52. 52.L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 5803–5812.
  53. 53.R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles, “Dense-captioning events in videos,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 706–715.
  54. 54.K. Li, Y. Wang, Y. Li, Y. Wang, Y. He, L. Wang, and Y. Qiao, “Unmasked teacher: Towards training-efficient video foundation models,” arXiv preprint arXiv:2303.16058, 2023.
  55. 55.D. Ko, J. Choi, H. K. Choi, K.-W. On, B. Roh, and H. J. Kim, “Meltr: Meta loss transformer for learning to fine-tune video foundation models,” arXiv preprint arXiv:2303.13009, 2023.
  56. 56.H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, and T. Li, “Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,” Neurocomputing, vol. 508, pp. 293–304, 2022.
  57. 57.Y. Wang, K. Li, Y. Li, Y. He, Y. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang et al., “Internvideo: General video foundation models via generative and discriminative learning,” arXiv preprint arXiv:2212.03191, 2022.
  58. 58.J. Lei, L. Yu, T. L. Berg, and M. Bansal, “Tvr: A large-scale dataset for video-subtitle moment retrieval,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16. Springer, 2020, pp. 447–463.
  59. 59.G. Li, Y. Wei, Y. Tian, C. Xu, J.-R. Wen, and D. Hu, “Learning to answer questions in dynamic audio-visual scenarios,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 19 108–19 118.
  60. 60.Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao, “Activitynet-qa: A dataset for understanding complex web videos via question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 9127–9134.
  61. 61.K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y. Lu, and L. Wang, “Swinbert: End-to-end transformers with sparse attention for video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 949–17 958.
  62. 62.M. Tang, Z. Wang, Z. Zeng, F. Rao, and D. Li, “Clip4caption++: Multi-clip for video caption,” arXiv preprint arXiv:2110.05204, 2021.
  63. 63.W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev et al., “The kinetics human action video dataset,” arXiv preprint arXiv:1705.06950, 2017.
  64. 64.S. Chen, Y. Zhao, Q. Jin, and Q. Wu, “Fine-grained video-text retrieval with hierarchical graph reasoning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 10 638–10 647.
  65. 65.A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” International Journal of Computer Vision, vol. 123, no. 1, pp. 94–120, 2017.
  66. 66.Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2758–2766.
  67. 67.A.-M. Oncescu, A. Koepke, J. F. Henriques, Z. Akata, and S. Albanie, “Audio retrieval with natural language queries,” arXiv preprint arXiv:2105.02192, 2021.
  68. 68.A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 3128–3137.
  69. 69.X. Cheng, H. Lin, X. Wu, F. Yang, and D. Shen, “Improving video-text retrieval by multi-stream corpus alignment and dual softmax loss,” arXiv preprint arXiv:2109.04290, 2021.
  70. 70.J. Lei, T. L. Berg, and M. Bansal, “Revealing single frame bias for video-and-language learning,” arXiv preprint arXiv:2206.03428, 2022.
  71. 71.J. Wang, D. Chen, Z. Wu, C. Luo, L. Zhou, Y. Zhao, Y. Xie, C. Liu, Y.-G. Jiang, and L. Yuan, “Omnivl: One foundation model for image-language and video-language tasks,” arXiv preprint arXiv:2209.07526, 2022.
  72. 72.Q. Ye, G. Xu, M. Yan, H. Xu, Q. Qian, J. Zhang, and F. Huang, “Hitea: Hierarchical temporal-aware video-language pre-training,” arXiv preprint arXiv:2212.14546, 2022.
  73. 73.F. Cheng, X. Wang, J. Lei, D. Crandall, M. Bansal, and G. Bertasius, “Vindlu: A recipe for effective video-and-language pretraining,” arXiv preprint arXiv:2212.05051, 2022.
  74. 74.L. Li, Z. Gan, K. Lin, C.-C. Lin, Z. Liu, C. Liu, and L. Wang, “Lavender: Unifying video-language understanding as masked language modeling,” arXiv preprint arXiv:2206.07160, 2022.
  75. 75.A. J. Wang, Y. Ge, R. Yan, Y. Ge, X. Lin, G. Cai, J. Wu, Y. Shan, Y. Qie, and M. Z. Shou, “All in one: Exploring unified video-language pre-training,” arXiv preprint arXiv:2203.07303, 2022.
  76. 76.Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, and R. Ji, “X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 638–647.
  77. 77.H. Xu, Q. Ye, M. Yan, Y. Shi, J. Ye, Y. Xu, C. Li, B. Bi, Q. Qian, W. Wang et al., “Mplug-2: A modularized multi-modal foundation model across text, image and video,” arXiv preprint arXiv:2302.00402, 2023.
  78. 78.H. Xue, Y. Sun, B. Liu, J. Fu, R. Song, H. Li, and J. Luo, “Clip-vip: Adapting pre-trained image-text model to video-language representation alignment,” arXiv preprint arXiv:2209.06430, 2022.
  79. 79.V. Gabeur, A. Nagrani, C. Sun, K. Alahari, and C. Schmid, “Masking modalities for cross-modal video retrieval,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 1766–1775.
  80. 80.Y.-B. Lin, J. Lei, M. Bansal, and G. Bertasius, “Eclipse: Efficient long-range video retrieval using sight and sound,” arXiv preprint arXiv:2204.02874, 2022.
  81. 81.M. Patrick, P.-Y. Huang, Y. Asano, F. Metze, A. Hauptmann, J. Henriques, and A. Vedaldi, “Support-set bottlenecks for video-text representation learning,” arXiv preprint arXiv:2010.02824, 2020.
  82. 82.Q. Wang, Y. Zhang, Y. Zheng, P. Pan, and X.-S. Hua, “Disentangled representation learning for text-video retrieval,” arXiv preprint arXiv:2203.07111, 2022.
  83. 83.H. Xu, G. Ghosh, P.-Y. Huang, P. Arora, M. Aminzadeh, C. Feichtenhofer, F. Metze, and L. Zettlemoyer, “Vlm: Task-agnostic video-language model pre-training for video understanding,” arXiv preprint arXiv:2105.09996, 2021.
  84. 84.D. Li, J. Li, H. Li, J. C. Niebles, and S. C. Hoi, “Align and prompt: Video-and-language pre-training with entity prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4953–4963.
  85. 85.T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y. Wang, L. Wang, and Z. Liu, “Violet: End-to-end video-language transformers with masked visual-token modeling,” arXiv preprint arXiv:2111.12681, 2021.
  86. 86.L. Yuan, D. Chen, Y.-L. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li et al., “Florence: A new foundation model for computer vision,” arXiv preprint arXiv:2111.11432, 2021.
  87. 87.J. Lei, L. Li, L. Zhou, Z. Gan, T. L. Berg, M. Bansal, and J. Liu, “Less is more: Clipbert for video-and-language learning via sparse sampling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 7331–7341.
  88. 88.T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y. Wang, L. Wang, and Z. Liu, “An empirical study of end-to-end video-language transformers with masked visual modeling,” arXiv preprint arXiv:2209.01540, 2022.
  89. 89.J. Huang, Y. Li, J. Feng, X. Sun, and R. Ji, “Clover: Towards a unified video-language alignment and fusion model,” arXiv preprint arXiv:2207.07885, 2022.
  90. 90.A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Just ask: Learning to answer questions from millions of narrated videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 1686–1697.
  91. 91.——, “Zero-shot video question answering via frozen bidirectional language models,” arXiv preprint arXiv:2206.08155, 2022.
  92. 92.J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al., “Flamingo: a visual language model for few-shot learning,” arXiv preprint arXiv:2204.14198, 2022.
  93. 93.S. Yan, T. Zhu, Z. Wang, Y. Cao, M. Zhang, S. Ghosh, Y. Wu, and J. Yu, “Video-text modeling with zero-shot transfer from contrastive captioners,” arXiv preprint arXiv:2212.04979, 2022.
  94. 94.A. Nagrani, P. H. Seo, B. Seybold, A. Hauth, S. Manen, C. Sun, and C. Schmid, “Learning audio-video modalities from image captions,” arXiv preprint arXiv:2204.00679, 2022.
  95. 95.P. Wang, S. Wang, J. Lin, S. Bai, X. Zhou, J. Zhou, X. Wang, and C. Zhou, “One-peace: Exploring one general representation model toward unlimited modalities,” arXiv preprint arXiv:2305.11172, 2023.
  96. 96.X. Xu, H. Dinkel, M. Wu, Z. Xie, and K. Yu, “Investigating local and global information for automated audio captioning with transfer learning,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 905–909.
  97. 97.Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “Simvlm: Simple visual language model pretraining with weak supervision,” arXiv preprint arXiv:2108.10904, 2021.
  98. 98.C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International Conference on Machine Learning. PMLR, 2021, pp. 4904–4916.
  99. 99.J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,” arXiv preprint arXiv:2205.01917, 2022.
  100. 100.M. Ding, B. Xiao, N. Codella, P. Luo, J. Wang, and L. Yuan, “Davit: Dual attention vision transformers,” arXiv preprint arXiv:2204.03645, 2022.

Citation

MLA
Chen, S., et al. “VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 72842–66, https://proceedings.neurips.cc/paper_files/paper/2023/file/e6b2b48b5ed90d07c305932729927781-Paper-Conference.pdf.
APA
Chen, S., Li, H., Wang, Q., Zhao, Z., Sun, M., Zhu, X., & Liu, J. (2023). VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset. Advances in Neural Information Processing Systems, 36, 72842–72866. https://proceedings.neurips.cc/paper_files/paper/2023/file/e6b2b48b5ed90d07c305932729927781-Paper-Conference.pdf
Chicago
Chen, S., H. Li, Q. Wang, et al. 2023. “VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset”. Advances in Neural Information Processing Systems 36: 72842–66. https://proceedings.neurips.cc/paper_files/paper/2023/file/e6b2b48b5ed90d07c305932729927781-Paper-Conference.pdf.
Harvard
Chen, S. et al. (2023) “VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 72842–72866. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/e6b2b48b5ed90d07c305932729927781-Paper-Conference.pdf.
Vancouver
1. Chen S, Li H, Wang Q, Zhao Z, Sun M, Zhu X, Liu J (2023) VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 72842–72866

BibTeX

@inproceedings{chen2023vast,
  title = {VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset},
  author = {Chen, Sihan and Li, Handong and Wang, Qunbo and Zhao, Zijia and Sun, Mingzhen and Zhu, Xinxin and Liu, Jing},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {72842-72866},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/e6b2b48b5ed90d07c305932729927781-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors