VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
Sihan ChenHandong LiQunbo WangZijia ZhaoMingzhen SunXinxin ZhuJing Liu
Presents VAST-27M, an automatically generated 27-million-clip omni-modality video caption dataset, along with a unified foundation model capable of processing vision, audio, subtitles, and text across diverse retrieval, captioning, and question-answering tasks.
Modern artificial intelligence systems increasingly rely on automated video understanding to power content search, description, and interactive question answering. However, conventional video-language models predominantly focus on pairing visual frames with text descriptions, often neglecting audio signals and spoken subtitles. Environmental sounds and speech provide essential context that reduces ambiguity and deepens comprehension. Existing training datasets either rely on raw transcriptions that fail to describe visual events or contain expensive, manually generated captions that cannot easily scale to meet real-world demands.
The article aims to introduce and evaluate VAST-27M, an automatically generated large-scale video dataset pairing omni-modality captions with video tracks, and VAST, a foundational model designed to perceive and integrate vision, audio, and subtitle modalities for cross-modal tasks.
To construct the dataset without expensive human annotation, the authors built an automated pipeline using open-domain video clips. Separate vision and audio captioning models were trained on established public corpora to produce single-modality descriptions. An off-the-shelf large language model then synthesized these descriptions alongside raw subtitles and instructional prompts into comprehensive omni-modality captions, yielding 27 million video clips with 297 million total captions. The VAST model was constructed using dedicated vision, audio, and text encoders, trained across combined contrastive matching and text generation objectives, and evaluated across numerous public benchmarks for retrieval, captioning, and question answering.
The experimental findings show that the omni-modality approach significantly outperforms prior specialized models across multiple domains. VAST achieved 22 new state-of-the-art results across various multimodal benchmarks. In text-to-video retrieval, the model improved retrieval accuracy over prior leading systems by roughly 5 to 17 percentage points on major benchmarks like MSRVTT and YouCook2. In audio-text retrieval, the model gained approximately 5 to 10 percentage points over previous baselines. Furthermore, on video captioning benchmarks, VAST outperformed larger models like GIT2 while using only about 22% of the parameter count and roughly 3% of the training data scale.
These results demonstrate that synthesizing vision, environmental sound, and spoken dialogue into a single pretraining framework dramatically improves model performance and efficiency without requiring massive increases in model size or costly human labeling. Integrating multiple sensory tracks mitigates modality gaps and enhances cross-domain transfer, offering organizations a more cost-effective blueprint for building high-performing multimodal AI systems.
Organizations developing video and audio intelligence should adopt automated caption synthesis pipelines and multimodal pretraining rather than relying strictly on visual-text alignments. Further engineering should explore integrating generative large language models directly into the architecture to expand contextual reasoning and deploying pilot evaluations in production settings.
The authors note that because the dataset generation relies on automated captioners and an off-the-shelf language model, the corpus may inherit the underlying biases of those base tools. While confidence in the benchmark improvements is high, decision-makers should account for potential dataset-specific noise and validate the pipeline across specialized or higher-stakes domains before wide operational deployment.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). Learn how Video-LLaMA pioneeringly integrates both video dynamics and auditory streams into large language models, establishing the multi-track video-audio understanding foundation that VAST expands to omni-modality captioning.
- Paper: HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips, Antoine Miech et al. (2019). Understand the foundational methodology of using massive video collections paired with speech transcripts to train scalable video-text embeddings without manual human annotation.
- Paper: Contrastive Audio-Visual Masked Autoencoder, Yuan Gong et al. (2023). Explore joint audio-visual representation learning via contrastive and masked modeling, providing the foundational principles for combining auditory and visual features in video models.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Discover the dual-encoder end-to-end retrieval paradigm and the WebVid dataset that set the standard for video-text pretraining architectures.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). Review the core crossmodal attention mechanisms designed to process unaligned language, visual, and acoustic streams simultaneously.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Examine how aligning visual modalities before language model projection enables unified multi-modal understanding across image and video domains.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). Study the architecture and instruction-tuning methodology for video-based conversational language models that VAST builds upon for omni-modal tasks.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Examine how modern multimodal LLMs are rigorously evaluated across complex, long-form video analysis tasks integrating visual, audio, and subtitle streams.
- Paper: LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding, Haoning Wu et al. (2024). Explore an extended evaluation benchmark specifically assessing fine-grained reasoning over interleaved, long-context video and subtitle sequences.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). See how video-language models are enhanced for long-video comprehension by explicitly grounding and aligning temporal timestamps with segment descriptions.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). Discover how unified multimodal representations are scaled across single images, multi-image sequences, and video understanding using synthetic data curricula.
- Paper: Video ReCap: Recursive Captioning of Hour-Long Videos, Md Mohaiminul Islam et al. (2024). Learn how video captioning is generalized to recursive, multi-tier summaries for hour-long untrimmed video streams.
- Paper: Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners, Yazhou Xing et al. (2024). Investigate how the joint alignment of visual and audio representations enables synchronized, open-domain cross-modal and joint video-audio generation.
