Built independently by an author, for readers. Read the story and support ChapterPal

keyword

omni-modality pretraining

Omni-modality pretraining is a machine learning training paradigm in which a foundational model is simultaneously trained on large-scale data encompassing an extensive array of sensory and data formats, such as visual streams, audio signals, text, and other sensory inputs. Unlike conventional multimodal methods that typically focus on pairwise alignments such as image-text pairs or operate through isolated modality pipelines, omni-modality pretraining integrates diverse information channels into a shared representational space to model complex inter-modal dependencies and unified semantics. This comprehensive approach enables the resulting foundation model to process, align, and reason across arbitrary combinations of modalities, creating transferable representations that support a broad spectrum of single-modal, cross-modal, and composite multimodal downstream tasks, such as content retrieval, automated captioning, and question answering.

1 item

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, Jing Liu

OrganizationsInstitute of Automation, Chinese Academy of SciencesUniversity of Chinese Academy of Sciences

Why you should read this

Presents VAST-27M, an automatically generated 27-million-clip omni-modality video caption dataset, along with a unified foundation model capable of processing vision, audio, subtitles, and text across diverse retrieval, captioning, and question-answering tasks.

Vision and text have been fully explored in contemporary video-text foudational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audio, and Subtitle, and Text by exploring an automatically generated large-scale omni-modality video caption dataset called VAST-27M. Specifically, we first collect 27 million open-domain video clips and separately train a vision and an audio captioner to generate vision and audio captions. Then, we employ an off-the-shelf Large Language Model (LLM) to integrate the generated captions, together with subtitles and instructional prompts into omni-modality captions. Based on the proposed VAST-27M dataset, we train an omni-modality video-text foundational model named VAST, which can perceive and process vision, audio, and subtitle modalities from video, and better support various tasks including vision-text, audio-text, and multi-modal video-text tasks (retrieval, captioning and QA). Extensive experiments have been conducted to demonstrate the effectiveness of our proposed VAST-27M corpus and VAST foundation model. VAST achieves 22 new state-of-the-art results on various cross-modality benchmarks. Code, model and dataset will be released at https://github.com/TXH-mercury/VAST.

Added

2026-09-26