An omni-modality video caption dataset is a curated collection of video clips paired with comprehensive textual descriptions that synthesize information across all primary video tracks, including visual imagery, audio signals, and linguistic elements such as subtitles or spoken dialogue. Unlike conventional video-text datasets that focus predominantly on describing visible actions and objects, an omni-modality dataset unifies visual scenes with acoustic events, environmental sounds, and speech transcripts into integrated, holistic annotations. These datasets serve as foundational resources for training and evaluating multimodal artificial intelligence models, enabling cross-modal understanding, retrieval, question answering, and caption generation across vision, audio, and text.