An omni-modality foundation model is an artificial intelligence model pre-trained across a comprehensive spectrum of data and sensory modalities, including visual imagery, audio signals, speech or subtitles, and written text, within a unified representation space. Unlike standard multimodal architectures that often operate strictly on pairs of inputs such as vision and text, an omni-modality foundation model perceives, aligns, and processes multiple heterogeneous media streams concurrently. This unified cross-modal understanding enables the system to generalize across diverse downstream tasks, including multimodal retrieval, content captioning, and question answering, across any combination of its supported sensory inputs and outputs.