Pre-trained uni-modal encoders are machine learning models that have been trained beforehand on a single data type, such as text, images, or audio, to convert raw input into rich numerical representations known as embeddings. Unlike multi-modal architectures that process multiple forms of data simultaneously, a uni-modal encoder specializes in capturing the intrinsic patterns, structures, and semantic relationships unique to one specific modality. Common examples include vision transformers and convolutional networks for visual data, as well as transformer-based language models for textual data. Once trained, these encoders serve as standardized feature extractors that can be fine-tuned for single-domain tasks or paired together within multi-modal learning frameworks to align representations across different sensory and linguistic domains.