Intra-modal self-similarity is a measure of the pairwise semantic, structural, or feature-level similarity among different data instances or elements within the same modality, such as between pairs of images or pairs of text passages. In multimodal machine learning and representation learning, it captures the internal relationships, shared semantic overlaps, and fine-grained correlations inherent to a single data type rather than comparing across different modalities. By modeling these within-modality relationships, systems can account for many-to-many correspondences and partial similarities across samples, providing richer supervisory signals to soften rigid one-to-one alignment targets during multimodal training.