keyword
joint representations
A joint representation is a unified mathematical embedding that combines information from multiple distinct data modalities, such as text, images, or audio, into a single shared feature space. By projecting disparate input streams into a common latent structure, this representation captures the correlations, interactions, and complementary semantics existing across different modalities. This unified modeling approach enables machine learning algorithms to process multimodal inputs holistically, supporting downstream tasks such as cross-modal retrieval, classification, and inference even when certain modalities are missing or incomplete.
1 item

