keyword
cross-modal code matching
Cross-modal code matching is a machine learning technique and optimization objective that aligns the discrete latent representations of different data modalities within a shared quantized embedding space. In this approach, inputs from diverse modalities, such as visual scenes and spoken or written language, are mapped into a common codebook of discrete embedding vectors through vector quantization. The matching objective encourages paired or corresponding inputs across modalities to produce similar probability distributions over these discrete codes. By standardizing representations to a shared categorical vocabulary, cross-modal code matching enables models to link fine-grained semantic concepts, objects, and actions across distinct sensory streams without requiring direct supervisory labels.
1 item

