Built independently by an author, for readers. Read the story and support ChapterPal

keyword

cross-modal code matching

Cross-modal code matching is a machine learning technique and optimization objective that aligns the discrete latent representations of different data modalities within a shared quantized embedding space. In this approach, inputs from diverse modalities, such as visual scenes and spoken or written language, are mapped into a common codebook of discrete embedding vectors through vector quantization. The matching objective encourages paired or corresponding inputs across modalities to produce similar probability distributions over these discrete codes. By standardizing representations to a shared categorical vocabulary, cross-modal code matching enables models to link fine-grained semantic concepts, objects, and actions across distinct sensory streams without requiring direct supervisory labels.

1 item

Cross-Modal Discrete Representation Learning

Cross-Modal Discrete Representation Learning

Alexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James R. Glass

OrganizationsMassachusetts Institute of Technology

Why you should read this

Presents a self-supervised framework that uses vector quantization and code matching across modalities to learn fine-grained discrete representations, enabling unsupervised concept localization and boosting retrieval performance.

In contrast to recent advances focusing on high-level representation learning across modalities, in this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events represented by visual objects or spoken words. Our framework relies on a discretized embedding space created via vector quantization that is shared across different modalities. Beyond the shared embedding space, we propose a Cross-Modal Code Matching objective that forces the representations from different views (modalities) to have a similar distribution over the discrete embedding space such that cross-modal objects/actions localization can be performed without direct supervision. We show that the proposed discretized multi-modal fine-grained representation (e.g., pixel/word/frame) can complement high-level summary representations (e.g., video/sentence/waveform) for improved performance on cross-modal retrieval tasks. We also observe that the discretized representation uses individual clusters to represent the same semantic concept across modalities.

Added

2026-09-26