Built independently by an author, for readers. Read the story and support ChapterPal

keyword

fine-grained representation

A fine-grained representation is a data encoding in machine learning that captures detailed, localized, or low-level features of an input rather than only its broad, global summary. While high-level representations compress an entire data instance, such as a complete document, image, or video, into a single generalized vector, fine-grained representations retain information about specific sub-components, such as individual words, image regions, or temporal frames. By preserving these distinct elements and their subtle semantic variations, fine-grained representations allow computational models to perform precise localized analysis, align specific concepts across different modalities, and distinguish subtle differences between closely related entities.

1 item

Cross-Modal Discrete Representation Learning

Cross-Modal Discrete Representation Learning

Alexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James R. Glass

OrganizationsMassachusetts Institute of Technology

Why you should read this

Presents a self-supervised framework that uses vector quantization and code matching across modalities to learn fine-grained discrete representations, enabling unsupervised concept localization and boosting retrieval performance.

In contrast to recent advances focusing on high-level representation learning across modalities, in this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events represented by visual objects or spoken words. Our framework relies on a discretized embedding space created via vector quantization that is shared across different modalities. Beyond the shared embedding space, we propose a Cross-Modal Code Matching objective that forces the representations from different views (modalities) to have a similar distribution over the discrete embedding space such that cross-modal objects/actions localization can be performed without direct supervision. We show that the proposed discretized multi-modal fine-grained representation (e.g., pixel/word/frame) can complement high-level summary representations (e.g., video/sentence/waveform) for improved performance on cross-modal retrieval tasks. We also observe that the discretized representation uses individual clusters to represent the same semantic concept across modalities.

Added

2026-09-26