keyword
similarity distribution matching
Similarity distribution matching is an optimization and representation alignment technique in multimodal machine learning that aligns feature embeddings across distinct modalities, such as vision and text, by framing cross-modal alignment as a probability distribution matching problem. Rather than relying strictly on pairwise distance ranking with fixed margins or hard-negative mining, this approach converts the pairwise similarity scores between samples across modalities into a normalized probability distribution and minimizes the statistical divergence, typically via Kullback-Leibler divergence, between the predicted distribution and the ground-truth label matching distribution. By operating symmetrically across modalities within a batch, similarity distribution matching facilitates robust global semantic alignment, effectively accommodates scenarios with multiple positive matches, and reduces sensitivity to noisy or ambiguous cross-modal correspondences.
1 item

