Noisy Correspondence Learning with Meta Similarity Correction
Haochen HanKaiyao MiaoQinghua ZhengMinnan Luo
Proposes a meta-learning framework that trains a correction network on clean and mismatched meta-data to rectify similarity scores and filter out mismatched cross-modal pairs during retrieval training.
Modern artificial intelligence systems increasingly rely on cross-modal retrieval to search and connect different types of data, such as finding images that match descriptive text. Standard algorithms assume that the image-text pairs used during training are accurately paired. In real-world applications, however, training data harvested from the web inevitably contains mismatched pairs, a challenge known as noisy correspondence. When models are trained on this misaligned data, they mistakenly learn to treat irrelevant concepts as similar, resulting in severe performance drops.
The article demonstrates a robust training framework called the Meta Similarity Correction Network (MSCN) to overcome noisy correspondence in cross-modal retrieval. The primary objective is to reliably estimate true multimodal similarity scores and cleanse noisy pairs using a small set of clean reference data.
The researchers evaluated their approach through extensive experiments on standard benchmark datasets (Flickr30K and MS-COCO) injected with synthetic noise ratios of 20%, 50%, and 70%, as well as a large-scale real-world dataset (Conceptual Captions CC152K). The framework frames similarity estimation as a binary classification task trained via bi-level optimization. It leverages both positive pairs and automatically constructed negative pairs from a small clean reference set—comprising only about 2% of the training size—as meta-knowledge. Additionally, the approach integrates an adaptive ranking margin and a two-component Beta Mixture Model to identify and remove mismatched pairs from the training pool before model updates occur.
The experimental findings show that MSCN consistently outperforms existing baseline methods across all noise levels. Under moderate noise conditions (20% to 50%), the method improves top-rank retrieval recall by 2.2% to 4.5% compared to the strongest noise-robust baseline. Under severe noise of 70%, where conventional methods experience catastrophic degradation, MSCN achieves dramatic recall improvements ranging from 27.3% to 52.9% over existing approaches, performing nearly on par with models trained strictly on clean data. On real-world noisy web data, it achieves overall sum recall improvements of 4.9% and 6.2% for text and image queries, respectively, while retaining strong performance even when the clean reference data is reduced to under 1% of the dataset.
These results indicate that organizations can train high-performing multimodal retrieval systems directly on inexpensive, uncurated web datasets without manual cleaning of the entire corpus. By relying on only a tiny fraction of verified data to guide the learning process, organizations can significantly lower data curation costs and project timelines while maintaining high search accuracy. Ablation analyses confirm that both the adaptive margin mechanism and the data purification step are critical to prevent models from overfitting to erroneous pairings.
Organizations developing cross-modal systems should adopt meta-correction and distribution-based sample purification strategies when working with unverified data. Given that the approach requires a small set of verified positive pairs (or a clean validation set) to establish baseline distributions, practitioners must ensure this minor reference set is available. The reported findings are supported by consistent results across standard benchmarks; however, future work could examine whether performance holds across other data types beyond image-text pairs, such as audio-video and multilingual multimodal tasks.
- Paper: Learning to Reweight Examples for Robust Deep Learning, Mengye Ren et al. (2018). Its meta-learning framework establishes the clean-reference reweighting principle that MSCN adapts to estimate trustworthy multimodal similarities.
- Paper: DivideMix: Learning with Noisy Labels as Semi-supervised Learning, Junnan Li et al. (2020). DivideMix provides the foundational loss-based partitioning and sample-purification strategy that motivates MSCN’s removal of unreliable training pairs.
- Paper: Selective-Supervised Contrastive Learning with Noisy Labels, Shikun Li et al. (2022). Its confidence-based selection of reliable pairs supplies a closely related precursor for MSCN’s filtering of noisy correspondence data.
- Paper: Robust Cross-Modal Representation Learning with Progressive Self-Distillation, Alex Andonian et al. (2022). This directly establishes robust cross-modal contrastive learning under web-scale image-text correspondence noise, making MSCN’s correction problem and design easier to situate.
- Paper: Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision, Chao Jia et al. (2021). ALIGN demonstrates the web-scale dual-encoder retrieval setting and noisy text supervision that MSCN explicitly seeks to make more reliable.
- Paper: Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID, Wentao Tan et al. (2024). It applies noise-aware masking to automatically generated image-text descriptions in transferable person retrieval, extending MSCN’s correspondence-cleaning idea to synthetic captions and a new domain.
- Paper: DataComp: In search of the next generation of multimodal datasets, Samir Yitzhak Gadre et al. (2023). DataComp continues MSCN’s data-centric direction by systematically evaluating multimodal filtering and curation strategies at web scale.
- Paper: Contrastive Audio-Visual Masked Autoencoder, Yuan Gong et al. (2023). CAV-MAE extends the clean-correspondence and robust cross-modal training agenda from image-text retrieval to audio-visual representation learning.
