Built independently by an author, for readers. Read the story and support ChapterPal

keyword

noisy correspondence learning

Noisy correspondence learning is a machine learning paradigm focused on training robust models when paired data from different modalities, such as image-text or video-audio collections, contain misaligned or mismatched associations. Unlike traditional noisy label learning that involves incorrect category assignments for single items, noisy correspondence involves erroneous relationships between paired samples, where unrelated items are mistakenly treated as matching pairs. This challenge commonly arises in large-scale datasets collected automatically from the internet without human verification, which can mislead alignment and cross-modal retrieval algorithms into associating semantically incompatible data. Methods in this domain address the problem by estimating the reliability of sample pairs, filtering out or down-weighting mismatched examples, and rectifying similarity scores to preserve accurate semantic representations across modalities.

1 item

Noisy Correspondence Learning with Meta Similarity Correction

Noisy Correspondence Learning with Meta Similarity Correction

Haochen Han, Kaiyao Miao, Qinghua Zheng, Minnan Luo

OrganizationsXi'an Jiaotong University

Why you should read this

Proposes a meta-learning framework that trains a correction network on clean and mismatched meta-data to rectify similarity scores and filter out mismatched cross-modal pairs during retrieval training.

Despite the success of multimodal learning in cross-modal retrieval task, the remarkable progress relies on the correct correspondence among multimedia data. However, collecting such ideal data is expensive and time-consuming. In practice, most widely used datasets are harvested from the Internet and inevitably contain mismatched pairs. Training on such noisy correspondence datasets causes performance degradation because the cross-modal retrieval methods can wrongly enforce the mismatched data to be similar. To tackle this problem, we propose a Meta Similarity Correction Network (MSCN) to provide reliable similarity scores. We view a binary classification task as the meta-process that encourages the MSCN to learn discrimination from positive and negative meta-data. To further alleviate the influence of noise, we design an effective data purification strategy using meta-data as prior knowledge to remove the noisy samples. Extensive experiments are conducted to demonstrate the strengths of our method in both synthetic and real-world noises, including Flickr30K, MS-COCO, and Conceptual Captions. Our code is publicly available.

Added

2026-09-26