Noisy Correspondence Learning with Meta Similarity Correction

Haochen HanKaiyao MiaoQinghua ZhengMinnan Luo

article2023CVPR67 citations

Proposes a meta-learning framework that trains a correction network on clean and mismatched meta-data to rectify similarity scores and filter out mismatched cross-modal pairs during retrieval training.

Listen

Modern artificial intelligence systems increasingly rely on cross-modal retrieval to search and connect different types of data, such as finding images that match descriptive text. Standard algorithms assume that the image-text pairs used during training are accurately paired. In real-world applications, however, training data harvested from the web inevitably contains mismatched pairs, a challenge known as noisy correspondence. When models are trained on this misaligned data, they mistakenly learn to treat irrelevant concepts as similar, resulting in severe performance drops.

The article demonstrates a robust training framework called the Meta Similarity Correction Network (MSCN) to overcome noisy correspondence in cross-modal retrieval. The primary objective is to reliably estimate true multimodal similarity scores and cleanse noisy pairs using a small set of clean reference data.

The researchers evaluated their approach through extensive experiments on standard benchmark datasets (Flickr30K and MS-COCO) injected with synthetic noise ratios of 20%, 50%, and 70%, as well as a large-scale real-world dataset (Conceptual Captions CC152K). The framework frames similarity estimation as a binary classification task trained via bi-level optimization. It leverages both positive pairs and automatically constructed negative pairs from a small clean reference set—comprising only about 2% of the training size—as meta-knowledge. Additionally, the approach integrates an adaptive ranking margin and a two-component Beta Mixture Model to identify and remove mismatched pairs from the training pool before model updates occur.

The experimental findings show that MSCN consistently outperforms existing baseline methods across all noise levels. Under moderate noise conditions (20% to 50%), the method improves top-rank retrieval recall by 2.2% to 4.5% compared to the strongest noise-robust baseline. Under severe noise of 70%, where conventional methods experience catastrophic degradation, MSCN achieves dramatic recall improvements ranging from 27.3% to 52.9% over existing approaches, performing nearly on par with models trained strictly on clean data. On real-world noisy web data, it achieves overall sum recall improvements of 4.9% and 6.2% for text and image queries, respectively, while retaining strong performance even when the clean reference data is reduced to under 1% of the dataset.

These results indicate that organizations can train high-performing multimodal retrieval systems directly on inexpensive, uncurated web datasets without manual cleaning of the entire corpus. By relying on only a tiny fraction of verified data to guide the learning process, organizations can significantly lower data curation costs and project timelines while maintaining high search accuracy. Ablation analyses confirm that both the adaptive margin mechanism and the data purification step are critical to prevent models from overfitting to erroneous pairings.

Organizations developing cross-modal systems should adopt meta-correction and distribution-based sample purification strategies when working with unverified data. Given that the approach requires a small set of verified positive pairs (or a clean validation set) to establish baseline distributions, practitioners must ensure this minor reference set is available. The reported findings are supported by consistent results across standard benchmarks; however, future work could examine whether performance holds across other data types beyond image-text pairs, such as audio-video and multilingual multimodal tasks.

Cover for Noisy Correspondence Learning with Meta Similarity Correction

Abstract

Despite the success of multimodal learning in cross-modal retrieval task, the remarkable progress relies on the correct correspondence among multimedia data. However, collecting such ideal data is expensive and time-consuming. In practice, most widely used datasets are harvested from the Internet and inevitably contain mismatched pairs. Training on such noisy correspondence datasets causes performance degradation because the cross-modal retrieval methods can wrongly enforce the mismatched data to be similar. To tackle this problem, we propose a Meta Similarity Correction Network (MSCN) to provide reliable similarity scores. We view a binary classification task as the meta-process that encourages the MSCN to learn discrimination from positive and negative meta-data. To further alleviate the influence of noise, we design an effective data purification strategy using meta-data as prior knowledge to remove the noisy samples. Extensive experiments are conducted to demonstrate the strengths of our method in both synthetic and real-world noises, including Flickr30K, MS-COCO, and Conceptual Captions. Our code is publicly available.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Cross-Modal Retrieval
  • 2.2. Noisy Correspondence Learning
  • 2.3. Meta-Learning
  • 3. The Proposed Method
  • 3.1. Problem Formulation
  • 3.2. Meta-Learning Based Similarity Correction
  • 3.3. Meta-Knowledge Guided Data Purification
  • 4. Experiment
  • 4.1. Datasets and Evaluation Protocol
  • 4.2. Implementation Details
  • 4.3. Comparison with the State-of-the-Art
  • 4.4. Ablation Study
  • 4.5. Progressive Comparison
  • 4.6. Detected Noisy Samples
  • 4.7. Visualization on Similarity Score
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Meta Similarity Correction Network Architecture and Formulation

    model/method

    The Meta Similarity Correction Network (MSCN) addresses noisy correspondence in cross-modal retrieval, where mismatched multimodal pairs (such as mismatched image-text pairs) are incorrectly labeled as matching pairs in the training dataset Dtrain={(Ii,Ti,yi)}i=1N\mathcal{D}_{train} = \{(I_i, T_i, y_i)\}_{i=1}^N.

    Let f(I;Wf)f(I; W_f) and g(T;Wg)g(T; W_g) denote modal-specific embedding networks mapping visual input II and textual input TT into a joint multimodal embedding space. The similarity feature between visual and textual representations is computed as:

    FW(I,T)=S(f(I;Wf),g(T;Wg);Ws)=Ws∣f(I)−g(T)∣2∥Ws∣f(I)−g(T)∣2∥2,\mathcal{F}_W(I, T) = S(f(I; W_f), g(T; W_g); W_s) = \frac{W_s |f(I) - g(T)|^2}{\lVert W_s |f(I) - g(T)|^2\rVert_2},

    where WsW_s is a learnable projection matrix, and W={Wf,Wg,Ws}W = \{W_f, W_g, W_s\} denotes the complete set of main network parameters.

    To prevent the main network from overfitting to mismatched pairs, a multilayer perceptron (MLP) VΘ(⋅)\mathcal{V}_\Theta(\cdot) parameterized by Θ\Theta serves as the meta similarity correction network. It takes the similarity representation FW(Ii,Ti)\mathcal{F}_W(I_i, T_i) as input and outputs a scalar similarity score si∈[0,1]s_i \in [0, 1] via a Sigmoid activation:

    si=VΘ(FW(Ii,Ti)).s_i = \mathcal{V}_\Theta(\mathcal{F}_W(I_i, T_i)).

    The meta-process is formulated as a binary classification task supervised by a small clean meta-dataset Dmeta={(Ii,Ti)}i=1M\mathcal{D}_{meta} = \{(I_i, T_i)\}_{i=1}^M (M≪NM \ll N). For each iteration, MM negative (mismatched) pairs (Ii,Tj≠i,yi=0)(I_i, T_{j\neq i}, y_i=0) are sampled from different training instances and combined with the MM positive clean pairs (Ii,Ti,yi=1)(I_i, T_i, y_i=1) to construct an extended meta-dataset Dmeta′={(Ii,Ti,yi)}i=12M\mathcal{D}'_{meta} = \{(I_i, T_i, y_i)\}_{i=1}^{2M}. The optimal meta parameters Θ∗\Theta^* minimize the meta cross-entropy loss:

    Θ∗=arg⁡min⁡ΘE(Ii,Ti,yi)∈Dmeta′[−yi⋅log⁡VΘ(FW∗(Θ)(Ii,Ti))−(1−yi)⋅log⁡(1−VΘ(FW∗(Θ)(Ii,Ti)))].\Theta^* = \arg\min_\Theta \mathbb{E}_{(I_i, T_i, y_i) \in \mathcal{D}'_{meta}} \left[ -y_i \cdot \log \mathcal{V}_\Theta(\mathcal{F}_{W^*(\Theta)}(I_i, T_i)) - (1 - y_i) \cdot \log (1 - \mathcal{V}_\Theta(\mathcal{F}_{W^*(\Theta)}(I_i, T_i))) \right].

  2. Knowl 2 — Self-Adaptive Margin Triplet Ranking Loss for Noisy Correspondence

    equation

    In cross-modal retrieval under noisy correspondence, training with fixed margin triplet loss causes the model to force mismatched pairs together. To mitigate this, the triplet ranking loss is adapted with dynamic margins computed from predicted similarity scores.

    For a multimodal sample pair (Ii,Ti)(I_i, T_i), the triplet ranking loss is defined as:

    ltrain(Ii,Ti)=[γ^i−VΘ(FW(Ii,Ti))+VΘ(FW(Ii,Ti−))]++[γ^i−VΘ(FW(Ii,Ti))+VΘ(FW(Ii−,Ti))]+,l^{train}(I_i, T_i) = [\hat{\gamma}_i - \mathcal{V}_\Theta(\mathcal{F}_W(I_i, T_i)) + \mathcal{V}_\Theta(\mathcal{F}_W(I_i, T_i^-))]_+ + [\hat{\gamma}_i - \mathcal{V}_\Theta(\mathcal{F}_W(I_i, T_i)) + \mathcal{V}_\Theta(\mathcal{F}_W(I_i^-, T_i))]_+,

    where [x]+=max⁡(x,0)[x]_+ = \max(x, 0), Ii−I_i^- and Ti−T_i^- are the hardest negative image and text within the mini-batch, FW(⋅,⋅)\mathcal{F}_W(\cdot, \cdot) is the cross-modal similarity feature extractor, and VΘ(⋅)\mathcal{V}_\Theta(\cdot) is the meta similarity correction network outputting similarity score si=VΘ(FW(Ii,Ti))∈[0,1]s_i = \mathcal{V}_\Theta(\mathcal{F}_W(I_i, T_i)) \in [0, 1].

    The self-adaptive margin γ^i\hat{\gamma}_i is defined by:

    γ^i=11+(si1−si)−τγ,\hat{\gamma}_i = \frac{1}{1 + \left(\frac{s_i}{1 - s_i}\right)^{-\tau}} \gamma,

    where γ>0\gamma > 0 is the base margin value (typically γ=0.2\gamma = 0.2) and τ>0\tau > 0 is a scaling hyperparameter (set to τ=2\tau = 2). When si→1s_i \to 1 (high confidence clean pair), γ^i→γ\hat{\gamma}_i \to \gamma; when si→0s_i \to 0 (high confidence noisy pair), γ^i→0\hat{\gamma}_i \to 0, suppressing loss contribution from mismatched pairs.

  3. Knowl 3 — Bi-Level Alternating Optimization for Meta Similarity Correction

    algorithm

    The main embedding network parameters WW and meta similarity correction network (MSCN) parameters Θ\Theta are optimized via a bi-level nested gradient descent framework.

    Input: Training dataset Dtrain={(Ii,Ti)}D_{train} = \{(I_i, T_i)\}, clean meta set Dmeta={(Ii,Ti,yi=1)}D_{meta} = \{(I_i, T_i, y_i=1)\}, learning rates α,β\alpha, \beta, batch sizes n,mn, m, max iterations TmaxT_{max}
    Output: Optimal main network parameters WW
    Initialize main network W(0)W^{(0)} and meta network Θ(0)\Theta^{(0)}
    for iteration t=0t = 0 to Tmax−1T_{max} - 1 do
        Sample training mini-batch Btrain={(Ii,Ti)}i=1nB_{train} = \{(I_i, T_i)\}_{i=1}^n from DtrainD_{train}
        Virtual step: Compute temporary main network parameters
            W^(t)(Θ)=W(t)−α∑i=1n∇Wltrain(Ii,Ti)∣W(t)\hat{W}^{(t)}(\Theta) = W^{(t)} - \alpha \sum_{i=1}^n \nabla_W l^{train}(I_i, T_i)\big|_{W^{(t)}}
        Sample clean meta mini-batch {(Ii,Ti,yi=1)}i=1m/2\{(I_i, T_i, y_i=1)\}_{i=1}^{m/2} from DmetaD_{meta}
        Construct negative meta pairs {(Ii,Tj≠i,yi=0)}i=1m/2\{(I_i, T_{j \neq i}, y_i=0)\}_{i=1}^{m/2} from DtrainD_{train}
        Combine to form meta mini-batch Bmeta={(Ii,Ti,yi)}i=1mB_{meta} = \{(I_i, T_i, y_i)\}_{i=1}^m
        Update meta network parameters using W^(t)(Θ)\hat{W}^{(t)}(\Theta):
            Θ(t+1)=Θ(t)−β1m∑i=1m∇Θlmeta(Ii,Ti,yi;W^(t)(Θ))∣Θ(t)\Theta^{(t+1)} = \Theta^{(t)} - \beta \frac{1}{m} \sum_{i=1}^m \nabla_\Theta l^{meta}(I_i, T_i, y_i; \hat{W}^{(t)}(\Theta))\big|_{\Theta^{(t)}}
        Update main network parameters using updated meta network Θ(t+1)\Theta^{(t+1)}:
            W(t+1)=W(t)−α∑i=1n∇Wltrain(Ii,Ti;Θ(t+1))∣W(t)W^{(t+1)} = W^{(t)} - \alpha \sum_{i=1}^n \nabla_W l^{train}(I_i, T_i; \Theta^{(t+1)})\big|_{W^{(t)}}
    end for
    return WW

    In standard implementations, the model is trained with the Adam optimizer using batch size n=64n=64, initial main network learning rate α=2×10−4\alpha = 2 \times 10^{-4}, and meta network learning rate β=1.7×10−5\beta = 1.7 \times 10^{-5}, with learning rates decayed by a factor of 0.10.1 at epoch 30 across 50 epochs.

  4. Knowl 4 — Meta-Knowledge Guided Beta Mixture Model Data Purification

    model/method

    Because cross-modal triplet loss can produce positive penalties even for pairs with near-zero similarity scores (V(F(Ii,Ti))≈0<V(F(Ii−,Ti))+γ^V(F(I_i, T_i)) \approx 0 < V(F(I_i^-, T_i)) + \hat{\gamma}), noisy pairs must be explicitly filtered from training. While Gaussian Mixture Models (GMMs) are conventional for loss modeling, the distribution of similarity scores si∈[0,1]s_i \in [0, 1] exhibits strong positive skew toward 1, making Beta Mixture Models (BMMs) substantially more accurate.

    The distribution of per-sample similarity scores si=VΘ(FW(Ii,Ti))s_i = \mathcal{V}_\Theta(\mathcal{F}_W(I_i, T_i)) is modeled with a two-component Beta Mixture Model:

    p(si)=∑k=12λkϕ(si∣αk,βk),p(s_i) = \sum_{k=1}^2 \lambda_k \phi(s_i | \alpha_k, \beta_k),

    where λk\lambda_k is the mixture weight with ∑kλk=1\sum_k \lambda_k = 1, and ϕ(s∣αk,βk)=Γ(αk+βk)Γ(αk)Γ(βk)sαk−1(1−s)βk−1\phi(s | \alpha_k, \beta_k) = \frac{\Gamma(\alpha_k + \beta_k)}{\Gamma(\alpha_k)\Gamma(\beta_k)} s^{\alpha_k - 1}(1 - s)^{\beta_k - 1} is the Beta probability density function with shape parameters αk,βk>0\alpha_k, \beta_k > 0.

    After optimizing the BMM with the Expectation-Maximization (EM) algorithm, the posterior probability of sample ii belonging to the clean component (k=cleank = \text{clean}) is evaluated as:

    pi=p(k=clean∣si)=λcleanϕ(si∣αclean,βclean)p(si).p_i = p(k = \text{clean} | s_i) = \frac{\lambda_{\text{clean}} \phi(s_i | \alpha_{\text{clean}}, \beta_{\text{clean}})}{p(s_i)}.

    The purified training subset is defined as:

    Dtrain′={(Ii,Ti)∈Dtrain∣pi>0.5}.\mathcal{D}'_{train} = \{(I_i, T_i) \in \mathcal{D}_{train} \mid p_i > 0.5\}.

    To prevent confirmation bias and error accumulation, dual networks {FW1,VΘ1}\{\mathcal{F}_{W^1}, \mathcal{V}_{\Theta^1}\} and {FW2,VΘ2}\{\mathcal{F}_{W^2}, \mathcal{V}_{\Theta^2}\} are maintained in a co-training scheme, where the purified dataset produced by one network is used to train the other network.

  5. Knowl 5 — Closed-Form Moment Matching Initialization for Beta Mixture Models via Meta-Data

    equation

    Standard Expectation-Maximization (EM) for Beta Mixture Models uses KK-Means clustering for parameter initialization, which has time complexity O(TKN)\mathcal{O}(TKN) (where TT is iteration count, KK is component count, and NN is total sample count) and often converges to poor local optima.

    By leveraging the clean meta-data and dynamically constructed negative meta-data as prior distributions, the initial parameters for the clean component (k=1k=1) and noisy component (k=2k=2) are calculated in closed form using method of moments directly from meta similarity score subsets Sp={si}i=1M\mathcal{S}_p = \{s_i\}_{i=1}^M (positive meta-pairs) and Sn={si}i=1M\mathcal{S}_n = \{s_i\}_{i=1}^M (negative meta-pairs):

    αk=(1−E(S))(E(S))2V(S)−E(S),\alpha_k = \frac{(1 - E(\mathcal{S})) (E(\mathcal{S}))^2}{V(\mathcal{S})} - E(\mathcal{S}),

    βk=αk(1−E(S))E(S),\beta_k = \frac{\alpha_k (1 - E(\mathcal{S}))}{E(\mathcal{S})},

    where E(S)E(\mathcal{S}) denotes the sample mean 1M∑j=1Msj\frac{1}{M} \sum_{j=1}^M s_j and V(S)V(\mathcal{S}) denotes the sample variance 1M−1∑j=1M(sj−E(S))2\frac{1}{M-1} \sum_{j=1}^M (s_j - E(\mathcal{S}))^2 computed over set S∈{Sp,Sn}\mathcal{S} \in \{\mathcal{S}_p, \mathcal{S}_n\}.

    This closed-form initialization reduces initialization computational complexity to O(KM)\mathcal{O}(KM) while providing accurate prior approximations for EM convergence.

  6. Knowl 6 — Cross-Modal Retrieval Performance under Synthetic Noise on Flickr30K and MS-COCO

    data/table

    The performance of MSCN is evaluated against state-of-the-art cross-modal retrieval methods on Flickr30K and MS-COCO 1K under synthetic noise ratios of 20%, 50%, and 70%. Noise is introduced by randomly shuffling image and text pairings. Results are reported using recall at rank KK (R@1, R@5, R@10) for both Image-to-Text and Text-to-Image retrieval.

    Flickr30K MS-COCO
    Image to Text Text to Image Image to Text Text to Image
    Noise Methods R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10 R@1 R@5 R@10
    20% SCAN 59.1 83.4 90.4 36.6 67.0 77.5 66.2 91.0 96.4 45.0 80.2 89.3
    VSRN 58.1 82.6 89.3 40.7 68.7 78.2 25.1 59.0 74.8 17.6 49.0 64.1
    IMRAM 63.0 86.0 91.3 41.4 71.2 80.5 68.6 92.8 97.6 55.7 85.0 91.0
    SAF 51.0 79.3 88.0 38.3 66.5 76.2 67.3 92.5 96.6 53.4 84.5 92.4
    SGR* 62.8 86.2 92.2 44.4 72.3 80.4 67.8 91.7 96.2 52.9 83.5 90.1
    SGR-C 72.8 90.8 95.4 56.4 82.1 88.6 75.4 95.2 97.9 60.1 88.5 94.8
    NCR 75.0 93.9 97.5 58.3 83.0 89.0 77.7 95.5 98.2 62.5 89.3 95.3
    MSCN 77.4 94.9 97.6 59.6 83.2 89.2 78.1 97.2 98.8 64.3 90.4 95.8
    50% SCAN 27.7 57.6 68.8 16.2 39.3 49.8 40.8 73.5 84.9 5.4 15.1 21.0
    VSRN 14.3 37.6 50.0 12.1 30.0 39.4 23.5 54.7 69.3 16.0 47.8 65.9
    IMRAM 9.1 26.6 38.2 2.7 8.4 12.7 21.3 60.2 75.9 22.3 52.8 64.3
    SAF 30.3 63.6 75.4 27.9 53.7 65.1 30.4 67.8 82.3 33.5 69.0 82.8
    SGR* 36.9 68.1 80.2 29.3 56.2 67.0 60.6 87.4 93.6 46.0 74.2 79.0
    SGR-C 69.8 90.3 94.8 50.1 77.5 85.2 71.7 94.1 97.7 57.0 86.6 93.7
    NCR 72.9 93.0 96.3 54.3 79.8 86.5 74.6 94.6 97.8 59.1 87.8 94.5
    MSCN 74.4 93.2 96.0 55.3 80.4 86.8 77.5 96.2 98.7 60.7 89.1 94.9
    70% SCAN 5.6 19.3 27.4 2.2 8.0 12.8 18.1 43.1 57.4 0.3 1.3 2.3
    VSRN 0.8 2.5 4.1 0.5 1.5 2.7 5.1 15.7 24.6 2.5 8.8 13.3
    IMRAM 1.3 3.1 3.9 0.3 1.2 2.8 7.1 20.0 33.4 5.3 15.2 22.0
    SAF 0.5 2.2 3.0 0.2 0.8 1.7 0.1 1.7 4.0 0.6 1.9 3.0
    SGR* 17.9 42.1 51.9 14.6 31.0 40.8 35.7 71.2 85.4 31.6 65.8 79.0
    SGR-C 65.0 89.3 94.7 48.1 74.5 81.1 69.8 93.6 97.5 56.5 86.0 93.4
    NCR 16.1 38.5 52.8 11.0 29.5 41.4 35.4 69.5 83.4 31.5 66.4 81.1
    MSCN 69.0 89.3 93.0 49.2 73.1 79.0 74.4 94.9 97.7 58.8 87.2 93.7

    Under high noise (70%), all baseline methods suffer severe degradation due to noise overfitting (NCR drops to 16.1% on Flickr30K R@1 I2T). In contrast, MSCN achieves 69.0% R@1 I2T on Flickr30K and 74.4% on MS-COCO, outperforming SGR-C (trained purely on clean data).

  7. Knowl 7 — Cross-Modal Retrieval Performance under Real-World Noise on CC152K

    data/table

    The Conceptual Captions 152K (CC152K) dataset contains natural, real-world web noise estimated between 3% and 20%. Evaluation is performed on 1,000 test pairs using recall metrics (R@1, R@5, R@10) and their cumulative sum (SUM). MSCN uses 3,000 clean pairs from validation as meta-data, whereas MSCN* uses a reduced meta set of only 1,000 pairs (~0.67% of total data).

    Image to Text Text to Image
    Methods R@1 R@5 R@10 SUM R@1 R@5 R@10 SUM
    SCAN (ECCV'18) 30.5 55.3 65.3 151.1 26.9 53.0 64.7 144.6
    VSRN (ICCV'19) 32.6 61.3 70.5 164.4 32.5 59.4 70.4 162.3
    IMRAM (CVPR'20) 33.1 57.6 68.1 158.8 29.0 56.8 67.4 153.2
    SAF (AAAI'21) 31.7 59.3 68.2 159.2 31.9 59.0 67.9 158.8
    SGR (AAAI'21) 11.3 29.7 39.6 80.6 13.1 30.1 41.6 84.8
    SGR* (AAAI'21) 35.0 63.4 73.3 171.7 34.9 63.0 72.8 170.7
    NCR (NIPS'21) 39.5 64.5 73.5 177.5 40.3 64.6 73.2 178.1
    MSCN* (Ours) 39.7 65.4 75.3 180.4 39.8 66.1 75.0 180.9
    MSCN (Ours) 40.1 65.7 76.6 182.4 40.6 67.4 76.3 184.3

    MSCN achieves state-of-the-art performance across all metrics on CC152K, surpassing NCR by +4.9 points on Image-to-Text SUM (182.4 vs. 177.5) and +6.2 points on Text-to-Image SUM (184.3 vs. 178.1).

  8. Knowl 8 — Ablation Study of Self-Adaptive Margin and Data Purification on Flickr30K

    data/table

    An ablation study evaluated the individual contributions of the self-adaptive margin γ^\hat{\gamma} and the meta-guided data purification subset Dtrain′\mathcal{D}'_{train} on Flickr30K with 20% synthetic noise rate. For models trained without data purification (w/o Dtrain′\mathcal{D}'_{train}), a single model is used instead of dual-network co-training.

    Method Image to Text Text to Image
    MSCN w/o γ^\hat{\gamma} w/o Dtrain′\mathcal{D}'_{train} R@1 R@5 R@10 R@1 R@5 R@10
    ✓ 77.4 94.9 97.6 59.6 83.2 89.2
    ✓ ✓ 75.3 94.5 97.2 58.3 83.1 88.9
    ✓ ✓ 75.8 93.8 96.2 55.8 74.5 77.4
    ✓ ✓ ✓ 74.1 91.5 94.7 53.4 71.2 72.2

    Removing the self-adaptive margin drops Image-to-Text R@1 by 2.1% (75.3 vs. 77.4) and Text-to-Image R@1 by 1.3% (58.3 vs. 59.6). Removing data purification drops Image-to-Text R@1 by 1.6% and Text-to-Image R@1 by 3.8%. Removing both components results in the largest performance drop (Image-to-Text R@1 drops to 74.1% and Text-to-Image R@1 drops to 53.4%).

Coverage note — None was omitted; all key contributions including MSCN architecture, self-adaptive margin, bi-level optimization, BMM initialization, and comprehensive experimental results on synthetic and real datasets have been extracted into knowls.

References

  1. 1.Relja Arandjelovic and Andrew Zisserman. Objects that sound. In Proceedings of the European conference on computer vision (ECCV), pages 435–451, 2018.
  2. 2.Eric Arazo, Diego Ortego, Paul Albert, Noel O’Connor, and Kevin McGuinness. Unsupervised label noise modeling and loss correction. In International conference on machine learning, pages 312–321. PMLR, 2019.
  3. 3.Devansh Arpit, Stanislaw K Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron C Courville, Yoshua Bengio, et al. A closer look at memorization in deep networks. In ICML, 2017.
  4. 4.Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12655–12663, 2020.
  5. 5.Arthur P Dempster, Nan M Laird, and Donald B Rubin. Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22, 1977.
  6. 6.Cheng Deng, Zhaojia Chen, Xianglong Liu, Xinbo Gao, and Dacheng Tao. Triplet-based deep hashing network for cross-modal retrieval. IEEE Transactions on Image Processing, 27(8):3893–3903, 2018.
  7. 7.Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. Similarity reasoning and filtration for image-text matching. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1218–1226, 2021.
  8. 8.Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. Vse++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612, 2017.
  9. 9.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning, pages 1126–1135. PMLR, 2017.
  10. 10.Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil. Bilevel programming for hyperparameter optimization and meta-learning. In International Conference on Machine Learning, pages 1568–1577. PMLR, 2018.
  11. 11.Dalu Guo, Chang Xu, and Dacheng Tao. Image-question-answer synergistic network for visual dialog. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10434–10443, 2019.
  12. 12.Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. Co-teaching: Robust training of deep neural networks with extremely noisy labels. Advances in neural information processing systems, 31, 2018.
  13. 13.Haochen Han, Qinghua Zheng, Minnan Luo, Kaiyao Miao, Feng Tian, and Yan Chen. Noise-tolerant learning for audiovisual action recognition. arXiv preprint arXiv:2205.07611, 2022.
  14. 14.Peng Hu, Xi Peng, Hongyuan Zhu, Liangli Zhen, and Jie Lin. Learning cross-modal retrieval with noisy labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5403–5413, 2021.
  15. 15.Zhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding, Xinyan Xiao, Hua Wu, and Xi Peng. Learning with noisy correspondence for cross-modal matching. Advances in Neural Information Processing Systems, 34:29406–29419, 2021.
  16. 16.Haoxuanye Ji, Le Wang, Sanping Zhou, Wei Tang, Nanning Zheng, and Gang Hua. Meta pairwise relationship distillation for unsupervised person re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3661–3670, 2021.
  17. 17.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR (Poster), 2015.
  18. 18.Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In Proceedings of the European conference on computer vision (ECCV), pages 201–216, 2018.
  19. 19.Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy Hospedales. Learning to generalize: Meta-learning for domain generalization. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018.
  20. 20.Junnan Li, Richard Socher, and Steven CH Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. In International Conference on Learning Representations, 2019.
  21. 21.Junnan Li, Yongkang Wong, Qi Zhao, and Mohan S Kankanhalli. Learning to learn from noisy labeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5051–5059, 2019.
  22. 22.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Visual semantic reasoning for image-text matching. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4654–4662, 2019.
  23. 23.Aristidis Likas, Nikos Vlassis, and Jakob J Verbeek. The global k-means clustering algorithm. Pattern recognition, 36(2):451–461, 2003.
  24. 24.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  25. 25.Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. Graph structured network for image-text matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10921–10930, 2020.
  26. 26.Lin Ma, Zhengdong Lu, and Hang Li. Learning to answer questions from image using convolutional neural network. In Thirtieth AAAI Conference on Artificial Intelligence, 2016.
  27. 27.Zhanyu Ma and Arne Leijon. Bayesian estimation of beta mixture models with variational inference. IEEE Transactions on Pattern Analysis and Machine Intelligence, 33(11):2160–2173, 2011.
  28. 28.Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. A simple neural attentive meta-learner. In International Conference on Learning Representations, 2018.
  29. 29.Yang Qin, Dezhong Peng, Xi Peng, Xu Wang, and Peng Hu. Deep evidential learning with noisy correspondence for cross-modal retrieval. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4948–4956, 2022.
  30. 30.Mengye Ren, Wenyuan Zeng, Bin Yang, and Raquel Urtasun. Learning to reweight examples for robust deep learning. In International conference on machine learning, pages 4334–4343. PMLR, 2018.
  31. 31.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018.
  32. 32.Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. Meta-weight-net: Learning an explicit mapping for sample weighting. Advances in neural information processing systems, 32, 2019.
  33. 33.Didac Surís, Amanda Duarte, Amaia Salvador, Jordi Torres, and Xavier Giró-i Nieto. Cross-modal embeddings for video and audio retrieval. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018.
  34. 34.Jun Tang, Ke Wang, and Ling Shao. Supervised matrix factorization hashing for cross-modal retrieval. IEEE Transactions on Image Processing, 25(7):3157–3166, 2016.
  35. 35.Bokun Wang, Yang Yang, Xing Xu, Alan Hanjalic, and Heng Tao Shen. Adversarial cross-modal retrieval. In Proceedings of the 25th ACM international conference on Multimedia, pages 154–162, 2017.
  36. 36.Liwei Wang, Yin Li, and Svetlana Lazebnik. Learning deep structure-preserving image-text embeddings. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5005–5013, 2016.
  37. 37.Yunchao Wei, Yao Zhao, Canyi Lu, Shikui Wei, Luoqi Liu, Zhenfeng Zhu, and Shuicheng Yan. Cross-modal retrieval with cnn visual features: A new baseline. IEEE transactions on cybernetics, 47(2):449–460, 2016.
  38. 38.Xing Xu, Li He, Huimin Lu, Lianli Gao, and Yanli Ji. Deep adversarial metric learning for cross-modal retrieval. World Wide Web, 22(2):657–672, 2019.
  39. 39.Mouxing Yang, Zhenyu Huang, Peng Hu, Taihao Li, Jiancheng Lv, and Xi Peng. Learning with twin noisy labels for visible-infrared person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14308–14317, 2022.
  40. 40.Mouxing Yang, Yunfan Li, Peng Hu, Jinfeng Bai, Jiancheng Lv, and Xi Peng. Robust multi-view clustering with incomplete information. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):1055–1069, 2022.
  41. 41.Mouxing Yang, Yunfan Li, Zhenyu Huang, Zitao Liu, Peng Hu, and Xi Peng. Partially view-aligned representation learning with noise-robust contrastive loss. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1134–1143, 2021.
  42. 42.Quanming Yao, Hansi Yang, Bo Han, Gang Niu, and James Tin-Yau Kwok. Searching to exploit memorization effect in learning with noisy labels. In International Conference on Machine Learning, pages 10789–10798. PMLR, 2020.
  43. 43.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  44. 44.Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z Li. Context-aware attention network for image-text retrieval. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3536–3545, 2020.
  45. 45.Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng. Deep supervised cross-modal retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10394–10403, 2019.
  46. 46.Guoqing Zheng, Ahmed Hassan Awadallah, and Susan Dumais. Meta label correction for noisy label learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 11053–11061, 2021.

Citation

MLA
Han, H., et al. “Noisy Correspondence Learning with Meta Similarity Correction”. arXiv, 2023, http://arxiv.org/abs/2304.06275v1.
APA
Han, H., Miao, K., Zheng, Q., & Luo, M. (2023). Noisy Correspondence Learning with Meta Similarity Correction. arXiv. http://arxiv.org/abs/2304.06275v1
Chicago
Han, H., K. Miao, Q. Zheng, and M. Luo. 2023. “Noisy Correspondence Learning with Meta Similarity Correction”. arXiv. http://arxiv.org/abs/2304.06275v1.
Harvard
Han, H. et al. (2023) “Noisy Correspondence Learning with Meta Similarity Correction”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.06275v1.
Vancouver
1. Han H, Miao K, Zheng Q, Luo M (2023) Noisy Correspondence Learning with Meta Similarity Correction. arXiv

BibTeX

@article{han2023noisy,
  title = {Noisy Correspondence Learning with Meta Similarity Correction},
  author = {Han, Haochen and Miao, Kaiyao and Zheng, Qinghua and Luo, Minnan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.06275v1},
  eprint = {2304.06275}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE