Learning to Reweight Examples for Robust Deep Learning

Mengye RenWenyuan ZengBin YangRaquel Urtasun

article2018ICML1,714 citations

Proposes a meta-learning method that dynamically weights training examples by matching their gradient directions to a clean validation set, eliminating manual hyperparameter tuning while effectively training deep networks on corrupted or imbalanced data.

Listen

Modern deep learning systems rely heavily on massive datasets that often contain significant flaws, such as extreme class imbalances or incorrect labels. Standard neural networks tend to memorize these errors and biases, leading to degraded performance when deployed in real-world environments. Existing remediation techniques—such as manual data resampling or weighting examples strictly by training loss—frequently require tedious trial-and-error tuning and fail when data is simultaneously noisy and imbalanced. Addressing these vulnerabilities is critical for high-stakes applications, such as autonomous driving and medical imaging, where rare classes must be recognized reliably despite imperfect training data.

The article evaluates an automated meta-learning method designed to dynamically reweight training examples during standard model optimization. The primary objective is to demonstrate that a deep learning model can achieve high robustness against severe label corruption and class imbalance by leveraging a tiny set of clean, unbiased validation data.

The evaluated approach optimizes training dynamically in an online manner. During every training step, the algorithm inspects the gradient direction of each training example and compares it to the gradient direction required to minimize loss on a small, clean validation set. If an example's gradient aligns with the validation objective, the algorithm assigns it a higher weight; if it conflicts, the algorithm assigns it a weight of zero. The researchers benchmarked this technique against standard baselines and prior specialized algorithms using standard image classification datasets (MNIST and CIFAR) under varied conditions of class imbalance and artificial label corruption.

The findings demonstrate substantial improvements in model robustness across multiple demanding scenarios. First, under extreme class imbalance (a 200:1 ratio on binary MNIST), the proposed approach limited the test error increase to roughly 2%, significantly outperforming conventional resampling and hard-example mining techniques. Second, under a severe 40% uniform label noise condition on CIFAR-100, the method achieved a 61.34% test accuracy, outperforming existing state-of-the-art models by over 3 percentage points. Third, when label noise was increased incrementally from 0% to 50% on CIFAR-10, the method's accuracy dropped by only 6%, compared to a catastrophic decline of more than 40% observed in standard baselines. Fourth, the experiments revealed that as few as 15 to 100 clean validation images across all classes are sufficient to guide the entire training process effectively.

These results demonstrate that an organization does not need perfectly curated massive datasets to train high-performing deep networks. Instead, investments can be focused on acquiring a very small, highly accurate validation set to supervise the learning from inexpensive, coarsely labeled data. Furthermore, the approach eliminates the operational risk of models overfitting to corrupted data over time, removing the need to carefully engineer early stopping schedules or fine-tuning pipelines. Because the clean validation data acts primarily as a dynamic regularizer rather than direct training data, the workflow avoids the severe overfitting typical of training directly on small sample sizes.

Organizations handling imperfect datasets should consider adopting gradient-based online reweighting pipelines to lower annotation costs and improve model reliability. Implementation introduces an approximate threefold increase in per-iteration computational training time due to nested differentiation passes; however, this trade-off is often offset by eliminating hyperparameter search time. Before full-scale deployment, teams should conduct pilot implementations to verify that the validation set is truly representative and unbiased, as the entire optimization trajectory aligns with this reference sample. Future evaluations should focus on extending this mechanism to other domains beyond image classification, such as natural language processing and multimodal systems.

Cover for Learning to Reweight Examples for Robust Deep Learning

Abstract

Deep neural networks have been shown to be very powerful modeling tools for many supervised learning tasks involving complex input patterns. However, they can also easily overfit to training set biases and label noises. In addition to various regularizers, example reweighting algorithms are popular solutions to these problems, but they require careful tuning of additional hyperparameters, such as example mining schedules and regularization hyperparameters. In contrast to past reweighting methods, which typically consist of functions of the cost value of each example, in this work we propose a novel meta-learning algorithm that learns to assign weights to training examples based on their gradient directions. To determine the example weights, our method performs a meta gradient descent step on the current mini-batch example weights (which are initialized from zero) to minimize the loss on a clean unbiased validation set. Our proposed method can be easily implemented on any type of deep network, does not require any additional hyperparameter tuning, and achieves impressive performance on class imbalance and corrupted label problems where only a small amount of clean validation data is available.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Learning to Reweight Examples
  • 3.1 From a meta-learning objective to an online approximation
  • 3.2 Example: learning to reweight examples in a multi-layer perceptron network
  • 3.3 Implementation using automatic differentiation
  • 3.4 Analysis: convergence of the reweighted training
  • 4 Experiments
  • 4.1 MNIST data imbalance experiments
  • 4.2 CIFAR noisy label experiments
  • 4.3 Results and Discussion
  • 5 Conclusion
  • References
  • A Reweighting in an MLP
  • B Convergence of our method
  • C Convergence rate of our method

Knowls

  1. Knowl 1 — Online Meta-Learning Objective for Example Reweighting

    model/method

    Let a deep neural network parametrized by θ\theta have training loss fi(θ)=C(Φ(xi,θ),yi)f_i(\theta) = C(\Phi(x_i, \theta), y_i) on training example (xi,yi)(x_i, y_i) from a training dataset {(xi,yi)}i=1N\{(x_i, y_i)\}_{i=1}^N, and validation loss fjv(θ)=C(Φ(xjv,θ),yjv)f_j^v(\theta) = C(\Phi(x_j^v, \theta), y_j^v) on clean validation example (xjv,yjv)(x_j^v, y_j^v) from a small unbiased validation dataset {(xjv,yjv)}j=1M\{(x_j^v, y_j^v)\}_{j=1}^M with M≪NM \ll N. The ideal sample reweighting problem is framed as a bi-level meta-learning objective:

    θ∗(w)=arg⁡min⁡θ∑i=1Nwifi(θ),subject tow∗=arg⁡min⁡w≥01M∑j=1Mfjv(θ∗(w))\theta^*(w) = \arg\min_\theta \sum_{i=1}^N w_i f_i(\theta), \quad \text{subject to} \quad w^* = \arg\min_{w \ge 0} \frac{1}{M} \sum_{j=1}^M f_j^v(\theta^*(w))

    To bypass the prohibitive cost of nested optimization loops, an online local approximation is performed at training step tt on a mini-batch of nn training examples. A local weighting perturbation vector ϵ=(ϵ1,…,ϵn)\epsilon = (\epsilon_1, \dots, \epsilon_n) is introduced around zero:

    θ^t+1(ϵ)=θt−α∇θ∑i=1nϵifi(θ)∣θ=θt\hat{\theta}_{t+1}(\epsilon) = \theta_t - \alpha \nabla_\theta \left. \sum_{i=1}^n \epsilon_i f_i(\theta) \right|_{\theta = \theta_t}

    where α\alpha is the step size. The meta-gradient of a mini-batch of mm clean validation examples with respect to ϵi\epsilon_i evaluated at ϵi=0\epsilon_i = 0 is computed as:

    ui,t=−η∂∂ϵi,t1m∑j=1mfjv(θ^t+1(ϵ))∣ϵi,t=0u_{i,t} = -\eta \left. \frac{\partial}{\partial \epsilon_{i,t}} \frac{1}{m} \sum_{j=1}^m f_j^v(\hat{\theta}_{t+1}(\epsilon)) \right|_{\epsilon_{i,t}=0}

    To ensure non-negativity and eliminate dependence on the meta-learning rate η\eta, the raw weights are rectified and batch-normalized:

    w~i,t=max⁡(ui,t,0),wi,t=w~i,t∑j=1nw~j,t+δ(∑j=1nw~j,t)\tilde{w}_{i,t} = \max(u_{i,t}, 0), \quad w_{i,t} = \frac{\tilde{w}_{i,t}}{\sum_{j=1}^n \tilde{w}_{j,t} + \delta\left(\sum_{j=1}^n \tilde{w}_{j,t}\right)}

    where δ(a)=1\delta(a) = 1 if a=0a = 0 and δ(a)=0\delta(a) = 0 otherwise, preventing division by zero when all weights in a batch are zero.

  2. Knowl 2 — Automatic Differentiation Algorithm for Online Example Reweighting

    algorithm

    The online example reweighting procedure updates model parameters θ\theta while adaptively computing example importance weights ww in every mini-batch iteration using second-order automatic differentiation.

    Input: Initial parameters θ0\theta_0, training set DfD_f, clean validation set DgD_g, training batch size nn, validation batch size mm, total steps TT, training learning rate α\alpha
    Output: Final trained parameters θT\theta_T
    for t=0t = 0 to T−1T-1 do
        {Xf,yf}←SampleMiniBatch(Df,n)\{X_f, y_f\} \leftarrow \text{SampleMiniBatch}(D_f, n)
        {Xg,yg}←SampleMiniBatch(Dg,m)\{X_g, y_g\} \leftarrow \text{SampleMiniBatch}(D_g, m)
        y^f←Forward(Xf,yf,θt)\hat{y}_f \leftarrow \text{Forward}(X_f, y_f, \theta_t)
        ϵ←0\epsilon \leftarrow 0
        lf←∑i=1nϵiC(yf,i,y^f,i)l_f \leftarrow \sum_{i=1}^n \epsilon_i C(y_{f,i}, \hat{y}_{f,i})
        ∇θt←BackwardAD(lf,θt)\nabla \theta_t \leftarrow \text{BackwardAD}(l_f, \theta_t)
        θ^t←θt−α∇θt\hat{\theta}_t \leftarrow \theta_t - \alpha \nabla \theta_t
        y^g←Forward(Xg,yg,θ^t)\hat{y}_g \leftarrow \text{Forward}(X_g, y_g, \hat{\theta}_t)
        lg←1m∑i=1mC(yg,i,y^g,i)l_g \leftarrow \frac{1}{m} \sum_{i=1}^m C(y_{g,i}, \hat{y}_{g,i})
        ∇ϵ←BackwardAD(lg,ϵ)\nabla \epsilon \leftarrow \text{BackwardAD}(l_g, \epsilon)
        w~←max⁡(−∇ϵ,0)\tilde{w} \leftarrow \max(-\nabla \epsilon, 0)
        w←w~∑jw~j+δ(∑jw~j)w \leftarrow \frac{\tilde{w}}{\sum_j \tilde{w}_j + \delta(\sum_j \tilde{w}_j)}
        l^f←∑i=1nwiC(yf,i,y^f,i)\hat{l}_f \leftarrow \sum_{i=1}^n w_i C(y_{f,i}, \hat{y}_{f,i})
        ∇θt←BackwardAD(l^f,θt)\nabla \theta_t \leftarrow \text{BackwardAD}(\hat{l}_f, \theta_t)
        θt+1←OptimizerStep(θt,∇θt)\theta_{t+1} \leftarrow \text{OptimizerStep}(\theta_t, \nabla \theta_t)
    end for
    return θT\theta_T

    The algorithm first unrolls the computational graph for one gradient descent step with dummy perturbation weights ϵ=0\epsilon = 0, evaluates the loss on the validation batch at the lookahead parameter θ^t\hat{\theta}_t, takes a backward-on-backward pass to obtain ∇ϵ\nabla \epsilon, rectifies and normalizes the weights, and finally updates θt\theta_t along the reweighted training loss gradient.

  3. Knowl 3 — Layerwise Gradient and Activation Alignment Decomposition

    equation

    For a multi-layer perceptron (MLP) with LL layers parametrized by θ={θl}l=1L\theta = \{\theta_l\}_{l=1}^L, where zl=θl⊤z~l−1z_l = \theta_l^\top \tilde{z}_{l-1} is the pre-activation vector, z~l=σ(zl)\tilde{z}_l = \sigma(z_l) is the post-activation vector, and gl=∇zlf(θ)g_l = \nabla_{z_l} f(\theta) is the loss gradient with respect to pre-activations, the unnormalized meta-gradient of the expected validation loss with respect to training sample perturbation ϵi,t\epsilon_{i,t} is given by:

    ∂∂ϵi,tE[fv(θ^t+1(ϵ))]∣ϵi,t=0∝−1m∑j=1m∑l=1L(z~j,l−1v⊤z~i,l−1)(gj,lv⊤gi,l)\left. \frac{\partial}{\partial \epsilon_{i,t}} \mathbb{E}[f^v(\hat{\theta}_{t+1}(\epsilon))] \right|_{\epsilon_{i,t}=0} \propto - \frac{1}{m} \sum_{j=1}^m \sum_{l=1}^L (\tilde{z}_{j,l-1}^{v \top} \tilde{z}_{i,l-1}) (g_{j,l}^{v \top} g_{i,l})

    where z~i,l−1\tilde{z}_{i,l-1} and gi,lg_{i,l} denote the activation and gradient vectors of training example ii, and z~j,l−1v\tilde{z}_{j,l-1}^v and gj,lvg_{j,l}^v denote the activation and gradient vectors of validation example jj.

    This shows that the meta-gradient is a sum of products of two scalar alignments across all layers: the feature similarity z~j,l−1v⊤z~i,l−1\tilde{z}_{j,l-1}^{v \top} \tilde{z}_{i,l-1} and the gradient direction similarity gj,lv⊤gi,lg_{j,l}^{v \top} g_{i,l}. Training examples whose features and descent directions align positively with validation samples receive positive weight updates, whereas conflicting examples receive negative updates and are subsequently zeroed out.

  4. Knowl 4 — Convergence Guarantees of Online Meta-Reweighted Training

    theoretical result

    Let the total validation loss G(θ)=1M∑i=1Mfiv(θ)G(\theta) = \frac{1}{M} \sum_{i=1}^M f_i^v(\theta) be Lipschitz-smooth with constant LL, meaning ∥∇G(x)−∇G(y)∥≤L∥x−y∥\|\nabla G(x) - \nabla G(y)\| \le L\|x - y\| for all x,y∈Rdx, y \in \mathbb{R}^d. Assume each training loss function fi(θ)f_i(\theta) has σ\sigma-bounded gradients, ∥∇fi(x)∥≤σ\|\nabla f_i(x)\| \le \sigma. If the training step size αt\alpha_t satisfies:

    αt≤2nLσ2\alpha_t \le \frac{2n}{L \sigma^2}

    where nn is the training mini-batch size, then the validation loss monotonically decreases at each step tt:

    G(θt+1)≤G(θt)G(\theta_{t+1}) \le G(\theta_t)

    Furthermore, the expected validation loss satisfies Et[G(θt+1)]=G(θt)\mathbb{E}_t[G(\theta_{t+1})] = G(\theta_t) if and only if ∇G(θt)=0\nabla G(\theta_t) = 0.

    Under these conditions, the algorithm converges to an ϵ\epsilon-critical point satisfying E[∥∇G(θt)∥2]≤ϵ\mathbb{E}[\|\nabla G(\theta_t)\|^2] \le \epsilon in O(1/ϵ2)O(1/\epsilon^2) steps. Specifically:

    min⁡0≤t<TE[∥∇G(θt)∥2]≤CT\min_{0 \le t < T} \mathbb{E}[\|\nabla G(\theta_t)\|^2] \le \frac{C}{\sqrt{T}}

    where CC is a constant independent of TT, matching the convergence rate of standard stochastic gradient descent.

  5. Knowl 5 — Classification Accuracy on CIFAR under Corrupted Labels

    data/table

    The online example reweighting method was evaluated on CIFAR-10 and CIFAR-100 under two 40% label corruption regimes: UniformFlip (uniform random flipping across all classes using WideResNet-28-10 with 1,000 clean validation images) and BackgroundFlip (all classes flipping to a single background class using ResNet-32 with 10 clean images per class). Comparisons include standard training (Baseline), bootstrapping (Reed), noise adaptation layers (S-Model), MentorNet, random Gaussian reweighting (Random), and clean validation fine-tuning (+FT) with early stopping (+ES).

    Model CIFAR-10 CIFAR-100
    UniformFlip (40% Noise, WRN-28-10)
    Baseline 67.97±0.6267.97 \pm 0.62 50.66±0.2450.66 \pm 0.24
    Reed-Hard 69.66±1.2169.66 \pm 1.21 51.34±0.1751.34 \pm 0.17
    S-Model 70.64±3.0970.64 \pm 3.09 49.10±0.5849.10 \pm 0.58
    MentorNet 76.676.6 56.956.9
    Random 86.06±0.3286.06 \pm 0.32 58.01±0.3758.01 \pm 0.37
    Using 1,000 Clean Images
    Clean Only 46.64±3.9046.64 \pm 3.90 9.94±0.829.94 \pm 0.82
    Baseline + FT 78.66±0.4478.66 \pm 0.44 54.52±0.4054.52 \pm 0.40
    MentorNet + FT 7878 5959
    Random + FT 86.55±0.2486.55 \pm 0.24 58.54±0.5258.54 \pm 0.52
    Ours 86.92±0.19\mathbf{86.92 \pm 0.19} 61.34±2.06\mathbf{61.34 \pm 2.06}
    BackgroundFlip (40% Noise, ResNet-32)
    Baseline 59.54±2.1659.54 \pm 2.16 37.82±0.6937.82 \pm 0.69
    Baseline + ES 64.96±1.1964.96 \pm 1.19 39.08±0.6539.08 \pm 0.65
    Random 69.51±1.3669.51 \pm 1.36 36.56±0.4436.56 \pm 0.44
    Weighted 79.17±1.3679.17 \pm 1.36 36.56±0.4436.56 \pm 0.44
    Reed Soft + ES 63.47±1.0563.47 \pm 1.05 38.44±0.9038.44 \pm 0.90
    Reed Hard + ES 65.22±1.0665.22 \pm 1.06 39.03±0.5539.03 \pm 0.55
    S-Model 58.60±2.3358.60 \pm 2.33 37.02±0.3437.02 \pm 0.34
    S-Model + Conf 68.93±1.0968.93 \pm 1.09 46.72±1.8746.72 \pm 1.87
    S-Model + Conf + ES 79.24±0.5679.24 \pm 0.56 54.50±2.5154.50 \pm 2.51
    Using 10 Clean Images Per Class
    Clean Only 15.90±3.3215.90 \pm 3.32 8.06±0.768.06 \pm 0.76
    Baseline + FT 82.82±0.9382.82 \pm 0.93 54.23±1.7554.23 \pm 1.75
    Baseline + ES + FT 85.19±0.4685.19 \pm 0.46 55.22±1.4055.22 \pm 1.40
    Weighted + FT 85.98±0.4785.98 \pm 0.47 53.99±1.6253.99 \pm 1.62
    S-Model + Conf + FT 81.90±0.8581.90 \pm 0.85 53.11±1.3353.11 \pm 1.33
    S-Model + Conf + ES + FT 85.86±0.6385.86 \pm 0.63 55.75±1.2655.75 \pm 1.26
    Ours 86.73±0.48\mathbf{86.73 \pm 0.48} 59.30±0.60\mathbf{59.30 \pm 0.60}

    The meta-reweighting algorithm outperforms all baseline and fine-tuned methods on both datasets, improving over previous state-of-the-art results on CIFAR-100 by over 3 percentage points in both noise settings without requiring early stopping or noise transition oracle knowledge.

  6. Knowl 6 — Robustness to Class Imbalance on MNIST

    empirical result

    The algorithm was evaluated on an imbalanced binary classification task between MNIST digits 4 and 9 (5,000 total images, size 28×2828 \times 28) using a LeNet trained with SGD (learning rate 1e-3, batch size 100, 8,000 steps). A balanced validation set of only 10 images was reserved from the training data. Proportion of majority class (digit 9) was varied from 90% to 99.5% (a 200:1 imbalance ratio).

    Compared against class-proportion weighting (inverse frequency), class-balanced resampling, hard negative mining, and random weighting, the proposed meta-reweighting method maintained a test error rate below 3% across all ratios, exhibiting only ~2% error degradation at the 200:1 ratio. In contrast, standard baseline error rates grew from ~5% at 90% majority proportion to over 15–20% at 99.5%.

  7. Knowl 7 — Effect of Validation Set Size on Meta-Reweighting vs Fine-Tuning

    empirical result

    On CIFAR-10 under 40% background flip noise using ResNet-32, varying the number of clean validation images from 15 to 1,000 demonstrated distinct behaviors between meta-reweighting and post-hoc fine-tuning:

    1. Meta-reweighting with only 15 clean validation images total (1.5 images per class) achieved over 84% test accuracy, representing only a ~2% decrease compared to using 1,000 clean validation images.
    2. Meta-reweighting performance saturated rapidly around 100 clean validation images (10 images per class).
    3. In contrast, baselines fine-tuned on the clean validation set (Baseline + FT, S-Model + Conf + FT) suffered severe performance drops when provided with 15 to 50 clean images (dropping below 75% accuracy) and required roughly 1,000 clean images to approach the meta-reweighting performance.

    This confirms that the clean validation set in online meta-reweighting operates primarily as a gradient-direction regularizer rather than as a primary data source for parameter fitting.

  8. Knowl 8 — Computational Complexity and Time Overhead of Meta-Reweighting

    model/method

    Each iteration of the automatic reweighting method requires four computational stages:

    1. A forward pass and backward pass on the training mini-batch to compute gradient directions under dummy perturbation weights ϵ=0\epsilon = 0.
    2. A forward pass and backward pass on the clean validation mini-batch at the one-step lookahead parameters θ^t\hat{\theta}_t.
    3. A backward-on-backward automatic differentiation pass to evaluate ∇ϵlg\nabla_\epsilon l_g, computing the gradient of validation loss with respect to training sample weights.
    4. A final backward pass on the training mini-batch using the rectified, normalized weights ww.

    Because a backward-on-backward pass has time complexity comparable to a standard forward pass, the total training time per step is approximately 3×3\times that of standard stochastic gradient descent training.

  9. Knowl 9 — Robustness to Increasing Label Noise Levels

    empirical result

    On CIFAR-10 and CIFAR-100 with ResNet-32 under background flip noise ratios ranging from 0% to 50%:

    1. As label noise increases from 0% to 50%, the test accuracy of the proposed meta-reweighting method decreases by only approximately 6% on CIFAR-10.
    2. Over the same range, standard baseline model accuracy drops by more than 40% due to overfitting label noise.
    3. At 0% label noise, meta-reweighting performs slightly below the unweighted baseline due to subsample variance from optimizing on the small validation subset.

Coverage note — None was omitted.

References

  1. 1.Abadi, Martın, Barham, Paul, Chen, Jianmin, Chen, Zhifeng, Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Irving, Geoffrey, Isard, Michael, Kudlur, Manjunath, Levenberg, Josh, Monga, Rajat, Moore, Sherry, Murray, Derek Gordon, Steiner, Benoit, Tucker, Paul A., Vasudevan, Vijay, Warden, Pete, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang. Tensorflow: A system for largescale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation, OSDI, 2016.
  2. 2.Andrychowicz, Marcin, Denil, Misha, Colmenarejo, Sergio Gomez, Hoffman, Matthew W., Pfau, David, Schaul, Tom, and de Freitas, Nando. Learning to learn by gradient descent by gradient descent. In Advances in Neural Information Processing Systems, NIPS, 2016.
  3. 3.Angluin, Dana and Laird, Philip. Learning from noisy examples. Machine Learning, 2(4):343–370, Apr 1988. ISSN 1573-0565.
  4. 4.Azadi, Samaneh, Feng, Jiashi, Jegelka, Stefanie, and Darrell, Trevor. Auxiliary image regularization for deep cnns with noisy labels. In Proceedings of the 4th International Conference on Learning Representation, ICLR, 2016.
  5. 5.Bengio, Yoshua, Louradour, Jerome, Collobert, Ronan, and Weston, Jason. Curriculum learning. In Proceedings of the 26th Annual International Conference on Machine Learning, ICML, 2009.
  6. 6.Chang, Haw-Shiuan, Learned-Miller, Erik G., and McCallum, Andrew. Active bias: Training more accurate neural networks by emphasizing high variance samples. In Advances in Neural Information Processing Systems, NIPS, 2017.
  7. 7.Chawla, Nitesh V., Bowyer, Kevin W., Hall, Lawrence O., and Kegelmeyer, W. Philip. SMOTE: synthetic minority over-sampling technique. J. Artif. Intell. Res., 16:321–357, 2002.
  8. 8.Chen, Xinlei and Gupta, Abhinav. Webly supervised learning of convolutional networks. In Proceedings of the 2015 IEEE International Conference on Computer Vision, ICCV, 2015.
  9. 9.Cordts, Marius, Omran, Mohamed, Ramos, Sebastian, Rehfeld, Timo, Enzweiler, Markus, Benenson, Rodrigo, Franke, Uwe, Roth, Stefan, and Schiele, Bernt. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2016.
  10. 10.Dong, Qi, Gong, Shaogang, and Zhu, Xiatian. Class rectification hard mining for imbalanced deep learning. In Proceedings of the IEEE International Conference on Computer Vision, ICCV, 2017.
  11. 11.Finn, Chelsea, Abbeel, Pieter, and Levine, Sergey. Modelagnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning, ICML, 2017.
  12. 12.Freund, Yoav and Schapire, Robert E. A decision-theoretic generalization of on-line learning and an application to boosting. J. Comput. Syst. Sci., 55(1):119–139, 1997.
  13. 13.Goldberger, Jacob and Ben-Reuven, Ehud. Training deep neural-networks using a noise adaptation layer. In Proceedings of the 5th International Conference on Learning Representation, ICLR, 2017.
  14. 14.He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2016.
  15. 15.Hendrycks, Dan, Mazeika, Mantas, Wilson, Duncan, and Gimpel, Kevin. Using trusted data to train deep networks on labels corrupted by severe noise. CoRR, abs/1802.05300, 2018.
  16. 16.Huang, Chen, Li, Yining, Loy, Chen Change, and Tang, Xiaoou. Learning deep representation for imbalanced classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2016.
  17. 17.Jiang, Lu, Meng, Deyu, Zhao, Qian, Shan, Shiguang, and Hauptmann, Alexander G. Self-paced curriculum learning. In Proceedings of the 29th AAAI Conference on Artificial Intelligence, 2015.
  18. 18.Jiang, Lu, Zhou, Zhengyuan, Leung, Thomas, Li, Li-Jia, and Fei-Fei, Li. Mentornet: Regularizing very deep neural networks on corrupted labels. CoRR, abs/1712.05055, 2017.
  19. 19.Kahn, Herman and Marshall, Andy W. Methods of reducing sample size in monte carlo computations. Journal of the Operations Research Society of America, 1(5):263–278, 1953.
  20. 20.Khan, Salman Hameed, Bennamoun, Mohammed, Sohel, Ferdous Ahmed, and Togneri, Roberto. Cost sensitive learning of deep feature representations from imbalanced data. CoRR, abs/1508.03422, 2015.
  21. 21.Koh, Pang Wei and Liang, Percy. Understanding black-box predictions via influence functions. In Proceedings of the 34th International Conference on Machine Learning, ICML, 2017.
  22. 22.Kumar, M. Pawan, Packer, Benjamin, and Koller, Daphne. Self-paced learning for latent variable models. In Advances in Neural Information Processing Systems, NIPS, 2010.
  23. 23.Lake, Brenden M., Ullman, Tomer D., Tenenbaum, Joshua B., and Gershman, Samuel J. Building machines that learn and think like people. Behav Brain Sci, 40: e253, Jan 2017.
  24. 24.Li, Yuncheng, Yang, Jianchao, Song, Yale, Cao, Liangliang, Luo, Jiebo, and Li, Li-Jia. Learning from noisy labels with distillation. In Proceedings of the IEEE International Conference on Computer Vision, ICCV, 2017.
  25. 25.Lin, Tsung-Yi, Goyal, Priya, Girshick, Ross B., He, Kaiming, and Dollar, Piotr. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, ICCV, 2017.
  26. 26.Lorraine, Jonathan and Duvenaud, David. Stochastic hyperparameter optimization through hypernetworks. CoRR, abs/1802.09419, 2018.
  27. 27.Ma, Fan, Meng, Deyu, Xie, Qi, Li, Zina, and Dong, Xuanyi. Self-paced co-training. In Proceedings of the 34th International Conference on Machine Learning, ICML, 2017.
  28. 28.Malisiewicz, Tomasz, Gupta, Abhinav, and Efros, Alexei A. Ensemble of exemplar-svms for object detection and beyond. In Proceedings of the IEEE International Conference on Computer Vision, ICCV, 2011.
  29. 29.Munoz-Gonzalez, Luis, Biggio, Battista, Demontis, Ambra, Paudice, Andrea, Wongrassamee, Vasin, Lupu, Emil C., and Roli, Fabio. Towards poisoning of deep learning algorithms with back-gradient optimization. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, AISec@CCS, 2017.
  30. 30.Natarajan, Nagarajan, Dhillon, Inderjit S., Ravikumar, Pradeep, and Tewari, Ambuj. Learning with noisy labels. In Advances in Neural Information Processing Systems, NIPS, 2013.
  31. 31.Ravi, Sachin and Larochelle, Hugo. Optimization as a model for few-shot learning. In Proceedings of the 5th International Conference on Learning Representations, ICLR, 2017.
  32. 32.Reddi, Sashank J., Hefny, Ahmed, Sra, Suvrit, Poczos, Barnabas, and Smola, Alexander J. Stochastic variance reduction for nonconvex optimization. In Proceedings of the 33rd International Conference on Machine Learning, ICML, 2016.
  33. 33.Reed, Scott E., Lee, Honglak, Anguelov, Dragomir, Szegedy, Christian, Erhan, Dumitru, and Rabinovich, Andrew. Training deep neural networks on noisy labels with bootstrapping. CoRR, abs/1412.6596, 2014.
  34. 34.Ren, Mengye, Triantafillou, Eleni, Ravi, Sachin, Snell, Jake, Swersky, Kevin, Tenenbaum, Joshua B., Larochelle, Hugo, and Zemel, Richard S. Meta learning for few-shot semi-supervised classification. In Proceedings of the 6th International Conference on Learning Representations, ICLR, 2018.
  35. 35.Russakovsky, Olga, Deng, Jia, Su, Hao, Krause, Jonathan, Satheesh, Sanjeev, Ma, Sean, Huang, Zhiheng, Karpathy, Andrej, Khosla, Aditya, Bernstein, Michael, Berg, Alexander C., and Fei-Fei, Li. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, IJCV, 115(3):211–252, 2015.
  36. 36.Sukhbaatar, Sainbayar and Fergus, Rob. Learning from noisy labels with deep neural networks. CoRR, abs/1406.2080, 2014.
  37. 37.Thrun, Sebastian and Pratt, Lorien. Learning to Learn. Springer, 1998.
  38. 38.Ting, Kai Ming. A comparative study of cost-sensitive boosting algorithms. In Proceedings of the 17th International Conference on Machine Learning, ICML, 2000.
  39. 39.Vahdat, Arash. Toward robustness against label noise in training deep discriminative neural networks. In Advances in Neural Information Processing Systems, NIPS, 2017.
  40. 40.Wang, Yixin, Kucukelbir, Alp, and Blei, David M. Robust probabilistic modeling with bayesian data reweighting. In Proceedings of the 34th International Conference on Machine Learning, ICML, 2017.
  41. 41.Wu, Yuhuai, Ren, Mengye, Liao, Renjie, and Grosse, Roger B. Understanding short-horizon bias in stochastic meta-optimization. In Proceedings of the 6th International Conference on Learning Representations, ICLR, 2018.
  42. 42.Xiao, Tong, Xia, Tian, Yang, Yi, Huang, Chang, and Wang, Xiaogang. Learning from massive noisy labeled data for image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2015.
  43. 43.Zagoruyko, Sergey and Komodakis, Nikos. Wide residual networks. In Proceedings of the British Machine Vision Conference, BMVC, 2016.
  44. 44.Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol. Understanding deep learning requires rethinking generalization. In Proceedings of the 5th International Conference on Learning Representations, ICLR, 2017.

Citation

MLA
Ren, M., et al. “Learning to Reweight Examples for Robust Deep Learning”. arXiv, 2018, http://arxiv.org/abs/1803.09050v3.
APA
Ren, M., Zeng, W., Yang, B., & Urtasun, R. (2018). Learning to Reweight Examples for Robust Deep Learning. arXiv. http://arxiv.org/abs/1803.09050v3
Chicago
Ren, M., W. Zeng, B. Yang, and R. Urtasun. 2018. “Learning to Reweight Examples for Robust Deep Learning”. arXiv. http://arxiv.org/abs/1803.09050v3.
Harvard
Ren, M. et al. (2018) “Learning to Reweight Examples for Robust Deep Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1803.09050v3.
Vancouver
1. Ren M, Zeng W, Yang B, Urtasun R (2018) Learning to Reweight Examples for Robust Deep Learning. arXiv

BibTeX

@article{ren2018learning,
  title = {Learning to Reweight Examples for Robust Deep Learning},
  author = {Ren, Mengye and Zeng, Wenyuan and Yang, Bin and Urtasun, Raquel},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1803.09050v3},
  eprint = {1803.09050}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/