On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning

Lorenzo BonicelliMatteo BoschiniAngelo PorrelloConcetto SpampinatoSimone Calderara

article2022NeurIPS62 citations

Proposes a surrogate objective that constrains layer-wise Lipschitz constants on replay data to prevent unstable decision boundaries and buffer overfitting across standard continual learning methods.

Listen

Continual learning enables artificial intelligence models to learn continuously from new data streams without forgetting previously mastered tasks. A widely used strategy is memory replay, which preserves a small buffer of past examples and periodically retrains the model on them. However, continually optimizing a model on a restricted pool of stored data causes severe memory overfitting. This problem distorts decision boundaries around past classes, making the model highly sensitive to minor input perturbations and hurting overall classification performance.

The article introduces and evaluates Lipschitz-Driven Rehearsal (LiDER), an optimization technique designed to prevent memory overfitting in continual learning. The main objective is to demonstrate that constraining the network’s mathematical smoothness—specifically by bounding its layer-wise Lipschitz constants with respect to replayed examples—stabilizes decision boundaries and improves generalization across diverse benchmark tasks.

The evaluation evaluated LiDER across several standard image classification benchmarks, including Split CIFAR-100, Split miniImageNet, and Split CUB-200. The testing suite covered standard neural network backbones (ResNet18, ResNet50, and EfficientNet-B2) under both randomly initialized and pre-trained settings. The researchers integrated LiDER into five existing replay methods, such as Dark Experience Replay (DER++) and Experience Replay with Asymmetric Cross-Entropy (ER-ACE), and compared its performance against parameter-level regularization baselines, as well as under buffer label poisoning and buffer detection tests.

The findings demonstrate consistent performance gains across all benchmarks and architectures. Applying LiDER delivered average accuracy improvements of approximately 2.32% on Split CIFAR-100, 2.08% on Split miniImageNet, and 4.36% on Split CUB-200. In specific configurations, gains reached over 8% in classification accuracy (such as DER++ on Split CUB-200 with smaller buffers). Second, LiDER outperformed alternative parameter-space regularization methods, especially on complex tasks and pre-trained models. Third, the method substantially improved robustness against corrupted data, retaining higher accuracy when memory buffer labels were poisoned. Fourth, analysis confirmed that applying smoothness constraints specifically to rehearsed past examples is far more effective than applying them to incoming new data streams.

These results establish that managing input-space smoothness directly counteracts the deterioration of learned decision surfaces in continual learning. For engineering and machine learning teams, LiDER offers a plug-and-play enhancement that boosts model retention and robustness with minimal computational overhead and without increasing memory buffer sizes. This reduces operational risk and hardware overhead in real-time or streaming machine learning systems that must learn continuously from evolving data.

Organizations deploying replay-based continual learning models should integrate smoothness regularization into their training pipelines, learning the target layer bounds dynamically rather than setting them as static constants. Before deploying in production environments, teams should run pilot evaluations on domain-specific data to tune regularizer weighting.

The evaluation has two main boundaries. The approach relies on an upper-bound mathematical approximation of the Lipschitz constant rather than an exact computation, and it cannot directly accommodate certain modern layer architectures like cross-attention mechanisms. Furthermore, because replay methods rely on storing raw past data, the approach may require additional safeguards in privacy-sensitive applications. Within standard computer vision and classification architectures, confidence in the reported performance gains remains high.

arXiv: 2210.06443
Cover for On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning

Abstract

Rehearsal approaches enjoy immense popularity with Continual Learning (CL) practitioners. These methods collect samples from previously encountered data distributions in a small memory buffer; subsequently, they repeatedly optimize on the latter to prevent catastrophic forgetting. This work draws attention to a hidden pitfall of this widespread practice: repeated optimization on a small pool of data inevitably leads to tight and unstable decision boundaries, which are a major hindrance to generalization. To address this issue, we propose Lipschitz-DrivEn Rehearsal (LiDER), a surrogate objective that induces smoothness in the backbone network by constraining its layer-wise Lipschitz constants w.r.t. replay examples. By means of extensive experiments, we show that applying LiDER delivers a stable performance gain to several state-of-the-art rehearsal CL methods across multiple datasets, both in the presence and absence of pre-training. Through additional ablative experiments, we highlight peculiar aspects of buffer overfitting in CL and better characterize the effect produced by LiDER. Code is available at https://github.com/aimagelab/LiDER.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Lipschitz-Driven Rehearsal
  • 3.2 Relation with Other Regularization Approaches
  • 4 Experiments
  • 4.1 Benchmarks
  • 4.2 Comparison with Rehearsal Approaches
  • 4.3 Comparison with Regularization Approaches
  • 5 Model Analysis
  • 5.1 On the Memory Buffer
  • 5.2 Applying LiDER on Current Task's Examples
  • 5.3 Generalization Measures
  • 5.4 Optimization with a Fixed Target
  • 6 Conclusions
  • Limitations & Societal Impact
  • References

Knowls

  1. Knowl 1 — Buffer overfitting and Lipschitz-Driven Rehearsal

    model/method

    Rehearsal-based continual learning stores a small memory buffer MM of examples from previously observed tasks and repeatedly trains on those examples together with data from the current task. The paper identifies a characteristic failure mode: repeated optimization makes the decision surface unusually tight around points in MM, while decision boundaries around non-rehearsed examples from old tasks progressively erode. This produces poor local robustness and generalization despite good performance on the stored examples.

    Lipschitz-DrivEn Rehearsal (LiDER) addresses this failure by adding an input-space smoothness regularizer to an existing rehearsal method. LiDER estimates layer-wise Lipschitz quantities using replay examples and encourages controlled, small values for those quantities. The regularizer requires no memory beyond the buffer already maintained by the base continual-learning method and can therefore be attached to methods such as ER-ACE, DER++, X-DER, iCaRL, and GDumb.

  2. Knowl 2 — Layer-wise Lipschitz upper bound for the backbone

    theoretical result

    For a function f:Rn→Rmf:\mathbb{R}^{n}\to\mathbb{R}^{m}, the Lipschitz constant is the smallest L∈R+L\in\mathbb{R}_{+} satisfying

    ∥f(x)−f(y)∥2≤L∥x−y∥2,∀x,y∈Rn.\|f(x)-f(y)\|_{2}\le L\|x-y\|_{2},\qquad \forall x,y\in\mathbb{R}^{n}.

    For the feed-forward backbone f=(HK∘σK∘⋯∘H1)f=(H_{K}\circ\sigma^{K}\circ\cdots\circ H_{1}), where Hk(h)=WkThH_{k}(h)=W_{k}^{\mathsf T}h is the kkth linear layer and each activation σk\sigma^{k} is ReLU, the paper uses the upper bound

    ∥f∥L≤∏k=1K∥Hk∥L=∏k=1K∥Wk∥SN,\|f\|_{L}\le\prod_{k=1}^{K}\|H_{k}\|_{L}=\prod_{k=1}^{K}\|W_{k}\|_{\mathrm{SN}},

    where ∥Wk∥SN\|W_{k}\|_{\mathrm{SN}} is the spectral norm, or largest singular value, of the weight matrix WkW_{k}. The bound follows from ReLU being globally 11-Lipschitz and from the Lipschitz constant of a composition being no larger than the product of the component constants. The exact Lipschitz constant of a deep network is computationally intractable in general, so LiDER regularizes this computable upper-bound structure instead.

  3. Knowl 3 — LiDER surrogate objective

    equation

    For each layer kk of a KK-layer backbone, let λ1k≥0\lambda_{1}^{k}\ge 0 denote the largest-eigenvalue estimate of the layer's transmitting matrix, used as a proxy for the squared spectral norm, and let ck>0c_{k}>0 be a learnable target for that layer. LiDER defines a target-matching penalty

    Lc-Lip=1K∑k=0K∣λ1k−ck∣\mathcal{L}_{c\text{-Lip}}=\frac{1}{K}\sum_{k=0}^{K}\left|\lambda_{1}^{k}-c_{k}\right|

    and a capacity-reduction penalty

    L0-Lip=1K∑k=0K∣λ1k∣.\mathcal{L}_{0\text{-Lip}}=\frac{1}{K}\sum_{k=0}^{K}\left|\lambda_{1}^{k}\right|.

    The complete surrogate objective added to the base rehearsal loss is

    LLiDER=αLc-Lip+βL0-Lip,\mathcal{L}_{\mathrm{LiDER}}=\alpha\mathcal{L}_{c\text{-Lip}}+\beta\mathcal{L}_{0\text{-Lip}},

    where α,β≥0\alpha,\beta\ge 0 are weighting coefficients. The first term makes each layer approach its learned smoothness target, while the second term discourages unnecessarily large upper bounds and avoids a solution in which the targets simply absorb excessive layer sensitivity. During estimation of λ1k\lambda_{1}^{k}, activation maps from current-task examples are discarded; the regularizer is computed from replay-buffer examples because current-task data are plentiful, whereas old-task knowledge is represented by only a few repeatedly reused samples. The targets ckc_k are optimized by gradient descent rather than fixed as manually chosen budgets.

  4. Knowl 4 — Differentiable estimation of layer sensitivity

    model/method

    LiDER avoids singular-value decomposition, which is inconvenient for convolutional and residual structures. For a minibatch of BB replay examples, it forms the paper's transmitting matrix for each layer from the L2L_{2}-normalized feature maps entering and leaving that layer. If Fk−1F^{k-1} and FkF^{k} denote the corresponding normalized feature-map matrices, the largest eigenvalue λ1k\lambda_{1}^{k} of this transmitting matrix is used as a differentiable proxy for the layer's squared spectral norm.

    The largest eigenvalue is estimated with power iteration rather than an explicit eigendecomposition. The resulting estimate can be backpropagated through during ordinary training, adds only modest computation, and applies to the structured layers used by the evaluated convolutional backbones. The estimates are computed on replay activations rather than current-task activations, making the smoothness constraint specifically target the regions most exposed to buffer overfitting.

  5. Knowl 5 — Continual-learning evaluation protocol

    experimental setup

    The experiments use class-incremental continual learning (CIL), in which the classifier must predict among all classes encountered so far without receiving a task identifier at test time. Performance is measured by Final Average Accuracy (FAA), the classification accuracy averaged over all tasks after the complete task stream.

    The evaluated streams are:

    • Split CIFAR-100: 100 classes divided into 10 tasks of 32×3232\times32 images, using randomly initialized ResNet-18 or ResNet-18 pretrained on Tiny ImageNet.
    • Split miniImageNet: 100 classes divided into 20 tasks of resized 84×8484\times84 images, using an unpretrained EfficientNet-B2.
    • Split CUB-200: 200 bird classes divided into ten 20-class tasks of 224×224224\times224 images, using an ImageNet-pretrained ResNet-50.

    LiDER is attached to five rehearsal baselines: ER-ACE, DER++, X-DER with RPC, iCaRL, and GDumb. Reservoir sampling is used for buffer updates unless the method specifies otherwise; iCaRL uses herding. Buffer sizes vary by benchmark, including 500 and 2,000 for randomly initialized Split CIFAR-100, 500 and 2,000 for its Tiny ImageNet-pretrained version, 2,000 and 5,000 for Split miniImageNet, and 400 and 1,000 for Split CUB-200. A jointly trained model supplies a non-continual upper bound, while sequential fine-tuning without forgetting countermeasures supplies a lower reference.

  6. Knowl 6 — LiDER improves rehearsal methods across benchmarks

    data/table

    The page-7 benchmark matrix compares Final Average Accuracy with and without LiDER. Each pair of values is baseline / LiDER, and the eight benchmark columns are ordered as follows: Split CIFAR-100 without pretraining with buffer sizes 500 and 2,000; Split CIFAR-100 with Tiny ImageNet pretraining with buffer sizes 500 and 2,000; Split miniImageNet without pretraining with buffer sizes 2,000 and 5,000; and Split CUB-200 with ImageNet pretraining with buffer sizes 400 and 1,000. The results show that LiDER improves every listed base method in every evaluated configuration.

    Could not parse LaTeX table

    The mean accuracy gains attributed to LiDER are 2.32 percentage points on Split CIFAR-100, 2.08 points on Split miniImageNet, and 4.36 points on Split CUB-200. The gain is especially large for settings that already use larger buffers or pretrained representations: GDumb on Split CIFAR-100 improves by about 0.98 points with a 500-example buffer, 6.46 points with a 2,000-example buffer, and 2.99 points with Tiny ImageNet pretraining.

  7. Knowl 7 — LiDER outperforms alternative regularizers

    data/table

    The page-7 regularization comparison combines ER-ACE or DER++ with Stable SGD (sSGD), online Elastic Weight Consolidation (oEwC), online Laplace (oLAP), or LiDER. The six columns are Split CIFAR-100 with buffer sizes 500 and 2,000, Split miniImageNet with buffer sizes 2,000 and 5,000, and Split CUB-200 with buffer sizes 400 and 1,000. LiDER is almost always the strongest regularizer. sSGD helps on CIFAR-100 but degrades strongly on the longer miniImageNet stream and the pretrained CUB-200 stream; oEwC and oLAP are more useful when pretraining is present.

    Could not parse LaTeX table
  8. Knowl 8 — Buffer-specific regularization improves robustness to poisoned and identifiable memories

    empirical result

    The paper evaluates whether LiDER specifically mitigates overfitting to replay-buffer items. In the Split CIFAR-100 experiment with DER++, each example entering the buffer is assigned a randomly selected current-task label with probability pp. As pp increases, accuracy falls for both methods, but LiDER remains more accurate:

    • p=0%p=0\%: DER++ 37.1437.14, DER++ + LiDER 39.2539.25;
    • p=0.01%p=0.01\%: 36.1336.13 versus 38.0838.08;
    • p=0.1%p=0.1\%: 31.3531.35 versus 35.5335.53;
    • p=0.25%p=0.25\%: 28.7428.74 versus 30.7830.78.

    The page-8 Buffer Guessing Game tests whether a trained model makes it possible to infer which examples from the first task are stored in its buffer. For ER-ACE, the ROC-AUC values at buffer sizes ∣M∣=500,1000,2000|M|=500,1000,2000 are respectively 0.67,0.58,0.550.67,0.58,0.55; with LiDER they fall to 0.56,0.50,0.540.56,0.50,0.54. Thus, replay examples become less distinguishable from non-replayed examples under LiDER, especially when the buffer is small. Together with the poisoning results, this supports the paper's interpretation that LiDER loosens the unusually specialized decision boundaries formed around replay items.

  9. Knowl 9 — Ablations support replay targeting, learned targets, and smoother decision surfaces

    empirical result

    LiDER is most effective when its Lipschitz penalty is computed on replay-buffer examples rather than on current-task examples. In the eight benchmark configurations used for the target comparison, buffer-targeted LiDER gives ER-ACE accuracies 38.43,50.32,48.97,57.39,24.13,30.00,50.89,60.9238.43,50.32,48.97,57.39,24.13,30.00,50.89,60.92, whereas current-task targeting gives 37.54,50.37,48.94,57.07,23.35,29.25,48.44,59.6037.54,50.37,48.94,57.07,23.35,29.25,48.44,59.60. For DER++, buffer targeting gives 39.25,53.27,45.37,60.88,28.33,35.04,57.90,67.9739.25,53.27,45.37,60.88,28.33,35.04,57.90,67.97, compared with 34.78,49.76,44.48,59.39,24.84,31.05,56.96,66.6334.78,49.76,44.48,59.39,24.84,31.05,56.96,66.63 for current-task targeting. The configurations are ordered as Split CIFAR-100 buffers 500 and 2,000, Tiny ImageNet-pretrained CIFAR-100 buffers 500 and 2,000, Split miniImageNet buffers 2,000 and 5,000, and Split CUB-200 buffers 400 and 1,000.

    Learning the layer targets also outperforms fixing them. Without pretraining, fixed-target versus learned-target FAA is 36.4236.42 versus 39.2539.25 for DER++ with buffer 500, 34.9934.99 versus 38.4338.43 for ER-ACE with buffer 500, 51.5251.52 versus 53.2753.27 for DER++ with buffer 2,000, and 46.7046.70 versus 48.9748.97 for ER-ACE with buffer 2,000. With pretraining, the corresponding pairs are 43.16/45.3743.16/45.37, 45.21/48.9745.21/48.97, 59.53/60.6859.53/60.68, and 54.82/57.3954.82/57.39.

    The visual decision-surface analysis on page 9 shows that ER-ACE's perturbation-tolerant region around replayed first-task examples shrinks substantially as later tasks are learned, whereas ER-ACE with LiDER exhibits little deterioration. The page-10 loss-landscape analysis further reports that adding LiDER to ER-ACE or DER++ improves resilience to weight perturbations and reduces the summed Hessian eigenvalues, indicating flatter parameter-space minima in addition to the intended input-space smoothing.

  10. Knowl 10 — Limitations of the Lipschitz approximation and deployment constraints

    limitation

    LiDER does not compute the exact network Lipschitz constant. It relies on a coarse, feature-based upper-bound estimate because exact computation is NP-hard; tighter local or global bounds exist but incur greater computational cost, which may be problematic in continual learning. The approximation also assumes that the network layers are Lipschitz continuous and cannot directly handle non-Lipschitz components such as cross-attention.

    The bound ∥σl∥L≤1\|\sigma^{l}\|_{L}\le 1 for each activation is a global assumption over the entire input domain. The paper notes that this can be overly restrictive because actual layer inputs occupy structured subspaces; local Lipschitz bounds could provide tighter regularization signals. Finally, rehearsal stores raw examples in the buffer, so LiDER inherits the privacy limitation of rehearsal methods and is unsuitable for settings in which retaining in-plain data is prohibited.

Coverage note — Supplementary-only optimizer details, variance/error bars, Final Forgetting values, efficiency–accuracy curves, and hyperparameter-sensitivity studies were omitted because they are not load-bearing contributions in the supplied main paper.

References

  1. 1.Davide Abati, Jakub Tomczak, Tijmen Blankevoort, Simone Calderara, Rita Cucchiara, and Babak Ehteshami Bejnordi. Conditional channel gated networks for task-aware continual learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2020.
  2. 2.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision, 2018.
  3. 3.Rahaf Aljundi, Eugene Belilovsky, Tinne Tuytelaars, Laurent Charlin, Massimo Caccia, Min Lin, and Lucas Page-Caccia. Online continual learning with maximal interfered retrieval. In Advances in Neural Information Processing Systems, 2019.
  4. 4.Rahaf Aljundi, Min Lin, Baptiste Goujaud, and Yoshua Bengio. Gradient based sample selection for online continual learning. In Advances in Neural Information Processing Systems, 2019.
  5. 5.Cem Anil, James Lucas, and Roger Grosse. Sorting out lipschitz function approximation. In International Conference on Machine Learning, 2019.
  6. 6.Elahe Arani, Fahad Sarfraz, and Bahram Zonooz. Learning fast, learning slow: A general continual learning method based on complementary learning system. In International Conference on Learning Representations Workshop, 2022.
  7. 7.Jihwan Bang, Heesu Kim, YoungJoon Yoo, Jung-Woo Ha, and Jonghyun Choi. Rainbow memory: Continual learning with a memory of diverse samples. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2021.
  8. 8.Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky. Spectrally-normalized margin bounds for neural networks. In Advances in Neural Information Processing Systems, 2017.
  9. 9.Giovanni Bellitto, Matteo Pennisi, Simone Palazzo, Lorenzo Bonicelli, Matteo Boschini, Simone Calderara, and Concetto Spampinato. Effects of auxiliary knowledge on continual learning. In International Conference on Pattern Recognition, 2022.
  10. 10.Ari S Benjamin, David Rolnick, and Konrad Kording. Measuring and regularizing networks in function space. In International Conference on Learning Representations Workshop, 2019.
  11. 11.David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A Raffel. Mixmatch: A holistic approach to semi-supervised learning. Advances in Neural Information Processing Systems, 2019.
  12. 12.Matteo Boschini, Lorenzo Bonicelli, Pietro Buzzega, Angelo Porrello, and Simone Calderara. Class-incremental continual learning into the extended der-verse. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  13. 13.Matteo Boschini, Lorenzo Bonicelli, Angelo Porrello, Giovanni Bellitto, Matteo Pennisi, Simone Palazzo, Concetto Spampinato, and Simone Calderara. Transfer without forgetting. In Proceedings of the European Conference on Computer Vision, 2022.
  14. 14.Matteo Boschini, Pietro Buzzega, Lorenzo Bonicelli, Angelo Porrello, and Simone Calderara. Continual semi-supervised learning through contrastive interpolation consistency. Pattern Recognition Letters, 2022.
  15. 15.Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. Dark Experience for General Continual Learning: a Strong, Simple Baseline. In Advances in Neural Information Processing Systems, 2020.
  16. 16.Pietro Buzzega, Matteo Boschini, Angelo Porrello, and Simone Calderara. Rethinking Experience Replay: a Bag of Tricks for Continual Learning. In International Conference on Pattern Recognition, 2020.
  17. 17.Lucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars, Joelle Pineau, and Eugene Belilovsky. New Insights on Reducing Abrupt Representation Change in Online Continual Learning. In International Conference on Learning Representations Workshop, 2022.
  18. 18.Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. Co2l: Contrastive continual learning. In IEEE International Conference on Computer Vision, 2021.
  19. 19.Arslan Chaudhry, Puneet K Dokania, Thalaiyasingam Ajanthan, and Philip HS Torr. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In Proceedings of the European Conference on Computer Vision, 2018.
  20. 20.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. Efficient Lifelong Learning with A-GEM. In International Conference on Learning Representations Workshop, 2019.
  21. 21.Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. On tiny episodic memories in continual learning. In International Conference on Machine Learning Workshop, 2019.
  22. 22.Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier. Parseval networks: Improving robustness to adversarial examples. In International Conference on Machine Learning, 2017.
  23. 23.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Greg Slabaugh, and Tinne Tuytelaars. A continual learning survey: Defying forgetting in classification tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  24. 24.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2009.
  25. 25.Sayna Ebrahimi, Suzanne Petryk, Akash Gokul, William Gan, Joseph E Gonzalez, Marcus Rohrbach, and Trevor Darrell. Remembering for the right reasons: Explanations reduce catastrophic forgetting. Applied AI Letters, 2021.
  26. 26.Sebastian Farquhar and Yarin Gal. Towards Robust Evaluations of Continual Learning. In International Conference on Machine Learning Workshop, 2018.
  27. 27.Enrico Fini, Victor G Turrisi da Costa, Xavier Alameda-Pineda, Elisa Ricci, Karteek Alahari, and Julien Mairal. Self-supervised models are continual learners. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 9621–9630, 2022.
  28. 28.Tommaso Furlanello, Zachary C Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. Born again neural networks. In International Conference on Machine Learning, 2018.
  29. 29.Henry Gouk, Eibe Frank, Bernhard Pfahringer, and Michael J Cree. Regularisation of neural networks by enforcing lipschitz continuity. Machine Learning, 2021.
  30. 30.Florian Graf, Sebastian Zeng, Marc Niethammer, and Roland Kwitt. On measuring excess capacity in neural networks. arXiv preprint arXiv:2202.08070, 2022.
  31. 31.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2016.
  32. 32.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2019.
  33. 33.Yujia Huang, Huan Zhang, Yuanyuan Shi, J Zico Kolter, and Anima Anandkumar. Training certifiably robust neural networks with efficient local lipschitz bounds. In Advances in Neural Information Processing Systems, 2021.
  34. 34.Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey. Three factors influencing minima in sgd. In International Conference on Artificial Neural Networks, 2018.
  35. 35.Matt Jordan and Alexandros G Dimakis. Exactly computing the local lipschitz constant of relu networks. Advances in Neural Information Processing Systems, 2020.
  36. 36.Sandesh Kamath, Amit Despande, and KV Subrahmanyam. On adversarial robustness of small vs large batch training. In International Conference on Machine Learning Workshop, 2019.
  37. 37.Zixuan Ke, Bing Liu, and Xingchang Huang. Continual learning of a mixed sequence of similar and dissimilar tasks. In Advances in Neural Information Processing Systems, 2020.
  38. 38.Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? Advances in Neural Information Processing Systems, 2017.
  39. 39.Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. In International Conference on Learning Representations Workshop, 2017.
  40. 40.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised Contrastive Learning. In Advances in Neural Information Processing Systems, 2020.
  41. 41.Hyunjik Kim, George Papamakarios, and Andriy Mnih. The lipschitz constant of self-attention. In International Conference on Machine Learning, 2021.
  42. 42.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 2017.
  43. 43.Alex Krizhevsky et al. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009.
  44. 44.Alexey Kurakin, Ian Goodfellow, Samy Bengio, et al. Adversarial examples in the physical world. In International Conference on Learning Representations Workshop, 2016.
  45. 45.Sungyoon Lee, Jaewook Lee, and Saerom Park. Lipschitz-certifiable training with a tight outer bound. In Advances in Neural Information Processing Systems, 2020.
  46. 46.Klas Leino, Zifan Wang, and Matt Fredrikson. Globally-robust neural networks. In International Conference on Machine Learning, 2021.
  47. 47.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017.
  48. 48.Hsueh-Ti Derek Liu, Francis Williams, Alec Jacobson, Sanja Fidler, and Or Litany. Learning smooth neural functions via lipschitz regularization. arXiv preprint arXiv:2202.08345, 2022.
  49. 49.David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems, 2017.
  50. 50.Divyam Madaan, Jaehong Yoon, Yuanchun Li, Yunxin Liu, and Sung Ju Hwang. Representational continuity for unsupervised continual learning. In International Conference on Learning Representations Workshop, 2022.
  51. 51.Arun Mallya and Svetlana Lazebnik. Packnet: Adding multiple tasks to a single network by iterative pruning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018.
  52. 52.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of learning and motivation, 1989.
  53. 53.Sanket Vaibhav Mehta, Darshan Patil, Sarath Chandar, and Emma Strubell. An empirical investigation of the role of pre-training in lifelong learning. In International Conference on Machine Learning, 2021.
  54. 54.Seyed Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, and Hassan Ghasemzadeh. Understanding the Role of Training Regimes in Continual Learning. In Advances in Neural Information Processing Systems, 2020.
  55. 55.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations Workshop, 2018.
  56. 56.Yuji Nakatsukasa and Nicholas J Higham. Stable and efficient spectral divide and conquer algorithms for the symmetric eigenvalue decomposition and the svd. SIAM Journal on Scientific Computing, 2013.
  57. 57.Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro. Exploring Generalization in Deep Learning. In Advances in Neural Information Processing Systems, 2017.
  58. 58.German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. Continual lifelong learning with neural networks: A review. Neural Networks, 2019.
  59. 59.Federico Pernici, Matteo Bruni, Claudio Baecchi, and Alberto Del Bimbo. Regular polytope networks. IEEE Transactions on Neural Networks and Learning Systems, 2021.
  60. 60.Quang Pham, Chenghao Liu, and Steven Hoi. Dualnet: Continual learning, fast and slow. In Advances in Neural Information Processing Systems, 2021.
  61. 61.Ameya Prabhu, Philip HS Torr, and Puneet K Dokania. GDumb: A simple approach that questions our progress in continual learning. In Proceedings of the European Conference on Computer Vision, 2020.
  62. 62.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2017.
  63. 63.Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro. Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference. In International Conference on Learning Representations Workshop, 2019.
  64. 64.Hippolyt Ritter, Aleksandar Botev, and David Barber. Online structured laplace approximations for overcoming catastrophic forgetting. Advances in Neural Information Processing Systems, 2018.
  65. 65.Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. Progress & compress: A scalable framework for continual learning. In International Conference on Machine Learning, 2018.
  66. 66.Joan Serra, Didac Suris, Marius Miron, and Alexandros Karatzoglou. Overcoming Catastrophic Forgetting with Hard Attention to the Task. In International Conference on Machine Learning, 2018.
  67. 67.Yuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie, and Yan Yan. Lipschitz continuity guided knowledge distillation. In IEEE International Conference on Computer Vision, 2021.
  68. 68.Konstantin Shmelkov, Cordelia Schmid, and Karteek Alahari. Incremental learning of object detectors without catastrophic forgetting. In IEEE International Conference on Computer Vision, 2017.
  69. 69.Yi-Fan Song, Zhang Zhang, Caifeng Shan, and Liang Wang. Constructing stronger and faster baselines for skeleton-based action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  70. 70.Stanford. Tiny ImageNet Challenge (CS231n), 2015. https://www.kaggle.com/c/tiny-imagenet.
  71. 71.David Stutz, Matthias Hein, and Bernt Schiele. Relating adversarially robust generalization to flat minima. In IEEE International Conference on Computer Vision, 2021.
  72. 72.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations Workshop, 2014.
  73. 73.Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. Lipschitz-margin training: Scalable certification of perturbation invariance for deep neural networks. In Advances in Neural Information Processing Systems, 2018.
  74. 74.Gido M van de Ven and Andreas S Tolias. Three continual learning scenarios. In Neural Information Processing Systems Workshops, 2018.
  75. 75.Eli Verwimp, Matthias De Lange, and Tinne Tuytelaars. Rehearsal revealed: The limits and merits of revisiting samples in continual learning. In IEEE International Conference on Computer Vision, 2021.
  76. 76.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in Neural Information Processing Systems, 2016.
  77. 77.Aladin Virmaux and Kevin Scaman. Lipschitz regularity of deep neural networks: analysis and efficient estimation. In Advances in Neural Information Processing Systems, 2018.
  78. 78.Jeffrey S Vitter. Random sampling with a reservoir. ACM Transactions on Mathematical Software, 1985.
  79. 79.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 Dataset. Technical report, California Institute of Technology, 2011.
  80. 80.Lily Weng, Huan Zhang, Hongge Chen, Zhao Song, Cho-Jui Hsieh, Luca Daniel, Duane Boning, and Inderjit Dhillon. Towards fast computation of certified robustness for relu networks. In International Conference on Machine Learning, 2018.
  81. 81.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2019.
  82. 82.Huan Xu and Shie Mannor. Robustness and generalization. Machine Learning, 2012.
  83. 83.Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney. Hessian-based analysis of large batch training and robustness to adversaries. Advances in Neural Information Processing Systems, 2018.
  84. 84.Dong Yin, Mehrdad Farajtabar, Ang Li, Nir Levine, and Alex Mott. Optimization and generalization of regularization-based continual learning: a loss approximation viewpoint. arXiv preprint arXiv:2006.10974, 2020.
  85. 85.Jaehong Yoon, Divyam Madaan, Eunho Yang, and Sung Ju Hwang. Online coreset selection for rehearsal-based continual learning. In International Conference on Learning Representations Workshop, 2022.
  86. 86.Yuichi Yoshida and Takeru Miyato. Spectral norm regularization for improving the generalizability of deep learning. arXiv preprint arXiv:1705.10941, 2017.
  87. 87.Fuxun Yu, Zhuwei Qin, Chenchen Liu, Liang Zhao, Yanzhi Wang, and Xiang Chen. Interpreting and evaluating neural network robustness. In International Joint Conference on Artificial Intelligence, 2019.
  88. 88.Lu Yu, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang, Yongmei Cheng, Shangling Jui, and Joost van de Weijer. Semantic drift compensation for class-incremental learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2020.
  89. 89.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In International Conference on Machine Learning, 2017.
  90. 90.Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning requires rethinking generalization. In International Conference on Learning Representations Workshop, 2017.
  91. 91.Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan. Theoretically principled trade-off between robustness and accuracy. In International Conference on Machine Learning, 2019.

Citation

MLA
Bonicelli, L., et al. “On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 31886–901, https://proceedings.neurips.cc/paper_files/paper/2022/file/cf10920ac985275845247f865b452529-Paper-Conference.pdf.
APA
Bonicelli, L., Boschini, M., Porrello, A., Spampinato, C., & CALDERARA, S. (2022). On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning. Advances in Neural Information Processing Systems, 35, 31886–31901. https://proceedings.neurips.cc/paper_files/paper/2022/file/cf10920ac985275845247f865b452529-Paper-Conference.pdf
Chicago
Bonicelli, L., M. Boschini, A. Porrello, C. Spampinato, and S. CALDERARA. 2022. “On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning”. Advances in Neural Information Processing Systems 35: 31886–901. https://proceedings.neurips.cc/paper_files/paper/2022/file/cf10920ac985275845247f865b452529-Paper-Conference.pdf.
Harvard
Bonicelli, L. et al. (2022) “On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 31886–31901. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/cf10920ac985275845247f865b452529-Paper-Conference.pdf.
Vancouver
1. Bonicelli L, Boschini M, Porrello A, Spampinato C, CALDERARA S (2022) On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 31886–31901

BibTeX

@inproceedings{bonicelli2022the,
  title = {On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning},
  author = {Bonicelli, Lorenzo and Boschini, Matteo and Porrello, Angelo and Spampinato, Concetto and CALDERARA, SIMONE},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {31886-31901},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/cf10920ac985275845247f865b452529-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors