Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Lipschitz regularization

Lipschitz regularization is a machine learning technique that constrains or penalizes the Lipschitz constant of a predictive model or its constituent layers to enforce mathematical smoothness and limit how rapidly the output can change relative to changes in the input. By bounding this rate of variation, typically through explicit penalty terms, gradient regularizers, or weight normalization methods such as spectral norm constraints, the approach prevents neural networks from forming excessively sharp or brittle decision boundaries. This constraint improves model generalization, mitigates overfitting on small or repeatedly sampled datasets, increases robustness against input perturbations, and promotes greater numerical stability during training.

1 item

On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning

On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning

Lorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato, Simone Calderara

OrganizationsPeRCeiVe LabUniversity of CataniaUniversity of Modena and Reggio Emilia

Why you should read this

Proposes a surrogate objective that constrains layer-wise Lipschitz constants on replay data to prevent unstable decision boundaries and buffer overfitting across standard continual learning methods.

Rehearsal approaches enjoy immense popularity with Continual Learning (CL) practitioners. These methods collect samples from previously encountered data distributions in a small memory buffer; subsequently, they repeatedly optimize on the latter to prevent catastrophic forgetting. This work draws attention to a hidden pitfall of this widespread practice: repeated optimization on a small pool of data inevitably leads to tight and unstable decision boundaries, which are a major hindrance to generalization. To address this issue, we propose Lipschitz-DrivEn Rehearsal (LiDER), a surrogate objective that induces smoothness in the backbone network by constraining its layer-wise Lipschitz constants w.r.t. replay examples. By means of extensive experiments, we show that applying LiDER delivers a stable performance gain to several state-of-the-art rehearsal CL methods across multiple datasets, both in the presence and absence of pre-training. Through additional ablative experiments, we highlight peculiar aspects of buffer overfitting in CL and better characterize the effect produced by LiDER. Code is available at https://github.com/aimagelab/LiDER.

Added

2026-09-26