Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Gradient Episodic Memory

Gradient Episodic Memory is a continual machine learning algorithm designed to prevent neural networks from forgetting previously learned information when acquiring knowledge across a sequence of tasks. The method maintains an episodic memory buffer containing a small subset of training examples from each past task. When updating model parameters on a current task, the algorithm checks whether the proposed parameter adjustment would increase the loss on the stored examples of earlier tasks. If an update would degrade past performance, the algorithm projects the gradient of the current task onto a feasible region where its dot product with the gradients of previous tasks is non-negative, using quadratic programming to prevent catastrophic forgetting while facilitating beneficial knowledge transfer.

6 items

Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation

Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation

Dingcheng Li, Zheng Chen, Eunah Cho, Jie Hao, Xiaohu Liu, Fan Xing, Chenlei Guo, Yang Liu

OrganizationsAmazon

Why you should read this

Proposes a framework combining adaptive parameter regularization with embedding-space domain drift estimation to prevent catastrophic forgetting in sequential sequence-to-sequence language generation without storing past task data.

Seq2seq language generation models that are trained offline with multiple domains in a sequential fashion often suffer from catastrophic forgetting. Lifelong learning has been proposed to handle this problem. However, existing work such as experience replay or elastic weighted consolidation requires incremental memory space. In this work, we propose an innovative framework, RMR_DSE that leverages a recall optimization mechanism to selectively memorize important parameters of previous tasks via regularization, and uses a domain drift estimation algorithm to compensate for the drift between different domains in the embedding space. These designs enable the model to be trained on the current task while keeping the memory of previous tasks, and avoid much additional data storage. Furthermore, RMR_DSE can be combined with existing lifelong learning approaches. Our experiments on two seq2seq language generation tasks, paraphrase and dialog response generation, show that RMR_DSE outperforms state-of-the-art models by a considerable margin and greatly reduces forgetting.

Added

2026-10-03

End-to-End Incremental Learning

End-to-End Incremental Learning

Francisco M. Castro, Manuel J. Marín-Jiménez, Nicolás Guil, Cordelia Schmid, Karteek Alahari

OrganizationsCNRSINRIALJKUniversité Grenoble AlpesUniversity of CórdobaUniversity of Málaga

Why you should read this

Proposes a fully end-to-end incremental learning framework that mitigates catastrophic forgetting by combining knowledge distillation with a small exemplar set to jointly optimize feature representations and classifiers on sequentially added image classes.

Although deep learning approaches have stood out in recent years due to their state-of-the-art results, they continue to suffer from catastrophic forgetting, a dramatic decrease in overall performance when training with new classes added incrementally. This is due to current neural network architectures requiring the entire dataset, consisting of all the samples from the old as well as the new classes, to update the model -a requirement that becomes easily unsustainable as the number of classes grows. We address this issue with our approach to learn deep neural networks incrementally, using new data and only a small exemplar set corresponding to samples from the old classes. This is based on a loss composed of a distillation measure to retain the knowledge acquired from the old classes, and a cross-entropy loss to learn the new classes. Our incremental training is achieved while keeping the entire framework end-to-end, i.e., learning the data representation and the classifier jointly, unlike recent methods with no such guarantees. We evaluate our method extensively on the CIFAR-100 and ImageNet (ILSVRC 2012) image classification datasets, and show state-of-the-art performance.

Added

2026-09-25

Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence

Riemannian Walk for Incremental Learning: Understanding Forgetting and Intransigence

Arslan Chaudhry, Puneet K. Dokania, Thalaiyasingam Ajanthan, Philip H. S. Torr

OrganizationsUniversity of Oxford

Why you should read this

Proposes Riemannian Walk (RWalk), a continual learning method based on KL-divergence that optimizes the trade-off between catastrophic forgetting and model intransigence, while establishing formal metrics to quantify both behaviors.

Incremental learning (IL) has received a lot of attention recently, however, the literature lacks a precise problem definition, proper evaluation settings, and metrics tailored specifically for the IL problem. One of the main objectives of this work is to fill these gaps so as to provide a common ground for better understanding of IL. The main challenge for an IL algorithm is to update the classifier whilst preserving existing knowledge. We observe that, in addition to forgetting, a known issue while preserving knowledge, IL also suffers from a problem we call intransigence, inability of a model to update its knowledge. We introduce two metrics to quantify forgetting and intransigence that allow us to understand, analyse, and gain better insights into the behaviour of IL algorithms. We present RWalk, a generalization of EWC++ (our efficient version of EWC [Kirkpatrick2016EWC]) and Path Integral [Zenke2017Continual] with a theoretically grounded KL-divergence based perspective. We provide a thorough analysis of various IL algorithms on MNIST and CIFAR-100 datasets. In these experiments, RWalk obtains superior results in terms of accuracy, and also provides a better trade-off between forgetting and intransigence.

Added

2026-09-25

Dark Experience for General Continual Learning: a Strong, Simple Baseline

Dark Experience for General Continual Learning: a Strong, Simple Baseline

Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, Simone Calderara

OrganizationsUniversity of Modena and Reggio Emilia

Why you should read this

Proposes Dark Experience Replay, a simple baseline that matches past network logits during replay to outperform complex state-of-the-art methods in realistic continual learning settings without explicit task boundaries.

Continual Learning has inspired a plethora of approaches and evaluation settings; however, the majority of them overlooks the properties of a practical scenario, where the data stream cannot be shaped as a sequence of tasks and offline training is not viable. We work towards General Continual Learning (GCL), where task boundaries blur and the domain and class distributions shift either gradually or suddenly. We address it through mixing rehearsal with knowledge distillation and regularization; our simple baseline, Dark Experience Replay, matches the network's logits sampled throughout the optimization trajectory, thus promoting consistency with its past. By conducting an extensive analysis on both standard benchmarks and a novel GCL evaluation setting (MNIST-360), we show that such a seemingly simple baseline outperforms consolidated approaches and leverages limited resources. We further explore the generalization capabilities of our objective, showing its regularization being beneficial beyond mere performance.

Added

2026-09-25

Continual Lifelong Learning with Neural Networks: A Review

Continual Lifelong Learning with Neural Networks: A Review

German I. Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter

OrganizationsHeriot-Watt UniversityRochester Institute of TechnologyUniversität Hamburg

Why you should read this

Synthesizes neural network approaches for preventing catastrophic forgetting in continual learning and connects these algorithmic strategies to biological mechanisms such as structural plasticity and memory replay.

Humans and animals have the ability to continually acquire, fine-tune, and transfer knowledge and skills throughout their lifespan. This ability, referred to as lifelong learning, is mediated by a rich set of neurocognitive mechanisms that together contribute to the development and specialization of our sensorimotor skills as well as to long-term memory consolidation and retrieval. Consequently, lifelong learning capabilities are crucial for autonomous agents interacting in the real world and processing continuous streams of information. However, lifelong learning remains a long-standing challenge for machine learning and neural network models since the continual acquisition of incrementally available information from non-stationary data distributions generally leads to catastrophic forgetting or interference. This limitation represents a major drawback for state-of-the-art deep neural network models that typically learn representations from stationary batches of training data, thus without accounting for situations in which information becomes incrementally available over time. In this review, we critically summarize the main challenges linked to lifelong learning for artificial learning systems and compare existing neural network approaches that alleviate, to different extents, catastrophic forgetting. We discuss well-established and emerging research motivated by lifelong learning factors in biological systems such as structural plasticity, memory replay, curriculum and transfer learning, intrinsic motivation, and multisensory integration.

Added

2026-09-11