keyword
progressive neural networks
A progressive neural network is a continual learning architecture designed to acquire a sequence of tasks over time while preventing catastrophic forgetting and enabling transfer learning. Instead of updating a single fixed-capacity model, the framework instantiates an independent neural network column for each new task while freezing the parameters of all previously trained columns. Knowledge is transferred across tasks via lateral connections that feed intermediate representations from older, frozen columns into the layers of the active column. Because existing network weights remain untouched, the architecture completely preserves previously learned capabilities while progressively expanding its capacity to leverage prior experience for new objectives.
5 items

Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning
Kai Zhu, Wei Zhai, Yang Cao, Jiebo Luo, Zhengjun Zha
Why you should read this
Develops a self-sustaining representation expansion framework for non-exemplar class-incremental learning that prevents catastrophic forgetting and parameter explosion by dynamically reorganizing network structures and selectively distilling knowledge using new class prototypes.
Non-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a novel self-sustaining representation expansion scheme. Our scheme consists of a structure reorganization strategy that fuses main-branch expansion and side-branch updating to maintain the old features, and a main-branch distillation scheme to transfer the invariant knowledge. Furthermore, a prototype selection mechanism is proposed to enhance the discrimination between the old and new classes by selectively incorporating new samples into the distillation process. Extensive experiments on three benchmarks demonstrate significant incremental performance, outperforming the state-of-the-art methods by a margin of 3%, 3% and 6%, respectively.
Added
2026-09-26

Dark Experience for General Continual Learning: a Strong, Simple Baseline
Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, Simone Calderara
Why you should read this
Proposes Dark Experience Replay, a simple baseline that matches past network logits during replay to outperform complex state-of-the-art methods in realistic continual learning settings without explicit task boundaries.
Continual Learning has inspired a plethora of approaches and evaluation settings; however, the majority of them overlooks the properties of a practical scenario, where the data stream cannot be shaped as a sequence of tasks and offline training is not viable. We work towards General Continual Learning (GCL), where task boundaries blur and the domain and class distributions shift either gradually or suddenly. We address it through mixing rehearsal with knowledge distillation and regularization; our simple baseline, Dark Experience Replay, matches the network's logits sampled throughout the optimization trajectory, thus promoting consistency with its past. By conducting an extensive analysis on both standard benchmarks and a novel GCL evaluation setting (MNIST-360), we show that such a seemingly simple baseline outperforms consolidated approaches and leverages limited resources. We further explore the generalization capabilities of our objective, showing its regularization being beneficial beyond mere performance.
Added
2026-09-25

PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning
Arun Mallya, Svetlana Lazebnik
Why you should read this
Proposes an iterative network pruning and weight-freezing technique that packs multiple sequential tasks into a single deep model to eliminate catastrophic forgetting without relying on proxy loss functions.
This paper presents a method for adding multiple tasks to a single deep neural network while avoiding catastrophic forgetting. Inspired by network pruning techniques, we exploit redundancies in large deep networks to free up parameters that can then be employed to learn new tasks. By performing iterative pruning and network re-training, we are able to sequentially "pack" multiple tasks into a single network while ensuring minimal drop in performance and minimal storage overhead. Unlike prior work that uses proxy losses to maintain accuracy on older tasks, we always optimize for the task at hand. We perform extensive experiments on a variety of network architectures and large-scale datasets, and observe much better robustness against catastrophic forgetting than prior work. In particular, we are able to add three fine-grained classification tasks to a single ImageNet-trained VGG-16 network and achieve accuracies close to those of separately trained networks for each task. Code available at this https URL
Added
2026-09-24

Progressive Neural Networks
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, Raia Hadsell
Why you should read this
Introduces an architecture that eliminates catastrophic forgetting in sequential reinforcement learning tasks by allocating new network columns and leveraging lateral connections to transfer previously learned features.
Learning to solve complex sequences of tasks--while both leveraging transfer and avoiding catastrophic forgetting--remains a key obstacle to achieving human-level intelligence. The progressive networks approach represents a step forward in this direction: they are immune to forgetting and can leverage prior knowledge via lateral connections to previously learned features. We evaluate this architecture extensively on a wide variety of reinforcement learning tasks (Atari and 3D maze games), and show that it outperforms common baselines based on pretraining and finetuning. Using a novel sensitivity measure, we demonstrate that transfer occurs at both low-level sensory and high-level control layers of the learned policy.
Added
2026-09-24

Continual Lifelong Learning with Neural Networks: A Review
German I. Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter
Why you should read this
Synthesizes neural network approaches for preventing catastrophic forgetting in continual learning and connects these algorithmic strategies to biological mechanisms such as structural plasticity and memory replay.
Humans and animals have the ability to continually acquire, fine-tune, and transfer knowledge and skills throughout their lifespan. This ability, referred to as lifelong learning, is mediated by a rich set of neurocognitive mechanisms that together contribute to the development and specialization of our sensorimotor skills as well as to long-term memory consolidation and retrieval. Consequently, lifelong learning capabilities are crucial for autonomous agents interacting in the real world and processing continuous streams of information. However, lifelong learning remains a long-standing challenge for machine learning and neural network models since the continual acquisition of incrementally available information from non-stationary data distributions generally leads to catastrophic forgetting or interference. This limitation represents a major drawback for state-of-the-art deep neural network models that typically learn representations from stationary batches of training data, thus without accounting for situations in which information becomes incrementally available over time. In this review, we critically summarize the main challenges linked to lifelong learning for artificial learning systems and compare existing neural network approaches that alleviate, to different extents, catastrophic forgetting. We discuss well-established and emerging research motivated by lifelong learning factors in biological systems such as structural plasticity, memory replay, curriculum and transfer learning, intrinsic motivation, and multisensory integration.
Added
2026-09-11
