keyword
representation drift
Representation drift refers to the continuous or abrupt shift in the internal feature representations learned by an artificial neural network as its parameters are updated over time. This phenomenon typically occurs in online, continual, and distributed reinforcement learning environments where a model processes non-stationary data distributions, newly introduced task classes, or asynchronous parameter updates. As network weights change to accommodate incoming inputs or experiences, the internal latent space encoding earlier information alters, causing stored historical representations and recurrent states to become stale or misaligned with the current model state. Consequently, representation drift can destabilize training dynamics, cause representational interference between new and old data, and accelerate catastrophic forgetting.
2 items

New Insights on Reducing Abrupt Representation Change in Online Continual Learning
Lucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars, Joelle Pineau, Eugene Belilovsky
Why you should read this
Demonstrates how Experience Replay causes disruptive representation shifts when new classes appear in online continual learning, and resolves this with an asymmetric update rule that forces incoming data to adapt to established representations.
In the online continual learning paradigm, agents must learn from a changing distribution while respecting memory and compute constraints. Experience Replay (ER), where a small subset of past data is stored and replayed alongside new data, has emerged as a simple and effective learning strategy. In this work, we focus on the change in representations of observed data that arises when previously unobserved classes appear in the incoming data stream, and new classes must be distinguished from previous ones. We shed new light on this question by showing that applying ER causes the newly added classes' representations to overlap significantly with the previous classes, leading to highly disruptive parameter updates. Based on this empirical analysis, we propose a new method which mitigates this issue by shielding the learned representations from drastic adaptation to accommodate new classes. We show that using an asymmetric update rule pushes new classes to adapt to the older ones (rather than the reverse), which is more effective especially at task boundaries, where much of the forgetting typically occurs. Empirical results show significant gains over strong baselines on standard continual learning benchmarks.
Added
2026-09-26

Recurrent Experience Replay in Distributed Reinforcement Learning
Steven Kapturowski, Georg Ostrovski, John Quan, Rémi Munos, Will Dabney
Why you should read this
Demonstrates how to effectively combine recurrent neural networks with distributed experience replay to overcome training instabilities caused by stale memory states and outdated parameters, resulting in the first agent to exceed human performance on 52 of 57 Atari games using a single architecture.
Building on the recent successes of distributed training of RL agents, in this paper we investigate the training of RNN-based RL agents from distributed prioritized experience replay. We study the effects of parameter lag resulting in representational drift and recurrent state staleness and empirically derive an improved training strategy. Using a single network architecture and fixed set of hyper-parameters, the resulting agent, Recurrent Replay Distributed DQN, quadruples the previous state of the art on Atari-57, and matches the state of the art on DMLab-30. It is the first agent to exceed human-level performance in 52 of the 57 Atari games.
Added
2026-03-29

