keyword
replay memory
Replay memory is a storage buffer used in machine learning, particularly within reinforcement learning and continual learning, to retain previously observed experiences or data samples for reuse during subsequent training cycles. In sequential decision-making tasks, it records past transitions composed of states, actions, rewards, and resulting states, allowing the algorithm to draw randomized batches of past data rather than learning solely from immediate feedback. This mechanism reduces temporal correlations across consecutive observations, stabilizes neural network parameter updates, and improves data efficiency. In continual learning settings, replay memory serves as an experience reservoir that allows the model to periodically rehearse previous tasks, thereby helping mitigate catastrophic forgetting as the system adapts to new data distributions.
3 items

Temporal-Difference Variational Continual Learning
Luckeciano Carvalho Melo, Alessandro Abate, Yarin Gal
Why you should read this
Introduces a temporal-difference-inspired variational continual learning objective that regularizes model updates using multiple past posterior estimates to prevent compounding approximation errors and reduce catastrophic forgetting.
Machine Learning models in real-world applications must continuously learn new tasks to adapt to shifts in the data-generating distribution. Yet, for Continual Learning (CL), models often struggle to balance learning new tasks (plasticity) with retaining previous knowledge (memory stability). Consequently, they are susceptible to Catastrophic Forgetting, which degrades performance and undermines the reliability of deployed systems. In the Bayesian CL literature, variational methods tackle this challenge by employing a learning objective that recursively updates the posterior distribution while constraining it to stay close to its previous estimate. Nonetheless, we argue that these methods may be ineffective due to compounding approximation errors over successive recursions. To mitigate this, we propose new learning objectives that integrate the regularization effects of multiple previous posterior estimations, preventing individual errors from dominating future posterior updates and compounding over time. We reveal insightful connections between these objectives and Temporal-Difference methods, a popular learning mechanism in Reinforcement Learning and Neuroscience. Experiments on challenging CL benchmarks show that our approach effectively mitigates Catastrophic Forgetting, outperforming strong Variational CL methods.
Added
2026-10-05

On the Effectiveness of Lipschitz-Driven Rehearsal in Continual Learning
Lorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato, Simone Calderara
Why you should read this
Proposes a surrogate objective that constrains layer-wise Lipschitz constants on replay data to prevent unstable decision boundaries and buffer overfitting across standard continual learning methods.
Rehearsal approaches enjoy immense popularity with Continual Learning (CL) practitioners. These methods collect samples from previously encountered data distributions in a small memory buffer; subsequently, they repeatedly optimize on the latter to prevent catastrophic forgetting. This work draws attention to a hidden pitfall of this widespread practice: repeated optimization on a small pool of data inevitably leads to tight and unstable decision boundaries, which are a major hindrance to generalization. To address this issue, we propose Lipschitz-DrivEn Rehearsal (LiDER), a surrogate objective that induces smoothness in the backbone network by constraining its layer-wise Lipschitz constants w.r.t. replay examples. By means of extensive experiments, we show that applying LiDER delivers a stable performance gain to several state-of-the-art rehearsal CL methods across multiple datasets, both in the presence and absence of pre-training. Through additional ablative experiments, we highlight peculiar aspects of buffer overfitting in CL and better characterize the effect produced by LiDER. Code is available at https://github.com/aimagelab/LiDER.
Added
2026-09-26

Deep Recurrent Q-Learning for Partially Observable MDPs
Matthew Hausknecht, Peter Stone
Why you should read this
Introduces Deep Recurrent Q-Networks by integrating an LSTM into standard Deep Q-Networks, demonstrating that recurrent memory allows reinforcement learning agents to handle partially observable environments and better adapt to degraded visual inputs than traditional frame-stacking methods.
Deep Reinforcement Learning has yielded proficient controllers for complex tasks. However, these controllers have limited memory and rely on being able to perceive the complete game screen at each decision point. To address these shortcomings, this article investigates the effects of adding recurrency to a Deep Q-Network (DQN) by replacing the first post-convolutional fully-connected layer with a recurrent LSTM. The resulting \textit{Deep Recurrent Q-Network} (DRQN), although capable of seeing only a single frame at each timestep, successfully integrates information through time and replicates DQN's performance on standard Atari games and partially observed equivalents featuring flickering game screens. Additionally, when trained with partial observations and evaluated with incrementally more complete observations, DRQN's performance scales as a function of observability. Conversely, when trained with full observations and evaluated with partial observations, DRQN's performance degrades less than DQN's. Thus, given the same length of history, recurrency is a viable alternative to stacking a history of frames in the DQN's input layer and while recurrency confers no systematic advantage when learning to play the game, the recurrent net can better adapt at evaluation time if the quality of observations changes.
Added
2026-09-17
