Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks
Zhiwei DengOlga Russakovsky
Proposes a dataset distillation framework that stores shared memory bases combined via learned addressing functions, breaking the linear scaling bottleneck with class count and substantially outperforming prior distillation and continual learning baselines.
Modern artificial intelligence models suffer from catastrophic forgetting, rapidly losing previously learned skills when trained on new tasks. Retaining past knowledge traditionally requires storing massive historical datasets, which creates significant data storage costs, memory bottlenecks, and privacy challenges. Dataset distillation compresses large training sets into tiny synthetic subsets that can rapidly retrain neural networks from scratch. However, existing distillation methods assign separate synthetic examples to each class, causing memory usage to scale linearly with the number of categories and creating redundant representations across related classes.
The main objective of the article is to demonstrate that large training datasets can be compressed into a shared, addressable memory space where cross-class patterns are reused. The authors evaluate whether this shared memory architecture can improve compression rates, enhance model accuracy upon retraining, and overcome catastrophic forgetting in continuous learning environments.
To achieve this, the authors restructured dataset distillation by separating shared fundamental components (memory bases) from task-specific retrieval instructions (addressing matrices). Instead of storing independent images per class, the system learns a common memory bank and combines these bases dynamically based on class queries. The authors trained this system using a bi-level optimization process with back-propagation through time, discovering that using momentum and long unrolled training trajectories significantly outperforms standard distillation baselines. They validated the approach across six standard image classification benchmarks and four continuous learning benchmarks.
The evaluation produced four primary findings. First, the method achieved state-of-the-art retained accuracy across all dataset distillation benchmarks under tight storage budgets; on CIFAR-10 with a budget of just one image per class, it reached 66.4% accuracy, outperforming the previous state-of-the-art by 16.5 percentage points. Second, the shared memory framework proved that classes naturally share information, such as related tree or vehicle categories reusing common visual bases. Third, a simple "compress-then-recall" approach achieved state-of-the-art results across four lifelong learning benchmarks, notably improving accuracy on the challenging MANY benchmark from 50.8% to 74.1% (a 23.3 percentage point gain). Fourth, the shared addressable architecture generalized beyond fixed class labels, successfully generating new classifiers for unseen task combinations and recalling synthetic training data directly from continuous image feature queries.
These findings demonstrate that separating shared visual concepts from retrieval mechanisms drastically cuts storage requirements while boosting retraining performance. For organizations deploying edge AI and continuous learning systems, this architecture reduces data storage footprints, mitigates catastrophic forgetting without complex dynamic network designs, and supports flexible retraining under shifting task demands.
Organizations managing constrained edge devices, streaming data, or frequent model updates should consider adopting addressable dataset distillation as a memory-efficient replay mechanism. Development teams can begin by implementing the publicly available codebase on pilot classification workflows to assess compression gains before scaling to production.
While the method shows strong empirical performance, the authors note computational constraints: the inner-loop optimization requires significant training time and processing power during the distillation phase, which may present scaling hurdles for massive models or very high-resolution datasets. Furthermore, extreme compression carries a potential risk of losing distribution diversity, requiring validation to ensure fairness and accuracy across underrepresented data subcategories.
- Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). Introduces continual learning via synthetic generative replay to prevent catastrophic forgetting, establishing the replay paradigm that dataset distillation and addressable memory build upon.
- Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). Establishes incremental learning combining exemplar memory with distillation loss, providing the foundational setup for retaining past category knowledge under strict sample budgets.
- Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Demonstrates the power of replaying continuous optimization dynamics and logits to mitigate catastrophic forgetting under tight buffer constraints.
- Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Formalizes continual learning with episodic memory buffers and gradient projection, motivating subsequent memory-efficient distillation replay architectures.
- Paper: Meta-Learning with Memory-Augmented Neural Networks, Adam Santoro et al. (2016). Pioneers external, addressable memory spaces for meta-learning and rapid adaptation, laying the architectural conceptual basis for querying shared synthetic memory bases.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Introduces the use of knowledge distillation objectives to adapt neural networks incrementally without accessing original legacy training data.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Provides the foundational framework and benchmark definitions for evaluating catastrophic forgetting in sequential neural network learning.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). Surveys the taxonomy and empirical baselines of continual classification methods, framing the benchmark challenges addressed by dataset distillation replay.
- Paper: Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation, Jiawei Du et al. (2023). Investigates and resolves accumulated trajectory errors in dataset distillation by regularizing optimization paths toward flat loss landscapes.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). Provides an expansive, updated survey synthesizing modern replay-based, optimization-based, and modular continual learning paradigms.
- Paper: Dense Network Expansion for Class Incremental Learning, Zhiyuan Hu et al. (2023). Applies cross-task intermediate feature reuse to vision transformers to curb architectural expansion costs in class-incremental learning.
