Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks

Zhiwei DengOlga Russakovsky

article2022NeurIPS139 citations

Proposes a dataset distillation framework that stores shared memory bases combined via learned addressing functions, breaking the linear scaling bottleneck with class count and substantially outperforming prior distillation and continual learning baselines.

Listen

Modern artificial intelligence models suffer from catastrophic forgetting, rapidly losing previously learned skills when trained on new tasks. Retaining past knowledge traditionally requires storing massive historical datasets, which creates significant data storage costs, memory bottlenecks, and privacy challenges. Dataset distillation compresses large training sets into tiny synthetic subsets that can rapidly retrain neural networks from scratch. However, existing distillation methods assign separate synthetic examples to each class, causing memory usage to scale linearly with the number of categories and creating redundant representations across related classes.

The main objective of the article is to demonstrate that large training datasets can be compressed into a shared, addressable memory space where cross-class patterns are reused. The authors evaluate whether this shared memory architecture can improve compression rates, enhance model accuracy upon retraining, and overcome catastrophic forgetting in continuous learning environments.

To achieve this, the authors restructured dataset distillation by separating shared fundamental components (memory bases) from task-specific retrieval instructions (addressing matrices). Instead of storing independent images per class, the system learns a common memory bank and combines these bases dynamically based on class queries. The authors trained this system using a bi-level optimization process with back-propagation through time, discovering that using momentum and long unrolled training trajectories significantly outperforms standard distillation baselines. They validated the approach across six standard image classification benchmarks and four continuous learning benchmarks.

The evaluation produced four primary findings. First, the method achieved state-of-the-art retained accuracy across all dataset distillation benchmarks under tight storage budgets; on CIFAR-10 with a budget of just one image per class, it reached 66.4% accuracy, outperforming the previous state-of-the-art by 16.5 percentage points. Second, the shared memory framework proved that classes naturally share information, such as related tree or vehicle categories reusing common visual bases. Third, a simple "compress-then-recall" approach achieved state-of-the-art results across four lifelong learning benchmarks, notably improving accuracy on the challenging MANY benchmark from 50.8% to 74.1% (a 23.3 percentage point gain). Fourth, the shared addressable architecture generalized beyond fixed class labels, successfully generating new classifiers for unseen task combinations and recalling synthetic training data directly from continuous image feature queries.

These findings demonstrate that separating shared visual concepts from retrieval mechanisms drastically cuts storage requirements while boosting retraining performance. For organizations deploying edge AI and continuous learning systems, this architecture reduces data storage footprints, mitigates catastrophic forgetting without complex dynamic network designs, and supports flexible retraining under shifting task demands.

Organizations managing constrained edge devices, streaming data, or frequent model updates should consider adopting addressable dataset distillation as a memory-efficient replay mechanism. Development teams can begin by implementing the publicly available codebase on pilot classification workflows to assess compression gains before scaling to production.

While the method shows strong empirical performance, the authors note computational constraints: the inner-loop optimization requires significant training time and processing power during the distillation phase, which may present scaling hurdles for massive models or very high-resolution datasets. Furthermore, extreme compression carries a potential risk of losing distribution diversity, requiring validation to ensure fairness and accuracy across underrepresented data subcategories.

  • Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). Introduces continual learning via synthetic generative replay to prevent catastrophic forgetting, establishing the replay paradigm that dataset distillation and addressable memory build upon.
  • Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). Establishes incremental learning combining exemplar memory with distillation loss, providing the foundational setup for retaining past category knowledge under strict sample budgets.
  • Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Demonstrates the power of replaying continuous optimization dynamics and logits to mitigate catastrophic forgetting under tight buffer constraints.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). Formalizes continual learning with episodic memory buffers and gradient projection, motivating subsequent memory-efficient distillation replay architectures.
  • Paper: Meta-Learning with Memory-Augmented Neural Networks, Adam Santoro et al. (2016). Pioneers external, addressable memory spaces for meta-learning and rapid adaptation, laying the architectural conceptual basis for querying shared synthetic memory bases.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Introduces the use of knowledge distillation objectives to adapt neural networks incrementally without accessing original legacy training data.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). Provides the foundational framework and benchmark definitions for evaluating catastrophic forgetting in sequential neural network learning.
  • Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). Surveys the taxonomy and empirical baselines of continual classification methods, framing the benchmark challenges addressed by dataset distillation replay.
Cover for Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks

Abstract

We propose an algorithm that compresses the critical information of a large dataset into compact addressable memories. These memories can then be recalled to quickly re-train a neural network and recover the performance (instead of storing and re-training on the full original dataset). Building upon the dataset distillation framework, we make a key observation that a shared common representation allows for more efficient and effective distillation. Concretely, we learn a set of bases (aka “memories”) which are shared between classes and combined through learned flexible addressing functions to generate a diverse set of training examples. This leads to several benefits: 1) the size of compressed data does not necessarily grow linearly with the number of classes; 2) an overall higher compression rate with more effective distillation is achieved; and 3) more generalized queries are allowed beyond recalling the original classes. We demonstrate state-of-the-art results on the dataset distillation task across six benchmarks, including up to 16.5% and 9.7% in retained accuracy improvement when distilling CIFAR10 and CIFAR100 respectively. We then leverage our framework to perform continual learning, achieving state-of-the-art results on four benchmarks, with 23.2% accuracy improvement on MANY. The code is released on our project webpage¹.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 Background: dataset distillation
  • 4 Model
  • 4.1 Dataset Distillation as memory addressing
  • 4.2 Learning framework: back-propagation through time
  • 5 Experiments
  • 5.1 Dataset Distillation
  • 5.2 Continual learning
  • 5.3 Synthesizing new classifiers after learning
  • 5.3.1 Extrapolating between tasks
  • 5.3.2 Dataset Distillation extension - recall the past with images
  • 6 Conclusion and limitations
  • 7 Acknowledgements
  • References
  • Checklist

Knowls

  1. Knowl 1 — Dataset Distillation via Shared Memory Addressing

    model/method

    Instead of synthesizing independent images for each individual class, dataset distillation can be formulated as a memory addressing process. A compact set of KK shared base vectors (memories) M={b1,…,bK}\mathcal{M} = \{b_1, \dots, b_K\} is learned, where each base bk∈Rdb_k \in \mathbb{R}^d corresponds to intrinsic spatial components characterizing the task mapping from input space X\mathcal{X} to label space Y\mathcal{Y}.

    To synthesize rr training examples for a given query y∈Rdyy \in \mathbb{R}^{d_y} (such as a one-hot label vector or a feature embedding), a set of rr learnable addressing matrices {A1,…,Ar}⊂Rdy×K\{A_1, \dots, A_r\} \subset \mathbb{R}^{d_y \times K} is defined. The ii-th synthesized example xi′∈Rdx'_i \in \mathbb{R}^d for query yy is computed via a linear combination of the shared bases:

    xi′T=yTAi[b1;… ;bK]Tx_i'^T = y^T A_i [b_1; \dots; b_K]^T

    where v=yTAi∈R1×Kv = y^T A_i \in \mathbb{R}^{1 \times K} is the coefficient vector for query yy. The reconstructed synthetic dataset for a set of queries {yi}\{y_i\} is Ds=⋃yi{(xj′,yi)}j=1r\mathcal{D}_s = \bigcup_{y_i} \{(x'_j, y_i)\}_{j=1}^r. This shared representation decouples the total parameter count from linear scaling with the number of classes and enables cross-class feature sharing.

    Standard dataset distillation is a special case of this formulation where K=N′K = N' (total number of synthetic images), r=N′/Cr = N'/C, and Ai∈{0,1}C×N′A_i \in \{0, 1\}^{C \times N'} are fixed binary indicator matrices such that Ai[m,n]=1A_i[m, n] = 1 if n=m(N′/C)+in = m(N'/C) + i and 00 otherwise.

  2. Knowl 2 — Storage Budget Formulation for Memory Distillation

    equation

    To evaluate compression fairly against standard dataset distillation that stores NN images per class across CC discrete classes, the combined parameter footprint of the memory bases and the addressing matrices is constrained to match the total storage of NCNC full-resolution images:

    size(bases)+size(addressing matrices)≈NC⋅size(image)\text{size}(\text{bases}) + \text{size}(\text{addressing matrices}) \approx N C \cdot \text{size}(\text{image})

    where size(⋅)\text{size}(\cdot) denotes the number of floating-point parameters in the corresponding tensor.

    To improve compression, base vectors can be stored in downsampled spatial resolution (e.g., downsampled by a factor of 2 in both height and width) and deterministically upsampled via bilinear interpolation before linear recombination. Given the chosen base count KK and downsampling factor, the maximum number of addressing matrices rr is set to the integer lower bound that satisfies this equality.

  3. Knowl 3 — Bi-Level Optimization for Distillation via Back-Propagation Through Time

    algorithm

    The parameters ϕ={M,A}\phi = \{\mathcal{M}, \mathcal{A}\} comprising the shared memory bases M={b1,…,bK}\mathcal{M} = \{b_1, \dots, b_K\} and addressing matrices A={A1,…,Ar}\mathcal{A} = \{A_1, \dots, A_r\} are trained end-to-end using bi-level optimization with back-propagation through time (BPTT). The inner loop simulates classifier training on synthesized data with momentum over long unrolling trajectories, and gradients of the outer generalization loss on real training batches are back-propagated through the unrolled steps to update ϕ\phi.

    Input: Training dataset Dtr\mathcal{D}_{tr}, memory bases M\mathcal{M}, addressing matrices A\mathcal{A}, loss function ℓ(⋅,⋅)\ell(\cdot, \cdot), inner learning rate α0\alpha_0, outer learning rate α1\alpha_1, inner momentum rate β0\beta_0, outer momentum rate β1\beta_1, inner trajectory steps TT
    Output: Optimized memory and addressing parameters ϕ={M,A}\phi = \{\mathcal{M}, \mathcal{A}\}
    repeat
        Sample a subset of class labels Y′\mathcal{Y}'
        Construct synthetic dataset DsY′=⋃y∈Y′{(xj′,y)}j=1r\mathcal{D}_s^{\mathcal{Y}'} = \bigcup_{y \in \mathcal{Y}'} \{(x'_j, y)\}_{j=1}^r where xj′T=yTAj[b1;… ;bK]Tx_j'^T = y^T A_j [b_1; \dots; b_K]^T
        Randomly initialize model parameters θ0\theta_0
        Initialize momentum buffer m0←0m_0 \leftarrow 0
        for t=1t = 1 to TT do
            Sample minibatch Bs={(xi′,yi)}B_s = \{(x'_i, y_i)\} from DsY′\mathcal{D}_s^{\mathcal{Y}'}
            Compute inner loss L←1∣Bs∣∑i=1∣Bs∣ℓ(fθt−1(xi′),yi)\mathcal{L} \leftarrow \frac{1}{|B_s|} \sum_{i=1}^{|B_s|} \ell(f_{\theta_{t-1}}(x'_i), y_i)
            Update inner momentum mt←β0mt−1+∇θt−1Lm_t \leftarrow \beta_0 m_{t-1} + \nabla_{\theta_{t-1}} \mathcal{L}
            Update classifier parameters θt←θt−1−α0mt\theta_t \leftarrow \theta_{t-1} - \alpha_0 m_t
        end for
        Sample a validation minibatch B={(xi,yi)}B = \{(x_i, y_i)\} from Dtr\mathcal{D}_{tr} with labels in Y′\mathcal{Y}'
        Compute outer generalization loss J(ϕ)←1∣B∣∑i=1∣B∣ℓ(fθT(xi),yi)J(\phi) \leftarrow \frac{1}{|B|} \sum_{i=1}^{|B|} \ell(f_{\theta_T}(x_i), y_i)
        Update ϕ←OPT-STEP(ϕ,∇ϕJ(ϕ),α1,β1)\phi \leftarrow \text{OPT-STEP}(\phi, \nabla_\phi J(\phi), \alpha_1, \beta_1)
    until converged
  4. Knowl 4 — Dataset Distillation Performance on Standard Classification Benchmarks

    data/table

    Memory-addressable dataset distillation consistently outperforms prior distillation techniques—including Dataset Condensation (DC), Differentiable Siamese Augmentation (DSA), Kernel Inducing Points (KIP), Aligning Features (CAFE), Trajectory Matching (TM), and Distribution Matching (DM)—when evaluated by training a standard ConvNet (with InstanceNorm, ReLU, and 2×22\times 2 average pooling) from scratch on the synthetic data for 300 epochs over 20 random initializations.

    Dataset IPC DC DSA KIP CAFE TM DM Ours
    MNIST 1 91.7±0.591.7\pm0.5 88.7±0.688.7\pm0.6 90.1±0.190.1\pm0.1 93.1±0.393.1\pm0.3 - 89.7±0.689.7\pm0.6 98.7±0.7\mathbf{98.7\pm0.7}
    10 97.4±0.297.4\pm0.2 97.8±0.197.8\pm0.1 97.5±0.097.5\pm0.0 97.5±0.197.5\pm0.1 - 97.5±0.197.5\pm0.1 99.3±0.5\mathbf{99.3\pm0.5}
    50 98.8±0.298.8\pm0.2 99.2±0.199.2\pm0.1 98.3±0.198.3\pm0.1 98.9±0.298.9\pm0.2 - 98.6±0.198.6\pm0.1 99.4±0.4\mathbf{99.4\pm0.4}
    FashionMNIST 1 70.5±0.670.5\pm0.6 70.6±0.670.6\pm0.6 73.5±0.573.5\pm0.5 77.1±0.977.1\pm0.9 - - 88.5±0.1\mathbf{88.5\pm0.1}
    10 82.3±0.482.3\pm0.4 84.6±0.384.6\pm0.3 86.8±0.186.8\pm0.1 83.0±0.383.0\pm0.3 - - 90.0±0.7\mathbf{90.0\pm0.7}
    50 83.6±0.483.6\pm0.4 88.7±0.288.7\pm0.2 88.0±0.188.0\pm0.1 88.2±0.388.2\pm0.3 - - 91.2±0.3\mathbf{91.2\pm0.3}
    SVHN 1 31.2±1.431.2\pm1.4 27.5±1.427.5\pm1.4 57.3±0.157.3\pm0.1 42.9±3.042.9\pm3.0 - - 87.3±0.1\mathbf{87.3\pm0.1}
    10 76.1±0.676.1\pm0.6 79.2±0.579.2\pm0.5 75.0±0.175.0\pm0.1 77.9±0.677.9\pm0.6 - - 89.1±0.2\mathbf{89.1\pm0.2}
    50 82.3±0.382.3\pm0.3 84.4±0.484.4\pm0.4 80.5±0.180.5\pm0.1 82.3±0.482.3\pm0.4 - - 89.5±0.2\mathbf{89.5\pm0.2}
    CIFAR-10 1 28.3±0.528.3\pm0.5 28.8±0.728.8\pm0.7 49.9±0.249.9\pm0.2 31.6±0.831.6\pm0.8 46.3±0.846.3\pm0.8 26.0±0.826.0\pm0.8 66.4±0.4\mathbf{66.4\pm0.4}
    10 44.9±0.544.9\pm0.5 52.1±0.552.1\pm0.5 62.7±0.362.7\pm0.3 50.9±0.550.9\pm0.5 65.3±0.765.3\pm0.7 48.9±0.648.9\pm0.6 71.2±0.4\mathbf{71.2\pm0.4}
    50 53.9±0.553.9\pm0.5 60.6±0.560.6\pm0.5 68.6±0.268.6\pm0.2 62.3±0.462.3\pm0.4 71.6±0.271.6\pm0.2 63.0±0.463.0\pm0.4 73.6±0.5\mathbf{73.6\pm0.5}
    CIFAR-100 1 12.8±0.312.8\pm0.3 13.9±0.313.9\pm0.3 15.7±0.215.7\pm0.2 14.0±0.314.0\pm0.3 24.3±0.324.3\pm0.3 11.4±0.311.4\pm0.3 34.0±0.4\mathbf{34.0\pm0.4}
    10 25.2±0.325.2\pm0.3 32.3±0.332.3\pm0.3 28.3±0.128.3\pm0.1 31.5±0.231.5\pm0.2 40.1±0.440.1\pm0.4 29.7±0.329.7\pm0.3 42.9±0.7\mathbf{42.9\pm0.7}
    TinyImageNet 1 - - - - 8.8±0.38.8\pm0.3 3.9±0.23.9\pm0.2 16.0±0.7\mathbf{16.0\pm0.7}

    IPC denotes images per class storage budget. Performance gains are especially pronounced in extreme compression settings (1 IPC), where the method improves accuracy over prior state-of-the-art by 30.0%30.0\% on SVHN (from 57.3%57.3\% to 87.3%87.3\%), 16.5%16.5\% on CIFAR-10 (from 49.9%49.9\% to 66.4%66.4\%), and 9.7%9.7\% on CIFAR-100 (from 24.3%24.3\% to 34.0%34.0\%).

  5. Knowl 5 — Ablation of BPTT Hyperparameters and Memory Addressing Components

    empirical result

    Ablation experiments demonstrate that the strong performance of memory distillation stems from both algorithmic choices in BPTT (momentum and long unrolling trajectories) and structural modeling choices (spatial downsampling and shared memory addressing):

    1. Inner-loop momentum and unroll length in BPTT: Unlike meta-learning where momentum is frequently omitted, setting inner-loop momentum β0=0.9\beta_0 = 0.9 yields consistent improvements on CIFAR-10 (+7.0%+7.0\% at 10 inner steps, +9.2%+9.2\% at 100 inner steps). Extending unrolled trajectories from 1 step to 150–200 steps increases recovered test accuracy on CIFAR-10 by 18.2%18.2\% and on SVHN by 42.3%42.3\%. A vanilla BPTT baseline with momentum and long unrolls achieves 49.1±0.6%49.1\pm0.6\% on CIFAR-10 (1 IPC) and 80.8%80.8\% on SVHN (1 IPC), substantially outperforming single-step gradient matching (28.8±0.7%28.8\pm0.7\% on CIFAR-10 and 31.2±1.4%31.2\pm1.4\% on SVHN).

    2. Component breakdown on CIFAR benchmarks:

    Dataset IPC Single-step GM Vanilla BPTT BPTT + ds Full w/o Aug Full
    CIFAR-10 1 28.8±0.728.8\pm0.7 49.1±0.649.1\pm0.6 55.2±0.555.2\pm0.5 64.2±0.664.2\pm0.6 66.4±0.466.4\pm0.4
    10 52.1±0.552.1\pm0.5 62.4±0.462.4\pm0.4 65.9±0.465.9\pm0.4 70.9±0.470.9\pm0.4 71.2±0.471.2\pm0.4
    50 60.6±0.560.6\pm0.5 70.5±0.470.5\pm0.4 71.1±0.571.1\pm0.5 72.1±0.572.1\pm0.5 73.8±0.473.8\pm0.4
    CIFAR-100 1 13.9±0.313.9\pm0.3 21.3±0.621.3\pm0.6 25.9±0.425.9\pm0.4 33.5±0.233.5\pm0.2 34.0±0.434.0\pm0.4
    10 32.3±0.332.3\pm0.3 34.7±0.534.7\pm0.5 36.5±0.436.5\pm0.4 40.6±0.340.6\pm0.3 42.9±0.742.9\pm0.7

    where ds denotes 2×2\times spatial downsampling of bases (allowing more base vectors under the fixed storage budget), and Full combines BPTT, downsampled shared memory addressing, and data augmentation. Data augmentation contributes a minor 1–2%1\text{--}2\% gain, confirming that shared memory addressing and long BPTT unrolls are the primary drivers of performance.

  6. Knowl 6 — Continual Learning Performance via Compress-Then-Recall

    data/table

    In continual learning streams, memory distillation enables a 'compress-then-recall' paradigm. Rather than updating a single continuous network across task streams, the streaming data of each task is distilled into memory bases and addressing matrices stored in a memory buffer. For each new task in a total of TT tasks, the distilled representation occupies 1/T1/T of the fixed buffer space. At test time, synthetic data is recalled from the stored memories to train a classifier from scratch for the requested task, completely avoiding catastrophic forgetting and task interference.

    MNIST Rotations MNIST Permutations MANY Permutations Incremental CIFAR-100
    Method RA (%) ↑\uparrow BTI ↓\downarrow RA (%) ↑\uparrow BTI ↓\downarrow RA (%) ↑\uparrow BTI ↓\downarrow RA (%) ↑\uparrow BTI ↓\downarrow
    ONLINE 53.38±1.5353.38\pm1.53 -5.44 55.42±0.6555.42\pm0.65 -13.76 32.62±0.4332.62\pm0.43 -19.06 32.62±0.4332.62\pm0.43 -19.06
    EWC 57.96±1.3357.96\pm1.33 -20.42 62.32±1.3462.32\pm1.34 -13.32 33.10±0.1433.10\pm0.14 -18.50 - -
    GEM 67.38±1.7567.38\pm1.75 -18.02 55.42±1.1055.42\pm1.10 -24.42 39.50±0.6239.50\pm0.62 -17.50 48.27±1.1048.27\pm1.10 -13.7
    MER 77.42±0.7877.42\pm0.78 -5.60 73.46±0.4573.46\pm0.45 -9.96 51.00±0.5451.00\pm0.54 -13.57 51.38±1.0551.38\pm1.05 -12.83
    La-MAML 77.42±0.6577.42\pm0.65 -8.64 74.34±0.6774.34\pm0.67 -7.60 50.43±0.2150.43\pm0.21 -10.00 61.18±1.4461.18\pm1.44 -9.00
    Sparse-LaMAML 77.77±0.5877.77\pm0.58 -8.16 76.88±0.7276.88\pm0.72 -8.39 50.81±0.7950.81\pm0.79 -13.73 - -
    Ours 80.32±0.28\mathbf{80.32\pm0.28} N/A 78.48±0.76\mathbf{78.48\pm0.76} N/A 74.07±0.51\mathbf{74.07\pm0.51} N/A 62.58±1.1\mathbf{62.58\pm1.1} N/A

    RA denotes retained accuracy (average accuracy across all tasks at conclusion) and BTI denotes backward transfer and interference. On MANY Permutations (100 tasks of 200 samples each), compress-then-recall achieves 74.07%74.07\%, outperforming prior methods by 23.26%23.26\%. In large-scale task regimes (60,000 samples per task), the method achieves 87.3±0.92%87.3\pm0.92\% on Permuted MNIST and 88.3±0.58%88.3\pm0.58\% on Rotated MNIST, surpassing both Kernel Continual Learning (85.5%85.5\% and 81.8%81.8\%) and the multi-task joint training upper bound (86.5%86.5\% and 87.3%87.3\%).

  7. Knowl 7 — Classifier Extrapolation and Continuous Image-Based Memory Recall

    empirical result

    The memory addressing formulation supports non-standard query mechanisms, including synthesizing classifiers for novel class combinations and addressing memories with continuous visual embeddings:

    1. Extrapolating across disjoint tasks: CIFAR-100 is split into 20 disjoint 5-way tasks, each distilled independently into memories. At test time, class combinations unseen during training are formed by sampling 1 class from each of kk distinct tasks (k=2k=2 or k=5k=5). Recalling synthetic data for these label subsets and training a new classifier yields 72.53±8.74%72.53\pm8.74\% accuracy on 2-way tasks and 46.54±6.42%46.54\pm6.42\% on 5-way tasks (evaluated over 1,000 sampled combinations at 1 IPC budget; real dataset upper bounds are 92.23±4.76%92.23\pm4.76\% and 82.72±4.29%82.72\pm4.29\%).

    2. Few-shot memory addressing with image queries: When categorical labels are absent at query time, memories can be recalled using continuous visual feature vectors y=h(x)y = h(x) extracted by a jointly trained feature extractor network h(⋅)h(\cdot). Evaluated on CIFAR-100 few-shot performance recovery:

    Method 1-shot 5-shot
    Nearest neighbor 48.55 61.72
    Classify-then-recall 50.58 58.46
    Image addressing (Ours) 55.74 71.20

    Direct image addressing outperforms both a nearest-neighbor baseline on pretrained features and a classify-then-recall baseline that predicts discrete classes before addressing.

  8. Knowl 8 — Computational and Representation Limitations of Memory Distillation

    limitation

    The memory-addressing distillation framework possesses two notable limitations:

    1. Inner-loop optimization overhead: Unrolling back-propagation through time for 150 to 200 inner SGD steps with momentum requires storing intermediate computational graphs, resulting in high GPU compute and memory costs during meta-gradient calculation. This inner loop unrolling becomes computationally expensive when scaling to large neural network backbones or large-scale high-resolution datasets.
    2. Diversity compression bias: Aggressive dataset compression may fail to capture the full distributional diversity of the original dataset, risking degraded classifier performance on underrepresented minority sub-populations.

Coverage note — None was omitted; all contributed models, equations, algorithms, empirical distillation benchmarks, ablation analyses, continual learning experiments, query generalization setups, and stated limitations are covered.

References

  1. 1.Timothy F Brady, Talia Konkle, and George A Alvarez. Compression in visual working memory: using statistical regularities to form more efficient memory representations. Journal of Experimental Psychology: General, 138(4):487, 2009.
  2. 2.Geoffrey R Loftus and Elizabeth F Loftus. Human memory: The processing of information. Psychology Press, 2019.
  3. 3.John R Anderson and Gordon H Bower. Human associative memory. Psychology press, 2014.
  4. 4.Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio. An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211, 2013.
  5. 5.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pages 109–165. Elsevier, 1989.
  6. 6.Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. Dataset distillation. arXiv preprint arXiv:1811.10959, 2018.
  7. 7.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014.
  8. 8.Ian Gemp, Brian McWilliams, Claire Vernade, and Thore Graepel. Eigengame: Pca as a nash equilibrium. In International Conference on Learning Representations, 2020.
  9. 9.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  10. 10.Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. Dataset condensation with gradient matching. In International Conference on Learning Representations, 2021.
  11. 11.Bo Zhao and Hakan Bilen. Dataset condensation with differentiable siamese augmentation. In International Conference on Machine Learning, 2021.
  12. 12.Timothy Nguyen, Roman Novak, Lechao Xiao, and Jaehoon Lee. Dataset distillation with infinitely wide convolutional networks. Advances in Neural Information Processing Systems, 34, 2021.
  13. 13.Timothy Nguyen, Zhourong Chen, and Jaehoon Lee. Dataset meta-learning from kernel ridge-regression. In International Conference on Learning Representations, 2021.
  14. 14.Bo Zhao and Hakan Bilen. Dataset condensation with distribution matching. arXiv preprint arXiv:2110.04181, 2021.
  15. 15.George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu. Dataset distillation by matching training trajectories. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4750–4759, 2022.
  16. 16.Ilia Sucholutsky and Matthias Schonlau. Soft-label dataset distillation and text dataset distillation. arXiv preprint arXiv:1910.02551, 2019.
  17. 17.Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro. Learning to learn without forgetting by maximizing transfer and minimizing interference. arXiv preprint arXiv:1810.11910, 2018.
  18. 18.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, pages 1126–1135. PMLR, 2017.
  19. 19.Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil. Forward and reverse gradient-based hyperparameter optimization. In International Conference on Machine Learning, pages 1165–1173. PMLR, 2017.
  20. 20.Aniruddh Raghu, Maithra Raghu, Simon Kornblith, David Duvenaud, and Geoffrey Hinton. Teaching with commentaries. In International Conference on Learning Representations, 2020.
  21. 21.Sid Reddy, Anca Dragan, and Sergey Levine. Pragmatic image compression for human-in-the-loop decision-making. Advances in Neural Information Processing Systems, 34, 2021.
  22. 22.Shengjia Zhao, Abhishek Sinha, Yutong He, Aidan Perreault, Jiaming Song, and Stefano Ermon. Comparing distributions by measuring differences that affect decision making. In International Conference on Learning Representations, 2021.
  23. 23.Roger Ratcliff. Connectionist models of recognition memory: constraints imposed by learning and forgetting functions. Psychological review, 97(2):285, 1990.
  24. 24.David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. Advances in neural information processing systems, 30:6467–6476, 2017.
  25. 25.Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. On tiny episodic memories in continual learning. arXiv preprint arXiv:1902.10486, 2019.
  26. 26.Tom Veniat, Ludovic Denoyer, and Marc’Aurelio Ranzato. Efficient continual learning with modular networks and task-driven priors. arXiv preprint arXiv:2012.12631, 2020.
  27. 27.Cuong V Nguyen, Yingzhen Li, Thang D Bui, and Richard E Turner. Variational continual learning. arXiv preprint arXiv:1710.10628, 2017.
  28. 28.Gunshi Gupta, Karmesh Yadav, and Liam Paull. La-maml: Look-ahead meta learning for continual learning. arXiv preprint arXiv:2007.13904, 2020.
  29. 29.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2001–2010, 2017.
  30. 30.Johannes Von Oswald, Dominic Zhao, Seijin Kobayashi, Simon Schug, Massimo Caccia, Nicolas Zucchet, and João Sacramento. Learning where to learn: Gradient sparsity in meta and continual learning. Advances in Neural Information Processing Systems, 34, 2021.
  31. 31.Ameya Prabhu, Philip HS Torr, and Puneet K Dokania. Gdumb: A simple approach that questions our progress in continual learning. In European conference on computer vision, pages 524–540. Springer, 2020.
  32. 32.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. Continual learning: A comparative study on how to defy forgetting in classification tasks. arXiv preprint arXiv:1909.08383, 2(6), 2019.
  33. 33.Raia Hadsell, Dushyant Rao, Andrei A Rusu, and Razvan Pascanu. Embracing change: Continual learning in deep neural networks. Trends in cognitive sciences, 24(12):1028–1040, 2020.
  34. 34.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  35. 35.Jaehong Yoon, Eunho Yang, Jeongtae Lee, and Sung Ju Hwang. Lifelong learning with dynamically expandable networks. In International Conference on Learning Representations, 2018.
  36. 36.Martial Mermillod, Aurélia Bugaiska, and Patrick Bonin. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects. Frontiers in psychology, 4:504, 2013.
  37. 37.Seyed Iman Mirzadeh, Mehrdad Farajtabar, Razvan Pascanu, and Hassan Ghasemzadeh. Understanding the role of training regimes in continual learning. Advances in Neural Information Processing Systems, 33:7308–7320, 2020.
  38. 38.Gobinda Saha, Isha Garg, and Kaushik Roy. Gradient projection memory for continual learning. In International Conference on Learning Representations, 2020.
  39. 39.Mohammad Mahdi Derakhshani, Xiantong Zhen, Ling Shao, and Cees Snoek. Kernel continual learning. In International Conference on Machine Learning, pages 2621–2631. PMLR, 2021.
  40. 40.Alex Nichol and John Schulman. Reptile: a scalable metalearning algorithm. arXiv preprint arXiv:1803.02999, 2(3):4, 2018.
  41. 41.Li Deng. The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine, 29(6):141–142, 2012.
  42. 42.Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  43. 43.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  44. 44.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  45. 45.Ya Le and Xuan Yang. Tiny imagenet visual recognition challenge. CS 231N, 7(7):3, 2015.
  46. 46.Kai Wang, Bo Zhao, Xiangyu Peng, Zheng Zhu, Shuo Yang, Shuo Wang, Guan Huang, Hakan Bilen, Xinchao Wang, and Yang You. Cafe learning to condense dataset by aligning features. In IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022, 2022.
  47. 47.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. Efficient lifelong learning with a-gem. arXiv preprint arXiv:1812.00420, 2018.
  48. 48.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.

Citation

MLA
Deng, Z., and O. Russakovsky. “Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 34391–404, https://proceedings.neurips.cc/paper_files/paper/2022/file/de3d2bb604cfc43c81edd2a31b257f03-Paper-Conference.pdf.
APA
Deng, Z., & Russakovsky, O. (2022). Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks. Advances in Neural Information Processing Systems, 35, 34391–34404. https://proceedings.neurips.cc/paper_files/paper/2022/file/de3d2bb604cfc43c81edd2a31b257f03-Paper-Conference.pdf
Chicago
Deng, Z., and O. Russakovsky. 2022. “Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks”. Advances in Neural Information Processing Systems 35: 34391–404. https://proceedings.neurips.cc/paper_files/paper/2022/file/de3d2bb604cfc43c81edd2a31b257f03-Paper-Conference.pdf.
Harvard
Deng, Z. and Russakovsky, O. (2022) “Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 34391–34404. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/de3d2bb604cfc43c81edd2a31b257f03-Paper-Conference.pdf.
Vancouver
1. Deng Z, Russakovsky O (2022) Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 34391–34404

BibTeX

@inproceedings{deng2022remember,
  title = {Remember the Past: Distilling Datasets into Addressable Memories for Neural Networks},
  author = {Deng, Zhiwei and Russakovsky, Olga},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {34391-34404},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/de3d2bb604cfc43c81edd2a31b257f03-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors