Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning

Nader AsadiMohammadReza DavariSudhir P. MudurRahaf AljundiEugene Belilovsky

article2023ICML74 citations

Proposes a replay-free continual learning framework that prevents catastrophic forgetting by distilling the relative similarities between incoming samples and past class prototypes, outperforming memory-buffer methods without storing historical data.

Listen

Continual learning requires machine learning systems to adapt continuously to incoming streams of new data without suffering from catastrophic forgetting, which occurs when a model forgets previously acquired knowledge. The most effective existing methods generally rely on rehearsal buffers that store past data samples to preserve performance. However, storing historical data poses serious regulatory, privacy, and technical challenges when handling sensitive medical or proprietary information, while also creating escalating storage demands as task sequences grow.

The article evaluates a new framework designed to eliminate the need for storing prior task data during continual learning. Specifically, it demonstrates that joint representation learning and prototype adaptation can maintain high predictive accuracy across evolving tasks without requiring access to previous samples.

The researchers developed an approach termed Prototype-Sample Relation Distillation (PRD). The framework learns rich data representations via supervised contrastive learning and maintains class prototypes—vector representations of classes—without using suppressive contrastive terms. To prevent older class prototypes from becoming obsolete when the underlying model updates on new tasks, PRD applies a relation distillation loss that preserves the relative similarity rankings between old prototypes and incoming data samples. The authors evaluated PRD across multiple benchmark datasets, including Split-CIFAR100, Split-MiniImageNet, ImageNet32, and the real-world autonomous driving dataset CLAD-C, assessing both task-incremental and class-incremental configurations over short and long sequences reaching up to 200 tasks.

The evaluation yielded several key findings. First, in task-incremental settings across 20 to 200 tasks, PRD consistently outperformed existing replay-free baselines and matched or exceeded experience replay methods using 50 stored samples per class. Second, in class-incremental settings without any stored data, PRD achieved 27.8% accuracy on CIFAR-100 and 20.0% on MiniImageNet, outperforming experience replay baselines limited to 5 stored samples per class. Third, when provided with replay samples, PRD reached 45.1% accuracy on CIFAR-100, surpassing standard experience replay at 38.3%. Fourth, on the real-world CLAD-C benchmark exhibiting day-night shifts and severe class imbalance, PRD attained an Average Mean Class Accuracy of 65.1%, outperforming alternative replay-free techniques. Finally, ablation analyses confirmed that relative relation distillation is critical for maintaining stability on older classes while preserving plasticity to learn new tasks.

These findings indicate that organizations can train adaptable machine learning systems without retaining historical user data or intellectual property. This significantly reduces data governance risks, storage costs, and compliance overhead associated with long-term data retention. The results challenge the conventional belief that competitive continual learning strictly requires rehearsal buffers, proving that high plasticity and retention can be attained through relative geometric alignment in representation space.

Organizations developing machine learning models for streaming environments should consider prototype-relation frameworks when data retention is restricted by privacy policies or memory limits. In scenarios where data storage is permissible, hybrid implementations combining PRD with small buffers provide substantial performance gains over conventional replay methods. Teams should conduct domain-specific pilot testing on their streaming pipelines to fine-tune distillation coefficients for target sequence lengths.

Confidence in these findings is supported by rigorous benchmarking across diverse image datasets and consistent outperformance over established baselines. However, limitations remain: the article evaluated the method primarily on computer vision classification tasks using standard convolutional architectures, and initial accuracy drops were observed in class-incremental settings. Practitioners should exercise caution before generalizing these results to non-visual modalities, such as natural language or tabular data, without further empirical validation.

arXiv: 2303.14771
Cover for Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning

Abstract

In Continual learning (CL) balancing effective adaptation while combating catastrophic forgetting is a central challenge. Many of the recent best-performing methods utilize various forms of prior task data, e.g. a replay buffer, to tackle the catastrophic forgetting problem. Having access to previous task data can be restrictive in many real-world scenarios, for example when task data is sensitive or proprietary. To overcome the necessity of using previous tasks' data, in this work, we start with strong representation learning methods that have been shown to be less prone to forgetting. We propose a holistic approach to jointly learn the representation and class prototypes while maintaining the relevance of old class prototypes and their embedded similarities. Specifically, samples are mapped to an embedding space where the representations are learned using a supervised contrastive loss. Class prototypes are evolved continually in the same latent space, enabling learning and prediction at any point. To continually adapt the prototypes without keeping any prior task data, we propose a novel distillation loss that constrains class prototypes to maintain relative similarities as compared to new task data. This method yields state-of-the-art performance in the task-incremental setting, outperforming methods relying on large amounts of data, and provides strong performance in the class-incremental setting without using any stored data points.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 4. Experiments
  • 4.1. Evaluations on Task-Incremental Setting
  • 4.2. Class-Incremental Setting
  • 4.3. Analysis and Ablations
  • 5. Conclusion
  • Acknowledgements
  • References
  • APPENDIX
  • A. Experimental Setup
  • B. Ablation on Prototypes-Samples Similarity Distillation
  • C. Ablation on Prototype Learning without Contrasts
  • D. Domain-Incremental Experiment

Knowls

  1. Knowl 1 — PRD combines contrastive representation learning with prototype classification without replay

    model/method

    Prototype-Sample Relation Distillation (PRD) is a continual-learning method for sequential training sessions in which raw data from earlier sessions cannot be accessed. An encoder maps inputs to features, and a projection head maps features to a latent space for supervised contrastive learning. Each class has a prototype in the encoder-feature space; new classes receive prototypes as they are learned, while prototypes for previously learned classes are retained. Training on a session combines three objectives: supervised contrastive learning on the incoming data, a tightness loss that fits current-class prototypes to their samples, and a relation-distillation loss that uses incoming samples to preserve the similarity relationships of old prototypes. No old examples are required for this distillation. At inference, a sample is assigned to the class whose prototype has the greatest similarity to its encoder feature.

  2. Knowl 2 — Prototype-sample relation distillation preserves old prototypes’ relative similarities

    equation

    At session tt, let XX be a minibatch of incoming samples, fhetatf_{ heta_t} the current encoder, and PotP_o^t the prototypes for classes learned before session tt. For an old-class prototype pktp_k^t, define a probability distribution over minibatch samples by applying a softmax over their similarities to that prototype. Here ii indexes samples in XX, h(p,z)=exp⁡(sim⁡(p,z)/τ)h(p,z)=\exp(\operatorname{sim}(p,z)/\tau), sim⁡\operatorname{sim} is cosine similarity, and τ>0\tau>0 is the temperature. The old model and prototypes, fθt−1f_{\theta_{t-1}} and Pot−1P_o^{t-1}, are the reference snapshot from the end of the preceding session.

    Pt(k)i=h(pkt,fθt(xi))∑xj∈Xh(pkt,fθt(xj)),Ld=∑pk∈PotKL ⁣(Pt(k) ∥ Pt−1(k)).P_t(k)_i=\frac{h(p_k^t,f_{\theta_t}(x_i))}{\sum_{x_j\in X}h(p_k^t,f_{\theta_t}(x_j))},\qquad L_d=\sum_{p_k\in P_o^t}\mathrm{KL}\!\left(P_t(k)\,\|\,P_{t-1}(k)\right).

    The KL divergence is between distributions over samples in the same incoming minibatch, not between class distributions for each sample. Minimizing this loss encourages each old prototype to retain its relative similarity pattern across current-task samples as the representation changes. The reference distribution is computed using the preceding-session model and prototypes; no previous-session samples are needed.

  3. Knowl 3 — Supervised contrastive learning trains the representation on incoming sessions

    equation

    For a minibatch XX of augmented inputs with labels yiy_i, let zi=g(fθ(xi))z_i=g(f_\theta(x_i)) be the projection of sample xix_i, where fθf_\theta is the encoder and gg is the projection head. Let A(i)A(i) contain the other views in the minibatch that share xix_i’s class, including its augmented views, and let τ>0\tau>0 be the temperature. With cosine similarity sim⁡\operatorname{sim} and h(a,b)=exp⁡(sim⁡(a,b)/τ)h(a,b)=\exp(\operatorname{sim}(a,b)/\tau), the supervised contrastive objective is

    LSC(X)=−∑i∈X1∣A(i)∣∑p∈A(i)log⁡h(zp,zi)∑a∈X, a≠ih(za,zi).L_{SC}(X)=-\sum_{i\in X}\frac{1}{|A(i)|}\sum_{p\in A(i)}\log\frac{h(z_p,z_i)}{\sum_{a\in X,\ a\ne i}h(z_a,z_i)}.

    Thus, each anchor is trained to have high similarity to same-class positives relative to the other samples in the minibatch. PRD applies this loss to incoming-session data to learn the representation.

  4. Knowl 4 — Current-class prototypes are learned with a contrast-free tightness loss

    equation

    For each class cc introduced in the current session, PRD initializes a prototype pc∈Rdp_c\in\mathbb{R}^d, where dd is the encoder-feature dimension. For a labeled minibatch XX, the prototype loss is

    Lp(X)=−1∣X∣∑(xi,yi)∈Xsim⁡ ⁣(pyi,sg⁡[fθ(xi)]),L_p(X)=-\frac{1}{|X|}\sum_{(x_i,y_i)\in X}\operatorname{sim}\!\left(p_{y_i},\operatorname{sg}[f_\theta(x_i)]\right),

    where fθ(xi)∈Rdf_\theta(x_i)\in\mathbb{R}^d, sim⁡\operatorname{sim} is cosine similarity, and sg⁡\operatorname{sg} stops gradients through the encoder feature. Consequently, this term updates current-class prototypes to be similar to their own samples without using negative prototype contrasts or directly changing the representation. The paper’s stated design is that omitting contrasts avoids suppressing prototypes of previously learned classes; representation learning and class separation are handled by the supervised contrastive objective.

  5. Knowl 5 — PRD outperforms replay baselines in task-incremental sequences, including 200 tasks

    empirical result

    In task-incremental experiments, the paper evaluates 20 five-class tasks on Split-CIFAR100 and Split-MiniImageNet, and 200 five-class tasks on ImageNet32. The reported metric is average observed accuracy: accuracy on each task’s test data after each training step, averaged over the observed tasks at the end of the sequence. The common image-experiment setup uses a ResNet-18, batch size 128, and 100 training epochs per task; results are reported as means and standard errors over three runs.

    On both 20-task datasets, the reported accuracy curves show PRD outperforming replay-free baselines and the experience-replay (ER) baseline with 50 stored samples per class, while approaching joint i.i.d. training. On the 200-task ImageNet32 sequence, PRD matches the large-buffer ER baseline and overtakes it in later stages, while its average observed accuracy rises during part of the sequence. These results are reported for PRD without storing previous-task data.

  6. Knowl 6 — Class-incremental accuracy with and without replay

    data/table

    The table reports average observed accuracy (%) on 20-task class-incremental Split-CIFAR100 and Split-MiniImageNet sequences. MM is the number of replay samples stored per class; PRD at M=0M=0 uses no replay. Entries are mean ±\pm standard error over three runs. The comparison shows that replay-free PRD exceeds the listed replay-based methods at M=5M=5 on both datasets. With M=50M=50, PRD also exceeds ER with the same buffer size on both datasets.

    Split-CIFAR100 Split-MiniImageNet
    Method M=0M=0 M=5M=5 M=20M=20 M=50M=50 M=0M=0 M=5M=5 M=20M=20 M=50M=50
    i.i.d. 65.3 65.3 65.3 65.3 54.5 54.5 54.5 54.5
    Fine-tuning 4.6 4.6 4.6 4.6 4.4 4.4 4.4 4.4
    ER – 15.2±0.715.2\pm0.7 29.6±0.829.6\pm0.8 38.3±0.938.3\pm0.9 – 13.4±0.213.4\pm0.2 21.7±0.821.7\pm0.8 28.7±0.528.7\pm0.5
    iCaRL – 19.8±0.519.8\pm0.5 28.6±0.728.6\pm0.7 32.9±0.532.9\pm0.5 – 16.2±0.116.2\pm0.1 22.8±0.322.8\pm0.3 26.1±0.226.1\pm0.2
    ER-AML – 21.4±0.821.4\pm0.8 35.3±0.635.3\pm0.6 42.4±0.842.4\pm0.8 – 17.1±0.317.1\pm0.3 26.3±0.726.3\pm0.7 32.3±0.232.3\pm0.2
    ER-ACE – 22.8±0.522.8\pm0.5 35.7±0.235.7\pm0.2 43.3±0.243.3\pm0.2 – 18.8±0.118.8\pm0.1 27.1±0.527.1\pm0.5 34.2±0.534.2\pm0.5
    PRD 27.8±0.227.8\pm0.2 32.0±0.432.0\pm0.4 39.5±0.439.5\pm0.4 45.1±0.545.1\pm0.5 20.0±0.120.0\pm0.1 25.7±0.325.7\pm0.3 31.3±0.531.3\pm0.5 35.8±0.435.8\pm0.4
  7. Knowl 7 — PRD improves class-incremental accuracy after pre-trained initialization

    data/table

    This experiment uses the class-incremental protocol with pre-training on half the classes, followed by K−1K-1 incremental phases. ImageNet-Subset contains 100 classes, with the first 50 used in the initial phase; the remaining classes are evenly divided among the later phases. Split-CIFAR100 and ImageNet-Subset are evaluated with K=6K=6 and K=11K=11. The metric is average cumulative incremental accuracy across phases. The reported values show PRD exceeding the listed baselines, including SPB, in all four dataset-and-phase configurations.

    Split-CIFAR100 ImageNet-Subset
    Method K=6K=6 K=11K=11 K=6K=6 K=11K=11
    i.i.d. 73.4 73.2 82.0 82.7
    Fine-tuning 22.3 12.6 23.6 13.2
    LwF-E 57.0 56.8 65.5 65.6
    EWC-E 56.3 55.4 65.2 64.1
    MAS-E 56.9 56.6 65.8 65.8
    SDC 57.1 56.8 65.6 65.7
    SPB 60.9 60.4 68.7 67.2
    PRD 64.3 63.7 71.8 70.3
  8. Knowl 8 — PRD’s task-incremental results indicate a stability–plasticity trade-off

    empirical result

    On task-incremental Split-CIFAR100, the paper separately compares accuracy on the current task and average accuracy on old tasks for tasks 5, 10, and 15. Relative to EWC and ER with 50 replay samples per class, PRD has higher current-task accuracy while largely maintaining old-task accuracy as training proceeds. The authors characterize PRD’s ability to learn new tasks as close to unconstrained fine-tuning, while retaining substantially better stability. This analysis supports their claim that PRD’s strong observed accuracy is not obtained solely by sacrificing adaptation to new tasks.

  9. Knowl 9 — PRD reduces representation forgetting on a 200-task sequence

    empirical result

    The paper evaluates representation retention on task-incremental ImageNet32 with 200 five-class tasks by tracking task 1 linear-probe accuracy as later tasks are learned. PRD’s task-1 probe accuracy remains relatively flat and increases at some later points in the sequence. The reported comparison shows a substantial improvement over naive supervised-contrastive fine-tuning; the authors also report that naive supervised-contrastive training exceeds ER with a five-sample-per-class buffer on this representation-forgetting measure. The analysis is presented as evidence that prototype-sample relation distillation benefits the representation itself, not only the prototype classifier.

  10. Knowl 10 — PRD obtains the highest reported AMCA on a domain-incremental driving benchmark

    data/table

    The domain-incremental experiment uses CLAD-C, a dashcam dataset with six day/night-shifted tasks and six object classes. The training set contains 22,249 objects, and the test set contains 69,881 objects spanning day and night. All methods use an ImageNet-pretrained ResNet-50 and batch size 32. The reported final Average Mean Class Accuracy (AMCA) averages accuracy over tasks and classes. PRD achieves 65.1, above LwF at 63.7, EWC at 62.5, and fine-tuning at 40.5, in a setting described as having severe class imbalance and distribution shifts.

    AMCA=1∣T∣∣C∣∑t∈T∑c∈CAct,\mathrm{AMCA}=\frac{1}{|T||C|}\sum_{t\in T}\sum_{c\in C}A_c^t,

    where TT is the set of tasks, CC is the set of classes, and ActA_c^t is accuracy for class cc on task tt.

    Method Fine-tuning EWC LwF PRD
    AMCA 40.5 62.5 63.7 65.1
  11. Knowl 11 — Relation-distillation weight matters most on long task sequences

    data/table

    This ablation varies β\beta, the coefficient on prototype-sample relation distillation, and reports task-incremental average observed accuracy as mean ±\pm standard error. The datasets have 20 tasks for Split-CIFAR100 and Split-MiniImageNet, and 200 tasks for ImageNet32. Setting β=0\beta=0 sharply reduces accuracy on all three datasets. The best tested values are β=8\beta=8 for both 20-task datasets and β=16\beta=16 for 200-task ImageNet32, consistent with the paper’s observation that stronger distillation can help on longer sequences.

    β\beta Split-CIFAR100 (K=20K=20) Split-MiniImageNet (K=20K=20) ImageNet32 (K=200K=200)
    0 39.4±1.539.4\pm1.5 31.2±0.931.2\pm0.9 21.3±0.821.3\pm0.8
    1 80.0±0.580.0\pm0.5 59.3±0.359.3\pm0.3 55.4±0.455.4\pm0.4
    2 82.1±0.382.1\pm0.3 63.7±0.363.7\pm0.3 59.2±0.459.2\pm0.4
    4 83.5±0.483.5\pm0.4 68.3±0.468.3\pm0.4 62.7±0.262.7\pm0.2
    8 83.1±0.483.1\pm0.4 70.9±0.570.9\pm0.5 65.1±0.365.1\pm0.3
    16 82.7±0.282.7\pm0.2 67.2±0.567.2\pm0.5 67.5±0.267.5\pm0.2
  12. Knowl 12 — Prototype-learning weight is necessary, but its tested value has limited effect

    data/table

    This ablation varies α\alpha, the coefficient on the current-class prototype tightness loss, and reports task-incremental average observed accuracy on the two 20-task datasets. With α=0\alpha=0, prototypes are not optimized and accuracy falls to near chance. For positive tested values, performance varies comparatively little; the best listed result is α=4\alpha=4 on Split-CIFAR100 and α=2\alpha=2 on Split-MiniImageNet.

    α\alpha Split-CIFAR100 (K=20K=20) Split-MiniImageNet (K=20K=20)
    0 20.6±0.220.6\pm0.2 20.2±0.220.2\pm0.2
    1 81.8±0.581.8\pm0.5 68.0±0.468.0\pm0.4
    2 82.2±0.482.2\pm0.4 69.8±0.369.8\pm0.3
    4 83.5±0.483.5\pm0.4 68.3±0.568.3\pm0.5
    8 82.7±0.582.7\pm0.5 67.9±0.367.9\pm0.3
    16 82.3±0.482.3\pm0.4 67.4±0.467.4\pm0.4

Coverage note — No substantial contributed method or experimental finding was deliberately omitted; detailed baseline-specific hyperparameter grids and implementation minutiae were left out because they support reproduction rather than constitute distinct contributions.

References

  1. 1.Ahn, H., Kwak, J., Lim, S., Bang, H., Kim, H., and Moon, T. Ss-il: Separated softmax for incremental learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 844–853, October 2021.
  2. 2.Aljundi, R., Chakravarty, P., and Tuytelaars, T. Expert gate: Lifelong learning with a network of experts. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  3. 3.Aljundi, R., Caccia, L., Belilovsky, E., Caccia, M., Charlin, L., and Tuytelaars, T. Online continual learning with maximally interfered retrieval. In Advances in Neural Information Processing (NeurIPS), 2019.
  4. 4.Asadi, N., Mudur, S., and Belilovsky, E. Tackling online one-class incremental learning by removing negative contrasts. arXiv preprint arXiv:2203.13307, 2022.
  5. 5.Barletti, T., Biondi, N., Pernici, F., Bruni, M., and Del Bimbo, A. Contrastive supervised distillation for continual representation learning. In International Conference on Image Analysis and Processing, pp. 597–609. Springer, 2022.
  6. 6.Boudiaf, M., Rony, J., Ziko, I. M., Granger, E., Pedersoli, M., Piantanida, P., and Ayed, I. B. A unifying mutual information view of metric learning: cross-entropy vs. pairwise losses. In European conference on computer vision, pp. 548–564. Springer, 2020.
  7. 7.Caccia, L., Aljundi, R., Asadi, N., Tuytelaars, T., Pineau, J., and Belilovsky, E. New insights on reducing abrupt representation change in online continual learning. arXiv preprint arXiv:2203.03798, 2022.
  8. 8.Caccia, M., Rodriguez, P., Ostapenko, O., Normandin, F., Lin, M., Caccia, L., Laradji, I., Rish, I., Lacoste, A., Vazquez, D., et al. Online fast adaptation and knowledge accumulation: a new approach to continual learning. arXiv preprint arXiv:2003.05856, 2020.
  9. 9.Cha, H., Lee, J., and Shin, J. Co2l: Contrastive continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9516–9525, 2021.
  10. 10.Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M. Efficient lifelong learning with a-gem. In ICLR 2019.
  11. 11.Chaudhry, A., Dokania, P. K., Ajanthan, T., and Torr, P. H. Riemannian walk for incremental learning: Understanding forgetting and intransigence. arXiv preprint arXiv:1801.10112, 2018.
  12. 12.Chaudhry, A., Rohrbach, M., Elhoseiny, M., Ajanthan, T., Dokania, P. K., Torr, P. H., and Ranzato, M. Continual learning with tiny episodic memories. arXiv preprint arXiv:1902.10486, 2019.
  13. 13.Chrabaszcz, P., Loshchilov, I., and Hutter, F. A downsampled variant of imagenet as an alternative to the cifar datasets. arXiv preprint arXiv:1707.08819, 2017.
  14. 14.Davari, M., Asadi, N., Mudur, S., Aljundi, R., and Belilovsky, E. Probing representation forgetting in supervised and unsupervised continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16712–16721, 2022.
  15. 15.De Lange, M. and Tuytelaars, T. Continual prototype evolution: Learning online from non-stationary data streams. arXiv preprint arXiv:2009.00919, 2020.
  16. 16.De Lange, M. and Tuytelaars, T. Continual prototype evolution: Learning online from non-stationary data streams. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8250–8259, 2021.
  17. 17.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  18. 18.Dohare, S., Mahmood, A. R., and Sutton, R. S. Continual backprop: Stochastic gradient descent with persistent randomness. arXiv preprint arXiv:2108.06325, 2021.
  19. 19.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  20. 20.Hinton, G., Vinyals, O., and Dean, J. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015.
  21. 21.Huszar, F. On quadratic penalties in elastic weight consolidation. arXiv preprint arXiv:1712.03847, 2017.
  22. 22.Javed, K. and Shafait, F. Revisiting distillation and incremental classifier learning. In Asian conference on computer vision, pp. 3–17. Springer, 2018.
  23. 23.Ji, X., Henriques, J., Tuytelaars, T., and Vedaldi, A. Automatic recall machines: Internal replay, continual learning and the brain. arXiv preprint arXiv:2006.12323, 2020.
  24. 24.Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., and Krishnan, D. Supervised contrastive learning. Advances in Neural Information Processing Systems, 33:18661–18673, 2020.
  25. 25.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. Overcoming catastrophic forgetting in neural networks. arXiv preprint arXiv:1612.00796, 2016.
  26. 26.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  27. 27.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097–1105, 2012.
  28. 28.Li, X., Zhou, Y., Wu, T., Socher, R., and Xiong, C. Learn to grow: A continual structure learning framework for overcoming catastrophic forgetting. In International Conference on Machine Learning, pp. 3925–3934. PMLR, 2019.
  29. 29.Li, Z. and Hoiem, D. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):2935–2947, 2017.
  30. 30.Lopez-Paz, D. et al. Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems, pp. 6467–6476, 2017.
  31. 31.McCloskey, M. and Cohen, N. J. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of learning and motivation, 24:109–165, 1989.
  32. 32.Mermillod, M., Bugaiska, A., and Bonin, P. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013.
  33. 33.Myung, I. J. Tutorial on maximum likelihood estimation. Journal of mathematical Psychology, 47(1):90–100, 2003.
  34. 34.Park, W., Kim, D., Lu, Y., and Cho, M. Relational knowledge distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3967–3976, 2019.
  35. 35.Rebuffi, S.-A., Kolesnikov, A., Sperl, G., and Lampert, C. H. icarl: Incremental classifier and representation learning. In Proc. CVPR, 2017.
  36. 36.Rosenfeld, A. and Tsotsos, J. K. Incremental learning through deep adaptation. IEEE transactions on pattern analysis and machine intelligence, 42(3):651–663, 2018.
  37. 37.Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  38. 38.Tian, Y., Krishnan, D., and Isola, P. Contrastive representation distillation. arXiv preprint arXiv:1910.10699, 2019.
  39. 39.Verwimp, E., Yang, K., Parisot, S., Lanqing, H., McDonagh, S., Perez-Pellitero, E., De Lange, M., and Tuytelaars, T. Clad: A realistic continual learning benchmark for autonomous driving. arXiv preprint arXiv:2210.03482, 2022.
  40. 40.Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al. Matching networks for one shot learning. Advances in neural information processing systems, 29, 2016.
  41. 41.Wu, G., Gong, S., and Li, P. Striking a balance between stability and plasticity for class-incremental learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1124–1133, 2021.
  42. 42.Yu, L., Twardowski, B., Liu, X., Herranz, L., Wang, K., Cheng, Y., Jui, S., and Weijer, J. v. d. Semantic drift compensation for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6982–6991, 2020.
  43. 43.Zhu, F., Zhang, X.-Y., Wang, C., Yin, F., and Liu, C.-L. Prototype augmentation and self-supervision for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5871–5880, 2021a.
  44. 44.Zhu, J., Tang, S., Chen, D., Yu, S., Liu, Y., Rong, M., Yang, A., and Wang, X. Complementary relation contrastive distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9260–9269, 2021b.

Citation

MLA
Asadi, N., et al. “Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 1093–106, https://proceedings.mlr.press/v202/asadi23a.html.
APA
Asadi, N., Davari, M., Mudur, S., Aljundi, R., & Belilovsky, E. (2023). Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning. International Conference on Machine Learning, 202, 1093–1106. https://proceedings.mlr.press/v202/asadi23a.html
Chicago
Asadi, N., M. Davari, S. Mudur, R. Aljundi, and E. Belilovsky. 2023. “Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning”. International Conference on Machine Learning 202: 1093–1106. https://proceedings.mlr.press/v202/asadi23a.html.
Harvard
Asadi, N. et al. (2023) “Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning”, International Conference on Machine Learning. PMLR, pp. 1093–1106. Available at: https://proceedings.mlr.press/v202/asadi23a.html.
Vancouver
1. Asadi N, Davari M, Mudur S, Aljundi R, Belilovsky E (2023) Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning. In: International Conference on Machine Learning. PMLR, pp 1093–1106

BibTeX

@InProceedings{pmlr-v202-asadi23a,
  title = 	 {Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning},
  author =       {Asadi, Nader and Davari, Mohammadreza and Mudur, Sudhir and Aljundi, Rahaf and Belilovsky, Eugene},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {1093--1106},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/asadi23a/asadi23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/asadi23a.html},
  abstract = 	 {In Continual learning (CL) balancing effective adaptation while combating catastrophic forgetting is a central challenge. Many of the recent best-performing methods utilize various forms of prior task data, e.g. a replay buffer, to tackle the catastrophic forgetting problem. Having access to previous task data can be restrictive in many real-world scenarios, for example when task data is sensitive or proprietary. To overcome the necessity of using previous tasks’ data, in this work, we start with strong representation learning methods that have been shown to be less prone to forgetting. We propose a holistic approach to jointly learn the representation and class prototypes while maintaining the relevance of old class prototypes and their embedded similarities. Specifically, samples are mapped to an embedding space where the representations are learned using a supervised contrastive loss. Class prototypes are evolved continually in the same latent space, enabling learning and prediction at any point. To continually adapt the prototypes without keeping any prior task data, we propose a novel distillation loss that constrains class prototypes to maintain relative similarities as compared to new task data. This method yields state-of-the-art performance in the task-incremental setting, outperforming methods relying on large amounts of data, and provides strong performance in the class-incremental setting without using any stored data points.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/