Built independently by an author, for readers. Read the story and support ChapterPal

keyword

prototype-sample relation distillation

Prototype-sample relation distillation is a machine learning technique used primarily in continual learning to prevent catastrophic forgetting by preserving the relational similarities between data sample embeddings and class prototypes across sequentially learned tasks. Rather than storing and replaying historical training data, this method evaluates the geometric relationships and relative distances between incoming data points and existing class representations within a shared feature space, enforcing consistency in these similarity distributions as the model updates. By transferring and constraining the relational structure between prototypes and samples across model iterations, the technique preserves previously established decision boundaries and class relationships while enabling the model to incorporate new classes without retaining past task samples.

1 item

Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning

Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning

Nader Asadi, MohammadReza Davari, Sudhir P. Mudur, Rahaf Aljundi, Eugene Belilovsky

OrganizationsConcordia UniversityMila – Québec Artificial Intelligence InstituteToyota Motor Europe

Why you should read this

Proposes a replay-free continual learning framework that prevents catastrophic forgetting by distilling the relative similarities between incoming samples and past class prototypes, outperforming memory-buffer methods without storing historical data.

In Continual learning (CL) balancing effective adaptation while combating catastrophic forgetting is a central challenge. Many of the recent best-performing methods utilize various forms of prior task data, e.g. a replay buffer, to tackle the catastrophic forgetting problem. Having access to previous task data can be restrictive in many real-world scenarios, for example when task data is sensitive or proprietary. To overcome the necessity of using previous tasks' data, in this work, we start with strong representation learning methods that have been shown to be less prone to forgetting. We propose a holistic approach to jointly learn the representation and class prototypes while maintaining the relevance of old class prototypes and their embedded similarities. Specifically, samples are mapped to an embedding space where the representations are learned using a supervised contrastive loss. Class prototypes are evolved continually in the same latent space, enabling learning and prediction at any point. To continually adapt the prototypes without keeping any prior task data, we propose a novel distillation loss that constrains class prototypes to maintain relative similarities as compared to new task data. This method yields state-of-the-art performance in the task-incremental setting, outperforming methods relying on large amounts of data, and provides strong performance in the class-incremental setting without using any stored data points.

Added

2026-10-03