Prototype-Sample Relation Distillation: Towards Replay-Free Continual Learning
Nader AsadiMohammadReza DavariSudhir P. MudurRahaf AljundiEugene Belilovsky
Proposes a replay-free continual learning framework that prevents catastrophic forgetting by distilling the relative similarities between incoming samples and past class prototypes, outperforming memory-buffer methods without storing historical data.
Continual learning requires machine learning systems to adapt continuously to incoming streams of new data without suffering from catastrophic forgetting, which occurs when a model forgets previously acquired knowledge. The most effective existing methods generally rely on rehearsal buffers that store past data samples to preserve performance. However, storing historical data poses serious regulatory, privacy, and technical challenges when handling sensitive medical or proprietary information, while also creating escalating storage demands as task sequences grow.
The article evaluates a new framework designed to eliminate the need for storing prior task data during continual learning. Specifically, it demonstrates that joint representation learning and prototype adaptation can maintain high predictive accuracy across evolving tasks without requiring access to previous samples.
The researchers developed an approach termed Prototype-Sample Relation Distillation (PRD). The framework learns rich data representations via supervised contrastive learning and maintains class prototypes—vector representations of classes—without using suppressive contrastive terms. To prevent older class prototypes from becoming obsolete when the underlying model updates on new tasks, PRD applies a relation distillation loss that preserves the relative similarity rankings between old prototypes and incoming data samples. The authors evaluated PRD across multiple benchmark datasets, including Split-CIFAR100, Split-MiniImageNet, ImageNet32, and the real-world autonomous driving dataset CLAD-C, assessing both task-incremental and class-incremental configurations over short and long sequences reaching up to 200 tasks.
The evaluation yielded several key findings. First, in task-incremental settings across 20 to 200 tasks, PRD consistently outperformed existing replay-free baselines and matched or exceeded experience replay methods using 50 stored samples per class. Second, in class-incremental settings without any stored data, PRD achieved 27.8% accuracy on CIFAR-100 and 20.0% on MiniImageNet, outperforming experience replay baselines limited to 5 stored samples per class. Third, when provided with replay samples, PRD reached 45.1% accuracy on CIFAR-100, surpassing standard experience replay at 38.3%. Fourth, on the real-world CLAD-C benchmark exhibiting day-night shifts and severe class imbalance, PRD attained an Average Mean Class Accuracy of 65.1%, outperforming alternative replay-free techniques. Finally, ablation analyses confirmed that relative relation distillation is critical for maintaining stability on older classes while preserving plasticity to learn new tasks.
These findings indicate that organizations can train adaptable machine learning systems without retaining historical user data or intellectual property. This significantly reduces data governance risks, storage costs, and compliance overhead associated with long-term data retention. The results challenge the conventional belief that competitive continual learning strictly requires rehearsal buffers, proving that high plasticity and retention can be attained through relative geometric alignment in representation space.
Organizations developing machine learning models for streaming environments should consider prototype-relation frameworks when data retention is restricted by privacy policies or memory limits. In scenarios where data storage is permissible, hybrid implementations combining PRD with small buffers provide substantial performance gains over conventional replay methods. Teams should conduct domain-specific pilot testing on their streaming pipelines to fine-tune distillation coefficients for target sequence lengths.
Confidence in these findings is supported by rigorous benchmarking across diverse image datasets and consistent outperformance over established baselines. However, limitations remain: the article evaluated the method primarily on computer vision classification tasks using standard convolutional architectures, and initial accuracy drops were observed in class-incremental settings. Practitioners should exercise caution before generalizing these results to non-visual modalities, such as natural language or tabular data, without further empirical validation.
- Paper: Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning, Kai Zhu et al. (2022). Its exemplar-free class-incremental approach uses class prototypes and distillation to preserve prior knowledge, providing the closest precursor to PRD’s prototype-based retention strategy.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL establishes how class prototypes and distillation can support incremental classification, making PRD’s move from exemplar-based prototypes to relation-preserving ones easier to follow.
- Paper: Probing Representation Forgetting in Supervised and Unsupervised Continual Learning, MohammadReza Davari et al. (2022). Its analysis of representation retention under supervised contrastive continual learning helps explain why PRD combines contrastive representation learning with a separate mechanism for preserving past-class relations.
- Paper: New Insights on Reducing Abrupt Representation Change in Online Continual Learning, Lucas Caccia et al. (2022). Its account of representation drift during continual learning motivates the stability problem that PRD addresses by preserving old-prototype similarity rankings.
- Paper: Continual Learning Mechanisms Compose for Long-Horizon Memorization, Zheyuan Zhang et al. (2026). By testing combinations of replay, self-distillation, and regularization over long task sequences, it extends the continual-learning design space beyond PRD’s single relation-distillation mechanism.
