Self-Supervised Models are Continual Learners

Enrico FiniVictor G. Turrisi da CostaXavier Alameda-PinedaElisa RicciKarteek AlahariJulien Mairal

article2022CVPR244 citations

Proposes CaSSLe, a general framework that mitigates catastrophic forgetting in continual self-supervised learning by using a predictor network to convert standard self-supervised loss functions into representation distillation mechanisms without requiring extra hyperparameter tuning.

Listen

Modern computer vision increasingly relies on self-supervised learning, a technique that trains artificial intelligence models on massive amounts of raw, unlabeled image data. While these models match or exceed the performance of models trained with expensive human annotations when trained in a single offline phase, they degrade catastrophically when exposed to continuous streams of new data over time. In real-world applications where data arrives progressively, retraining models from scratch on the entire accumulated dataset is computationally prohibitive, costly, and often impossible due to data retention limits. Standard continual learning techniques designed to prevent this forgetting typically depend heavily on human labels, leaving sequential unsupervised learning largely unsolved.

The article introduces and evaluates "CaSSLe," a framework designed to enable continual self-supervised visual representation learning across sequential tasks without requiring human annotations or extensive parameter tuning.

The authors develop a distillation architecture where a small helper prediction network projects the model's current internal representations back into the feature space of its previous, frozen state. Rather than introducing artificial regularization terms, the framework repurposes the underlying self-supervised learning objectives as distillation mechanisms. The approach was evaluated across six major self-supervised learning methods on small, medium, and large image datasets (CIFAR100, ImageNet100, and DomainNet) across three sequential learning scenarios: learning new classes over time, processing batches of new images, and handling shifts in image style or domain.

Across all experiments, the framework substantially outperformed existing continual learning baselines and standard sequential fine-tuning. First, in class-incremental benchmarks, the framework improved representation accuracy by an average of 6.8% on CIFAR100 and 4.0% on ImageNet100 compared to standard sequential fine-tuning, successfully matching or outperforming supervised fine-tuning. Second, in domain-shift environments, it improved model accuracy by 4.4% on average across all tested methods, bringing sequential performance close to the upper-bound benchmark of full offline training. Third, representations learned sequentially with the framework transferred better to completely unseen downstream tasks and performed robustly in semi-supervised settings with as little as 1% to 10% labeled evaluation data. Traditional replay methods, which store exemplars in memory buffers, proved ineffective in this setup because the prolonged training required by self-supervised models caused severe overfitting on the stored samples.

These findings demonstrate that self-supervised vision models possess an inherent capacity for sequential learning when equipped with appropriate temporal projection mechanisms. For organizations deploying computer vision, this framework offers a practical path to continuously update models on streaming unlabeled data, sharply reducing data annotation expenses and avoiding the high computational costs of retraining models from scratch. Moreover, the self-supervised approach eliminates the reliance on task-specific labels, yielding general-purpose visual representations that adapt effectively to shifting data environments.

Organizations seeking to implement continuous model updates should adopt temporal prediction architectures rather than buffer-based sample replay. Technical teams can implement the framework directly alongside existing self-supervised pipelines, as it requires no extra hyperparameter tuning. Prior to deployment in production environments, teams should evaluate computational trade-offs, since the dual-network distillation process increases per-task training time and memory requirements by approximately 30%. Furthermore, because the framework relies on clearly defined task boundaries, further research is required to adapt the methodology to unstructured, continuous data streams where task transitions are unannounced.

While confidence in the reported experimental improvements is high across multiple benchmarks, practitioners should remain cautious regarding inherent biases in uncurated streaming data, which can propagate into downstream applications without human oversight. Additionally, because the framework generates general feature extractors rather than direct category classifiers, deploying teams must retain a small labeled dataset or a clustering mechanism to map learned visual features to end-user business categories.

Cover for Self-Supervised Models are Continual Learners

Abstract

Self-supervised models have been shown to produce comparable or better visual representations than their supervised counterparts when trained offline on unlabeled data at scale. However, their efficacy is catastrophically reduced in a Continual Learning (CL) scenario where data is presented to the model sequentially. In this paper, we show that self-supervised loss functions can be seamlessly converted into distillation mechanisms for CL by adding a predictor network that maps the current state of the representations to their past state. This enables us to devise a framework for Continual self-supervised visual representation Learning that (i) significantly improves the quality of the learned representations, (ii) is compatible with several state-of-the-art self-supervised objectives, and (iii) needs little to no hyperparameter tuning. We demonstrate the effectiveness of our approach empirically by training six popular self-supervised models in various CL settings. Code: github.com/DonkeyShot21/cassle.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Continual Self-Supervised Learning
  • 5. The CaSSLe Framework
  • 5.1. Compatibility of SSL methods with CaSSLe
  • 6. Experiments
  • 6.1. Experimental Protocol
  • 6.2. Results
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Continual self-supervised learning problem

    definition

    Continual Self-Supervised Learning (CSSL) trains a visual encoder on an ordered sequence of unlabeled task distributions D1,…,DT\mathcal{D}_1,\ldots,\mathcal{D}_T. The model receives task tt only after completing task t−1t-1, discards data from earlier tasks, receives no class labels, and is assumed to know the task boundaries. For an image x∼Dtx\sim\mathcal{D}_t, stochastic augmentations produce two views xAx^A and xBx^B, and an encoder fθf_\theta produces representations zA=fθ(xA)z^A=f_\theta(x^A) and zB=fθ(xB)z^B=f_\theta(x^B). The CSSL objective is

    argmin⁡θ  ∑t=1TEx∼Dt[LSSL(zA,zB)],\underset{\theta}{\operatorname{argmin}}\;\sum_{t=1}^{T}\mathbb{E}_{x\sim\mathcal{D}_t}\left[\mathcal{L}_{SSL}(z^A,z^B)\right],

    where LSSL\mathcal{L}_{SSL} is the native self-supervised loss of the chosen method. The three evaluated task organizations are: class-incremental, in which tasks contain disjoint class sets; data-incremental, in which tasks contain disjoint image sets but may contain all classes; and domain-incremental, in which every task contains the same classes but comes from a different image domain.

  2. Knowl 2 — CaSSLe temporal distillation framework

    model/method

    CaSSLe converts a self-supervised objective into a continual-learning distillation mechanism. When task tt begins, the current encoder is copied to a frozen encoder ft−1f^{t-1}, while the trainable encoder is denoted ftf^t. For an augmented view, the current encoder produces z=ft(x)z=f^t(x) and the frozen encoder produces zˉ=ft−1(x)\bar z=f^{t-1}(x); the frozen representation is detached so that no gradient reaches ft−1f^{t-1}. A trainable predictor gg maps the current representation into the previous representation space. The generic distillation loss is

    LD(z,zˉ)=LSSL(g(z),zˉ).\mathcal{L}_D(z,\bar z)=\mathcal{L}_{SSL}(g(z),\bar z).

    The predictor is optimized so that the current representation retains enough information to reproduce the old representation without forcing the current encoder itself to remain unchanged. For one pair of views, CaSSLe trains with

    L=LSSL(zA,zB)+LSSL(g(zA),zˉA).\mathcal{L}=\mathcal{L}_{SSL}(z^A,z^B)+\mathcal{L}_{SSL}\bigl(g(z^A),\bar z^A\bigr).

    The loss can be symmetrized by adding the corresponding term for view BB, and it can be extended to multi-crop training. CaSSLe does not introduce a coefficient weighting the distillation term against the original self-supervised term and requires no additional hyperparameter tuning relative to the underlying self-supervised method.

  3. Knowl 3 — Reuse of diverse self-supervised losses

    model/method

    CaSSLe is compatible with multiple families of self-supervised objectives because the distillation term reuses the chosen method's own loss rather than imposing a separate feature-matching loss. For InfoNCE methods such as SimCLR and MoCoV2+, the predictor learns to discriminate current-task instances in the feature space established by the frozen encoder, preserving similarity to the corresponding past representation and separation from negatives. For MSE-based methods such as BYOL, SimSiam, and VICReg, the predictor learns consistency with the frozen representation; BYOL and SimSiam use normalized representations, whereas VICReg does not and uses its native regularization structure. For cross-entropy methods such as SwAV, the frozen encoder produces assignments to frozen cluster prototypes, and the predictor learns to reproduce those assignments. For cross-correlation methods such as Barlow Twins, the same cross-correlation objective is applied between predicted current features and frozen past features, which also encourages decorrelation among dimensions of the predicted representation. This loss matching reduces interference between the continual-learning mechanism and method-specific normalization or anti-collapse mechanisms.

  4. Knowl 4 — Evaluation protocol and experimental scope

    experimental setup

    The experiments use linear evaluation: after each task, a linear classifier is trained on the learned backbone and its top-1 accuracy is measured. Let Aj,kA_{j,k} be the accuracy on task kk after the model has observed task jj, let TT be the number of tasks, and let RiR_i be the accuracy of a random network on task ii. The reported average accuracy, forgetting, and forward transfer are

    A=1T∑i=1TAT,i,F=1T−1∑i=1T−1max⁡t∈{1,…,T}(At,i−AT,i),A=\frac{1}{T}\sum_{i=1}^{T}A_{T,i},\qquad F=\frac{1}{T-1}\sum_{i=1}^{T-1}\max_{t\in\{1,\ldots,T\}}(A_{t,i}-A_{T,i}), FT=1T−1∑i=2T(Ai−1,i−Ri).FT=\frac{1}{T-1}\sum_{i=2}^{T}(A_{i-1,i}-R_i).

    Class- and data-incremental evaluations use task-agnostic classification, while domain-incremental evaluation primarily reports task-aware accuracy. Experiments use CIFAR100 with 100 classes and 60,000 32×3232\times32 color images, ImageNet100 with approximately 130,000 high-resolution images resized to 224×224224\times224, and DomainNet with 345 classes, approximately 600,000 images, and six domains. Class- and data-incremental experiments use five tasks; domain-incremental experiments use six tasks. The backbone is ResNet-18, batch size is 256, optimization uses LARS, and each task is trained for 500 epochs on CIFAR100, 400 epochs on ImageNet100, or 200 epochs on DomainNet. CaSSLe assumes known task boundaries, increases training memory and time by roughly 30%, and does not itself learn a data-to-class mapping; supervised linear evaluation or a separate clustering method is still required.

  5. Knowl 5 — Comparison with continual-learning baselines

    empirical result

    On class-incremental CIFAR100 with five tasks, CaSSLe substantially outperforms fine-tuning and the evaluated continual-learning baselines when combined with SimCLR, Barlow Twins, or BYOL. Each tuple reports average accuracy AA, forgetting FF, and forward transfer FTFT in percentage points; higher AA and FTFT and lower FF are better.

    SimCLR results were: fine-tuning (48.9,1.0,33.5)(48.9,1.0,33.5), EWC (53.6,0.0,33.3)(53.6,0.0,33.3), ER (50.3,0.1,32.7)(50.3,0.1,32.7), DER (50.7,0.4,33.2)(50.7,0.4,33.2), LUMP (52.3,0.3,34.5)(52.3,0.3,34.5), Less-Forget (52.5,0.2,33.8)(52.5,0.2,33.8), POD (51.3,0.1,33.8)(51.3,0.1,33.8), and CaSSLe (58.3,0.2,36.4)(58.3,0.2,36.4); the offline upper bound was 65.865.8 accuracy. Barlow Twins results were: fine-tuning (54.3,0.4,39.2)(54.3,0.4,39.2), EWC (56.7,0.2,39.1)(56.7,0.2,39.1), ER (54.6,3.0,39.4)(54.6,3.0,39.4), DER (55.3,2.5,39.6)(55.3,2.5,39.6), LUMP (57.8,0.3,41.0)(57.8,0.3,41.0), Less-Forget (56.4,0.2,40.1)(56.4,0.2,40.1), POD (55.9,0.3,40.3)(55.9,0.3,40.3), and CaSSLe (60.4,0.4,42.2)(60.4,0.4,42.2); the offline upper bound was 70.970.9. BYOL results were: fine-tuning (52.7,0.1,35.9)(52.7,0.1,35.9), EWC (56.4,0.0,39.9)(56.4,0.0,39.9), ER (54.7,0.4,36.3)(54.7,0.4,36.3), DER (54.8,1.1,36.7)(54.8,1.1,36.7), LUMP (56.4,0.2,37.9)(56.4,0.2,37.9), Less-Forget (58.6,0.2,41.1)(58.6,0.2,41.1), POD (57.9,0.0,41.1)(57.9,0.0,41.1), and CaSSLe (62.2,0.0,43.6)(62.2,0.0,43.6); the offline upper bound was 70.570.5.

    The largest gains are in representation accuracy and forward transfer rather than forgetting, because self-supervised models already forget relatively little on CIFAR100. Replay methods ER and DER do not improve CSSL substantially, whereas EWC and feature-distillation baselines are stronger but remain below CaSSLe.

  6. Knowl 6 — Class-incremental performance across six self-supervised methods

    data/table

    On five-task class-incremental CIFAR100 and ImageNet100, CaSSLe improves every evaluated self-supervised method over ordinary fine-tuning. Each tuple is (A,F,FT)(A,F,FT), with accuracy, forgetting, and forward transfer reported in percentage points.

    For CIFAR100, the results are: Barlow Twins fine-tuning (54.3,0.4,39.2)(54.3,0.4,39.2) and CaSSLe (60.4,0.4,42.2)(60.4,0.4,42.2), with offline accuracy 70.970.9; SwAV fine-tuning (55.5,0.0,32.8)(55.5,0.0,32.8) and CaSSLe (57.8,0.0,34.5)(57.8,0.0,34.5), with offline accuracy 64.964.9; BYOL fine-tuning (52.7,0.1,35.9)(52.7,0.1,35.9) and CaSSLe (62.2,0.0,42.2)(62.2,0.0,42.2), with offline accuracy 70.570.5; VICReg fine-tuning (51.5,0.9,36.4)(51.5,0.9,36.4) and CaSSLe (53.6,0.2,41.1)(53.6,0.2,41.1), with offline accuracy 68.568.5; MoCoV2+ fine-tuning (47.3,0.2,33.4)(47.3,0.2,33.4) and CaSSLe (59.5,0.0,39.6)(59.5,0.0,39.6), with offline accuracy 69.969.9; and SimCLR fine-tuning (48.9,1.0,33.5)(48.9,1.0,33.5) and CaSSLe (58.3,0.2,36.4)(58.3,0.2,36.4), with offline accuracy 65.865.8. Supervised fine-tuning obtained (54.1,6.8,36.5)(54.1,6.8,36.5) and supervised offline training obtained accuracy 75.675.6.

    For ImageNet100, Barlow Twins improved from fine-tuning (63.1,10.7,44.4)(63.1,10.7,44.4) to CaSSLe (68.2,1.3,47.9)(68.2,1.3,47.9), with offline accuracy 80.480.4; SwAV improved from (64.4,4.3,42.8)(64.4,4.3,42.8) to (66.0,0.2,43.6)(66.0,0.2,43.6), with offline accuracy 74.374.3; BYOL improved from (66.0,2.9,43.2)(66.0,2.9,43.2) to (66.4,1.1,46.6)(66.4,1.1,46.6), with offline accuracy 80.380.3; VICReg improved from (61.3,7.9,42.0)(61.3,7.9,42.0) to (64.8,4.3,45.3)(64.8,4.3,45.3), with offline accuracy 79.479.4; MoCoV2+ improved from (62.0,8.4,41.6)(62.0,8.4,41.6) to (68.8,1.5,46.8)(68.8,1.5,46.8), with offline accuracy 79.379.3; and SimCLR improved from (61.5,8.1,40.3)(61.5,8.1,40.3) to (68.0,2.2,45.8)(68.0,2.2,45.8), with offline accuracy 77.577.5. Supervised fine-tuning obtained (63.1,5.6,42.5)(63.1,5.6,42.5) and supervised offline training obtained accuracy 81.981.9. Averaged over the six self-supervised methods, CaSSLe raises accuracy by about 6.86.8 points on CIFAR100 and 44 points on ImageNet100 relative to fine-tuning, while particularly reducing forgetting on ImageNet100.

  7. Knowl 7 — Data- and domain-incremental results

    data/table

    CaSSLe remains effective when tasks differ by images or domains rather than by disjoint classes. On five-task data-incremental ImageNet100, the linear-evaluation accuracies for fine-tuning, CaSSLe, and offline training were respectively: Barlow Twins 71.371.3, 74.974.9, and 80.480.4; SwAV 70.870.8, 71.371.3, and 74.374.3; BYOL 74.074.0, 73.373.3, and 80.380.3; VICReg 70.270.2, 72.372.3, and 79.479.4; MoCoV2+ 69.569.5, 71.971.9, and 78.278.2; and SimCLR 68.968.9, 72.172.1, and 77.577.5. Supervised fine-tuning achieved 75.975.9 and supervised offline training achieved 81.981.9. CaSSLe improves the self-supervised methods by about 22 points on average except for BYOL, whose momentum encoder already preserves some information from earlier data.

    On six-task domain-incremental DomainNet, the corresponding accuracies were: Barlow Twins 50.350.3, 55.555.5, and 57.257.2; SwAV 49.649.6, 54.354.3, and 54.654.6; BYOL 50.650.6, 55.155.1, and 56.656.6; VICReg 49.349.3, 52.952.9, and 56.756.7; MoCoV2+ 43.243.2, 46.746.7, and 53.753.7; and SimCLR 45.145.1, 50.050.0, and 52.652.6. Supervised fine-tuning achieved 55.955.9 and supervised offline training achieved 66.466.4. CaSSLe improves all six self-supervised methods by 4.44.4 points on average and brings most of them close to their offline performance despite the domain shifts.

  8. Knowl 8 — Predictor and view-ablation findings

    empirical result

    The predictor gg is essential to CaSSLe's performance, whereas distilling swapped views is not beneficial in the evaluated configuration. On five-task class-incremental CIFAR100, the average accuracies for the swapped-view variant, the variant without a predictor, and full CaSSLe were respectively: SimCLR 49.349.3, 52.652.6, and 58.358.3; Barlow Twins 57.457.4, 57.357.3, and 60.460.4; and BYOL 52.052.0, 58.658.6, and 62.262.2. The results support mapping current features into the frozen encoder's feature space instead of directly forcing current and past representations to coincide. The authors attribute the lack of benefit from swapped-view distillation to the frozen encoder not necessarily being invariant to the current task's views.

    Against a contrastive continual-learning method, CaSSLe also obtained higher CIFAR100 class-incremental accuracy. With two tasks, Lin et al.'s method reached 55.755.7 for SimCLR and 56.156.1 for MoCoV2, whereas CaSSLe reached 61.861.8 and 63.363.3 for SimCLR and MoCoV2+, respectively. With five tasks, Lin et al.'s MoCoV2 result was 53.853.8, while CaSSLe obtained 58.358.3 with SimCLR and 59.559.5 with MoCoV2+.

  9. Knowl 9 — Downstream and semi-supervised transfer

    empirical result

    Representations trained continually on ImageNet100 transfer to DomainNet's Real domain. Linear downstream accuracy for fine-tuning versus CaSSLe was respectively: Barlow Twins 56.256.2 versus 60.360.3, SwAV 55.955.9 versus 56.956.9, BYOL 55.055.0 versus 56.956.9, VICReg 54.054.0 versus 56.356.3, MoCoV2+ 52.452.4 versus 58.758.7, and SimCLR 51.651.6 versus 56.556.5. The supervised fine-tuning baseline was 54.354.3. CaSSLe improves downstream accuracy by 3.43.4 points on average over self-supervised fine-tuning, and every CaSSLe representation exceeds the supervised baseline.

    With a frozen backbone and only 10%10\% of ImageNet100 labels, fine-tuning accuracies for Barlow Twins, SwAV, BYOL, VICReg, MoCoV2+, and SimCLR were 56.656.6, 57.657.6, 55.755.7, 53.653.6, 54.954.9, and 52.552.5, while CaSSLe gave 60.360.3, 58.258.2, 56.556.5, 56.556.5, 61.761.7, and 58.958.9. The supervised baseline was 60.860.8, so CaSSLe MoCoV2+ exceeded it. With 1%1\% of labels, fine-tuning gave 42.642.6, 42.542.5, 42.342.3, 40.440.4, 40.940.9, and 39.739.7, whereas CaSSLe gave 47.047.0, 43.143.1, 43.443.4, 43.243.2, 47.847.8, and 46.846.8; the supervised baseline was 48.148.1. Thus CaSSLe consistently improves the self-supervised representations under limited-label evaluation, although it does not universally surpass fully supervised learning at 1%1\% labels.

  10. Knowl 10 — Continual training versus longer training on less data

    empirical result

    The study compares continual training on five tasks with training five times longer on only one-fifth of the data. On class-incremental ImageNet100, the accuracies for ordinary fine-tuning, offline training on one-fifth of the data, and CaSSLe were respectively: SimCLR 61.561.5, 63.163.1, and 68.068.0; Barlow Twins 63.163.1, 63.563.5, and 68.268.2; and BYOL 66.066.0, 60.660.6, and 66.466.4. On data-incremental ImageNet100, the corresponding values were: SimCLR 68.968.9, 67.267.2, and 72.172.1; Barlow Twins 71.371.3, 70.270.2, and 74.974.9; and BYOL 74.074.0, 66.766.7, and 73.373.3.

    The comparison shows that the preferred training strategy depends on the self-supervised method and continual-learning setting. Offline training on a smaller class subset can outperform ordinary fine-tuning for some class-incremental methods, while fine-tuning can outperform longer training on a smaller dataset in the data-incremental setting, especially for BYOL. CaSSLe gives the best result among the listed strategies except for BYOL in the data-incremental case, where ordinary fine-tuning is higher.

Coverage note — Supplementary-only results for alternate task counts and domain-agnostic evaluation, detailed baseline implementation choices, and broader-impact discussion were omitted because they do not add load-bearing contributed methods or main-paper findings.

References

  1. 1.Alessandro Achille, Tom Eccles, Loic Matthey, Christopher P Burgess, Nick Watters, Alexander Lerchner, and Irina Higgins. Life-long disentangled representation learning with cross-domain latent homologies. NeurIPS, 2018. 2, 6
  2. 2.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. Memory aware synapses: Learning what (not) to forget. In ECCV, 2018. 2
  3. 3.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv preprint arXiv:2105.04906, 2021. 1, 2, 5
  4. 4.Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. Dark experience for general continual learning: a strong, simple baseline. In NeurIPS, 2020. 2, 6
  5. 5.Lucas Caccia and Joelle Pineau. Special: Self-supervised pretraining for continual learning. arXiv preprint arXiv:2106.09065, 2021. 2
  6. 6.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, pages 132–149, 2018. 2
  7. 7.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. NeurIPS, 2020. 1, 2, 4, 5
  8. 8.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. ICCV, 2021. 1, 2, 5
  9. 9.Francisco M Castro, Manuel J Marin-Jimenez, Nicolas Guil, Cordelia Schmid, and Karteek Alahari. End-to-end incremental learning. In ECCV, 2018. 2
  10. 10.Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. Co2l: Contrastive continual learning. In CVPR, pages 9516–9525, 2021. 2
  11. 11.Arslan Chaudhry, Puneet K Dokania, Thalaiyasingam Ajanthan, and Philip HS Torr. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In ECCV, 2018. 2
  12. 12.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. Efficient lifelong learning with a-gem. In ICLR, 2019. 2
  13. 13.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 1, 2, 4, 5
  14. 14.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020. 1, 2
  15. 15.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In ICCV, 2021. 2, 4, 5, 6
  16. 16.Victor G. Turrisi da Costa, Enrico Fini, Moin Nabi, Nicu Sebe, and Elisa Ricci. Solo-learn: A library of self-supervised methods for visual representation learning. arXiv 2108.01775, 2021. 5, 6
  17. 17.Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. Continual learning: A comparative study on how to defy forgetting in classification tasks. IEEE TPAMI, 2(6), 2019. 1, 2, 6
  18. 18.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In ECCV, 2020. 2, 3, 6
  19. 19.Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman. With a little help from my friends: Nearest-neighbor contrastive learning of visual representations. ICCV, 2021. 2, 5
  20. 20.Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, and Nicu Sebe. Whitening for self-supervised representation learning. In ICML, 2021. 2, 5
  21. 21.Enrico Fini, Stephane Lathuilière, Enver Sangineto, Moin Nabi, and Elisa Ricci. Online continual learning under extreme memory constraints. In ECCV, 2020. 2
  22. 22.Robert M. French. Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4):128–135, 1999. 1
  23. 23.Jhair Gallardo, Tyler L Hayes, and Christopher Kanan. Self-supervised training enhances online continual learning. arXiv preprint arXiv:2103.14010, 2021. 2
  24. 24.I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio. An empirical investigation of catastrophic forgetting in gradient-based neural networks. arXiv preprint arXiv:1312.6211, 2013. 1
  25. 25.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. NeurIPS, 2020. 1, 2, 4, 5
  26. 26.Michael Gutmann and Aapo Hyvarinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In AISTATS, 2010. 2
  27. 27.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020. 1, 2, 4, 5
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 5
  29. 29.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 3
  30. 30.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In CVPR, pages 831–839, 2019. 2, 3, 6
  31. 31.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proc. of the national academy of sciences, 2017. 2, 6
  32. 32.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, Univ. Toronto, 2009. 5
  33. 33.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE TPAMI, 2017. 2, 3
  34. 34.Zhiwei Lin, Yongtao Wang, and Hongxiang Lin. Continual contrastive self-supervised learning for image classification. arXiv preprint arXiv:2107.01776, 2021. 2, 6, 7
  35. 35.David Lopez-Paz and Marc-Aurelio Ranzato. Gradient episodic memory for continual learning. In NeurIPS, 2017. 2, 5
  36. 36.Divyam Madaan, Jaehong Yoon, Yuanchun Li, Yunxin Liu, and Sung Ju Hwang. Rethinking the representational continuity: Towards unsupervised continual learning. arXiv preprint arXiv:2110.06976, 2021. 2, 6
  37. 37.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pages 109–165. Elsevier, 1989. 1
  38. 38.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 2
  39. 39.Oleksiy Ostapenko, Mihai Puscas, Tassilo Klein, Patrick Jähnichen, and Moin Nabi. Learning to remember: A synaptic plasticity driven framework for continual learning. In CVPR, 2019. 2
  40. 40.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In ICCV, 2019. 5
  41. 41.Ameya Prabhu, Philip HS Torr, and Puneet K Dokania. Gdumb: A simple approach that questions our progress in continual learning. In ECCV, pages 524–540. Springer, 2020. 2
  42. 42.Dushyant Rao, Francesco Visin, Andrei A Rusu, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. Continual unsupervised representation learning. NeurIPS, 2019. 2, 6
  43. 43.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In CVPR, 2017. 2
  44. 44.Anthony Robins. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science, 1995. 2, 6
  45. 45.A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell. Progressive neural networks. arXiv 1606.04671, 2016. 2
  46. 46.Joan Serra, Didac Suris, Marius Miron, and Alexandros Karatzoglou. Overcoming catastrophic forgetting with hard attention to the task. ICML, 2018. 2
  47. 47.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. Continual learning with deep generative replay. In NeurIPS, 2017. 2
  48. 48.James Smith, Cameron Taylor, Seth Baer, and Constantine Dovrolis. Unsupervised progressive learning and the stam architecture. arXiv preprint arXiv:1904.02021, 2019. 2
  49. 49.Yonglong Tian, Olivier J Henaff, and Aaron van den Oord. Divide and contrast: Self-supervised learning from uncurated data. arXiv preprint arXiv:2105.08054, 2021. 2
  50. 50.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In ECCV, pages 776–794. Springer, 2020. 5
  51. 51.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In CVPR, 2019. 2
  52. 52.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, 2018. 2
  53. 53.Yang You, Igor Gitman, and Boris Ginsburg. Large batch training of convolutional networks. arXiv preprint arXiv:1708.03888, 2017. 5
  54. 54.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. ICML, 2021. 1, 2, 5
  55. 55.Friedeman Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In ICML, 2017. 2

Citation

MLA
Fini, E., et al. “Self-Supervised Models Are Continual Learners”. arXiv, 2021, http://arxiv.org/abs/2112.04215v2.
APA
Fini, E., Costa, V. G. T. da ., Alameda-Pineda, X., Ricci, E., Alahari, K., & Mairal, J. (2021). Self-Supervised Models are Continual Learners. arXiv. http://arxiv.org/abs/2112.04215v2
Chicago
Fini, E., V. G. T. da . Costa, X. Alameda-Pineda, E. Ricci, K. Alahari, and J. Mairal. 2021. “Self-Supervised Models Are Continual Learners”. arXiv. http://arxiv.org/abs/2112.04215v2.
Harvard
Fini, E. et al. (2021) “Self-Supervised Models are Continual Learners”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.04215v2.
Vancouver
1. Fini E, Costa VGT da, Alameda-Pineda X, Ricci E, Alahari K, Mairal J (2021) Self-Supervised Models are Continual Learners. arXiv

BibTeX

@article{fini2021self,
  title = {Self-Supervised Models are Continual Learners},
  author = {Fini, Enrico and Costa, Victor G. Turrisi da and Alameda-Pineda, Xavier and Ricci, Elisa and Alahari, Karteek and Mairal, Julien},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.04215v2},
  eprint = {2112.04215}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE