End-to-End Incremental Learning

Francisco M. CastroManuel J. Marín-JiménezNicolás GuilCordelia SchmidKarteek Alahari

article2018ECCV1,348 citations

Proposes a fully end-to-end incremental learning framework that mitigates catastrophic forgetting by combining knowledge distillation with a small exemplar set to jointly optimize feature representations and classifiers on sequentially added image classes.

Listen

Practical visual recognition systems frequently need to learn new categories over time without retraining from scratch on entire historical datasets. However, standard deep learning models suffer from catastrophic forgetting, where integrating new information causes a steep drop in accuracy on previously learned categories. Storing all historical data to retrain models is computationally expensive and quickly becomes unsustainable as the number of classes grows.

The main objective of the article is to demonstrate an end-to-end deep learning framework that learns new image classes incrementally while preserving past knowledge using only a minimal set of preserved reference images.

To accomplish this, the authors developed a convolutional neural network architecture trained using a combined objective called cross-distilled loss. This combines a cross-entropy loss to learn incoming categories with a distillation loss that transfers and locks in knowledge from older categories. The system retains a small representative memory of past categories selected via a mean-distance ranking method (herding). The process incorporates data augmentation and a balanced fine-tuning stage to correct for the numerical imbalance between new data and limited historical exemplars. The method was evaluated using ResNet architectures on standard image classification benchmarks (CIFAR-100 and ImageNet ILSVRC 2012) across varying incremental step sizes and memory configurations.

The framework achieved state-of-the-art results across major benchmarks. On CIFAR-100 with a fixed 2,000-exemplar memory, it reached an average incremental accuracy of roughly 63.4% to 63.8% across 2-to-20 class steps, significantly outperforming prior methods like iCaRL (54.1% to 62.0%) and LwF.MC (9.6% to 47.6%). On large-scale ImageNet benchmarks, the method attained top-5 accuracies of 90.4% (for 10-class steps) and 69.4% (for 100-class steps), beating previous techniques by over 5 percentage points. Furthermore, on fine-grained recognition tests with highly similar classes (such as distinguishing vehicle types or dog breeds), the end-to-end framework outperformed competing baselines by roughly 10 to 25 percentage points.

These findings demonstrate that jointly learning feature representations and classifiers avoids the performance bottlenecks of decoupled, prototype-based classifiers. By retaining only a small representative memory and adjusting for data imbalances during fine-tuning, organizations can continually expand visual recognition systems at stable memory footprints and reduced training overhead, making lifelong machine learning practical in real-world deployment.

Teams implementing incremental computer vision systems should adopt joint end-to-end training paired with distillation loss and balanced fine-tuning rather than freezing network layers or relying on external nearest-mean classifiers. Systems should utilize herding selection to manage exemplar storage. As next steps, the authors recommend exploring dynamic exemplar allocation to further optimize memory efficiency across classes.

The primary limitation of the approach arises when the number of retained past samples is severely restricted relative to large batches of new classes (e.g., adding 50 classes at once with a small fixed memory), which can degrade performance due to extreme data imbalance. Confidence in these results remains high across standard and fine-grained classification tasks, though practitioners should remain cautious in extreme class-imbalance scenarios without adequate fine-tuning.

arXiv: 1807.09536
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). It introduces the exemplar-based class-incremental learning framework with distillation and classification losses that this work directly builds upon and turns into an end-to-end architecture.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). It establishes the foundational distillation loss strategy for preserving prior-task capabilities in neural networks without retraining on the full original dataset.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). It formalizes the replay-based continual learning paradigm and benchmark setups on CIFAR-100 that contextualize memory-constrained incremental training.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). It provides the essential foundational baseline for parameter-regularization approaches to catastrophic forgetting in sequential neural network learning.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). It introduces online synapse-importance estimation to protect prior knowledge during sequential task training, serving as key background in continual learning.
  • Paper: Continual Learning with Deep Generative Replay, Hanul Shin et al. (2017). It establishes generative replay mechanisms as an alternative replay paradigm for mitigating catastrophic forgetting in sequential image classification.
Cover for End-to-End Incremental Learning

Abstract

Although deep learning approaches have stood out in recent years due to their state-of-the-art results, they continue to suffer from catastrophic forgetting, a dramatic decrease in overall performance when training with new classes added incrementally. This is due to current neural network architectures requiring the entire dataset, consisting of all the samples from the old as well as the new classes, to update the model -a requirement that becomes easily unsustainable as the number of classes grows. We address this issue with our approach to learn deep neural networks incrementally, using new data and only a small exemplar set corresponding to samples from the old classes. This is based on a loss composed of a distillation measure to retain the knowledge acquired from the old classes, and a cross-entropy loss to learn the new classes. Our incremental training is achieved while keeping the entire framework end-to-end, i.e., learning the data representation and the classifier jointly, unlike recent methods with no such guarantees. We evaluate our method extensively on the CIFAR-100 and ImageNet (ILSVRC 2012) image classification datasets, and show state-of-the-art performance.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Model
  • 3.1 Representative memory
  • 3.2 Deep network
  • 4 Incremental Learning
  • 5 Implementation Details
  • 6 Evaluation on CIFAR-100
  • 6.1 Fixed memory size
  • 6.2 Fixed number of samples
  • 6.3 Ablation studies
  • 7 Evaluation on ImageNet
  • 8 Summary
  • References

Knowls

  1. Knowl 1 — End-to-End Incremental Learning Architecture and Pipeline

    model/method

    The End-to-End Incremental Learning (EEIL) model trains deep convolutional neural networks continually on newly emerging classes without catastrophic forgetting, learning feature representations and classifiers jointly end-to-end.

    The deep architecture consists of a shared feature extractor followed by a succession of fully-connected classification layers, one for each added task/set of classes. When training incremental step tt, a new classification layer CLtCL_t is attached to the feature extractor alongside the previous classification layers CL1,…,CLt−1CL_1, \dots, CL_{t-1}.

    Each incremental step comprises four sequential stages:

    1. Construction of the Training Set: The training batch is formed by combining all available samples of the new classes with a representative exemplar set of the old classes stored in memory. For each image, two types of training targets are generated: standard one-hot classification labels across all observed classes, and F=t−1F = t - 1 distillation targets consisting of the logits produced by each of the FF classification layers corresponding to old classes.
    2. Training Process: All network parameters (both feature extractor and classification heads) are updated end-to-end by minimizing a cross-distilled loss combining multi-class cross-entropy and distillation losses.
    3. Balanced Fine-Tuning: An additional fine-tuning phase is executed with a reduced learning rate on a class-balanced subset containing equal numbers of samples per class, preventing bias toward the more numerous new-class samples.
    4. Representative Memory Updating: Exemplars of the newly learned classes are selected using herding and stored in the representative memory, while older exemplar sets are pruned to maintain budget constraints.
  2. Knowl 2 — Cross-Distilled Loss Function

    equation

    The cross-distilled loss function L(ω)L(\omega) governs the joint optimization of network parameters ω\omega during incremental training by combining cross-entropy classification loss across all observed classes with distillation loss across all historical classification heads:

    L(ω)=LC(ω)+∑f=1FLDf(ω)L(\omega) = L_C(\omega) + \sum_{f=1}^F L_{D_f}(\omega)

    where FF is the total number of previous classification heads corresponding to old classes, LC(ω)L_C(\omega) is the multi-class cross-entropy loss applied to all training samples (both old exemplars and new classes), and LDf(ω)L_{D_f}(\omega) is the distillation loss associated with classification head ff.

    The multi-class cross-entropy loss is given by:

    LC(ω)=−1N∑i=1N∑j=1Cpijlog⁡qijL_C(\omega) = -\frac{1}{N} \sum_{i=1}^N \sum_{j=1}^C p_{ij} \log q_{ij}

    where NN is the number of training samples in the batch, CC is the total number of classes seen so far, pij∈{0,1}p_{ij} \in \{0, 1\} is the one-hot ground-truth indicator for sample ii and class jj, and qijq_{ij} is the softmax probability computed from the network logits for sample ii at class jj.

    The distillation loss LDf(ω)L_{D_f}(\omega) on classification layer ff is defined as:

    LDf(ω)=−1N∑i=1N∑j=1Cfpdistijlog⁡qdistijL_{D_f}(\omega) = -\frac{1}{N} \sum_{i=1}^N \sum_{j=1}^{C_f} pdist_{ij} \log qdist_{ij}

    where CfC_f is the number of classes handled by classification layer ff, and pdistijpdist_{ij} and qdistijqdist_{ij} are temperature-scaled target and predicted probabilities:

    pdistij=exp⁡(zold,ij/T)∑k=1Cfexp⁡(zold,ik/T),qdistij=exp⁡(znew,ij/T)∑k=1Cfexp⁡(znew,ik/T)pdist_{ij} = \frac{\exp(z_{old, ij} / T)}{\sum_{k=1}^{C_f} \exp(z_{old, ik} / T)}, \quad qdist_{ij} = \frac{\exp(z_{new, ij} / T)}{\sum_{k=1}^{C_f} \exp(z_{new, ik} / T)}

    where zold,ijz_{old, ij} denotes the saved logit of sample ii on class jj from the model before the current increment, znew,ijz_{new, ij} is the current model logit, and TT is the distillation temperature hyperparameter (T=2T=2 in all experiments). Distillation loss is evaluated on all training samples (new classes as well as old exemplars) to preserve past representations.

  3. Knowl 3 — Balanced Fine-Tuning Stage for Incremental Learning

    model/method

    Because the exemplar memory only stores a limited number of samples from old classes, incremental training sets are heavily imbalanced (e.g., up to 12.5×12.5\times more training samples for new classes than old classes). This causes the classifier to become severely biased toward new classes.

    To correct this bias, a balanced fine-tuning stage is performed immediately after the main training phase:

    1. A balanced training subset is constructed containing an identical number of samples for each class (both old and new). For each new class, the number of samples is subsampled down to the per-class exemplar count nn by choosing the most representative samples according to the herding selection rule.
    2. To prevent the network from catastrophic forgetting of the new classes during this aggressive data reduction, a temporary distillation loss is added to the classification layer of the new classes using the logits predicted at the end of the main training phase.
    3. The network is fine-tuned on this balanced subset for 30 epochs using a reduced learning rate starting at 0.01 (compared to 0.1 in the main phase), divided by 10 every 10 epochs.
  4. Knowl 4 — Representative Memory Management and Herding Selection

    model/method

    Representative memory maintains a bounded set of exemplars for previously trained classes to mitigate catastrophic forgetting.

    Two memory budget protocols are defined:

    1. Fixed Memory Capacity (KK): The total memory across all observed classes is capped at KK samples. When cc classes are currently stored, each class is allocated n=⌊K/c⌋n = \lfloor K/c \rfloor exemplars. As new classes are added, nn decreases.
    2. Fixed Number of Samples (nn): A constant number of nn exemplars is stored for each old class, allowing total memory capacity to scale linearly with the number of classes.

    Exemplar selection and pruning operate as follows:

    • Herding Selection: For a given class with training feature vectors {x1,…,xM}\{\mathbf{x}_1, \dots, \mathbf{x}_M\} and class mean μ=1M∑i=1Mxi\boldsymbol{\mu} = \frac{1}{M} \sum_{i=1}^M \mathbf{x}_i, exemplars are greedily selected into an ordered list p1,…,pn\mathbf{p}_1, \dots, \mathbf{p}_n such that each subsequent exemplar pk\mathbf{p}_k minimizes the distance between the running exemplar mean and the full class mean:

    pk=arg⁡min⁡x∈X∖{p1,…,pk−1}∥μ−1k(x+∑j=1k−1pj)∥\mathbf{p}_k = \arg\min_{\mathbf{x} \in \mathcal{X} \setminus \{\mathbf{p}_1, \dots, \mathbf{p}_{k-1}\}} \left\| \boldsymbol{\mu} - \frac{1}{k}\left(\mathbf{x} + \sum_{j=1}^{k-1} \mathbf{p}_j\right) \right\|

    • Memory Pruning: When nn decreases due to fixed total capacity KK, the memory unit removes excess samples from the tail of each class's sorted exemplar list. Removed exemplars are never reused.
  5. Knowl 5 — Comparative Incremental Classification Accuracy on CIFAR-100 and ImageNet

    data/table

    The performance of the end-to-end incremental CNN (Our-CNN) was evaluated under a fixed memory budget (K=2000K=2000 for CIFAR-100 with ResNet-32; K=20000K=20000 for ImageNet ILSVRC 2012 with ResNet-18). Evaluations measured the average multi-class test accuracy over all incremental batches (excluding the initial non-incremental step). On CIFAR-100, results report the mean and standard deviation over 5 random class orders; on ImageNet, top-5 accuracy is reported.

    (a) CIFAR-100 (K=2000K=2000 exemplars)
    Method 2 classes 5 classes 10 classes 20 classes 50 classes
    Our-CNN 63.8 ±\pm 1.9 63.4 ±\pm 1.6 63.6 ±\pm 1.3 63.7 ±\pm 1.1 60.8 ±\pm 0.3
    iCaRL 54.1 ±\pm 2.5 57.8 ±\pm 2.6 60.5 ±\pm 1.6 62.0 ±\pm 1.2 61.8 ±\pm 0.4
    Hybrid1 34.9 ±\pm 4.5 48.4 ±\pm 2.6 55.8 ±\pm 1.8 60.4 ±\pm 1.0 60.8 ±\pm 0.7
    LwF.MC 9.6 ±\pm 1.5 29.5 ±\pm 2.2 40.4 ±\pm 2.0 47.6 ±\pm 1.5 52.9 ±\pm 0.6
    (b) ImageNet ILSVRC 2012 (K=20000K=20000 exemplars, Top-5 Accuracy)
    Method 10 classes 100 classes
    Our-CNN 90.4 69.4
    iCaRL 85.0 62.5
    Hybrid1 83.5 46.1
    LwF.MC 79.1 43.8

    Our-CNN achieves stability across all step sizes on CIFAR-100 (ranging between 63.4% and 63.8% for 2 to 20 classes/step), with improvements over iCaRL being statistically significant (p<0.01p < 0.01 via paired tt-test) for 2, 5, 10, and 20 class increments. On ImageNet, Our-CNN surpasses iCaRL by +5.4%+5.4\% (10 classes/step) and +6.9%+6.9\% (100 classes/step).

  6. Knowl 6 — Incremental Classification Performance with Fixed Exemplars Per Class

    data/table

    When representative memory is allocated a constant number of exemplars per class (n∈{50,75,100}n \in \{50, 75, 100\}) rather than a fixed global budget, memory size grows proportionally with the total number of classes. Average incremental classification accuracy on CIFAR-100 is summarized below:

    # Classes / step 5 10 20
    # Exemplars / class (nn) 50 75 100 50 75 100 50 75 100
    Our-CNN 62.4 66.9 68.6 62.7 65.7 68.5 63.3 65.4 67.3
    iCaRL 56.5 59.9 62.2 60.0 62.3 63.7 61.9 63.0 64.0
    Hybrid1 45.7 49.2 50.9 55.3 56.5 57.4 60.4 61.5 62.2

    Our-CNN outperforms both iCaRL and Hybrid1 across all configurations. Retaining 100 exemplars per class (which represents 20% of the CIFAR-100 training set) achieves 68.6% accuracy, closely approaching the non-incremental upper-bound baseline of 69.5%.

  7. Knowl 7 — Ablation of Incremental Learning Components on CIFAR-100

    data/table

    An ablation study on CIFAR-100 with a fixed memory capacity of K=2000K=2000 exemplars isolates the contributions of data augmentation (DA) and balanced fine-tuning (BF) relative to the baseline incremental CNN (Our-CNN-Base):

    Configuration 5 classes / step 10 classes / step 20 classes / step
    Our-CNN-Base (no DA, no BF) 57.0 53.7 50.1
    Our-CNN-DA (DA only, no BF) 59.2 57.9 57.2
    Our-CNN-BF (BF only, no DA) 57.9 58.1 57.1
    Our-CNN-Full (DA + BF) 63.8 64.0 63.2
    iCaRL 58.8 60.9 61.2
    Hybrid1 48.7 55.1 59.8

    Adding data augmentation improves accuracy across all step sizes. Balanced fine-tuning provides large improvements for wider step increments (e.g., +7.0%+7.0\% on 20 classes/step over Our-CNN-Base) where class imbalance between new data and old exemplars is most pronounced. Combining both components yields the highest accuracy (63.2%–64.0%).

    An ablation of exemplar selection strategies on CIFAR-100 (10 classes/step) shows that Herding achieves 63.6%, Random selection achieves 63.1%, and Distance Histogram selection achieves 59.1%.

  8. Knowl 8 — Performance on Incremental Learning of Visually Similar Classes

    empirical result

    The effectiveness of end-to-end feature and classifier learning was evaluated on subsets of visually similar (fine-grained) classes:

    1. CIFAR-100 Vehicles: A subset of 10 vehicle classes trained over 5 incremental steps of 2 classes each, with a memory budget of K=200K=200 exemplars.

      • Our-CNN: 73.3%73.3\% average incremental top-1 accuracy.
      • iCaRL: 47.5%47.5\% average incremental top-1 accuracy.
      • Hybrid1: 46.0%46.0\% average incremental top-1 accuracy.
    2. ImageNet Dog Breeds: A subset of 120 dog breeds trained over 12 incremental steps of 10 classes each, with a memory budget of K=2400K=2400 exemplars.

      • Our-CNN: 57.0%57.0\% average incremental top-5 accuracy.
      • iCaRL: 47.0%47.0\% average incremental top-5 accuracy.
      • Hybrid1: 40.7%40.7\% average incremental top-5 accuracy.

    Decoupled nearest-mean classifiers (NMC) like iCaRL struggle on fine-grained distributions because class means overlap closely in feature space, whereas end-to-end classifier optimization learns subtle inter-class boundaries.

  9. Knowl 9 — Comparison with Gradient Episodic Memory

    empirical result

    The proposed end-to-end incremental learning framework was compared against Gradient Episodic Memory (GEM) on CIFAR-100 split into 5-class incremental steps:

    1. Task-Aware Evaluation (Task known a priori at test time): Test prediction is restricted only to the 5 classes belonging to the sample's known task.
    Memory Size (samples) 200 1280 2560 5120
    GEM 48.0 62.6 64.6 67.8
    Our-CNN 77.9 86.1 88.4 90.7

    Our-CNN outperforms GEM by +29.9%+29.9\% at 200 memory samples and +22.9%+22.9\% at 5120 memory samples.

    1. Task-Agnostic Evaluation (Standard incremental learning, task unknown at test time): When the task ID is unknown and the model must classify over all seen classes (K=2000K=2000 exemplars), GEM suffers severe catastrophic forgetting and drops to 1.6%1.6\% average incremental accuracy, whereas Our-CNN achieves 64.8%64.8\%.

Coverage note — None was omitted; all key architectural components, equations, training stages, memory mechanisms, and experimental comparisons across CIFAR-100 and ImageNet benchmarks from the paper and its appendices are covered.

References

  1. 1.Supplementary material. Also available in the arXiv technical report. https://arxiv.org/abs/1807.09536
  2. 2.Ans, B., Rousset, S., French, R.M., Musca, S.: Self-refreshing memory in artificial neural networks: Learning temporal sequences without catastrophic forgetting. Connection Science 16(2), 71–99 (2004)
  3. 3.Bengio, Y., Courville, A., Vincent, P.: Representation learning: A review and new perspectives. PAMI 35(8), 1798–1828 (2013)
  4. 4.Cauwenberghs, G., Poggio, T.: Incremental and decremental support vector machine learning. In: NIPS (2000)
  5. 5.Chen, X., Shrivastava, A., Gupta, A.: NEIL: Extracting visual knowledge from web data. In: ICCV (2013)
  6. 6.Cortes, C., Vapnik, V.: Support-vector networks. Machine Learning 20(3), 273–297 (1995)
  7. 7.Divvala, S., Farhadi, A., Guestrin, C.: Learning everything about anything: Webly-supervised visual concept learning. In: CVPR (2014)
  8. 8.French, R.M.: Dynamically constraining connectionist networks to produce distributed, orthogonal representations to reduce catastrophic interference. In: Cognitive Science Society Conf. (1994)
  9. 9.Furlanello, T., Zhao, J., Saxe, A.M., Itti, L., Tjan, B.S.: Active long term memory networks. ArXiv e-prints, arXiv 1606.02355 (2016)
  10. 10.Goodfellow, I., Mirza, M., Xiao, D., Courville, A., Bengio, Y.: An empirical investigation of catastrophic forgetting in gradient-based neural networks. ArXiv e-prints, arXiv 1312.6211 (2013)
  11. 11.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
  12. 12.Hinton, G., Vinyals, O., Dean, J.: Distilling the knowledge in a neural network. In: NIPS workshop (2014)
  13. 13.Jung, H., Ju, J., Jung, M., Kim, J.: Less-forgetting learning in deep neural networks. ArXiv e-prints, arXiv 1607.00122 (2016)
  14. 14.Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A.A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., Hadsell, R.: Overcoming catastrophic forgetting in neural networks. Proc. National Academy of Sciences 114(13), 3521–3526 (2017)
  15. 15.Krizhevsky, A.: Learning multiple layers of features from tiny images. Tech. rep., University of Toronto (2009)
  16. 16.Li, Z., Hoiem, D.: Learning without forgetting. PAMI (2018)
  17. 17.Lopez-Paz, D., Ranzato, M.A.: Gradient episodic memory for continual learning. In: NIPS (2017)
  18. 18.McCloskey, M., Cohen, N.J.: Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of Learning and Motivation 24, 109 – 165 (1989)
  19. 19.Mensink, T., Verbeek, J., Perronnin, F., Csurka, G.: Distance-based image classification: Generalizing to new classes at near-zero cost. PAMI 35(11), 2624–2637 (2013)
  20. 20.Mitchell, T., Cohen, W., Hruschka, E., Talukdar, P., Betteridge, J., Carlson, A., Mishra, B.D., Gardner, M., Kisiel, B., Krishnamurthy, J., Lao, N., Mazaitis, K., Mohamed, T., Nakashole, N., Platanios, E., Ritter, A., Samadi, M., Settles, B., Wang, R., Wijaya, D., Gupta, A., Chen, X., Saparov, A., Greaves, M., Welling, J.: Never-ending learning. In: AAAI (2015)
  21. 21.Neelakantan, A., Vilnis, L., Le, Q.V., Sutskever, I., Kaiser, L., Kurach, K., Martens, J.: Adding gradient noise improves learning for very deep networks. ArXiv e-prints, arXiv 1511.06807 (2017)
  22. 22.Ratcliff, R.: Connectionist models of recognition memory: constraints imposed by learning and forgetting functions. Psychological review 97(2), 285 (1990)
  23. 23.Rebuffi, S.A., Kolesnikov, A., Sperl, G., Lampert, C.H.: iCaRL: Incremental classifier and representation learning. In: CVPR (2017)
  24. 24.Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards real-time object detection with region proposal networks. In: NIPS (2015)
  25. 25.Ristin, M., Guillaumin, M., Gall, J., Gool, L.V.: Incremental learning of ncm forests for large-scale image classification. In: CVPR (2014)
  26. 26.Ruping, S.: Incremental learning with support vector machines. In: ICDM (2001)
  27. 27.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet Large Scale Visual Recognition Challenge. IJCV 115(3), 211–252 (2015)
  28. 28.Rusu, A.A., Rabinowitz, N.C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., Hadsell, R.: Progressive neural networks. ArXiv e-prints, arXiv 1606.04671 (2016)
  29. 29.Ruvolo, P., Eaton, E.: ELLA: An efficient lifelong learning algorithm. In: ICML (2013)
  30. 30.Shmelkov, K., Schmid, C., Alahari, K.: Incremental learning of object detectors without catastrophic forgetting. In: ICCV (2017)
  31. 31.Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: NIPS (2014)
  32. 32.Terekhov, A.V., Montone, G., O’Regan, J.K.: Knowledge transfer in deep block-modular neural networks. In: Biomimetic and Biohybrid Systems (2015)
  33. 33.Thrun, S.: Lifelong Learning Algorithms, pp. 181–209. Springer US (1998)
  34. 34.Triki, A.R., Aljundi, R., Blaschko, M.B., Tuytelaars, T.: Encoder based lifelong learning. In: ICCV (2017)
  35. 35.Vedaldi, A., Lenc, K.: MatConvNet – Convolutional Neural Networks for MATLAB. In: ACM Multimedia (2015)
  36. 36.Welling, M.: Herding dynamical weights to learn. In: ICML (2009)
  37. 37.Xiao, T., Zhang, J., Yang, K., Peng, Y., Zhang, Z.: Error-driven incremental learning in deep convolutional neural network for large-scale image classification. In: ACM Multimedia (2014)

Citation

MLA
Castro, F. M., et al. “End-to-End Incremental Learning”. arXiv, 2018, http://arxiv.org/abs/1807.09536v2.
APA
Castro, F. M., Marín-Jiménez, M. J., Guil, N., Schmid, C., & Alahari, K. (2018). End-to-End Incremental Learning. arXiv. http://arxiv.org/abs/1807.09536v2
Chicago
Castro, F. M., M. J. Marín-Jiménez, N. Guil, C. Schmid, and K. Alahari. 2018. “End-to-End Incremental Learning”. arXiv. http://arxiv.org/abs/1807.09536v2.
Harvard
Castro, F.M. et al. (2018) “End-to-End Incremental Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1807.09536v2.
Vancouver
1. Castro FM, Marín-Jiménez MJ, Guil N, Schmid C, Alahari K (2018) End-to-End Incremental Learning. arXiv

BibTeX

@article{castro2018end,
  title = {End-to-End Incremental Learning},
  author = {Castro, Francisco M. and Marín-Jiménez, Manuel J. and Guil, Nicolás and Schmid, Cordelia and Alahari, Karteek},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1807.09536v2},
  eprint = {1807.09536}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF