Efficient Lifelong Learning with A-GEM

Arslan ChaudhryMarc'Aurelio RanzatoMarcus RohrbachMohamed Elhoseiny

article2018ICLR1,880 citations

Develops Averaged Gradient Episodic Memory (A-GEM) and a realistic single-pass benchmark protocol, achieving the high accuracy of memory-based continual learning at a fraction of the computational and storage cost.

Listen

Real-world intelligent systems require the ability to continually adapt to changing environments and learn new skills from streaming data without catastrophically forgetting previously acquired knowledge. While lifelong learning seeks to achieve this, standard evaluation protocols and existing methods rely heavily on unrealistic assumptions, such as observing data across multiple passes, sweeping hyper-parameters over the entire task sequence, or scaling memory and computational requirements linearly with every new task.

The article sets out to establish a realistic, resource-efficient evaluation protocol for lifelong learning and introduces Averaged Gradient Episodic Memory (A-GEM), an algorithm designed to maintain high predictive accuracy while drastically cutting computation time and memory usage during single-pass learning from task streams.

To evaluate this framework, the authors implemented a rigorous protocol where models tune hyper-parameters on a separate set of cross-validation tasks and subsequently process target task streams in a strict single pass. They benchmarked A-GEM against established baseline models—including standard unregularized models, regularization methods (such as Elastic Weight Consolidation), expanding architectural models, and the original Gradient Episodic Memory (GEM)—across four standard image classification datasets (Permuted MNIST, Split CIFAR, Split CUB, and Split AWA). Additionally, they introduced a metric called Learning Curve Area to capture how rapidly a system acquires new skills, and incorporated compositional task descriptors through joint-embedding architectures to facilitate rapid knowledge transfer.

The experimental findings demonstrate three primary outcomes: First, A-GEM delivers accuracy comparable to or better than the original GEM (achieving 89.1% accuracy on Permuted MNIST and 62.3% on Split CIFAR) while executing approximately 100 times faster and consuming roughly 10 times less memory during training. Second, conventional regularization approaches perform only slightly better than unregularized baselines in a single-pass streaming setting because they require multi-epoch training and heavily over-parameterized models to avoid forgetting. Third, introducing compositional task descriptors systematically accelerates learning across all tested models, substantially boosting zero-shot performance over time (for example, raising A-GEM's accuracy on Split CUB from 62% to 71%).

These findings indicate that memory-efficient, constraint-based gradient methods like A-GEM offer a viable pathway to deploy continually adapting machine learning systems on resource-constrained hardware in real-time environments. Unlike dynamically expanding networks that face out-of-memory failures on larger-scale setups, A-GEM maintains fixed-capacity efficiency while mitigating catastrophic forgetting.

Organizations developing real-time, streaming AI applications should adopt single-pass evaluation standards and consider episodic gradient projection methods like A-GEM when deploying models under fixed compute and memory budgets. Furthermore, practitioners should integrate task descriptors where metadata is available to enhance rapid transfer learning. However, decision-makers should note that a significant performance gap remains between single-pass sequential learning and traditional multi-task models that train on fully aggregated data simultaneously, warranting ongoing research into improved positive backward and forward transfer mechanisms.

arXiv: 1812.00420facebookresearch/agem
Cover for Efficient Lifelong Learning with A-GEM

Abstract

In lifelong learning, the learner is presented with a sequence of tasks, incrementally building a data-driven prior which may be leveraged to speed up learning of a new task. In this work, we investigate the efficiency of current lifelong approaches, in terms of sample complexity, computational and memory cost. Towards this end, we first introduce a new and a more realistic evaluation protocol, whereby learners observe each example only once and hyper-parameter selection is done on a small and disjoint set of tasks, which is not used for the actual learning experience and evaluation. Second, we introduce a new metric measuring how quickly a learner acquires a new skill. Third, we propose an improved version of GEM (Lopez-Paz & Ranzato, 2017), dubbed Averaged GEM (A-GEM), which enjoys the same or even better performance as GEM, while being almost as computationally and memory efficient as EWC (Kirkpatrick et al., 2016) and other regularization-based methods. Finally, we show that all algorithms including A-GEM can learn even more quickly if they are provided with task descriptors specifying the classification tasks under consideration. Our experiments on several standard lifelong learning benchmarks demonstrate that A-GEM has the best trade-off between accuracy and efficiency.

Table of Contents

  • 1 Introduction
  • 2 Learning Protocol
  • 3 Metrics
  • 4 Averaged Gradient Episodic Memory (a-gem)
  • 5 Joint Embedding Model Using Compositional Task Descriptors
  • 6 Experiments
  • 6.1 Results
  • 7 Related Work
  • 8 Conclusion
  • References
  • A Dataset Statistics
  • B a-gem Algorithm
  • C a-gem Update Rule
  • D Analysis of gem and a-gem
  • D.1 Frequency of Constraint Violations
  • D.2 Average Accuracy and Worst-Case Forgetting
  • D.3 Stochastic gem (s-gem)
  • E Result Tables
  • F Analysis of EWC
  • G Hyper-parameter Selection
  • H Pictorial Description of Joint Embedding Model

Knowls

  1. Knowl 1 — Averaged Gradient Episodic Memory Formulation and Projected Update Rule

    model/method

    Averaged Gradient Episodic Memory (A-GEM) is a continual learning method designed for streaming settings where tasks arrive sequentially and training data is observed in a single pass. For a model fθf_\theta parameterized by θ∈RP\theta \in \mathbb{R}^P, when training on the current task tt, A-GEM constrains parameter updates such that the average loss across all previous tasks stored in an episodic memory M=⋃k<tMk\mathcal{M} = \bigcup_{k < t} \mathcal{M}_k does not increase.

    Let g=∇θℓ(fθ(x,t),y)g = \nabla_\theta \ell(f_\theta(x, t), y) be the gradient computed on the current training batch (x,y)∈Dttrain(x, y) \in \mathcal{D}_t^{train}, and let gref=∇θℓ(fθ(xref,t),yref)g_{ref} = \nabla_\theta \ell(f_\theta(x_{ref}, t), y_{ref}) be the reference gradient computed on a random batch (xref,yref)∼M(x_{ref}, y_{ref}) \sim \mathcal{M} sampled uniformly from the accumulated episodic memory of all previous tasks. The parameter update gradient g~\tilde{g} is obtained by solving:

    min⁡g~12∥g−g~∥22subject tog~⊤gref≥0\min_{\tilde{g}} \frac{1}{2} \|g - \tilde{g}\|_2^2 \quad \text{subject to} \quad \tilde{g}^\top g_{ref} \ge 0

    When the unconstrained gradient violates the constraint (i.e., g⊤gref<0g^\top g_{ref} < 0), the closed-form projected gradient solution is:

    g~=g−g⊤grefgref⊤grefgref\tilde{g} = g - \frac{g^\top g_{ref}}{g_{ref}^\top g_{ref}} g_{ref}

    If g⊤gref≥0g^\top g_{ref} \ge 0, the gradient requires no modification and g~=g\tilde{g} = g. Compared to the standard Gradient Episodic Memory (GEM) formulation—which maintains t−1t-1 individual task inequality constraints and solves a quadratic program with t−1t-1 variables at every step—A-GEM aggregates past task knowledge into a single constraint, eliminating the quadratic program solver.

  2. Knowl 2 — A-GEM Training and Episodic Memory Management

    algorithm

    The A-GEM training procedure updates network parameters θ∈RP\theta \in \mathbb{R}^P across a sequence of TT tasks using a single pass per task dataset Dttrain\mathcal{D}_t^{train}. It populates and samples from an episodic memory buffer M\mathcal{M} to project gradients whenever parameter updates conflict with past performance.

    procedure TRAIN(f_\theta, D_train, D_test)
        M = {}
        A = 0 in R^{T x T}
        for t = 1 to T do
            for (x, y) in D_train[t] do
                (x_ref, y_ref) ~ M
                g_ref = grad_theta(loss(f_\theta(x_ref, t), y_ref))
                g = grad_theta(loss(f_\theta(x, t), y))
                if dot(g, g_ref) >= 0 then
                    g_tilde = g
                else
                    g_tilde = g - (dot(g, g_ref) / dot(g_ref, g_ref)) * g_ref
                end if
                theta = theta - alpha * g_tilde
            end for
            M = UPDATE_EPISODIC_MEMORY(M, D_train[t], |M_total| / T)
            A[t, :] = EVALUATE(f_\theta, D_test)
        end for
        return f_\theta, A
    end procedure
    procedure UPDATE_EPISODIC_MEMORY(M, D_train_t, s)
        for i = 1 to s do
            (x, y) ~ D_train_t
            M = M union {(x, y)}
        end for
        return M
    end procedure
    procedure EVALUATE(f_\theta, D_test)
        a = 0 in R^T
        for t = 1 to T do
            a[t] = compute_accuracy(f_\theta, D_test[t])
        end for
        return a
    end procedure

    Here, α>0\alpha > 0 is the learning rate, ∣Mtotal∣|M_{total}| is the total episodic memory budget across all tasks, and s=∣Mtotal∣/Ts = |M_{total}| / T is the allocated memory sample slots per task populated uniformly at random at the conclusion of training on task tt.

  3. Knowl 3 — Disjoint Cross-Validation Protocol for Single-Pass Lifelong Learning

    experimental setup

    To evaluate lifelong learning under strict streaming conditions where learners only observe each training example once, hyperparameter selection is detached from the evaluation tasks.

    The full task sequence is divided into two disjoint subsets drawn from the same underlying task distribution: a cross-validation task stream DCV={D1,…,DTCV}\mathcal{D}^{CV} = \{\mathcal{D}_1, \dots, \mathcal{D}_{T^{CV}}\} and an evaluation task stream DEV={DTCV+1,…,DT}\mathcal{D}^{EV} = \{\mathcal{D}_{T^{CV}+1}, \dots, \mathcal{D}_T\}, where TCV<TT^{CV} < T (e.g., TCV=3T^{CV} = 3 and T=20T = 20).

    1. Cross-Validation Phase: Hyperparameters (such as learning rate and regularization weight) are swept over DCV\mathcal{D}^{CV}. Multiple passes and resets over DCV\mathcal{D}^{CV} are permitted solely to select the optimal hyperparameter configuration h∗=arg⁡max⁡hATCV(h)h^* = \arg\max_h A_{T^{CV}}(h) based on average test accuracy across DCV\mathcal{D}^{CV}.
    2. Evaluation Phase: The network parameters θ\theta and all metric trackers are completely reset. Training proceeds on the unseen task sequence DEV\mathcal{D}^{EV} using h∗h^*, where the learner makes exactly one pass over each task dataset in sequential order. All reported test metrics are evaluated on DEV\mathcal{D}^{EV}.
  4. Knowl 4 — Learning Curve Area (LCA) Metric

    definition

    The Learning Curve Area (LCA∈[0,1]\text{LCA} \in [0, 1]) measures the sample efficiency and acquisition speed of a lifelong learning model during the early stages of learning new tasks.

    Let ak,b,j∈[0,1]a_{k,b,j} \in [0, 1] denote the classification accuracy on the test set of task jj after the model has been trained on the bb-th mini-batch of task kk. The average bb-shot accuracy across all TT evaluation tasks is:

    Zb=1T∑k=1Tak,b,kZ_b = \frac{1}{T} \sum_{k=1}^T a_{k,b,k}

    The Learning Curve Area up to mini-batch step β\beta is the normalized area under the convergence curve ZbZ_b:

    LCAβ=1β+1∑b=0βZb\text{LCA}_\beta = \frac{1}{\beta + 1} \sum_{b=0}^\beta Z_b

    Properties:

    • When β=0\beta = 0, LCA0=Z0\text{LCA}_0 = Z_0 corresponds to average zero-shot accuracy (forward transfer).
    • For small values of β\beta (e.g., β=10\beta = 10), LCAβ\text{LCA}_\beta differentiates between models that acquire competence rapidly from few gradient steps versus models that converge slowly to the same final average accuracy ATA_T.
  5. Knowl 5 — Joint Embedding Model with Compositional Task Descriptors

    model/method

    To enhance forward transfer and few-shot learning in lifelong learning, tasks are represented by compositional task descriptors rather than arbitrary discrete identifiers. A task descriptor tk∈RCk×At^k \in \mathbb{R}^{C_k \times A} is a matrix of class attribute representations, where CkC_k is the number of classes in task kk and AA is the total number of attribute dimensions across the dataset.

    The framework consists of two modules:

    1. A feature extractor ϕθ:X→RD\phi_\theta: \mathcal{X} \to \mathbb{R}^D mapping an input xkx^k to a DD-dimensional feature vector.
    2. A task embedding module ψω:RCk×A→RCk×D\psi_\omega: \mathbb{R}^{C_k \times A} \to \mathbb{R}^{C_k \times D}, parameterized by an attribute look-up matrix ω∈RA×D\omega \in \mathbb{R}^{A \times D}. The embedding of class cc is computed as a linear combination of its active attributes, producing a class embedding vector in RD\mathbb{R}^D.

    For an input example (xik,tk,yik)(x_i^k, t^k, y_i^k) with true class label c∈{1,…,Ck}c \in \{1, \dots, C_k\}, the model predicts class probabilities via a softmax over inner products between the image embedding and the task class embeddings:

    p(c∣xik,tk;θ,ω)=exp⁡([ϕθ(xik)ψω(tk)⊤]c)∑j=1Ckexp⁡([ϕθ(xik)ψω(tk)⊤]j)p(c \mid x_i^k, t^k; \theta, \omega) = \frac{\exp\left(\left[\phi_\theta(x_i^k) \psi_\omega(t^k)^\top\right]_c\right)}{\sum_{j=1}^{C_k} \exp\left(\left[\phi_\theta(x_i^k) \psi_\omega(t^k)^\top\right]_j\right)}

    The parameters θ\theta and ω\omega are trained jointly via cross-entropy loss:

    ℓk(θ,ω)=−1N∑i=1Nlog⁡p(yik∣xik,tk;θ,ω)\ell_k(\theta, \omega) = -\frac{1}{N} \sum_{i=1}^N \log p(y_i^k \mid x_i^k, t^k; \theta, \omega)

    This architecture enables zero-shot recognition of unseen attribute combinations by leveraging shared representations learned in prior tasks.

  6. Knowl 6 — Computational and Memory Complexity of Continual Learning Methods

    data/table

    Computational cost (measured as training time in seconds on a single GPU) and theoretical memory complexity for continual learning methods across 20-task benchmarks under the single-pass protocol:

    Methods Training Time [s] Memory
    MNIST CIFAR CUB AWA Training Testing
    VAN 186 105 54 4123 P+B⋅HP + B \cdot H P+B⋅HP + B \cdot H
    EWC 403 250 72 4136 4P+B⋅H4P + B \cdot H P+B⋅HP + B \cdot H
    PROG-NN 510 409 ∞\infty ∞\infty 2PT+BHT2PT + BHT 2PT+BHT2PT + BHT
    GEM 3442 5238 - - PT+(B+M)HPT + (B+M)H P+B⋅HP + B \cdot H
    A-GEM 477 449 420 5221 2P+(B+M)H2P + (B+M)H P+B⋅HP + B \cdot H

    Notation:

    • PP: Number of model parameters.
    • BB: Mini-batch size.
    • HH: Size of network hidden state.
    • MM: Number of episodic memory samples per task.
    • TT: Total number of tasks.
    • ∞\infty: Out-of-memory failure during training.

    Key comparisons:

    1. A-GEM trains roughly 7 to 12 times faster than GEM on Permuted MNIST and Split CIFAR while eliminating the O(PT)O(PT) scaling in training memory.
    2. Progressive Networks (PROG-NN) allocate O(PT)O(PT) parameters and O(BHT)O(BHT) activations, leading to out-of-memory failures on high-resolution image datasets (Split CUB and Split AWA).
    3. Regularization methods (EWC) require 4P4P training memory (storing Fisher information matrices and anchor parameters) but exhibit low runtime comparable to vanilla SGD (VAN).
  7. Knowl 7 — Empirical Performance of Continual Learning Approaches across Benchmarks

    empirical result

    Performance across continual learning benchmarks evaluated on final Average Accuracy (ATA_T, higher is better), Average Forgetting (FTF_T, lower is better), and Learning Curve Area at step 10 (LCA10\text{LCA}_{10}, higher is better) under single-pass training on the evaluation tasks:

    1. Permuted MNIST (20 tasks, 5 runs):

      • VAN: AT=47.9%A_T = 47.9\%, FT=0.51F_T = 0.51, LCA10=0.26\text{LCA}_{10} = 0.26
      • EWC: AT=68.3%A_T = 68.3\%, FT=0.29F_T = 0.29, LCA10=0.27\text{LCA}_{10} = 0.27
      • PROG-NN: AT=93.5%A_T = 93.5\%, FT=0.00F_T = 0.00, LCA10=0.19\text{LCA}_{10} = 0.19
      • GEM: AT=89.5%A_T = 89.5\%, FT=0.06F_T = 0.06, LCA10=0.23\text{LCA}_{10} = 0.23
      • A-GEM: AT=89.1%A_T = 89.1\%, FT=0.06F_T = 0.06, LCA10=0.29\text{LCA}_{10} = 0.29
      • Multi-Task Upper Bound: AT=95.3%A_T = 95.3\%
    2. Split CIFAR (20 tasks, 5 runs):

      • VAN: AT=42.9%A_T = 42.9\%, FT=0.25F_T = 0.25, LCA10=0.30\text{LCA}_{10} = 0.30
      • EWC: AT=42.4%A_T = 42.4\%, FT=0.26F_T = 0.26, LCA10=0.33\text{LCA}_{10} = 0.33
      • PROG-NN: AT=59.2%A_T = 59.2\%, FT=0.00F_T = 0.00, LCA10=0.21\text{LCA}_{10} = 0.21
      • GEM: AT=61.2%A_T = 61.2\%, FT=0.06F_T = 0.06, LCA10=0.36\text{LCA}_{10} = 0.36
      • A-GEM: AT=62.3%A_T = 62.3\%, FT=0.07F_T = 0.07, LCA10=0.35\text{LCA}_{10} = 0.35
      • Multi-Task Upper Bound: AT=68.3%A_T = 68.3\%
    3. Split CUB (20 tasks, 10 runs, Standard / Joint-Embedding (-JE)):

      • VAN: AT=54.3%/67.1%A_T = 54.3\% / 67.1\%, FT=0.13/0.10F_T = 0.13 / 0.10, LCA10=0.29/0.52\text{LCA}_{10} = 0.29 / 0.52
      • EWC: AT=54.0%/68.4%A_T = 54.0\% / 68.4\%, FT=0.13/0.09F_T = 0.13 / 0.09, LCA10=0.29/0.52\text{LCA}_{10} = 0.29 / 0.52
      • A-GEM: AT=62.0%/71.0%A_T = 62.0\% / 71.0\%, FT=0.07/0.07F_T = 0.07 / 0.07, LCA10=0.30/0.54\text{LCA}_{10} = 0.30 / 0.54
      • Multi-Task Upper Bound: AT=65.6%/73.8%A_T = 65.6\% / 73.8\%
    4. Split AWA (20 tasks, 10 runs, Standard / Joint-Embedding (-JE)):

      • VAN: AT=30.3%/42.8%A_T = 30.3\% / 42.8\%, FT=0.04/0.07F_T = 0.04 / 0.07, LCA10=0.21/0.37\text{LCA}_{10} = 0.21 / 0.37
      • EWC: AT=33.9%/43.3%A_T = 33.9\% / 43.3\%, FT=0.08/0.07F_T = 0.08 / 0.07, LCA10=0.26/0.37\text{LCA}_{10} = 0.26 / 0.37
      • A-GEM: AT=44.0%/50.0%A_T = 44.0\% / 50.0\%, FT=0.05/0.03F_T = 0.05 / 0.03, LCA10=0.29/0.39\text{LCA}_{10} = 0.29 / 0.39
      • Multi-Task Upper Bound: AT=64.8%/66.8%A_T = 64.8\% / 66.8\%

    A-GEM achieves the best trade-off between accuracy and resource efficiency among fixed-capacity architectures, and incorporating compositional task descriptors (-JE) substantially boosts few-shot learning and final accuracy across all methods.

  8. Knowl 8 — Stochastic Gradient Episodic Memory Formulation

    model/method

    Stochastic Gradient Episodic Memory (S-GEM) is a baseline variant of GEM designed to reduce the computational cost of the multi-constraint quadratic program. At each training step during task tt, S-GEM randomly samples a single past task index k∼{1,…,t−1}k \sim \{1, \dots, t-1\} and enforces an inequality constraint exclusively with respect to that sampled task's gradient gk=∇θℓ(fθ,Mk)g_k = \nabla_\theta \ell(f_\theta, \mathcal{M}_k):

    min⁡g~12∥g−g~∥22subject to⟨g~,gk⟩≥0\min_{\tilde{g}} \frac{1}{2} \|g - \tilde{g}\|_2^2 \quad \text{subject to} \quad \langle \tilde{g}, g_k \rangle \ge 0

    While S-GEM reduces the quadratic program to a single inner product check per step, it enforces alignment with only one random past task at a time. In empirical evaluations, S-GEM underperforms both full GEM and A-GEM on Permuted MNIST (AT=88.2%A_T = 88.2\% vs 89.5%89.5\% GEM and 89.1%89.1\% A-GEM) and Split CIFAR (AT=56.2%A_T = 56.2\% vs 61.2%61.2\% GEM and 62.3%62.3\% A-GEM), showing that constraining updates against the average gradient over all past memories (grefg_{ref}) provides superior regularization compared to sampling single task constraints.

  9. Knowl 9 — Optimization Trade-Offs and Constraint Violations in A-GEM versus GEM

    empirical result

    GEM and A-GEM optimize different criteria with distinct theoretical guarantees and empirical convergence behaviors:

    1. Worst-Case vs. Average Loss: GEM enforces t−1t-1 separate constraints ℓ(fθ,Mk)≤ℓ(fθt−1,Mk)\ell(f_\theta, \mathcal{M}_k) \le \ell(f_\theta^{t-1}, \mathcal{M}_k) for every k<tk < t, optimizing for lower worst-case forgetting across individual tasks on episodic memory samples (Fwst=0.00F_{wst} = 0.00 on MNIST memory, 0.050.05 on CIFAR memory for GEM, compared to 0.0080.008 and 0.150.15 for A-GEM). Conversely, A-GEM constrains the average loss over the union M=⋃k<tMk\mathcal{M} = \bigcup_{k<t} \mathcal{M}_k, which permits updates that reduce overall average loss even if a single task's loss increases slightly, leading to equal or higher average test accuracy (AT=62.3%A_T = 62.3\% on CIFAR for A-GEM vs 61.2%61.2\% for GEM).
    2. Constraint Violation Frequency: As the task sequence grows, GEM's individual constraints are violated at nearly every gradient step (approaching the total step limit of 5,500 on Permuted MNIST and 250 on Split CIFAR). In contrast, A-GEM's single average constraint violation frequency plateaus at a significantly lower number (under 2,000 steps on MNIST and under 100 steps on CIFAR). This reduced violation frequency contributes directly to A-GEM's training speedup.
  10. Knowl 10 — Limitations of Weight-Regularization Continual Learning Methods in Single-Pass Settings

    empirical result

    Weight-regularization methods (e.g., Elastic Weight Consolidation / EWC, Path Integral / PI, Synaptic Intelligence) fail to significantly outperform unregularized continual fine-tuning (VAN) when evaluated in single-pass, streaming lifelong learning settings with constrained network capacities.

    On Split CIFAR under single-pass training, EWC obtains an average accuracy of 42.4%42.4\%, which is comparable to the unregularized VAN baseline (42.9%42.9\%).

    Empirical analysis reveals two preconditions required for regularization methods like EWC to succeed:

    1. Over-Parameterized Architectures: Increasing network capacity (e.g., expanding from a 2-layer MLP with 256 units per layer to 2000 units per layer on MNIST, or moving from a reduced ResNet-18 to standard ResNet-18 on CIFAR) provides orthogonal parameter subspaces that mitigate interference between tasks.
    2. Multi-Epoch Optimization: When the number of training passes per task is increased from 1 to 10 epochs on MNIST and from 1 to 30 epochs on CIFAR, EWC's average accuracy increases significantly (reaching >80%>80\% on MNIST and >60%>60\% on CIFAR) while forgetting drops dramatically (FT<0.1F_T < 0.1). In single-pass regimes with limited data, Fisher information matrices cannot be reliably estimated from unconverged models, preventing effective regularization.

Coverage note — No substantial contributed material was omitted. Full algorithmic details, mathematical formulations, evaluation protocols, metrics (LCA), dataset statistics, complexity tables, empirical results, and analytical ablation studies (S-GEM, EWC failure mode analysis, GEM vs A-GEM constraint analysis) are fully covered.

References

  1. 1.Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars. Expert gate: Lifelong learning with a network of experts. In CVPR, pp. 7120–7129, 2017.
  2. 2.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. Memory aware synapses: Learning what (not) to forget. In ECCV, 2018.
  3. 3.Marco Baroni, Armand Joulin, Allan Jabri, Germàn Kruszewski, Angeliki Lazaridou, Klemen Simonic, and Tomas Mikolov. Commai: Evaluating the first steps towards a useful general ai. arXiv preprint arXiv:1701.08954, 2017.
  4. 4.Michael Chang, Abhishek Gupta, Sergey Levine, and Thomas L. Griffiths. Automatically composing representation transformations as a means for generalization. In ICML workshop Neural Abstract Machines and Program Induction v2, 2018.
  5. 5.Arslan Chaudhry, Puneet K Dokania, Thalaiyasingam Ajanthan, and Philip HS Torr. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In ECCV, 2018.
  6. 6.Mohamed Elhoseiny, Ahmed Elgammal, and Babak Saleh. Write a classifier: Predicting visual classifiers from unstructured text. IEEE TPAMI, 39(12):2539–2553, 2017.
  7. 7.Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A Rusu, Alexander Pritzel, and Daan Wierstra. Pathnet: Evolution channels gradient descent in super neural networks. arXiv preprint arXiv:1701.08734, 2017.
  8. 8.Leslie P. Kaelbling Ferran Alet, Tomas Lozano-Perez. Modular meta-learning. arXiv preprint arXiv:1806.10166v1, 2018.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pp. 770–778, 2016.
  10. 10.David Isele, Mohammad Rostami, and Eric Eaton. Using task features for zero-shot knowledge transfer in lifelong learning. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, IJCAI’16, pp. 1620–1626. AAAI Press, 2016. ISBN 978-1-57735-770-4.
  11. 11.James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences of the United States of America (PNAS), 2016.
  12. 12.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. https://www.cs.toronto.edu/ kriz/cifar.html, 2009.
  13. 13.Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. Learning to detect unseen object classes by between-class attribute transfer. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pp. 951–958. IEEE, 2009.
  14. 14.Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. Attribute-based classification for zero-shot visual object categorization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(3):453–465, 2014.
  15. 15.Yann LeCun. The mnist database of handwritten digits. http://yann.lecun.com/exdb/mnist/, 1998.
  16. 16.David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continuum learning. In NIPS, 2017.
  17. 17.Cuong V Nguyen, Yingzhen Li, Thang D Bui, and Richard E Turner. Variational continual learning. ICLR, 2018.
  18. 18.S-V. Rebuffi, A. Kolesnikov, and C. H. Lampert. iCaRL: Incremental classifier and representation learning. In CVPR, 2017.
  19. 19.Mark B Ring. Child: A first step towards continual learning. Machine Learning, 28(1):77–104, 1997.
  20. 20.Clemens Rosenbaum, Tim Klinger, and Matthew Riemer. Routing networks: Adaptive selection of non-linear functions for multi-task learning. In International Conference on Learning Representations, 2018.
  21. 21.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  22. 22.T. Schaul, D. Horgan, K. Gregor, and D. Silver. Universal value function approximators. ICML, 2015.
  23. 23.Jonathan Schwarz, Jelena Luketina, Wojciech M. Czarnecki, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. Progress and compress: A scalable framework for continual learning. In International Conference in Machine Learning, 2018.
  24. 24.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. Continual learning with deep generative replay. In NIPS, 2017.
  25. 25.R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup. Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction. The 10th International Conference on Autonomous Agents and Multiagent Systems, 2011.
  26. 26.Sebastian Thrun. Lifelong learning algorithms. In Learning to learn, pp. 181–209. Springer, 1998.
  27. 27.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The caltech-ucsd birds-200-2011 dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011.
  28. 28.Yongqin Xian, Christoph H Lampert, Bernt Schiele, and Zeynep Akata. Zero-shot learning-a comprehensive evaluation of the good, the bad and the ugly. IEEE transactions on pattern analysis and machine intelligence, 2018.
  29. 29.Ju Xu and Zhanxing Zhu. Reinforced continual learning. In arXiv preprint arXiv:1805.12369v1, 2018.
  30. 30.F. Zenke, B. Poole, and S. Ganguli. Continual learning through synaptic intelligence. In ICML, 2017.
  31. 31.Ji Zhang, Yannis Kalantidis, Marcus Rohrbach, Manohar Paluri, Ahmed Elgammal, and Mohamed Elhoseiny. Large-scale visual relationship understanding. arXiv preprint arXiv:1804.10660, 2018.

Citation

MLA
Chaudhry, A., et al. “Efficient Lifelong Learning with A-GEM”. arXiv, 2018, https://doi.org/10.48550/arxiv.1812.00420.
APA
Chaudhry, A., Ranzato, M., Rohrbach, M., & Elhoseiny, M. (2018). Efficient Lifelong Learning with A-GEM. arXiv. https://doi.org/10.48550/arxiv.1812.00420
Chicago
Chaudhry, A., M. Ranzato, M. Rohrbach, and M. Elhoseiny. 2018. “Efficient Lifelong Learning with A-GEM”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1812.00420.
Harvard
Chaudhry, A. et al. (2018) “Efficient Lifelong Learning with A-GEM”. arXiv. Available at: https://doi.org/10.48550/arxiv.1812.00420.
Vancouver
1. Chaudhry A, Ranzato M, Rohrbach M, Elhoseiny M (2018) Efficient Lifelong Learning with A-GEM. https://doi.org/10.48550/arxiv.1812.00420

BibTeX

@misc{https://doi.org/10.48550/arxiv.1812.00420,
  doi = {10.48550/ARXIV.1812.00420},
  url = {https://arxiv.org/abs/1812.00420},
  author = {Chaudhry, Arslan and Ranzato, Marc'Aurelio and Rohrbach, Marcus and Elhoseiny, Mohamed},
  keywords = {Machine Learning (cs.LG), Machine Learning (stat.ML), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Efficient Lifelong Learning with A-GEM},
  publisher = {arXiv},
  year = {2018},
  copyright = {Creative Commons Attribution 4.0 International}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/