Energy-based Latent Aligner for Incremental Learning

K. J. JosephSalman KhanFahad Shahbaz KhanRao Muhammad AnwerVineeth N. Balasubramanian

article2022CVPR58 citations

Presents a plug-and-play energy-based model that counteracts feature drift in incremental learning by mapping shifted representations back to lower-energy regions of earlier tasks, boosting performance across standard image classification and object detection benchmarks.

Listen

Modern artificial intelligence deployments frequently encounter dynamic environments where models must acquire new capabilities sequentially. When deep learning models are updated on new tasks without retaining access to previous data—often restricted due to privacy, storage, or computing limitations—they suffer from catastrophic forgetting. This degradation occurs because parameter updates optimized for new tasks cause internal feature representations to drift, disrupting the model's ability to recall earlier knowledge.

The article introduces and evaluates an Energy-based Latent Aligner for Incremental Learning (ELI). The primary objective is to demonstrate that an energy-based model can measure and undo the internal representational shifts that cause forgetting, functioning as a modular, plug-and-play component for existing continual learning architectures.

The authors develop a framework that models energy distributions directly within the intermediate feature space rather than the high-dimensional input image space. The system trains a lightweight energy function to assign lower energy values to representations produced by the previous model and higher energy values to those produced by the updated model, using only current task data as a proxy. During inference, sampling techniques adjust shifted features toward lower energy states, thereby realigning them to preserve past knowledge. This approach was evaluated across extensive image classification and object detection benchmarks, including CIFAR-100, ImageNet subsets, the full 1,000-class ImageNet dataset, and Pascal VOC.

The experimental findings show consistent, substantial performance gains across all evaluated settings. First, integrating ELI into leading incremental classification frameworks improved average accuracy across the board, yielding average gains of 8.17% on iCaRL, 3.05% on LUCIR, and 2.53% on AANet on the challenging ImageNet-1k benchmark. Second, ELI enhanced state-of-the-art incremental object detection by 3.0% to 7.0% in mean average precision across various task configurations on Pascal VOC. Third, aligning intermediate latent features proved significantly more effective than attempting to align final output classification scores. Finally, the energy aligner proved computationally lightweight, requiring only a small neural network that converged rapidly and introduced minimal sampling latency during inference.

These results indicate that internal representation alignment provides effective implicit regularization without requiring complex modifications to underlying learning algorithms or storing vast exemplar datasets. For engineering and product leaders, this translates to reduced memory footprints, enhanced data privacy compliance, and higher accuracy when deploying continuously updating vision systems in production.

Organizations maintaining vision models in evolving environments should consider adopting latent energy alignment as an add-on to their existing continual learning pipelines. The primary limitation noted in the article is that the aligner assumes high-level task knowledge is available during inference to determine whether a given sample belongs to the current task or a prior one. Despite this operational boundary condition, confidence in the reported performance gains is high given the rigorous evaluation across more than 45 experiments spanning diverse architectures and datasets.

arXiv: 2203.14952
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). iCaRL establishes the foundational class-incremental learning framework combining representation learning, exemplar management, and distillation that ELI directly builds upon and regularizes.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). Learning without Forgetting introduces the core distillation loss for preserving prior capabilities in sequential neural network adaptation, which ELI enhances by aligning latent representations.
  • Paper: End-to-End Incremental Learning, Francisco M. Castro et al. (2018). This work establishes end-to-end class-incremental learning via cross-distillation and balanced fine-tuning, providing essential context for representation drift mitigation.
  • Paper: Learning a Unified Classifier Incrementally via Rebalancing, Saihui Hou et al. (2019). This paper analyzes the representational imbalance and boundary confusion between old and new classes, motivating ELI's latent-space manifold alignment approach.
  • Paper: Large Scale Incremental Learning, Yue Wu et al. (2019). Bias Correction identifies how prediction and feature biases degrade large-scale incremental learning, providing key background for ELI's plug-and-play latent aligner.
  • Paper: Gradient Episodic Memory for Continual Learning, David Lopez-Paz et al. (2017). GEM defines gradient-alignment constraints across sequential tasks, formulating the fundamental optimization tension that ELI addresses via energy manifolds.
  • Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). Dark Experience Replay demonstrates how preserving continuous logit dynamics counters forgetting, offering a relevant baseline for latent-level regularization techniques.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). EWC introduces the seminal parameter-regularization approach for continual learning, laying the foundational concepts of implicit regularization against catastrophic forgetting.
Cover for Energy-based Latent Aligner for Incremental Learning

Abstract

Deep learning models tend to forget their earlier knowledge while incrementally learning new tasks. This behavior emerges because the parameter updates optimized for the new tasks may not align well with the updates suitable for older tasks. The resulting latent representation mismatch causes forgetting. In this work, we propose ELI: Energy-based Latent Aligner for Incremental Learning, which first learns an energy manifold for the latent representations such that previous task latents will have low energy and the current task latents have high energy values. This learned manifold is used to counter the representational shift that happens during incremental learning. The implicit regularization that is offered by our proposed methodology can be used as a plug-and-play module in existing incremental learning methodologies. We validate this through extensive evaluation on CIFAR-100, ImageNet subset, ImageNet 1k and Pascal VOC datasets. We observe consistent improvement when ELI is added to three prominent methodologies in class-incremental learning, across multiple incremental settings. Further, when added to the state-of-the-art incremental object detector, ELI provides over 5% improvement in detection accuracy, corroborating its effectiveness and complementary advantage to the existing art. Code is available at: https://github.com/JosephKJ/ELI.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Energy-based Latent Aligner
  • 3.1. Problem Setting
  • 3.2. Latent Aligner
  • 3.3. Toy Example
  • 4. Experiments and Results
  • 4.1. Incremental Classification Results
  • 4.2. Incremental Object Detection Results
  • 5. Discussions and Analysis
  • 5.1. ELI as an Implicit Regularizer
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Energy-based Latent Aligner Framework for Incremental Learning

    model/method

    The Energy-based Latent Aligner (ELI) is a plug-and-play framework designed to counteract catastrophic forgetting in incremental learning by rectifying the representational shift in the model's latent feature space.

    Let an incremental learning model at task Tt={τ1,τ2,…,τt}\mathcal{T}_t = \{\tau_1, \tau_2, \dots, \tau_t\} be decomposed as MTt(x)=(FϕTt∘FθTt)(x)M^{\mathcal{T}_t}(x) = (F_\phi^{\mathcal{T}_t} \circ F_\theta^{\mathcal{T}_t})(x), where FθTtF_\theta^{\mathcal{T}_t} is a feature extractor parameterized by θ\theta mapping an input xx to a latent vector z∈RDz \in \mathbb{R}^D, and FϕTtF_\phi^{\mathcal{T}_t} is the task-solving head parameterized by ϕ\phi (e.g., a classifier or detection sub-network). When the model is updated from task τt−1\tau_{t-1} to task τt\tau_t, parameter updates perturb the latent representations of earlier task data, causing catastrophic forgetting.

    ELI addresses this by modeling an energy manifold Eψ(z)E_\psi(z) over the latent space:

    1. Energy Manifold Learning: Without requiring stored exemplars from previous tasks, current task data x∼pdataτtx \sim p_{data}^{\tau_t} is passed through both the previous feature extractor FθTt−1F_\theta^{\mathcal{T}_{t-1}} and the current feature extractor FθTtF_\theta^{\mathcal{T}_t}. The previous model latents zTt−1=FθTt−1(x)z^{\mathcal{T}_{t-1}} = F_\theta^{\mathcal{T}_{t-1}}(x) serve as a proxy for the desired in-distribution manifold and are assigned low energy. The current model latents zTt=FθTt(x)z^{\mathcal{T}_t} = F_\theta^{\mathcal{T}_t}(x) are treated as out-of-distribution and assigned high energy.

    2. Inference-Time Latent Alignment: At test time, when evaluating a sample from a previous task, its extracted latent representation from the current model zTt=FθTt(x)z^{\mathcal{T}_t} = F_\theta^{\mathcal{T}_t}(x) is realigned by taking gradient descent steps along the learned energy function ∇zEψ(z)\nabla_z E_\psi(z) to minimize its energy, shifting the latent vector back toward the preserved previous-task distribution before passing it to FϕTtF_\phi^{\mathcal{T}_t}.

  2. Knowl 2 — Objective and Gradient Formulation for Latent Energy-Based Modeling

    equation

    ELI parameterizes the continuous energy function Eψ(z):RD→RE_\psi(z): \mathbb{R}^D \to \mathbb{R} using a neural network with parameters ψ\psi. The corresponding probability density over latent representations z∈RDz \in \mathbb{R}^D is defined by the Gibbs distribution:

    pψ(z)=exp⁡(−Eψ(z))∫zexp⁡(−Eψ(z))dzp_\psi(z) = \frac{\exp(-E_\psi(z))}{\int_z \exp(-E_\psi(z)) dz}

    The energy-based model is trained by maximizing the log-likelihood of latent samples drawn from the target in-distribution representations ptrue(z)p_{true}(z) (approximated by passing current-task inputs through the previous-task feature extractor FθTt−1F_\theta^{\mathcal{T}_{t-1}}):

    L(ψ)=Ez∼ptrue[log⁡pψ(z)]\mathcal{L}(\psi) = \mathbb{E}_{z \sim p_{true}}[\log p_\psi(z)]

    The analytical gradient of this log-likelihood objective with respect to ψ\psi is:

    ∂ψL(ψ)=Ez∼ptrue[−∂ψEψ(z)]+Ez∼pψ[∂ψEψ(z)]\partial_\psi \mathcal{L}(\psi) = \mathbb{E}_{z \sim p_{true}}[-\partial_\psi E_\psi(z)] + \mathbb{E}_{z \sim p_\psi}[\partial_\psi E_\psi(z)]

    Because exact sampling from pψ(z)p_\psi(z) is intractable due to the partition function, negative samples zsampledz_{sampled} are generated via stochastic gradient Langevin dynamics initialized at current-task latents zTt=FθTt(x)z^{\mathcal{T}_t} = F_\theta^{\mathcal{T}_t}(x):

    zi+1=zi−λ2∂zEψ(zi)+λωi,ωi∼N(0,I)z_{i+1} = z_i - \frac{\lambda}{2} \partial_z E_\psi(z_i) + \sqrt{\lambda}\omega_i, \quad \omega_i \sim \mathcal{N}(0, I)

    where λ\lambda denotes the Langevin step size and ωi\omega_i injects Gaussian noise capturing data uncertainty.

  3. Knowl 3 — Learning the Energy Manifold on Latent Features

    algorithm

    Algorithm for training the energy-based model EψE_\psi on intermediate representations during incremental learning without accessing previous task datasets:

    Input: Feature extractor trained till current task FθTtF_\theta^{\mathcal{T}_t}, Feature extractor trained till previous task FθTt−1F_\theta^{\mathcal{T}_{t-1}}, Current task data distribution pdataτtp_{data}^{\tau_t}
    Output: Trained energy function parameters ψ\psi
    Initialize the energy function EψE_\psi with Kaiming initialization
    while not converged do
        Sample mini-batch x∼pdataτtx \sim p_{data}^{\tau_t}
        zTt−1←FθTt−1(x)z^{\mathcal{T}_{t-1}} \leftarrow F_\theta^{\mathcal{T}_{t-1}}(x)
        zTt←FθTt(x)z^{\mathcal{T}_t} \leftarrow F_\theta^{\mathcal{T}_t}(x)
        zsampledTt←z_{sampled}^{\mathcal{T}_t} \leftarrow Sample from EψE_\psi using Langevin dynamics with zTtz^{\mathcal{T}_t} as initialization
        in_dist_energy ←Eψ(zTt−1)\leftarrow E_\psi(z^{\mathcal{T}_{t-1}})
        out_of_dist_energy ←Eψ(zsampledTt)\leftarrow E_\psi(z_{sampled}^{\mathcal{T}_t})
        Loss \leftarrow - \text{in_dist_energy} + \text{out_of_dist_energy}
        Update ψ\psi by optimizing Loss with RMSprop optimizer
    end while
    return EψE_\psi

    The algorithm iterates for 1500 iterations using mini-batches of size 128 and an initial learning rate of 0.00010.0001. Sampling from EψE_\psi uses 30 Langevin iterations.

  4. Knowl 4 — Latent Feature Alignment via Langevin Dynamics

    algorithm

    Procedure for realigning drifted latent feature representations of previous-task test instances at inference time using the learned energy manifold:

    Input: Latent vector z∈RDz \in \mathbb{R}^D extracted from current model FθTtF_\theta^{\mathcal{T}_t}, Learned EBM EψE_\psi, Number of Langevin steps LstepsL_{steps}, Learning rate λ\lambda
    Output: Aligned latent vector zz
    for step = 1 to LstepsL_{steps} do
        grad←∇zEψ(z)grad \leftarrow \nabla_z E_\psi(z)
        z←z−λ⋅gradz \leftarrow z - \lambda \cdot grad
    end for
    return zz

    For each test sample identified as belonging to an older task, its latent vector zz is iteratively adjusted by following the negative gradient of the energy function Eψ(z)E_\psi(z). This shifts the feature representation from a high-energy (drifted) location to a lower-energy (historically aligned) region of the latent space before passing it to the downstream task head.

  5. Knowl 5 — Task Identification Requirement at Inference

    assumption

    The inference alignment mechanism in ELI assumes that high-level task information is available at test time. Specifically, the system requires an indicator determining whether a given test instance's latent vector zz originates from the current task τt\tau_t or from any previous task in Tt−1\mathcal{T}_{t-1}. Latent alignment via ∇zEψ(z)\nabla_z E_\psi(z) is applied selectively to representations belonging to previous tasks, while representations from the current task are processed without energy-based adaptation.

  6. Knowl 6 — Performance on Class-Incremental Classification Benchmarks

    data/table

    ELI was integrated into three class-incremental classification baselines (iCaRL, LUCIR, AANet) across CIFAR-100 and ImageNet Subset (100 classes). Evaluations were performed under two protocols: (1) training on 50 classes initially and then adding remaining classes over 5, 10, or 25 tasks; and (2) splitting all classes equally across 5, 10, or 20 tasks. Average accuracy across tasks is reported.

    Settings Half of all classes in first task Same number of classes for each task
    Datasets CIFAR-100 ImageNet subset CIFAR-100 ImageNet subset
    Methods 5 Tasks 10 Tasks 25 Tasks 5 Tasks 10 Tasks 25 Tasks 5 Tasks 10 Tasks 20 Tasks 5 Tasks 10 Tasks 20 Tasks
    iCaRL 56.97 53.28 50.98 58.24 51.60 49.02 61.59 60.05 57.81 71.46 65.25 60.21
    iCaRL + ELI 63.68 58.92 54.00 68.94 61.48 56.11 70.13 67.81 63.06 78.51 71.66 66.77
    LUCIR 64.37 62.57 59.91 71.38 68.99 64.65 62.01 58.95 54.20 74.22 67.97 62.20
    LUCIR + ELI 66.06 63.50 60.30 74.58 71.62 66.35 64.55 59.51 54.98 75.38 70.28 65.51
    AANet 67.53 66.25 64.28 70.84 70.30 69.07 63.89 60.94 56.88 65.86 54.13 44.96
    AANet + ELI 68.78 66.62 64.72 73.54 71.82 70.32 66.36 61.72 57.65 67.43 55.47 46.93

    Adding ELI improves classification accuracy in every evaluated setting. On iCaRL, ELI provides gains of up to +10.73%+10.73\% on ImageNet subset (5 tasks) and +8.54%+8.54\% on CIFAR-100 (5 tasks). Under equal class distribution across tasks, iCaRL + ELI outperforms more complex explicit regularizers such as baseline LUCIR and AANet.

  7. Knowl 7 — Incremental Object Detection Evaluation on Pascal VOC 2007

    data/table

    ELI was integrated into the incremental Object Detector (iOD) framework based on Faster R-CNN on the Pascal VOC 2007 dataset under three incremental splits: adding 10 classes to 10 base classes (10+10), 5 classes to 15 base classes (15+5), and 1 class to 19 base classes (19+1). Results report mean Average Precision (mAP in %) on all 20 classes.

    Method 10 + 10 Setting (mAP %) 15 + 5 Setting (mAP %) 19 + 1 Setting (mAP %)
    All 20 (Joint Upper Bound) 75.2 75.2 75.2
    First Task Only 73.4 53.2 67.5
    Standard Fine-tuning 37.3 16.8 14.3
    Shmelkov et al. 63.1 65.9 68.3
    Faster ILOD 62.2 67.9 68.6
    ORE 64.6 68.5 68.9
    iOD 66.3 67.8 70.2
    iOD + ELI 71.7 74.8 73.2

    ELI improves the state-of-the-art detector iOD by +5.4%+5.4\% mAP in the 10+10 setting (66.3%→71.7%66.3\% \to 71.7\%), by +7.0%+7.0\% mAP in the 15+5 setting (67.8%→74.8%67.8\% \to 74.8\%), and by +3.0%+3.0\% mAP in the 19+1 setting (70.2%→73.2%70.2\% \to 73.2\%), demonstrating that energy-based latent realignment generalizes to object detection without structural modifications to the detector.

  8. Knowl 8 — Superiority of Latent Feature Alignment over Logit Alignment

    data/table

    To evaluate whether energy-based alignment should be performed in the latent feature space z=FθTt(x)z = F_\theta^{\mathcal{T}_t}(x) or in the final classification logit space FϕTt(FθTt(x))F_\phi^{\mathcal{T}_t}(F_\theta^{\mathcal{T}_t}(x)), incremental CIFAR-100 experiments were conducted using iCaRL, LUCIR, and AANet across 5, 10, and 25 task splits.

    Method 5 Tasks 10 Tasks 25 Tasks
    iCaRL 56.97 53.28 50.98
    iCaRL + Logit Aligner 57.97 (+1.00) 54.42 (+1.14) 51.49 (+0.51)
    iCaRL + ELI (Latent Aligner) 63.68 (+6.71) 58.92 (+5.64) 54.00 (+3.02)
    LUCIR 64.37 62.57 59.91
    LUCIR + Logit Aligner 62.50 (-1.87) 61.67 (-0.90) 59.22 (-0.69)
    LUCIR + ELI (Latent Aligner) 66.06 (+1.69) 63.50 (+0.93) 60.30 (+0.39)
    AANet 67.53 66.25 64.28
    AANet + Logit Aligner 66.16 (-1.37) 65.29 (-0.96) 63.81 (-0.47)
    AANet + ELI (Latent Aligner) 68.78 (+1.25) 66.62 (+0.37) 64.72 (+0.44)

    Aligning the output logits degrades performance on LUCIR (drops of 0.69%0.69\% to 1.87%1.87\%) and AANet (drops of 0.47%0.47\% to 1.37%1.37\%), while ELI (latent alignment) consistently improves performance across all methods and task configurations. This indicates that output logits are task-specific, whereas intermediate latent representations capture shared, transferable feature structures across incremental tasks.

  9. Knowl 9 — Implicit Latent Regularization Property of ELI

    empirical result

    ELI operates as an implicit regularizer of latent space distributions, capable of replacing explicit feature-space regularization losses.

    In experiments on ImageNet Subset (100 classes):

    1. When the explicit latent regularization term (ER) is removed from AANet, baseline performance drops from 70.8%→69.1%70.8\% \to 69.1\% (5 tasks), 70.3%→68.9%70.3\% \to 68.9\% (10 tasks), and 69.1%→63.0%69.1\% \to 63.0\% (25 tasks).
    2. Adding ELI to this unregularized model (AANet - ER + ELI) improves accuracy by +5.41%+5.41\% on 5 tasks (70.7%→76.1%70.7\% \to 76.1\%), +3.58%+3.58\% on 10 tasks (69.3%→72.9%69.3\% \to 72.9\%), and +2.57%+2.57\% on 25 tasks (68.9%→71.5%68.9\% \to 71.5\%).

    The accuracy gains provided by ELI are larger on the unregularized model than on the explicitly regularized model, demonstrating that ELI's post-hoc energy alignment effectively compensates for the absence of explicit penalization during backbone training.

  10. Knowl 10 — Implementation Specifications, Architecture, and Computational Complexity of ELI

    experimental setup

    ELI specifies the following architectural and computational parameters across tasks:

    • EBM Architecture: A multi-layer perceptron with two hidden layers containing 64 neurons each and a single output unit computing scalar energy (i-64-64-oi\text{-}64\text{-}64\text{-}o). For 64-dimensional latent vectors (CIFAR-100), the EBM contains 8,385 parameters and requires 1.057 MFLOPs. Maintaining an exponential moving average (EMA) of the EBM weights is used during training.
    • Latent Dimensions: D=32D = 32 and D=512D = 512 for MNIST; D=64D = 64 for CIFAR-100 (ResNet-18); D=512D = 512 for ImageNet-subset and ImageNet-1k (ResNet-32); D=2048D = 2048 (penultimate RoI Head feature vector of Faster R-CNN) for Pascal VOC.
    • Training Settings: Trained for 1500 iterations with mini-batch size 128 using the RMSprop optimizer with a learning rate of 10−410^{-4}.
    • Inference Cost: Sampling/alignment utilizes 30 Langevin iterations, requiring 0.039±0.0030.039 \pm 0.003 seconds per mini-batch of 128 samples on an Nvidia Tesla K80 GPU.

Coverage note — The illustrative two-task toy experiment on MNIST (Section 3.3 / Figure 3) and the full ImageNet-1k multi-task curves (Figure 4) were summarized conceptually and within the empirical results rather than as isolated knowls to preserve focus on the core methodology, analytical benchmarks, and quantitative ablations.

References

  1. 1.Davide Abati, Jakub Tomczak, Tijmen Blankevoort, Simone Calderara, Rita Cucchiara, and Babak Ehteshami Bejnordi. Conditional channel gated networks for task-aware continual learning. In CVPR, pages 3931–3940, 2020.
  2. 2.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision (ECCV), pages 139–154, 2018.
  3. 3.Michael Arbel, Liang Zhou, and Arthur Gretton. Generalized energy based models. In ICLR, 2021.
  4. 4.Anton Bakhtin, Yuntian Deng, Sam Gross, Myle Ott, Marc’Aurelio Ranzato, and Arthur Szlam. Residual energy-based models for text. Journal of Machine Learning Research, 22(40):1–41, 2021.
  5. 5.David Belanger and Andrew McCallum. Structured prediction energy networks. pages 983–992. PMLR, 2016.
  6. 6.Eden Belouadah and Adrian Popescu. Il2m: Class incremental learning with dual memory. In ICCV, pages 583–592, 2019.
  7. 7.Francisco M Castro, Manuel J Marín-Jiménez, Nicolás Guil, Cordelia Schmid, and Karteek Alahari. End-to-end incremental learning. In ECCV, pages 233–248, 2018.
  8. 8.Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elhoseiny. Efficient lifelong learning with a-gem. In ICLR, 2019.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  10. 10.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In ICCV, pages 86–102. Springer, 2020.
  11. 11.Yilun Du and Igor Mordatch. Implicit generation and modeling with energy based models. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, NeurIPS, volume 32, 2019.
  12. 12.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010.
  13. 13.Robert M French. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135, 1999.
  14. 14.Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. Your classifier is secretly an energy based model and you should treat it like one. In ICLR, 2019.
  15. 15.Akshita Gupta, Sanath Narayan, KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Mubarak Shah. Ow-detr: Open-world detection transformer. CVPR, 2021.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In European conference on computer vision, pages 630–645. Springer, 2016.
  17. 17.Mitch Hill, Jonathan Craig Mitchell, and Song-Chun Zhu. Stochastic security: Adversarial defense using long-run dynamics of energy-based models. In ICLR, 2021.
  18. 18.Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky. Neural networks for machine learning lecture 6a overview of mini-batch gradient descent. Cited on, 14(8):2, 2012.
  19. 19.Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. Distilling the knowledge in a neural network. In NIPS Deep Learning and Representation Learning Workshop, 2015.
  20. 20.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In CVPR, pages 831–839, 2019.
  21. 21.KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Vineeth N Balasubramanian. Towards open world object detection. In CVPR, pages 5830–5840, 2021.
  22. 22.KJ Joseph, Jathushan Rajasegaran, Salman Khan, Fahad Khan, and Vineeth N Balasubramanian. Incremental object detection via meta-learning. IEEE Transactions on Pattern Analysis & Machine Intelligence, Nov 2021.
  23. 23.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017.
  24. 24.Joseph KJ and Vineeth Nallure Balasubramanian. Meta-consolidation for continual learning. NeurIPS, 33, 2020.
  25. 25.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Citeseer, 2009.
  26. 26.Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang. A tutorial on energy-based learning. Predicting structured data, 1(0), 2006.
  27. 27.Shuang Li, Yilun Du, Gido M van de Ven, and Igor Mordatch. Energy-based models for continual learning. arXiv preprint arXiv:2011.12216, 2020.
  28. 28.Zhizhong Li and Derek Hoiem. Learning without forgetting. PAMI, 40(12):2935–2947, 2017.
  29. 29.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):2935–2947, 2018.
  30. 30.Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based out-of-distribution detection. NeurIPS, 2020.
  31. 31.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Adaptive aggregation networks for class-incremental learning. In CVPR, pages 2544–2553, 2021.
  32. 32.Yaoyao Liu, Yuting Su, An-An Liu, Bernt Schiele, and Qianru Sun. Mnemonics training: Multi-class incremental learning without forgetting. In CVPR, pages 12245–12254, 2020.
  33. 33.David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for continual learning. In NeurIPS, pages 6467–6476, 2017.
  34. 34.Arun Mallya and Svetlana Lazebnik. Packnet: Adding multiple tasks to a single network by iterative pruning. In CVPR, pages 7765–7773, 2018.
  35. 35.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pages 109–165. Elsevier, 1989.
  36. 36.Radford M Neal et al. Mcmc using hamiltonian dynamics. Handbook of markov chain monte carlo, 2(11):2, 2011.
  37. 37.Can Peng, Kun Zhao, and Brian C Lovell. Faster ilod: Incremental learning for object detectors based on faster rcnn. Pattern Recognition Letters, 140:109–115, 2020.
  38. 38.Jathushan Rajasegaran, Munawar Hayat, Salman Khan, Fahad Shahbaz Khan, and Ling Shao. Random path selection for incremental learning. NeurIPS, 2019.
  39. 39.Jathushan Rajasegaran, Munawar Hayat, Salman Khan, Fahad Shahbaz Khan, Ling Shao, and Ming-Hsuan Yang. An adaptive random path selection approach for incremental learning. arXiv preprint arXiv:1906.01120, 2019.
  40. 40.Jathushan Rajasegaran, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Mubarak Shah. itaml: An incremental task-agnostic meta-learning approach. In CVPR, pages 13588–13597, 2020.
  41. 41.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In CVPR, pages 2001–2010, 2017.
  42. 42.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 28:91–99, 2015.
  43. 43.Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, , and Gerald Tesauro. Learning to learn without forgetting by maximizing transfer and minimizing interference. In ICLR, 2019.
  44. 44.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671, 2016.
  45. 45.Joan Serra, Didac Suris, Marius Miron, and Alexandros Karatzoglou. Overcoming catastrophic forgetting with hard attention to the task. pages 4548–4557, 2018.
  46. 46.Konstantin Shmelkov, Cordelia Schmid, and Karteek Alahari. Incremental learning of object detectors without catastrophic forgetting. In Proceedings of the IEEE international conference on computer vision, pages 3400–3409, 2017.
  47. 47.Francesco Tonin, Arun Pandey, Panagiotis Patrinos, and Johan A. K. Suykens. Unsupervised energy-based out-of-distribution detection using stiefel-restricted kernel machine. In 2021 International Joint Conference on Neural Networks (IJCNN), pages 1–8, 2021.
  48. 48.Lifu Tu and Kevin Gimpel. Learning approximate inference networks for structured prediction. In ICLR, 2018.
  49. 49.Yezhen Wang, Bo Li, Tong Che, Kaiyang Zhou, Ziwei Liu, and Dongsheng Li. Energy-based open-world uncertainty modeling for confidence calibration. In ICCV, pages 9302–9311, 2021.
  50. 50.Max Welling and Yee W Teh. Bayesian learning via stochastic gradient langevin dynamics. pages 681–688, 2011.
  51. 51.Oliver Woodford. Notes on contrastive divergence. Department of Engineering Science, University of Oxford, Tech. Rep, 2006.
  52. 52.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In CVPR, pages 374–382, 2019.
  53. 53.Zhisheng Xiao, Karsten Kreis, Jan Kautz, and Arash Vahdat. Vaebm: A symbiosis between variational autoencoders and energy-based models. In ICLR, 2021.
  54. 54.Jianwen Xie, Yang Lu, Song-Chun Zhu, and Yingnian Wu. A theory of generative convnet. pages 2635–2644. PMLR, 2016.
  55. 55.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. pages 3987–3995, 2017.
  56. 56.Yang Zhao and Changyou Chen. Unpaired image-to-image translation via latent energy transport. In CVPR, pages 16418–16427, 2021.
  57. 57.Yang Zhao, Jianwen Xie, and Ping Li. Learning energy-based generative models via coarse-to-fine expanding and sampling. In ICLR, 2020.

Citation

MLA
Joseph, K. J., et al. “Energy-based Latent Aligner for Incremental Learning”. arXiv, 2022, http://arxiv.org/abs/2203.14952v1.
APA
Joseph, K. J., Khan, S., Khan, F. S., Anwer, R. M., & Balasubramanian, V. N. (2022). Energy-based Latent Aligner for Incremental Learning. arXiv. http://arxiv.org/abs/2203.14952v1
Chicago
Joseph, K. J., S. Khan, F. S. Khan, R. M. Anwer, and V. N. Balasubramanian. 2022. “Energy-based Latent Aligner for Incremental Learning”. arXiv. http://arxiv.org/abs/2203.14952v1.
Harvard
Joseph, K.J. et al. (2022) “Energy-based Latent Aligner for Incremental Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14952v1.
Vancouver
1. Joseph KJ, Khan S, Khan FS, Anwer RM, Balasubramanian VN (2022) Energy-based Latent Aligner for Incremental Learning. arXiv

BibTeX

@article{joseph2022energy,
  title = {Energy-based Latent Aligner for Incremental Learning},
  author = {Joseph, K J and Khan, Salman and Khan, Fahad Shahbaz and Anwer, Rao Muhammad and Balasubramanian, Vineeth N},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14952v1},
  eprint = {2203.14952}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE