Class-Incremental Exemplar Compression for Class-Incremental Learning

Zilin LuoYaoyao LiuBernt SchieleQianru Sun

article2023CVPR75 citations

Proposes an adaptive masking strategy optimized through bilevel learning that selectively compresses non-discriminative background pixels in exemplars, enabling memory-constrained class-incremental learning systems to store substantially more training samples and improve classification accuracy across benchmarks.

Listen

Artificial intelligence systems deployed in dynamic real-world environments must continually learn new object classes over time while retaining previously learned knowledge. A standard approach to prevent the forgetting of earlier classes is to store a small budget of past representative images, known as exemplars, to replay during future training. However, strict memory budgets typically restrict models to keeping only a handful of exemplars per old class, creating severe data imbalance between massive new data and scarce old data, which leads to substantial performance degradation.

The article evaluates and demonstrates a novel compression approach called class-incremental masking. The main objective is to overcome the severe sample shortage of old classes by downsampling non-essential background pixels while preserving critical object features at original resolutions, enabling systems to store substantially more exemplars within the exact same memory footprint without requiring manual localization annotations.

The researchers designed an adaptive mask generation framework that leverages the model's own visual attention mechanisms, known as class activation maps, to locate discriminative features. Because optimal visual cues shift as new classes are introduced, the method embeds learnable activation functions and uses a nested, two-level optimization routine to dynamically adjust masks across incremental learning stages. The framework was evaluated across standard high-resolution benchmark datasets—including Food-101, ImageNet-100, and large-scale ImageNet-1000—integrated as a plug-in component alongside top-performing continuous learning architectures under both fixed and expanding memory conditions.

The evaluation yielded several key findings. First, the proposed approach achieved state-of-the-art accuracy across all benchmark settings, consistently outperforming leading baselines like FOSTER and DER. Second, performance gains widened significantly under tighter memory constraints; on the 1,000-class benchmark with a restrictive 5,000-sample memory budget, the method boosted average accuracy by 4.8 percentage points in a 10-phase setting. Third, the framework delivered outsized benefits over longer training horizons with more incremental phases and achieved the highest relative accuracy gains on smaller target objects, where background downsampling yields greater memory savings and allows more exemplars to be preserved.

These findings indicate that strategic, selective image downsampling effectively mitigates the trade-off between sample quality and sample diversity in memory-constrained artificial intelligence systems. Rather than uniformly degrading image resolution or relying on fragile pixel synthesis, preserving fine-grained foreground cues while discarding irrelevant background details provides a cost-effective, high-performing pathway to continuous model updating without increasing physical memory costs or infrastructure footprint.

Engineering and research teams managing continuous machine learning pipelines should evaluate this masking approach as a modular plug-in to alleviate memory bottlenecks and catastrophic forgetting. Organizations should prioritize its application in scenarios characterized by high-resolution visual inputs, long operational deployment horizons, and tight edge-device storage constraints.

The methodology possesses clear boundary conditions. It is specifically tailored for high-resolution visual data and is ineffective on very low-resolution images, where compression parameter overhead outweighs the memory saved from downsampling. Additionally, the framework cannot retroactively modify previously saved exemplars once their original training data has been discarded. Nevertheless, given the thorough empirical validation and consistent improvements across standardized benchmarks, confidence in the reported performance gains remains high for high-resolution vision applications.

arXiv: 2303.14042

No sufficiently relevant recommendations were found.

Cover for Class-Incremental Exemplar Compression for Class-Incremental Learning

Abstract

Exemplar-based class-incremental learning (CIL) [36] finetunes the model with all samples of new classes but few-shot exemplars of old classes in each incremental phase, where the “few-shot” abides by the limited memory budget. In this paper, we break this “few-shot” limit based on a simple yet surprisingly effective idea: compressing exemplars by downsampling non-discriminative pixels and saving “many-shot” compressed exemplars in the memory. Without needing any manual annotation, we achieve this compression by generating 0-1 masks on discriminative pixels from class activation maps (CAM) [49]. We propose an adaptive mask generation model called class-incremental masking (CIM) to explicitly resolve two difficulties of using CAM: 1) transforming the heatmaps of CAM to 0-1 masks with an arbitrary threshold leads to a trade-off between the coverage on discriminative pixels and the quantity of exemplars, as the total memory is fixed; and 2) optimal thresholds vary for different object classes, which is particularly obvious in the dynamic environment of CIL. We optimize the CIM model alternatively with the conventional CIL model through a bilevel optimization problem [40]. We conduct extensive experiments on high-resolution CIL benchmarks including Food-101, ImageNet-100, and ImageNet-1000, and show that using the compressed exemplars by CIM can achieve a new state-of-the-art CIL accuracy, e.g., 4.8 percentage points higher than FOSTER [42] on 10-phase ImageNet-1000. Our code is available at https://github.com/xflz/CIM-CIL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminary
  • 4. Methodology
  • 4.1. CAM-based Compression Pipeline
  • 4.2. Class-Incremental Masking (CIM)
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Results and Analyses
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — Pixel-selective exemplar compression for CIL

    model/method

    The paper introduces a class-incremental learning (CIL) strategy that stores more old-class exemplars under the same memory budget by compressing only pixels that are judged non-discriminative. For an image exemplar, pixels inside a discriminative object region remain at the original resolution, while pixels outside that region are downsampled. The discriminative region is obtained without additional pixel-level annotations: a class activation map (CAM) provides a baseline mask, and the proposed class-incremental masking (CIM) model learns masks that adapt to the class and incremental phase. The resulting compressed exemplars are replayed together with all training data from the newly arriving classes.

  2. Knowl 2 — CAM masks and bounding-box localization

    equation

    For an image xx from incremental phase ii with class label yy, let F(x;θi)F(x;\theta_i) be the feature map produced by feature extractor parameters θi\theta_i, and let ωi,y\omega_{i,y} be the classifier weight vector for class yy. At each spatial feature-map location, the CAM activation is A=ωi,y⊤F(x;θi)A=\omega_{i,y}^{\top}F(x;\theta_i). After spatial upsampling to the resolution of xx, the normalized CAM is

    MCAM=A−min⁡(A)max⁡(A)−min⁡(A).M^{\mathrm{CAM}}=\frac{A-\min(A)}{\max(A)-\min(A)}.

    Here min⁡(A)\min(A) and max⁡(A)\max(A) are taken over spatial locations. Given a threshold τ∈[0,1]\tau\in[0,1], the binary discriminative mask is Mτ=I(MCAM>τ)M^{\tau}=\mathbb{I}(M^{\mathrm{CAM}}>\tau), where I\mathbb{I} is the indicator function. Instead of storing this image-sized mask, the method stores the tight rectangular bounding box B=[min⁡h,min⁡w;max⁡h,max⁡w]B=[\min h,\min w;\max h,\max w] over all locations (h,w)(h,w) for which Mτ(h,w)=1M^{\tau}(h,w)=1. The box is represented by four integer coordinates, so its storage cost is negligible compared with storing the mask itself.

  3. Knowl 3 — Bounding-box-based image compression and memory savings

    equation

    Let xx be an H×WH\times W RGB image, let BB have height HBH_B and width WBW_B, and let MBM^B be the binary image mask that is 11 inside BB and 00 outside. Let xηx_\eta be the full image downsampled by ratio η>1\eta>1. The compressed image is

    x~=MB⊙x+(1−MB)⊙xη,\widetilde{x}=M^B\odot x+(1-M^B)\odot x_\eta,

    where ⊙\odot and addition operate independently on the three RGB channels. Relative to the memory needed for one original image, the compressed image uses

    mx~=HBWBHW+1η(1−HBWBHW)=1−(1−1η)(1−HBWBHW).m_{\widetilde{x}}=\frac{H_BW_B}{HW}+\frac{1}{\eta}\left(1-\frac{H_BW_B}{HW}\right)=1-\left(1-\frac{1}{\eta}\right)\left(1-\frac{H_BW_B}{HW}\right).

    Thus, whenever the bounding box does not cover the entire image, mx~<1m_{\widetilde{x}}<1, allowing more exemplars to be stored in a fixed memory budget. The method stores the bounding-box coordinates together with the compressed pixels rather than storing the irregular pixel mask.

  4. Knowl 4 — Class-incremental masking network

    model/method

    CIM extends the CIL backbone with a second logical branch for mask generation. The ordinary CIL branch retains the original ReLU activation functions and produces the classifier used for the task. The CIM branch copies the parameters of the backbone's weight layers from the ordinary branch but replaces its activations with learnable activation functions, such as Padé Activation Units (PAUs). A PAU with numerator degree mm and denominator degree nn can be written as

    g(z)=a0+a1z+⋯+amzm1+b1z+⋯+bnzn,g(z)=\frac{a_0+a_1z+\cdots+a_mz^m}{1+b_1z+\cdots+b_nz^n},

    where zz is an activation input and a0,…,am,b1,…,bna_0,\ldots,a_m,b_1,\ldots,b_n are learnable parameters. The parameters of these learnable activations in phase ii are denoted by ϕi\phi_i. The CIM branch produces activation maps that are thresholded and converted to bounding boxes, so its masks can vary with both the input class and the current incremental phase. In the experiments, PAUs use degrees m=5m=5 and n=4n=4.

  5. Knowl 5 — Bilevel optimization of CIL and CIM

    model/method

    The CIL parameters and CIM parameters are optimized jointly through a local bilevel optimization problem in every incremental phase. Let E~0:i−1\widetilde{E}_{0:i-1} be the compressed exemplars retained from previous phases, DiD_i be the uncompressed new-class data in phase ii, and D~i(ϕi)\widetilde{D}_i(\phi_i) be that data compressed using CIM parameters ϕi\phi_i. The inner problem represents training a temporary CIL model on compressed data:

    θi∗=arg⁡min⁡θ  Ltrain ⁣(E~0:i−1∪D~i(ϕi);θ,ωi),\theta_i^*=\arg\min_{\theta}\;\mathcal{L}_{\mathrm{train}}\!\left(\widetilde{E}_{0:i-1}\cup\widetilde{D}_i(\phi_i);\theta,\omega_i\right),

    where θ\theta and ωi\omega_i are feature-extractor and classifier parameters, respectively. The outer problem selects mask parameters whose temporary model performs well on the original, uncompressed new data:

    min⁡ϕi  LCE(Di;θi∗,ωi)+μR(ϕi).\min_{\phi_i}\;\mathcal{L}_{\mathrm{CE}}(D_i;\theta_i^*,\omega_i)+\mu R(\phi_i).

    Here LCE\mathcal{L}_{\mathrm{CE}} is softmax cross-entropy, R(ϕi)R(\phi_i) is an ℓ2\ell_2 penalty on the generated mask, and μ\mu controls the pressure to reduce mask coverage and therefore memory use. In practice, the inner problem is approximated by one update θi+=θi−β1∇θLCIL(E~0:i−1∪D~i(ϕi);θi,ωi)\theta_i^+=\theta_i-\beta_1\nabla_{\theta}\mathcal{L}_{\mathrm{CIL}}(\widetilde{E}_{0:i-1}\cup\widetilde{D}_i(\phi_i);\theta_i,\omega_i). CIM is then updated with learning rate β2\beta_2 using the cross-entropy on uncompressed DiD_i, the memory regularizer, and an auxiliary cross-entropy term with weight μ′\mu' on E~0:i−1∪Di\widetilde{E}_{0:i-1}\cup D_i. The auxiliary term prevents different images from collapsing to identical activation maps.

  6. Knowl 6 — Phase-wise CIM-CIL training procedure

    algorithm

    The phase-wise procedure takes new data DiD_i, previously stored compressed exemplars E~0:i−1\widetilde{E}_{0:i-1}, the previous CIL parameters (θi−1,ωi−1)(\theta_{i-1},\omega_{i-1}), and the previous CIM parameters ϕi−1\phi_{i-1}. It returns updated models and new compressed exemplars.

    Input: New data DiD_i; previous compressed exemplars E~0:i−1\widetilde{E}_{0:i-1}; previous CIL model (θi−1,ωi−1)(\theta_{i-1},\omega_{i-1}); previous CIM parameters ϕi−1\phi_{i-1}
    Output: Updated CIL model (θi,ωi)(\theta_i,\omega_i); updated CIM parameters ϕi\phi_i; compressed exemplars E~i\widetilde{E}_i
    Initialize (θi,ωi)(\theta_i,\omega_i) from (θi−1,ωi−1)(\theta_{i-1},\omega_{i-1})
    Initialize ϕi\phi_i from ϕi−1\phi_{i-1}; use ReLU for ϕ0\phi_0
    for each training epoch do
        Update (θi,ωi)(\theta_i,\omega_i) on E~0:i−1∪Di\widetilde{E}_{0:i-1}\cup D_i using the selected baseline CIL loss
        Use ϕi\phi_i to produce activation maps, threshold them, form tight bounding boxes, and compress DiD_i outside the boxes
        Take one temporary CIL update to obtain θi+\theta_i^+ using E~0:i−1\widetilde{E}_{0:i-1} and the temporarily compressed data
        Update ϕi\phi_i using cross-entropy on the original DiD_i, mask-size regularization, and the auxiliary anti-collapse loss
    end for
    Compress all of DiD_i using the learned ϕi\phi_i
    Select E~i\widetilde{E}_i from the compressed data by feature herding
    return (θi,ωi)(\theta_i,\omega_i), ϕi\phi_i, and E~i\widetilde{E}_i

    For the reported experiments, the first phase uses 200 epochs and later phases use 170 epochs. The task learning rate starts at 0.10.1, β1\beta_1 starts at 0.10.1, and β2\beta_2 starts at 0.010.01; all three are cosine-annealed to zero. The regularization weights are μ=0.1\mu=0.1 and μ′=0.2\mu'=0.2, and the CIM gradient norm is clipped to at most 11.

  7. Knowl 7 — Compression-artifact augmentation

    model/method

    Selective compression creates a resolution discontinuity at the bounding-box boundary, which can introduce noisy high-frequency artifacts. To reduce sensitivity to these artifacts, each training epoch randomly converts a subset of the current new-class data DiD_i into compressed images using CAM-derived bounding boxes and the same downsampling ratio used for exemplar storage. These artificial compressed samples are mixed into training so that the CIL model learns to remain invariant to the boundary artifacts it will encounter when replaying compressed exemplars.

  8. Knowl 8 — Experimental protocols and implementation

    experimental setup

    The method is evaluated on high-resolution Food-101, ImageNet-100, and ImageNet-1000. Food-101 has 101 categories with 750 training and 250 test images per category; its images have maximum side length 512 pixels. ImageNet-1000 has 1,000 classes with approximately 1,300 training and 50 test images per class. ImageNet-100 is a 100-class subset of ImageNet-1000 sampled with NumPy seed 1993.

    Two CIL protocols are used. In learning from scratch (LFS), the same number of classes arrives in each phase, with N∈{5,10,20}N\in\{5,10,20\}. In learning from half (LFH), the first phase contains half of all classes and the remaining classes arrive evenly over N∈{5,10,25}N\in\{5,10,25\} later phases. Accuracy is evaluated after every phase on all classes seen so far; the reported metrics are mean accuracy over phases and final-phase accuracy. Each experiment is run three times and averaged.

    The fixed-memory setting uses 2,020 image units for Food-101, 2,000 for ImageNet-100, and either 5,000 or 20,000 for ImageNet-1000. The growing-memory setting allocates 20 image units per class. LFS uses fixed memory and LFH uses growing memory. All experiments use an 18-layer ResNet backbone and a fully connected classifier, SGD with momentum 0.90.9, weight decay 0.00050.0005, and cosine learning-rate annealing. Compression uses threshold τ=0.6\tau=0.6 and downsampling ratio η=4.0\eta=4.0.

  9. Knowl 9 — Consistent gains over DER and FOSTER on Food-101 and ImageNet-100

    empirical result

    Plugging CIM-CIL into the strong CIL baselines DER and FOSTER improves accuracy across the reported Food-101 and ImageNet-100 settings. The following values are average accuracies in percent; entries within each tuple correspond to the listed phase counts.

    • Food-101, LFS with N=(5,10,20)N=(5,10,20): DER obtains (73.88,70.76,64.39)(73.88,70.76,64.39) and DER with CIM obtains (75.63,73.09,69.17)(75.63,73.09,69.17); FOSTER obtains (75.03,72.72,66.73)(75.03,72.72,66.73) and FOSTER with CIM obtains (76.44,74.85,70.20)(76.44,74.85,70.20).
    • ImageNet-100, LFS with N=(5,10,20)N=(5,10,20): DER obtains (78.50,76.12,73.79)(78.50,76.12,73.79) and DER with CIM obtains (79.63,77.57,75.36)(79.63,77.57,75.36); FOSTER obtains (79.93,76.55,74.49)(79.93,76.55,74.49) and FOSTER with CIM obtains (80.58,77.94,75.23)(80.58,77.94,75.23).
    • Food-101, LFH with N=(5,10,25)N=(5,10,25): DER obtains (78.13,73.45,−)(78.13,73.45,-) and DER with CIM obtains (79.25,75.76,−)(79.25,75.76,-); FOSTER obtains (79.08,75.07,68.08)(79.08,75.07,68.08) and FOSTER with CIM obtains (79.76,76.86,70.50)(79.76,76.86,70.50).
    • ImageNet-100, LFH with N=(5,10,25)N=(5,10,25): DER obtains (79.08,77.73,−)(79.08,77.73,-) and DER with CIM obtains (80.30,79.05,−)(80.30,79.05,-); FOSTER obtains (80.07,77.54,72.40)(80.07,77.54,72.40) and FOSTER with CIM obtains (80.93,78.66,75.74)(80.93,78.66,75.74).

    The dashes indicate settings not reported for DER. Relative to FOSTER, CIM gives especially large gains when the number of phases is large: for example, the improvement is 3.473.47 percentage points on Food-101 LFS with N=20N=20 and 3.343.34 points on ImageNet-100 LFH with N=25N=25.

  10. Knowl 10 — ImageNet-1000 performance under fixed memory

    empirical result

    On ImageNet-1000 under the LFS protocol, CIM improves FOSTER with both 20,000 and 5,000 image-unit memory budgets. Each pair is (average accuracy,last-phase accuracy)(\text{average accuracy},\text{last-phase accuracy}) in percent.

    • With M=20,000M=20{,}000 and N=5N=5, iCaRL is (44.36,27.78)(44.36,27.78), WA is (58.37,50.62)(58.37,50.62), DER is (67.49,59.75)(67.49,59.75), FOSTER is (69.21,64.88)(69.21,64.88), and FOSTER with CIM is (69.93,66.05)(69.93,66.05).
    • With M=20,000M=20{,}000 and N=10N=10, iCaRL is (38.40,22.70)(38.40,22.70), WA is (54.10,45.66)(54.10,45.66), DER is (66.73,58.62)(66.73,58.62), FOSTER is (68.34,60.14)(68.34,60.14), and FOSTER with CIM is (69.53,62.07)(69.53,62.07).
    • With M=5,000M=5{,}000 and N=5N=5, FOSTER is (57.19,49.42)(57.19,49.42) and FOSTER with CIM is (61.37,54.46)(61.37,54.46).
    • With M=5,000M=5{,}000 and N=10N=10, FOSTER is (54.72,44.96)(54.72,44.96) and FOSTER with CIM is (59.48,50.83)(59.48,50.83).

    The tighter 5,000-image budget produces the largest gains: CIM raises FOSTER's average accuracy by 4.184.18 points for N=5N=5 and 4.764.76 points for N=10N=10, compared with gains of 0.720.72 and 1.191.19 points under the 20,000-image budget.

  11. Knowl 11 — Ablation evidence for adaptive masking and bilevel training

    empirical result

    An LFS ablation compares the baseline FOSTER pipeline with compression choices and CIM optimization variants. Values are average accuracies in percent for Food-101 and ImageNet-100 at N=10N=10 and N=20N=20, respectively:

    • Baseline: Food-101 (72.72,66.73)(72.72,66.73); ImageNet-100 (76.55,72.37)(76.55,72.37).
    • Artifact augmentation: (71.38,66.03)(71.38,66.03); (75.63,71.45)(75.63,71.45).
    • Full-image compression: (73.03,67.38)(73.03,67.38); (76.92,73.26)(76.92,73.26).
    • Random activation regions: (73.10,67.54)(73.10,67.54); (76.88,73.54)(76.88,73.54).
    • Center activation region: (73.29,67.88)(73.29,67.88); (76.78,73.82)(76.78,73.82).
    • Naive CAM activation: (73.76,68.65)(73.76,68.65); (77.21,74.67)(77.21,74.67).
    • Phase-wise manually selected threshold: (73.83,69.17)(73.83,69.17); (77.06,74.78)(77.06,74.78).
    • Joint CIL-CIM training: (73.44,69.01)(73.44,69.01); (77.34,74.59)(77.34,74.59).
    • Proposed bilevel optimization: (74.85,70.20)(74.85,70.20); (77.94,75.23)(77.94,75.23).
    • CIM with only the last backbone block learnable: (74.55,69.87)(74.55,69.87); (77.72,74.86)(77.72,74.86).
    • Weakly compressing the discriminative region as well, using downsampling ratio η′=2.0\eta'=2.0: (75.02,70.13)(75.02,70.13); (77.87,75.46)(77.87,75.46).

    The results indicate that model-derived activation regions are more useful than uniform, random, or center regions, and that the proposed bilevel optimization is stronger than either manual threshold selection or joint per-batch training. The additional weak compression of the foreground gives comparable accuracy but increases compression cost.

  12. Knowl 12 — Comparison with other exemplar-compression methods

    empirical result

    Using the same LUCIR CIL baseline, CIM is compared with Mnemonics, which distills data without increasing exemplar count, and MRDC, which applies uniform image compression. Average accuracies in percent are:

    • ImageNet-100 with N=(6,11,26)N=(6,11,26): LUCIR baseline (71.22,69.67,67.45)(71.22,69.67,67.45); LUCIR with Mnemonics (73.30,72.17,71.50)(73.30,72.17,71.50); LUCIR with MRDC (73.62,72.81,70.44)(73.62,72.81,70.44); LUCIR with CIM (74.05,73.76,72.84)(74.05,73.76,72.84).
    • ImageNet-1000 with N=(6,11,26)N=(6,11,26): LUCIR baseline (65.23,62.43,59.88)(65.23,62.43,59.88); LUCIR with Mnemonics (66.15,63.12,63.08)(66.15,63.12,63.08); LUCIR with MRDC (67.67,65.60,62.74)(67.67,65.60,62.74); LUCIR with CIM (68.03,66.54,63.77)(68.03,66.54,63.77).

    CIM is the best of the compared compression strategies in every listed setting. The paper attributes this to retaining high-resolution discriminative regions while increasing exemplar quantity and adapting the compression behavior to classes and phases, unlike fixed-count distillation or uniform compression.

  13. Knowl 13 — Effect of object size on compression benefit

    empirical result

    On ImageNet-100 with LFS and N=10N=10, classes are divided into 30 small, 40 middle, and 30 large objects according to bounding-box coverage: the smallest-covering classes are small and the largest-covering classes are large. CIM stores, on average, 39.40, 38.30, and 34.77 exemplars for small, middle, and large objects, respectively, whereas the uncompressed baseline stores 20 per class. The corresponding final-phase accuracies are:

    • Small objects: baseline 66.13%66.13\%, CIM 70.00%70.00\%, improvement +3.87+3.87 percentage points.
    • Middle objects: baseline 68.40%68.40\%, CIM 71.10%71.10\%, improvement +3.65+3.65 percentage points.
    • Large objects: baseline 69.93%69.93\%, CIM 72.26%72.26\%, improvement +2.33+2.33 percentage points.

    The largest gain occurs for small objects, which have more background pixels available for downsampling while preserving the smaller discriminative object region.

  14. Knowl 14 — Stated limitations of CIM-CIL

    limitation

    The method has three stated limitations. First, CIM cannot revise exemplars retained from earlier phases because their original validation data are no longer accessible; it adapts compression only for data arriving in the current phase. Second, the learnable activation branches add hundreds of activation parameters to the CIL model, although the paper reports that this is small relative to the total model size. Third, image compression is less useful for low-resolution datasets such as 32×3232\times32 CIFAR-100 because the storage for compression-related parameters and RGB pixels is comparable, making it more effective to spend memory on saving additional images directly.

Coverage note — The qualitative CAM-versus-CIM visualization examples and supplementary confidence intervals, preprocessing details, and hyperparameter-sensitivity analyses were omitted because they support the main method and results rather than adding separate load-bearing contributions.

References

  1. 1.Davide Abati, Jakub Tomczak, Tijmen Blankevoort, Simone Calderara, Rita Cucchiara, and Babak Ehteshami Bejnordi. Conditional channel gated networks for task-aware continual learning. In CVPR, 2020.
  2. 2.Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. Yolov4: Optimal speed and accuracy of object detection. arXiv, 2020.
  3. 3.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with random forests. In ECCV, 2014.
  4. 4.Gary Bradski. The opencv library. Dr. Dobb’s Journal: Software Tools for the Professional Programmer, 2000.
  5. 5.Kenneth R Castleman. Digital image processing. Prentice Hall Press, 1996.
  6. 6.George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A Efros, and Jun-Yan Zhu. Dataset distillation by matching training trajectories. In CVPR, 2022.
  7. 7.Can Chen, Xi Chen, Chen Ma, Zixuan Liu, and Xue Liu. Gradient-based bi-level optimization for deep learning: A survey. arXiv, 2022.
  8. 8.Zhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua, Hanwang Zhang, and Qianru Sun. Class re-activation maps for weakly-supervised semantic segmentation. In CVPR, 2022.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  10. 10.Prithviraj Dhar, Rajat Vikram Singh, Kuan-Chuan Peng, Ziyan Wu, and Rama Chellappa. Learning without memorizing. In CVPR, 2019.
  11. 11.Zhang Dong, Zhang Hanwang, Tang Jinhui, Hua Xiansheng, and Sun Qianru. Causal intervention for weakly supervised semantic segmentation. In NeurIPS, 2020.
  12. 12.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In ECCV, 2020.
  13. 13.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML, 2017.
  14. 14.Charles R Harris, K Jarrod Millman, Stefan J Van Der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al. Array programming with numpy. Nature, 2020.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV, 2015.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  17. 17.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. Learning a unified classifier incrementally via rebalancing. In CVPR, 2019.
  18. 18.Shenyang Huang, Vincent François-Lavet, and Guillaume Rabusseau. Neural architecture search for class-incremental learning. arXiv, 2019.
  19. 19.Imagenet object localization challenge. https://www.kaggle.com/competitions/imagenet-object-localization-challenge/.
  20. 20.Anil K Jain. Fundamentals of digital image processing. Prentice-Hall, Inc., 1989.
  21. 21.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. PNAS, 2017.
  22. 22.Jungbeom Lee, Eunji Kim, and Sungroh Yoon. Anti-adversarially manipulated attributions for weakly and semi-supervised semantic segmentation. In CVPR, 2021.
  23. 23.Zhizhong Li and Derek Hoiem. Learning without forgetting. PAMI, 2017.
  24. 24.Yaoyao Liu, Yingying Li, Bernt Schiele, and Qianru Sun. Online hyperparameter optimization for class-incremental learning. In AAAi, 2023.
  25. 25.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Adaptive aggregation networks for class-incremental learning. In CVPR, 2021.
  26. 26.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Rmm: Reinforced memory management for class-incremental learning. In NeurIPS, 2021.
  27. 27.Yaoyao Liu, Yuting Su, An-An Liu, Bernt Schiele, and Qianru Sun. Mnemonics training: Multi-class incremental learning without forgetting. In CVPR, 2020.
  28. 28.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv, 2016.
  29. 29.Dougal Maclaurin, David Duvenaud, and Ryan Adams. Gradient-based hyperparameter optimization through reversible learning. In ICML, 2015.
  30. 30.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of Learning and Motivation. Elsevier, 1989.
  31. 31.Ken McRae and Phil A Hetherington. Catastrophic interference is eliminated in pretrained networks. In Proceedings of the 15h Annual Conference of the Cognitive Science Society, 1993.
  32. 32.Alejandro Molina, Patrick Schramowski, and Kristian Kersting. Pade activation units: End-to-end learning of flexible activation functions in deep networks. In ICLR, 2019.
  33. 33.Vinod Nair and Geoffrey E Hinton. Rectified linear units improve restricted boltzmann machines. In ICML, 2010.
  34. 34.J Anthony Parker, Robert V Kenyon, and Donald E Troxel. Comparison of interpolating methods for image resampling. IEEE Transactions on Medical Imaging, 1983.
  35. 35.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019.
  36. 36.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. iCaRL: Incremental classifier and representation learning. In CVPR, 2017.
  37. 37.Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv, 2016.
  38. 38.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV, 2017.
  39. 39.Christian Simon, Piotr Koniusz, and Mehrtasht Harandi. On learning the geodesic path for incremental learning. In CVPR, 2021.
  40. 40.Ankur Sinha, Pekka Malo, and Kalyanmoy Deb. A review on bilevel optimization: from classical to evolutionary approaches and applications. IEEE Transactions on Evolutionary Computation, 2017.
  41. 41.Gregory K Wallace. The jpeg still picture compression standard. Communications of the ACM, 1991.
  42. 42.Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. Foster: Feature boosting and compression for class-incremental learning. arXiv, 2022.
  43. 43.Liyuan Wang, Xingxing Zhang, Kuo Yang, Longhui Yu, Chongxuan Li, Lanqing Hong, Shifeng Zhang, Zhenguo Li, Yi Zhong, and Jun Zhu. Memory replay with data compression for continual learning. In ICLR, 2022.
  44. 44.Yue Wu, Yinpeng Chen, Lijuan Wang, Yuancheng Ye, Zicheng Liu, Yandong Guo, and Yun Fu. Large scale incremental learning. In CVPR, 2019.
  45. 45.Ju Xu and Zhanxing Zhu. Reinforced continual learning. In NeurIPS, 2018.
  46. 46.Shipeng Yan, Jiangwei Xie, and Xuming He. Der: Dynamically expandable representation for class incremental learning. In CVPR, 2021.
  47. 47.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In ICML, 2017.
  48. 48.Bowen Zhao, Xi Xiao, Guojun Gan, Bin Zhang, and Shu-Tao Xia. Maintaining discrimination and fairness in class incremental learning. In CVPR, 2020.
  49. 49.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In CVPR, 2016.
  50. 50.Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. Co-transport for class-incremental learning. In Proceedings of the 29th ACM International Conference on Multimedia, 2021.

Citation

MLA
Luo, Z., et al. “Class-Incremental Exemplar Compression for Class-Incremental Learning”. arXiv, 2023, http://arxiv.org/abs/2303.14042v2.
APA
Luo, Z., Liu, Y., Schiele, B., & Sun, Q. (2023). Class-Incremental Exemplar Compression for Class-Incremental Learning. arXiv. http://arxiv.org/abs/2303.14042v2
Chicago
Luo, Z., Y. Liu, B. Schiele, and Q. Sun. 2023. “Class-Incremental Exemplar Compression for Class-Incremental Learning”. arXiv. http://arxiv.org/abs/2303.14042v2.
Harvard
Luo, Z. et al. (2023) “Class-Incremental Exemplar Compression for Class-Incremental Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.14042v2.
Vancouver
1. Luo Z, Liu Y, Schiele B, Sun Q (2023) Class-Incremental Exemplar Compression for Class-Incremental Learning. arXiv

BibTeX

@article{luo2023class,
  title = {Class-Incremental Exemplar Compression for Class-Incremental Learning},
  author = {Luo, Zilin and Liu, Yaoyao and Schiele, Bernt and Sun, Qianru},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.14042v2},
  eprint = {2303.14042}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE