Endpoints Weight Fusion for Class Incremental Semantic Segmentation

Jia-Wen XiaoChang-Bin ZhangJiekang FengXialei LiuJoost van de WeijerMing-Ming Cheng

article2023CVPR68 citations

Proposes a dynamic parameter-averaging strategy that fuses starting and ending network weights from each incremental step to mitigate catastrophic forgetting in semantic segmentation without increasing model size or requiring extra training.

Listen

Visual applications increasingly require computer vision systems to learn new object categories over time. However, updating these deep learning models with newly labeled data causes catastrophic forgetting, where the system overwrites and forgets previously learned classes. Storing previous data to retrain models from scratch creates severe data storage burdens and privacy risks, while existing mathematical constraints alone fail to retain strong historical memory.

The article demonstrates a model parameter aggregation framework called Endpoints Weight Fusion. The primary objective is to evaluate whether dynamically fusing the model weights from the start of a training stage with the weights at the end of that stage can successfully preserve past knowledge without expanding model size, requiring extra training steps, or relying on stored historical data.

The evaluation utilized benchmark visual segmentation datasets, specifically PASCAL VOC 2012 and ADE20K, across diverse incremental learning sequences ranging from two to eleven sequential steps. The framework pairs a standard knowledge distillation scheme—which constrains the models to remain close in parameter space during training—with a post-training parameter fusion calculation based on the ratio of new to total seen classes.

The analysis revealed three key findings. First, integrating the fusion strategy with existing baseline methods substantially boosted overall segmentation accuracy, improving baseline mean intersection over union by 21.4% to 33.3% in long sequential tasks. Second, the method exhibited significant performance gains on historical classes while maintaining high plasticity on newly introduced classes. Third, dynamic fusion factor calculation outperformed fixed weighting parameters across all tested scenarios and proved far superior to standard moving average updates, which overemphasize early collapsed gradients.

These findings indicate that teams deploying computer vision systems can continuously incorporate new object categories without inflating operational inference costs or expanding model footprints. The parameter-level fusion bypasses costly full-dataset retraining cycles, thereby reducing computational overhead and mitigating regulatory privacy concerns tied to long-term data storage.

Organizations developing continuous learning systems should adopt endpoint weight fusion on top of existing regularization techniques rather than relying on distillation constraints alone. Development teams should employ dynamic weighting ratios derived from category counts rather than manual hyperparameter tuning. Further research is recommended to explore the theoretical mechanics of weight fusion and evaluate its viability across other computer vision domains.

The findings are supported by consistent results across standard academic benchmarks and varied class arrival sequences. However, empirical testing remains bounded by standard supervised segmentation frameworks and Convolutional Neural Network architectures, meaning performance should be verified when deploying across alternative network designs or drastically different operational environments.

  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This foundational work introduces knowledge distillation as a regularization mechanism to preserve old task representations without historical data, which EWF directly builds upon and enhances via parameter fusion.
  • Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). This paper establishes the standard class-incremental learning framework combining classification loss and distillation loss, providing essential context for the catastrophic forgetting challenges tackled in EWF.
  • Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). This foundational paper formulates parameter-space regularization to mitigate catastrophic forgetting in neural networks, contextualizing EWF's strategy of parameter-space fusion and distance optimization.
  • Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey provides a comprehensive taxonomy of continual learning strategies and stability-plasticity trade-offs, establishing the conceptual background for regularization- and distillation-based incremental methods.
Cover for Endpoints Weight Fusion for Class Incremental Semantic Segmentation

Abstract

Class incremental semantic segmentation (CISS) focuses on alleviating catastrophic forgetting to improve discrimination. Previous work mainly exploits regularization (e.g., knowledge distillation) to maintain previous knowledge in the current model. However, distillation alone often yields limited gain to the model since only the representations of old and new models are restricted to be consistent. In this paper, we propose a simple yet effective method to obtain a model with a strong memory of old knowledge, named Endpoints Weight Fusion (EWF). In our method, the model containing old knowledge is fused with the model retaining new knowledge in a dynamic fusion manner, strengthening the memory of old classes in ever-changing distributions. In addition, we analyze the relationship between our fusion strategy and a popular moving average technique EMA, which reveals why our method is more suitable for class-incremental learning. To facilitate parameter fusion with closer distance in the parameter space, we use distillation to enhance the optimization process. Furthermore, we conduct experiments on two widely used datasets, achieving state-of-the-art performance.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • Class Incremental Learning.
  • Class Incremental Semantic Segmentation.
  • Weight Fusion Method.
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Endpoints Weight Fusion (EWF)
  • Knowledge distillation enhanced for EWF.
  • Discussion on EMA v.s. EWF.
  • 3.3. Overall Framework
  • 4. Experiments
  • 4.1. Experimental setups
  • Protocols.
  • Implementation Details.
  • 4.2. Comparison to competing methods
  • PASCAL VOC 2012.
  • Comparison with methods that introduce auxiliary data.
  • Visualization.
  • ADE20K.
  • 4.3. Ablation Study
  • Fusion Strategy.
  • Fusion Factor Selection.
  • Robustness of Class Order.
  • 5. Conclusions
  • Acknowledgment.
  • References

Knowls

  1. Knowl 1 — Endpoints Weight Fusion

    model/method

    Endpoints Weight Fusion (EWF) integrates old and newly learned knowledge by interpolating the parameters of two models at the end of each incremental task. For task tt, let θold\theta_{\mathrm{old}} be the final parameters from task t−1t-1, let θnew\theta_{\mathrm{new}} be the parameters after training on the current task, and let θbalanced\theta_{\mathrm{balanced}} be the parameters used for inference after fusion. EWF computes

    θbalanced=αtθnew+(1−αt)θold.\theta_{\mathrm{balanced}}=\alpha_t\theta_{\mathrm{new}}+(1-\alpha_t)\theta_{\mathrm{old}}.

    The fusion coefficient is selected dynamically from the number of newly learned classes NnewN_{\mathrm{new}} and previously learned classes NoldN_{\mathrm{old}}:

    αt=NnewNnew+Nold.\alpha_t=\sqrt{\frac{N_{\mathrm{new}}}{N_{\mathrm{new}}+N_{\mathrm{old}}}}.

    Here NnewN_{\mathrm{new}} and NoldN_{\mathrm{old}} are nonnegative class counts for the current incremental step, and αt∈[0,1]\alpha_t\in[0,1]. The same interpolation is applied to corresponding parameters throughout the network, including convolution and normalization layers. Unlike ensemble approaches, EWF retains one network at inference time; unlike compression or re-parameterization approaches, it requires no post-fusion training or special layer structure, keeps the model size constant, and introduces no additional inference computation. The framework diagram on page 3 depicts training two endpoint models and directly interpolating their weights into a balanced model.

  2. Knowl 2 — Class-Incremental Semantic Segmentation Setting

    definition

    Class-incremental semantic segmentation (CISS) learns a sequence of TT supervised segmentation tasks with model fθf_\theta. Task tt contains dataset Dt={(xi,yi)}D_t=\{(x_i,y_i)\}, where xix_i is an image and yiy_i is its pixel-level label map. The task-specific foreground classes form a set CtC_t, and the background label is denoted by cbc_b. Foreground class sets are disjoint across tasks, Ci∩Cj=∅C_i\cap C_j=\varnothing for i≠ji\ne j.

    Only the classes introduced at the current task are annotated. Consequently, cbc_b includes both genuine background pixels and pixels belonging to classes learned in previous tasks or reserved for future tasks. This overlapped-label setting makes the background label semantically different across tasks and causes current-task training to favor new classes while forgetting old ones. The model must therefore preserve predictions for all previously learned classes while acquiring discrimination for CtC_t, without access to earlier training data.

  3. Knowl 3 — Distillation-Enhanced EWF Training

    model/method

    EWF uses knowledge distillation during current-task optimization to keep the old and new endpoint models sufficiently close in parameter and representation space for interpolation to be effective. Let DtD_t be the current task dataset. For an old model with feature extractor Ψold\Psi_{\mathrm{old}} and classifier Φold\Phi_{\mathrm{old}}, and a new model with Ψnew\Psi_{\mathrm{new}} and Φnew\Phi_{\mathrm{new}}, the paper describes feature- and logit-based distillation losses as

    LFD=1∣Dt∣∑(xi,yi)∈Dt∥Ψold(xi)−Ψnew(xi)∥22,L_{\mathrm{FD}}=\frac{1}{|D_t|}\sum_{(x_i,y_i)\in D_t}\left\|\Psi_{\mathrm{old}}(x_i)-\Psi_{\mathrm{new}}(x_i)\right\|_2^2, LLD=1∣Dt∣∑(xi,yi)∈DtKL ⁣(Φold(Ψold(xi)),Φnew(Ψnew(xi))),L_{\mathrm{LD}}=\frac{1}{|D_t|}\sum_{(x_i,y_i)\in D_t}\mathrm{KL}\!\left(\Phi_{\mathrm{old}}(\Psi_{\mathrm{old}}(x_i)),\Phi_{\mathrm{new}}(\Psi_{\mathrm{new}}(x_i))\right),

    where ∣Dt∣|D_t| is the number of current-task samples and KL\mathrm{KL} is the Kullback–Leibler divergence between the old and new output distributions. Existing CISS losses such as UNKD for logits and POD for features can be substituted into EWF. During task tt, the model parameters are optimized with cross-entropy for current-class learning plus a distillation loss relative to the previous model:

    min⁡θt  LCE(θt)+LKD(θt;θt−1),\min_{\theta_t}\;L_{\mathrm{CE}}(\theta_t)+L_{\mathrm{KD}}(\theta_t;\theta_{t-1}),

    where θt−1\theta_{t-1} is the previous task's final model. Distillation prevents the unconstrained training trajectory from moving the two EWF endpoints too far apart and thereby improves the usefulness of their interpolation.

  4. Knowl 4 — EWF Versus Exponential Moving Average

    theoretical result

    The paper analyzes EWF against exponential moving average (EMA) under unit-learning-rate SGD. EMA updates a stored model viv^i after iteration ii using the current parameters θi\theta^i:

    vi=βvi−1+(1−β)θi,v^i=\beta v^{i-1}+(1-\beta)\theta^i,

    where β\beta is the EMA coefficient, typically 0.90.9 or 0.990.99. For a training trajectory beginning at θ1\theta^1, the coefficient on the gradient at iteration kk in the EMA endpoint is proportional to 1−βn−k1-\beta^{n-k} at iteration nn, so early gradients receive more influence and later gradients are progressively attenuated. In contrast, EWF combines the start and end parameters with a single coefficient αt\alpha_t, giving the gradients accumulated between the endpoints approximately uniform influence.

    The paper's representation-similarity experiment on page 4 observes that similarity between the initial old model and the current model first decreases, then recovers, and finally stabilizes during incremental training. The authors interpret the initial decrease as representation collapse while the model learns new classes and the later increase as recovery induced by cross-entropy and distillation. Because EMA emphasizes the early collapsed part of the trajectory whereas EWF gives relatively greater weight to the later recovered part, the paper argues that EWF is better suited to CISS than EMA.

  5. Knowl 5 — EWF Incremental Training Procedure

    algorithm

    The EWF procedure receives an initial model fθ0f_{\theta_0}, the number of tasks TT, task datasets DtD_t, and learning rate γ\gamma. At each task it starts from the previous balanced model, trains with cross-entropy plus knowledge distillation, and fuses the initial and final models without another training phase.

    Input: initial parameters θ0\theta_0, task count TT, datasets D1,…,DTD_1,\ldots,D_T, learning rate γ\gamma
    Output: final parameters θT\theta_T
    for task tt from 1 to TT do
        Count current and previous classes as NnewN_{\mathrm{new}} and NoldN_{\mathrm{old}}
        Set θ1←θt−1\theta^1 \leftarrow \theta_{t-1}
        Set αt←Nnew/(Nnew+Nold)\alpha_t \leftarrow \sqrt{N_{\mathrm{new}}/(N_{\mathrm{new}}+N_{\mathrm{old}})}
        Set iteration index i←1i \leftarrow 1
        while training has not converged do
            Sample a mini-batch from DtD_t
            Update θi+1←θi−γ∇θi[LCE+LKD]\theta^{i+1} \leftarrow \theta^i-\gamma\nabla_{\theta^i}[L_{\mathrm{CE}}+L_{\mathrm{KD}}]
            Increment ii
        end while
        Set θold←θ1\theta_{\mathrm{old}}\leftarrow\theta^1 and θnew←θi\theta_{\mathrm{new}}\leftarrow\theta^i
        Set θbalanced←αtθnew+(1−αt)θold\theta_{\mathrm{balanced}}\leftarrow\alpha_t\theta_{\mathrm{new}}+(1-\alpha_t)\theta_{\mathrm{old}}
        Set θt←θbalanced\theta_t\leftarrow\theta_{\mathrm{balanced}}
    end for
    return θT\theta_T
  6. Knowl 6 — Experimental Protocol and Implementation

    experimental setup

    The experiments use the overlapped CISS protocol: each task has disjoint newly labeled classes, current images may contain old classes marked as background, and only current-task data is available during training. PASCAL VOC 2012 has 10,582 training images, 1,449 validation images, 20 object classes, and background. ADE20K has 20,210 training images, 2,000 validation images, and 150 classes. An XX-YY setting starts with XX classes and introduces YY new classes per later task. The evaluated settings are PASCAL VOC 15-1, 10-1, 5-3, and 19-1, and ADE20K 100-5, 100-10, and 100-50.

    The segmentation model is DeepLab-v3 with a ResNet-101 backbone and in-place activated batch normalization. Training uses horizontal flipping and random cropping, SGD, batch size 24, and a polynomial learning-rate schedule. The model is trained for 30 epochs per PASCAL VOC task and 60 epochs per ADE20K task; the initial learning rate is 0.010.01 for the first task and 0.0010.001 for subsequent tasks. Twenty percent of the training data is used as an internal validation split, while the reported metric is mean intersection-over-union (mIoU) on the original validation set. Experiments use four RTX 2080Ti GPUs.

  7. Knowl 7 — PASCAL VOC Incremental Segmentation Results

    data/table

    The page-6 quantitative comparison reports final-task mIoU on PASCAL VOC 2012 for four overlapped CISS sequences. Each setting reports mIoU for old classes, newly introduced classes, and all classes. EWF is applied to the MiB and PLOP baselines. The largest gains occur in long or difficult sequences: MiB+EWF improves the all-class score over MiB by 33.3 points in 15-1 and 24.7 points in 10-1, while PLOP+EWF improves over PLOP by 12.4 and 21.4 points in those settings.

    Could not parse LaTeX table
  8. Knowl 8 — ADE20K Incremental Segmentation Results

    data/table

    The page-8 comparison evaluates final-task mIoU on the larger ADE20K dataset. EWF improves MiB from 25.9 to 32.1 in the challenging 100-5 sequence and from 29.2 to 33.2 in 100-10, corresponding to gains of 6.2 and 4.0 points respectively; the paper highlights the 6.2-point and 3.0-point improvements in its discussion using the selected comparison context. EWF also exceeds RC-IL in the 100-5 all-class result.

    Could not parse LaTeX table
  9. Knowl 9 — Fusion Strategy Ablation

    data/table

    The fusion ablation on PASCAL VOC 2012 with the 15-1 sequence compares no fusion, EMA, prediction-level model ensembling, and EWF. The values are mIoU after each incremental step. EWF is already substantially better at the first incremental step and retains the advantage as more tasks arrive; at step 5 it reaches 65.6 mIoU, compared with 37.3 for EMA, 37.2 for model ensembling, and 32.2 without fusion.

    Could not parse LaTeX table

    Relative to the no-fusion baseline, the paper reports improvements of 5.1 points for EMA, 5.0 points for model ensembling, and 33.3 points for EWF on the final 15-1 result, supporting the claimed benefit of parameter-space endpoint fusion.

  10. Knowl 10 — Dynamic Fusion-Factor Ablation and Class-Order Robustness

    data/table

    The paper compares the dynamic coefficient αt\alpha_t with fixed coefficients 0.20.2, 0.40.4, 0.60.6, and 0.80.8 on three PASCAL VOC settings. The dynamic rule obtains the best average mIoU, without manually tuning a coefficient for each class-incremental sequence.

    Could not parse LaTeX table

    The best fixed coefficient changes across settings: 0.20.2 is best for 10-1, 0.60.6 is best for 5-3, and the dynamic rule ties the best fixed value for 15-1 while producing the highest cross-setting average. In a separate experiment with five different class orders, the page-8 robustness chart reports that MiB+EWF has the strongest average performance and improved robustness relative to ILT, MiB, PLOP, and RC-IL.

  11. Knowl 11 — Integration with Auxiliary-Data Methods

    empirical result

    Although EWF is designed for settings without additional data, the paper integrates it with SSUL, an auxiliary-data method that uses saliency maps. Because SSUL freezes its backbone, the experiment makes the second backbone stage trainable so that its parameters can participate in EWF fusion. On PASCAL VOC 2012, SSUL+EWF improves SSUL for both old and new classes in the evaluated sequences.

    Could not parse LaTeX table

    The all-class score increases by 1.0 point for SSUL+EWF in both 15-1 and 10-1, showing that EWF can be layered on top of a method that already uses auxiliary information.

Coverage note — No substantial contributed method or result was omitted; qualitative prediction examples and the conclusion's proposed future investigations were not made separate knowls because they provide supporting visualization or plans rather than additional quantified contributions.

References

  1. 1.Jihwan Bang, Heesu Kim, YoungJoon Yoo, Jung-Woo Ha, and Jonghyun Choi. Rainbow memory: Continual learning with a memory of diverse samples. In CVPR, 2021. 2
  2. 2.Eden Belouadah and Adrian Popescu. Il2m: Class incremental learning with dual memory. In ICCV, pages 583–592, 2019. 2
  3. 3.Fabio Cermelli, Massimiliano Mancini, Samuel Rota Bulo, Elisa Ricci, and Barbara Caputo. Modeling the background for incremental learning in semantic segmentation. In CVPR, pages 9233–9242, 2020. 1, 2, 4, 5, 6, 7, 8
  4. 4.Fabio Cermelli, Massimiliano Mancini, Samuel Rota Bulo, Elisa Ricci, and Barbara Caputo. Modeling the background for incremental learning in semantic segmentation. In CVPR, pages 9233–9242, 2020. 2
  5. 5.Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. Co2l: Contrastive continual learning. In ICCV, pages 9516–9525, 2021. 2
  6. 6.Sungmin Cha, YoungJoon Yoo, Taesup Moon, et al. Ssul: Semantic segmentation with unknown label for exemplar-based class-incremental learning. NeurIPS, 34, 2021. 2, 5, 6, 7
  7. 7.Arslan Chaudhry, Puneet K Dokania, Thalaiyasingam Ajanthan, and Philip HS Torr. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In ECCV, 2018. 2
  8. 8.Arslan Chaudhry, Albert Gordo, Puneet Dokania, Philip Torr, and David Lopez-Paz. Using hindsight to anchor past knowledge in continual learning. In AAAI, volume 35, pages 6993–7001, 2021. 2
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE TPAMI, 40(4):834–848, 2017. 5
  10. 10.Yi-Hsin Chen, Wei-Yu Chen, Yu-Ting Chen, Bo-Cheng Tsai, Yu-Chiang Frank Wang, and Min Sun. No more discrimination: Cross city adaptation of road scene segmenters. In ICCV, pages 1992–2001, 2017. 1
  11. 11.Matthias Delange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Greg Slabaugh, and Tinne Tuytelaars. A continual learning survey: Defying forgetting in classification tasks. IEEE TPAMI, 2021. 2
  12. 12.Prithviraj Dhar, Rajat Vikram Singh, Kuan-Chuan Peng, Ziyan Wu, and Rama Chellappa. Learning without memorizing. In CVPR, 2019. 2
  13. 13.Thomas G Dietterich. Ensemble methods in machine learning. In International workshop on multiple classifier systems, pages 1–15. Springer, 2000. 8
  14. 14.Xiaohan Ding, Yuchen Guo, Guiguang Ding, and Jungong Han. Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks. In ICCV, October 2019. 3
  15. 15.Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun. Repvgg: Making vgg-style convnets great again. In CVPR, 2021. 3
  16. 16.Arthur Douillard, Yifu Chen, Arnaud Dapogny, and Matthieu Cord. Plop: Learning without forgetting for continual semantic segmentation. In CVPR, 2021. 1, 2, 4, 5, 6, 8
  17. 17.Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. Podnet: Pooled outputs distillation for small-tasks incremental learning. In ECCV, volume 12365, pages 86–102, 2020. 1, 2
  18. 18.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. http://www.pascal-network.org/challenges/VOC/voc2012/workshop/index.html. 5
  19. 19.Alberto Garcia-Garcia, Sergio Orts-Escolano, Sergiu Oprea, Victor Villena-Martinez, and Jose Garcia-Rodriguez. A review on deep learning techniques applied to semantic segmentation. arXiv preprint arXiv:1704.06857, 2017. 2
  20. 20.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin ´ Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Ghesh-laghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. NeurIPS, 33:21271–21284, 2020. 3
  21. 21.Tyler L Hayes, Kushal Kafle, Robik Shrestha, Manoj Acharya, and Christopher Kanan. Remind your neural network to prevent catastrophic forgetting. In ECCV, pages 466–483, 2020. 2
  22. 22.Kaiming He, Ross Girshick, and Piotr Dollar. Rethinking imagenet pre-training. In ICCV, pages 4918–4927, 2019. 1
  23. 23.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 5
  24. 24.Zilong Huang, Wentian Hao, Xinggang Wang, Mingyuan Tao, Jianqiang Huang, Wenyu Liu, and Xian-Sheng Hua. Half-real half-fake distillation for class-incremental semantic segmentation. arXiv preprint arXiv:2104.00875, 2021. 2
  25. 25.Christian Hane, Christopher Zach, Andrea Cohen, and Marc ¨ Pollefeys. Dense semantic 3d reconstruction. IEEE TPAMI, 39(9):1730–1743, 2017. 1
  26. 26.Ahmet Iscen, Jeffrey Zhang, Svetlana Lazebnik, and Cordelia Schmid. Memory-efficient incremental learning through feature adaptation. In ECCV, pages 699–715, 2020. 2
  27. 27.Menelaos Kanakis, David Bruggemann, Suman Saha, Stamatios Georgoulis, Anton Obukhov, and Luc Van Gool. Reparameterizing convolutions for incremental multi-task learning without task interference. In ECCV, pages 689–707, 2020. 2, 3
  28. 28.Chris Dongjoo Kim, Jinseo Jeong, Sangwoo Moon, and Gunhee Kim. Continual learning on noisy data streams via self-purified replay. In ICCV, pages 537–547, 2021. 2
  29. 29.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017. 2
  30. 30.Simon Kornblith, Jonathon Shlens, and Quoc V Le. Do better imagenet models transfer better? In CVPR, pages 2661–2671, 2019. 1
  31. 31.Zhizhong Li and Derek Hoiem. Learning without forgetting. IEEE TPAMI, 40(12):2935–2947, 2017. 5, 6
  32. 32.Tie Liu, Zejian Yuan, Jian Sun, Jingdong Wang, Nanning Zheng, Xiaoou Tang, and Heung-Yeung Shum. Learning to detect a salient object. IEEE TPAMI, 33(2):353–367, 2010. 6
  33. 33.Yaoyao Liu, Bernt Schiele, and Qianru Sun. Adaptive aggregation networks for class-incremental learning. In CVPR, 2021. 2
  34. 34.Andrea Maracani, Umberto Michieli, Marco Toldo, and Pietro Zanuttigh. Recall: Replay-based continual learning in semantic segmentation. In ICCV, pages 7026–7035, 2021. 6, 7
  35. 35.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In PsychologLearniny of learning and motivation, volume 24, pages 109–165. Elsevier, 1989. 1
  36. 36.Umberto Michieli and Pietro Zanuttigh. Incremental learning techniques for semantic segmentation. In ICCVW, 2019. 5, 6, 8
  37. 37.Umberto Michieli and Pietro Zanuttigh. Continual semantic segmentation via repulsion-attraction of sparse and disentangled latent representations. In CVPR, 2021. 1, 2, 5, 6
  38. 38.Boris T Polyak and Anatoli B Juditsky. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 30(4):838–855, 1992. 3, 4, 7, 8
  39. 39.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In CVPR, 2017. 2
  40. 40.Samuel Rota Bulo, Lorenzo Porzi, and Peter Kontschieder. In-place activated batchnorm for memory-optimized training of dnns. In CVPR, 2018. 5
  41. 41.Christian Simon, Piotr Koniusz, and MehrtasH Harandi. On learning the geodesic path for incremental learning. In CVPR, 2021. 2
  42. 42.Pravendra Singh, Pratik Mazumder, Piyush Rai, and Vinay P Namboodiri. Rectification-based knowledge retention for continual learning. In CVPR, pages 15282–15291, 2021. 2
  43. 43.Pravendra Singh, Vinay Kumar Verma, Pratik Mazumder, Lawrence Carin, and Piyush Rai. Calibrating cnns for lifelong learning. In NeurIPS, volume 33, 2020. 2
  44. 44.James Smith, Yen-Chang Hsu, Jonathan Balloch, Yilin Shen, Hongxia Jin, and Zsolt Kira. Always be dreaming: A new approach for data-free class-incremental learning. In ICCV, 2021. 2
  45. 45.Vinay Kumar Verma, Kevin J Liang, Nikhil Mehta, Piyush Rai, and Lawrence Carin. Efficient feature transformations for discriminative and generative continual learning. In CVPR, 2021. 2
  46. 46.Fu-Yun Wang, Da-Wei Zhou, Han-Jia Ye, and De-Chuan Zhan. Foster: Feature boosting and compression for class-incremental learning. arXiv preprint arXiv:2204.04662, 2022. 2
  47. 47.Shipeng Yan, Jiangwei Xie, and Xuming He. Der: Dynamically expandable representation for class incremental learning. In CVPR, 2021. 2
  48. 48.Shipeng Yan, Jiale Zhou, Jiangwei Xie, Songyang Zhang, and Xuming He. An em framework for online incremental learning of semantic segmentation. In ACM MM, pages 3052–3060, 2021. 2
  49. 49.Lu Yu, Xialei Liu, and Joost Van de Weijer. Self-training for class-incremental semantic segmentation. TNNLS, 2022. 6, 7
  50. 50.Chang-Bin Zhang, Jia-Wen Xiao, Xialei Liu, Ying-Cong Chen, and Ming-Ming Cheng. Representation compensation networks for continual semantic segmentation. In CVPR, 2022. 1, 2, 3, 5, 6, 8
  51. 51.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017. 5
  52. 52.Fei Zhu, Xu-Yao Zhang, Chuang Wang, Fei Yin, and Cheng-Lin Liu. Prototype augmentation and self-supervision for incremental learning. In CVPR, pages 5871–5880, 2021. 2
  53. 53.Kai Zhu, Wei Zhai, Yang Cao, Jiebo Luo, and Zheng-Jun Zha. Self-sustaining representation expansion for non-exemplar class-incremental learning. In CVPR, pages 9296–9305, 2022. 2

Citation

MLA
Xiao, J., et al. “Endpoints Weight Fusion for Class Incremental Semantic Segmentation”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7204–13, https://doi.org/10.1109/CVPR52729.2023.00696.
APA
Xiao, J., Zhang, C., Feng, J., Liu, X., van de Weijer, J., & Cheng, M. (2023). Endpoints Weight Fusion for Class Incremental Semantic Segmentation. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7204–7213. https://doi.org/10.1109/CVPR52729.2023.00696
Chicago
Xiao, J., C. Zhang, J. Feng, X. Liu, J. van de Weijer, and M. Cheng. 2023. “Endpoints Weight Fusion for Class Incremental Semantic Segmentation”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7204–13. https://doi.org/10.1109/CVPR52729.2023.00696.
Harvard
Xiao, J. et al. (2023) “Endpoints Weight Fusion for Class Incremental Semantic Segmentation”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 7204–7213. Available at: https://doi.org/10.1109/CVPR52729.2023.00696.
Vancouver
1. Xiao J, Zhang C, Feng J, Liu X, van de Weijer J, Cheng M (2023) Endpoints Weight Fusion for Class Incremental Semantic Segmentation. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 7204–7213

BibTeX

@inproceedings{Xiao_2023, title={Endpoints Weight Fusion for Class Incremental Semantic Segmentation}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00696}, DOI={10.1109/cvpr52729.2023.00696}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Xiao, Jia–Wen and Zhang, Chang–Bin and Feng, Jiekang and Liu, Xialei and van de Weijer, Joost and Cheng, Ming–Ming}, year={2023}, month=June, pages={7204–7213} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE