Endpoints Weight Fusion for Class Incremental Semantic Segmentation
Jia-Wen XiaoChang-Bin ZhangJiekang FengXialei LiuJoost van de WeijerMing-Ming Cheng
Proposes a dynamic parameter-averaging strategy that fuses starting and ending network weights from each incremental step to mitigate catastrophic forgetting in semantic segmentation without increasing model size or requiring extra training.
Visual applications increasingly require computer vision systems to learn new object categories over time. However, updating these deep learning models with newly labeled data causes catastrophic forgetting, where the system overwrites and forgets previously learned classes. Storing previous data to retrain models from scratch creates severe data storage burdens and privacy risks, while existing mathematical constraints alone fail to retain strong historical memory.
The article demonstrates a model parameter aggregation framework called Endpoints Weight Fusion. The primary objective is to evaluate whether dynamically fusing the model weights from the start of a training stage with the weights at the end of that stage can successfully preserve past knowledge without expanding model size, requiring extra training steps, or relying on stored historical data.
The evaluation utilized benchmark visual segmentation datasets, specifically PASCAL VOC 2012 and ADE20K, across diverse incremental learning sequences ranging from two to eleven sequential steps. The framework pairs a standard knowledge distillation scheme—which constrains the models to remain close in parameter space during training—with a post-training parameter fusion calculation based on the ratio of new to total seen classes.
The analysis revealed three key findings. First, integrating the fusion strategy with existing baseline methods substantially boosted overall segmentation accuracy, improving baseline mean intersection over union by 21.4% to 33.3% in long sequential tasks. Second, the method exhibited significant performance gains on historical classes while maintaining high plasticity on newly introduced classes. Third, dynamic fusion factor calculation outperformed fixed weighting parameters across all tested scenarios and proved far superior to standard moving average updates, which overemphasize early collapsed gradients.
These findings indicate that teams deploying computer vision systems can continuously incorporate new object categories without inflating operational inference costs or expanding model footprints. The parameter-level fusion bypasses costly full-dataset retraining cycles, thereby reducing computational overhead and mitigating regulatory privacy concerns tied to long-term data storage.
Organizations developing continuous learning systems should adopt endpoint weight fusion on top of existing regularization techniques rather than relying on distillation constraints alone. Development teams should employ dynamic weighting ratios derived from category counts rather than manual hyperparameter tuning. Further research is recommended to explore the theoretical mechanics of weight fusion and evaluate its viability across other computer vision domains.
The findings are supported by consistent results across standard academic benchmarks and varied class arrival sequences. However, empirical testing remains bounded by standard supervised segmentation frameworks and Convolutional Neural Network architectures, meaning performance should be verified when deploying across alternative network designs or drastically different operational environments.
- Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This foundational work introduces knowledge distillation as a regularization mechanism to preserve old task representations without historical data, which EWF directly builds upon and enhances via parameter fusion.
- Paper: iCaRL: Incremental Classifier and Representation Learning, Sylvestre-Alvise Rebuffi et al. (2016). This paper establishes the standard class-incremental learning framework combining classification loss and distillation loss, providing essential context for the catastrophic forgetting challenges tackled in EWF.
- Paper: Overcoming catastrophic forgetting in neural networks, James Kirkpatrick et al. (2017). This foundational paper formulates parameter-space regularization to mitigate catastrophic forgetting in neural networks, contextualizing EWF's strategy of parameter-space fusion and distance optimization.
- Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This survey provides a comprehensive taxonomy of continual learning strategies and stability-plasticity trade-offs, establishing the conceptual background for regularization- and distillation-based incremental methods.
- Paper: Re-basin via implicit Sinkhorn differentiation, Fidel A. Guerrero-Peña et al. (2023). This work explores model parameter fusion in continual learning by optimizing permutation alignments to lower interpolation barriers between independently updated models.
- Paper: Localizing Task Information for Improved Model Merging and Compression, Ke Wang et al. (2024). This study advances the analysis of weight interpolation and parameter merging across tasks by diagnosing cross-task interference and proposing localized task masks.
