Understanding Imbalanced Semantic Segmentation Through Neural Collapse
Zhisheng ZhongJiequan CuiYibo YangXiaoyang WuXiaojuan QiXiangyu ZhangJiaya Jia
Explains how class imbalance and contextual correlation disrupt neural collapse in semantic segmentation and introduces a feature center regularizer to recover maximally separated representations, achieving state-of-the-art results on challenging 2D and 3D benchmarks including ScanNet200.
Semantic segmentation—the task of classifying every pixel in an image or point in a 3D scene—is a core technology for autonomous systems, robotics, and spatial computing. However, models systematically struggle with severe class imbalance, as common categories like walls or roads dominate visual scenes while rare, high-consequence categories occupy only minor fractions of data. In standard image recognition, training models on balanced datasets leads to an ideal geometric state called neural collapse, where learned class representations separate maximally and evenly in feature space. In complex real-world segmentation tasks, this balanced separation fails, leading to poor identification of minority classes.
The article investigates why this balanced geometric separation breaks down in semantic segmentation and evaluates a new training framework designed to enforce structured feature separation, specifically aimed at improving accuracy on underrepresented classes.
The authors conducted empirical analyses across 2D image datasets (ADE20K and COCO-Stuff164K) and 3D point cloud datasets (ScanNetv2 and ScanNet200). They demonstrated that spatial correlations and massive data imbalances cause the gradients of dominant classes to overwhelm minor classes, pushing their representations into clustered, indistinguishable positions. To resolve this, the authors developed a dual-branch training method called the Center Collapse Regularizer. This approach maintains a standard learnable classifier for dense pixel/point predictions while adding an auxiliary branch during training that extracts class-level feature centers and forces them into a fixed, maximally separated geometric structure (a simplex equiangular tight frame).
Experimental results show substantial performance gains across both 2D and 3D benchmarks without adding computational overhead during deployment. On the 3D ScanNet200 benchmark, the proposed method achieved a new record on the competitive test leaderboard, improving mean Intersection-over-Union (a standard segmentation overlap metric) by up to 6.8 percentage points over previous baselines, driven largely by gains of 4.6 to 7.5 percentage points in rare 'tail' classes. On 2D image benchmarks, the regularizer consistently improved segmentation accuracy across diverse neural network architectures by 0.5 to 6.4 percentage points. The method also proved complementary to existing loss functions, providing an additional 1.0 to 2.0 percentage point gain when combined with standard specialized segmentation losses.
These findings show that regularizing feature representations at the class-center level resolves the fundamental gradient imbalance that degrades standard pixel-level training. In practice, this enables computer vision systems to reliably detect rare objects without sacrificing the flexibility required to interpret correlated visual contexts. Because the auxiliary regularization branch is discarded after training, organizations can deploy these higher-performing models directly into production without increasing inference latency, memory usage, or compute costs.
Technical leaders and engineering teams deploying dense visual recognition models should consider integrating class-center regularization into their training pipelines, particularly for safety-critical applications with long-tailed distributions. Given that the technique is compatible with standard network backbones and loss functions, teams can run low-risk pilot evaluations on existing datasets to validate accuracy gains. Future engineering efforts should explore applying this geometric framework to other imbalanced dense prediction tasks, such as panoptic and instance segmentation.
- Paper: Decoupling Representation and Classifier for Long-Tailed Recognition, Bingyi Kang et al. (2019). This seminal work decouples representation learning from classifier adjustment in long-tailed visual recognition, establishing the foundation for analyzing how feature space geometry interacts with class imbalance.
- Paper: Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss, Kaidi Cao et al. (2019). It introduces margin-based regularization and optimization strategies for long-tailed distributions, providing essential context for mitigating minority-class collapse during deep network training.
- Paper: Class-Balanced Loss Based on Effective Number of Samples, Yin Cui et al. (2019). This paper establishes the theoretical framework of sample volume overlap and class-balanced loss formulation, which serves as a standard baseline and conceptual antecedent for addressing class imbalance.
- Paper: COCO-Stuff: Thing and Stuff Classes in Context, Holger Caesar et al. (2016). It introduces the COCO-Stuff dataset and formalizes the distinct challenges of contextual 'stuff' and 'thing' segmentation under heavy spatial and frequency imbalances evaluated in the source.
- Paper: Large-Scale Point Cloud Semantic Segmentation with Superpoint Graphs, Loic Landrieu et al. (2017). This work establishes core methodology for large-scale 3D point cloud segmentation, providing the foundational spatial representation principles utilized in 3D dense recognition benchmarks.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It introduces end-to-end fully convolutional pixel-level classification, which forms the fundamental architectural paradigm analyzed and regularized in the source.
- Paper: Feature learning in deep classifiers through Intermediate Neural Collapse, Akshay Rangamani et al. (2023). This work investigates intermediate neural collapse dynamics across deep classifier layers, directly complementing and extending geometric collapse analyses beyond the final dense prediction head.
- Paper: Global and Local Mixture Consistency Cumulative Learning for Long-tailed Visual Recognitions, Fei Du et al. (2023). It builds on representation consistency and class rebalancing strategies to improve feature robustness under severe long-tailed imbalances in visual recognition tasks.
- Paper: Superpoint Transformer for 3D Scene Instance Segmentation, Jiahao Sun et al. (2023). This paper applies structured point clustering to 3D scene instance segmentation, exploring downstream dense prediction in complex spatial point clouds.
