Understanding Imbalanced Semantic Segmentation Through Neural Collapse

Zhisheng ZhongJiequan CuiYibo YangXiaoyang WuXiaojuan QiXiangyu ZhangJiaya Jia

article2023CVPR75 citations

Explains how class imbalance and contextual correlation disrupt neural collapse in semantic segmentation and introduces a feature center regularizer to recover maximally separated representations, achieving state-of-the-art results on challenging 2D and 3D benchmarks including ScanNet200.

Listen

Semantic segmentation—the task of classifying every pixel in an image or point in a 3D scene—is a core technology for autonomous systems, robotics, and spatial computing. However, models systematically struggle with severe class imbalance, as common categories like walls or roads dominate visual scenes while rare, high-consequence categories occupy only minor fractions of data. In standard image recognition, training models on balanced datasets leads to an ideal geometric state called neural collapse, where learned class representations separate maximally and evenly in feature space. In complex real-world segmentation tasks, this balanced separation fails, leading to poor identification of minority classes.

The article investigates why this balanced geometric separation breaks down in semantic segmentation and evaluates a new training framework designed to enforce structured feature separation, specifically aimed at improving accuracy on underrepresented classes.

The authors conducted empirical analyses across 2D image datasets (ADE20K and COCO-Stuff164K) and 3D point cloud datasets (ScanNetv2 and ScanNet200). They demonstrated that spatial correlations and massive data imbalances cause the gradients of dominant classes to overwhelm minor classes, pushing their representations into clustered, indistinguishable positions. To resolve this, the authors developed a dual-branch training method called the Center Collapse Regularizer. This approach maintains a standard learnable classifier for dense pixel/point predictions while adding an auxiliary branch during training that extracts class-level feature centers and forces them into a fixed, maximally separated geometric structure (a simplex equiangular tight frame).

Experimental results show substantial performance gains across both 2D and 3D benchmarks without adding computational overhead during deployment. On the 3D ScanNet200 benchmark, the proposed method achieved a new record on the competitive test leaderboard, improving mean Intersection-over-Union (a standard segmentation overlap metric) by up to 6.8 percentage points over previous baselines, driven largely by gains of 4.6 to 7.5 percentage points in rare 'tail' classes. On 2D image benchmarks, the regularizer consistently improved segmentation accuracy across diverse neural network architectures by 0.5 to 6.4 percentage points. The method also proved complementary to existing loss functions, providing an additional 1.0 to 2.0 percentage point gain when combined with standard specialized segmentation losses.

These findings show that regularizing feature representations at the class-center level resolves the fundamental gradient imbalance that degrades standard pixel-level training. In practice, this enables computer vision systems to reliably detect rare objects without sacrificing the flexibility required to interpret correlated visual contexts. Because the auxiliary regularization branch is discarded after training, organizations can deploy these higher-performing models directly into production without increasing inference latency, memory usage, or compute costs.

Technical leaders and engineering teams deploying dense visual recognition models should consider integrating class-center regularization into their training pipelines, particularly for safety-critical applications with long-tailed distributions. Given that the technique is compatible with standard network backbones and loss functions, teams can run low-risk pilot evaluations on existing datasets to validate accuracy gains. Future engineering efforts should explore applying this geometric framework to other imbalanced dense prediction tasks, such as panoptic and instance segmentation.

Cover for Understanding Imbalanced Semantic Segmentation Through Neural Collapse

Abstract

A recent study has shown a phenomenon called neural collapse in that the within-class means of features and the classifier weight vectors converge to the vertices of a simplex equiangular tight frame at the terminal phase of training for classification. In this paper, we explore the corresponding structures of the last-layer feature centers and classifiers in semantic segmentation. Based on our empirical and theoretical analysis, we point out that semantic segmentation naturally brings contextual correlation and imbalanced distribution among classes, which breaks the equiangular and maximally separated structure of neural collapse for both feature centers and classifiers. However, such a symmetric structure is beneficial to discrimination for the minor classes. To preserve these advantages, we introduce a regularizer on feature centers to encourage the network to learn features closer to the appealing structure in imbalanced semantic segmentation. Experimental results show that our method can bring significant improvements on both 2D and 3D semantic segmentation benchmarks. Moreover, our method ranks 1st and sets a new record (+6.8% mIoU) on the ScanNet200 test leaderboard.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Neural Collapse Observations
  • 3.1. Neural Collapse in Recognition
  • 3.2. Neural Collapse in Semantic Segmentation
  • 4. Main Approach
  • 4.1. Motivation
  • 4.2. Center Collapse Regularizer
  • 4.3. Empirical Support
  • 4.4. Theoretical Support
  • 5. Experiments
  • 5.1. Datasets and Implementation Details
  • 5.2. Ablation Study
  • 5.3. Main Results
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Simplex Equiangular Tight Frame and the Maximal Separation Property

    definition

    A simplex Equiangular Tight Frame (ETF) is a collection of KK vectors M=[m1,…,mK]∈Rd×KM = [m_1, \dots, m_K] \in \mathbb{R}^{d \times K} in a feature space of dimension d≥Kd \ge K defined by:

    M=KK−1U(IK−1K1K1K⊤)M = \sqrt{\frac{K}{K - 1}} U \left( I_K - \frac{1}{K} \mathbf{1}_K \mathbf{1}_K^\top \right)

    where IK∈RK×KI_K \in \mathbb{R}^{K \times K} is the identity matrix, 1K∈RK\mathbf{1}_K \in \mathbb{R}^K is an all-ones vector, and U∈Rd×KU \in \mathbb{R}^{d \times K} is a matrix with orthonormal columns satisfying U⊤U=IKU^\top U = I_K.

    All vectors in a simplex ETF have an identical ℓ2\ell_2 norm and equal pairwise inner products:

    mi⊤mj=KK−1δi,j−1K−1,∀i,j∈{1,…,K}m_i^\top m_j = \frac{K}{K - 1} \delta_{i,j} - \frac{1}{K - 1}, \quad \forall i, j \in \{1, \dots, K\}

    where δi,j\delta_{i,j} is the Kronecker delta (equal to 11 when i=ji=j and 00 otherwise).

    Furthermore, for any collection of KK normalized vectors W=[w1,…,wK]∈Rd×KW = [w_1, \dots, w_K] \in \mathbb{R}^{d \times K} with d≥Kd \ge K and wk⊤wk=1w_k^\top w_k = 1 for all k∈{1,…,K}k \in \{1, \dots, K\}, the maximal separation value satisfies:

    max⁡k≠k′cos⁡(wk,wk′)≥−1K−1\max_{k \neq k'} \cos(w_k, w_{k'}) \ge -\frac{1}{K - 1}

    Equality holds if and only if WW is a simplex ETF, meaning all distinct vector pairs attain the minimal possible pairwise cosine similarity of −1/(K−1)-1/(K - 1) simultaneously.

  2. Knowl 2 — Center Collapse Regularizer for Imbalanced Semantic Segmentation

    model/method

    The Center Collapse Regularizer (CeCo) is a dual-branch training framework designed to mitigate class imbalance in 2D pixel-wise and 3D point-wise semantic segmentation while preserving contextual adaptability.

    1. Point/Pixel Recognition Branch: Given an input x∈RN×sx \in \mathbb{R}^{N \times s} (NN pixels or points, ss color/spatial channels), a segmentation backbone extracts representations Z=[z1,…,zN]⊤∈RN×dZ = [z_1, \dots, z_N]^\top \in \mathbb{R}^{N \times d}. A learnable point/pixel linear classifier produces predictions supervised by the standard per-element cross-entropy loss LPR(Z,y)\mathcal{L}_{PR}(Z, y), where y∈{1,…,K}Ny \in \{1, \dots, K\}^N denotes the ground-truth labels for KK classes.

    2. Center Regularization Branch: For each semantic class k∈{1,…,K}k \in \{1, \dots, K\}, the feature center zˉk\bar{z}_k is computed over all representations assigned to class kk:

    zˉk=1nk∑i:yi=kzi\bar{z}_k = \frac{1}{n_k} \sum_{i: y_i = k} z_i

    where nkn_k is the number of samples in ZZ belonging to class kk. The class centers Zˉ=[zˉ1,…,zˉK]∈Rd×K\bar{Z} = [\bar{z}_1, \dots, \bar{z}_K] \in \mathbb{R}^{d \times K} are evaluated against a fixed simplex Equiangular Tight Frame (ETF) classifier W∗=[w1∗,…,wK∗]∈Rd×KW^* = [w_1^*, \dots, w_K^*] \in \mathbb{R}^{d \times K}, whose columns satisfy:

    wk∗⊤wk′∗=α2(Kδk,k′K−1−1K−1),∀k,k′∈{1,…,K}w_k^{*\top} w_{k'}^* = \alpha^2 \left( \frac{K \delta_{k,k'}}{K - 1} - \frac{1}{K - 1} \right), \quad \forall k, k' \in \{1, \dots, K\}

    where α>0\alpha > 0 is a weight scaling hyperparameter. The center collapse loss LCR(Zˉ,W∗)\mathcal{L}_{CR}(\bar{Z}, W^*) is computed via cross-entropy over the class centers:

    LCR(Zˉ,W∗)=−∑k=1Klog⁡(exp⁡(zˉk⊤wk∗)∑k′=1Kexp⁡(zˉk⊤wk′∗))\mathcal{L}_{CR}(\bar{Z}, W^*) = - \sum_{k=1}^K \log \left( \frac{\exp(\bar{z}_k^\top w_k^*)}{\sum_{k'=1}^K \exp(\bar{z}_k^\top w_{k'}^*)} \right)

    1. Total Objective and Inference: The overall loss function is:

    Ltotal=LPR(Z,y)+λLCR(Zˉ,W∗)\mathcal{L}_{total} = \mathcal{L}_{PR}(Z, y) + \lambda \mathcal{L}_{CR}(\bar{Z}, W^*)

    where λ>0\lambda > 0 is a loss balancing hyperparameter. During inference and evaluation, the center regularization branch is completely discarded, and only the point/pixel recognition branch is evaluated, adding no extra computational cost over the backbone.

  3. Knowl 3 — Gradient Analysis of Fixed Simplex ETF Center Regularization

    theoretical result

    Gradient analysis explains why adopting a fixed simplex Equiangular Tight Frame (ETF) classifier in the Center Collapse Regularizer prevents minority class degradation in imbalanced semantic segmentation:

    1. Center Classifier Gradient with Learnable Weights: If the center classifier W=[w1,…,wK]∈Rd×KW = [w_1, \dots, w_K] \in \mathbb{R}^{d \times K} were learnable under the cross-entropy loss LCR\mathcal{L}_{CR}, the gradient with respect to class vector wkw_k is:

    ∂LCR∂wk=∑i:yi=k(pk(zˉk)−1)zink⏟within-class+∑k′≠kK−1∑j:yj=k′pk(zˉk′)zjnk′⏟between-class\frac{\partial \mathcal{L}_{CR}}{\partial w_k} = \underbrace{\sum_{i: y_i = k} (p_k(\bar{z}_k) - 1) \frac{z_i}{n_k}}_{\text{within-class}} + \underbrace{\sum_{k' \neq k}^{K-1} \sum_{j: y_j = k'} p_k(\bar{z}_{k'}) \frac{z_j}{n_{k'}}}_{\text{between-class}}

    where pk(zˉ)=exp⁡(zˉ⊤wk)/∑m=1Kexp⁡(zˉ⊤wm)p_k(\bar{z}) = \exp(\bar{z}^\top w_k) / \sum_{m=1}^K \exp(\bar{z}^\top w_m), and nkn_k is the number of samples in class kk. The between-class component accumulates push terms across all other classes against the within-class pull terms. For minor classes (nk≪nk′n_k \ll n_{k'}), between-class forces dominate, collapsing and corrupting decision boundaries of minor classes. Fixing W=W∗W = W^* eliminates classifier gradient updates and avoids this gradient imbalance.

    1. Feature Gradient with Fixed ETF Weights: The gradient of LCR\mathcal{L}_{CR} with respect to a point/pixel feature ziz_i belonging to class kk is:

    ∂LCR∂zi=1nk∑k′=1Kpk′(zˉk)(wk′∗−wk∗)\frac{\partial \mathcal{L}_{CR}}{\partial z_i} = \frac{1}{n_k} \sum_{k'=1}^K p_{k'}(\bar{z}_k) (w_{k'}^* - w_k^*)

    In unconstrained imbalanced learning, minority collapse occurs such that ∥wk′−wk∥→0\|w_{k'} - w_k\| \to 0 under extreme imbalance. Under a fixed simplex ETF classifier W∗W^*, the distance between distinct classifier vectors is strictly preserved at ∥wk′∗−wk∗∥=2KK−1\|w_{k'}^* - w_k^*\| = \sqrt{\frac{2K}{K - 1}} for all k≠k′k \neq k'. This maintains gradient separation between minor classes and prevents minor class representations from collapsing into identical directions.

  4. Knowl 4 — Breakdown of Natural Neural Collapse in Semantic Segmentation

    empirical result

    In balanced image recognition, deep neural networks exhibit the neural collapse phenomenon at convergence: normalized within-class feature means and normalized linear classifier vectors converge to the same simplex Equiangular Tight Frame (ETF), driving the standard deviation of pairwise cosine similarities Stdk≠k′(cos⁡(z^k,z^k′))\text{Std}_{k \neq k'}(\cos(\hat{z}_k, \hat{z}_{k'})) and Stdk≠k′(cos⁡(w^k,w^k′))\text{Std}_{k \neq k'}(\cos(\hat{w}_k, \hat{w}_{k'})) toward zero.

    In contrast, empirical measurements on 2D image segmentation (ADE20K, COCO-Stuff164K) and 3D point cloud segmentation (ScanNetv2, ScanNet200) demonstrate that natural neural collapse fails to occur:

    • The standard deviations of pairwise cosine similarities for both feature centers and classifier weights remain high (0.10.1 to 0.40.4) throughout training, in contrast to the near-zero values (≈0.02\approx 0.02) reached in image classification (e.g., CIFAR-100).
    • Unlike classification benchmarks where classes have low mutual correlation and balanced sample sizes, semantic segmentation inherently features strong contextual correlation among neighboring classes and severe sample frequency imbalances (e.g., imbalanced factor nmax⁡/nmin⁡=37,256n_{\max}/n_{\min} = 37{,}256 for ScanNet200 and 827827 for ADE20K).
    • Class point/pixel frequencies correlate weakly with per-class segmentation accuracy (Pearson r=0.32r = 0.32 on ScanNet200, r=0.31r = 0.31 on ADE20K), whereas class feature center frequencies show higher correlation with accuracy (r=0.51r = 0.51 on ScanNet200, r=0.49r = 0.49 on ADE20K), indicating that rebalancing in the feature center space is better suited for segmentation than point/pixel-level reweighting.
  5. Knowl 5 — Ablation Study on Classifier Learnability Configurations

    data/table

    An ablation study on ScanNet200 validation (MinkowskiNet) and ADE20K validation (Swin-T backbone, single-scale inference) compares different combinations of learnable versus fixed simplex Equiangular Tight Frame (ETF) configurations for the Point/Pixel Classifier (PC) and the Center Regularization Classifier (CC).

    PC CC ScanNet200 (mIoU %) ADE20K (mIoU %)
    Learned – (Baseline) 27.8 44.5
    Fixed Fixed 28.3 44.8
    Fixed Learned 27.2 44.0
    Learned Fixed (CeCo) 30.3 (+2.5) 45.7 (+1.2)
    Learned Learned 26.9 43.7

    The experiments show:

    1. Fixing the point/pixel classifier (PC) to an ETF degrades performance compared to a learnable classifier because point/pixel recognition requires adaptive modeling of contextual correlations.
    2. Making the center classifier (CC) learnable degrades accuracy (26.9%26.9\% and 43.7%43.7\%) below the baseline due to imbalanced gradient interference from dominant classes.
    3. Combining a learnable point/pixel classifier with a fixed ETF center classifier (the CeCo architecture) yields the highest mIoU (30.3%30.3\% on ScanNet200, 45.7%45.7\% on ADE20K).
  6. Knowl 6 — Regularization Weight Sensitivity and Loss Function Orthogonality

    data/table

    The sensitivity of the Center Collapse Regularizer (CeCo) to the loss weighting hyperparameter λ\lambda and its compatibility with other specialized segmentation losses (Dice loss and Lovász-Softmax loss) were evaluated on ScanNet200 (MinkowskiNet) and ADE20K (Swin-T, single-scale).

    Sensitivity to λ\lambda:

    Loss Weight λ\lambda ScanNet200 (mIoU %) ADE20K (mIoU %)
    λ=0.0\lambda = 0.0 (Baseline) 27.8 44.5
    λ=0.1\lambda = 0.1 28.6 44.8
    λ=0.2\lambda = 0.2 29.4 45.0
    λ=0.3\lambda = 0.3 30.1 45.2
    λ=0.4\lambda = 0.4 30.3 (+2.5) 45.3
    λ=0.5\lambda = 0.5 29.7 45.7 (+1.2)
    λ=0.6\lambda = 0.6 28.7 44.9

    Optimal results are achieved at λ=0.4\lambda = 0.4 on ScanNet200 and λ=0.5\lambda = 0.5 on ADE20K, with gains observed across the entire range λ∈[0.1,0.6]\lambda \in [0.1, 0.6].

    Orthogonality to Dice and Lovász loss:

    Loss Type ScanNet200 (mIoU %) ADE20K (mIoU %)
    CE (Baseline) 27.8 44.5
    + Dice 28.8 44.8
    + Dice + CeCo 30.6 (+1.8) 45.8 (+1.0)
    + Lovász 30.0 45.0
    + Lovász + CeCo 32.0 (+2.0) 46.5 (+1.5)

    CeCo provides consistent 1.0%1.0\% to 2.0%2.0\% mIoU gains when combined with Dice loss or Lovász loss, confirming its orthogonality to standard segmentation losses.

  7. Knowl 7 — 3D Point Cloud Semantic Segmentation Benchmark Results on ScanNet200

    data/table

    Performance of the Center Collapse Regularizer (CeCo) evaluated on the 200-class ScanNet200 3D indoor point cloud segmentation benchmark using the MinkowskiNet backbone, stratified across Head, Common, and Tail categories according to class point frequency.

    ScanNet200 Validation Set Results:

    Method Head (mIoU %) Common (mIoU %) Tail (mIoU %) All (mIoU %)
    MinkowskiNet 48.3 19.1 7.9 25.1
    Instance Sampling 48.2 18.9 9.2 25.4
    Class-Balanced Focal 48.1 20.2 9.3 25.8
    SupCon 48.5 19.1 10.3 26.0
    CSC 49.4 19.5 10.3 26.5
    LG (CLIP) 50.4 22.8 10.1 27.7
    LG 51.5 22.7 12.5 28.9
    CeCo 51.2 22.9 (+0.2) 17.1 (+4.6) 30.3 (+1.4)
    CeCo (Lovász) 52.4 (+0.9) 26.2 (+3.5) 17.9 (+5.4) 32.0 (+3.1)

    ScanNet200 Test Set Results:

    Method Head (mIoU %) Common (mIoU %) Tail (mIoU %) All (mIoU %)
    MinkowskiNet 46.3 15.4 10.2 25.3
    CSC 45.5 17.1 7.9 24.9
    LG 48.5 18.4 10.6 27.2
    CeCo (Lovász) 52.1 (+3.6) 23.6 (+5.2) 15.2 (+4.6) 31.7 (+4.5)
    CeCo* (Lovász) 55.1 (+6.6) 24.7 (+6.3) 18.1 (+7.5) 34.0 (+6.8)

    Note: CeCo∗\text{CeCo}^* denotes an ensemble of three CeCo models. On the official ScanNet200 test benchmark, CeCo achieves 34.0%34.0\% overall mIoU (+6.8% over the MinkowskiNet baseline), with the largest improvements concentrated in the Tail (+7.5%) and Common (+6.3%) subsets.

  8. Knowl 8 — 2D Semantic Segmentation Benchmark Results on ADE20K

    data/table

    Evaluation of the Center Collapse Regularizer (CeCo) across multiple semantic segmentation heads and convolutional/transformer backbones on the 150-class ADE20K dataset under single-scale (s.s.) and multi-scale (m.s.) inference protocols.

    Overall ADE20K Results:

    Method Backbone mIoU (s.s. %) mIoU (m.s. %)
    OCRNet HRNet-W18 39.3 40.8
    + CeCo HRNet-W18 41.8 (+2.5) 43.5 (+2.7)
    DLV3P ResNet-50 43.9 44.9
    + CeCo ResNet-50 45.0 (+1.1) 46.4 (+1.5)
    OCRNet HRNet-W48 43.2 44.9
    + CeCo HRNet-W48 44.5 (+1.3) 46.1 (+1.2)
    UperNet ResNet-101 43.8 44.8
    + CeCo ResNet-101 44.8 (+1.0) 46.1 (+1.3)
    DLV3P ResNet-101 45.5 46.4
    + CeCo ResNet-101 46.7 (+1.2) 48.0 (+1.6)
    UperNet Swin-T 44.5 45.8
    + CeCo Swin-T 45.7 (+1.2) 47.6 (+1.8)
    UperNet Swin-B 50.0 51.7
    + CeCo Swin-B 51.2 (+1.2) 52.9 (+1.2)
    UperNet BEiT-L 56.7 57.0
    + CeCo BEiT-L 57.3 (+0.6) 57.7 (+0.7)

    Imbalanced Category Breakdown on ADE20K (with DeepLabV3+):

    Method Head (mIoU %) Common (mIoU %) Tail (mIoU %) All (mIoU %)
    DLV3P (ResNet-50) 67.7 48.3 36.4 44.9
    + DisAlign 67.7 48.6 37.8 45.7
    + CeCo 67.7 48.7 (+0.1) 39.0 (+1.2) 46.4 (+0.7)
    DLV3P (ResNet-101) 68.7 49.0 38.4 46.4
    + DisAlign 68.7 49.4 39.6 47.1
    + CeCo 68.8 (+0.1) 49.4 40.9 (+1.3) 48.0 (+0.9)

    CeCo consistently improves performance across both CNN (ResNet, HRNet) and vision transformer (Swin, BEiT) backbones, yielding superior improvements on the Tail subset compared to long-tailed frameworks like DisAlign.

  9. Knowl 9 — 2D Semantic Segmentation Benchmark Results on COCO-Stuff164K

    data/table

    Evaluation of the Center Collapse Regularizer (CeCo) across CNN and transformer architectures on the 171-class COCO-Stuff164K benchmark under single-scale (s.s.) and multi-scale (m.s.) inference:

    Method Backbone mIoU (s.s. %) mIoU (m.s. %)
    OCRNet HRNet-W18 31.6 32.4
    + CeCo HRNet-W18 38.0 (+6.4) 38.8 (+6.4)
    UperNet ResNet-50 39.9 40.3
    + CeCo ResNet-50 41.2 (+1.3) 41.7 (+1.4)
    DLV3P ResNet-50 40.9 41.5
    + CeCo ResNet-50 42.7 (+1.8) 43.5 (+2.0)
    OCRNet HRNet-W48 40.4 41.7
    + CeCo HRNet-W48 41.8 (+1.4) 43.3 (+1.6)
    UperNet ResNet-101 41.2 41.5
    + CeCo ResNet-101 42.3 (+1.1) 42.8 (+1.3)
    DLV3P ResNet-101 42.4 43.0
    + CeCo ResNet-101 43.9 (+1.5) 44.6 (+1.6)
    UperNet Swin-T 43.8 44.6
    + CeCo Swin-T 44.5 (+0.7) 45.2 (+0.6)
    UperNet Swin-B 47.7 48.6
    + CeCo Swin-B 48.2 (+0.5) 49.2 (+0.6)

    Integrating CeCo during training achieves consistent mIoU gains across all backbones, notably improving OCRNet with HRNet-W18 by +6.4%+6.4\% mIoU and DLV3P with ResNet-50 by +2.0%+2.0\% multi-scale mIoU.

Coverage note — No substantial contributed material was omitted; the knowls cover the foundational ETF properties, the CeCo method, theoretical gradient analysis, natural neural collapse failure observations in segmentation, and empirical evaluations across all benchmarks (ScanNet200, ADE20K, and COCO-Stuff164K).

References

  1. 1.MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020.
  2. 2.Md Amirul Islam, Mrigank Rochan, Neil DB Bruce, and Yang Wang. Gated feedback refinement network for dense image labeling. In CVPR, 2017.
  3. 3.Iro Armeni, Ozan Sener, Amir R. Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3D semantic parsing of large-scale indoor spaces. In CVPR, 2016.
  4. 4.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. SegNet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE TPAMI, 2017.
  5. 5.Hangbo Bao, Li Dong, and Furu Wei. BEiT: BERT pre-training of image transformers. In ICLR, 2022.
  6. 6.J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, C. Stachniss, and J. Gall. SemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences. In ICCV, 2019.
  7. 7.Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. Pearson correlation coefficient. In Noise reduction in speech processing, pages 1–4. Springer, 2009.
  8. 8.Maxim Berman, Amal Rannen Triki, and Matthew B Blaschko. The Lovász-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. In CVPR, pages 4413–4421, 2018.
  9. 9.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. COCO-Stuff: Thing and stuff classes in context. In CVPR, 2018.
  10. 10.Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma. Learning imbalanced datasets with label-distribution-aware margin loss. In NeurIPS, volume 32, 2019.
  11. 11.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected CRFs. In ICLR, 2015.
  12. 12.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE TPAMI, 2017.
  13. 13.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017.
  14. 14.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  15. 15.Christopher Choy, JunYoung Gwak, and Silvio Savarese. 4D spatio-temporal ConvNets: Minkowski convolutional neural networks. In CVPR, pages 3075–3084, 2019.
  16. 16.Ruihang Chu, Yukang Chen, Tao Kong, Lu Qi, and Lei Li. ICM-3D: Instantiated category modeling for 3D instance segmentation. IEEE Robotics and Automation Letters, 7(1):57–64, 2021.
  17. 17.Ruihang Chu, Xiaoqing Ye, Zhengzhe Liu, Xiao Tan, Xiaojuan Qi, Chi-Wing Fu, and Jiaya Jia. TWIST: Two-way inter-label self-training for semi-supervised 3D instance segmentation. In CVPR, pages 1100–1109, 2022.
  18. 18.Jiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu, and Jiaya Jia. Parametric contrastive learning. In ICCV, 2021.
  19. 19.Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie. Class-balanced loss based on effective number of samples. In CVPR, pages 9268–9277, 2019.
  20. 20.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3D reconstructions of indoor scenes. In CVPR, 2017.
  21. 21.Cong Fang, Hangfeng He, Qi Long, and Weijie J Su. Exploring deep neural networks via layer-peeled model: Minority collapse in imbalanced training. Proceedings of the National Academy of Sciences, 118(43), 2021.
  22. 22.Florian Graf, Christoph Hofer, Marc Niethammer, and Roland Kwitt. Dissecting supervised contrastive learning. In ICML, pages 3821–3830. PMLR, 2021.
  23. 23.Benjamin Graham, Martin Engelcke, and Laurens van der Maaten. 3D semantic segmentation with submanifold sparse convolutional networks. In CVPR, 2018.
  24. 24.Benjamin Graham, Martin Engelcke, and Laurens Van Der Maaten. 3D semantic segmentation with submanifold sparse convolutional networks. In CVPR, pages 9224–9232, 2018.
  25. 25.Benjamin Graham and Laurens van der Maaten. Submanifold sparse convolutional networks. arXiv:1706.01307, 2017.
  26. 26.Agrim Gupta, Piotr Dollár, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. In CVPR, 2019.
  27. 27.XY Han, Vardan Papyan, and David L Donoho. Neural collapse under MSE loss: Proximity to and dynamics on the central path. In ICLR, 2022.
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  29. 29.Ji Hou, Benjamin Graham, Matthias Nießner, and Saining Xie. Exploring data-efficient 3D scene understanding with contrastive scene contexts. In CVPR, pages 15587–15597, 2021.
  30. 30.Chen Huang, Yining Li, Change Loy Chen, and Xiaoou Tang. Deep imbalanced learning for face recognition and attribute prediction. IEEE TPAMI, 2019.
  31. 31.Muhammad Abdullah Jamal, Matthew Brown, Ming-Hsuan Yang, Liqiang Wang, and Boqing Gong. Rethinking class-balanced methods for long-tailed visual recognition from a domain adaptation perspective. In CVPR, 2020.
  32. 32.Wenlong Ji, Yiping Lu, Yiliang Zhang, Zhun Deng, and Weijie J Su. An unconstrained layer-peeled perspective on neural collapse. In ICLR, 2022.
  33. 33.Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, Jiashi Feng, and Yannis Kalantidis. Decoupling representation and classifier for long-tailed recognition. In ICLR, 2020.
  34. 34.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In NIPS, volume 33, pages 18661–18673, 2020.
  35. 35.Xin Lai, Jianhui Liu, Li Jiang, Liwei Wang, Hengshuang Zhao, Shu Liu, Xiaojuan Qi, and Jiaya Jia. Stratified transformer for 3D point cloud segmentation. In CVPR, pages 8500–8509, 2022.
  36. 36.Tianhong Li, Peng Cao, Yuan Yuan, Lijie Fan, Yuzhe Yang, Rogerio S Feris, Piotr Indyk, and Dina Katabi. Targeted supervised contrastive learning for long-tailed recognition. In CVPR, pages 6918–6928, 2022.
  37. 37.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. PointCNN: Convolution on x-transformed points. NeurIPS, 2018.
  38. 38.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. RefineNet: Multi-path refinement networks for high-resolution semantic segmentation. In CVPR, 2017.
  39. 39.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. In CVPR, pages 2980–2988, 2017.
  40. 40.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In ECCV, 2014.
  41. 41.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021.
  42. 42.Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X Yu. Large-scale long-tailed recognition in an open world. In CVPR, pages 2537–2546, 2019.
  43. 43.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015.
  44. 44.Jianfeng Lu and Stefan Steinerberger. Neural collapse with cross-entropy loss. arXiv preprint arXiv:2012.08465, 2020.
  45. 45.Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar. Long-tail learning via logit adjustment. In ICLR, 2020.
  46. 46.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-Net: Fully convolutional neural networks for volumetric medical image segmentation. In 3DV, pages 565–571. IEEE, 2016.
  47. 47.Dustin G Mixon, Hans Parshall, and Jianzong Pi. Neural collapse with unconstrained features. arXiv preprint arXiv:2011.11619, 2020.
  48. 48.Alexey Nekrasov, Jonas Schult, Or Litany, Bastian Leibe, and Francis Engelmann. Mix3D: Out-of-context data augmentation for 3D scenes. In 3DV, pages 116–125. IEEE, 2021.
  49. 49.Vardan Papyan, XY Han, and David L Donoho. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences, 117(40):24652–24663, 2020.
  50. 50.Federico Pernici, Matteo Bruni, Claudio Baecchi, and Alberto Del Bimbo. Regular polytope networks. IEEE Transactions on Neural Networks and Learning Systems, 2021.
  51. 51.Tomaso Poggio and Qianli Liao. Explicit regularization and implicit bias in deep network classifiers trained with the square loss. arXiv preprint arXiv:2101.00072, 2020.
  52. 52.Tobias Pohlen, Alexander Hermans, Markus Mathias, and Bastian Leibe. Full-resolution residual networks for semantic segmentation in street scenes. In CVPR, 2017.
  53. 53.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. PointNet: Deep learning on point sets for 3D classification and segmentation. In CVPR, 2017.
  54. 54.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. PointNet++: Deep hierarchical feature learning on point sets in a metric space. In NeurIPS, 2017.
  55. 55.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021.
  56. 56.Md Atiqur Rahman and Yang Wang. Optimizing intersection-over-union in deep neural networks for image segmentation. In International Symposium on Visual Computing, pages 234–244. Springer, 2016.
  57. 57.Jiawei Ren, Cunjun Yu, shunan sheng, Xiao Ma, Haiyu Zhao, Shuai Yi, and hongsheng Li. Balanced meta-softmax for long-tailed visual recognition. In NeurIPS, pages 4175–4186, 2020.
  58. 58.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015.
  59. 59.David Rozenberszki, Or Litany, and Angela Dai. Language-grounded indoor 3D semantic segmentation in the wild. In ECCV, 2022.
  60. 60.Seyed Sadegh Mohseni Salehi, Deniz Erdogmus, and Ali Gholipour. Tversky loss function for image segmentation using 3D fully convolutional deep networks. In International Workshop on Machine Learning in Medical Imaging, pages 379–387. Springer, 2017.
  61. 61.Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. Meta-weight-Net: Learning an explicit mapping for sample weighting. In NeurIPS, 2019.
  62. 62.Christos Thrampoulidis, Ganesh R Kini, Vala Vakilian, and Tina Behnia. Imbalance trouble: Revisiting neural-collapse geometry. arXiv preprint arXiv:2208.05512, 2022.
  63. 63.Zhuotao Tian, Pengguang Chen, Xin Lai, Li Jiang, Shu Liu, Hengshuang Zhao, Bei Yu, Ming-Chang Yang, and Jiaya Jia. Adaptive perspective distillation for semantic segmentation. IEEE TPAMI, 45(2):1372–1387, 2023.
  64. 64.Tom Tirer and Joan Bruna. Extended unconstrained features model for exploring deep neural collapse. In ICML, 2022.
  65. 65.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. IEEE TPAMI, 2020.
  66. 66.E Weinan and Stephan Wojtowytsch. On the emergence of tetrahedral symmetry in the final and penultimate layers of neural network classifiers. arXiv preprint arXiv:2012.05420, 2020.
  67. 67.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, 2018.
  68. 68.Liang Xie, Yibo Yang, Deng Cai, and Xiaofei He. Neural collapse inspired attraction-repulsion-balanced loss for imbalanced learning. arXiv preprint arXiv:2204.08735, 2022.
  69. 69.Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan. Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models. arXiv preprint arXiv:2208.06677, 2022.
  70. 70.Yibo Yang, Liang Xie, Shixiang Chen, Xiangtai Li, Zhouchen Lin, and Dacheng Tao. Do we really need a learnable classifier at the end of deep neural network? In NeurIPS, 2022.
  71. 71.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016.
  72. 72.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In ECCV, 2020.
  73. 73.Songyang Zhang, Zeming Li, Shipeng Yan, Xuming He, and Jian Sun. Distribution alignment: A unified framework for long-tail visual recognition. In CVPR, pages 2361–2370, 2021.
  74. 74.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun. Point transformer. In ICCV, 2021.
  75. 75.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, 2017.
  76. 76.Zhisheng Zhong, Jiequan Cui, Shu Liu, and Jiaya Jia. Improving calibration for long-tailed recognition. In CVPR, pages 16489–16498, 2021.
  77. 77.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. In CVPR, pages 633–641, 2017.
  78. 78.Jinxin Zhou, Xiao Li, Tianyu Ding, Chong You, Qing Qu, and Zhihui Zhu. On the optimization landscape of neural collapse under MSE loss: Global optimality with unconstrained features. arXiv preprint arXiv:2203.01238, 2022.
  79. 79.Jianggang Zhu, Zheng Wang, Jingjing Chen, Yi-Ping Phoebe Chen, and Yu-Gang Jiang. Balanced contrastive learning for long-tailed visual recognition. In CVPR, pages 6908–6917, 2022.
  80. 80.Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu. A geometric analysis of neural collapse with unconstrained features. In NeurIPS, 2021.

Citation

MLA
Zhong, Z., et al. “Understanding Imbalanced Semantic Segmentation Through Neural Collapse”. arXiv, 2023, http://arxiv.org/abs/2301.01100v1.
APA
Zhong, Z., Cui, J., Yang, Y., Wu, X., Qi, X., Zhang, X., & Jia, J. (2023). Understanding Imbalanced Semantic Segmentation Through Neural Collapse. arXiv. http://arxiv.org/abs/2301.01100v1
Chicago
Zhong, Z., J. Cui, Y. Yang, et al. 2023. “Understanding Imbalanced Semantic Segmentation Through Neural Collapse”. arXiv. http://arxiv.org/abs/2301.01100v1.
Harvard
Zhong, Z. et al. (2023) “Understanding Imbalanced Semantic Segmentation Through Neural Collapse”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.01100v1.
Vancouver
1. Zhong Z, Cui J, Yang Y, Wu X, Qi X, Zhang X, Jia J (2023) Understanding Imbalanced Semantic Segmentation Through Neural Collapse. arXiv

BibTeX

@article{zhong2023understanding,
  title = {Understanding Imbalanced Semantic Segmentation Through Neural Collapse},
  author = {Zhong, Zhisheng and Cui, Jiequan and Yang, Yibo and Wu, Xiaoyang and Qi, Xiaojuan and Zhang, Xiangyu and Jia, Jiaya},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.01100v1},
  eprint = {2301.01100}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE