Feature learning in deep classifiers through Intermediate Neural Collapse
Akshay RangamaniMarius LindegaardTomer GalantiTomaso A. Poggio
Demonstrates that Neural Collapse extends beyond the final classification layer into intermediate representations, showing how deep networks progressively compress within-class variance and align weight subspaces across layers to form geometric simplex structures.
Deep learning models achieve strong empirical results across image recognition, language processing, and automated decision-making, yet the internal mechanics of how deep networks transform representations across intermediate hidden layers remain poorly understood. A known training phenomenon termed Neural Collapse demonstrates that the final layer of a trained classifier compresses representations of the same class into rigid, highly symmetric geometric structures. The article evaluates whether these collapse properties extend backward into intermediate hidden layers, aiming to characterize the multi-layer feature learning process from input to output.
The authors conducted an empirical investigation across four benchmark image datasets—MNIST, FashionMNIST, CIFAR-10, and SVHN—using three representative neural network architectures: multilayer perceptrons, deep convolutional networks, and residual networks. The models were trained to zero classification error using regularized gradient-based optimization. Across every layer and training stage, the study evaluated the statistical spread of internal features, the subspace alignment between internal features and layer weights, the effective rank of representations, and the accuracy of simple nearest-center decision rules applied directly to intermediate layers.
The article delivers four primary findings. First, beyond a specific threshold layer in the network, internal representations consistently exhibit feature collapse, where within-class variance drops to less than 20% of the total variance while between-class variance dominates all subsequent layers. Second, class centers in these collapsed layers converge toward an equiangular geometric layout that aligns directly with the dominant singular components of the corresponding weight matrices. Third, feature dimensionality follows an expansion-then-compression trajectory, expanding in early layers to facilitate class separation before compressing into low-rank representations in deeper layers. Fourth, fixing the weight matrices of these deeper collapsed layers to predetermined geometric templates does not degrade classification accuracy, whereas fixing early layers causes substantial performance loss.
These findings indicate that deep classifiers operate in two distinct functional stages: lower layers extract discriminative features by expanding inputs into a high-dimensional space, while upper collapsed layers act essentially as associative memories that eliminate within-class noise. This mechanism suggests significant opportunities to improve training efficiency, reduce computational costs, and streamline network compression by constraining or pruning parameters in deeper layers without sacrificing model accuracy.
Organizations developing or deploying deep classification models should test practical efficiency measures, such as freezing deeper network layers into predefined geometric matrices to accelerate training and lower parameter counts. However, because the article observed that intermediate collapse does not emerge under every hyperparameter setting or network configuration, practitioners should run pilot evaluations on their specific tasks before standardizing these architectural constraints in production workflows.
Confidence in these findings is high for standard supervised image classification settings across the tested architectures. Nevertheless, stakeholders should recognize key limitations: the evaluations rely on balanced benchmark datasets and specific optimization routines using mean squared error loss. Further research is required to confirm whether intermediate collapse holds across unbalanced data, complex multimodal models, and generative or self-supervised architectures.
- Paper: On the Role of Neural Collapse in Transfer Learning, Tomer Galanti et al. (2022). This paper establishes foundational metrics and theoretical limits for neural collapse in representation spaces, providing the formal vocabulary and baseline concepts of class-covariance geometry that the source expands to intermediate layers.
- Paper: Understanding intermediate layers using linear classifier probes, Guillaume Alain et al. (2016). This work introduces the methodology of tracking representation quality and linear separability layer-by-layer across deep architectures, directly motivating the source's empirical study of intermediate-layer feature dynamics.
- Paper: Similarity of Neural Network Representations Revisited, Simon Kornblith et al. (2019). This paper presents principled methods for analyzing and comparing representations across intermediate hidden layers, establishing the analytical backdrop for layer-wise covariance evolution in deep networks.
- Paper: Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, Andrew M. Saxe et al. (2014). This study analyzes the layer-wise singular value and subspace alignment dynamics under gradient descent, laying the theoretical foundation for understanding how weight matrices align with data covariance across depths.
- Paper: Understanding Imbalanced Semantic Segmentation Through Neural Collapse, Zhisheng Zhong et al. (2023). This paper investigates how the neural collapse phenomenon behaves under class imbalance and spatial contexts in dense prediction tasks, directly complementing the source's exploration of representation geometry in classification.
- Paper: Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs, Amirhesam Abedsoltan et al. (2026). This work analyzes how individual neurons specialize and structure representations for multi-cluster data, offering a fine-grained, neuron-level mechanism that extends the source's findings on layer-wise class-mean subspace alignment.
- Paper: Low-dimensional topology of deep neural networks, Junyu Ren et al. (2026). This work analyzes how intermediate layer mappings geometrically and topologically untangle representations across depth, complementing the source's covariance-reduction view of intermediate feature learning.
