Feature learning in deep classifiers through Intermediate Neural Collapse

Akshay RangamaniMarius LindegaardTomer GalantiTomaso A. Poggio

article2023ICML64 citations

Demonstrates that Neural Collapse extends beyond the final classification layer into intermediate representations, showing how deep networks progressively compress within-class variance and align weight subspaces across layers to form geometric simplex structures.

Listen

Deep learning models achieve strong empirical results across image recognition, language processing, and automated decision-making, yet the internal mechanics of how deep networks transform representations across intermediate hidden layers remain poorly understood. A known training phenomenon termed Neural Collapse demonstrates that the final layer of a trained classifier compresses representations of the same class into rigid, highly symmetric geometric structures. The article evaluates whether these collapse properties extend backward into intermediate hidden layers, aiming to characterize the multi-layer feature learning process from input to output.

The authors conducted an empirical investigation across four benchmark image datasets—MNIST, FashionMNIST, CIFAR-10, and SVHN—using three representative neural network architectures: multilayer perceptrons, deep convolutional networks, and residual networks. The models were trained to zero classification error using regularized gradient-based optimization. Across every layer and training stage, the study evaluated the statistical spread of internal features, the subspace alignment between internal features and layer weights, the effective rank of representations, and the accuracy of simple nearest-center decision rules applied directly to intermediate layers.

The article delivers four primary findings. First, beyond a specific threshold layer in the network, internal representations consistently exhibit feature collapse, where within-class variance drops to less than 20% of the total variance while between-class variance dominates all subsequent layers. Second, class centers in these collapsed layers converge toward an equiangular geometric layout that aligns directly with the dominant singular components of the corresponding weight matrices. Third, feature dimensionality follows an expansion-then-compression trajectory, expanding in early layers to facilitate class separation before compressing into low-rank representations in deeper layers. Fourth, fixing the weight matrices of these deeper collapsed layers to predetermined geometric templates does not degrade classification accuracy, whereas fixing early layers causes substantial performance loss.

These findings indicate that deep classifiers operate in two distinct functional stages: lower layers extract discriminative features by expanding inputs into a high-dimensional space, while upper collapsed layers act essentially as associative memories that eliminate within-class noise. This mechanism suggests significant opportunities to improve training efficiency, reduce computational costs, and streamline network compression by constraining or pruning parameters in deeper layers without sacrificing model accuracy.

Organizations developing or deploying deep classification models should test practical efficiency measures, such as freezing deeper network layers into predefined geometric matrices to accelerate training and lower parameter counts. However, because the article observed that intermediate collapse does not emerge under every hyperparameter setting or network configuration, practitioners should run pilot evaluations on their specific tasks before standardizing these architectural constraints in production workflows.

Confidence in these findings is high for standard supervised image classification settings across the tested architectures. Nevertheless, stakeholders should recognize key limitations: the evaluations rely on balanced benchmark datasets and specific optimization routines using mean squared error loss. Further research is required to confirm whether intermediate collapse holds across unbalanced data, complex multimodal models, and generative or self-supervised architectures.

Rangamani et al (2023).pdf
  • Paper: On the Role of Neural Collapse in Transfer Learning, Tomer Galanti et al. (2022). This paper establishes foundational metrics and theoretical limits for neural collapse in representation spaces, providing the formal vocabulary and baseline concepts of class-covariance geometry that the source expands to intermediate layers.
  • Paper: Understanding intermediate layers using linear classifier probes, Guillaume Alain et al. (2016). This work introduces the methodology of tracking representation quality and linear separability layer-by-layer across deep architectures, directly motivating the source's empirical study of intermediate-layer feature dynamics.
  • Paper: Similarity of Neural Network Representations Revisited, Simon Kornblith et al. (2019). This paper presents principled methods for analyzing and comparing representations across intermediate hidden layers, establishing the analytical backdrop for layer-wise covariance evolution in deep networks.
  • Paper: Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, Andrew M. Saxe et al. (2014). This study analyzes the layer-wise singular value and subspace alignment dynamics under gradient descent, laying the theoretical foundation for understanding how weight matrices align with data covariance across depths.
  • Paper: Understanding Imbalanced Semantic Segmentation Through Neural Collapse, Zhisheng Zhong et al. (2023). This paper investigates how the neural collapse phenomenon behaves under class imbalance and spatial contexts in dense prediction tasks, directly complementing the source's exploration of representation geometry in classification.
  • Paper: Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs, Amirhesam Abedsoltan et al. (2026). This work analyzes how individual neurons specialize and structure representations for multi-cluster data, offering a fine-grained, neuron-level mechanism that extends the source's findings on layer-wise class-mean subspace alignment.
  • Paper: Low-dimensional topology of deep neural networks, Junyu Ren et al. (2026). This work analyzes how intermediate layer mappings geometrically and topologically untangle representations across depth, complementing the source's covariance-reduction view of intermediate feature learning.
Cover for Feature learning in deep classifiers through Intermediate Neural Collapse

Abstract

In this paper, we conduct an empirical study of the feature learning process in deep classifiers. Recent research has identified a training phenomenon called Neural Collapse (NC), in which the top-layer feature embeddings of samples from the same class tend to concentrate around their means, and the top layer's weights align with those features. Our study aims to investigate if these properties extend to intermediate layers. We empirically study the evolution of the covariance and mean of representations across different layers and show that as we move deeper into a trained neural network, the within-class covariance decreases relative to the between-class covariance. Additionally, we find that in the top layers, where the between-class covariance is dominant, the subspace spanned by the class means aligns with the subspace spanned by the most significant singular vector components of the weight matrix in the corresponding layer. Finally, we discuss the relationship between NC and Associative Memories (Willshaw et al., 1969).

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Problem Setup
  • 4. Intermediate Neural Collapse
  • 5. Results
  • 5.1. Intermediate Neural Collapse
  • 5.2. Stable Rank of intermediate features and weights
  • 5.3. Fixing all Collapsed Layers with Simplex ETFs
  • 5.4. Solutions without Intermediate Neural Collapse
  • 6. Neural Collapse and Associative Memories
  • 7. Conclusion and Future Work
  • Acknowledgements
  • References
  • A. Experimental details
  • B. Solutions without intermediate Neural Collapse
  • C. Additional Figures establishing Intermediate Neural Collapse

Knowls

  1. Knowl 1 — Conditions of Intermediate Neural Collapse

    definition

    Let a deep classifier fW:X→RCf_W : \mathcal{X} \to \mathbb{R}^C be a composition of LL parametric layers fW(x)=TL∘⋯∘T1(x)f_W(x) = T_L \circ \dots \circ T_1(x) with parameters W={W1,…,WL}W = \{W_1, \dots, W_L\}, operating on a balanced dataset of CC classes with NN samples per class, denoted {(xi,c,yi,c)}i=1,c=1N,C\{(x_{i,c}, y_{i,c})\}_{i=1,c=1}^{N,C}. The feature representation at layer ℓ∈{1,…,L}\ell \in \{1, \dots, L\} for sample xi,cx_{i,c} is hℓ(xi,c)=Tℓ∘⋯∘T1(xi,c)∈Rpℓh^\ell(x_{i,c}) = T_\ell \circ \dots \circ T_1(x_{i,c}) \in \mathbb{R}^{p_\ell}.

    The layer-wise first- and second-order statistics are defined as:

    μcℓ:=1N∑i=1Nhℓ(xi,c),μGℓ:=1C∑c=1Cμcℓ\mu_c^\ell := \frac{1}{N} \sum_{i=1}^N h^\ell(x_{i,c}), \quad \mu_G^\ell := \frac{1}{C} \sum_{c=1}^C \mu_c^\ell

    ΣWℓ:=1NC∑c=1C∑i=1N(hℓ(xi,c)−μcℓ)(hℓ(xi,c)−μcℓ)⊤\Sigma_W^\ell := \frac{1}{NC} \sum_{c=1}^C \sum_{i=1}^N (h^\ell(x_{i,c}) - \mu_c^\ell)(h^\ell(x_{i,c}) - \mu_c^\ell)^\top

    ΣBℓ:=1C∑c=1C(μcℓ−μGℓ)(μcℓ−μGℓ)⊤\Sigma_B^\ell := \frac{1}{C} \sum_{c=1}^C (\mu_c^\ell - \mu_G^\ell)(\mu_c^\ell - \mu_G^\ell)^\top

    ΣTℓ:=ΣWℓ+ΣBℓ=1NC∑c=1C∑i=1N(hℓ(xi,c)−μGℓ)(hℓ(xi,c)−μGℓ)⊤\Sigma_T^\ell := \Sigma_W^\ell + \Sigma_B^\ell = \frac{1}{NC} \sum_{c=1}^C \sum_{i=1}^N (h^\ell(x_{i,c}) - \mu_G^\ell)(h^\ell(x_{i,c}) - \mu_G^\ell)^\top

    Intermediate Neural Collapse (NC) at layer ℓ\ell is characterized by four conditions:

    1. NC1 (Feature variability suppression): The normalized within-class variance is bounded below a threshold ϵ\epsilon (empirically ϵ≈0.2\epsilon \approx 0.2):

    Tr⁡(ΣWℓ)Tr⁡(ΣTℓ)<ϵ\frac{\operatorname{Tr}(\Sigma_W^\ell)}{\operatorname{Tr}(\Sigma_T^\ell)} < \epsilon

    1. NC2 (Simplex Equiangular Tight Frame (ETF) structure): The centered class means μcℓ−μGℓ\mu_c^\ell - \mu_G^\ell become equinorm and maximally equiangular:

    ∣∥μcℓ−μGℓ∥2−∥μc′ℓ−μGℓ∥2∣→0,∀c,c′\big| \|\mu_c^\ell - \mu_G^\ell\|_2 - \|\mu_{c'}^\ell - \mu_G^\ell\|_2 \big| \to 0, \quad \forall c, c'

    ⟨μcℓ−μGℓ∥μcℓ−μGℓ∥2,μc′ℓ−μGℓ∥μc′ℓ−μGℓ∥2⟩→−1C−1,∀c≠c′\left\langle \frac{\mu_c^\ell - \mu_G^\ell}{\|\mu_c^\ell - \mu_G^\ell\|_2}, \frac{\mu_{c'}^\ell - \mu_G^\ell}{\|\mu_{c'}^\ell - \mu_G^\ell\|_2} \right\rangle \to -\frac{1}{C-1}, \quad \forall c \neq c'

    1. NC3 (Feature-weight subspace alignment): Let Mℓ=[μ1ℓ−μGℓ,…,μCℓ−μGℓ]∈Rpℓ×CM_\ell = [\mu_1^\ell - \mu_G^\ell, \dots, \mu_C^\ell - \mu_G^\ell] \in \mathbb{R}^{p_\ell \times C} be the matrix of centered class means and Wℓ∈Rpℓ+1×pℓW_\ell \in \mathbb{R}^{p_{\ell+1} \times p_\ell} be the weight matrix. Let θ1,…,θC\theta_1, \dots, \theta_C be the Principal Angles Between Subspaces (PABS) between the range space of MℓM_\ell and the top-CC right singular subspace (input space) of WℓW_\ell. Alignment occurs when:

    1C∑k=1Ccos⁡(θk)→1\frac{1}{C} \sum_{k=1}^C \cos(\theta_k) \to 1

    and the top CC singular values of WℓW_\ell are approximately equal.

    1. NC4 (Behavioral equivalence to nearest class center classification): The network classification decision converges to the nearest class center (NCC) decision rule evaluated on intermediate representations:

    arg⁡max⁡c⟨WcL,hL(x)⟩→arg⁡min⁡c∥hℓ(x)−μcℓ∥2\arg\max_{c} \langle W_c^L, h^L(x) \rangle \to \arg\min_{c} \|h^\ell(x) - \mu_c^\ell\|_2

  2. Knowl 2 — Emergence of Intermediate Neural Collapse Across Deep Layers

    empirical result

    When overparameterized deep neural networks (Multilayer Perceptrons, Convolutional Networks, and Residual Networks) are trained with weight decay and Mean Squared Error (MSE) loss into the terminal phase of training (after training classification error reaches zero), the Neural Collapse phenomenon propagates backward from the final classification layer into a continuum of intermediate hidden layers.

    In models exhibiting intermediate collapse, there exists a specific hidden layer beyond which all subsequent downstream layers (the collapsed layers) simultaneously display the four collapse properties:

    1. Feature Variability Suppression (NC1): Beyond the threshold layer, the within-class covariance trace Tr⁡(ΣWℓ)\operatorname{Tr}(\Sigma_W^\ell) drops to a small fraction (<0.2< 0.2) of the total feature covariance Tr⁡(ΣTℓ)\operatorname{Tr}(\Sigma_T^\ell), while the between-class covariance trace Tr⁡(ΣBℓ)\operatorname{Tr}(\Sigma_B^\ell) accounts for the vast majority of total feature variance.
    2. Convergence to Simplex ETF (NC2): The normalized standard deviations of class mean norms and the variance of pairwise inner products drop toward zero in the collapsed layers, with the closest match to a canonical Simplex ETF occurring in the deepest layers.
    3. Subspace Alignment (NC3): The average cosine of the principal angles between the range space of centered class means MℓM_\ell and the top-CC input singular subspace of WℓW_\ell approaches 1.01.0. In Residual Networks, this alignment is sharpest at layers immediately preceding residual addition junctions, while intermediate convolutions within residual blocks show lower alignment.
    4. Nearest Class Center Equivalence (NC4): The training and test accuracy of the nearest class center rule applied directly to activations at collapsed layers ℓ\ell matches the accuracy of the deep classifier's final output.
  3. Knowl 3 — Hunchback Feature Dimensionality and Weight Near-Orthogonality in Deep Networks

    empirical result

    In deep networks trained into the intermediate neural collapse regime, the singular value spectra of the weight matrices and the rank of internal activations exhibit characteristic layer-wise transformations:

    1. Low-Rank and Near-Orthogonal Weights: In the collapsed layers of MLPs and ResNets, the top CC singular values of the weight matrix Wℓ∈Rpℓ+1×pℓW_\ell \in \mathbb{R}^{p_{\ell+1} \times p_\ell} are orders of magnitude larger than the remaining singular values and are tightly clustered at equal magnitudes. This demonstrates that weight transformations in collapsed layers act as low-rank, near-orthogonal projection operators.
    2. Hunchback Profile of Feature Stable Rank: Let Hcℓ=[hi,cℓ−μcℓ]i=1,c=1N,C∈Rpℓ×NCH_c^\ell = [h_{i,c}^\ell - \mu_c^\ell]_{i=1,c=1}^{N,C} \in \mathbb{R}^{p_\ell \times NC} be the matrix of centered within-class activations. The squared stable rank of HcℓH_c^\ell, defined as:

    srank⁡(Hcℓ):=∥Hcℓ∥F2∥Hcℓ∥22\operatorname{srank}(H_c^\ell) := \frac{\|H_c^\ell\|_F^2}{\|H_c^\ell\|_2^2}

    follows a non-monotonic "hunchback" trajectory across depth. The stable rank expands significantly in the early-to-intermediate layers (where the network projects inputs into a high-dimensional space to enable linear separability) and then decreases monotonically in the collapsed layers (where within-class variation is suppressed and discriminative features are compressed onto low-dimensional class subspaces).

  4. Knowl 4 — Effect of Replacing Collapsed Layers with Fixed Simplex Equiangular Tight Frames

    empirical result

    In deep architectures (such as 10-layer MLPs trained on MNIST, FashionMNIST, or CIFAR-10), layers that exhibit intermediate neural collapse can have their weight matrices fixed to static, unlearned canonical simplex Equiangular Tight Frames (ETFs) without degrading training convergence or test classification accuracy.

    In this setup, the bottom LL layers are trained while the remaining 10−L10 - L upper layers are fixed to canonical simplex ETFs. For a layer with width HH, the canonical simplex ETF is given by:

    Wℓ=HH−1(IH−1H1H1H⊤)W_\ell = \sqrt{\frac{H}{H-1}} \left( I_H - \frac{1}{H}\mathbf{1}_H \mathbf{1}_H^\top \right)

    and the final layer WL⊤∈Rd×CW_L^\top \in \mathbb{R}^{d \times C} is set to:

    WL⊤=CC−1P(IC−1C1C1C⊤)W_L^\top = \sqrt{\frac{C}{C-1}} P \left( I_C - \frac{1}{C}\mathbf{1}_C \mathbf{1}_C^\top \right)

    where P∈Rd×CP \in \mathbb{R}^{d \times C} contains the first CC columns of the d×dd \times d identity matrix.

    Empirical results demonstrate that:

    • Fixing layers in the collapsed regime (e.g., layers 7–10 for MNIST/FashionMNIST or layers 8–10 for CIFAR-10) to fixed simplex ETFs yields classification performance identical to fully trained networks.
    • Fixing non-collapsed earlier layers (e.g., layers 2–6) to simplex ETFs causes a severe drop in both training and test accuracy.

    This indicates that feature extraction and representation learning occur predominantly in the lower half of the network, while deeper collapsed layers primarily execute fixed geometric dimension reduction.

  5. Knowl 5 — Equivalence Between Neural Collapse Classifier Weights and Associative Memories

    theoretical result

    Let HL∈Rp×NCH_L \in \mathbb{R}^{p \times NC} denote the matrix of centered penultimate feature representations across CC balanced classes with NN samples per class, and let Y∈RC×NCY \in \mathbb{R}^{C \times NC} denote the corresponding matrix of one-hot target vectors. Under variability collapse (NC1), all feature vectors within class cc collapse exactly to their centered class mean, yielding the factorization:

    HL=MLYH_L = M_L Y

    where ML=[μ1L−μGL,…,μCL−μGL]∈Rp×CM_L = [\mu_1^L - \mu_G^L, \dots, \mu_C^L - \mu_G^L] \in \mathbb{R}^{p \times C} is the matrix of centered class feature means.

    A linear associative memory matrix W^L\widehat{W}_L constructed from outer products of stimulus-response pairs (hi,cL,yi,c)(h_{i,c}^L, y_{i,c}) is given by:

    W^L=YHL⊤=Y(MLY)⊤=YY⊤ML⊤\widehat{W}_L = Y H_L^\top = Y (M_L Y)^\top = Y Y^\top M_L^\top

    Since classes are balanced with NN one-hot samples each, YY⊤=NICY Y^\top = N I_C, which simplifies the associative memory matrix to:

    W^L=NICML⊤=NML⊤\widehat{W}_L = N I_C M_L^\top = N M_L^\top

    This matrix is directly proportional to ML⊤M_L^\top, exactly matching the classifier weights predicted by the feature-weight alignment property (NC3) of Neural Collapse. Furthermore, because MLM_L converges to a simplex Equiangular Tight Frame (NC2), the stimulus keys are maximally separated and near-orthogonal, which provides optimal conditions for robust associative pattern recall.

  6. Knowl 6 — Non-Universality of Intermediate Neural Collapse

    limitation

    Intermediate Neural Collapse is not a universal property of all networks that achieve terminal-phase Neural Collapse at the final classification layer. Under certain hyperparameter configurations and architectures (such as specific deep convolutional networks trained on CIFAR-10 or FashionMNIST with Mean Squared Error loss):

    1. Within-class feature covariance Tr⁡(ΣWℓ)\operatorname{Tr}(\Sigma_W^\ell) remains high relative to between-class covariance Tr⁡(ΣBℓ)\operatorname{Tr}(\Sigma_B^\ell) across all intermediate layers, only collapsing sharply at the final hidden layer.
    2. Intermediate class feature means do not organize into simplex Equiangular Tight Frames (ETFs).
    3. Nearest class center (NCC) classification accuracy on intermediate representations remains poor and does not track the accuracy of the overall deep classifier until the penultimate layer.

    Thus, final-layer Neural Collapse can occur without the emergence of intermediate Neural Collapse in upstream layers.

  7. Knowl 7 — Experimental Protocol for Measuring Intermediate Neural Collapse

    experimental setup

    The empirical evaluation of intermediate Neural Collapse uses the following configurations and measurement procedures:

    1. Datasets: MNIST, FashionMNIST, CIFAR-10, and SVHN. Images are centered and standardized using pixel-wise mean and standard deviation without data augmentation.
    2. Architectures:
      • Multilayer Perceptron (MLP): L=10L = 10 hidden layers of width H=1024H = 1024, each composed of Linear + BatchNorm + ReLU, followed by a linear classification head.
      • Deep ConvNet: Two initial 2×22 \times 2 convolutions (stride 2, BatchNorm, ReLU) followed by L=20L = 20 convolutional stacks (3×33 \times 3 kernels, H=128H = 128 channels, stride 1, padding 1, BatchNorm, ReLU), terminated with a linear head.
      • ResNets: ResNet-18 for MNIST and FashionMNIST, ResNet-34 for SVHN, and ResNet-50 for CIFAR-10.
    3. Optimization: Networks are trained to minimize Mean Squared Error (MSE) loss using SGD with momentum 0.9, weight decay 5×10−45 \times 10^{-4}, batch size 128, for 350 epochs. Initial learning rates are selected via a logarithmic sweep in [0.0001,0.25][0.0001, 0.25] with a 1-epoch linear warmup, followed by two step decays by a factor of 0.1 (or 0.2 for MLPs and ConvNets).
    4. Measurement Metrics:
      • NC1: Computed directly via Tr⁡(ΣWℓ)/Tr⁡(ΣTℓ)\operatorname{Tr}(\Sigma_W^\ell) / \operatorname{Tr}(\Sigma_T^\ell) from full dataset class means.
      • NC2: Computed via the relative standard deviation of class mean norms ∥μcℓ−μGℓ∥2\|\mu_c^\ell - \mu_G^\ell\|_2 and the mean/standard deviation of pairwise angles cos⁡(∠(μcℓ−μGℓ,μc′ℓ−μGℓ))+1C−1\cos(\angle(\mu_c^\ell - \mu_G^\ell, \mu_{c'}^\ell - \mu_G^\ell)) + \frac{1}{C-1}.
      • NC3: Computed using Singular Value Decompositions Wℓ=UWSWVW⊤W_\ell = U_W S_W V_W^\top and Mℓ=UMSMVM⊤M_\ell = U_M S_M V_M^\top. The cosines of the Principal Angles Between Subspaces (PABS) are obtained from the singular values of VW⊤UMV_W^\top U_M, and their mean 1C∑k=1Ccos⁡(θk)\frac{1}{C}\sum_{k=1}^C \cos(\theta_k) serves as the alignment measure. For convolutional layers, alignment is evaluated across input channels.
      • NC4: Evaluated as the classification accuracy of the rule arg⁡min⁡c∥hℓ(x)−μcℓ∥2\arg\min_c \|h^\ell(x) - \mu_c^\ell\|_2 across training and test splits.

Coverage note — None was omitted; all contributed definitions, empirical findings across architectures and datasets, rank and dimensionality analysis, ETF layer-fixing experiments, associative memory connections, and negative cases have been covered.

References

  1. 1.Aizerman, M. A., Braverman, E. M., and Rozonoer, L. I. Theoretical foundation of potential functions method in pattern recognition. Avtomatika i Telemekhanika, 25(6): 917–936, 1964.
  2. 2.Alain, G. and Bengio, Y. Understanding intermediate layers using linear classifier probes. ArXiv, abs/1610.01644, 2017.
  3. 3.Allen-Zhu, Z. and Li, Y. What can resnet learn efficiently, going beyond kernels? Advances in Neural Information Processing Systems, 32, 2019.
  4. 4.Allen-Zhu, Z. and Li, Y. Backward feature correction: How deep learning performs deep learning. arXiv preprint arXiv:2001.04413, 2020.
  5. 5.Anderson, J. A. A simple neural network generating an interactive memory. Mathematical biosciences, 14(3-4): 197–220, 1972.
  6. 6.Ansuini, A., Laio, A., Macke, J. H., and Zoccolan, D. Intrinsic dimension of data representations in deep neural networks. Advances in Neural Information Processing Systems, 32, 2019.
  7. 7.Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R. On exact computation with an infinitely wide neural net. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/dbc4d84bfcfe2284ba11beffb853a8c4-Paper.pdf.
  8. 8.Bau, D., Liu, S., Wang, T., Zhu, J.-Y., and Torralba, A. Rewriting a deep generative model. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.
  9. 9.Ben-Shaul, I. and Dekel, S. Nearest class-center simplification through intermediate layers. arXiv preprint arXiv:2201.08924, 2022.
  10. 10.Björck, A. and Golub, G. H. Numerical methods for computing angles between linear subspaces. Mathematics of computation, 27(123):579–594, 1973.
  11. 11.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020.
  12. 12.Chen, M., Bai, Y., Lee, J. D., Zhao, T., Wang, H., Xiong, C., and Socher, R. Towards understanding hierarchical learning: Benefits of neural representations. Advances in Neural Information Processing Systems, 33:22134–22145, 2020.
  13. 13.Cohen, G., Sapiro, G., and Giryes, R. Dnn or k-nn: That is the generalize vs. memorize question. arXiv preprint arXiv:1805.06822, 2018.
  14. 14.Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. Knowledge neurons in pretrained transformers. arXiv preprint arXiv:2104.08696, 2021.
  15. 15.Du, S., Lee, J., Li, H., Wang, L., and Zhai, X. Gradient descent finds global minima of deep neural networks. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 1675–1685. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/du19c.html.
  16. 16.Du, S. S., Zhai, X., Póczos, B., and Singh, A. Gradient descent provably optimizes over-parameterized neural networks. ArXiv, abs/1810.02054, 2018.
  17. 17.Ergen, T. and Pilanci, M. Revealing the structure of deep neural networks via convex duality. arXiv preprint arXiv:2002.09773, 2020.
  18. 18.Fang, C., He, H., Long, Q., and Su, W. J. Layer-peeled model: Toward understanding well-trained deep neural networks. CoRR, abs/2101.12699, 2021. URL https://arxiv.org/abs/2101.12699.
  19. 19.Galanti, T., Galanti, L., and Ben-Shaul, I. On the implicit bias towards minimal depth of deep neural networks. arXiv preprint arXiv:2202.09028, 2022a.
  20. 20.Galanti, T., György, A., and Hutter, M. Improved generalization bounds for transfer learning via neural collapse. In First Workshop on Pre-training: Perspectives, Pitfalls, and Paths Forward at ICML 2022, 2022b. URL https://openreview.net/forum?id=VrK7pKwOhT_.
  21. 21.Galanti, T., György, A., and Hutter, M. On the role of neural collapse in transfer learning. In International Conference on Learning Representations, 2022c. URL https://openreview.net/forum?id=SwIp410B6aQ.
  22. 22.Galanti, T., György, A., and Hutter, M. Generalization bounds for transfer learning with pretrained classifiers, 2022d. URL https://arxiv.org/abs/2212.12532.
  23. 23.Geva, M., Schuster, R., Berant, J., and Levy, O. Transformer feed-forward layers are key-value memories. arXiv preprint arXiv:2012.14913, 2020.
  24. 24.Goldfeld, Z., Van Den Berg, E., Greenewald, K., Melnyk, I., Nguyen, N., Kingsbury, B., and Polyanskiy, Y. Estimating information flow in deep neural networks. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 2299–2308. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/goldfeld19a.html.
  25. 25.Han, X., Papyan, V., and Donoho, D. L. Neural collapse under MSE loss: Proximity to and dynamics on the central path. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=w1UbdvWH_R3.
  26. 26.He, H. and Su, W. J. A law of data separation in deep learning. arXiv preprint arXiv:2210.17020, 2022.
  27. 27.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016. doi: 10.1109/CVPR.2016.90.
  28. 28.Hopfield, J. J. Neural networks and physical systems with emergent collective computational abilities. Proceedings of the national academy of sciences, 79(8):2554–2558, 1982.
  29. 29.Irie, K., Csordás, R., and Schmidhuber, J. The dual form of neural networks revisited: Connecting test time predictions to training patterns via spotlights of attention. arXiv preprint arXiv:2202.05798, 2022.
  30. 30.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS, pp. 8580–8589, Red Hook, NY, USA, 2018. Curran Associates Inc.
  31. 31.Ji, W., Lu, Y., Zhang, Y., Deng, Z., and Su, W. J. An unconstrained layer-peeled perspective on neural collapse. arXiv preprint arXiv:2110.02796, 2021.
  32. 32.Jordan, C. Essai sur la géométrie à nn dimensions. Bulletin de la Société Mathématique de France, 3:103–174.
  33. 33.Kanerva, P. Sparse distributed memory and related models. Technical report, 1992.
  34. 34.Kohonen, T. Self-organization and associative memory, 1989.
  35. 35.Krizhevsky, A. and Hinton, G. Learning multiple layers of features from tiny images. Master’s thesis, Department of Computer Science, University of Toronto, 2009.
  36. 36.LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  37. 37.Littwin, E., Galanti, T., Wolf, L., and Yang, G. On infinite-width hypernetworks. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 13226–13237. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/999df4ce78b966de17aee1dc87111044-Paper.pdf.
  38. 38.Lu, J. and Steinerberger, S. Neural collapse with cross-entropy loss. CoRR, abs/2012.08465, 2020. URL https://arxiv.org/abs/2012.08465.
  39. 39.Malach, E., Kamath, P., Abbe, E., and Srebro, N. Quantifying the benefit of using differentiable learning over tangent kernels. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 7379–7389. PMLR, 18–24 Jul 2021.
  40. 40.Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Locating and editing factual knowledge in gpt. arXiv preprint arXiv:2202.05262, 2022.
  41. 41.Mixon, D. G., Parshall, H., and Pi, J. Neural collapse with unconstrained features. CoRR, abs/2011.11619, 2020. URL https://arxiv.org/abs/2011.11619.
  42. 42.Netzer, Y., Wang, T., Coates, A., Bissacco, B., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. Advances in neural information processing systems, pp. 557–565, 2011.
  43. 43.Papyan, V. Traces of class/cross-class structure pervade deep learning spectra. Journal of Machine Learning Research, 21(252):1–64, 2020. URL http://jmlr.org/papers/v21/20-933.html.
  44. 44.Papyan, V., Han, X., and Donoho, D. L. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences, 117(40):24652–24663, 2020.
  45. 45.Rangamani, A. and Banburski-Fahey, A. Neural collapse in deep homogeneous classifiers and the role of weight decay. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4243–4247. IEEE, 2022.
  46. 46.Recanatesi, S., Farrell, M., Advani, M., Moore, T., Lajoie, G., and Shea-Brown, E. Dimensionality compression and expansion in deep neural networks. arXiv preprint arXiv:1906.00443, 2019.
  47. 47.Santurkar, S., Tsipras, D., Elango, M., Bau, D., Torralba, A., and Madry, A. Editing a classifier by rewriting its prediction rules. Advances in Neural Information Processing Systems, 34:23359–23373, 2021.
  48. 48.Saxe, A. M., Bansal, Y., Dapello, J., Advani, M., Kolchinsky, A., Tracey, B. D., and Cox, D. D. On the information bottleneck theory of deep learning. Journal of Statistical Mechanics: Theory and Experiment, 2019(12):124020, 2019.
  49. 49.Shwartz-Ziv, R. and Tishby, N. Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810, 2017.
  50. 50.Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. Mastering the game of Go with deep neural networks and tree search. Nature, 529:484–489, 2016. ISSN 0028-0836. doi: 10.1038/nature16961.
  51. 51.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014.
  52. 52.Tirer, T. and Bruna, J. Extended unconstrained features model for exploring deep neural collapse. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 21478–21505. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/tirer22a.html.
  53. 53.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  54. 54.Willshaw, D. J., Buneman, O. P., and Longuet-Higgins, H. C. Non-holographic associative memory. Nature, 222(5197): 960–962, 1969.
  55. 55.Wojtowytsch, S. et al. On the emergence of tetrahedral symmetry in the final and penultimate layers of neural network classifiers. arXiv preprint arXiv:2012.05420, 2020.
  56. 56.Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N. Kernel and rich regimes in overparametrized models. In Conference on Learning Theory, pp. 3635–3673. PMLR, 2020.
  57. 57.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  58. 58.Xu, M., Rangamani, A., Liao, Q., Galanti, T., and Poggio, T. Dynamics in deep classifiers trained with the square loss: normalization, low rank, neural collapse and generalization bounds. Research, 2023.
  59. 59.Yang, G. Tensor programs ii: Neural tangent kernel for any architecture, 2020. URL https://arxiv.org/abs/2006.14548.
  60. 60.Yang, G. and Littwin, E. Tensor programs iib: Architectural universality of neural tangent kernel training dynamics. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 11762–11772. PMLR, 18–24 Jul 2021.
  61. 61.Zhou, J., Li, X., Ding, T., You, C., Qu, Q., and Zhu, Z. On the optimization landscape of neural collapse under mse loss: Global optimality with unconstrained features. arXiv preprint arXiv:2203.01238, 2022.
  62. 62.Zhu, Z., Ding, T., Zhou, J., Li, X., You, C., Sulam, J., and Qu, Q. A geometric analysis of neural collapse with unconstrained features. Advances in Neural Information Processing Systems, 34:29820–29834, 2021.

Citation

MLA
Rangamani, A., et al. “Feature Learning in Deep Classifiers Through Intermediate Neural Collapse”. International Conference on Machine Learning, vol. 202, 2023, pp. 28729–45, https://proceedings.mlr.press/v202/rangamani23a.html.
APA
Rangamani, A., Lindegaard, M., Galanti, T., & Poggio, T. A. (2023). Feature learning in deep classifiers through Intermediate Neural Collapse. International Conference on Machine Learning, 202, 28729–28745. https://proceedings.mlr.press/v202/rangamani23a.html
Chicago
Rangamani, A., M. Lindegaard, T. Galanti, and T. A. Poggio. 2023. “Feature Learning in Deep Classifiers Through Intermediate Neural Collapse”. International Conference on Machine Learning 202: 28729–45. https://proceedings.mlr.press/v202/rangamani23a.html.
Harvard
Rangamani, A. et al. (2023) “Feature learning in deep classifiers through Intermediate Neural Collapse”, International Conference on Machine Learning. PMLR, pp. 28729–28745. Available at: https://proceedings.mlr.press/v202/rangamani23a.html.
Vancouver
1. Rangamani A, Lindegaard M, Galanti T, Poggio TA (2023) Feature learning in deep classifiers through Intermediate Neural Collapse. In: International Conference on Machine Learning. PMLR, pp 28729–28745

BibTeX

@InProceedings{pmlr-v202-rangamani23a,
  title = 	 {Feature learning in deep classifiers through Intermediate Neural Collapse},
  author =       {Rangamani, Akshay and Lindegaard, Marius and Galanti, Tomer and Poggio, Tomaso A},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {28729--28745},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/rangamani23a/rangamani23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/rangamani23a.html},
  abstract = 	 {In this paper, we conduct an empirical study of the feature learning process in deep classifiers. Recent research has identified a training phenomenon called Neural Collapse (NC), in which the top-layer feature embeddings of samples from the same class tend to concentrate around their means, and the top layer’s weights align with those features. Our study aims to investigate if these properties extend to intermediate layers. We empirically study the evolution of the covariance and mean of representations across different layers and show that as we move deeper into a trained neural network, the within-class covariance decreases relative to the between-class covariance. Additionally, we find that in the top layers, where the between-class covariance is dominant, the subspace spanned by the class means aligns with the subspace spanned by the most significant singular vector components of the weight matrix in the corresponding layer. Finally, we discuss the relationship between NC and Associative Memories (Willshaw et. al. 1969).}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/