On the Role of Neural Collapse in Transfer Learning

Tomer GalantiAndrás GyörgyMarcus Hutter

article2022ICLR127 citations

Demonstrates that the geometric phenomenon of neural collapse extends to unseen classes, providing a theoretical foundation for why standard classifiers transfer effectively to few-shot learning tasks.

Listen

Transfer learning using large, pretrained neural network representations—often termed foundation models—has emerged as a highly effective approach across modern machine learning. In many real-world scenarios, organizations must deploy models into new operational settings where labeled target data is scarce and expensive to obtain. While specialized few-shot algorithms were long assumed necessary to adapt models from minimal data, recent practical results demonstrate that simple classifiers built on general-purpose foundation models match or exceed specialized methods. Until now, theoretical justification explaining why standard classification pretraining transfers so effectively to unseen tasks has been lacking.

The article establishes both theoretical foundations and empirical evidence explaining why foundation models transfer effectively to new classes in low-data regimes. The authors evaluate how a geometric training phenomenon known as neural collapse—where intermediate representations within the same class tightly cluster around their mean while maximizing the distance between different classes—generalizes beyond the source training data to completely new, unseen categories.

The research combines mathematical generalization bounds with systematic empirical evaluations. The authors formalized a clustering metric termed class-distance normalized variance, which quantifies within-class feature spread relative to between-class separation. They derived statistical bounds establishing how this variance metric behaves when encountering new samples from known classes as well as novel classes drawn from the same underlying distribution. To validate the theory, the authors trained standard convolutional and residual network architectures across four benchmark datasets—Mini-ImageNet, CIFAR-FS, FC-100, and EMNIST—and evaluated downstream performance using a simple linear ridge regression classifier without fine-tuning the underlying feature extractor.

The primary findings demonstrate a clear mechanism behind transfer learning success. First, neural collapse successfully generalizes to unseen samples of source classes as sample size increases, and more crucially, generalizes to completely new target classes when the number of source training classes grows. Second, mathematical bounds prove that stronger neural collapse directly guarantees a lower upper bound on classification error for simple downstream classifiers, requiring very few samples per class to achieve high accuracy. Third, empirical tests show that as the number of source classes increases, target clustering variance steadily decreases and few-shot accuracy improves, with standard architectures and simple linear heads remaining highly competitive with complex meta-learning methods.

These results have substantial practical implications for engineering, computational cost, and risk management. Machine learning teams do not need to invest engineering resources into fragile or specialized few-shot meta-learning pipelines; standard supervised pretraining on diverse datasets inherently produces transferable, well-clustered representations. However, the findings reveal a key operational trade-off: over-optimizing or training with excessively small learning rates can lead to overfitting on source classes, which increases target variance and degrades adaptation performance. Deploying standard learning-rate schedules combined with validation-based checkpoint selection effectively mitigates this transfer degradation.

Decision-makers should focus pretraining strategies on maximizing class diversity and dataset breadth rather than pursuing complex meta-learning architectures. When adapting pretrained models to downstream tasks with limited data, teams should use simple linear or nearest-mean classifiers as strong, low-cost default solutions. Furthermore, practitioners should track feature-clustering metrics on validation data during pretraining to identify optimal model checkpoints before source overfitting occurs.

The primary limitations of this work center on the assumption that source and target classes originate from the same broader distribution of categories. In applications with severe domain shift where target data diverges fundamentally from the pretraining domain, or where feature means collapse together, the theoretical guarantees diminish. Nevertheless, for standard domain transfer, the findings provide strong confidence that standard classification training at scale serves as a reliable, cost-effective engine for few-shot adaptation.

arXiv: 2112.15121
  • Paper: A Closer Look at Few-shot Classification, Wei-Yu Chen et al. (2019). This benchmark paper establishes the empirical baseline where simple representations from standard classifiers rival complex meta-learning methods in few-shot settings, providing the direct empirical puzzle that the source paper theoretically explains via neural collapse.
  • Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). This work introduces prototype-based metric classification for few-shot learning, establishing the class-mean representation structure that neural collapse provides a geometric foundation for.
  • Paper: Do Better ImageNet Models Transfer Better?, Simon Kornblith et al. (2018). This study systematically demonstrates that standard classification pretraining yields highly transferable representations for downstream tasks, serving as a core foundation for analyzing transferability in overparameterized networks.
  • Paper: Matching Networks for One Shot Learning, Oriol Vinyals et al. (2016). This foundational paper frames the metric-based few-shot learning paradigm against which standard multi-class classifier representations are compared and analyzed.
  • Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). This paper offers the primary empirical exploration of how layer representations in deep networks transition from general to task-specific features during transfer.

No sufficiently relevant recommendations were found.

Cover for On the Role of Neural Collapse in Transfer Learning

Abstract

We study the ability of foundation models to learn representations for classification that are transferable to new, unseen classes. Recent results in the literature show that representations learned by a single classifier over many classes are competitive on few-shot learning problems with representations learned by special-purpose algorithms designed for such problems. In this paper we provide an explanation for this behavior based on the recently observed phenomenon that the features learned by overparameterized classification networks show an interesting clustering property, called neural collapse. We demonstrate both theoretically and empirically that neural collapse generalizes to new samples from the training classes, and -- more importantly -- to new classes as well, allowing foundation models to provide feature maps that work well in transfer learning and, specifically, in the few-shot setting.

Table of Contents

  • 1 Introduction
  • 1.1 Other Related Work
  • 2 Problem Setup
  • 3 Neural Collapse
  • 4 Neural Collapse on Unseen Data
  • 4.1 Generalization to New Samples from the Same Class
  • 4.2 Neural Collapse Generalizes to New Classes
  • 4.3 CDNV and Classification Error
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Results
  • 6 Conclusions
  • References
  • A Experimental Details
  • B Additional Experiments
  • B.1 Varying Hyperparameters
  • B.2 Neural Collapse and Lower Layers
  • B.3 Dynamics of the Class-Embedding Distances
  • B.4 Class-Covariance Normalized Variance
  • C Proof of Proposition
  • D Proof of Propositions
  • E Analysis for Sections -
  • F Analysis for Section
  • G Analysis for Section
  • H List of Notation

Knowls

  1. Knowl 1 — Class-Distance Normalized Variance and Within-Class Variation Collapse

    definition

    Let X⊂Rd\mathcal{X} \subset \mathbb{R}^d denote the input instance space, and let f:Rd→Rpf: \mathbb{R}^d \to \mathbb{R}^p be a feature map (such as the penultimate layer representation of a deep neural network). For any distribution QQ over X\mathcal{X}, let μf(Q)=Ex∼Q[f(x)]∈Rp\mu_f(Q) = \mathbb{E}_{x \sim Q}[f(x)] \in \mathbb{R}^p denote the feature mean and Varf(Q)=Ex∼Q[∥f(x)−μf(Q)∥2]\text{Var}_f(Q) = \mathbb{E}_{x \sim Q}[\|f(x) - \mu_f(Q)\|^2] denote the total feature variance.

    For two class-conditional distributions Q1,Q2Q_1, Q_2 over X\mathcal{X}, their Class-Distance Normalized Variance (CDNV) is defined as: Vf(Q1,Q2)=Varf(Q1)+Varf(Q2)2∥μf(Q1)−μf(Q2)∥2V_f(Q_1, Q_2) = \frac{\text{Var}_f(Q_1) + \text{Var}_f(Q_2)}{2\|\mu_f(Q_1) - \mu_f(Q_2)\|^2}

    For finite sample sets S1,S2⊂XS_1, S_2 \subset \mathcal{X}, the empirical CDNV is defined by Vf(S1,S2)=Vf(U[S1],U[S2])V_f(S_1, S_2) = V_f(U[S_1], U[S_2]), where U[S]U[S] is the uniform distribution over SS.

    In an ll-class classification problem with training sets S~1,…,S~l\tilde{S}_1, \dots, \tilde{S}_l, within-class variation collapse (neural collapse during training) asserts that the average pairwise CDNV between classes converges to zero asymptotically with training time tt: lim⁡t→∞Avgi≠j∈[l][Vft(S~i,S~j)]=0\lim_{t \to \infty} \text{Avg}_{i \neq j \in [l]} \left[ V_{f_t}(\tilde{S}_i, \tilde{S}_j) \right] = 0

  2. Knowl 2 — Generalization of Neural Collapse to New Classes

    theoretical result

    Let DC\mathcal{D}_C be a distribution over class-conditional distributions in a collection of classes C\mathcal{C}. Assume ll source class distributions P~={P~i}i=1l\tilde{\mathcal{P}} = \{\tilde{P}_i\}_{i=1}^l are drawn conditionally independent according to DC(P~1,…,P~l∣P~i≠P~j for all i≠j∈[l])\mathcal{D}_C(\tilde{P}_1, \dots, \tilde{P}_l \mid \tilde{P}_i \neq \tilde{P}_j \text{ for all } i \neq j \in [l]).

    For a feature map f∈Ff \in \mathcal{F}, define Hf(Pc)=(μf(Pc),Varf(Pc))∈Rp+1H_f(P_c) = (\mu_f(P_c), \text{Var}_f(P_c)) \in \mathbb{R}^{p+1}, and for a finite set of candidate functions F∗⊂F\mathcal{F}^* \subset \mathcal{F}, let HF∗(P~)={(Hf(P~c))c=1l:f∈F∗}H_{\mathcal{F}^*}(\tilde{\mathcal{P}}) = \{ (H_f(\tilde{P}_c))_{c=1}^l : f \in \mathcal{F}^* \}. Define the minimal class separation margin: Δ(F∗)=inf⁡f∈F∗inf⁡Pc≠Pc′∥μf(Pc)−μf(Pc′)∥>0\Delta(\mathcal{F}^*) = \inf_{f \in \mathcal{F}^*} \inf_{P_c \neq P_{c'}} \|\mu_f(P_c) - \mu_f(P_{c'})\| > 0

    With probability at least 1−δ1 - \delta over the random draw of source distributions P~\tilde{\mathcal{P}}, every f∈F∗f \in \mathcal{F}^* satisfies: EPc≠Pc′[Vf(Pc,Pc′)]≤Avgi≠j∈[l][Vf(P~i,P~j)]+(8+16sup⁡f′∈F∗,P′∈CVarf′(P′)Δ(F∗))2πlog⁡(l)E[R(HF∗(P~))](l−1)Δ(F∗)2+(1+4sup⁡x∈X,f′∈F∗∥f′(x)∥Δ(F∗))2log⁡(1/δ)sup⁡f′∈F∗,P′∈CVarf′(P′)lΔ(F∗)2\mathbb{E}_{P_c \neq P_{c'}} [V_f(P_c, P_{c'})] \le \text{Avg}_{i \neq j \in [l]} \left[ V_f(\tilde{P}_i, \tilde{P}_j) \right] + \left( 8 + \frac{16 \sup_{f' \in \mathcal{F}^*, P' \in \mathcal{C}} \text{Var}_{f'}(P')}{\Delta(\mathcal{F}^*)} \right) \frac{\sqrt{2\pi \log(l)} \mathbb{E}[\mathcal{R}(H_{\mathcal{F}^*}(\tilde{\mathcal{P}}))]}{(l - 1) \Delta(\mathcal{F}^*)^2} + \left( 1 + \frac{4 \sup_{x \in \mathcal{X}, f' \in \mathcal{F}^*} \|f'(x)\|}{\Delta(\mathcal{F}^*)} \right) \frac{2 \sqrt{\log(1/\delta)} \sup_{f' \in \mathcal{F}^*, P' \in \mathcal{C}} \text{Var}_{f'}(P')}{\sqrt{l} \Delta(\mathcal{F}^*)^2} where R(⋅)\mathcal{R}(\cdot) denotes the Rademacher complexity and Vf(Q1,Q2)=Varf(Q1)+Varf(Q2)2∥μf(Q1)−μf(Q2)∥2V_f(Q_1, Q_2) = \frac{\text{Var}_f(Q_1) + \text{Var}_f(Q_2)}{2\|\mu_f(Q_1) - \mu_f(Q_2)\|^2} is the class-distance normalized variance.

  3. Knowl 3 — Generalization of CDNV to Unseen Samples from the Same Class

    theoretical result

    Fix two source classes ii and jj with true class-conditional distributions P~i\tilde{P}_i and P~j\tilde{P}_j, and let S~c∼P~cmc\tilde{S}_c \sim \tilde{P}_c^{m_c} be sets of mcm_c i.i.d. samples for c∈{i,j}c \in \{i, j\}. Let f:X→Rpf: \mathcal{X} \to \mathbb{R}^p be a learned feature extractor. For any δ∈(0,1)\delta \in (0, 1) and c∈{i,j}c \in \{i, j\}, let ϵ1c(δ)\epsilon_1^c(\delta) and ϵ2c(δ)\epsilon_2^c(\delta) be upper bounds on the generalization gap satisfying with probability at least 1−δ1 - \delta: ∥Ex∼P~c[f(x)]−Avgx∈S~c[f(x)]∥≤ϵ1c(δ)\|\mathbb{E}_{x \sim \tilde{P}_c}[f(x)] - \text{Avg}_{x \in \tilde{S}_c}[f(x)]\| \le \epsilon_1^c(\delta) ∣Ex∼P~c[∥f(x)∥2]−Avgx∈S~c[∥f(x)∥2]∣≤ϵ2c(δ)\left| \mathbb{E}_{x \sim \tilde{P}_c}[\|f(x)\|^2] - \text{Avg}_{x \in \tilde{S}_c}[\|f(x)\|^2] \right| \le \epsilon_2^c(\delta)

    Define the relative perturbation terms: A=ϵ1i(δ/4)+ϵ1j(δ/4)∥μf(P~i)−μf(P~j)∥A = \frac{\epsilon_1^i(\delta/4) + \epsilon_1^j(\delta/4)}{\|\mu_f(\tilde{P}_i) - \mu_f(\tilde{P}_j)\|} B=Avgc∈{i,j}[ϵ2c(δ/4)+2∥μf(P~c)∥⋅ϵ1c(δ/4)+(ϵ1c(δ/4))2]∥μf(S~i)−μf(S~j)∥2B = \frac{\text{Avg}_{c \in \{i, j\}} \left[ \epsilon_2^c(\delta/4) + 2\|\mu_f(\tilde{P}_c)\| \cdot \epsilon_1^c(\delta/4) + (\epsilon_1^c(\delta/4))^2 \right]}{\|\mu_f(\tilde{S}_i) - \mu_f(\tilde{S}_j)\|^2}

    Then, with probability at least 1−δ1 - \delta over the training samples S~i,S~j\tilde{S}_i, \tilde{S}_j: Vf(P~i,P~j)≤(Vf(S~i,S~j)+B)(1+A)2V_f(\tilde{P}_i, \tilde{P}_j) \le \left( V_f(\tilde{S}_i, \tilde{S}_j) + B \right) (1 + A)^2 where VfV_f denotes the class-distance normalized variance.

  4. Knowl 4 — Upper Bounds on Nearest Class-Mean Error via CDNV

    theoretical result

    Consider a balanced kk-class classification problem with class-conditional distributions P={Pc}c=1k\mathcal{P} = \{P_c\}_{c=1}^k and uniform class priors. A dataset S=⋃c=1kScS = \bigcup_{c=1}^k S_c contains ncn_c i.i.d. samples Sc∼PcncS_c \sim P_c^{n_c} per class. The nearest empirical class-mean classifier on a feature representation f:Rd→Rpf: \mathbb{R}^d \to \mathbb{R}^p is hf,S(x)=arg⁡min⁡c∈[k]∥f(x)−μf(Sc)∥h_{f,S}(x) = \arg\min_{c \in [k]} \|f(x) - \mu_f(S_c)\|. The expected classification error E[Err]=ESE(x,y)∼P[I(hf,S(x)≠y)]\mathbb{E}[\text{Err}] = \mathbb{E}_S \mathbb{E}_{(x,y) \sim P} [\mathbb{I}(h_{f,S}(x) \neq y)] satisfies:

    1. General and Spherically Symmetric Cases: E[Err]≤16(k−1)(1s(f,P)+1nc)Avgi≠j[Vf(Pi,Pj)]\mathbb{E}[\text{Err}] \le 16(k - 1) \left( \frac{1}{s(f, \mathcal{P})} + \frac{1}{n_c} \right) \text{Avg}_{i \neq j} \left[ V_f(P_i, P_j) \right] where s(f,P)=ps(f, \mathcal{P}) = p if the distributions {f∘Pc}c=1k\{f \circ P_c\}_{c=1}^k are spherically symmetric, and s(f,P)=1s(f, \mathcal{P}) = 1 otherwise.

    2. Spherical Gaussian Case: If {f∘Pc}c=1k\{f \circ P_c\}_{c=1}^k are spherical Gaussians with Vfmax⁡:=max⁡i≠jVarf(Pi)∥μf(Pi)−μf(Pj)∥2≤116V_f^{\max} := \max_{i \neq j} \frac{\text{Var}_f(P_i)}{\|\mu_f(P_i) - \mu_f(P_j)\|^2} \le \frac{1}{16}, then: E[Err]≤3(k−1)exp⁡(−p/(32Vfmax⁡))(5Vfmax⁡)p/2\mathbb{E}[\text{Err}] \le 3(k - 1) \frac{\exp\left(-p / (32 V_f^{\max})\right)}{(5 V_f^{\max})^{p/2}} More generally, for any γ∈(0,1)\gamma \in (0, 1) such that Vfmax⁡≤γ2nc/4V_f^{\max} \le \gamma^2 n_c / 4: E[Err]≤(k−1)(2Vfmax⁡(1−γ)2πpexp⁡(−(1−γ)2p4Vfmax⁡)+(ncγ2e4Vfmax⁡)pexp⁡(−γ2ncp4Vfmax⁡))\mathbb{E}[\text{Err}] \le (k - 1) \left( \sqrt{\frac{2 V_f^{\max}}{(1 - \gamma)^2 \pi p}} \exp\left( -\frac{(1 - \gamma)^2 p}{4 V_f^{\max}} \right) + \left( \frac{n_c \gamma^2 e}{4 V_f^{\max}} \right)^p \exp\left( -\frac{\gamma^2 n_c p}{4 V_f^{\max}} \right) \right)

  5. Knowl 5 — Moment Concentration and Class-Space Rademacher Complexity for ReLU Networks

    theoretical result

    Let F\mathcal{F} be the class of qq-layer ReLU neural feature maps f(x)=Wqσ(Wq−1…σ(W1x)):Rd→Rpf(x) = W_q \sigma(W_{q-1} \dots \sigma(W_1 x)): \mathbb{R}^d \to \mathbb{R}^p where σ\sigma is the ReLU activation and Wi∈Rdi+1×diW_i \in \mathbb{R}^{d_{i+1} \times d_i} with d1=dd_1 = d and dq+1=pd_{q+1} = p. The spectral complexity is defined as C(f)=max⁡j∈[p]∥Wq,j∥∏r=1q−1∥Wr∥\mathcal{C}(f) = \max_{j \in [p]} \|W_{q,j}\| \prod_{r=1}^{q-1} \|W_r\|, where ∥Wr∥\|W_r\| is the matrix spectral norm and ∥Wq,j∥\|W_{q,j}\| is the Euclidean norm of row jj.

    Let X⊂Rd\mathcal{X} \subset \mathbb{R}^d be bounded. Given ll class-conditional distributions P~={P~c}c=1l\tilde{\mathcal{P}} = \{\tilde{P}_c\}_{c=1}^l and sample sets Sc∼P~cmcS_c \sim \tilde{P}_c^{m_c}:

    1. Moment Concentration: With probability at least 1−δ1 - \delta, for all c∈[l]c \in [l] and all f∈Ff \in \mathcal{F}: ∥Ex∼P~c[f(x)]−Avgx∈Sc[f(x)]∥≤p(C(f)+1)sup⁡x∈X∥x∥mc(3q+2+log⁡(4pl/δ)2+log⁡(C(f)+1))\|\mathbb{E}_{x \sim \tilde{P}_c}[f(x)] - \text{Avg}_{x \in S_c}[f(x)]\| \le \frac{p(\mathcal{C}(f) + 1) \sup_{x \in \mathcal{X}} \|x\|}{\sqrt{m_c}} \left( 3\sqrt{q} + 2 + \sqrt{\frac{\log(4pl/\delta)}{2}} + \sqrt{\log(\mathcal{C}(f) + 1)} \right) ∣Ex∼P~c[∥f(x)∥2]−Avgx∈Sc[∥f(x)∥2]∣≤pM2sup⁡x∈X∥x∥2mc(6q+4+3log⁡(4l/δ)2+3log⁡(C(f)+1))\left| \mathbb{E}_{x \sim \tilde{P}_c}[\|f(x)\|^2] - \text{Avg}_{x \in S_c}[\|f(x)\|^2] \right| \le \frac{p M^2 \sup_{x \in \mathcal{X}} \|x\|^2}{\sqrt{m_c}} \left( 6\sqrt{q} + 4 + 3\sqrt{\frac{\log(4l/\delta)}{2}} + 3\sqrt{\log(\mathcal{C}(f) + 1)} \right) where M≥C(f)M \ge \mathcal{C}(f). Both moment gaps scale as O(log⁡(1/δ)/mc)O(\sqrt{\log(1/\delta)/m_c}).

    2. Rademacher Complexity: For F∗={f∈F:C(f)≤M}\mathcal{F}^* = \{f \in \mathcal{F} : \mathcal{C}(f) \le M\} and HF∗(P~)={(μf(P~c),Varf(P~c))c=1l:f∈F∗}H_{\mathcal{F}^*}(\tilde{\mathcal{P}}) = \{(\mu_f(\tilde{P}_c), \text{Var}_f(\tilde{P}_c))_{c=1}^l : f \in \mathcal{F}^*\}: EP~[R(HF∗(P~))]≤l(1.5q+1)Msup⁡x∈X∥x∥(1+4pMsup⁡x∈X∥x∥)=O(l)\mathbb{E}_{\tilde{\mathcal{P}}}[\mathcal{R}(H_{\mathcal{F}^*}(\tilde{\mathcal{P}}))] \le \sqrt{l}(1.5\sqrt{q} + 1) M \sup_{x \in \mathcal{X}} \|x\| \left( 1 + 4p M \sup_{x \in \mathcal{X}} \|x\| \right) = O(\sqrt{l})

  6. Knowl 6 — Two-Stage Transfer Learning with Frozen Penultimate Representations and Ridge Regression

    model/method

    The transfer learning framework consists of two sequential phases:

    1. Source Pretraining: A network h~=g~∘f\tilde{h} = \tilde{g} \circ f is trained on an ll-class source dataset S~={(x~i,y~i)}i=1m\tilde{S} = \{(\tilde{x}_i, \tilde{y}_i)\}_{i=1}^m by minimizing cross-entropy loss between network logits and one-hot labels using SGD with learning rate η\eta, momentum 0.90.9, and batch size 6464. The network decomposes into a penultimate feature map f:Rd→Rpf: \mathbb{R}^d \to \mathbb{R}^p and a linear output classification layer g~:Rp→Rl\tilde{g}: \mathbb{R}^p \to \mathbb{R}^l.

    2. Target Adaptation: On a target kk-class classification task with n=k⋅ncn = k \cdot n_c training samples S={(xi,yi)}i=1nS = \{(x_i, y_i)\}_{i=1}^n (ncn_c samples per class), feature extractor ff is kept fixed without fine-tuning. The target classifier g(z)=WS,f⊤zg(z) = W_{S,f}^\top z is solved via closed-form linear ridge regression: WS,f=(f(X)⊤f(X)+λnIp)−1f(X)⊤YW_{S,f} = \left( f(X)^\top f(X) + \lambda_n I_p \right)^{-1} f(X)^\top Y where f(X)∈Rn\pf(X) \in \mathbb{R}^{n \p} is the matrix of embedded samples, Y∈Rn×kY \in \mathbb{R}^{n \times k} is the one-hot label matrix, IpI_p is the p×pp \times p identity matrix, and the regularization parameter is λn=αn\lambda_n = \alpha \sqrt{n} (default α=1\alpha = 1). Predictions for a query instance xx are given by arg⁡max⁡c∈[k](WS,f⊤f(x))c\arg\max_{c \in [k]} (W_{S,f}^\top f(x))_c.

  7. Knowl 7 — Empirical Generalization of Neural Collapse Across Samples, Classes, and Source Diversity

    empirical result

    Experimental evaluations using Wide ResNets (WRN-28-4) and convolutional networks (Conv-28-4, Conv-16-2) across CIFAR-FS, Mini-ImageNet, and EMNIST demonstrate:

    1. Within-class variation collapse (measured by both CDNV and Papyan et al.'s class-covariance normalized variance CCNV) generalizes to unseen test samples from the source classes and to completely unseen target classes during SGD training.

    2. Increasing the number of source training classes l∈{5,10,20,30,40,50,60}l \in \{5, 10, 20, 30, 40, 50, 60\} systematically decreases target CDNV (stronger neural collapse on target classes) and monotonically improves target few-shot accuracy across 1-, 5-, 10-, and 20-shot target tasks.

    3. Minimal class-mean distances min⁡i≠j∥μf(Pi)−μf(Pj)∥2\min_{i \neq j} \|\mu_f(P_i) - \mu_f(P_j)\|^2 between target classes grow with the number of source classes ll.

    4. Target few-shot accuracy peaks when the target CDNV reaches its minimum (around step 16,000 for CIFAR-FS and step 20,000 for Mini-ImageNet), after which late-stage source overfitting causes target CDNV to rise and target accuracy to decline slightly.

  8. Knowl 8 — Few-Shot Target Classification Accuracy Benchmark Results

    data/table

    The table reports 5-way 1-shot and 5-shot test target classification accuracies (%) on Mini-ImageNet, CIFAR-FS, and FC-100 benchmarks. The evaluated methods comprise standard meta-learning approaches and target-agnostic feature representations trained with cross-entropy and adapted via ridge regression on top of penultimate embeddings (WRN-28-4).

    Mini-ImageNet CIFAR-FS FC-100
    Method Architecture 1-shot 5-shot 1-shot 5-shot 1-shot 5-shot
    Matching Networks 64-64-64-64 43.56±0.8443.56 \pm 0.84 55.31±0.7355.31 \pm 0.73 - - - -
    LSTM Meta-Learner 64-64-64-64 43.44±0.7743.44 \pm 0.77 60.60±0.7160.60 \pm 0.71 - - - -
    MAML 32-32-32-32 48.70±1.8448.70 \pm 1.84 63.11±0.9263.11 \pm 0.92 58.9±1.958.9 \pm 1.9 71.5±1.071.5 \pm 1.0 - -
    Prototypical Networks 64-64-64-64 49.42±0.7849.42 \pm 0.78 68.20±0.6668.20 \pm 0.66 55.5±0.755.5 \pm 0.7 72.0±0.672.0 \pm 0.6 35.3±0.635.3 \pm 0.6 48.6±0.648.6 \pm 0.6
    Relation Networks 64-96-128-256 50.44±0.8250.44 \pm 0.82 65.32±0.765.32 \pm 0.7 55.0±1.055.0 \pm 1.0 69.3±0.869.3 \pm 0.8 - -
    SNAIL ResNet-12 55.71±0.9955.71 \pm 0.99 68.88±0.9268.88 \pm 0.92 - - - -
    TADAM ResNet-12 58.50±0.3058.50 \pm 0.30 76.7±0.376.7 \pm 0.3 - - 40.1±0.440.1 \pm 0.4 56.1±0.456.1 \pm 0.4
    AdaResNet ResNet-12 56.88±0.6256.88 \pm 0.62 71.94±0.5771.94 \pm 0.57 - - - -
    Dynamics Few-Shot 64-64-128-128 56.20±0.8656.20 \pm 0.86 73.0±0.6473.0 \pm 0.64 - - - -
    Activation to Parameter WRN-28-10 59.60±0.4159.60 \pm 0.41 73.74±0.1973.74 \pm 0.19 - - - -
    R2D2 96-192-384-512 51.2±0.651.2 \pm 0.6 68.8±0.168.8 \pm 0.1 65.3±0.265.3 \pm 0.2 79.4±0.179.4 \pm 0.1 - -
    Shot-Free ResNet-12 59.04±n/a59.04 \pm \text{n/a} 77.64±n/a77.64 \pm \text{n/a} 69.2±n/a69.2 \pm \text{n/a} 84.7±n/a84.7 \pm \text{n/a} - -
    TEWAM ResNet-12 60.07±n/a60.07 \pm \text{n/a} 75.90±n/a75.90 \pm \text{n/a} 70.4±n/a70.4 \pm \text{n/a} 81.3±n/a81.3 \pm \text{n/a} - -
    TPN ResNet-12 55.51±0.8655.51 \pm 0.86 75.64±n/a75.64 \pm \text{n/a} - - - -
    LEO WRN-28-10 61.76±0.0861.76 \pm 0.08 77.59±0.1277.59 \pm 0.12 - - - -
    MTL ResNet-12 61.20±1.8061.20 \pm 1.80 75.50±0.8075.50 \pm 0.80 - - - -
    OptNet-RR ResNet-12 61.41±0.6161.41 \pm 0.61 77.88±0.4677.88 \pm 0.46 72.6±0.772.6 \pm 0.7 84.3±0.584.3 \pm 0.5 40.5±0.640.5 \pm 0.6 57.6±0.957.6 \pm 0.9
    MetaOptNet ResNet-12 62.64±0.6162.64 \pm 0.61 78.63±0.4678.63 \pm 0.46 72.0±0.772.0 \pm 0.7 84.2±0.584.2 \pm 0.5 41.1±0.641.1 \pm 0.6 55.3±0.655.3 \pm 0.6
    Transductive Fine-Tuning WRN-28-10 65.73±0.6865.73 \pm 0.68 78.40±0.5278.40 \pm 0.52 76.58±0.6876.58 \pm 0.68 85.79±0.585.79 \pm 0.5 43.16±0.5943.16 \pm 0.59 57.57±0.5557.57 \pm 0.55
    Distill-simple ResNet-12 62.02±0.6362.02 \pm 0.63 79.64±0.4479.64 \pm 0.44 71.5±0.871.5 \pm 0.8 86.0±0.586.0 \pm 0.5 42.6±0.742.6 \pm 0.7 59.1±0.659.1 \pm 0.6
    Distill ResNet-12 64.82±0.6064.82 \pm 0.60 82.14±0.4382.14 \pm 0.43 73.9±0.873.9 \pm 0.8 86.9±0.586.9 \pm 0.5 44.6±0.744.6 \pm 0.7 60.9±0.660.9 \pm 0.6
    Ours (simple) WRN-28-4 58.12±1.1958.12 \pm 1.19 72.0±0.9972.0 \pm 0.99 68.81±1.2068.81 \pm 1.20 81.49±0.9881.49 \pm 0.98 44.96±1.1444.96 \pm 1.14 57.21±10.8957.21 \pm 10.89
    Ours (lr scheduling) WRN-28-4 60.37±1.2560.37 \pm 1.25 72.35±0.9972.35 \pm 0.99 70.0±1.2970.0 \pm 1.29 81.39±0.9681.39 \pm 0.96 43.42±1.043.42 \pm 1.0 54.14±1.154.14 \pm 1.1
    Ours (lr sched. + model sel.) WRN-28-4 61.27±1.1461.27 \pm 1.14 74.74±0.7674.74 \pm 0.76 72.37±1.1272.37 \pm 1.12 82.94±0.8982.94 \pm 0.89 45.81±1.2745.81 \pm 1.27 56.85±1.3056.85 \pm 1.30

    The results show that standard classification pretraining followed by ridge regression is competitive with complex meta-learning algorithms and achieves state-of-the-art 1-shot performance (45.81%45.81\%) on FC-100.

  9. Knowl 9 — Layer-wise Comparison of Neural Collapse in Deep Networks

    empirical result

    In a WRN-28-4 network trained on CIFAR-FS with l=64l = 64 source classes, comparing the penultimate (top embedding) layer with the second-to-last embedding layer reveals:

    1. Both layers exhibit neural collapse across source training data, source test data, and unseen target classes.

    2. The penultimate embedding layer exhibits markedly lower CDNV than the second-to-last embedding layer across all three datasets throughout training.

    3. The 5-shot 5-class target accuracy obtained from the penultimate layer is consistently higher than that obtained from the second-to-last layer, indicating that within-class variance collapse becomes more pronounced in deeper layers and directly tracks few-shot generalization performance.

Coverage note — Omitted mathematical derivations and intermediate lemma calculations (such as the proof of Lemma 1 and the Ramanujan Gamma approximation in Lemma 2) as they serve solely to establish the main stated propositions.

References

  1. 1.Peter L. Bartlett and Shahar Mendelson. Rademacher and gaussian complexities: Risk bounds and structural results. J. Mach. Learn. Res., 3:463–482, 2002.
  2. 2.Peter L. Bartlett, Dylan J. Foster, and Matus Telgarsky. Spectrally-normalized margin bounds for neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, pp. 6241–6250, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN 9781510860964.
  3. 3.Jonathan Baxter. A model of inductive bias learning. J. Artif. Int. Res., 12(1):149–198, March 2000.
  4. 4.Shai Ben-David, John Blitzer, Koby Crammer, and Fernando Pereira. Analysis of representations for domain adaptation. In Advances in Neural Information Processing Systems, pp. 137–144, 2006.
  5. 5.Yoshua Bengio. Deep learning of representations for unsupervised and transfer learning. In Proceedings of ICML Workshop on Unsupervised and Transfer Learning, volume 27 of Proceedings of Machine Learning Research, pp. 17–36, Bellevue, Washington, USA, 02 Jul 2012. PMLR.
  6. 6.Luca Bertinetto, Joao F. Henriques, Philip Torr, and Andrea Vedaldi. Meta-learning with differentiable closed-form solvers. In International Conference on Learning Representations, 2019.
  7. 7.Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al. On the opportunities and risks of foundation models. CoRR, abs/2108.07258, 2021.
  8. 8.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020.
  9. 9.Rich Caruana. Learning many related tasks at the same time with backpropagation. In Advances in Neural Information Processing Systems, volume 7. MIT Press, 1995.
  10. 10.Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su. This looks like that: Deep learning for interpretable image recognition. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019a.
  11. 11.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV (7), volume 11211 of Lecture Notes in Computer Science, pp. 833–851. Springer, 2018.
  12. 12.Sihong Chen, Kai Ma, and Yefeng Zheng. Med3d: Transfer learning for 3d medical image analysis. CoRR, abs/1904.00625, 2019b.
  13. 13.Gregory Cohen, Saeed Afshar, Jonathan Tapson, and Andre van Schaik. Emnist: an extension of mnist to handwritten letters. arXiv preprint arXiv:1702.05373, 2017.
  14. 14.Guneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, and Stefano Soatto. A baseline for few-shot image classification. In International Conference on Learning Representations, 2020.
  15. 15.Simon Shaolei Du, Wei Hu, Sham M. Kakade, Jason D. Lee, and Qi Lei. Few-shot learning via learning the representation, provably. In International Conference on Learning Representations, 2021.
  16. 16.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pp. 1126–1135. PMLR, 06–11 Aug 2017.
  17. 17.Tomer Galanti, Lior Wolf, and Tamir Hazan. A theoretical framework for deep transfer learning. Information and Inference: A Journal of the IMA, 5:159–209, 2016.
  18. 18.Spyros Gidaris and Nikos Komodakis. Dynamic few-shot visual learning without forgetting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  19. 19.Noah Golowich, Alexander Rakhlin, and Ohad Shamir. Size-independent sample complexity of neural networks. Information and Inference: A Journal of the IMA, 9, 12 2017. doi: 10.1093/imaiai/iaz007.
  20. 20.X. Y. Han, Vardan Papyan, and David L. Donoho. Neural collapse under mse loss: Proximity to and dynamics on the central path, 2021.
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770–778, 2016. doi: 10.1109/CVPR.2016.90.
  22. 22.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct 2017.
  23. 23.Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications, 2017.
  24. 24.Zixuan Huang and Yin Li. Interpretable and accurate fine-grained recognition via region grouping. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8662–8672, 2020.
  25. 25.Sham Kakade and Ambuj Tewari. Learning theory, 2008.
  26. 26.Ekatherina A. Karatsuba. On the asymptotic representation of the euler gamma function by ramanujan. J. Comput. Appl. Math., 135(2):225–240, October 2001.
  27. 27.Mikhail Khodak, Maria-Florina F Balcan, and Ameet S Talwalkar. Adaptive gradient-based meta-learning methods. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019.
  28. 28.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 05 2012.
  29. 29.Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. Meta-learning with differentiable convex optimization. In CVPR, 2019.
  30. 30.Yanbin Liu, Juho Lee, Minseop Park, Saehoon Kim, Eunho Yang, Sungju Hwang, and Yi Yang. Learning to propagate labels: Transductive propagation network for few-shot learning. In International Conference on Learning Representations, 2019.
  31. 31.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3431–3440, 2015. doi: 10.1109/CVPR.2015.7298965.
  32. 32.Yishay Mansour, Mehryar Mohri, and Afshin Rostamizadeh. Domain adaptation: Learning bounds and algorithms. In Proceedings of The 22nd Annual Conference on Learning Theory (COLT 2009), Montreal, Canada, 2009.
  33. 33.Andreas Maurer and M. Pontil. Uniform concentration and symmetrization for weak interactions. In COLT, 2019.
  34. 34.Andreas Maurer, Massimiliano Pontil, and Bernardino Romera-Paredes. The benefit of multitask representation learning. J. Mach. Learn. Res., 17(1):2853–2884, January 2016.
  35. 35.Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. A simple neural attentive meta-learner. In International Conference on Learning Representations, 2018.
  36. 36.Dustin G. Mixon, Hans Parshall, and Jianzong Pi. Neural collapse with unconstrained features, 2020.
  37. 37.Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. Foundations of Machine Learning. Adaptive Computation and Machine Learning. MIT Press, Cambridge, MA, 2nd edition, 2018.
  38. 38.Tsendsuren Munkhdalai, Xingdi Yuan, Soroush Mehri, and Adam Trischler. Rapid adaptation with conditionally shifted neurons. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 3664–3673. PMLR, 2018.
  39. 39.Boris Oreshkin, Pau Rodríguez Lopez, and Alexandre Lacoste. Tadam: Task dependent adaptive metric for improved few-shot learning. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018.
  40. 40.Vardan Papyan, X. Y. Han, and David L. Donoho. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences, 117(40): 24652–24663, 2020.
  41. 41.Anastasia Pentina and Christoph Lampert. A pac-bayesian bound for lifelong learning. In Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pp. 991–999, Bejing, China, 22–24 Jun 2014. PMLR.
  42. 42.T. Poggio and Q. Liao. Implicit dynamic regularization in deep networks. Technical report, Center for Brains, Minds and Machines (CBMM), 2020a.
  43. 43.Tomaso Poggio and Qianli Liao. Explicit regularization and implicit bias in deep network classifiers trained with the square loss, 2020b.
  44. 44.Tomaso Poggio, Andrzej Banburski, and Qianli Liao. Theoretical issues in deep networks. Proceedings of the National Academy of Sciences, 117(48):30039–30045, 2020.
  45. 45.Limeng Qiao, Yemin Shi, Jia Li, Yaowei Wang, Tiejun Huang, and Yonghong Tian. Transductive episodic-wise adaptive metric for few-shot learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019.
  46. 46.Siyuan Qiao, Chenxi Liu, Wei Shen, and Alan L. Yuille. Few-shot image recognition by predicting parameters from activations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  47. 47.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. CoRR, abs/2102.12092, 2021.
  48. 48.Akshay Rangamani, Mengjia Xu, Andrzej Banburski, Qianli Liao, and Tomaso Poggio. Dynamics and neural collapse in deep classifiers trained with the square loss. Technical report, Center for Brains, Minds and Machines (CBMM), 2021.
  49. 49.S. Ravi and H. Larochelle. Optimization as a model for few-shot learning. In International Conference on Learning Representations, 2017.
  50. 50.Avinash Ravichandran, Rahul Bhotika, and Stefano Soatto. Few-shot learning with embedded class models and shot-free meta training. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 331–339, 2019.
  51. 51.Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016.
  52. 52.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015.
  53. 53.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015.
  54. 54.Andrei A. Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Hadsell. Meta-learning with latent embedding optimization. In International Conference on Learning Representations, 2019.
  55. 55.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. CoRR, abs/1409.1556, 2014.
  56. 56.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.
  57. 57.Qianru Sun, Yaoyao Liu, Tat-Seng Chua, and Bernt Schiele. Meta-transfer learning for few-shot learning. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  58. 58.Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip H.S. Torr, and Timothy M. Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018.
  59. 59.Yonglong Tian, Yue Wang, Dilip Krishnan, Joshua B. Tenenbaum, and Phillip Isola. Rethinking few-shot image classification: A good embedding is all you need? In Proceedings of the European Conference on Computer Vision (ECCV), volume 12359 of Lecture Notes in Computer Science, pp. 266–282. Springer, 2020.
  60. 60.Oriol Vinyals, Charles Blundell, Timothy Lillicrap, koray kavukcuoglu, and Daan Wierstra. Matching networks for one shot learning. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc., 2016.
  61. 61.Ze Yang, Tiange Luo, Dong Wang, Zhiqiang Hu, Jun Gao, and Liwei Wang. Learning to navigate for fine-grained classification. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.
  62. 62.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014.
  63. 63.Sergey Zagoruyko and N. Komodakis. Wide residual networks. ArXiv, abs/1605.07146, 2016.
  64. 64.Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu. A geometric analysis of neural collapse with unconstrained features, 2021.

Citation

MLA
Galanti, T., et al. “On the Role of Neural Collapse in Transfer Learning”. arXiv, 2021, http://arxiv.org/abs/2112.15121v2.
APA
Galanti, T., György, A., & Hutter, M. (2021). On the Role of Neural Collapse in Transfer Learning. arXiv. http://arxiv.org/abs/2112.15121v2
Chicago
Galanti, T., A. György, and M. Hutter. 2021. “On the Role of Neural Collapse in Transfer Learning”. arXiv. http://arxiv.org/abs/2112.15121v2.
Harvard
Galanti, T., György, A. and Hutter, M. (2021) “On the Role of Neural Collapse in Transfer Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.15121v2.
Vancouver
1. Galanti T, György A, Hutter M (2021) On the Role of Neural Collapse in Transfer Learning. arXiv

BibTeX

@article{galanti2021the,
  title = {On the Role of Neural Collapse in Transfer Learning},
  author = {Galanti, Tomer and György, András and Hutter, Marcus},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.15121v2},
  eprint = {2112.15121}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/