High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

Jimmy BaMurat A. ErdogduTaiji SuzukiZhichao WangDenny WuGreg Yang

article2022NeurIPS172 citations

Proves that taking a single gradient descent step on the first layer of a two-layer neural network with a sufficiently large learning rate enables the resulting kernel to outperform fixed random features and surpass linear estimators in high dimensions.

Listen

Modern machine learning increasingly relies on deep neural networks because they automatically extract and adapt useful internal representations from data, a capability known as feature learning. In contrast, traditional kernel methods and fixed random feature models rely on static, unlearned transformations that often perform no better than simple linear predictors in high-dimensional tasks. While empirical observations show that critical representation learning occurs in the very earliest training iterations, theoretical frameworks have struggled to precisely capture when, how, and by how much gradient-based training breaks free from the limitations of fixed kernels.

The article establishes a precise mathematical framework to evaluate how taking a single gradient descent step on the first layer of a two-layer neural network improves representation quality and prediction risk. Specifically, it analyzes whether one feature learning step enables the model to outperform fixed kernel benchmarks in high-dimensional settings where the dataset size, input dimensionality, and hidden neuron count scale proportionally.

To conduct this evaluation, the authors combine random matrix theory and operator-valued free probability with high-dimensional statistical simulations. They study a student-teacher setup where the teacher is a single-index model with Gaussian inputs. The first-layer network weights undergo one step of gradient descent on a training set, and the quality of the updated features is then evaluated by computing the prediction risk of kernel ridge regression trained on a fresh sample of independent data. The analysis focuses on two distinct learning rate regimes: standard moderate step sizes and large step sizes that scale proportionally with the square root of the network width.

The investigation yields several key findings regarding model behavior. First, the article proves that the initial gradient update matrix is approximately rank-1 and aligns directly with the underlying target signal. Second, under a moderate learning rate, the model satisfies a Gaussian Equivalence property; the single gradient step consistently reduces prediction risk compared to the initial random features, yet the model remains confined to a linear regime and cannot defeat the best linear estimator on the raw input. Third, when the learning rate is scaled up proportionally to the square root of the network width, the neurons lose their near-orthogonality and break the linear barrier. In this large step-size regime, the updated features allow the model to achieve substantially lower prediction risk than standard fixed kernel lower bounds for certain nonlinear activation functions, achieving risk improvements that scale directly with the ratio of data dimensions to sample size.

These findings provide fundamental insights for practical machine learning engineering, training efficiency, and computational resource allocation. The results demonstrate that neural networks can realize substantial representational benefits immediately at the onset of optimization rather than requiring thousands of steps to escape the kernel regime. Crucially, the analysis reveals that learning rate magnitude during the initial phase acts as a structural switch: small learning rates restrict networks to linear approximations, whereas sufficiently large initial updates unlock true non-linear feature adaptation. This validates aggressive initial learning rate schedules and parameterization frameworks designed for large models.

For practitioners and engineering teams, the analysis suggests using sufficiently large learning rates during initial training phases to ensure the network exits the restrictive kernel regime and forms task-adapted features early. Organizations developing pretraining and transfer learning pipelines should prioritize capturing this initial phase of representation learning before fine-tuning readout layers. Furthermore, future technical investigations should explore intermediate learning rates to pinpoint the exact transition boundary between regimes, evaluate multi-step training dynamics, and extend the theoretical guarantees to settings where representation updates and regression are trained simultaneously on identical data samples.

Readers should interpret these findings within the scope of the article's foundational assumptions. The rigorous results rely on proportional asymptotic scaling, Gaussian input distributions, and single-index target functions, with feature updating and final readout regression performed on independent data splits. While the core qualitative behaviors are robustly supported by both rigorous mathematical proofs and empirical simulations, caution is warranted when extrapolating these specific quantitative error bounds directly to complex, non-Gaussian data architectures.

Cover for High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

Abstract

We study the first gradient descent step on the first-layer parameters W in a two-layer neural network: f(x) = 1/√N a^⊤ σ(W^⊤ x), where W ∈ ℝ^{d×N}, a ∈ ℝ^N are randomly initialized, and the training objective is the empirical MSE loss: 1/n ∑_{i=1}^n (f(x_i) - y_i)^2. In the proportional asymptotic limit where n, d, N → ∞ at the same rate, and an idealized student-teacher setting where the teacher f^* is a single-index model, we compute the prediction risk of ridge regression on the conjugate kernel after one gradient step on W with learning rate η. We consider two scalings of the first step learning rate η. For small η, we establish a Gaussian equivalence property for the trained feature map, and prove that the learned kernel improves upon the initial random feature model, but cannot defeat the best linear model on the input. Whereas for sufficiently large η, we prove that for certain f^*, the same ridge estimator on trained features can go beyond this “linear regime” and outperform a wide range of (fixed) kernels. Our results demonstrate that even one gradient step can lead to a considerable advantage over random features, and highlight the role of learning rate scaling in the initial phase of training.

Table of Contents

  • 1 Introduction
  • 1.1 Contributions
  • 1.2 Related works
  • 2 Problem setup and assumptions
  • 2.1 Training procedure
  • 2.2 Student-teacher setting and main assumptions
  • 3 Preliminary results
  • 3.1 Lower bound for kernel ridge regression
  • 3.2 Almost rank-1 property of the gradient matrix
  • 4 η = Θ(1) : improvement over the initial CK
  • 4.1 The Gaussian equivalence property
  • 4.2 Precise asymptotics of CK ridge regression
  • 5 η = Θ( √ N ) : improvement over the kernel lower bound
  • 6 Conclusion
  • Acknowledgement
  • References
  • Checklist

Knowls

  1. Knowl 1 — Two-Stage Feature Learning and Ridge Regression Training Setup

    model/method

    Consider a fully-connected two-layer neural network parameterized as

    fNN(x)=1Na⊤σ(W⊤x)=1N∑j=1Najσ(⟨x,wj⟩),f_{\text{NN}}(x) = \frac{1}{\sqrt{N}} a^\top \sigma(W^\top x) = \frac{1}{\sqrt{N}} \sum_{j=1}^N a_j \sigma(\langle x, w_j \rangle),

    where x∈Rdx \in \mathbb{R}^d, W=[w1,…,wN]∈Rd×NW = [w_1, \dots, w_N] \in \mathbb{R}^{d \times N}, a∈RNa \in \mathbb{R}^N, and σ:R→R\sigma: \mathbb{R} \to \mathbb{R} is an activation function applied entrywise. The learning process consists of two stages:

    1. First-layer gradient update: Starting from random initialization W0∈Rd×NW_0 \in \mathbb{R}^{d \times N} and fixed a∈RNa \in \mathbb{R}^N, the first layer is trained for one gradient descent step with learning rate η>0\eta > 0 on empirical squared loss over dataset X=[x1,…,xn]⊤∈Rn×dX = [x_1, \dots, x_n]^\top \in \mathbb{R}^{n \times d} and y=[y1,…,yn]⊤∈Rny = [y_1, \dots, y_n]^\top \in \mathbb{R}^n:

      W1=W0+ηNG0,W_1 = W_0 + \eta \sqrt{N} G_0,

      where the gradient matrix G0∈Rd×NG_0 \in \mathbb{R}^{d \times N} is given by

      G0:=1nX⊤[(1N(y−1Nσ(XW0)a)a⊤)⊙σ′(XW0)],G_0 := \frac{1}{n} X^\top \left[ \left( \frac{1}{\sqrt{N}} \left(y - \frac{1}{\sqrt{N}} \sigma(X W_0) a\right) a^\top \right) \odot \sigma'(X W_0) \right],

      with ⊙\odot denoting the entrywise Hadamard product and σ′\sigma' the derivative of σ\sigma.

    2. Second-layer ridge regression on fresh data: Given the updated representation map x↦1Nσ(W1⊤x)x \mapsto \frac{1}{\sqrt{N}} \sigma(W_1^\top x), the second-layer weights are fitted on an independent fresh dataset (X~,y~)(\tilde{X}, \tilde{y}) of size nn, with X~∈Rn×d\tilde{X} \in \mathbb{R}^{n \times d} and y~∈Rn\tilde{y} \in \mathbb{R}^n. Defining Φ=1Nσ(X~W1)∈Rn×N\Phi = \frac{1}{\sqrt{N}} \sigma(\tilde{X} W_1) \in \mathbb{R}^{n \times N}, the ridge estimator is

      a^λ=arg⁡min⁡a∈RN{1n∥y~−Φa∥22+λN∥a∥22}=(Φ⊤Φ+λnNIN)−1Φ⊤y~,\hat{a}_\lambda = \arg\min_{a \in \mathbb{R}^N} \left\{ \frac{1}{n} \|\tilde{y} - \Phi a\|_2^2 + \frac{\lambda}{N} \|a\|_2^2 \right\} = \left(\Phi^\top \Phi + \frac{\lambda n}{N} I_N\right)^{-1} \Phi^\top \tilde{y},

      with regularization parameter λ>0\lambda > 0. The test prediction risk against a teacher function f∗f^* is evaluated under x∼N(0,Id)x \sim \mathcal{N}(0, I_d) as R1(λ)=Ex[(1Na^λ⊤σ(W1⊤x)−f∗(x))2]\mathcal{R}_1(\lambda) = \mathbb{E}_x [(\frac{1}{\sqrt{N}} \hat{a}_\lambda^\top \sigma(W_1^\top x) - f^*(x))^2].

  2. Knowl 2 — High-Dimensional Single-Index Student-Teacher Model Assumptions

    assumption

    The analysis of the two-layer neural network operates under the following set of assumptions:

    1. Proportional asymptotic scaling: The sample size nn, input dimensionality dd, and hidden width NN jointly tend to infinity with asymptotic aspect ratios

      nd→ψ1∈(0,∞)andNd→ψ2∈(0,∞).\frac{n}{d} \to \psi_1 \in (0, \infty) \quad \text{and} \quad \frac{N}{d} \to \psi_2 \in (0, \infty).
    2. Gaussian initialization: The weights are initialized independently as d[W0]ij∼i.i.d.N(0,1)\sqrt{d} [W_0]_{ij} \stackrel{\text{i.i.d.}}{\sim} \mathcal{N}(0, 1) for i∈[d],j∈[N]i \in [d], j \in [N], and N[a]j∼i.i.d.N(0,1)\sqrt{N} [a]_j \stackrel{\text{i.i.d.}}{\sim} \mathcal{N}(0, 1) for j∈[N]j \in [N].

    3. Normalized activation function: The activation function σ:R→R\sigma: \mathbb{R} \to \mathbb{R} has bounded first three derivatives almost surely and satisfies the Hermite coefficients

      μ0:=Ez∼N(0,1)[σ(z)]=0,μ1:=Ez∼N(0,1)[zσ(z)]≠0,μ2:=Ez∼N(0,1)[σ(z)2]−μ12≠0.\mu_0 := \mathbb{E}_{z \sim \mathcal{N}(0, 1)}[\sigma(z)] = 0, \quad \mu_1 := \mathbb{E}_{z \sim \mathcal{N}(0, 1)}[z \sigma(z)] \neq 0, \quad \mu_2 := \sqrt{\mathbb{E}_{z \sim \mathcal{N}(0, 1)}[\sigma(z)^2] - \mu_1^2} \neq 0.
    4. Single-index teacher model: Labels are generated by yi=f∗(xi)+εiy_i = f^*(x_i) + \varepsilon_i, where xi∼i.i.d.N(0,Id)x_i \stackrel{\text{i.i.d.}}{\sim} \mathcal{N}(0, I_d) and εi\varepsilon_i is independent sub-Gaussian noise with mean zero and variance σε2\sigma_\varepsilon^2. The teacher function is f∗(x)=σ∗(⟨x,β∗⟩)f^*(x) = \sigma^*(\langle x, \beta^* \rangle) with unknown unit vector β∗∈Rd\beta^* \in \mathbb{R}^d (∥β∗∥2=1\|\beta^*\|_2 = 1) and Lipschitz link σ∗\sigma^*, decomposed in L2(Rd,N(0,Id))L_2(\mathbb{R}^d, \mathcal{N}(0, I_d)) as

      f∗(x)=μ0∗+μ1∗⟨x,β∗⟩+P>1f∗(x),with μ0∗=0, μ1∗≠0,f^*(x) = \mu_0^* + \mu_1^* \langle x, \beta^* \rangle + P_{>1} f^*(x), \quad \text{with } \mu_0^* = 0, \, \mu_1^* \neq 0,

      where P>1P_{>1} is the orthogonal projection operator onto functions orthogonal to degree-≤1\le 1 polynomials, with ∥P>1f∗∥L2→μ2∗\|P_{>1} f^*\|_{L_2} \to \mu_2^* as d→∞d \to \infty.

  3. Knowl 3 — Linear Risk Lower Bound for Fixed Kernels and Random Features

    theoretical result

    In the proportional asymptotic limit where n,d,N→∞n, d, N \to \infty with n/d→ψ1∈(0,∞)n/d \to \psi_1 \in (0, \infty) and N/d→ψ2∈(0,∞)N/d \to \psi_2 \in (0, \infty), consider Gaussian input vectors x∼N(0,Id)x \sim \mathcal{N}(0, I_d) and a single-index teacher f∗(x)=σ∗(⟨x,β∗⟩)f^*(x) = \sigma^*(\langle x, \beta^* \rangle) with ∥β∗∥2=1\|\beta^*\|_2=1.

    For ridge regression estimators based on:

    • the initialized conjugate kernel (CK) feature map ϕCK(x)=1Nσ(W0⊤x)\phi_{\text{CK}}(x) = \frac{1}{\sqrt{N}} \sigma(W_0^\top x) with prediction risk RCK(λ)\mathcal{R}_{\text{CK}}(\lambda);
    • the neural tangent kernel (NTK) feature map ϕNTK(x)=1Ndvec(σ′(W0⊤x)x⊤)\phi_{\text{NTK}}(x) = \frac{1}{\sqrt{Nd}} \text{vec}(\sigma'(W_0^\top x) x^\top) with prediction risk RNTK(λ)\mathcal{R}_{\text{NTK}}(\lambda);
    • any rotationally invariant kernel k(x,y)=g(⟨x,y⟩/d)k(x, y) = g(\langle x, y \rangle / d) or k(x,y)=g(∥x−y∥22/d)k(x, y) = g(\|x - y\|_2^2 / d) with smooth gg, with prediction risk Rker(λ)\mathcal{R}_{\text{ker}}(\lambda);

    the prediction risk for any regularization parameter λ>0\lambda > 0 is asymptotically lower bounded by the L2L_2-norm of the nonlinear part of the target function:

    inf⁡λ>0min⁡{RCK(λ),RNTK(λ),Rker(λ)}≥∥P>1f∗∥L22+od,P(1),\inf_{\lambda > 0} \min \left\{ \mathcal{R}_{\text{CK}}(\lambda), \mathcal{R}_{\text{NTK}}(\lambda), \mathcal{R}_{\text{ker}}(\lambda) \right\} \ge \|P_{>1} f^*\|_{L_2}^2 + o_{d,P}(1),

    where P>1P_{>1} is the projector orthogonal to degree-≤1\le 1 polynomials under the Gaussian input measure. Thus, fixed kernel models cannot outperform the best linear estimator on the input features in the proportional regime.

  4. Knowl 4 — Rank-One Approximation of the Initial Gradient Step Matrix

    theoretical result

    Under proportional scaling n,d,N→∞n, d, N \to \infty, Gaussian initialization d[W0]ij∼N(0,1)\sqrt{d}[W_0]_{ij} \sim \mathcal{N}(0, 1), N[a]j∼N(0,1)\sqrt{N}[a]_j \sim \mathcal{N}(0, 1), and single-index teacher yi=σ∗(⟨xi,β∗⟩)+εiy_i = \sigma^*(\langle x_i, \beta^* \rangle) + \varepsilon_i, let the gradient update matrix on the first layer be

    G0:=1ηN(W1−W0)=1nX⊤[(1N(y−1Nσ(XW0)a)a⊤)⊙σ′(XW0)].G_0 := \frac{1}{\eta \sqrt{N}} (W_1 - W_0) = \frac{1}{n} X^\top \left[ \left( \frac{1}{\sqrt{N}} \left(y - \frac{1}{\sqrt{N}} \sigma(X W_0) a\right) a^\top \right) \odot \sigma'(X W_0) \right].

    Define the deterministic-structure rank-1 matrix

    A:=μ1nNX⊤ya⊤∈Rd×N,A := \frac{\mu_1}{n \sqrt{N}} X^\top y a^\top \in \mathbb{R}^{d \times N},

    where μ1=Ez∼N(0,1)[zσ(z)]\mu_1 = \mathbb{E}_{z \sim \mathcal{N}(0, 1)}[z \sigma(z)]. There exist positive constants c,C>0c, C > 0 such that for all large n,d,Nn, d, N,

    ∥G0−A∥2≤Clog⁡2nn∥G0∥2\|G_0 - A\|_2 \le C \frac{\log^2 n}{\sqrt{n}} \|G_0\|_2

    with probability at least 1−ne−clog⁡2n1 - n e^{-c \log^2 n}, where ∥⋅∥2\|\cdot\|_2 denotes the spectral operator norm. Consequently, the leading directional update of W1W_1 concentrates along the rank-1 tensor product of the data response vector X⊤yX^\top y and the initial readout weights aa.

  5. Knowl 5 — Gaussian Equivalence Theorem for One-Step Trained Features at Small Learning Rate

    theoretical result

    Let W1=W0+ηNG0W_1 = W_0 + \eta \sqrt{N} G_0 be the first-layer weight matrix trained for one gradient descent step with step size η=Θ(1)\eta = \Theta(1) on dataset (X,y)(X, y) in the proportional limit (n,d,N→∞n, d, N \to \infty with n/d→ψ1n/d \to \psi_1 and N/d→ψ2N/d \to \psi_2). Assume the activation function σ\sigma is odd and has bounded first three derivatives, and the ridge regression estimator a^λ\hat{a}_\lambda is fitted on an independent fresh dataset (X~,y~)(\tilde{X}, \tilde{y}).

    Define the nonlinear conjugate feature map ϕCK(x)=1Nσ(W1⊤x)\phi_{\text{CK}}(x) = \frac{1}{\sqrt{N}} \sigma(W_1^\top x) and the linear Gaussian equivalent (GE) feature map

    ϕGE(x)=1N(μ1W1⊤x+μ2z)∈RN,\phi_{\text{GE}}(x) = \frac{1}{\sqrt{N}} (\mu_1 W_1^\top x + \mu_2 z) \in \mathbb{R}^N,

    where μ1=Ez0∼N(0,1)[z0σ(z0)]\mu_1 = \mathbb{E}_{z_0 \sim \mathcal{N}(0, 1)}[z_0 \sigma(z_0)], μ2=E[σ(z0)2]−μ12\mu_2 = \sqrt{\mathbb{E}[\sigma(z_0)^2] - \mu_1^2}, and z∼N(0,IN)z \sim \mathcal{N}(0, I_N) is independent of xx and W1W_1.

    For any λ>0\lambda > 0, the prediction risk RCK(λ)\mathcal{R}_{\text{CK}}(\lambda) of the nonlinear model equals that of the Gaussian equivalent model RGE(λ)\mathcal{R}_{\text{GE}}(\lambda) up to vanishing error:

    ∣RCK(λ)−RGE(λ)∣=od,P(1).|\mathcal{R}_{\text{CK}}(\lambda) - \mathcal{R}_{\text{GE}}(\lambda)| = o_{d, P}(1).

    Because the Gaussian equivalent model cannot learn nonlinear components of f∗f^*, its risk remains lower-bounded by RCK(λ)≥∥P>1f∗∥L22+od,P(1)\mathcal{R}_{\text{CK}}(\lambda) \ge \|P_{>1} f^*\|_{L_2}^2 + o_{d,P}(1).

  6. Knowl 6 — Exact Asymptotic Risk Reduction from One Feature Learning Step at Small Step Size

    theoretical result

    Let R0(λ)\mathcal{R}_0(\lambda) be the prediction risk of kernel ridge regression on the randomly initialized conjugate kernel ϕCK,0(x)=1Nσ(W0⊤x)\phi_{\text{CK}, 0}(x) = \frac{1}{\sqrt{N}} \sigma(W_0^\top x), and let R1(λ)\mathcal{R}_1(\lambda) be the prediction risk on the trained feature map ϕCK,1(x)=1Nσ(W1⊤x)\phi_{\text{CK}, 1}(x) = \frac{1}{\sqrt{N}} \sigma(W_1^\top x) after one gradient descent step with learning rate η=Θ(1)\eta = \Theta(1) under proportional scaling (n/d→ψ1n/d \to \psi_1, N/d→ψ2N/d \to \psi_2) and odd activation σ\sigma.

    The difference in prediction risk converges in probability to a non-negative deterministic limit:

    R0(λ)−R1(λ)→Pδ(η,λ,ψ1,ψ2)≥0,\mathcal{R}_0(\lambda) - \mathcal{R}_1(\lambda) \xrightarrow{P} \delta(\eta, \lambda, \psi_1, \psi_2) \ge 0,

    where δ\delta is strictly positive if and only if μ1≠0\mu_1 \neq 0, μ1∗≠0\mu_1^* \neq 0, and η>0\eta > 0. The asymptotic improvement satisfies:

    1. Strict benefit across all data and width regimes: δ>0\delta > 0 holds for any finite ψ1,ψ2∈(0,∞)\psi_1, \psi_2 \in (0, \infty), meaning one gradient step strictly reduces generalization risk even when the sample size nn is small and even when σ≠σ∗\sigma \neq \sigma^*.
    2. Linear constraint: The total risk reduction is bounded by δ≤R0(λ)−(μ2∗)2\delta \le \mathcal{R}_0(\lambda) - (\mu_2^*)^2, maintaining the estimator within the linear regime.
    3. Asymptotic limits: In the large-sample limit ψ1→∞\psi_1 \to \infty, δ\delta increases monotonically with the learning rate η\eta. In the large-width limit ψ2→∞\psi_2 \to \infty, δ→0\delta \to 0, showing that the advantage of η=Θ(1)\eta = \Theta(1) feature learning over random features vanishes in ultra-wide networks.
  7. Knowl 7 — Asymptotic Risk Upper Bound for Maximal Update Parameterization Feature Learning

    theoretical result

    Let the first-layer weights W1=W0+ηNG0W_1 = W_0 + \eta \sqrt{N} G_0 be updated with a large learning rate η=Θ(N)\eta = \Theta(\sqrt{N}), matching the maximal update parameterization scaling where ∥W1−W0∥F≍∥W0∥F\|W_1 - W_0\|_F \asymp \|W_0\|_F. Assume that the activation function σ\sigma is bounded with bounded derivatives, and the target is a single-index model f∗(x)=σ∗(⟨x,β∗⟩)f^*(x) = \sigma^*(\langle x, \beta^* \rangle).

    Define the optimal scalar approximation error τ∗\tau^* between σ∗\sigma^* and σ\sigma:

    τ∗:=inf⁡κ∈REξ1∼N(0,1)[(σ∗(ξ1)−Eξ2∼N(0,1)[σ(κξ1+ξ2)])2].\tau^* := \inf_{\kappa \in \mathbb{R}} \mathbb{E}_{\xi_1 \sim \mathcal{N}(0, 1)} \left[ \left( \sigma^*(\xi_1) - \mathbb{E}_{\xi_2 \sim \mathcal{N}(0, 1)} [\sigma(\kappa \xi_1 + \xi_2)] \right)^2 \right].

    Then there exist positive constants C>0C > 0 and ψ1∗>0\psi_1^* > 0 such that for any sample-to-dimension ratio ψ1=n/d>ψ1∗\psi_1 = n/d > \psi_1^*, the second-layer ridge regression estimator a^λ\hat{a}_\lambda trained on fresh data with regularization nε−1<N−1λ<n−εn^{\varepsilon - 1} < N^{-1}\lambda < n^{-\varepsilon} for small ε>0\varepsilon > 0 satisfies

    R1(λ)≤10τ∗+C(τ∗dn+dn)\mathcal{R}_1(\lambda) \le 10 \tau^* + C \left( \sqrt{\tau^*} \sqrt{\frac{d}{n}} + \frac{d}{n} \right)

    with probability 1 as n,d,N→∞n, d, N \to \infty proportionally.

  8. Knowl 8 — Non-Linear Separation of Trained Conjugate Kernels from Fixed Kernel Methods

    theoretical result

    Under proportional scaling n≍d≍Nn \asymp d \asymp N, when the two-layer network is trained for one gradient descent step with large learning rate η=Θ(N)\eta = \Theta(\sqrt{N}), if the target link σ∗\sigma^* and activation σ\sigma satisfy ∥P>1f∗∥L22≥10τ∗\|P_{>1} f^*\|_{L_2}^2 \ge 10 \tau^*, where

    τ∗:=inf⁡κ∈REξ1∼N(0,1)[(σ∗(ξ1)−Eξ2∼N(0,1)[σ(κξ1+ξ2)])2],\tau^* := \inf_{\kappa \in \mathbb{R}} \mathbb{E}_{\xi_1 \sim \mathcal{N}(0, 1)} \left[ \left( \sigma^*(\xi_1) - \mathbb{E}_{\xi_2 \sim \mathcal{N}(0, 1)} [\sigma(\kappa \xi_1 + \xi_2)] \right)^2 \right],

    then for sufficiently large sample-to-dimension ratio ψ1=n/d>ψ1∗\psi_1 = n/d > \psi_1^*, the conjugate kernel ridge regression estimator satisfies

    R1(λ)<∥P>1f∗∥L22,\mathcal{R}_1(\lambda) < \|P_{>1} f^*\|_{L_2}^2,

    strictly outperforming the fundamental lower bound of all fixed rotation-invariant kernels and random feature models.

    Two concrete examples demonstrate this separation:

    • Error function activation and target (σ=σ∗=erf\sigma = \sigma^* = \text{erf}): τ∗=0\tau^* = 0 is achieved exactly at κ∗=3\kappa^* = \sqrt{3}, leading to parametric prediction risk R1(λ)=O(d/n)\mathcal{R}_1(\lambda) = \mathcal{O}(d/n), while fixed kernel models are lower bounded by a constant strictly greater than zero.
    • Hyperbolic tangent activation and target (σ=σ∗=tanh⁡\sigma = \sigma^* = \tanh): τ∗\tau^* is strictly smaller than 110∥P>1f∗∥L22\frac{1}{10} \|P_{>1} f^*\|_{L_2}^2, guaranteeing R1(λ)<∥P>1f∗∥L22\mathcal{R}_1(\lambda) < \|P_{>1} f^*\|_{L_2}^2.
  9. Knowl 9 — Learning Rate Scaling Regimes and Feature Representation Shifts

    definition

    For a two-layer neural network f(x)=1Na⊤σ(W⊤x)f(x) = \frac{1}{\sqrt{N}} a^\top \sigma(W^\top x) trained on empirical MSE loss in the proportional regime n≍d≍Nn \asymp d \asymp N, the magnitude of the weight update W1−W0=ηNG0W_1 - W_0 = \eta \sqrt{N} G_0 is classified by the learning rate scaling η=Θ(Nα)\eta = \Theta(N^\alpha):

    1. Sub-learning regime (α<0\alpha < 0, η=od(1)\eta = o_d(1)): The weight update satisfies ∥W1−W0∥F=od,P(1)\|W_1 - W_0\|_F = o_{d,P}(1). The representation remains effectively unperturbed, and test performance matches the initial random feature model.
    2. Small step-size regime (α=0\alpha = 0, η=Θ(1)\eta = \Theta(1)): The update satisfies ∥W1−W0∥2≍∥W0∥2=Θd,P(1)\|W_1 - W_0\|_2 \asymp \|W_0\|_2 = \Theta_{d,P}(1), while ∥W1−W0∥F=Θd,P(1)≪∥W0∥F=Θd,P(d)\|W_1 - W_0\|_F = \Theta_{d,P}(1) \ll \|W_0\|_F = \Theta_{d,P}(\sqrt{d}). Neuron preactivations change by od,P(1)o_{d,P}(1), preserving pairwise near-orthogonality among weight vectors and ensuring the Gaussian Equivalence Theorem holds (linear regime).
    3. Large step-size / Maximal update regime (α=1/2\alpha = 1/2, η=Θ(N)\eta = \Theta(\sqrt{N})): The Frobenius norm of the update matches initialization: ∥W1−W0∥F≍∥W0∥F=Θd,P(d)\|W_1 - W_0\|_F \asymp \|W_0\|_F = \Theta_{d,P}(\sqrt{d}). The preactivation change satisfies ∣σ(W1⊤x)i−σ(W0⊤x)i∣=Θ~d,P(1)|\sigma(W_1^\top x)_i - \sigma(W_0^\top x)_i| = \tilde{\Theta}_{d,P}(1), breaking near-orthogonality and Gaussian equivalence, and allowing the network to fit nonlinear target components.
    4. Instability regime (α>1/2\alpha > 1/2): The gradient step dominates initialization, causing preactivations ⟨x,wi⟩\langle x, w_i \rangle to diverge as N→∞N \to \infty.
  10. Knowl 10 — Theoretical and Practical Limitations of One-Step Feature Learning Analysis

    limitation

    The high-dimensional analysis of one-step feature learning is subject to three primary theoretical boundaries:

    1. Data split requirement: The second-layer kernel ridge regression estimator a^λ\hat{a}_\lambda must be evaluated on an independent dataset (X~,y~)(\tilde{X}, \tilde{y}) drawn from the same distribution, rather than the dataset (X,y)(X, y) used to update W1W_1. While this rigorously models transfer learning and feature pretraining, it does not analyze single-dataset end-to-end training.
    2. Single gradient step scope: The exact risk asymptotics via the Gaussian Equivalence Theorem and operator-valued free probability apply exclusively to a single gradient descent step. Subsequent gradient steps introduce higher-order dependencies across neurons, causing the Gaussian Equivalence Theorem to fail.
    3. Dependence on target-student nonlinearity match: Outperforming the kernel lower bound ∥P>1f∗∥L22\|P_{>1} f^*\|_{L_2}^2 in a single large step requires the target link σ∗\sigma^* and activation σ\sigma to have a sufficiently small single-neuron Gaussian approximation error τ∗\tau^*. For general activations or complex targets where τ∗\tau^* is large, one feature learning step is insufficient to beat fixed kernels, necessitating multi-step optimization.

Coverage note — Deliberately omitted the technical proof derivations, the explicit operator-valued free probability resolvent formulas from Appendix C, and the multi-step empirical simulations from Appendix A, as they support the stated theoretical and empirical results without introducing independent standalone concepts.

References

  1. 1.Emmanuel Abbe, Enric Boix Adsera, Matthew Brennan, Guy Bresler, and Dheeraj Nagaraj, The staircase property: How hierarchical structure can guide deep learning, Advances in Neural Information Processing Systems 34 (2021).
  2. 2.Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz, The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks, arXiv preprint arXiv:2202.08658 (2022).
  3. 3.Radoslaw Adamczak, A note on the hanson-wright inequality for random vectors with dependencies, Electronic Communications in Probability 20 (2015), 1–13.
  4. 4.Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang, On exact computation with an infinitely wide neural net, Advances in Neural Information Processing Systems 32 (2019).
  5. 5.Ben Adlam and Jeffrey Pennington, The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization, International Conference on Machine Learning, PMLR, 2020, pp. 74–84.
  6. 6.Zeyuan Allen-Zhu and Yuanzhi Li, What can resnet learn efficiently, going beyond kernels?, Advances in Neural Information Processing Systems 32 (2019).
  7. 7.Zeyuan Allen-Zhu and Yuanzhi Li, Backward feature correction: How deep learning performs deep learning, arXiv preprint arXiv:2001.04413 (2020).
  8. 8.Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang, Learning and generalization in overparameterized neural networks, going beyond two layers, Advances in neural information processing systems 32 (2019).
  9. 9.Francis Bach, Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research 18 (2017), no. 1, 629–681.
  10. 10.Francis Bach, Learning theory from first principles, MIT Press, 2023.
  11. 11.Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang, High-dimensional asymptotics of feature learning in the early phase of neural network training, In Preparation (2022).
  12. 12.Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off, Proceedings of the National Academy of Sciences 116 (2019), no. 32, 15849–15854.
  13. 13.Yu Bai and Jason D. Lee, Beyond linearization: On quadratic and higher-order approximation of wide neural networks, International Conference on Learning Representations, 2020.
  14. 14.Antoine Bodin and Nicolas Macris, Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature model, Advances in Neural Information Processing Systems 34 (2021).
  15. 15.Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin, Deep learning: a statistical viewpoint, Acta numerica 30 (2021), 87–201.
  16. 16.Lucas Benigni and Sandrine Péché, Eigenvalue distribution of some nonlinear models of random matrices, Electronic Journal of Probability 26 (2021), 1–37.
  17. 17.Lucas Benigni and Sandrine Péché, Largest eigenvalues of the conjugate kernel of single-layered neural networks, arXiv preprint arXiv:2201.04753 (2022).
  18. 18.Zhi-Dong Bai and Jack W Silverstein, No eigenvalues outside the support of the limiting spectral distribution of large-dimensional sample covariance matrices, The Annals of Probability 26 (1998), no. 1, 316–345.
  19. 19.Zhidong Bai and Jack W Silverstein, Spectral analysis of large dimensional random matrices, vol. 20, Springer, 2010.
  20. 20.Lenaic Chizat and Francis Bach, On the global convergence of gradient descent for over-parameterized models using optimal transport, Advances in neural information processing systems, 2018, pp. 3036–3046.
  21. 21.Lenaic Chizat and Francis Bach, Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss, Conference on Learning Theory, PMLR, 2020, pp. 1305–1338.
  22. 22.Lénaïc Chizat, Mean-field langevin dynamics: Exponential convergence and annealing, arXiv preprint arXiv:2202.01009 (2022).
  23. 23.Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar, Gradient descent on neural networks typically occurs at the edge of stability, International Conference on Learning Representations, 2021.
  24. 24.Niladri S Chatterji, Philip M Long, and Peter L Bartlett, When does gradient descent with logistic loss find interpolating two-layer networks?, Journal of Machine Learning Research 22 (2021), no. 159, 1–48.
  25. 25.Lenaic Chizat, Edouard Oyallon, and Francis Bach, On lazy training in differentiable programming, Advances in Neural Information Processing Systems 32 (2019).
  26. 26.Xiuyuan Cheng and Amit Singer, The spectrum of random inner-product kernel matrices, Random Matrices: Theory and Applications 2 (2013), no. 04, 1350010.
  27. 27.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, Bert: Pretraining of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018).
  28. 28.Ethan Dyer and Guy Gur-Ari, Asymptotics of wide networks from feynman diagrams, International Conference on Learning Representations, 2020.
  29. 29.Oussama Dhifallah and Yue M Lu, A precise performance analysis of learning with random features, arXiv preprint arXiv:2008.11904 (2020).
  30. 30.Amit Daniely and Eran Malach, Learning parities with neural networks, Advances in Neural Information Processing Systems 33 (2020), 20356–20365.
  31. 31.Yen Do and Van Vu, The spectrum of random kernel matrices: universality results for rough and varying kernels, Random Matrices: Theory and Applications 2 (2013), no. 03, 1350005.
  32. 32.Edgar Dobriban and Stefan Wager, High-dimensional asymptotics of prediction: Ridge regression and classification, The Annals of Statistics 46 (2018), no. 1, 247–279.
  33. 33.Konstantin Donhauser, Mingqi Wu, and Fanny Yang, How rotational invariance of common kernels prevents generalization in high dimensions, International Conference on Machine Learning, PMLR, 2021, pp. 2804–2814.
  34. 34.Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient descent provably optimizes over-parameterized neural networks, International Conference on Learning Representations, 2019.
  35. 35.Noureddine El Karoui, The spectrum of kernel random matrices, The Annals of Statistics 38 (2010), no. 1, 1–50.
  36. 36.Noureddine El Karoui, On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators, Probability Theory and Related Fields 170 (2018), no. 1, 95–175.
  37. 37.Spencer Frei, Niladri S Chatterji, and Peter L Bartlett, Random feature amplification: Feature learning and generalization in neural networks, arXiv preprint arXiv:2202.07626 (2022).
  38. 38.Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M Roy, and Surya Ganguli, Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel, Advances in Neural Information Processing Systems 33 (2020), 5850–5861.
  39. 39.Zhou Fan and Andrea Montanari, The spectral norm of random inner-product kernel matrices, Probability Theory and Related Fields 173 (2019), no. 1-2, 27–85.
  40. 40.Reza Rashidi Far, Tamer Oraby, Wlodzimierz Bryc, and Roland Speicher, Spectra of large block matrices, arXiv preprint cs/0610045 (2006).
  41. 41.Zhou Fan and Zhichao Wang, Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks, Advances in neural information processing systems 33 (2020), 7710–7721.
  42. 42.Aditya Sharad Golatkar, Alessandro Achille, and Stefano Soatto, Time matters in regularizing deep networks: Weight decay and data augmentation affect early learning dynamics, matter little near convergence, Advances in Neural Information Processing Systems 32 (2019).
  43. 43.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik, Rich feature hierarchies for accurate object detection and semantic segmentation, Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587.
  44. 44.Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, Generalisation error in learning with random features and the hidden manifold model, International Conference on Machine Learning, PMLR, 2020, pp. 3452–3462.
  45. 45.Sebastian Goldt, Bruno Loureiro, Galen Reeves, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, The gaussian equivalence of generative models for learning with shallow neural networks, Proceedings of Machine Learning Research vol 145 (2021), 1–46.
  46. 46.Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová, Modeling the influence of data structure on learning in neural networks: The hidden manifold model, Physical Review X 10 (2020), no. 4, 041044.
  47. 47.Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Limitations of lazy training of two-layers neural network, Advances in Neural Information Processing Systems 32 (2019).
  48. 48.Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, When do neural networks outperform kernel methods?, Advances in Neural Information Processing Systems 33 (2020), 14820–14830.
  49. 49.Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Linearized two-layers neural networks in high dimension, The Annals of Statistics 49 (2021), no. 2, 1029–1054.
  50. 50.Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart, Disentangling feature and lazy training in deep neural networks, Journal of Statistical Mechanics: Theory and Experiment 2020 (2020), no. 11, 113301.
  51. 51.Karl Hajjar, Lénaïc Chizat, and Christophe Giraud, Training integrable parameterizations of deep neural networks in the infinite-width limit, arXiv preprint arXiv:2110.15596 (2021).
  52. 52.J William Helton, Reza Rashidi Far, and Roland Speicher, Operator-valued semicircular elements: solving a quadratic matrix equation with positivity constraints, International Mathematics Research Notices 2007 (2007), no. 9, rnm086–rnm086.
  53. 53.Hong Hu and Yue M Lu, Universality laws for high-dimensional learning with random features, arXiv preprint arXiv:2009.07669 (2020).
  54. 54.J William Helton, Tobias Mai, and Roland Speicher, Applications of realizations (aka linearizations) to free probability, Journal of Functional Analysis 274 (2018), no. 1, 1–79.
  55. 55.Jiaoyang Huang and Horng-Tzer Yau, Dynamics of deep neural networks and neural tangent hierarchy, International conference on machine learning, PMLR, 2020, pp. 4542–4551.
  56. 56.Masaaki Imaizumi and Kenji Fukumizu, Deep neural networks learn non-smooth functions effectively, The 22nd international conference on artificial intelligence and statistics, PMLR, 2019, pp. 869–878.
  57. 57.Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks, Advances in neural information processing systems, 2018, pp. 8571–8580.
  58. 58.Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras, The break-even point on optimization trajectories of deep neural networks, International Conference on Learning Representations, 2020.
  59. 59.Ziwei Ji and Matus Telgarsky, Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks, International Conference on Learning Representations, 2020.
  60. 60.Stefani Karp, Ezra Winston, Yuanzhi Li, and Aarti Singh, Local signal adaptivity: Provable feature learning in neural networks beyond kernels, Advances in Neural Information Processing Systems 34 (2021).
  61. 61.Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari, The large learning rate phase of deep learning: the catapult mechanism, arXiv preprint arXiv:2003.02218 (2020).
  62. 62.Zhenyu Liao, Romain Couillet, and Michael W Mahoney, A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent, Advances in Neural Information Processing Systems 33 (2020), 13939–13950.
  63. 63.Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborová, Learning curves of generic features maps for realistic datasets with a teacher-student model, Advances in Neural Information Processing Systems 34 (2021).
  64. 64.Cosme Louart, Zhenyu Liao, and Romain Couillet, A random matrix approach to neural networks, The Annals of Applied Probability 28 (2018), no. 2, 1190–1248.
  65. 65.Guillaume Leclerc and Aleksander Madry, The two regimes of deep network training, arXiv preprint arXiv:2002.10376 (2020).
  66. 66.Yuanzhi Li, Tengyu Ma, and Hongyang R Zhang, Learning over-parametrized two-layer neural networks beyond ntk, Conference on learning theory, PMLR, 2020, pp. 2613–2682.
  67. 67.Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize, The Annals of Statistics 48 (2020), no. 3, 1329–1347.
  68. 68.Yuanzhi Li, Colin Wei, and Tengyu Ma, Towards explaining the regularization effect of initial large learning rate in training neural networks, Advances in Neural Information Processing Systems, 2019, pp. 11674–11685.
  69. 69.Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro, Quantifying the benefit of using differentiable learning over tangent kernels, International Conference on Machine Learning, PMLR, 2021, pp. 7379–7389.
  70. 70.Song Mei and Andrea Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve, Communications on Pure and Applied Mathematics 75 (2022), no. 4, 667–766.
  71. 71.Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration, Applied and Computational Harmonic Analysis (2021).
  72. 72.Song Mei, Andrea Montanari, and Phan-Minh Nguyen, A mean field view of the landscape of two-layer neural networks, Proceedings of the National Academy of Sciences 115 (2018), no. 33, E7665–E7671.
  73. 73.James A Mingo and Roland Speicher, Free probability and random matrices, vol. 35, Springer, 2017.
  74. 74.Andrea Montanari and Basil N Saeed, Universality of empirical risk minimization, Conference on Learning Theory, PMLR, 2022, pp. 4310–4312.
  75. 75.Andrea Montanari and Yiqiao Zhong, The interpolation phase transition in neural networks: Memorization and generalization under lazy training, arXiv preprint arXiv:2007.12826v1 (2020).
  76. 76.Radford M Neal, Bayesian learning for neural networks, vol. 118, Springer Science & Business Media, 1995.
  77. 77.Phan-Minh Nguyen, Analysis of feature learning in weight-tied autoencoders via the mean field lens, arXiv preprint arXiv:2102.08373 (2021).
  78. 78.Atsushi Nitanda and Taiji Suzuki, Stochastic particle gradient descent for infinite ensembles, arXiv preprint arXiv:1712.05438 (2017).
  79. 79.Atsushi Nitanda, Denny Wu, and Taiji Suzuki, Convex analysis of the mean field langevin dynamics, arXiv preprint arXiv:2201.10469 (2022).
  80. 80.S Péché, A note on the pennington-worah distribution, Electronic Communications in Probability 24 (2019), 1–7.
  81. 81.Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion, Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity, Advances in Neural Information Processing Systems 34 (2021).
  82. 82.Jeffrey Pennington and Pratik Worah, Nonlinear random matrix theory for deep learning, Advances in Neural Information Processing Systems, 2017, pp. 2637–2646.
  83. 83.Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová, Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed, International Conference on Machine Learning, PMLR, 2021, pp. 8936–8947.
  84. 84.Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines, Advances in neural information processing systems, 2008, pp. 1177–1184.
  85. 85.Taiji Suzuki and Shunta Akiyama, Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods, arXiv preprint arXiv:2012.03224 (2020).
  86. 86.Johannes Schmidt-Hieber, Nonparametric regression using deep neural networks with relu activation function, The Annals of Statistics 48 (2020), no. 4, 1875–1897.
  87. 87.Taiji Suzuki, Adaptivity of deep relu network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality, arXiv preprint arXiv:1810.08033 (2018).
  88. 88.Nilesh Tripuraneni, Ben Adlam, and Jeffrey Pennington, Covariate shift in high-dimensional random feature regression, arXiv preprint arXiv:2111.08234 (2021).
  89. 89.Roman Vershynin, High-dimensional probability: An introduction with applications in data science, vol. 47, Cambridge university press, 2018.
  90. 90.Rodrigo Veiga, Ludovic Stephan, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová, Phase diagram of stochastic gradient descent in high-dimensional two-layer neural networks, arXiv preprint arXiv:2202.00293 (2022).
  91. 91.Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro, Kernel and rich regimes in overparametrized models, Conference on Learning Theory, PMLR, 2020, pp. 3635–3673.
  92. 92.Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma, Regularization matters: Generalization and optimization of neural nets vs their induced kernel, Advances in Neural Information Processing Systems, 2019, pp. 9712–9724.
  93. 93.Denny Wu and Ji Xu, On the optimal weighted ℓ2 regularization in overparameterized linear regression, Advances in Neural Information Processing Systems 33 (2020), 10112–10123.
  94. 94.Zhichao Wang and Yizhe Zhu, Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks, arXiv preprint arXiv:2109.09304 (2021).
  95. 95.Greg Yang, Tensor programs iii: Neural matrix laws, arXiv preprint arXiv:2009.10685 (2020).
  96. 96.Greg Yang and Edward J Hu, Feature learning in infinite-width neural networks, arXiv preprint arXiv:2011.14522 (2020).
  97. 97.Gilad Yehudai and Ohad Shamir, On the power and limitations of random features for understanding neural networks, Advances in Neural Information Processing Systems 32 (2019).

Citation

MLA
Ba, J., et al. “High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 37932–46, https://proceedings.neurips.cc/paper_files/paper/2022/file/f7e7fabd73b3df96c54a320862afcb78-Paper-Conference.pdf.
APA
Ba, J., Erdogdu, M., Suzuki, T., Wang, Z., Wu, D., & Yang, G. (2022). High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation. Advances in Neural Information Processing Systems, 35, 37932–37946. https://proceedings.neurips.cc/paper_files/paper/2022/file/f7e7fabd73b3df96c54a320862afcb78-Paper-Conference.pdf
Chicago
Ba, J., M. Erdogdu, T. Suzuki, Z. Wang, D. Wu, and G. Yang. 2022. “High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation”. Advances in Neural Information Processing Systems 35: 37932–46. https://proceedings.neurips.cc/paper_files/paper/2022/file/f7e7fabd73b3df96c54a320862afcb78-Paper-Conference.pdf.
Harvard
Ba, J. et al. (2022) “High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 37932–37946. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/f7e7fabd73b3df96c54a320862afcb78-Paper-Conference.pdf.
Vancouver
1. Ba J, Erdogdu M, Suzuki T, Wang Z, Wu D, Yang G (2022) High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 37932–37946

BibTeX

@inproceedings{ba2022high,
  title = {High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation},
  author = {Ba, Jimmy and Erdogdu, Murat and Suzuki, Taiji and Wang, Zhichao and Wu, Denny and Yang, Greg},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {37932-37946},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/f7e7fabd73b3df96c54a320862afcb78-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission