Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks

Blake BordelonCengiz Pehlevan

article2022NeurIPS123 citations

Develops a self-consistent dynamical field theory and an efficient sampling algorithm to track kernel evolution and feature learning in infinite-width neural networks beyond the static neural tangent kernel regime.

Listen

Deep learning models have achieved widespread success across modern technology, yet theoretical understanding of how these networks learn features from data during training remains incomplete. Existing foundational theories often assume an infinite-width limit where internal data representations remain static throughout optimization—known as the lazy training regime. However, practical deep neural networks succeed precisely because they adapt their internal representations to the data. Bridging this gap between theory and practical feature adaptation is essential for systematically understanding, tuning, and scaling modern artificial intelligence systems.

The main objective of the article is to establish and validate an exact dynamical field theory that characterizes how internal feature representations and gradient signals evolve over time in wide neural networks. It demonstrates how these dynamics can be captured through a reduced set of deterministic kernel parameters that smoothly interpolate between static and adaptive learning regimes.

To achieve this, the article adapts mathematical tools from statistical physics known as dynamical mean field theory. The authors formulate the gradient-based training trajectory as a functional path integral in the infinite-width limit, deriving a set of self-consistent equations that govern internal network features and training dynamics. For deep linear architectures, these equations simplify into closed algebraic systems, whereas for nonlinear networks, the authors introduce a polynomial-time alternating sampling algorithm to compute the solutions. The resulting theoretical predictions were validated against numerical simulations of finite-width fully connected networks and convolutional neural networks trained on image classification tasks.

The investigation yields several key findings. First, the article demonstrates that network activation and gradient distributions over training can be exactly summarized by deterministic kernel order parameters, recovering previous stochastic descriptions from tensor program frameworks while providing a more flexible formulation. Second, deep linear networks yield exact algebraic matrix solutions showing that deeper architectures experience substantially larger shifts in internal representations during training. Third, existing approximation methods—such as the static neural tangent kernel, gradient independence assumptions, and leading-order perturbation theory—break down significantly as the strength of feature learning and network depth increase, whereas the dynamical mean field theory remains accurate. Finally, experiments on image classification confirm that when the feature learning strength parameter is held constant, loss and representation dynamics remain invariant across different network widths, matching the theory once width reaches moderately large sizes (e.g., around 500 hidden units).

These findings have practical implications for designing and scaling deep learning architectures. By showing that representation dynamics are invariant under specific width and learning parameter scalings, the theory supports principled hyperparameter transfer from small pilot models to large-scale production architectures without costly trial-and-error retraining. Furthermore, the framework explains why regularized networks can preserve expressive, non-trivial predictive solutions rather than decaying to zero when feature adaptation is sufficiently strong.

Organizations developing large-scale neural network architectures should consider utilizing these scaling relationships to guide model expansion and hyperparameter tuning across widths. While the exact numerical solution currently incurs a computational complexity that scales cubically with time steps and sample size, practitioners can use the qualitative scaling laws to inform model depth and initialization schedules. For decision-makers evaluating theoretical predictive tooling, additional work should focus on exploring accelerated projected gradient methods or data-averaged approximations before deploying these theoretical solvers across massive-scale production datasets.

Confidence in these theoretical results is high within the defined boundaries of wide networks and gradient-flow training. Readers should exercise caution, however, when extrapolating to small sample-to-width ratios or extreme training horizons where the infinite-width assumptions and physics-based heuristic saddle-point derivations may require further finite-size corrections.

arXiv: 2205.09653
Cover for Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks

Abstract

We analyze feature learning in infinite-width neural networks trained with gradient flow through a self-consistent dynamical field theory. We construct a collection of deterministic dynamical order parameters which are inner-product kernels for hidden unit activations and gradients in each layer at pairs of time points, providing a reduced description of network activity through training. These kernel order parameters collectively define the hidden layer activation distribution, the evolution of the neural tangent kernel, and consequently output predictions. We show that the field theory derivation recovers the recursive stochastic process of infinite-width feature learning networks obtained from Yang & Hu with Tensor Programs [1]. For deep linear networks, these kernels satisfy a set of algebraic matrix equations. For nonlinear networks, we provide an alternating sampling procedure to self-consistently solve for the kernel order parameters. We provide comparisons of the self-consistent solution to various approximation schemes including the static NTK approximation, gradient independence assumption, and leading order perturbation theory, showing that each of these approximations can break down in regimes where general self-consistent solutions still provide an accurate description. Lastly, we provide experiments in more realistic settings which demonstrate that the loss and kernel dynamics of CNNs at fixed feature learning strength is preserved across different widths on a CIFAR classification task.

Table of Contents

  • 1 Introduction
  • 1.1 Related Works
  • 2 Problem Setup and Definitions
  • 3 Self-consistent DMFT
  • 3.1 Path Integral Construction
  • 3.2 Deriving the DMFT Equations from the Path Integral Saddle Point
  • 4 Solving the Self-Consistent DMFT
  • 4.1 Deep Linear Networks: Closed Form Self-Consistent Equations
  • 4.2 Feature Learning with L2 Regularization
  • 5 Approximation Schemes
  • 5.1 Gradient Independence Ansatz
  • 5.2 Perturbation theory in 0 at infinite-width
  • 6 Feature Learning Dynamics is Preserved at Fixed 0
  • 7 Discussion
  • Acknowledgments and Disclosure of Funding
  • References
  • Checklist

Knowls

  1. Knowl 1 — Self-Consistent Dynamical Mean Field Theory for Infinite-Width Networks

    theoretical result

    In the infinite-width limit (N→∞N \to \infty with feature learning parameter γ0≡γ/N=ON(1)\gamma_0 \equiv \gamma/\sqrt{N} = \mathcal{O}_N(1) and training set size P=ON(1)P = \mathcal{O}_N(1)), the preactivations hiℓ(t)h_i^\ell(t) and pre-gradients ziℓ(t)z_i^\ell(t) of hidden units i∈{1,…,N}i \in \{1, \dots, N\} across layers ℓ∈{1,…,L}\ell \in \{1, \dots, L\} become independent and identically distributed draws from an effective single-site stochastic process. The single-site fields satisfy the self-consistent causal integro-differential equations:

    hμℓ(t)=uμℓ(t)+γ0∫0tds∑α=1P[Aμαℓ−1(t,s)+Δα(s)Φμαℓ−1(t,s)]zαℓ(s)ϕ˙(hαℓ(s)),h_\mu^\ell(t) = u_\mu^\ell(t) + \gamma_0 \int_0^t ds \sum_{\alpha=1}^P \left[ A_{\mu\alpha}^{\ell-1}(t, s) + \Delta_\alpha(s) \Phi_{\mu\alpha}^{\ell-1}(t, s) \right] z_\alpha^\ell(s) \dot{\phi}(h_\alpha^\ell(s)),

    zμℓ(t)=rμℓ(t)+γ0∫0tds∑α=1P[Bμαℓ(t,s)+Δα(s)Gμαℓ+1(t,s)]ϕ(hαℓ(s)),z_\mu^\ell(t) = r_\mu^\ell(t) + \gamma_0 \int_0^t ds \sum_{\alpha=1}^P \left[ B_{\mu\alpha}^\ell(t, s) + \Delta_\alpha(s) G_{\mu\alpha}^{\ell+1}(t, s) \right] \phi(h_\alpha^\ell(s)),

    where Δμ(t)=−∂L∂fμ∣fμ(t)\Delta_\mu(t) = -\frac{\partial \mathcal{L}}{\partial f_\mu}\big|_{f_\mu(t)} is the prediction error under loss L\mathcal{L}, ϕ\phi is the activation function, and the driving fields {uμℓ(t)}\{u_\mu^\ell(t)\} and {rμℓ(t)}\{r_\mu^\ell(t)\} are zero-mean Gaussian processes with covariances:

    ⟨uμℓ(t)uαℓ(s)⟩=Φμαℓ−1(t,s),⟨rμℓ(t)rαℓ(s)⟩=Gμαℓ+1(t,s),\langle u_\mu^\ell(t) u_\alpha^\ell(s) \rangle = \Phi_{\mu\alpha}^{\ell-1}(t, s), \quad \langle r_\mu^\ell(t) r_\alpha^\ell(s) \rangle = G_{\mu\alpha}^{\ell+1}(t, s),

    with boundary conditions Φμα0(t,s)=Kμαx=1Dxμ⋅xα\Phi_{\mu\alpha}^0(t, s) = K_{\mu\alpha}^x = \frac{1}{D} x_\mu \cdot x_\alpha, GμαL+1(t,s)=1G_{\mu\alpha}^{L+1}(t, s) = 1, A0(t,s)=0A^0(t, s) = 0, and BL(t,s)=0B^L(t, s) = 0.

    The deterministic order parameter kernels are determined self-consistently by averaging over the single-site process distribution:

    Φμαℓ(t,s)=⟨ϕ(hμℓ(t))ϕ(hαℓ(s))⟩,Gμαℓ(t,s)=⟨gμℓ(t)gαℓ(s)⟩,\Phi_{\mu\alpha}^\ell(t, s) = \langle \phi(h_\mu^\ell(t)) \phi(h_\alpha^\ell(s)) \rangle, \quad G_{\mu\alpha}^\ell(t, s) = \langle g_\mu^\ell(t) g_\alpha^\ell(s) \rangle,

    Aμαℓ(t,s)=γ0−1⟨δϕ(hμℓ(t))δrαℓ(s)⟩,Bμαℓ(t,s)=γ0−1⟨δgμℓ+1(t)δuαℓ+1(s)⟩,A_{\mu\alpha}^\ell(t, s) = \gamma_0^{-1} \left\langle \frac{\delta \phi(h_\mu^\ell(t))}{\delta r_\alpha^\ell(s)} \right\rangle, \quad B_{\mu\alpha}^\ell(t, s) = \gamma_0^{-1} \left\langle \frac{\delta g_\mu^{\ell+1}(t)}{\delta u_\alpha^{\ell+1}(s)} \right\rangle,

    where gμℓ(t)=ϕ˙(hμℓ(t))zμℓ(t)g_\mu^\ell(t) = \dot{\phi}(h_\mu^\ell(t)) z_\mu^\ell(t). The Neural Tangent Kernel (NTK) and network outputs fμ(t)f_\mu(t) evolve according to:

    KμαNTK(t,s)=∑ℓ=0LGμαℓ+1(t,s)Φμαℓ(t,s),ddtfμ(t)=∑α=1PKμαNTK(t,t)Δα(t).K_{\mu\alpha}^{\text{NTK}}(t, s) = \sum_{\ell=0}^L G_{\mu\alpha}^{\ell+1}(t, s) \Phi_{\mu\alpha}^\ell(t, s), \quad \frac{d}{dt} f_\mu(t) = \sum_{\alpha=1}^P K_{\mu\alpha}^{\text{NTK}}(t, t) \Delta_\alpha(t).

  2. Knowl 2 — Feature Learning Scaling Parameterization for Wide Neural Networks

    model/method

    Consider a fully connected neural network with LL hidden layers processing input vectors xμ∈RDx_\mu \in \mathbb{R}^D for sample indices μ∈{1,…,P}\mu \in \{1, \dots, P\}. The preactivation vectors hμℓ∈RNh_\mu^\ell \in \mathbb{R}^N for layers ℓ∈{1,…,L}\ell \in \{1, \dots, L\} and the network output fμ∈Rf_\mu \in \mathbb{R} are defined by:

    fμ=1γNwL⋅ϕ(hμL),hμℓ+1=1NWℓϕ(hμℓ),hμ1=1DW0xμ,f_\mu = \frac{1}{\gamma \sqrt{N}} w^L \cdot \phi(h_\mu^L), \quad h_\mu^{\ell+1} = \frac{1}{\sqrt{N}} W^\ell \phi(h_\mu^\ell), \quad h_\mu^1 = \frac{1}{\sqrt{D}} W^0 x_\mu,

    where ϕ:R→R\phi: \mathbb{R} \to \mathbb{R} is a twice-differentiable activation function applied elementwise, θ=Vec{W0,W1,…,WL−1,wL}\theta = \text{Vec}\{W^0, W^1, \dots, W^{L-1}, w^L\} are trainable weights initialized as independent standard Gaussian variables Wijℓ∼N(0,1)W_{ij}^\ell \sim \mathcal{N}(0, 1), and γ>0\gamma > 0 parameterizes training dynamics. Weights evolve under gradient flow:

    dθdt=−γ2∇θL,\frac{d\theta}{dt} = -\gamma^2 \nabla_\theta \mathcal{L},

    where L\mathcal{L} is the empirical loss function. Under the infinite-width limit where N,γ→∞N, \gamma \to \infty with fixed feature learning strength:

    γ0≡γN=ON(1),\gamma_0 \equiv \frac{\gamma}{\sqrt{N}} = \mathcal{O}_N(1),

    the parameterization allows non-trivial feature adaptation throughout training. When γ0→0\gamma_0 \to 0, the parameterization recovers the static Neural Tangent Kernel (NTK) lazy training limit. Backpropagation signals are characterized by gμℓ≡γN∂fμ∂hμℓ=ϕ˙(hμℓ)⊙zμℓg_\mu^\ell \equiv \gamma \sqrt{N} \frac{\partial f_\mu}{\partial h_\mu^\ell} = \dot{\phi}(h_\mu^\ell) \odot z_\mu^\ell, where zμℓ≡1N(Wℓ)⊤gμℓ+1∈RNz_\mu^\ell \equiv \frac{1}{\sqrt{N}} (W^\ell)^\top g_\mu^{\ell+1} \in \mathbb{R}^N is the pre-gradient vector.

  3. Knowl 3 — Closed-Form Algebraic DMFT Equations for Deep Linear Networks

    theoretical result

    For deep linear neural networks where ϕ(h)=h\phi(h) = h, the dynamical mean field theory saddle-point equations reduce to closed algebraic matrix equations without requiring Monte Carlo sampling. Vectorizing the hidden preactivation trajectories across training samples μ∈{1,…,P}\mu \in \{1, \dots, P\} and continuous time t∈R+t \in \mathbb{R}_+ as hℓ=Vec{hμℓ(t)}h^\ell = \text{Vec}\{h_\mu^\ell(t)\}, and defining the kernel operators Hℓ=Mat{⟨hμℓ(t)hαℓ(s)⟩}H^\ell = \text{Mat}\{\langle h_\mu^\ell(t) h_\alpha^\ell(s) \rangle\} and Gℓ=Mat{⟨gμℓ(t)gαℓ(s)⟩}G^\ell = \text{Mat}\{\langle g_\mu^\ell(t) g_\alpha^\ell(s) \rangle\} acting under the continuous inner product a⋅b=∫0∞dt∑μ=1Paμ(t)bμ(t)a \cdot b = \int_0^\infty dt \sum_{\mu=1}^P a_\mu(t) b_\mu(t):

    The stochastic fields are linear functionals of independent Gaussian processes uℓ∼GP(0,Hℓ−1)u^\ell \sim \mathcal{GP}(0, H^{\ell-1}) and rℓ∼GP(0,Gℓ+1)r^\ell \sim \mathcal{GP}(0, G^{\ell+1}):

    (I−γ02CℓDℓ)hℓ=uℓ+γ0Cℓrℓ,(I−γ02DℓCℓ)gℓ=rℓ+γ0Dℓuℓ,(I - \gamma_0^2 C^\ell D^\ell) h^\ell = u^\ell + \gamma_0 C^\ell r^\ell, \quad (I - \gamma_0^2 D^\ell C^\ell) g^\ell = r^\ell + \gamma_0 D^\ell u^\ell,

    where CℓC^\ell and DℓD^\ell are causal integral operators determined by {Aℓ−1,Hℓ−1}\{A^{\ell-1}, H^{\ell-1}\} and {Bℓ,Gℓ+1}\{B^\ell, G^{\ell+1}\} respectively. Consequently, the feature and gradient kernels satisfy the self-consistent algebraic matrix system:

    Hℓ=(I−γ02CℓDℓ)−1[Hℓ−1+γ02CℓGℓ+1(Cℓ)⊤][(I−γ02CℓDℓ)−1]⊤,H^\ell = (I - \gamma_0^2 C^\ell D^\ell)^{-1} \left[ H^{\ell-1} + \gamma_0^2 C^\ell G^{\ell+1} (C^\ell)^\top \right] \left[ (I - \gamma_0^2 C^\ell D^\ell)^{-1} \right]^\top,

    Gℓ=(I−γ02DℓCℓ)−1[Gℓ+1+γ02DℓHℓ−1(Dℓ)⊤][(I−γ02DℓCℓ)−1]⊤.G^\ell = (I - \gamma_0^2 D^\ell C^\ell)^{-1} \left[ G^{\ell+1} + \gamma_0^2 D^\ell H^{\ell-1} (D^\ell)^\top \right] \left[ (I - \gamma_0^2 D^\ell C^\ell)^{-1} \right]^\top.

  4. Knowl 4 — Error and Kernel Dynamics of Two-Layer Linear Networks with Whitened Data

    theoretical result

    In a two-layer (L=1L=1) linear network trained on whitened input data (Kx=I∈RP×PK^x = I \in \mathbb{R}^{P \times P}) with target label vector y∈RPy \in \mathbb{R}^P under mean squared error loss, the DMFT equations yield an exact differential equation for the norm of the prediction error Δ(t)=∥Δ(t)∥2\Delta(t) = \|\Delta(t)\|_2, where Δμ(t)=yμ−fμ(t)\Delta_\mu(t) = y_\mu - f_\mu(t):

    ddtΔ(t)=−21+γ02(∥y∥2−Δ(t))2Δ(t).\frac{d}{dt}\Delta(t) = -2 \sqrt{1 + \gamma_0^2 (\|y\|_2 - \Delta(t))^2} \Delta(t).

    In the lazy limit γ0→0\gamma_0 \to 0, this reproduces the linear convergence rate of the static Neural Tangent Kernel. In the rich limit γ0→∞\gamma_0 \to \infty, it approaches logistic (sigmoidal) dynamics.

    The feature correlation matrix H(t)=⟨h1(t)h1(t)⊤⟩∈RP×PH(t) = \langle h^1(t) h^1(t)^\top \rangle \in \mathbb{R}^{P \times P} evolves exclusively along the rank-one target subspace yy⊤y y^\top:

    Hy(t)≡1∥y∥22y⊤H(t)y=1+γ02(∥y∥2−Δ(t))2.H_y(t) \equiv \frac{1}{\|y\|_2^2} y^\top H(t) y = \sqrt{1 + \gamma_0^2 (\|y\|_2 - \Delta(t))^2}.

    At the end of training (t→∞t \to \infty where Δ(t)→0\Delta(t) \to 0), the feature kernel reaches the fixed point:

    H(∞)=I+1∥y∥22[1+γ02∥y∥22−1]yy⊤,H(\infty) = I + \frac{1}{\|y\|_2^2} \left[ \sqrt{1 + \gamma_0^2 \|y\|_2^2} - 1 \right] y y^\top,

    recovering a rank-one spike aligned with the target labels.

  5. Knowl 5 — Leading-Order Perturbation Theory for NTK Evolution with Depth and Feature Learning Strength

    theoretical result

    Expanding the DMFT observables in powers of the feature learning parameter γ0\gamma_0 around the lazy limit γ0=0\gamma_0 = 0, the odd-order kernel corrections vanish identically, giving asymptotic expansions Φ=Φ0+γ02Φ2+O(γ04)\Phi = \Phi_0 + \gamma_0^2 \Phi_2 + \mathcal{O}(\gamma_0^4) and G=G0+γ02G2+O(γ04)G = G_0 + \gamma_0^2 G_2 + \mathcal{O}(\gamma_0^4).

    For an infinite-width deep linear network of depth LL, the leading-order correction to the Neural Tangent Kernel is given by:

    KμνNTK(t,s)=(L+1)Kμνx+γ02L(L+1)2Kμνx∑α,β=1PKαβx[vαβ(t)+vβα(s)+vα(t)vβ(s)]K_{\mu\nu}^{\text{NTK}}(t, s) = (L+1) K_{\mu\nu}^x + \gamma_0^2 \frac{L(L+1)}{2} K_{\mu\nu}^x \sum_{\alpha,\beta=1}^P K_{\alpha\beta}^x \left[ v_{\alpha\beta}(t) + v_{\beta\alpha}(s) + v_\alpha(t) v_\beta(s) \right]

    +γ02L(L+1)2(∑α,β=1PKμαxKνβx[vαβ(t)+vβα(s)]+∑α,β=1PKμαxKνβxvα(t)vβ(s))+O(γ04),+ \gamma_0^2 \frac{L(L+1)}{2} \left( \sum_{\alpha,\beta=1}^P K_{\mu\alpha}^x K_{\nu\beta}^x \left[ v_{\alpha\beta}(t) + v_{\beta\alpha}(s) \right] + \sum_{\alpha,\beta=1}^P K_{\mu\alpha}^x K_{\nu\beta}^x v_\alpha(t) v_\beta(s) \right) + \mathcal{O}(\gamma_0^4),

    where Kμνx=1Dxμ⋅xνK_{\mu\nu}^x = \frac{1}{D} x_\mu \cdot x_\nu, and the temporal dependence is parameterized by scalar integrated error functions:

    vα(t)=∫0tds Δα(s),vαβ(t)=∫0tds Δα(s)∫0sds′ Δβ(s′).v_\alpha(t) = \int_0^t ds \, \Delta_\alpha(s), \quad v_{\alpha\beta}(t) = \int_0^t ds \, \Delta_\alpha(s) \int_0^s ds' \, \Delta_\beta(s').

    The relative kernel variation scales linearly with depth and quadratically with feature learning strength:

    ∥KNTK(t,s)−KNTK(0,0)∥∥KNTK(0,0)∥=O(γ02L)=O(γ2LN).\frac{\|K^{\text{NTK}}(t, s) - K^{\text{NTK}}(0, 0)\|}{\|K^{\text{NTK}}(0, 0)\|} = \mathcal{O}\left(\gamma_0^2 L\right) = \mathcal{O}\left(\frac{\gamma^2 L}{N}\right).

  6. Knowl 6 — Non-Trivial Fixed Points in Infinite-Width Networks with L2 Regularization

    theoretical result

    Consider a neural network with positive homogeneous activation functions satisfying f(cθ)=cκf(θ)f(c \theta) = c^\kappa f(\theta) for κ>0\kappa > 0 (such as linear, ReLU, or quadratic activations) trained with weight decay under the gradient flow:

    dθdt=−γ2∇θL−λθ,\frac{d\theta}{dt} = -\gamma^2 \nabla_\theta \mathcal{L} - \lambda \theta,

    where λ>0\lambda > 0 is the weight decay parameter. In the infinite-width feature learning limit, the network predictor at convergence t→∞t \to \infty acts as a kernel regressor with the asymptotic Neural Tangent Kernel K∈RP×PK \in \mathbb{R}^{P \times P}:

    lim⁡t→∞f(x,t)=k(x)⊤[K+λκI]−1y,\lim_{t \to \infty} f(x, t) = k(x)^\top [K + \lambda \kappa I]^{-1} y,

    where [k(x)]μ=K(x,xμ)[k(x)]_\mu = K(x, x_\mu), [K]μα=K(xμ,xα)[K]_{\mu\alpha} = K(x_\mu, x_\alpha), and y∈RPy \in \mathbb{R}^P is the target vector.

    In the static NTK limit (γ0→0\gamma_0 \to 0), weight decay drives homogeneous networks to a trivial fixed point with K(x,x′)→0K(x, x') \to 0 and predictor f→0f \to 0. In contrast, for non-zero feature learning strength γ0>0\gamma_0 > 0, feature adaptation prevents collapse to the zero kernel, yielding a non-trivial, non-zero fixed point for KK and ff.

  7. Knowl 7 — Breakdown of the Gradient Independence Ansatz in Feature Learning

    theoretical result

    The gradient independence ansatz assumes that the feedforward weight matrices Wℓ(0)W^\ell(0) and their transposed feedback counterparts Wℓ(0)⊤W^\ell(0)^\top act as statistically independent Gaussian matrices throughout training. In the DMFT framework, this assumption is formally equivalent to setting the response kernels to zero:

    Aμαℓ(t,s)=0,Bμαℓ(t,s)=0.A_{\mu\alpha}^\ell(t, s) = 0, \quad B_{\mu\alpha}^\ell(t, s) = 0.

    Under this condition, the stochastic fields χℓ\chi^\ell and ξℓ\xi^\ell become conditionally independent Gaussian processes given the kernel order parameters {Φℓ,Gℓ}\{\Phi^\ell, G^\ell\}. While the gradient independence ansatz provides an accurate approximation near initialization (t≈0t \approx 0) or in the lazy training regime (γ0≈0\gamma_0 \approx 0) where δhδr,δzδu=O(γ0)\frac{\delta h}{\delta r}, \frac{\delta z}{\delta u} = \mathcal{O}(\gamma_0), it systematically breaks down in rich feature learning regimes at large γ0\gamma_0 and later training times tt, failing to track the empirical loss trajectory and kernel dynamics.

  8. Knowl 8 — Computational Complexity Comparison for Neural Network Feature Dynamics

    data/table

    Computing feature kernel evolution and trained network predictions across PP training samples over a grid of TT time steps for a network of depth LL and width NN has the following asymptotic memory and time complexities across different theoretical frameworks and finite networks:

    Metric Width-NN NN Static NTK Perturbative Full DMFT
    Memory for Kernels O(N2)\mathcal{O}(N^2) O(P2)\mathcal{O}(P^2) O(P4T)\mathcal{O}(P^4 T) O(P2T2)\mathcal{O}(P^2 T^2)
    Time for Kernels O(PN2T)\mathcal{O}(P N^2 T) O(P2)\mathcal{O}(P^2) O(P4T)\mathcal{O}(P^4 T) O(P3T3)\mathcal{O}(P^3 T^3)
    Time for Final Outputs O(PN2T)\mathcal{O}(P N^2 T) O(P3)\mathcal{O}(P^3) O(P4)\mathcal{O}(P^4) O(P3T3)\mathcal{O}(P^3 T^3)

    The static NTK computes an initial P×PP \times P kernel without time dependence. Perturbative corrections in γ02\gamma_0^2 compute P4P^4 four-point tensors over TT time steps. Full DMFT performs matrix operations on PT×PTP T \times P T operators, resulting in O(P3T3)\mathcal{O}(P^3 T^3) runtime and O(P2T2)\mathcal{O}(P^2 T^2) memory. DMFT is more computationally and memory efficient than simulating a finite width-NN network when N≫PTN \gg P T, and it is faster than leading-order perturbation theory when T≪PT \ll \sqrt{P}.

  9. Knowl 9 — Invariance of Loss and Kernel Dynamics Under Width Rescaling at Constant Feature Strength

    empirical result

    In convolutional neural networks (depth-5 CNNs with L=4L=4 hidden layers) trained on CIFAR-10 binary classification (boat vs. plane) with mean squared error loss, scaling the channel count N→RNN \to R N while simultaneously scaling the learning rate multiplier γ→γ/R\gamma \to \gamma / \sqrt{R} leaves the feature learning strength parameter γ0=γ/N\gamma_0 = \gamma / \sqrt{N} unchanged.

    Empirical evaluation at channel widths N∈{250,500}N \in \{250, 500\} demonstrates that test mean squared error trajectories, test classification accuracy, and the alignment of the final-layer feature kernel ΦL\Phi^L with the target matrix yy⊤y y^\top collapse onto identical dynamical curves for fixed γ0\gamma_0. Furthermore, increasing γ0\gamma_0 accelerates training convergence and produces significantly greater alignment of ΦL\Phi^L with yy⊤y y^\top.

  10. Knowl 10 — Limitations of the Infinite-Width DMFT Framework

    limitation

    The self-consistent dynamical mean field theory of kernel evolution has several specific limitations:

    1. Asymptotic Scaling Assumptions: The theory is derived in the asymptotic regime where sample size P=ON(1)P = \mathcal{O}_N(1) and training time T=ON(1)T = \mathcal{O}_N(1) while width N→∞N \to \infty. It may fail to describe other scaling regimes, such as proportional data-to-width ratios P/N=ON(1)P/N = \mathcal{O}_N(1) or logarithmic time horizons T/log⁡N=ON(1)T / \log N = \mathcal{O}_N(1).
    2. Computational Complexity with Time and Sample Size: Solving the full non-perturbative self-consistent equations requires O(P3T3)\mathcal{O}(P^3 T^3) time and O(P2T2)\mathcal{O}(P^2 T^2) memory due to matrix operations on PT×PTP T \times P T kernel operators, which restricts direct evaluation on very large datasets and long training trajectories.
    3. Mathematical Rigor: The derivation relies on Martin-Siggia-Rose-De Dominicis-Janssen path integral formulations and saddle-point approximations from statistical physics rather than rigorous probabilistic limit theorems.

Coverage note — Omitted supplementary derivations of the path integral action in Appendix D, procedural details of the alternating Monte Carlo algorithm in Appendix B, and architecture-specific extensions to momentum, discrete-time training, and Langevin sampling.

References

  1. 1.Greg Yang and Edward J Hu. Tensor programs iv: Feature learning in infinite-width neural networks. In International Conference on Machine Learning, pages 11727–11737. PMLR, 2021.
  2. 2.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. MIT press, 2016.
  3. 3.Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553):436–444, 2015.
  4. 4.Radford M Neal. Bayesian learning for neural networks, volume 118. Springer Science & Business Media, 2012.
  5. 5.Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri. Deep neural networks as gaussian processes. In International Conference on Learning Representations, 2018.
  6. 6.Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani. Gaussian process behaviour in wide deep neural networks. In International Conference on Learning Representations, 2018.
  7. 7.Arthur Jacot, Franck Gabriel, and Clement Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31, pages 8571–8580. Curran Associates, Inc., 2018.
  8. 8.Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. Advances in neural information processing systems, 32, 2019.
  9. 9.Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang. On exact computation with an infinitely wide neural net. Advances in Neural Information Processing Systems, 32, 2019.
  10. 10.Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient descent finds global minima of deep neural networks. In International conference on machine learning, pages 1675–1685. PMLR, 2019.
  11. 11.B. Bordelon, A. Canatar, and C. Pehlevan. Spectrum dependent learning curves in kernel regression and wide neural networks. International Conference of Machine Learning, 2020.
  12. 12.Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan. Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks. Nature communications, 12(1):1–12, 2021.
  13. 13.Omry Cohen, Or Malka, and Zohar Ringel. Learning curves for overparametrized deep neural networks: A field theory perspective. Physical Review Research, 3(2):023034, 2021.
  14. 14.Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, and Franck Gabriel. Kernel alignment risk estimator: Risk prediction from training data. Advances in Neural Information Processing Systems, 33:15568–15578, 2020.
  15. 15.Bruno Loureiro, Cedric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mezard, and Lenka Zdeborova. Learning curves of generic features maps for realistic datasets with a teacher-student model. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021.
  16. 16.James B Simon, Madeline Dickens, and Michael R DeWeese. Neural tangent kernel eigenvalues accurately predict generalization. arXiv preprint arXiv:2110.03922, 2021.
  17. 17.Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A convergence theory for deep learning via over-parameterization. In International Conference on Machine Learning, pages 242–252. PMLR, 2019.
  18. 18.Lenaic Chizat, Edouard Oyallon, and Francis Bach. On lazy training in differentiable programming. Advances in Neural Information Processing Systems, 32, 2019.
  19. 19.Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart. Disentangling feature and lazy training in deep neural networks. Journal of Statistical Mechanics: Theory and Experiment, 2020(11):113301, 2020.
  20. 20.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  22. 22.Song Mei, Andrea Montanari, and Phan-Minh Nguyen. A mean field view of the landscape of two-layer neural networks. Proceedings of the National Academy of Sciences, 115(33):E7665–E7671, 2018.
  23. 23.Laurence Aitchison. Why bigger is not always better: on finite and infinite neural networks. In International Conference on Machine Learning, pages 156–164. PMLR, 2020.
  24. 24.Sho Yaida. Non-gaussian processes and neural networks at finite widths. In Mathematical and Scientific Machine Learning, pages 165–192. PMLR, 2020.
  25. 25.Jacob Zavatone-Veth, Abdulkadir Canatar, Ben Ruben, and Cengiz Pehlevan. Asymptotics of representation learning in finite bayesian neural networks. Advances in Neural Information Processing Systems, 34, 2021.
  26. 26.Gadi Naveh, Oded Ben David, Haim Sompolinsky, and Zohar Ringel. Predicting the outputs of finite deep neural networks trained with noisy gradients. Physical Review E, 104(6):064301, 2021.
  27. 27.Daniel A Roberts, Sho Yaida, and Boris Hanin. The principles of deep learning theory. arXiv preprint arXiv:2106.10165, 2021.
  28. 28.Boris Hanin. Correlation functions in random fully connected neural networks at finite width. arXiv preprint arXiv:2204.01058, 2022.
  29. 29.Kai Segadlo, Bastian Epping, Alexander van Meegen, David Dahmen, Michael Krämer, and Moritz Helias. Unified field theory for deep and recurrent neural networks, 2021.
  30. 30.Gadi Naveh and Zohar Ringel. A self consistent theory of gaussian processes captures feature learning effects in finite cnns. Advances in Neural Information Processing Systems, 34, 2021.
  31. 31.Inbar Seroussi and Zohar Ringel. Separation of scales and a thermodynamic description of feature learning in some cnns. arXiv preprint arXiv:2112.15383, 2021.
  32. 32.Qianyi Li and Haim Sompolinsky. Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization. Physical Review X, 11(3):031059, 2021.
  33. 33.Jacob A Zavatone-Veth and Cengiz Pehlevan. Depth induces scale-averaging in overparameterized linear bayesian neural networks. 55th Asilomar Conference on Signals, Systems, and Computers, 2021.
  34. 34.Jiaoyang Huang and Horng-Tzer Yau. Dynamics of deep neural networks and neural tangent hierarchy. In International conference on machine learning, pages 4542–4551. PMLR, 2020.
  35. 35.Ethan Dyer and Guy Gur-Ari. Asymptotics of wide networks from feynman diagrams. arXiv preprint arXiv:1909.11304, 2019.
  36. 36.Anders Andreassen and Ethan Dyer. Asymptotics of wide convolutional neural networks. arXiv preprint arXiv:2008.08675, 2020.
  37. 37.Jacob A Zavatone-Veth, William L Tong, and Cengiz Pehlevan. Contrasting random and learned features in deep bayesian linear regression. arXiv preprint arXiv:2203.00573, 2022.
  38. 38.Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari. The large learning rate phase of deep learning: the catapult mechanism. arXiv preprint arXiv:2003.02218, 2020.
  39. 39.Dyego Araújo, Roberto I Oliveira, and Daniel Yukimura. A mean-field limit for certain deep neural networks. arXiv preprint arXiv:1906.00193, 2019.
  40. 40.Lenaic Chizat and Francis Bach. On the global convergence of gradient descent for overparameterized models using optimal transport. Advances in neural information processing systems, 31, 2018.
  41. 41.Grant M Rotskoff and Eric Vanden-Eijnden. Trainability and accuracy of neural networks: An interacting particle system approach. arXiv preprint arXiv:1805.00915, 2018.
  42. 42.Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit. In Conference on Learning Theory, pages 2388–2464. PMLR, 2019.
  43. 43.Phan-Minh Nguyen. Mean field limit of the learning dynamics of multilayer neural networks. arXiv preprint arXiv:1902.02880, 2019.
  44. 44.Cong Fang, Jason Lee, Pengkun Yang, and Tong Zhang. Modeling from features: a mean-field framework for over-parameterized deep neural networks. In Conference on learning theory, pages 1887–1936. PMLR, 2021.
  45. 45.Greg Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tuning large neural networks via zero-shot hyperparameter transfer. Advances in Neural Information Processing Systems, 34, 2021.
  46. 46.Greg Yang, Michael Santacroce, and Edward J Hu. Efficient computation of deep nonlinear infinite-width neural networks that learn features. In International Conference on Learning Representations, 2022.
  47. 47.Paul Cecil Martin, ED Siggia, and HA Rose. Statistical dynamics of classical systems. Physical Review A, 8(1):423, 1973.
  48. 48.C De Dominicis. Dynamics as a substitute for replicas in systems with quenched random impurities. Physical Review B, 18(9):4913, 1978.
  49. 49.Haim Sompolinsky and Annette Zippelius. Dynamic theory of the spin-glass phase. Physical Review Letters, 47(5):359, 1981.
  50. 50.Haim Sompolinsky and Annette Zippelius. Relaxational dynamics of the edwards-anderson model and the mean-field theory of spin-glasses. Physical Review B, 25(11):6860, 1982.
  51. 51.G Ben Arous and Alice Guionnet. Large deviations for langevin spin glass dynamics. Probability Theory and Related Fields, 102(4):455–509, 1995.
  52. 52.G Ben Arous and Alice Guionnet. Symmetric langevin spin glass dynamics. The Annals of Probability, 25(3):1367–1422, 1997.
  53. 53.Gérard Ben Arous, Amir Dembo, and Alice Guionnet. Cugliandolo-kurchan equations for dynamics of spin-glasses. Probability theory and related fields, 136(4):619–660, 2006.
  54. 54.A Crisanti and H Sompolinsky. Path integral approach to random neural networks. Physical Review E, 98(6):062120, 2018.
  55. 55.Haim Sompolinsky, Andrea Crisanti, and Hans-Jurgen Sommers. Chaos in random neural networks. Physical review letters, 61(3):259, 1988.
  56. 56.Moritz Helias and David Dahmen. Statistical Field Theory for Neural Networks. Springer International Publishing, 2020.
  57. 57.Lutz Molgedey, J Schuchhardt, and Heinz G Schuster. Suppressing chaos in neural networks by noise. Physical review letters, 69(26):3717, 1992.
  58. 58.M Samuelides and Bruno Cessac. Random recurrent neural networks dynamics. The European Physical Journal Special Topics, 142(1):89–122, 2007.
  59. 59.Kanaka Rajan, LF Abbott, and Haim Sompolinsky. Stimulus-dependent suppression of chaos in recurrent neural networks. Physical review e, 82(1):011903, 2010.
  60. 60.Stefano Sarao Mannelli, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborova. Passed & spurious: Descent algorithms and local minima in spiked matrix-tensor models. In international conference on machine learning, pages 4333–4342. PMLR, 2019.
  61. 61.Stefano Sarao Mannelli, Giulio Biroli, Chiara Cammarota, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová. Marvels and pitfalls of the langevin algorithm in noisy high-dimensional inference. Physical Review X, 10(1):011057, 2020.
  62. 62.Francesca Mignacco, Pierfrancesco Urbani, and Lenka Zdeborová. Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem. Machine Learning: Science and Technology, 2(3):035029, 2021.
  63. 63.Elisabeth Agoritsas, Giulio Biroli, Pierfrancesco Urbani, and Francesco Zamponi. Out-of-equilibrium dynamical mean-field equations for the perceptron model. Journal of Physics A: Mathematical and Theoretical, 51(8):085002, 2018.
  64. 64.Francesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová. Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification. Advances in Neural Information Processing Systems, 33:9540–9550, 2020.
  65. 65.Michael Celentano, Chen Cheng, and Andrea Montanari. The high-dimensional asymptotics of first order methods with random data. arXiv preprint arXiv:2112.07572, 2021.
  66. 66.Francesca Mignacco and Pierfrancesco Urbani. The effective noise of stochastic gradient descent. arXiv preprint arXiv:2112.10852, 2021.
  67. 67.Yizhang Lou, Chris E Mingard, and Soufiane Hayou. Feature learning and signal propagation in deep neural networks. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 14248–14282. PMLR, 17–23 Jul 2022.
  68. 68.Alessandro Manacorda, Grégory Schehr, and Francesco Zamponi. Numerical solution of the dynamical mean field theory of infinite-dimensional equilibrium liquids. The Journal of chemical physics, 152(16):164506, 2020.
  69. 69.Kenji Fukumizu. Dynamics of batch learning in multilayer neural networks. In International Conference on Artificial Neural Networks, pages 189–194. Springer, 1998.
  70. 70.Andrew M Saxe, James L McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv preprint arXiv:1312.6120, 2013.
  71. 71.Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu. A convergence analysis of gradient descent for deep linear neural networks. In International Conference on Learning Representations, 2019.
  72. 72.Madhu S Advani, Andrew M Saxe, and Haim Sompolinsky. High-dimensional dynamics of generalization error in neural networks. Neural Networks, 132:428–446, 2020.
  73. 73.Arthur Jacot, François Ged, Franck Gabriel, Berfin Şimşek, and Clément Hongler. Deep linear networks dynamics: Low-rank biases induced by initialization scale and l2 regularization. arXiv preprint arXiv:2106.15933, 2021.
  74. 74.Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan. Neural networks as kernel learners: The silent alignment effect. In International Conference on Learning Representations, 2022.
  75. 75.Aitor Lewkowycz and Guy Gur-Ari. On the training dynamics of deep networks with l_2 regularization. Advances in Neural Information Processing Systems, 33:4790–4799, 2020.
  76. 76.Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli. Exponential expressivity in deep neural networks through transient chaos. Advances in neural information processing systems, 29, 2016.
  77. 77.Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein. Deep information propagation. International Conference of Learning Representations, 2017.
  78. 78.Greg Yang and Samuel Schoenholz. Mean field residual networks: On the edge of chaos. Advances in neural information processing systems, 30, 2017.
  79. 79.Greg Yang. Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation. arXiv preprint arXiv:1902.04760, 2019.
  80. 80.Greg Yang and Etai Littwin. Tensor programs iib: Architectural universality of neural tangent kernel training dynamics. In International Conference on Machine Learning, pages 11762–11772. PMLR, 2021.
  81. 81.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  82. 82.Haozhe Shan and Blake Bordelon. A theory of neural tangent kernel alignment and its influence on training, 2021.
  83. 83.Stéphane d’Ascoli, Maria Refinetti, and Giulio Biroli. Optimal learning rate schedules in high-dimensional non-convex optimization problems, 2022.
  84. 84.James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. JAX: composable transformations of Python+NumPy programs, 2018.
  85. 85.Juha Honkonen. Ito and stratonovich calculuses in stochastic field theory. arXiv preprint arXiv:1102.1581, 2011.
  86. 86.Crispin W Gardiner et al. Handbook of stochastic methods, volume 3. springer Berlin, 1985.
  87. 87.Carl M Bender and Steven Orszag. Advanced mathematical methods for scientists and engineers I: Asymptotic methods and perturbation theory, volume 1. Springer Science & Business Media, 1999.
  88. 88.John Hubbard. Calculation of partition functions. Physical Review Letters, 3(2):77, 1959.
  89. 89.Charles Stein. A bound for the error in the normal approximation to the distribution of a sum of dependent random variables. In Proceedings of the sixth Berkeley symposium on mathematical statistics and probability, volume 2: Probability theory, volume 6, pages 583–603. University of California Press, 1972.
  90. 90.Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein. Bayesian deep convolutional networks with many channels are gaussian processes. arXiv preprint arXiv:1810.05148, 2018.
  91. 91.Greg Yang. Wide feedforward or recurrent neural networks of any architecture are gaussian processes. Advances in Neural Information Processing Systems, 32, 2019.
  92. 92.Adam X. Yang, Maxime Robeyns, Edward Milsom, Nandi Schoots, and Laurence Aitchison. A theory of representation learning in deep neural networks gives a deep generalisation of kernel methods, 2021.
  93. 93.Yurii E Nesterov. A method for solving the convex programming problem with convergence rate o (1/kˆ 2). In Dokl. akad. nauk Sssr, volume 269, pages 543–547, 1983.
  94. 94.Yurii Nesterov and Boris T Polyak. Cubic regularization of newton method and its global performance. Mathematical Programming, 108(1):177–205, 2006.
  95. 95.Gabriel Goh. Why momentum really works. Distill, 2017.
  96. 96.Michael Muehlebach and Michael I Jordan. Optimization with momentum: Dynamical, control-theoretic, and symplectic perspectives. Journal of Machine Learning Research, 22(73):1–50, 2021.
  97. 97.Mehran Kardar. Statistical physics of fields. Cambridge University Press, 2007.

Citation

MLA
Bordelon, B., and C. Pehlevan. “Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 32240–56, https://proceedings.neurips.cc/paper_files/paper/2022/file/d027a5c93d484a4312cc486d399c62c1-Paper-Conference.pdf.
APA
Bordelon, B., & Pehlevan, C. (2022). Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks. Advances in Neural Information Processing Systems, 35, 32240–32256. https://proceedings.neurips.cc/paper_files/paper/2022/file/d027a5c93d484a4312cc486d399c62c1-Paper-Conference.pdf
Chicago
Bordelon, B., and C. Pehlevan. 2022. “Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks”. Advances in Neural Information Processing Systems 35: 32240–56. https://proceedings.neurips.cc/paper_files/paper/2022/file/d027a5c93d484a4312cc486d399c62c1-Paper-Conference.pdf.
Harvard
Bordelon, B. and Pehlevan, C. (2022) “Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 32240–32256. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/d027a5c93d484a4312cc486d399c62c1-Paper-Conference.pdf.
Vancouver
1. Bordelon B, Pehlevan C (2022) Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 32240–32256

BibTeX

@inproceedings{bordelon2022self,
  title = {Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks},
  author = {Bordelon, Blake and Pehlevan, Cengiz},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {32240-32256},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/d027a5c93d484a4312cc486d399c62c1-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors