Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks
Blake BordelonCengiz Pehlevan
Develops a self-consistent dynamical field theory and an efficient sampling algorithm to track kernel evolution and feature learning in infinite-width neural networks beyond the static neural tangent kernel regime.
Deep learning models have achieved widespread success across modern technology, yet theoretical understanding of how these networks learn features from data during training remains incomplete. Existing foundational theories often assume an infinite-width limit where internal data representations remain static throughout optimization—known as the lazy training regime. However, practical deep neural networks succeed precisely because they adapt their internal representations to the data. Bridging this gap between theory and practical feature adaptation is essential for systematically understanding, tuning, and scaling modern artificial intelligence systems.
The main objective of the article is to establish and validate an exact dynamical field theory that characterizes how internal feature representations and gradient signals evolve over time in wide neural networks. It demonstrates how these dynamics can be captured through a reduced set of deterministic kernel parameters that smoothly interpolate between static and adaptive learning regimes.
To achieve this, the article adapts mathematical tools from statistical physics known as dynamical mean field theory. The authors formulate the gradient-based training trajectory as a functional path integral in the infinite-width limit, deriving a set of self-consistent equations that govern internal network features and training dynamics. For deep linear architectures, these equations simplify into closed algebraic systems, whereas for nonlinear networks, the authors introduce a polynomial-time alternating sampling algorithm to compute the solutions. The resulting theoretical predictions were validated against numerical simulations of finite-width fully connected networks and convolutional neural networks trained on image classification tasks.
The investigation yields several key findings. First, the article demonstrates that network activation and gradient distributions over training can be exactly summarized by deterministic kernel order parameters, recovering previous stochastic descriptions from tensor program frameworks while providing a more flexible formulation. Second, deep linear networks yield exact algebraic matrix solutions showing that deeper architectures experience substantially larger shifts in internal representations during training. Third, existing approximation methods—such as the static neural tangent kernel, gradient independence assumptions, and leading-order perturbation theory—break down significantly as the strength of feature learning and network depth increase, whereas the dynamical mean field theory remains accurate. Finally, experiments on image classification confirm that when the feature learning strength parameter is held constant, loss and representation dynamics remain invariant across different network widths, matching the theory once width reaches moderately large sizes (e.g., around 500 hidden units).
These findings have practical implications for designing and scaling deep learning architectures. By showing that representation dynamics are invariant under specific width and learning parameter scalings, the theory supports principled hyperparameter transfer from small pilot models to large-scale production architectures without costly trial-and-error retraining. Furthermore, the framework explains why regularized networks can preserve expressive, non-trivial predictive solutions rather than decaying to zero when feature adaptation is sufficiently strong.
Organizations developing large-scale neural network architectures should consider utilizing these scaling relationships to guide model expansion and hyperparameter tuning across widths. While the exact numerical solution currently incurs a computational complexity that scales cubically with time steps and sample size, practitioners can use the qualitative scaling laws to inform model depth and initialization schedules. For decision-makers evaluating theoretical predictive tooling, additional work should focus on exploring accelerated projected gradient methods or data-averaged approximations before deploying these theoretical solvers across massive-scale production datasets.
Confidence in these theoretical results is high within the defined boundaries of wide networks and gradient-flow training. Readers should exercise caution, however, when extrapolating to small sample-to-width ratios or extreme training horizons where the infinite-width assumptions and physics-based heuristic saddle-point derivations may require further finite-size corrections.
- Paper: Neural Tangent Kernel: Convergence and Generalization in Neural Networks, Arthur Jacot et al. (2018). Introduces the Neural Tangent Kernel and the static infinite-width training regime that the source article builds upon and extends into an adaptive dynamical field theory.
- Paper: Wide neural networks of any depth evolve as linear models under gradient descent, Jaehoon Lee et al. (2019). Establishes how wide neural networks behave as linear models under gradient descent, providing the exact lazy training baseline that the source paper seeks to generalize beyond.
- Paper: Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, Andrew M. Saxe et al. (2014). Derives the analytical nonlinear learning dynamics in deep linear networks that the source paper directly adapts and generalizes to wide nonlinear architectures via field theory.
- Book: The Principles of Deep Learning Theory, Daniel A. Roberts et al. (2022). Provides the foundational statistical physics, field-theoretic path integral methods, and large-width perturbation expansions required to understand the source's mathematical derivations.
- Paper: Gradient Descent Provably Optimizes Over-parameterized Neural Networks, Simon S. Du et al. (2018). Analyzes the convergence of over-parameterized networks by tracking Gram matrix dynamics near initialization, laying the groundwork for continuous kernel tracking during training.
- Paper: There Will Be a Scientific Theory of Deep Learning, Jamie Simon et al. (2026). Synthesizes emerging scientific theories of learning mechanics, contextualizing infinite-width dynamical limits and the transition from lazy kernels to rich feature representations.
- Paper: High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation, Jimmy Ba et al. (2022). Examines how initial gradient descent steps induce high-dimensional feature learning beyond static kernels, providing a complementary concrete analysis of early representation evolution.
- Paper: Exact learning dynamics of deep linear networks with prior knowledge, Lukas Braun et al. (2022). Extends the analytical study of representation dynamics and kernel metrics in deep linear networks to structured initializations and transfer learning regimes.
- Paper: Feature learning in deep classifiers through Intermediate Neural Collapse, Akshay Rangamani et al. (2023). Explores how intermediate layer representations compress and collapse into geometric structures throughout training, extending the empirical understanding of feature evolution.
