Bayes-optimal Learning of Deep Random Networks of Extensive-width

Hugo CuiFlorent KrzakalaLenka Zdeborová

article2023ICML53 citations

Establishes closed-form theoretical bounds on Bayes-optimal test errors for learning deep extensive-width random neural networks, revealing that simple kernel and ridge methods match optimal Bayesian performance when sample size scales linearly with input dimension but fail when sample size grows quadratically.

Listen

Understanding the fundamental data requirements and performance limits of deep learning architectures remains a major challenge in artificial intelligence theory. Machine learning practitioners often deploy highly parameterized neural networks without a precise theoretical understanding of the minimum sample size required to learn a given function or whether standard optimization algorithms can match theoretical efficiency bounds. Characterizing these fundamental limits is critical for designing cost-effective data collection strategies and selecting appropriate model architectures.

The article establishes the information-theoretic performance limits—specifically the Bayes-optimal test error—for learning deep non-linear neural networks with extensive width from random Gaussian data. It evaluates whether standard empirical risk minimization methods, including linear models, random features, and kernel methods, can achieve these theoretical optimality bounds across both classification and regression tasks.

To conduct this evaluation, the authors combine statistical physics techniques, specifically the replica method, with high-dimensional probability theory. The analysis models a target network consisting of multiple non-linear layers with random Gaussian weights in the proportional asymptotic regime, where the sample size, input dimension, and network layer widths grow large at proportional rates. The analytical framework relies on the Bayesian Gaussian Equivalence Property, which approximates the layer-to-layer activations of deep architectures using equivalent Gaussian representations.

The investigation yields several key findings regarding model capabilities across data regimes. First, in the proportional data regime where the number of training samples scales linearly with the input dimension, no algorithm can extract more than a linear approximation of the deep target network. Second, simple and computationally inexpensive techniques—specifically optimally regularized ridge regression and kernel regression—match the Bayes-optimal error exactly for regression tasks, while standard logistic regression approaches near-optimal performance for classification. Third, finite-width random feature models remain suboptimal compared to full kernel methods due to representation mismatch. Finally, when the sample size increases quadratically relative to the input dimension, linear and kernel methods become strictly suboptimal, while gradient-trained deep neural networks learn the non-linear target almost perfectly, reducing error by several orders of magnitude.

These findings have direct implications for system design, training costs, and computational resource allocation. When training data is scarce and scales proportionally with the feature dimension, complex deep learning pipelines offer negligible performance advantages over optimally tuned linear or kernel baselines, making the added training and maintenance costs unjustifiable. However, when large datasets are available at super-linear sample complexities, deep neural networks demonstrate a definitive performance advantage through genuine non-linear feature learning that simpler models cannot replicate.

Organizations should adopt a tiered modeling approach based on available data scale. In low-data environments where sample size is comparable to feature dimension, engineering teams should deploy regularized linear or kernel baselines using the explicit noise-to-signal regularizers derived in the analysis to minimize computational overhead. Transitioning to computationally demanding deep architectures should be prioritized primarily when dataset sizes exceed linear scaling thresholds. Further theoretical work should focus on formally establishing mathematical proofs for the deep Gaussian equivalence conjecture and quantitatively mapping the sample complexity transitions in super-linear data regimes.

The conclusions are subject to certain boundary conditions. The analysis assumes Gaussian input distributions, random Gaussian network weights, and asymptotic scaling limits. While theoretical predictions are supported by extensive numerical simulations across various network depths and activation functions, leaders should exercise caution when extrapolating these quantitative bounds directly to highly structured, non-Gaussian real-world datasets.

Cover for Bayes-optimal Learning of Deep Random Networks of Extensive-width

Abstract

We consider the problem of learning a target function corresponding to a deep, extensive-width, non-linear neural network with random Gaussian weights. We consider the asymptotic limit where the number of samples, the input dimension and the network width are proportionally large. We propose a closed-form expression for the Bayes-optimal test error, for regression and classification tasks. We further compute closed-form expressions for the test errors of ridge regression, kernel and random features regression. We find, in particular, that optimally regularized ridge regression, as well as kernel regression, achieve Bayes-optimal performances, while the logistic loss yields a near-optimal test error for classification. We further show numerically that when the number of samples grows faster than the dimension, ridge & kernel methods become suboptimal, while neural networks achieve test error close to zero from quadratically many samples.

Table of Contents

  • 1. Introduction
  • 1.1. Related works
  • 2. Setting
  • 3. Bayes-optimal Error
  • 3.1. The Bayesian Gaussian Equivalence Property
  • 3.2. Deep (Bayesian) Gaussian Equivalence Property
  • 3.3. Bayes-optimal errors
  • 4. ERM with Linear Methods
  • 4.1. Ridge regression
  • 4.2. Random Features
  • 4.3. Kernels
  • 4.4. Logistic and ridge classification
  • 5. Beyond the Proportional Regime
  • Conclusion
  • Acknowledgements
  • References
  • A. Deep Gaussian Equivalence Principles
  • A.1. Setting
  • A.2. Closed-form formulas for second order population statistics
  • A.3. Derivation for Ω , Ψ
  • A.4. Derivation sketch for Φ
  • A.5. Numerical evidence of the closed-form recursions
  • A.6. Numerical test of the 1dCLT
  • A.7. Towards rigorous proofs
  • B. Replica computation for shallow networks
  • B.1. Replica trick
  • B.2. GET-linearization of the hidden layer
  • B.3. Entropic potential
  • B.4. Replica-symmetric ansatz
  • B.5. RS free energy
  • B.6. Finite temperature free energy
  • B.7. Bayes-Optimal setting
  • B.8. Nishimori identities
  • B.9. Bayes-optimal saddle points
  • C. Replica computation for deep networks
  • C.1. Generalization to multi-layer nets
  • C.2. Tower of order parameters
  • C.3. Replica Symmetry
  • C.4. Multilayer Loss potential
  • C.5. Layer-wise entropy
  • C.6. Finite T multilayer free energy
  • C.7. Bayes-Optimal setting
  • C.8. Bayes-Optimal Saddle point equations
  • D. Equivalent shallow network
  • E. ERM : the shallow case
  • F. Optimality of kernel ERM on shallow targets
  • F.1. Ridge ERM achieves Bayes optimal error
  • F.2. Kernels achieve Bayes optimal error
  • F.3. A short RMT argument for finite width RF
  • G. ERM : the deep case
  • G.1. Ridge regression
  • G.2. Random nets / features
  • G.3. GP kernels
  • H. Classification
  • H.1. Bayes-optimal error for classification
  • H.2. Bayes optimal classification error
  • H.3. ERM

Knowls

  1. Knowl 1 — Extensive-width teacher–student setting

    definition

    The target is a depth-LL random neural network evaluated on Gaussian inputs x∈Rdx\in\mathbb{R}^d with x∼N(0,Σ)x\sim\mathcal{N}(0,\Sigma). Set h0(x)=xh_0(x)=x and, for 1≤ℓ≤L1\leq\ell\leq L, define hℓ(x)=σℓ(Wℓhℓ−1(x)/kℓ−1)h_\ell(x)=\sigma_\ell(W_\ell h_{\ell-1}(x)/\sqrt{k_{\ell-1}}), where k0=dk_0=d, Wℓ∈Rkℓ×kℓ−1W_\ell\in\mathbb{R}^{k_\ell\times k_{\ell-1}} has independent Gaussian entries of variance Δℓ\Delta_\ell, and the readout vector has distribution a∼N(0,ΔaIkL)a\sim\mathcal{N}(0,\Delta_a I_{k_L}). The label is y=f∗(aThL(x)/kL+ξ)y=f_*(a^T h_L(x)/\sqrt{k_L}+\xi), with output noise ξ∼N(0,Δ)\xi\sim\mathcal{N}(0,\Delta); f∗(u)=uf_*(u)=u for regression and f∗(u)=sign⁡(u)f_*(u)=\operatorname{sign}(u) for classification. The learner knows the architecture, activations, and parameter distributions but not the sampled weights. The main asymptotic regime sends n,d,k1,…,kLn,d,k_1,\ldots,k_L to infinity with fixed ratios α=n/d\alpha=n/d and γℓ=kℓ/d\gamma_\ell=k_\ell/d. The input covariance has a limiting spectral distribution μ\mu with finite, nonzero first and second moments; the paper assumes odd activations for simplicity.

  2. Knowl 2 — Deep Bayesian Gaussian equivalence

    model/method

    The paper conjectures that in the extensive-width proportional limit, the outputs of any finite number of networks independently sampled from the matched Bayes posterior are jointly Gaussian over the Gaussian input xx. For replica indices a,ba,b, let hℓa(x)h_\ell^a(x) be the post-activation vector at layer ℓ\ell and define Ωℓab=Ex[hℓa(x)hℓb(x)T]\Omega_\ell^{ab}=\mathbb{E}_x[h_\ell^a(x)h_\ell^b(x)^T], with Ω0ab=Σ\Omega_0^{ab}=\Sigma. The conjectured population-covariance recursion is

    Ωℓab=(κ1(ℓ))2WℓaΩℓ−1ab(Wℓb)Tkℓ−1+δab(κ∗(ℓ))2Ikℓ.\Omega_\ell^{ab}=(\kappa_1^{(\ell)})^2\frac{W_\ell^a\Omega_{\ell-1}^{ab}(W_\ell^b)^T}{k_{\ell-1}}+\delta_{ab}(\kappa_*^{(\ell)})^2 I_{k_\ell}.

    Here WℓaW_\ell^a is the layer-ℓ\ell weight matrix in replica aa, and δab\delta_{ab} is one when a=ba=b and zero otherwise. To define the coefficients, let M=∫z dμ(z)M=\int z\,d\mu(z), r1=Δ1Mr_1=\Delta_1M, rℓ+1=Δℓ+1EZ∼N(0,rℓ)[σℓ(Z)2]r_{\ell+1}=\Delta_{\ell+1}\mathbb{E}_{Z\sim\mathcal{N}(0,r_\ell)}[\sigma_\ell(Z)^2], κ1(ℓ)=E[Zσℓ(Z)]/rℓ\kappa_1^{(\ell)}=\mathbb{E}[Z\sigma_\ell(Z)]/r_\ell, and (κ∗(ℓ))2=E[σℓ(Z)2]−rℓ(κ1(ℓ))2(\kappa_*^{(\ell)})^2=\mathbb{E}[\sigma_\ell(Z)^2]-r_\ell(\kappa_1^{(\ell)})^2, where the expectations use Z∼N(0,rℓ)Z\sim\mathcal{N}(0,r_\ell). The claim applies when the activations have zero mean under these Gaussian pre-activations, including odd activations. This is a conjecture, not a proved theorem. Numerical support includes close agreement of theoretical and empirically estimated activation covariances for tanh, sign, and erf networks at d=500d=500 using 10510^5 inputs (relative squared Frobenius discrepancies about 0.0050.005, 0.0080.008, and 0.0040.004, respectively; covariance panels on pages 19–20). Histograms and quantile–quantile plots for depth 3 and 9 networks, with dimensions up to about 1000, support Gaussianity of the scaled output (pages 21–23).

  3. Knowl 3 — Bayes-optimal regression error

    theoretical result

    Under the deep Bayesian Gaussian-equivalence conjecture, the asymptotic Bayes-optimal mean-squared error for the extensive-width teacher is

    ϵg,regBO=S(ΔaMP−q)+ϵr,q=12∫αSz2Δa2P2ϵg,regBO+αSzΔaP dμ(z),\epsilon_{g,\mathrm{reg}}^{\mathrm{BO}}=S\left(\Delta_a M P-q\right)+\epsilon_r, \qquad q=\frac12\int\frac{\alpha S z^2\Delta_a^2P^2}{\epsilon_{g,\mathrm{reg}}^{\mathrm{BO}}+\alpha S z\Delta_aP}\,d\mu(z),

    where S=∏ℓ=1L(κ1(ℓ))2S=\prod_{\ell=1}^L(\kappa_1^{(\ell)})^2, P=∏ℓ=1LΔℓP=\prod_{\ell=1}^L\Delta_\ell, M=∫z dμ(z)M=\int z\,d\mu(z), and zz is a spectral variable distributed according to the limiting input-covariance spectral measure μ\mu. The residual error is

    ϵr=∑ℓ0=1L−1(κ∗(ℓ0))2Δa∏ℓ=ℓ0+1L(κ1(ℓ))2Δℓ+(κ∗(L))2Δa+Δ.\epsilon_r=\sum_{\ell_0=1}^{L-1}(\kappa_*^{(\ell_0)})^2\Delta_a\prod_{\ell=\ell_0+1}^{L}(\kappa_1^{(\ell)})^2\Delta_\ell+(\kappa_*^{(L)})^2\Delta_a+\Delta.

    The coefficients κ1(ℓ)\kappa_1^{(\ell)} and κ∗(ℓ)\kappa_*^{(\ell)} are defined from the layer activations and Gaussian pre-activation variances in the deep Gaussian-equivalence recursion. The scalar qq is the Bayes posterior self-overlap. This error is the information-theoretic minimum achievable from nn samples in the proportional regime, conditional on the paper’s Gaussian-equivalence conjecture.

  4. Knowl 4 — Bayes-optimal classification error

    theoretical result

    For the sign-readout classification task, the conjectured asymptotic Bayes misclassification probability is

    ϵg,classBO=1πarccos⁡(SqST+ϵr),\epsilon_{g,\mathrm{class}}^{\mathrm{BO}}=\frac{1}{\pi}\arccos\left(\sqrt{\frac{S q}{S T+\epsilon_r}}\right),

    where S=∏ℓ=1L(κ1(ℓ))2S=\prod_{\ell=1}^L(\kappa_1^{(\ell)})^2, T=ΔaMPT=\Delta_a M P, M=∫z dμ(z)M=\int z\,d\mu(z), P=∏ℓ=1LΔℓP=\prod_{\ell=1}^L\Delta_\ell, and ϵr\epsilon_r is the residual variance defined by the layerwise activation coefficients and output noise. The self-overlap qq and its conjugate scalar q^\widehat q are determined by

    q=∫q^ 2Δa2P2z21+q^ΔaPz dμ(z),q=\int\frac{\widehat q^{\,2}\Delta_a^2P^2z^2}{1+\widehat q\Delta_aPz}\,d\mu(z), \widehat q=\frac{2\alpha S}{ST+\epsilon_r-Sq}\int_{-\infty}^{\infty}\frac{2\exp\!\left[-\frac12\frac{ST+\epsilon_r+Sq}{ST+\epsilon_r-Sq}\,u^2\right]}{(2\pi)^{3/2}\left[1-\operatorname{erf}\!\left(\frac{\sqrt{Sq}\,u}{\sqrt{2(ST+\epsilon_r-Sq)}}\right)\right]}}\,du.

    Here uu is a real integration variable, erf⁡\operatorname{erf} is the Gaussian error function, and α=n/d\alpha=n/d. The result is the paper’s Bayes-optimal classification prediction in the proportional regime, relying on the conjectured deep Bayesian Gaussian equivalence.

  5. Knowl 5 — Equivalent noisy linear teacher for Bayes errors

    model/method

    Although the target is a nonlinear deep network, its Bayes errors in the proportional regime are conjectured to equal those of a single-layer noisy teacher

    yeq(x)=f∗(ρ θTxd+ϵr ζ),ρ=Δa∏ℓ=1L(κ1(ℓ))2Δℓ.y_{\mathrm{eq}}(x)=f_*\left(\frac{\sqrt{\rho}\,\theta^Tx}{\sqrt d}+\sqrt{\epsilon_r}\,\zeta\right), \qquad \rho=\Delta_a\prod_{\ell=1}^L(\kappa_1^{(\ell)})^2\Delta_\ell.

    Here θ∈Rd\theta\in\mathbb{R}^d has independent standard Gaussian entries, ζ∼N(0,1)\zeta\sim\mathcal{N}(0,1) is independent output noise, f∗f_* is the regression identity or classification sign function, and ϵr\epsilon_r is the residual variance generated by the deep nonlinearities and any teacher output noise. The coefficients κ1(ℓ)\kappa_1^{(\ell)} and κ∗(ℓ)\kappa_*^{(\ell)} are those of the deep Gaussian-equivalence recursion. The effective noise represents nonlinear components that are not learned from a proportional number of samples; it does not assert that the original fixed-weight deep teacher is itself stochastic.

  6. Knowl 6 — Optimally regularized ridge regression is Bayes-optimal

    theoretical result

    For regression in the proportional regime, ridge regression reaches the Bayes-optimal MSE when its regularization is set to

    λ∗=ϵrρ,ρ=Δa∏ℓ=1L(κ1(ℓ))2Δℓ.\lambda^*=\frac{\epsilon_r}{\rho}, \qquad \rho=\Delta_a\prod_{\ell=1}^L(\kappa_1^{(\ell)})^2\Delta_\ell.

    Here ϵr\epsilon_r is the effective residual variance of the equivalent noisy linear teacher, and the activation coefficients κ1(ℓ)\kappa_1^{(\ell)} define its signal strength ρ\rho. With this regularization, the ridge test-error equations reduce to the Bayes-optimal regression equations. Thus, under the paper’s asymptotic characterization, no estimator can improve on optimally regularized linear regression using only proportionally many samples. The formula matches the interpretation that the ridge penalty should equal the effective teacher noise-to-signal ratio.

  7. Knowl 7 — Optimally regularized kernel regression is Bayes-optimal

    theoretical result

    The infinite-width kernel-regression limit also attains the Bayes-optimal regression MSE when regularized with

    λ∗=κ12(ϵrρ−κ∗2κ12).\lambda^*=\kappa_1^2\left(\frac{\epsilon_r}{\rho}-\frac{\kappa_*^2}{\kappa_1^2}\right).

    Here ρ=Δa∏ℓ=1L(κ1(ℓ))2Δℓ\rho=\Delta_a\prod_{\ell=1}^L(\kappa_1^{(\ell)})^2\Delta_\ell and ϵr\epsilon_r describe the effective deep teacher; κ1\kappa_1 and κ∗\kappa_* are the linear and residual Gaussian-equivalence coefficients of the kernel’s activation. The kernel’s implicit regularization is accounted for in the formula, so the explicit optimal λ∗\lambda^* can be negative. In that case the total risk remains convex because of the kernel’s implicit ℓ2\ell_2 regularization. At the stated regularization, the kernel test-error equations reduce to the Bayes-optimal regression equations.

  8. Knowl 8 — Finite-width random features are suboptimal for regression

    empirical result

    For isotropic inputs and a random-feature width kk proportional to dimension dd, the paper’s asymptotic analysis and experiments find that optimally regularized random-features regression has higher MSE than the Bayes-optimal benchmark at finite width ratio γ=k/d\gamma=k/d. In the equivalent teacher, the target signal is linear in the original input coordinates, whereas random features perform linear readout in a transformed feature basis; this basis mismatch prevents finite-width random features from attaining the Bayes error. In the infinite-width limit γ→∞\gamma\to\infty, random features converge to kernel regression, which can attain the Bayes error with its optimal regularization. The regression learning-curve comparisons, including the finite-width random-feature gap, are shown for one- and two-hidden-layer tanh teachers on page 6.

  9. Knowl 9 — Classification ERM is close to, but not exactly, Bayes-optimal

    empirical result

    In the noiseless classification experiments, optimally regularized logistic regression and ridge classification have test errors very close to the conjectured Bayes error, but neither coincides with it. For the reported three-layer target with tanh⁡(2x)\tanh(2x) activations, widths 700, input dimension 500, and 30 trials, the theoretical gaps are approximately 10−410^{-4} for logistic regression and 10−310^{-3} for ridge classification. The learning curves and a magnified view of the gaps appear on page 8. These results concern the specified proportional-regime setup and should not be read as exact Bayes optimality for either classifier.

  10. Knowl 10 — Neural networks outperform linear and kernel methods with quadratic sample counts

    empirical result

    The paper numerically explores regression when the sample count scales quadratically with input dimension, where the proportional-regime Gaussian-equivalence analysis no longer applies. For one-hidden-layer ReLU and erf teachers with d=30d=30 and width k1=20k_1=20, the comparison on page 9 shows that a fully trained neural network can learn the target nearly perfectly, reaching MSE of order 10−510^{-5} at quadratic sample counts. Optimally regularized kernels learn only a quadratic approximation and have errors about three orders of magnitude larger; ridge regression, which captures only a linear approximation, performs worse still. The neural-network experiments used Adam for 2000 epochs, batch size n/3n/3, learning rate 3×10−33\times10^{-3}, and 10 trials. This is numerical evidence for the quadratic regime, not a closed-form asymptotic theory.

Coverage note — The detailed replica derivations and full fixed-point systems for the ERM learning curves are omitted because they are supporting calculations rather than separate conclusions; the paper also leaves rigorous proofs of the deep Bayesian Gaussian-equivalence conjecture and a quantitative theory beyond the proportional regime open.

References

  1. 1.Advani, M. and Ganguli, S. Statistical mechanics of optimal convex inference in high dimensions. Physical Review X, 6(3):031034, 2016.
  2. 2.Ariosto, S., Pacelli, R., Pastore, M., Ginelli, F., Gherardi, M., and Rotondo, P. Statistical mechanics of deep learning beyond the infinite-width limit. ArXiv, abs/2209.04882, 2022.
  3. 3.Arora, S., Du, S. S., Li, Z., Salakhutdinov, R., Wang, R., and Yu, D. Harnessing the power of infinitely wide deep nets on small-data tasks. Proc. Int. Conf. Learning Rep. (ICLR), 2020.
  4. 4.Aubin, B., Maillard, A., Barbier, J., Krzakala, F., Macris, N., and Zdeborova, L. The committee machine: computational to statistical gaps in learning a two-layers neural network. Journal of Statistical Mechanics: Theory and Experiment, 2019, 2018.
  5. 5.Aubin, B., Krzakala, F., Lu, Y. M., and Zdeborova, L. Generalization error in high-dimensional perceptrons: Approaching bayes error with convex optimization. Advances in Neural Information Processing Systems, 33:12199–12210, 2020.
  6. 6.Barbier, J., Krzakala, F., Macris, N., Miolane, L., and Zdeborova, L. Optimal errors and phase transitions in high-dimensional generalized linear models. Proceedings of the National Academy of Sciences of the United States of America, 116:5451 – 5460, 2017.
  7. 7.Bosch, D., Panahi, A., and Hassibi, B. Precise asymptotic analysis of deep random feature models. ArXiv, abs/2302.06210, 2023.
  8. 8.Canatar, A., Bordelon, B., and Pehlevan, C. Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks. Nature Communications, 12, 2020.
  9. 9.Clarté, L., Loureiro, B., Krzakala, F., and Zdeborova, L. A study of uncertainty quantification in overparametrized high-dimensional models. ArXiv, abs/2210.12760, 2022.
  10. 10.Cui, H., Saglietti, L., and Zdeborova, L. Large deviations for the perceptron model and consequences for active learning. Mach. Learn. Sci. Technol., 2:45001, 2019.
  11. 11.Cui, H., Loureiro, B., Krzakala, F., and Zdeborova, L. Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime. In NeurIPS, 2021.
  12. 12.Cui, H., Loureiro, B., Krzakala, F., and Zdeborova, L. Error rates for kernel classification under source and capacity conditions. ArXiv, abs/2201.12655, 2022.
  13. 13.d’Ascoli, S., Gabrie, M., Sagun, L., and Biroli, G. On the interplay between data structure and loss function in classification problems. In NeurIPS, 2021.
  14. 14.de G. Matthews, A. G., Rowland, M., Hron, J., Turner, R. E., and Ghahramani, Z. Gaussian process behaviour in wide deep neural networks. NeurIPS Workshop on Advances in Approximate Bayesian Inference, 2018.
  15. 15.El Karoui, N. Spectrum estimation for large dimensional covariance matrices using random matrix theory. The Annals of Statistics, 36(6):2757–2790, 2008.
  16. 16.Fan, Z. and Wang, Z. Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks. ArXiv, abs/2005.11879, 2020.
  17. 17.Fischer, K., Ren’e, A., Keup, C., Layer, M., Dahmen, D., and Helias, M. Decomposing neural networks as mappings of correlation functions. Physical Review Research, 2022.
  18. 18.Gabrie, M. Mean-field inference methods for neural networks. Journal of Physics A: Mathematical and Theoretical, 53, 2019.
  19. 19.Gerace, F., Loureiro, B., Krzakala, F., Mezard, M., and Zdeborova, L. Generalisation error in learning with random features and the hidden manifold model. Proceedings of Machine Learning Research, 37:3452–3462, 2020.
  20. 20.Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A. Linearized two-layers neural networks in high dimension. ArXiv, abs/1904.12191, 2019.
  21. 21.Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A. When do neural networks outperform kernel methods? Journal of Statistical Mechanics: Theory and Experiment, 2021, 2020.
  22. 22.Goldt, S., Mezard, M., Krzakala, F., and Zdeborova, L. Modelling the influence of data structure on learning in neural networks. Physical Review X, 4:041044, 2020.
  23. 23.Goldt, S., Loureiro, B., Reeves, G., Krzakala, F., Mezard, M., and Zdeborova, L. The gaussian equivalence of generative models for learning with shallow neural networks. In MSML, 2021.
  24. 24.Goldt, S., Loureiro, B., Reeves, G., Krzakala, F., Mezard, M., and Zdeborova, L. The gaussian equivalence of generative models for learning with shallow neural networks. In Mathematical and Scientific Machine Learning, pp. 426–471. PMLR, 2022.
  25. 25.Hanin, B. Correlation functions in random fully connected neural networks at finite width. ArXiv, abs/2204.01058, 2022.
  26. 26.Hanin, B. and Zlokapa, A. Bayesian interpolation with deep linear networks. ArXiv, abs/2212.14457, 2022.
  27. 27.Hastie, T. J., Montanari, A., Rosset, S., and Tibshirani, R. J. Surprises in high-dimensional ridgeless least squares interpolation. Annals of statistics, 50 2:949–986, 2019.
  28. 28.Hron, J., Bahri, Y., Novak, R., Pennington, J., and Sohl-Dickstein, J. N. Exact posterior distributions of wide bayesian neural networks. ArXiv, abs/2006.10541, 2020.
  29. 29.Hu, H. and Lu, Y. M. Universality laws for high-dimensional learning with random features. IEEE Transactions on Information Theory, 2022a.
  30. 30.Hu, H. and Lu, Y. M. Sharp asymptotics of kernel ridge regression beyond the linear regime. ArXiv, abs/2205.06798, 2022b.
  31. 31.Iba, Y. The nishimori line and bayesian statistics. Journal of Physics A, 32:3875–3888, 1998.
  32. 32.Lee, J., Bahri, Y., Novak, R., Schoenholz, S. S., Pennington, J., and Sohl-Dickstein, J. N. Deep neural networks as gaussian processes. International Conference on Learning Representations, 2018.
  33. 33.Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Novak, R., Sohl-Dickstein, J. N., and Pennington, J. S. Wide neural networks of any depth evolve as linear models under gradient descent. Journal of Statistical Mechanics: Theory and Experiment, 2020, 2019.
  34. 34.Lee, J., Schoenholz, S. S., Pennington, J., Adlam, B., Xiao, L., Novak, R., and Sohl-Dickstein, J. N. Finite versus infinite neural networks: an empirical study. NeurIPS, 2020.
  35. 35.Li, Q. and Sompolinsky, H. Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization. Physical Review X, 2021.
  36. 36.Louart, C., Liao, Z., and Couillet, R. A random matrix approach to neural networks. The Annals of Applied Probability, 28:1190–1248, 2017.
  37. 37.Loureiro, B., Gerbelot, C., Cui, H., Goldt, S., Krzakala, F., Mezard, M., and Zdeborova, L. Learning curves of generic features maps for realistic datasets with a teacher-student model. Advances in Neural Information Processing Systems, 34:18137–18151, 2021.
  38. 38.Maillard, A., Loureiro, B., Krzakala, F., and Zdeborova, L. Phase retrieval in high dimensions: Statistical and computational phase transitions. Advances in Neural Information Processing Systems, 33:11071–11082, 2020.
  39. 39.Mei, S. and Montanari, A. The generalization error of random features regression: Precise asymptotics and the double descent curve. Communications on Pure and Applied Mathematics, 75, 2019.
  40. 40.Mei, S., Misiakiewicz, T., and Montanari, A. Generalization error of random feature and kernel methods: hypercontractivity and kernel matrix concentration. Applied and Computational Harmonic Analysis, 2021.
  41. 41.Mezard, M. and Montanari, A. Information,Physics and computation. Oxford University Press, 2002.
  42. 42.Misiakiewicz, T. Spectrum of inner-product kernel matrices in the polynomial regime and multiple descent phenomenon in kernel ridge regression. ArXiv, abs/2204.10425, 2022.
  43. 43.Montanari, A. and Saeed, B. Universality of empirical risk minimization. Conference on Learning Theory, pp. 4310–4312, 2022.
  44. 44.Montanari, A., Ruan, F., Sohn, Y., and Yan, J. The generalization error of max-margin linear classifiers: High-dimensional asymptotics in the overparametrized regime. ArXiv, abs/1911.01544, 2019.
  45. 45.Neal, R. M. Priors for infinite networks (tech. rep. no. crg-tr-94-1). University of Toronto, 1994.
  46. 46.Nishimori, H. Statistical Physics of Spin Glasses and Information Processing. Oxford:Clarendon, 2001.
  47. 47.Opper and Haussler. Generalization performance of bayes optimal classification algorithm for learning a perceptron. Physical review letters, 66 20:2677–2680, 1991.
  48. 48.Parisi, G. Towards a mean field theory for spin glasses. Phys. Lett, 73(A):203–205, 1979.
  49. 49.Parisi, G. Order parameter for spin glasses. Phys. Rev. Lett, 50:1946–1948, 1983.
  50. 50.Pennington, J. and Worah, P. Nonlinear random matrix theory for deep learning. Journal of Statistical Mechanics: Theory and Experiment, 2019, 2019.
  51. 51.Rahimi, A. and Recht, B. Random features for large-scale kernel machines. In NIPS, 2007.
  52. 52.Roberts, D. A., Yaida, S., and Hanin, B. The principles of deep learning theory. ArXiv, abs/2106.10165, 2021.
  53. 53.Sahraee-Ardakan, M., Emami, M., Pandit, P., Rangan, S., and Fletcher, A. K. Kernel methods and multi-layer perceptrons learn linear models in high dimensions. ArXiv, abs/2201.08082, 2022.
  54. 54.Schroder, D., Cui, H., Dmitriev, D., and Loureiro, B. Deterministic equivalent and error universality of deep random features learning. ArXiv, abs/2302.00401, 2023.
  55. 55.Schwarze, H. Learning a rule in a multilayer neural network. Journal of Physics A: Mathematical and General, 26(21):5781, 1993.
  56. 56.Seung, H. S., Sompolinsky, H., and Tishby, N. Statistical mechanics of learning from examples. Physical review A, 45(8):6056, 1992.
  57. 57.Talagrand, M. The parisi formula. Annals of mathematics, pp. 221–263, 2006.
  58. 58.Thrampoulidis, C., Abbasi, E., and Hassibi, B. Precise error analysis of regularized m-estimators in high dimensions. IEEE Transactions on Information Theory, 64(8):5592–5628, 2018.
  59. 59.Watkin, T. L., Rau, A., and Biehl, M. The statistical mechanics of learning a rule. Reviews of Modern Physics, 65(2):499, 1993.
  60. 60.Wu, D. and Xu, J. On the optimal weighted ℓ2 regularization in overparameterized linear regression. Advances in Neural Information Processing Systems, 33:10112–10123, 2020.
  61. 61.Xiao, L. and Pennington, J. Precise learning curves and higher-order scaling limits for dot product kernel regression. ArXiv, abs/2205.14846, 2022.
  62. 62.Yaida, S. Non-gaussian processes and neural networks at finite widths. In Mathematical and Scientific Machine Learning, 2019.
  63. 63.Zavatone-Veth, J. A., Canatar, A., and Pehlevan, C. Asymptotics of representation learning in finite bayesian neural networks. Journal of Statistical Mechanics: Theory and Experiment, 2022, 2021.
  64. 64.Zavatone-Veth, J. A., Tong, W., and Pehlevan, C. Contrasting random and learned features in deep bayesian linear regression. Physical review. E, 105 6-1:064118, 2022.
  65. 65.Zdeborova, L. and Krzakala, F. Statistical physics of inference: thresholds and algorithms. Advances in Physics, 65:453 – 552, 2015.

Citation

MLA
Cui, H., et al. “Bayes-optimal Learning of Deep Random Networks of Extensive-width”. International Conference on Machine Learning, vol. 202, 2023, pp. 6468–521, https://proceedings.mlr.press/v202/cui23b.html.
APA
Cui, H., Krzakala, F., & Zdeborova, L. (2023). Bayes-optimal Learning of Deep Random Networks of Extensive-width. International Conference on Machine Learning, 202, 6468–6521. https://proceedings.mlr.press/v202/cui23b.html
Chicago
Cui, H., F. Krzakala, and L. Zdeborova. 2023. “Bayes-optimal Learning of Deep Random Networks of Extensive-width”. International Conference on Machine Learning 202: 6468–6521. https://proceedings.mlr.press/v202/cui23b.html.
Harvard
Cui, H., Krzakala, F. and Zdeborova, L. (2023) “Bayes-optimal Learning of Deep Random Networks of Extensive-width”, International Conference on Machine Learning. PMLR, pp. 6468–6521. Available at: https://proceedings.mlr.press/v202/cui23b.html.
Vancouver
1. Cui H, Krzakala F, Zdeborova L (2023) Bayes-optimal Learning of Deep Random Networks of Extensive-width. In: International Conference on Machine Learning. PMLR, pp 6468–6521

BibTeX

@InProceedings{pmlr-v202-cui23b,
  title = 	 {{B}ayes-optimal Learning of Deep Random Networks of Extensive-width},
  author =       {Cui, Hugo and Krzakala, Florent and Zdeborova, Lenka},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {6468--6521},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/cui23b/cui23b.pdf},
  url = 	 {https://proceedings.mlr.press/v202/cui23b.html},
  abstract = 	 {We consider the problem of learning a target function corresponding to a deep, extensive-width, non-linear neural network with random Gaussian weights. We consider the asymptotic limit where the number of samples, the input dimension and the network width are proportionally large and propose a closed-form expression for the Bayes-optimal test error, for regression and classification tasks. We further compute closed-form expressions for the test errors of ridge regression, kernel and random features regression. We find, in particular, that optimally regularized ridge regression, as well as kernel regression, achieve Bayes-optimal performances, while the logistic loss yields a near-optimal test error for classification. We further show numerically that when the number of samples grows faster than the dimension, ridge and kernel methods become suboptimal, while neural networks achieve test error close to zero from quadratically many samples.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/