Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Xiang ChengYuxin ChenSuvrit Sra

article2024ICML75 citations

Proves both theoretically and empirically that non-linear Transformers learn to execute functional gradient descent during in-context learning, converging to Bayes-optimal predictors when their attention activations match the underlying data distribution.

Listen

Modern transformer models exhibit a remarkable ability to learn from prompt demonstrations without updating their weights, a capability known as in-context learning. Prior theoretical studies explained this behavior in simplified linear settings by showing that transformers internally execute standard gradient descent on linear tasks. However, real-world applications rely heavily on nonlinear activation functions—such as softmax and rectified linear units—and process complex, nonlinear data distributions. Understanding the algorithmic mechanics that enable nonlinear transformers to learn complex functions in context has remained an open challenge.

The article investigates the algorithmic mechanisms implemented by nonlinear transformers and determines how they successfully learn nonlinear functions in context. Specifically, the authors evaluate whether transformer forward passes can execute optimization algorithms in function space and assess whether these mechanisms naturally emerge during standard model training.

To address these questions, the authors combine rigorous mathematical analysis with controlled empirical simulations. They formulate attention modules with arbitrary nonlinear activations and analyze data generated by generalized nonlinear processes, such as Gaussian processes. The theoretical work characterizes the loss landscape and stationary points of multi-layer transformers during in-context training, while the empirical evaluations track parameter convergence across various architectures and task types.

The findings establish that nonlinear transformers naturally implement functional gradient descent—an optimization method that updates predictive functions directly in a reproducing kernel Hilbert space. First, when a transformer's nonlinear activation matches the kernel governing the underlying data distribution, its layer-by-layer forward pass converges to the Bayes-optimal predictor as the number of layers increases. Second, mathematical analysis shows that functional gradient descent represents an exact stationary point of the in-context training loss, and optimization experiments confirm that standard training consistently drives the model parameters toward this configuration. Third, in unconstrained value-matrix settings, the transformer learns an advanced algorithm that alternates between transforming input covariates and taking functional gradient descent steps. Fourth, multi-head attention architectures with diverse activations can learn complex composite kernels, matching Bayes-optimal accuracy across varied function classes.

These results provide a solid mathematical foundation for the empirical success of transformers, establishing that they operate as principled meta-optimizers for nonlinear relationships rather than simple pattern matchers. The insights directly inform architectural design, demonstrating that the optimal choice of activation function is dictated by the functional structure of the target data. This understanding reduces the empirical trial-and-error traditionally required when configuring attention mechanisms for domain-specific tasks.

Organizations developing or applying transformer architectures should align activation choices with the data domain and consider multi-head designs with diverse activations to enhance expressive power across heterogeneous tasks. Further research should focus on extending global optimality guarantees for training dynamics, analyzing the exact algorithmic benefits of sequential layer composition, and conducting larger-scale empirical pilots on real-world datasets.

While the theoretical guarantees and controlled experiments provide high confidence in these mechanisms, the analysis relies on specific distributional assumptions regarding input symmetry and parameter structures. Stakeholders should note that performance on complex, real-world data distributions may introduce additional variables not fully captured by idealized Gaussian process settings.

  • Paper: Transformers Learn In-Context by Gradient Descent, Johannes von Oswald et al. (2023). This earlier analysis shows how transformers implement gradient descent on linear tasks, providing the linear in-context learning foundation that the source generalizes to nonlinear functions.

No sufficiently relevant recommendations were found.

Cover for Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Abstract

Many neural network architectures are known to be Turing Complete, and can thus, in principle implement arbitrary algorithms. However, Transformers are unique in that they can implement gradient-based learning algorithms under simple parameter configurations. This paper provides theoretical and empirical evidence that (non-linear) Transformers naturally learn to implement gradient descent in function space, which in turn enable them to learn non-linear functions in context. Our results apply to a broad class of combinations of non-linear architectures and non-linear in-context learning tasks. Additionally, we show that the optimal choice of non-linear activation depends in a natural way on the class of functions that need to be learned.

Table of Contents

  • 1. Introduction
  • 1.1. Summary of Contributions
  • 1.2. Related Work
  • 2. Setup: ICL with non-linear Transformers
  • Input Data for In-Context Learning
  • Transformers with general non-linear attention
  • The In-Context Loss
  • 2.1. Examples of Attention Modules
  • 3. Transformers can implement gradient descent in function space.
  • 3.1. Transformers can implement gradient descent in function space.
  • 3.1.1. CASE STUDY: LINEAR KERNEL
  • 3.1.2. CASE STUDY: EXPONENTIAL KERNEL, AND CONNECTION TO SOFTMAX ACTIVATION
  • 3.2. Optimality of ˜h for matching K.
  • 3.3. Experiments for Proposition 3.4
  • 3.4. Composing multiple attention heads
  • 4. Optimization Landscape Results
  • 4.1. Distributional Assumptions
  • 4.2. Architectural Assumptions
  • 4.3. Theorem 4.5: Functional gradient descent is a stationary point of (constrained) in-context loss.
  • 4.4. Theorem 4.6: Characterizing the stationary points of unconstrained in-context loss.
  • 4.5. Experiments for Theorems 4.5 and 4.6
  • 5. Future Directions
  • Acknowledgement
  • Impact Statement
  • References
  • A. Reformulating the In-Context Loss
  • B. Proof of Proposition 3.1
  • C. Proof of Proposition 3.4
  • D. Composing multiple attention heads with different activations
  • D.1. Proof of Proposition D.1
  • E. Theorem E.1: Functional Gradient Descent is locally optimal under A ℓ = 0 constraint.
  • E.1. Proof of Theorem E.1
  • E.2. Key Lemmas
  • F. Theorem F.1: characterizing local optimum when A ℓ are unconstrained.
  • F.1. Proof of Theorem F.1
  • F.2. Key Lemmas
  • G. Background on RKHS
  • H. Experiments
  • H.1. Experiment Details
  • Covariate Distribution
  • Label Distribution
  • Transformer Architecture
  • Training Algorithm
  • H.2. Experiment for Theorem 4.5
  • H.3. Experiments for Theorem 4.6
  • I. Miscellaneous Proofs
  • I.1. Verification of Example 8
  • I.2. Functional Gradient Descent for Euclidean Inner Product Kernel

Knowls

  1. Knowl 1 — A Transformer attention layer can implement RKHS functional gradient descent

    model/method

    Let nn demonstrations be (x(i),y(i))(x^{(i)},y^{(i)}), with x(i)∈Rdx^{(i)}\in\mathbb{R}^d and scalar y(i)y^{(i)}, and let xx be a query. For a positive-semidefinite kernel KK and its reproducing-kernel Hilbert space (RKHS) H\mathcal{H}, consider the empirical squared loss L(f)=∑i=1n(f(x(i))−y(i))2L(f)=\sum_{i=1}^n(f(x^{(i)})-y^{(i)})^2. Starting at f0=0f_0=0, RKHS gradient descent has updates fℓ+1(x)=fℓ(x)+rℓ∑i=1n(y(i)−fℓ(x(i)))K(x(i),x)f_{\ell+1}(x)=f_\ell(x)+r_\ell\sum_{i=1}^n(y^{(i)}-f_\ell(x^{(i)}))K(x^{(i)},x), where rℓr_\ell is a scalar step size. A residual Transformer with attention activation [h~(U,W)]ij=K(U:,i,W:,j)[\tilde h(U,W)]_{ij}=K(U_{:,i},W_{:,j}), query and key matrices Bℓ=Cℓ=IdB_\ell=C_\ell=I_d, and value matrix Vℓ=[000−rℓ]V_\ell=\begin{bmatrix}0&0\\0&-r_\ell\end{bmatrix} implements these iterates layer by layer. Here U,W∈Rd×(n+1)U,W\in\mathbb{R}^{d\times(n+1)} contain the demonstration covariates and query, and the Transformer readout uses the paper's convention of storing the negative of the function estimate. Thus a single layer's attention update adds kernel-weighted training residuals, and the construction applies to any kernel whose pairwise values can be used as the attention activation.

  2. Knowl 2 — Kernel-matched functional descent converges to the Bayes predictor for Gaussian-process labels

    theoretical result

    Fix covariates x(1),…,x(n+1)x^{(1)},\ldots,x^{(n+1)}, where x(n+1)x^{(n+1)} is the query, and let KK be a positive-semidefinite kernel. Suppose the labels are conditionally jointly Gaussian with mean zero and covariance matrix KXK_X, whose entries are (KX)ij=K(x(i),x(j))(K_X)_{ij}=K(x^{(i)},x^{(j)}). A Transformer that implements RKHS functional gradient descent using the same kernel KK approaches the Bayes conditional-mean predictor as its number of layers tends to infinity. In particular, if K^∈Rn×n\widehat K\in\mathbb{R}^{n\times n} is the training Gram matrix, ν∈Rn\nu\in\mathbb{R}^n has entries νi=K(x(i),x(n+1))\nu_i=K(x^{(i)},x^{(n+1)}), and K^\widehat K is invertible, the Bayes prediction is ν⊤K^−1y^\nu^\top\widehat K^{-1}\widehat y, where y^=(y(1),…,y(n))⊤\widehat y=(y^{(1)},\ldots,y^{(n)})^\top. The result is asymptotic in depth: it does not claim that kernel matching is best at every finite depth.

  3. Knowl 3 — Multi-head attention implements descent for sums of transformed kernels

    model/method

    Let a Transformer have HH attention heads. For head ss, choose a positive-semidefinite kernel KsK_s and a matrix Gs∈Rd×dG_s\in\mathbb{R}^{d\times d}, and define the composite kernel K⋄(u,v)=∑s=1HKs(Gsu,Gsv)K_\diamond(u,v)=\sum_{s=1}^H K_s(G_su,G_sv), requiring K⋄K_\diamond itself to be positive semidefinite. Give head ss the activation [h~s(U,W)]ij=Ks(U:,i,W:,j)[\tilde h_s(U,W)]_{ij}=K_s(U_{:,i},W_{:,j}), query and key matrices Bℓs=Cℓs=GsB_\ell^s=C_\ell^s=G_s, and a value matrix with only its bottom-right entry nonzero, set to the negative of that head's scalar step size. With compatible head step sizes, the summed attention updates implement functional gradient descent in the RKHS induced by K⋄K_\diamond. If labels are conditionally Gaussian with covariance given by the Gram matrix of K⋄K_\diamond, the resulting prediction approaches the Bayes predictor as depth tends to infinity. This construction lets different heads represent different kernel components, including components that use different feature transformations.

  4. Knowl 4 — With zero covariate updates, the kernel-descent parameter family is stationary under symmetry assumptions

    theoretical result

    Consider the expected in-context squared prediction loss for a residual Transformer whose value matrices are block diagonal, Vℓ=[Aℓ00rℓ]V_\ell=\begin{bmatrix}A_\ell&0\\0&r_\ell\end{bmatrix}. Assume covariates have a rotationally invariant distribution after rescaling by a symmetric positive-definite matrix Σ\Sigma: for every orthogonal UU, Σ1/2UΣ−1/2X\Sigma^{1/2}U\Sigma^{-1/2}X has the same distribution as XX. Assume also that the conditional label second-moment matrix is invariant under the same transformations, and that the attention activation obeys h~(W,V)=h~(S⊤W,S−1V)\tilde h(W,V)=\tilde h(S^\top W,S^{-1}V) for every invertible SS. When the covariate blocks are constrained to Aℓ=0A_\ell=0, the parameter family Bℓ=bℓΣ−1/2B_\ell=b_\ell\Sigma^{-1/2} and Cℓ=cℓΣ−1/2C_\ell=c_\ell\Sigma^{-1/2}, with scalar bℓ,cℓ,rℓb_\ell,c_\ell,r_\ell, has zero infimum of the in-context-loss gradient norm over that family. The paper interprets this family as a stationary solution. When [h~(U,W)]ij=K(U:,i,W:,j)[\tilde h(U,W)]_{ij}=K(U_{:,i},W_{:,j}) for a kernel KK, its forward-pass update is functional gradient descent for the rescaled kernel K~(u,v)=K(Σ−1/2u,Σ−1/2v)\widetilde K(u,v)=K(\Sigma^{-1/2}u,\Sigma^{-1/2}v). The result establishes stationarity under the stated symmetry and parameter restrictions; it does not establish global optimality.

  5. Knowl 5 — Allowing scalar identity covariate updates yields stationary points that interleave feature transformations and descent

    theoretical result

    Under the same distributional and attention-invariance assumptions as the rotationally symmetric Transformer setting, leave the covariate blocks of the block-diagonal value matrices unconstrained. The in-context loss has a stationary solution family with Aℓ=aℓIdA_\ell=a_\ell I_d, Bℓ=bℓΣ−1/2B_\ell=b_\ell\Sigma^{-1/2}, and Cℓ=cℓΣ−1/2C_\ell=c_\ell\Sigma^{-1/2} for scalar aℓ,bℓ,cℓa_\ell,b_\ell,c_\ell. The formal result is that the infimum of the squared gradient norm over this restricted family is zero. Its forward pass transforms the entire covariate matrix at each layer according to Xℓ+1=Xℓ+aℓXℓMh~(bℓΣ−1/2Xℓ,cℓΣ−1/2Xℓ)X_{\ell+1}=X_\ell+a_\ell X_\ell M\tilde h(b_\ell\Sigma^{-1/2}X_\ell,c_\ell\Sigma^{-1/2}X_\ell), where M=diag⁡(1,…,1,0)M=\operatorname{diag}(1,\ldots,1,0) masks the query-label position. The label row is updated by the same masked attention weights, scaled by rℓr_\ell. When the activation is a kernel, the layerwise predictions therefore interleave covariate transformations with kernel-based functional-gradient updates. The theorem identifies stationary structure, not a convergence guarantee or a global optimum.

  6. Knowl 6 — Matching the attention activation to the label-generating kernel usually gives the best test loss

    empirical result

    The experiments compared linear, ReLU, exponential, and masked-softmax attention activations on labels sampled from Gaussian processes with linear, ReLU, or exponential kernels. Covariates were sampled from the unit sphere, and test in-context loss was measured after training had converged, with the Bayes predictor used as a reference. Across the tested context lengths and layer depths, the lowest loss was generally obtained when the attention activation matched the kernel generating the labels; the paper reports this pattern for the linear and ReLU cases and for deeper exponential-kernel models. This supports the prediction that kernel-matched functional descent approaches the Bayes predictor with increasing depth. Softmax attention is an important exception for exponential-kernel data in shallow or long-context settings, so the empirical result is a general trend rather than a universal ranking.

  7. Knowl 7 — Softmax can outperform exponential-kernel attention at shallow depth on exponential-kernel data

    empirical result

    For exponential-kernel labels, the paper's softmax-attention update is similar to the unnormalized exponential-kernel update but multiplies the kernel-weighted residual sum at query xx by τ(x)=1/∑j=1nK(x,x(j))\tau(x)=1/\sum_{j=1}^n K(x,x^{(j)}). In three-layer comparisons, softmax attention had lower test loss than exponential-kernel attention for context lengths n∈{6,8,10,12,14}n\in\{6,8,10,12,14\}. At n=14n=14, their losses became close by depth 66 or greater; at n=6n=6, exponential-kernel attention was best once depth reached at least 55. The authors conjecture that softmax implements a more iteration-efficient but less statistically efficient algorithm: its normalization may help with few layers, while the kernel-matched exponential update improves with more layers or fewer demonstrations. This explanation is a conjecture, not a proved mechanism.

  8. Knowl 8 — Training empirically moves attention parameters toward the predicted stationary structure

    empirical result

    The authors trained three-layer Transformers using Adam with gradient clipping on in-context tasks with n=30n=30 demonstrations, comparing linear, ReLU, and softmax attention against linear-, ReLU-, and exponential-kernel Gaussian-process labels. The covariate dimension was d=5d=5; experiments used three runs with different covariance matrices and data-sampling seeds. Minibatches contained 30,00030{,}000 examples and were resampled every 1010 optimization steps. With covariate updates constrained to zero, the rescaled products Σ1/2Bℓ⊤CℓΣ1/2\Sigma^{1/2}B_\ell^\top C_\ell\Sigma^{1/2} approached scaled identity matrices in most tested settings. With covariate updates allowed, those products and the covariate blocks AℓA_\ell generally approached scaled identities, as predicted by the corresponding stationary-point characterizations. Some runs remained about 0.20.2 normalized Frobenius distance from identity; the authors could not determine whether this reflected optimization difficulty or convergence to a different stationary point. These experiments support, but do not prove, that training finds the characterized structures.

  9. Knowl 9 — The landscape assumptions include nonlinear labels from random two-layer ReLU networks

    assumption

    The stationary-point results apply when the joint covariate distribution is invariant under transformations X↦Σ1/2UΣ−1/2XX\mapsto\Sigma^{1/2}U\Sigma^{-1/2}X for orthogonal UU, and the conditional second moment of the labels shares this invariance. The covariates need not be independent: examples include Gaussian or sphere-based rotationally invariant samples after rescaling, as well as certain Gaussian-mixture constructions. A nonlinear label process covered by the assumptions is y(i)=⟨θ2,ReLU⁡(θ1x(i))⟩y^{(i)}=\langle\theta_2,\operatorname{ReLU}(\theta_1x^{(i)})\rangle, where θ1∈Rd×m\theta_1\in\mathbb{R}^{d\times m} and θ2∈Rm\theta_2\in\mathbb{R}^m have independent standard-normal entries. Its conditional label second moments are invariant under common orthogonal rotations because the random first-layer weights are rotationally invariant. Thus the analysis is not restricted to linear regression or to labels generated by a fixed kernel process.

  10. Knowl 10 — The theory does not guarantee global optimality or that training converges to the characterized points

    limitation

    The kernel-matching result guarantees Bayes-optimal prediction only in the infinite-depth limit; at finite depth, another attention activation may achieve lower loss by implementing a more iteration-efficient procedure. The optimization-landscape theorems identify stationary structure under symmetry and parameter restrictions, but do not prove that the characterized points are global minima or that gradient-based training converges to them. The training experiments provide empirical evidence of parameter movement toward the predicted structures, with some runs remaining measurably distant or possibly reaching other stationary points. The paper also leaves open the interpretation and benefits of the covariate-transforming algorithm in the unconstrained case.

Coverage note — The paper's exact mapping from masked linear, ReLU, and softmax attention to the generalized attention notation is omitted as architectural setup rather than an independent result; the main constructions, landscape claims, experiments, assumptions, and stated limitations are represented.

References

  1. 1.Ahn, K., Cheng, X., Daneshmand, H., and Sra, S. Transformers learn to implement preconditioned gradient descent for in-context learning. arXiv preprint arXiv:2306.00297, 2023.
  2. 2.Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D. What learning algorithm is in-context learning? investigations with linear models. International Conference on Learning Representations, 2022.
  3. 3.Ali, A., Touvron, H., Caron, M., Bojanowski, P., Douze, M., Joulin, A., Laptev, I., Neverova, N., Synnaeve, G., Verbeek, J., et al. Xcit: Cross-covariance image transformers. Advances in neural information processing systems, 34: 20014–20027, 2021.
  4. 4.Bai, Y., Chen, F., Wang, H., Xiong, C., and Mei, S. Transformers as statisticians: Provable in-context learning with in-context algorithm selection. arXiv preprint arXiv:2306.04637, 2023.
  5. 5.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Neural Information Processing Systems, 2020.
  6. 6.Chen, Y., Tao, Q., Tonin, F., and Suykens, J. A. Primal-attention: Self-attention through asymmetric kernel svd in primal representation. arXiv preprint arXiv:2305.19798, 2023.
  7. 7.Chi, T.-C., Fan, T.-H., Ramadge, P. J., and Rudnicky, A. Kerple: Kernelized relative positional embedding for length extrapolation. Advances in Neural Information Processing Systems, 35:8386–8399, 2022.
  8. 8.Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020.
  9. 9.Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F. Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers. arXiv preprint arXiv:2212.10559, 2022.
  10. 10.Garg, S., Tsipras, D., Liang, P. S., and Valiant, G. What can transformers learn in-context? a case study of simple function classes. Advances in Neural Information Processing Systems, 35:30583–30598, 2022.
  11. 11.Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D. Looped transformers as programmable computers. arXiv preprint arXiv:2301.13196, 2023.
  12. 12.Huang, Y., Cheng, Y., and Liang, Y. In-context convergence of transformers. arXiv preprint arXiv:2310.05249, 2023.
  13. 13.Lin, L., Bai, Y., and Mei, S. Transformers as decision makers: Provable in-context reinforcement learning via supervised pretraining. arXiv preprint arXiv:2310.08566, 2023.
  14. 14.Mahankali, A., Hashimoto, T. B., and Ma, T. One step of gradient descent is provably the optimal in-context learner with one layer of linear self-attention. arXiv preprint arXiv:2307.03576, 2023.
  15. 15.Nguyen, T., Pham, M., Nguyen, T., Nguyen, K., Osher, S., and Ho, N. Fourierformer: Transformer meets generalized fourier integral theorem. Advances in Neural Information Processing Systems, 35:29319–29335, 2022a.
  16. 16.Nguyen, T. M., Nguyen, T. M., Le, D. D., Nguyen, D. K., Tran, V.-A., Baraniuk, R., Ho, N., and Osher, S. Improving transformers with probabilistic attention keys. In International Conference on Machine Learning, pp. 16595–16621. PMLR, 2022b.
  17. 17.Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, K., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. In-context learning and induction heads. Transformer Circuits Thread, 2022.
  18. 18.Pérez, J., Barceló, P., and Marinkovic, J. Attention is turing complete. The Journal of Machine Learning Research, 2021.
  19. 19.Schlag, I., Irie, K., and Schmidhuber, J. Linear transformers are secretly fast weight programmers. In International Conference on Machine Learning, pp. 9355–9366. PMLR, 2021.
  20. 20.Schölkopf, B., Herbrich, R., and Smola, A. J. A generalized representer theorem. In International conference on computational learning theory, pp. 416–426. Springer, 2001.
  21. 21.Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R. Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel. arXiv preprint arXiv:1908.11775, 2019.
  22. 22.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 2017.
  23. 23.von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. Transformers learn in-context by gradient descent. In International Conference on Machine Learning, pp. 35151–35174. PMLR, 2023a.
  24. 24.von Oswald, J., Niklasson, E., Schlegel, M., Kobayashi, S., Zucchet, N., Scherrer, N., Miller, N., Sandler, M., Vladymyrov, M., Pascanu, R., et al. Uncovering mesa-optimization algorithms in transformers. arXiv preprint arXiv:2309.05858, 2023b.
  25. 25.Wainwright, M. J. High-dimensional statistics: A non-asymptotic viewpoint, volume 48. Cambridge university press, 2019.
  26. 26.Wei, C., Chen, Y., and Ma, T. Statistically meaningful approximation: a case study on approximating turing machines with transformers. Advances in Neural Information Processing Systems, 35:12071–12083, 2022.
  27. 27.Wortsman, M., Lee, J., Gilmer, J., and Kornblith, S. Replacing softmax with relu in vision transformers. arXiv preprint arXiv:2309.08586, 2023.
  28. 28.Wright, M. A. and Gonzalez, J. E. Transformers are deep infinite-dimensional non-mercer binary kernel machines. arXiv preprint arXiv:2106.01506, 2021.
  29. 29.Wu, J., Zou, D., Chen, Z., Braverman, V., Gu, Q., and Bartlett, P. L. How many pretraining tasks are needed for in-context learning of linear regression? arXiv preprint arXiv:2310.08391, 2023.
  30. 30.Zhang, R., Frei, S., and Bartlett, P. L. Trained transformers learn linear models in-context. arXiv preprint arXiv:2306.09927, 2023.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/