A Kernel-Based View of Language Model Fine-Tuning

Sadhika MalladiAlexander WettigDingli YuDanqi ChenSanjeev Arora

article2023ICML133 citations

Extends Neural Tangent Kernel theory to Adam optimizer dynamics and pre-trained initializations, explaining why prompt-based fine-tuning and parameter-efficient adaptation methods succeed in low-data regimes without overfitting.

Listen

Adapting large pre-trained language models to downstream tasks using very small datasets has become standard practice across language processing applications. However, a foundational theoretical explanation for why massive models with hundreds of millions of parameters do not severely overfit in these low-data regimes has remained elusive. Furthermore, practitioners frequently observe that fine-tuning success is highly sensitive to implementation choices, such as using natural language prompts or restricting parameter updates to low-rank subspaces, without a formal understanding of why these techniques work.

The article aims to evaluate whether the Neural Tangent Kernel framework—a mathematical lens originally developed to study the optimization of infinitely wide, randomly initialized neural networks—can describe and explain the fine-tuning of pre-trained language models. Specifically, the authors seek to determine the precise conditions under which fine-tuning behaves like kernel regression, characterize the dynamics of standard optimizers such as Adam, and explain the success of parameter-efficient fine-tuning methods.

To investigate these questions, the authors mathematically extended kernel theory to account for non-random, pre-trained initializations and derived a new sign-based kernel formulation that captures early-stage training dynamics under adaptive optimizers like Adam. They then evaluated this theoretical framework empirically across 14 diverse natural language understanding tasks—including sentiment classification, topic identification, natural language inference, and paraphrase detection—using few-shot datasets (16 and 64 training examples per class) and a pre-trained RoBERTa model. The empirical testing examined whether fine-tuning trajectories satisfy two defining criteria of kernel behavior: function linearization and fixed gradient features.

The investigation produced four central findings. First, formulating downstream tasks via natural language prompts is critical for inducing kernel behavior; prompt-based empirical kernels closely matched actual fine-tuning performance, whereas standard classification heads without prompts exhibited performance deficits of up to 16 percentage points. Second, the empirical kernel accurately solved 12 out of the 14 downstream tasks, matching fine-tuning performance within a 10% margin, with 8 tasks strictly satisfying all mathematical conditions of kernel behavior throughout training. Third, standard stochastic gradient descent performed within 4 percentage points of Adam in prompt-based settings, confirming that prompted optimization landscapes are sufficiently benign that optimizer differences diminish. Fourth, the kernel perspective mathematically explains the effectiveness of low-rank adaptation methods, proving that random subspace projections preserve the underlying kernel matrix and deliver accuracy comparable to full-parameter tuning.

These findings imply that successful prompt-based fine-tuning requires only minimal, linear parameter adjustments rather than complex feature re-learning. This mechanism explains why large language models generalize effectively from only a few dozen examples without catastrophic overfitting. The results also validate parameter-efficient adaptation strategies like low-rank adaptation as mathematically sound alternatives to full fine-tuning, offering substantial computational and storage cost savings with minimal risk to downstream accuracy.

Practitioners and decision-makers should prioritize natural, well-formatted prompt designs when adapting pre-trained models to ensure stable optimization within the kernel regime. Teams seeking to reduce computing overhead should confidently adopt low-rank adaptation techniques for prompt-based workflows. For complex tasks where standard prompts struggle—such as intricate textual entailment—organizations should conduct pilot evaluations and refine prompt phrasing before deploying fine-tuned models to production.

Confidence in these findings is high for few-shot classification using masked language models, supported by rigorous proofs and broad empirical testing across multiple task benchmarks. However, key limitations remain: the empirical validation focuses on the RoBERTa architecture, the theoretical guarantees for Adam apply primarily to early-stage training dynamics, and the kernel approach does not fully capture performance on tasks where prompt phrasing is unnatural. Further research is needed to extend this framework to generative decoder models and longer training durations.

No sufficiently relevant recommendations were found.

Cover for A Kernel-Based View of Language Model Fine-Tuning

Abstract

It has become standard to solve NLP tasks by fine-tuning pre-trained language models (LMs), especially in low-data settings. There is minimal theoretical understanding of empirical success, e.g., why fine-tuning a model with 10⁸ or more parameters on a couple dozen training points does not result in overfitting. We investigate whether the Neural Tangent Kernel (NTK)—which originated as a model to study the gradient descent dynamics of infinitely wide networks with suitable random initialization—describes fine-tuning of pre-trained LMs. This study was inspired by the decent performance of NTK for computer vision tasks (Wei et al., 2022). We extend the NTK formalism to Adam and use Tensor Programs (Yang, 2020b) to characterize conditions under which the NTK lens may describe fine-tuning updates to pre-trained language models. Extensive experiments on 14 NLP tasks validate our theory and show that formulating the downstream task as a masked word prediction problem through prompting often induces kernel-based dynamics during fine-tuning. Finally, we use this kernel view to propose an explanation for the success of parameter-efficient subspace-based fine-tuning methods.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 3.1. Pre-Training and Fine-Tuning Paradigm
  • 3.2. Kernel Behavior
  • 4. Kernel Derivation for Adam
  • 5. Theory: Prompt-Based Fine-Tuning Can Exhibit Kernel Behavior
  • 6. Experiments
  • 6.1. Kernel Performance on Downstream Tasks
  • 6.2. Measuring Kernel Behavior
  • 6.3. Tasks without Kernel Behavior
  • 7. Efficacy of Subspace-Based Fine-Tuning Methods
  • 8. Conclusion
  • Acknowledgements
  • References
  • A. Experimental Details
  • A.1. Datasets and Prompts
  • A.2. Computing the Kernel
  • A.3. Solving the Kernel
  • B. Additional Experimental Results
  • B.1. Solvable Task Experiments
  • B.2. Robustness to Choice of Prompt
  • C. Kernel Behavior and the Parametrization
  • C.1. Preliminaries
  • C.2. SignGD Kernel Derivation
  • C.3. Prompt-based Fine-Tuning
  • C.4. µP for SGD and SignGD
  • C.5. Prompt-based Fine-Tuning: A Linear Example
  • C.6. LoRA FT Exhibits Kernel Behavior
  • D. Subspace-Based Fine-Tuning Methods
  • D.1. Intrinsic Dimension FT
  • D.2. Proofs

Knowls

  1. Knowl 1 — Prompted fine-tuning is kernel-like when the pretrained model already solves the task in the width limit

    theoretical result

    Consider a family of pretrained networks of width nn, with downstream examples (ξ,y)(\xi,y) and loss ℓ\ell. Define the output derivative as χ(ξ,y,f0n)=∂ℓ(f0n(ξ),y)/∂f\chi(\xi,y,f_0^n)=\partial \ell(f_0^n(\xi),y)/\partial f, where f0nf_0^n is the pretrained model before fine-tuning. The downstream task is natural for the pretraining scheme if this derivative tends to zero as width grows for every example: lim⁡n→∞χ(ξ,y,f0n)=0\lim_{n\to\infty}\chi(\xi,y,f_0^n)=0. Under this condition, if the network is stable, nontrivial, and expressible as a Tensor Program, prompt-based fine-tuning exhibits both linearization and fixed features in the infinite-width limit. Stability and nontriviality mean, respectively, that outputs remain bounded with width and that training can change the output. The result concerns prompt-based fine-tuning that reuses the pretrained model, and relies on the stated Tensor Program assumptions.

  2. Knowl 2 — The kernel analog for SignGD uses signed training gradients

    theoretical result

    For a scalar-output model f(ξ;θ)f(\xi;\theta), SignGD updates parameters by θt=θt−1−η sign⁡(∇θℓt)\theta_t=\theta_{t-1}-\eta\,\operatorname{sign}(\nabla_\theta \ell_t), with the sign applied coordinate-wise and learning rate η\eta. Let χt=∂ℓ(f(ξt;θt−1),yt)/∂f\chi_t=\partial\ell(f(\xi_t;\theta_{t-1}),y_t)/\partial f be the loss derivative for the training example (ξt,yt)(\xi_t,y_t) at step tt, and let θ0\theta_0 be the parameters before fine-tuning. If training exhibits kernel behavior, its function updates satisfy f(ξ;θt)−f(ξ;θt−1)≈−η sign⁡(χt)K(A-SignGD)(ξ,ξt)f(\xi;\theta_t)-f(\xi;\theta_{t-1})\approx-\eta\,\operatorname{sign}(\chi_t)K^{(A\text{-SignGD})}(\xi,\xi_t), where K(A-SignGD)(ξ,ξ′)=⟨∇θf(ξ;θ0),sign⁡(∇θf(ξ′;θ0))⟩K^{(A\text{-SignGD})}(\xi,\xi')=\langle\nabla_\theta f(\xi;\theta_0),\operatorname{sign}(\nabla_\theta f(\xi';\theta_0))\rangle. This asymmetric kernel is the theoretically derived analog for SignGD. The paper uses SignGD to approximate early-stage Adam dynamics, when Adam’s coordinate-wise normalization is approximated by a sign update; the theorem does not claim that this kernel describes general or late-stage Adam training. The paper also considers the symmetric kernel formed by taking the inner product of the signed gradients at both inputs as a practical alternative.

  3. Knowl 3 — Kernel behavior means linearized updates with nearly fixed gradients

    definition

    A training process for a model f(ξ;θ)f(\xi;\theta) exhibits kernel behavior when, for each fixed input ξ\xi, two approximations hold: the change in output at step tt is captured by the first-order Taylor term, f(ξ;θt)−f(ξ;θt−1)≈⟨∇θf(ξ;θt−1),θt−θt−1⟩f(\xi;\theta_t)-f(\xi;\theta_{t-1})\approx\langle\nabla_\theta f(\xi;\theta_{t-1}),\theta_t-\theta_{t-1}\rangle, and the parameter gradient remains close to its pretrained value, ∇θf(ξ;θt)≈∇θf(ξ;θ0)\nabla_\theta f(\xi;\theta_t)\approx\nabla_\theta f(\xi;\theta_0). Here θ0\theta_0 is the initial parameter vector and θt\theta_t is the vector after tt training steps; for vector outputs, the conditions apply to each output component. Under these conditions, SGD has the fixed neural tangent kernel K(SGD)(ξ,ξ′)=⟨∇θf(ξ;θ0),∇θf(ξ′;θ0)⟩K^{(\mathrm{SGD})}(\xi,\xi')=\langle\nabla_\theta f(\xi;\theta_0),\nabla_\theta f(\xi';\theta_0)\rangle, so its function updates can be expressed through this kernel and the loss derivative.

  4. Knowl 4 — Random low-dimensional updates preserve the SGD kernel at sufficient dimension

    theoretical result

    The paper relates LoRA and random-subspace fine-tuning to preservation of the full fine-tuning SGD kernel. For a fully connected layer with input vectors xix_i, preactivation gradients dh(i)d_h(i), and LoRA matrix AA, the layer’s full SGD kernel entry is K(i,j)=⟨dh(i),dh(j)⟩⟨xi,xj⟩K(i,j)=\langle d_h(i),d_h(j)\rangle\langle x_i,x_j\rangle, whereas the LoRA kernel entry is KLoRA(i,j)=⟨dh(i),dh(j)⟩⟨Axi,Axj⟩K_{\mathrm{LoRA}}(i,j)=\langle d_h(i),d_h(j)\rangle\langle Ax_i,Ax_j\rangle. With the LoRA factor initialized as a suitably scaled random projection, the Johnson–Lindenstrauss guarantee implies that these entries are close with high probability when the rank is sufficiently large; for bounded layer gradients and inputs, the paper gives a failure probability at most 4N2exp⁡(−(ϵ2−ϵ3)k/4)4N^2\exp(- (\epsilon^2-\epsilon^3)k/4) for any entry error of at least c2ϵc^2\epsilon, where NN is the number of downstream examples, kk is the projection dimension, and the gradient and input squared norms are at most cc. Thus rank on the order of log⁡N/ϵ2\log N/\epsilon^2, up to the stated constants, preserves the kernel. The paper gives an analogous bound for intrinsic-dimension fine-tuning using a random projection of the full parameter gradient, assuming the full kernel entries are bounded. These results explain preservation of kernel dynamics when the corresponding full fine-tuning process itself exhibits kernel behavior.

  5. Knowl 5 — Few-shot experiments find kernel behavior on eight of fourteen prompted tasks

    empirical result

    The experiments compare empirical neural tangent kernels (eNTKs) with fine-tuning of RoBERTa-base (125M parameters) on 14 few-shot NLP classification tasks: eight single-sentence tasks and six sentence-pair tasks. They use manual prompt templates, five sampled datasets per task, and 16 or 64 examples per class; accuracy is reported except for MRPC and QQP, where F1 is used. The study evaluates SGD kernels against SGD fine-tuning and sign-based kernels against Adam fine-tuning. Across these settings, eNTK regression reaches at least 90% of fine-tuning performance on 12 of the 14 tasks, while eight tasks consistently meet the paper’s empirical criteria for kernel behavior across the sampled datasets. The results show that an eNTK can solve a task without necessarily showing that the fine-tuning trajectory itself follows kernel dynamics.

  6. Knowl 6 — Meaningful prompts sharply improve kernel performance over standard fine-tuning

    empirical result

    In comparisons across SST-2, MR, CR, QNLI, QQP, and RTE, the eNTK matches fine-tuning substantially better when the downstream task is phrased with a natural-language prompt and solved through masked-token prediction than when standard fine-tuning uses a learned classifier on the [CLS] representation. In the standard setting, the SGD fine-tuning versus SGD-kernel performance gap reaches 16 percentage points on tasks where the prompted setting’s gap is only about 3 points. Prompt choice also matters: on sentiment tasks, minimal null prompts yield a substantial gap between fine-tuning and the kernel, whereas manual prompts generally give closer results. These observations support the paper’s claim that a suitable prompt can make the downstream task more compatible with the pretrained language-model objective.

  7. Knowl 7 — Direct trajectory measurements support kernel behavior when the eNTK succeeds

    empirical result

    The experiments test fine-tuning dynamics separately from eNTK task performance. For a test input ξ\xi, the linearized model is f(ξ;θPT)+⟨∇θf(ξ;θPT),θFT−θPT⟩f(\xi;\theta_{\mathrm{PT}})+\langle\nabla_\theta f(\xi;\theta_{\mathrm{PT}}),\theta_{\mathrm{FT}}-\theta_{\mathrm{PT}}\rangle, where θPT\theta_{\mathrm{PT}} and θFT\theta_{\mathrm{FT}} are the pretrained and fine-tuned parameters. For every task the eNTK solves under the paper’s criterion, this linearized model recovers at least 50% of the improvement from the pretrained model to the fine-tuned model. The study also measures the element-wise relative distance between kernels before and after fine-tuning; distances below 2.0 are counted as evidence for fixed features, and tasks solved by the eNTK show low distances under this measure. The thresholds were selected manually because the theoretical definition does not prescribe numerical cutoffs.

  8. Knowl 8 — Several tasks fail to show kernel behavior, with prompt mismatch offered as an explanation

    empirical result

    TREC, MNLI, SNLI, QNLI, and MPQA consistently fail to meet the paper’s empirical criteria for kernel behavior. The authors conjecture that these outcomes reflect prompts or label words that do not make the task a natural fill-in-the-blank problem for the pretrained model: TREC’s colon prompt may provide too little signal; the label word Maybe can produce ungrammatical neutral MNLI and SNLI examples; QNLI premises and yes/no label words may not form natural statements; and MPQA inputs can be too short for its sentence-style sentiment prompt. QNLI is a qualified case: its eNTK sometimes solves the task, but the measurements suggest that its fine-tuning trajectory does not strongly satisfy linearization. These explanations are proposed interpretations rather than proven causes.

  9. Knowl 9 — SignGD fine-tuning provides empirical support for its use as an early-Adam proxy

    empirical result

    The paper compares prompt-based SignGD fine-tuning—coordinate-wise gradient signs, without momentum—with Adam fine-tuning using the same hyperparameter search. SignGD gives strong results and is often comparable to Adam, particularly on sentence-pair tasks. For example, at 64 shots per class, SignGD versus Adam accuracy is 69.3% versus 67.9% on MNLI, 77.4% versus 76.9% on SNLI, and 76.8% versus 74.2% on QNLI; on MRPC, F1 is 84.1% versus 80.9%. This evidence supports studying SignGD as an approximation to early-stage Adam in the paper’s kernel analysis, but does not establish equivalence for all training stages or settings.

  10. Knowl 10 — The evidence is limited to few-shot classification and early-stage Adam dynamics

    limitation

    The experiments cover few-shot classification, one masked language model (RoBERTa-base), and selected prompts; extending the study to other models, tasks, or larger training sets is computationally costly because eNTK computation is expensive. The theoretical account of Adam applies to early-stage training through its approximation by SignGD, and the paper does not establish how well the kernel description applies over longer training. The findings also show that kernel behavior depends on the downstream task and prompt, so the proposed account does not describe every fine-tuning trajectory.

Coverage note — Detailed Tensor Program proofs and secondary parametrization analyses are omitted because they support the stated main results rather than adding standalone findings.

References

  1. 1.Abnar, S., Dehghani, M., Neyshabur, B., and Sedghi, H. Exploring the limits of large scale pre-training, 2021. URL https://arxiv.org/abs/2110.02095.
  2. 2.Achille, A., Golatkar, A., Ravichandran, A., Polito, M., and Soatto, S. Lqf: Linear quadratic fine-tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15729–15739, 2021.
  3. 3.Aghajanyan, A., Gupta, S., and Zettlemoyer, L. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 7319–7328, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.568. URL https://aclanthology.org/2021.acl-long.568.
  4. 4.Allen-Zhu, Z., Li, Y., and Liang, Y. Learning and generalization in overparameterized neural networks, going beyond two layers. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019a. URL https://proceedings.neurips.cc/paper/2019/file/62dad6e273d32235ae02b7d321578ee8-Paper.pdf.
  5. 5.Allen-Zhu, Z., Li, Y., and Song, Z. A convergence theory for deep learning via over-parameterization. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 242–252. PMLR, 09–15 Jun 2019b. URL https://proceedings.mlr.press/v97/allen-zhu19a.html.
  6. 6.Arora, S., Du, S., Hu, W., Li, Z., and Wang, R. Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 322–332. PMLR, 09–15 Jun 2019a. URL https://proceedings.mlr.press/v97/arora19a.html.
  7. 7.Arora, S., Du, S. S., Hu, W., Li, Z., Salakhutdinov, R. R., and Wang, R. On exact computation with an infinitely wide neural net. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019b. URL https://proceedings.neurips.cc/paper/2019/file/dbc4d84bfcfe2284ba11beffb853a8c4-Paper.pdf.
  8. 8.Arora, S., Du, S. S., Li, Z., Salakhutdinov, R., Wang, R., and Yu, D. Harnessing the power of infinitely wide deep nets on small-data tasks. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rkl8sJBYvH.
  9. 9.Bar Haim, R., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I. The second PASCAL recognising textual entailment challenge. 2006. URL https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.60.8552&rep=rep1&type=pdf.
  10. 10.Ben Zaken, E., Goldberg, Y., and Ravfogel, S. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 1–9, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-short.1. URL https://aclanthology.org/2022.acl-short.1.
  11. 11.Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D. The fifth PASCAL recognizing textual entailment challenge. In TAC, 2009. URL https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.232.1231&rep=rep1&type=pdf.
  12. 12.Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 632–642, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1075. URL https://aclanthology.org/D15-1075.
  13. 13.Cao, Y. and Gu, Q. Generalization bounds of stochastic gradient descent for wide and deep neural networks. Advances in neural information processing systems, 32, 2019.
  14. 14.Chen, X., Liang, C., Huang, D., Real, E., Liu, Y., Wang, K., Hsieh, C.-J., Lu, Y., and Le, Q. V. Evolved optimizer for vision. In First Conference on Automated Machine Learning (Late-Breaking Workshop), 2022. URL https://openreview.net/forum?id=jK_eS5BxOuu.
  15. 15.Chua, K., Lei, Q., and Lee, J. D. How fine-tuning allows for effective meta-learning. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 8871–8884. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/4a533591763dfa743a13affab1a85793-Paper.pdf.
  16. 16.Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D. Electra: Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=r1xMH1BtvB.
  17. 17.Dagan, I., Glickman, O., and Magnini, B. The PASCAL recognising textual entailment challenge. In the First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment, 2005. URL https://kdd.cs.ksu.edu/Courses/Fall-2008/CIS798/Handouts/06-dagan05pascal.pdf.
  18. 18.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics (NAACL), pp. 4171–4186, 2019.
  19. 19.Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In the Third International Workshop on Paraphrasing (IWP2005), 2005. URL https://aclanthology.org/I05-5002.pdf.
  20. 20.Du, S., Lee, J., Li, H., Wang, L., and Zhai, X. Gradient descent finds global minima of deep neural networks. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 1675–1685. PMLR, 09–15 Jun 2019a. URL https://proceedings.mlr.press/v97/du19c.html.
  21. 21.Du, S. S., Zhai, X., Poczos, B., and Singh, A. Gradient descent provably optimizes over-parameterized neural networks. In International Conference on Learning Representations, 2019b. URL https://openreview.net/forum?id=S1eK3i09YQ.
  22. 22.Du, S. S., Hu, W., Kakade, S. M., Lee, J. D., and Lei, Q. Few-shot learning via learning the representation, provably. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=pW2Q2xLwIMD.
  23. 23.Gao, T., Fisch, A., and Chen, D. Making pre-trained language models better few-shot learners. In Association for Computational Linguistics (ACL), pp. 3816–3830, 2021.
  24. 24.Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, B. The third PASCAL recognizing textual entailment challenge. In the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, 2007. URL https://aclanthology.org/W07-1401.pdf.
  25. 25.He, H. and Zou, R. functorch: Jax-like composable function transforms for pytorch. https://github.com/pytorch/functorch, 2021.
  26. 26.He, M., He, F., Shi, L., Huang, X., and Suykens, J. A. K. Learning with asymmetric kernels: Least squares and feature interpretation, 2022. URL https://arxiv.org/abs/2202.01397.
  27. 27.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models, 2021. URL https://arxiv.org/abs/2106.09685.
  28. 28.Hu, M. and Liu, B. Mining and summarizing customer reviews. In ACM SIGKDD international conference on Knowledge discovery and data mining, 2004.
  29. 29.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper/2018/file/5a4be1fa34e62bb8a6ec6b91d2462f5a-Paper.pdf.
  30. 30.Johnson, W. B. Extensions of lipschitz mappings into a hilbert space. Contemp. Math., 26:189–206, 1984.
  31. 31.Joshi, M., Chen, D., Liu, Y., Weld, D. S., Zettlemoyer, L., and Levy, O. SpanBERT: Improving pre-training by representing and predicting spans. Transactions of the Association of Computational Linguistics (TACL), 2020.
  32. 32.Lee, J. D., Lei, Q., Saunshi, N., and ZHUO, J. Predicting what you already know helps: Provable self-supervised learning. In Advances in Neural Information Processing Systems, volume 34, pp. 309–323. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/02e656adee09f8394b402d9958389b7d-Paper.pdf.
  33. 33.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. In Empirical Methods in Natural Language Processing (EMNLP), pp. 3045–3059, 2021.
  34. 34.Li, C., Farkhoor, H., Liu, R., and Yosinski, J. Measuring the intrinsic dimension of objective landscapes. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=ryup8-WCW.
  35. 35.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 4582–4597, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.353. URL https://aclanthology.org/2021.acl-long.353.
  36. 36.Li, Y. and Liang, Y. Learning overparameterized neural networks via stochastic gradient descent on structured data. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper/2018/file/54fe976ba170c19ebae453679b362263-Paper.pdf.
  37. 37.Li, Z., Bhojanapalli, S., Zaheer, M., Reddi, S., and Kumar, S. Robust training of neural networks using scale invariant architectures. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 12656–12684. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/li22b.html.
  38. 38.Littwin, E. and Yang, G. Adaptive optimization in the ∞\infty-width limit. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=zgVDqw9ZUES.
  39. 39.Liu, L., Liu, X., Gao, J., Chen, W., and Han, J. Understanding the difficulty of training transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5747–5763, Online, November 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.463. URL https://aclanthology.org/2020.emnlp-main.463.
  40. 40.Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv., aug 2022. ISSN 0360-0300. doi: 10.1145/3560815. URL https://doi.org/10.1145/3560815. Just Accepted.
  41. 41.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Ro{bert}a: A robustly optimized {bert} pretraining approach, 2020b. URL https://openreview.net/forum?id=SyxS0T4tvS.
  42. 42.Logan IV, R., Balazevic, I., Wallace, E., Petroni, F., Singh, S., and Riedel, S. Cutting down on prompts and parameters: Simple few-shot learning with language models. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2824–2835, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.222. URL https://aclanthology.org/2022.findings-acl.222.
  43. 43.Ma, C., Wu, L., and E, W. A qualitative study of the dynamic behavior for adaptive gradient algorithms. In Proceedings of the 2nd Mathematical and Scientific Machine Learning Conference, volume 145 of Proceedings of Machine Learning Research, pp. 671–692. PMLR, 16–19 Aug 2022. URL https://proceedings.mlr.press/v145/ma22a.html.
  44. 44.Maddox, W., Tang, S., Moreno, P., Gordon Wilson, A., and Damianou, A. Fast adaptation with linearized neural networks. In Banerjee, A. and Fukumizu, K. (eds.), Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, volume 130 of Proceedings of Machine Learning Research, pp. 2737–2745. PMLR, 13–15 Apr 2021. URL https://proceedings.mlr.press/v130/maddox21a.html.
  45. 45.Malladi, S., Lyu, K., Panigrahi, A., and Arora, S. On the sdes and scaling rules for adaptive gradient algorithms, 2022. URL https://arxiv.org/abs/2205.10287.
  46. 46.Mu, F., Liang, Y., and Li, Y. Gradients as features for deep representation learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=BkeoaeHKDS.
  47. 47.Novak, R., Sohl-Dickstein, J., and Schoenholz, S. S. Fast finite width neural tangent kernel. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 17018–17044. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/novak22a.html.
  48. 48.Pang, B. and Lee, L. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Association for Computational Linguistics (ACL), 2004.
  49. 49.Pang, B. and Lee, L. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Association for Computational Linguistics (ACL), 2005.
  50. 50.Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. Improving language understanding by generative pre-training. 2018.
  51. 51.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  52. 52.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. SQuAD: 100,000+ questions for machine comprehension of text. In Empirical Methods in Natural Language Processing (EMNLP), 2016. URL https://aclanthology.org/D16-1264/.
  53. 53.Saunshi, N., Plevrakis, O., Arora, S., Khodak, M., and Khandeparkar, H. A theoretical analysis of contrastive unsupervised representation learning. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 5628–5637. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/saunshi19a.html.
  54. 54.Saunshi, N., Malladi, S., and Arora, S. A mathematical exploration of why language models help solve downstream tasks. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=vVjIW3sEc1s.
  55. 55.Saunshi, N., Ash, J., Goel, S., Misra, D., Zhang, C., Arora, S., Kakade, S., and Krishnamurthy, A. Understanding contrastive learning requires incorporating inductive biases. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 19250–19286. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/saunshi22a.html.
  56. 56.Schick, T. and Schutze, H. Exploiting cloze-questions for few-shot text classification and natural language inference. In European Chapter of the Association for Computational Linguistics (EACL), pp. 255–269, 2021.
  57. 57.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Empirical Methods in Natural Language Processing (EMNLP), 2013. URL https://aclanthology.org/D13-1170.pdf.
  58. 58.Tosh, C., Krishnamurthy, A., and Hsu, D. Contrastive learning, multi-view redundancy, and linear models. In Proceedings of the 32nd International Conference on Algorithmic Learning Theory, volume 132 of Proceedings of Machine Learning Research, pp. 1179–1206. PMLR, 16–19 Mar 2021a. URL https://proceedings.mlr.press/v132/tosh21a.html.
  59. 59.Tosh, C., Krishnamurthy, A., and Hsu, D. Contrastive estimation reveals topic posterior information to linear models. Journal of Machine Learning Research, 22(281):1–31, 2021b. URL http://jmlr.org/papers/v22/21-0089.html.
  60. 60.Tripuraneni, N., Jordan, M., and Jin, C. On the theory of transfer learning: The importance of task diversity. In Advances in Neural Information Processing Systems, volume 33, pp. 7852–7862. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/59587bffec1c7846f3e34230141556ae-Paper.pdf.
  61. 61.Tsai, Y.-H. H., Wu, Y., Salakhutdinov, R., and Morency, L.-P. Self-supervised learning from a multi-view perspective. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=-bdp_8Itjwp.
  62. 62.Voorhees, E. M. and Tice, D. M. Building a question answering test collection. In the 23rd annual international ACM SIGIR conference on Research and development in information retrieval, 2000.
  63. 63.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations (ICLR), 2019. URL https://openreview.net/forum?id=rJ4km2R5t7.
  64. 64.Wei, A., Hu, W., and Steinhardt, J. More than a toy: Random matrix models predict how real-world neural representations generalize. In Proceedings of the 39th International Conference on Machine Learning, volume 162, pp. 23549–23588. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/wei22a.html.
  65. 65.Wiebe, J., Wilson, T., and Cardie, C. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39(2-3), 2005.
  66. 66.Williams, A., Nangia, N., and Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), 2018. URL https://aclanthology.org/N18-1101.pdf.
  67. 67.Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N. Kernel and rich regimes in overparametrized models. In Proceedings of Thirty Third Conference on Learning Theory, volume 125 of Proceedings of Machine Learning Research, pp. 3635–3673. PMLR, 09–12 Jul 2020. URL https://proceedings.mlr.press/v125/woodworth20a.html.
  68. 68.Wu, S., Zhang, H. R., and Ré, C. Understanding and improving information transfer in multi-task learning. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SylzhkBtDB.
  69. 69.Yang, G. Wide feedforward or recurrent neural networks of any architecture are gaussian processes. Advances in Neural Information Processing Systems, 32, 2019.
  70. 70.Yang, G. Tensor programs ii: Neural tangent kernel for any architecture. arXiv preprint arXiv:2006.14548, 2020a.
  71. 71.Yang, G. Tensor programs iii: Neural matrix laws. arXiv preprint arXiv:2009.10685, 2020b.
  72. 72.Yang, G. and Hu, E. J. Tensor programs iv: Feature learning in infinite-width neural networks. In International Conference on Machine Learning, pp. 11727–11737. PMLR, 2021.
  73. 73.Yang, G. and Littwin, E. Tensor programs iib: Architectural universality of neural tangent kernel training dynamics. In International Conference on Machine Learning, pp. 11762–11772. PMLR, 2021.
  74. 74.Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J. Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer. arXiv preprint arXiv:2203.03466, 2022.
  75. 75.Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S. Why are adaptive methods good for attention models? In Advances in Neural Information Processing Systems, volume 33, pp. 15383–15393. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper/2020/file/b05b57f6add810d3b7490866d74c0053-Paper.pdf.
  76. 76.Zhang, X., Zhao, J., and LeCun, Y. Character-level convolutional networks for text classification. In Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL https://proceedings.neurips.cc/paper/2015/file/250cf8b51c773f3f8dc8b4be867a9a02-Paper.pdf.
  77. 77.Zou, D., Cao, Y., Zhou, D., and Gu, Q. Stochastic gradient descent optimizes over-parameterized deep relu networks, 2018. URL https://arxiv.org/abs/1811.08888.

Citation

MLA
Malladi, S., et al. “A Kernel-Based View of Language Model Fine-Tuning”. International Conference on Machine Learning, vol. 202, 2023, pp. 23610–41, https://proceedings.mlr.press/v202/malladi23a.html.
APA
Malladi, S., Wettig, A., Yu, D., Chen, D., & Arora, S. (2023). A Kernel-Based View of Language Model Fine-Tuning. International Conference on Machine Learning, 202, 23610–23641. https://proceedings.mlr.press/v202/malladi23a.html
Chicago
Malladi, S., A. Wettig, D. Yu, D. Chen, and S. Arora. 2023. “A Kernel-Based View of Language Model Fine-Tuning”. International Conference on Machine Learning 202: 23610–41. https://proceedings.mlr.press/v202/malladi23a.html.
Harvard
Malladi, S. et al. (2023) “A Kernel-Based View of Language Model Fine-Tuning”, International Conference on Machine Learning. PMLR, pp. 23610–23641. Available at: https://proceedings.mlr.press/v202/malladi23a.html.
Vancouver
1. Malladi S, Wettig A, Yu D, Chen D, Arora S (2023) A Kernel-Based View of Language Model Fine-Tuning. In: International Conference on Machine Learning. PMLR, pp 23610–23641

BibTeX

@InProceedings{pmlr-v202-malladi23a,
  title = 	 {A Kernel-Based View of Language Model Fine-Tuning},
  author =       {Malladi, Sadhika and Wettig, Alexander and Yu, Dingli and Chen, Danqi and Arora, Sanjeev},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {23610--23641},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/malladi23a/malladi23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/malladi23a.html},
  abstract = 	 {It has become standard to solve NLP tasks by fine-tuning pre-trained language models (LMs), especially in low-data settings. There is minimal theoretical understanding of empirical success, e.g., why fine-tuning a model with $10^8$ or more parameters on a couple dozen training points does not result in overfitting. We investigate whether the Neural Tangent Kernel (NTK)—which originated as a model to study the gradient descent dynamics of infinitely wide networks with suitable random initialization—describes fine-tuning of pre-trained LMs. This study was inspired by the decent performance of NTK for computer vision tasks (Wei et al., 2022). We extend the NTK formalism to Adam and use Tensor Programs (Yang, 2020) to characterize conditions under which the NTK lens may describe fine-tuning updates to pre-trained language models. Extensive experiments on 14 NLP tasks validate our theory and show that formulating the downstream task as a masked word prediction problem through prompting often induces kernel-based dynamics during fine-tuning. Finally, we use this kernel view to propose an explanation for the success of parameter-efficient subspace-based fine-tuning methods.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/