A Kernel-Based View of Language Model Fine-Tuning
Sadhika MalladiAlexander WettigDingli YuDanqi ChenSanjeev Arora
Extends Neural Tangent Kernel theory to Adam optimizer dynamics and pre-trained initializations, explaining why prompt-based fine-tuning and parameter-efficient adaptation methods succeed in low-data regimes without overfitting.
Adapting large pre-trained language models to downstream tasks using very small datasets has become standard practice across language processing applications. However, a foundational theoretical explanation for why massive models with hundreds of millions of parameters do not severely overfit in these low-data regimes has remained elusive. Furthermore, practitioners frequently observe that fine-tuning success is highly sensitive to implementation choices, such as using natural language prompts or restricting parameter updates to low-rank subspaces, without a formal understanding of why these techniques work.
The article aims to evaluate whether the Neural Tangent Kernel framework—a mathematical lens originally developed to study the optimization of infinitely wide, randomly initialized neural networks—can describe and explain the fine-tuning of pre-trained language models. Specifically, the authors seek to determine the precise conditions under which fine-tuning behaves like kernel regression, characterize the dynamics of standard optimizers such as Adam, and explain the success of parameter-efficient fine-tuning methods.
To investigate these questions, the authors mathematically extended kernel theory to account for non-random, pre-trained initializations and derived a new sign-based kernel formulation that captures early-stage training dynamics under adaptive optimizers like Adam. They then evaluated this theoretical framework empirically across 14 diverse natural language understanding tasks—including sentiment classification, topic identification, natural language inference, and paraphrase detection—using few-shot datasets (16 and 64 training examples per class) and a pre-trained RoBERTa model. The empirical testing examined whether fine-tuning trajectories satisfy two defining criteria of kernel behavior: function linearization and fixed gradient features.
The investigation produced four central findings. First, formulating downstream tasks via natural language prompts is critical for inducing kernel behavior; prompt-based empirical kernels closely matched actual fine-tuning performance, whereas standard classification heads without prompts exhibited performance deficits of up to 16 percentage points. Second, the empirical kernel accurately solved 12 out of the 14 downstream tasks, matching fine-tuning performance within a 10% margin, with 8 tasks strictly satisfying all mathematical conditions of kernel behavior throughout training. Third, standard stochastic gradient descent performed within 4 percentage points of Adam in prompt-based settings, confirming that prompted optimization landscapes are sufficiently benign that optimizer differences diminish. Fourth, the kernel perspective mathematically explains the effectiveness of low-rank adaptation methods, proving that random subspace projections preserve the underlying kernel matrix and deliver accuracy comparable to full-parameter tuning.
These findings imply that successful prompt-based fine-tuning requires only minimal, linear parameter adjustments rather than complex feature re-learning. This mechanism explains why large language models generalize effectively from only a few dozen examples without catastrophic overfitting. The results also validate parameter-efficient adaptation strategies like low-rank adaptation as mathematically sound alternatives to full fine-tuning, offering substantial computational and storage cost savings with minimal risk to downstream accuracy.
Practitioners and decision-makers should prioritize natural, well-formatted prompt designs when adapting pre-trained models to ensure stable optimization within the kernel regime. Teams seeking to reduce computing overhead should confidently adopt low-rank adaptation techniques for prompt-based workflows. For complex tasks where standard prompts struggle—such as intricate textual entailment—organizations should conduct pilot evaluations and refine prompt phrasing before deploying fine-tuned models to production.
Confidence in these findings is high for few-shot classification using masked language models, supported by rigorous proofs and broad empirical testing across multiple task benchmarks. However, key limitations remain: the empirical validation focuses on the RoBERTa architecture, the theoretical guarantees for Adam apply primarily to early-stage training dynamics, and the kernel approach does not fully capture performance on tasks where prompt phrasing is unnatural. Further research is needed to extend this framework to generative decoder models and longer training durations.
- Paper: Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning, Armen Aghajanyan et al. (2021). Its intrinsic-dimension account of why fine-tuning avoids overfitting motivates this paper’s complementary kernel-based explanation of low-data adaptation.
- Paper: XLNet: Generalized Autoregressive Pretraining for Language Understanding, Zhilin Yang et al. (2019). Its account of XLNet’s permutation-based pretraining clarifies the pretrained-model objectives and architectures that form the setting for this fine-tuning analysis.
No sufficiently relevant recommendations were found.
