Unveiling the Generalization Power of Fine-Tuned Large Language Models
Haoran YangYumeng ZhangJiaqi XuHongyuan LuPheng-Ann HengWai Lam
Reveals how task-specific fine-tuning alters the out-of-domain and cross-task generalization of large language models differently across classification and generation tasks, demonstrating that incorporating in-context demonstrations during fine-tuning prevents loss of generalizability.
Organizations increasingly customize large language models for specialized applications, yet the broader consequences of task-specific training remain poorly understood. While fine-tuning often yields high accuracy on familiar test data, practitioners risk degrading a model's intrinsic ability to generalize across new domains or adapt to different business tasks.
The article systematically evaluates how task-specific fine-tuning impacts a model's generalization capabilities across familiar domains, unfamiliar out-of-domain environments, and entirely different language tasks.
To assess these dynamics, the researchers conducted extensive empirical experiments using the open-source Llama-2-7B foundation model across five standard natural language tasks: text summarization, question generation, sentiment classification, paraphrase detection, and natural language inference. The study tested model performance across 14 distinct benchmark datasets under varying training sample sizes (2,000, 4,000, and 6,000 samples) and different inference prompting conditions, including zero-shot and multi-example in-context learning.
The analysis revealed several critical findings. First, task-specific fine-tuning consistently delivered strong zero-shot performance on familiar in-domain test data, but the models benefited very little from additional example prompts during inference. Second, generalization to new domains diverged sharply between task types: models fine-tuned on classification tasks successfully generalized to out-of-domain data, whereas models fine-tuned on text generation suffered noticeable performance drops compared to the un-fine-tuned base model. Third, cross-task transfer proved highly asymmetric; models fine-tuned on classification tasks entirely failed when applied to text generation because they collapsed into outputting isolated classification labels. Finally, prepending example demonstrations during the fine-tuning stage of generation tasks—termed fine-tuning with in-context learning—substantially improved out-of-domain generalization and cross-task adaptability by keeping model weights closer to the base model.
These results demonstrate that task-specific fine-tuning presents distinct trade-offs between specialization and generalization. Deploying specialized classification models carries minimal risk of domain degradation within similar tasks, but fine-tuning generative models carries a significant risk of catastrophic forgetting and reduced domain versatility. Furthermore, simply increasing training dataset sizes beyond 2,000 to 4,000 samples offered diminishing or even negative returns depending on the task, underscoring that more training data does not automatically improve generalization.
Organizations should adopt targeted deployment strategies based on task type. For classification tasks, teams should utilize standard fine-tuning workflows. For text generation tasks, practitioners seeking broader domain resilience should adopt fine-tuning with in-context learning by embedding example demonstrations directly into the training data. Additionally, engineering prompt structures to avoid uniform formatting across tasks can partially alleviate rigid output specialization.
Confidence in these findings is high for standard open-source transformer architectures operating under supervised fine-tuning. However, stakeholders should exercise caution, as the experiments were conducted on a single 7-billion-parameter model architecture and evaluated without advanced alignment techniques such as reinforcement learning from human feedback. Further research is necessary to confirm whether these generalization patterns hold across larger model scales and reinforcement-tuned architectures.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). It establishes instruction fine-tuning as a route to zero-shot transfer, giving useful context for this paper’s tests of how task-specific tuning affects performance beyond the training task.
- Paper: Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation, Tu Vu et al. (2022). Its experiments show how generative fine-tuning can cause catastrophic forgetting across languages, framing the generalization risks this paper measures across domains and tasks.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). Its unified text-to-text framework treats classification and generation in one model, clarifying the task-format distinction central to this paper’s asymmetric transfer findings.
- Paper: Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning, Lorenzo Jaime Flores et al. (2026). It carries the fine-tuning analysis into uncertainty estimation, showing why practitioners must reassess confidence calibration after adapting a model.
- Paper: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?, Zorik Gekhman et al. (2024). It extends the generative fine-tuning risk analysis to factual knowledge, testing how learning new facts affects hallucinations and retained knowledge.
- Paper: LESS: Selecting Influential Data for Targeted Instruction Tuning, Mengzhou Xia et al. (2024). It turns the paper’s warning that more training data can yield diminishing returns into a targeted data-selection method for instruction tuning.
- Paper: In-Context Learning with Long-Context Models: An In-Depth Exploration, Amanda Bertsch et al. (2025). It extends the comparison between parameter updates and in-context examples by testing whether thousands of demonstrations can rival or outperform fine-tuning.
