Unveiling the Generalization Power of Fine-Tuned Large Language Models

Haoran YangYumeng ZhangJiaqi XuHongyuan LuPheng-Ann HengWai Lam

article2024NAACL83 citations

Reveals how task-specific fine-tuning alters the out-of-domain and cross-task generalization of large language models differently across classification and generation tasks, demonstrating that incorporating in-context demonstrations during fine-tuning prevents loss of generalizability.

Listen

Organizations increasingly customize large language models for specialized applications, yet the broader consequences of task-specific training remain poorly understood. While fine-tuning often yields high accuracy on familiar test data, practitioners risk degrading a model's intrinsic ability to generalize across new domains or adapt to different business tasks.

The article systematically evaluates how task-specific fine-tuning impacts a model's generalization capabilities across familiar domains, unfamiliar out-of-domain environments, and entirely different language tasks.

To assess these dynamics, the researchers conducted extensive empirical experiments using the open-source Llama-2-7B foundation model across five standard natural language tasks: text summarization, question generation, sentiment classification, paraphrase detection, and natural language inference. The study tested model performance across 14 distinct benchmark datasets under varying training sample sizes (2,000, 4,000, and 6,000 samples) and different inference prompting conditions, including zero-shot and multi-example in-context learning.

The analysis revealed several critical findings. First, task-specific fine-tuning consistently delivered strong zero-shot performance on familiar in-domain test data, but the models benefited very little from additional example prompts during inference. Second, generalization to new domains diverged sharply between task types: models fine-tuned on classification tasks successfully generalized to out-of-domain data, whereas models fine-tuned on text generation suffered noticeable performance drops compared to the un-fine-tuned base model. Third, cross-task transfer proved highly asymmetric; models fine-tuned on classification tasks entirely failed when applied to text generation because they collapsed into outputting isolated classification labels. Finally, prepending example demonstrations during the fine-tuning stage of generation tasks—termed fine-tuning with in-context learning—substantially improved out-of-domain generalization and cross-task adaptability by keeping model weights closer to the base model.

These results demonstrate that task-specific fine-tuning presents distinct trade-offs between specialization and generalization. Deploying specialized classification models carries minimal risk of domain degradation within similar tasks, but fine-tuning generative models carries a significant risk of catastrophic forgetting and reduced domain versatility. Furthermore, simply increasing training dataset sizes beyond 2,000 to 4,000 samples offered diminishing or even negative returns depending on the task, underscoring that more training data does not automatically improve generalization.

Organizations should adopt targeted deployment strategies based on task type. For classification tasks, teams should utilize standard fine-tuning workflows. For text generation tasks, practitioners seeking broader domain resilience should adopt fine-tuning with in-context learning by embedding example demonstrations directly into the training data. Additionally, engineering prompt structures to avoid uniform formatting across tasks can partially alleviate rigid output specialization.

Confidence in these findings is high for standard open-source transformer architectures operating under supervised fine-tuning. However, stakeholders should exercise caution, as the experiments were conducted on a single 7-billion-parameter model architecture and evaluated without advanced alignment techniques such as reinforcement learning from human feedback. Further research is necessary to confirm whether these generalization patterns hold across larger model scales and reinforcement-tuned architectures.

Cover for Unveiling the Generalization Power of Fine-Tuned Large Language Models

Abstract

While Large Language Models (LLMs) have demonstrated exceptional multitasking abilities, fine-tuning these models on downstream, domain-specific datasets is often necessary to yield superior performance on test sets compared to their counterparts without fine-tuning. However, the comprehensive effects of fine-tuning on the LLMs’ generalization ability are not fully understood. This paper delves into the differences between original, unmodified LLMs and their fine-tuned variants. Our primary investigation centers on whether fine-tuning affects the generalization ability intrinsic to LLMs. To elaborate on this, we conduct extensive experiments across five distinct language tasks on various datasets. Our main findings reveal that models fine-tuned on generation and classification tasks exhibit dissimilar behaviors in generalizing to different domains and tasks. Intriguingly, we observe that integrating the in-context learning strategy during fine-tuning on generation tasks can enhance the model’s generalization ability. Through this systematic investigation, we aim to contribute valuable insights into the evolving landscape of fine-tuning practices for LLMs. The code and data are available at https://github.com/LHRYANG/Generalization_of_FT-LLM.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Large Language Models
  • 2.2 Fine-tuning vs. In-Context Learning
  • 3 Evaluation Design
  • 3.1 Evaluation Taxonomy
  • 3.2 Evaluation Benchmarks
  • 3.3 Experimental Setup
  • 4 Results and Findings
  • 4.1 Same Task, In-domain Datasets
  • 4.2 Same Task, Out-of-domain Datasets
  • 4.3 Different Tasks
  • 5 Fine-tuning with In-context Learning on Generation Tasks Helps Improve the Generalization Ability of LLMs
  • 5.1 Same Task
  • 5.2 Different Tasks
  • 5.3 Potential Reason
  • 5.4 FTICL on Classificaion Tasks
  • 6 Conclusion
  • Limitations
  • References
  • A Dataset Information
  • B Classification Labels
  • C Prompt Format

Knowls

  1. Knowl 1 — Fine-tuning has opposite out-of-domain effects for generation and classification

    empirical result

    In experiments with Llama-2-7B, task-specific fine-tuning had different effects on out-of-domain test sets depending on whether the task required generation or classification. On generation tasks, fine-tuned models underperformed the original Llama-2 baseline on PeerRead, CNN/DailyMail, Tweetqa, and Sciqa; this gap remained across the evaluated numbers of in-context examples, and adding more fine-tuning samples did not consistently close it. On classification tasks, fine-tuned models generally matched the baseline on sentiment datasets SST2 and Yelp, and consistently outperformed it on the paraphrase and natural-language-inference test sets QQP, STS-B, RTE, and GPTNLI. The authors suggest that the constrained label space of classification may make transfer to new domains easier than transfer to the open-ended output spaces of generation tasks.

  2. Knowl 2 — Fine-tuning with demonstrations improves generation-task generalization

    empirical result

    The paper’s fine-tuning with in-context learning (FTICL) procedure prepends one or two task examples to each training input. For generation tasks, models trained this way on 2,000 examples preserved or slightly improved performance on the training dataset’s test set compared with ordinary fine-tuning, and improved same-task out-of-domain performance. For example, XSum-trained FTICL with one training demonstration outperformed both vanilla fine-tuning and the baseline Llama-2 evaluated with one or two demonstrations on PeerRead; Socialqa-trained FTICL with two demonstrations beat vanilla fine-tuning on Sciqa, though it remained below the baseline. FTICL also improved cross-task results: Socialqa-trained FTICL was at least comparable to the baseline on Amazon and outperformed both the baseline and vanilla fine-tuning on MNLI-1 and Paws; FTICL improved generation-task transfer between Socialqa and XSum relative to vanilla fine-tuning. The authors hypothesize that FTICL helps retain the base model’s general knowledge because it changes the original model’s parameters less: the reported average parameter-weight differences from the original Llama-2 were 7.95×10−57.95\times10^{-5} for Socialqa FTICL versus 8.54×10−58.54\times10^{-5} for vanilla fine-tuning, and 8.03×10−58.03\times10^{-5} for XSum FTICL versus 1.0×10−41.0\times10^{-4} for vanilla fine-tuning. These measurements are consistent with the hypothesis but do not establish it as the cause of the performance gains.

  3. Knowl 3 — Benchmark design spans five tasks and three kinds of generalization

    experimental setup

    The evaluation distinguishes transfer to an in-domain dataset of the same task, transfer to an out-of-domain dataset of the same task, and transfer to a different task. The benchmark assignments are: (1) summarization, trained on XSum, tested in-domain on XSum and XLSum and out-of-domain on PeerRead and CNN/DailyMail; (2) question generation, trained on Socialqa, tested in-domain on Socialqa and out-of-domain on Tweetqa and Sciqa; (3) sentiment classification, trained on Amazon reviews, tested in-domain on Amazon and AmazonFood and out-of-domain on SST2 and Yelp; (4) paraphrase detection, trained on Paws, tested in-domain on Paws and out-of-domain on QQP and STS-B; and (5) natural-language inference, trained on MNLI, tested in-domain on MNLI matched and mismatched splits and out-of-domain on RTE and GPTNLI. Yelp’s 1–5 ratings were converted to binary sentiment labels, with scores below 3.5 negative and above 3.5 positive; STS-B’s similarity ratings were converted to paraphrase labels using a 3.5 threshold.

  4. Knowl 4 — Fine-tuning and inference protocol for the Llama-2-7B comparisons

    experimental setup

    All experiments used open-source Llama-2-7B. Each task-specific model was fine-tuned for two epochs with AdamW at a learning rate of 0.002, using training subsets of 2,000, 4,000, or 6,000 examples. Classification was trained as text generation: the model generated task labels, and inference selected the predefined label with the highest probability. Generation was evaluated with ROUGE-L and classification with accuracy. Same-task evaluations included zero-shot inference and in-context learning (ICL); generation used 1, 2, or 4 inference demonstrations, binary classification used 2, 4, or 6, and three-label NLI used 3, 5, or 7, with every label represented at least once. Cross-task evaluations used ICL to specify the new task; the paper also reports zero-shot classification scores because those predictions are based on label probabilities.

  5. Knowl 5 — Cross-task transfer depends strongly on the fine-tuning dataset

    empirical result

    Cross-task results were not uniform across fine-tuning sources. In the reported comparisons using models fine-tuned on 2,000 examples, Amazon fine-tuning harmed accuracy on MNLI-1, whereas XSum fine-tuning substantially improved it. For generation, Socialqa fine-tuning hurt performance on XSum, while XSum fine-tuning had little effect on Socialqa. These results indicate that cross-task transfer depends on the selected training dataset and is not reliably predicted by whether the source and target tasks are both generative or both classificatory.

  6. Knowl 6 — Classification fine-tuning can prevent generation, with prompt format affecting recovery

    empirical result

    When models fine-tuned on classification data were evaluated on generation tasks, they often produced classification labels instead of coherent text, resulting in nearly zero ROUGE-L. The paper identifies output-space specialization as one possible explanation and also investigates prompt-format effects: the default training and testing prompts for different tasks shared a leading sequence of hash marks, whereas an alternative prompt set used distinct natural-language instructions. Cross-task classification evaluations were more sensitive to this prompt change than generation evaluations. With the alternative prompts, an Amazon-fine-tuned model began to generate on XSum, but it still failed on Socialqa. Thus, prompt design can affect cross-task transfer, but changing the format did not resolve the failure generally.

  7. Knowl 7 — Fine-tuning often improves in-domain zero-shot performance, while inference demonstrations may not help

    empirical result

    On in-domain test sets, task-specific fine-tuned Llama-2 models generally achieved better zero-shot performance than the original Llama-2 evaluated with ICL, including on XLSum, Socialqa, MNLI matched and mismatched, and Paws. The reported exception was sentiment classification on Amazon and AmazonFood, where fine-tuned models slightly underperformed the baseline with ICL. Fine-tuned models often gained little from adding demonstrations at inference, and could perform worse with ICL than in the zero-shot setting; the relationship between shot count and fine-tuned performance was unclear. By contrast, the baseline model generally improved as more demonstrations were provided. The authors suggest that fine-tuning may reduce the model’s ability to benefit from new prompt contexts.

  8. Knowl 8 — In-domain gains from more fine-tuning examples are task-dependent

    empirical result

    Increasing the fine-tuning subset from 2,000 to 4,000 or 6,000 examples did not yield a consistent improvement in in-domain performance. On Paws, accuracy rose from 81.6% with 2,000 examples to 93.2% with 4,000, while the increase from 4,000 to 6,000 examples could be small or negative. On XSum, performance changed only slightly between the 2,000- and 6,000-example models. The experiments therefore show that the effect of training-set size depended on the task and dataset rather than increasing monotonically.

  9. Knowl 9 — FTICL did not improve classification fine-tuning in the same way as generation fine-tuning

    empirical result

    For classification tasks, FTICL models trained with demonstrations generally performed worse than vanilla fine-tuned models on the corresponding in-domain test sets. On out-of-domain same-task tests, FTICL usually outperformed the unfine-tuned baseline but lagged behind vanilla fine-tuning. Cross-task classification transfer was inconsistent; for example, Amazon fine-tuning helped Paws, whereas MNLI fine-tuning harmed Amazon. FTICL did outperform vanilla fine-tuning on cross-task generation evaluations from classification-trained models, suggesting it can partly alleviate output-space specialization. The authors hypothesize that classification demonstrations may distract the model because their labels can be copied instead of inferred from the relevant input; they also observed higher final training loss for classification FTICL than for vanilla fine-tuning. These explanations are proposed rather than established.

  10. Knowl 10 — The study leaves mechanisms and several fine-tuning settings unresolved

    limitation

    The paper identifies three limits to the scope of its findings: it does not adequately explain why fine-tuning has different generalization effects for classification and generation; it does not fully explain the operating mechanisms behind FTICL; and it does not evaluate more advanced tuning strategies such as reinforcement learning from human feedback. The reported conclusions therefore concern the tested task-specific fine-tuning and FTICL setups, rather than establishing how these mechanisms or other tuning methods behave.

Coverage note — Exact dataset examples and verbatim prompt templates are omitted because they illustrate the benchmark and formats rather than add separate experimental findings; the prompt-format effect itself is included.

References

  1. 1.
    1. Llama 2: Open foundation and fine-tuned chat models.
  2. 2.Cem Anil, Yuhuai Wu, Anders Johan Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Venkatesh Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022. Exploring length generalization in large language models. In Advances in Neural Information Processing Systems.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  4. 4.Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu. 2020. Recall and learn: Fine-tuning deep pretrained language models with less forgetting. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7870–7881, Online. Association for Computational Linguistics.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Sebastian Gehrmann, Zachary Ziegler, and Alexander Rush. 2019. Generating abstractive summaries with finetuned language models. In Proceedings of the 12th International Conference on Natural Language Generation, pages 516–522, Tokyo, Japan. Association for Computational Linguistics.
  8. 8.Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. The false promise of imitating proprietary llms.
  9. 9.Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, and Tianlu Wang. 2023. Understanding in-context learning via supportive pretraining data. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12660–12673, Toronto, Canada. Association for Computational Linguistics.
  10. 10.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XLsum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  11. 11.Karl Moritz Hermann, Tomáš Kocisk ˇ y, Edward Grefen- ` stette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Proceedings of the 28th International Conference on Neural Information Processing Systems-Volume 1, pages 1693–1701.
  12. 12.Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, Shuming Shi, and Zhaopeng Tu. 2023. Is chatgpt a good translator? yes with gpt-4 as the engine.
  13. 13.Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz. 2018. A dataset of peer reviews (peerread): Collection, insights and nlp applications. In Meeting of the North American Chapter of the Association for Computational Linguistics (NAACL), New Orleans, USA.
  14. 14.Phillip Keung, Yichao Lu, György Szarvas, and Noah A. Smith. 2020. The multilingual Amazon reviews corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4563–4568, Online. Association for Computational Linguistics.
  15. 15.Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2024. Understanding the effects of RLHF on LLM generalisation and diversity. In The Twelfth International Conference on Learning Representations.
  16. 16.Jannik Kossen, Yarin Gal, and Tom Rainforth. 2023. In-context learning learns label relationships but is not conventional learning.
  17. 17.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv., 55(9).
  18. 18.Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 61–68, Dublin, Ireland. Association for Computational Linguistics.
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  20. 20.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations.
  21. 21.Julian John McAuley and Jure Leskovec. 2013. From amateurs to connoisseurs: Modeling the evolution of user expertise through online reviews. In Proceedings of the 22nd International Conference on World Wide Web, WWW ’13, page 897–908, New York, NY, USA. Association for Computing Machinery.
  22. 22.Marius Mosbach, Tiago Pimentel, Shauli Ravfogel, Dietrich Klakow, and Yanai Elazar. 2023. Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation. In Findings of the Association for Computational Linguistics: ACL 2023, pages 12284–12314, Toronto, Canada. Association for Computational Linguistics.
  23. 23.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.
  24. 24.OpenAI. 2023. Gpt-4 technical report.
  25. 25.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  26. 26.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671, Seattle, United States. Association for Computational Linguistics.
  27. 27.Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. 2019. SocialIQA: Commonsense reasoning about social interactions. In EMNLP.
  28. 28.Timo Schick and Hinrich Schütze. 2021. It’s not just size that matters: Small language models are also few-shot learners. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2339–2352, Online. Association for Computational Linguistics.
  29. 29.Chufan Shi, Yixuan Su, Cheng Yang, Yujiu Yang, and Deng Cai. 2023. Specialist or generalist? instruction tuning for specific nlp tasks.
  30. 30.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020. Learning to summarize from human feedback. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA. Curran Associates Inc.
  31. 31.Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Ding, Hongyang Li, Mengzhe Geng, et al. 2023. A survey of reasoning with foundation models. arXiv preprint arXiv:2312.11562.
  32. 32.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, Brussels, Belgium. Association for Computational Linguistics.
  33. 33.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations.
  34. 34.Xiao Wang, Weikang Zhou, Can Zu, Han Xia, Tianze Chen, Yuansen Zhang, Rui Zheng, Junjie Ye, Qi Zhang, Tao Gui, Jihua Kang, Jingsheng Yang, Siyuan Li, and Chunsai Du. 2023a. Instructuie: Multi-task instruction tuning for unified information extraction.
  35. 35.Yifan Wang, Qingyan Guo, Xinzhe Ni, Chufan Shi, Lemao Liu, Haiyun Jiang, and Yujiu Yang. 2023b. Hint-enhanced in-context learning wakes large language models up for knowledge-intensive tasks. arXiv preprint arXiv:2311.01949.
  36. 36.Yihan Wang, Si Si, Daliang Li, Michal Lukasik, Felix Yu, Cho-Jui Hsieh, Inderjit S Dhillon, and Sanjiv Kumar. 2023c. Two-stage llm fine-tuning with less specialization and more generalization.
  37. 37.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  39. 39.Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. ArXiv, abs/1707.06209.
  40. 40.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  41. 41.Wenhan Xiong, Jiawei Wu, Hong Wang, Vivek Kulkarni, Mo Yu, Xiaoxiao Guo, Shiyu Chang, and William Yang Wang. 2019. Tweetqa: A social media focused question answering dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  42. 42.Chen Xu, Piji Li, Wei Wang, Haoran Yang, Siyun Wang, and Chuangbai Xiao. 2022. Cosplay: Concept set guided personalized dialogue generation across both party personas. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 201–211.
  43. 43.Haoran Yang, Wai Lam, and Piji Li. 2021. Contrastive representation learning for exemplar-guided paraphrase generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4754–4761, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  44. 44.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308, Minneapolis, Minnesota. Association for Computational Linguistics.
  45. 45.Yue Zhang, Leyang Cui, Deng Cai, Xinting Huang, Tao Fang, and Wei Bi. 2023. Multi-task instruction tuning of llama for specific scenarios: A preliminary study on writing assistance.
  46. 46.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. Lima: Less is more for alignment.

Citation

MLA
Yang, H., et al. “Unveiling the Generalization Power of Fine-Tuned Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 884–99, https://doi.org/10.18653/v1/2024.naacl-long.51.
APA
Yang, H., Zhang, Y., Xu, J., Lu, H., Heng, P.-A., & Lam, W. (2024). Unveiling the Generalization Power of Fine-Tuned Large Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 884–899. https://doi.org/10.18653/v1/2024.naacl-long.51
Chicago
Yang, H., Y. Zhang, J. Xu, H. Lu, P.-A. Heng, and W. Lam. 2024. “Unveiling the Generalization Power of Fine-Tuned Large Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 884–99. https://doi.org/10.18653/v1/2024.naacl-long.51.
Harvard
Yang, H. et al. (2024) “Unveiling the Generalization Power of Fine-Tuned Large Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 884–899. Available at: https://doi.org/10.18653/v1/2024.naacl-long.51.
Vancouver
1. Yang H, Zhang Y, Xu J, Lu H, Heng P-A, Lam W (2024) Unveiling the Generalization Power of Fine-Tuned Large Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 884–899

BibTeX

@inproceedings{yang-etal-2024-unveiling,
    title = "Unveiling the Generalization Power of Fine-Tuned Large Language Models",
    author = "Yang, Haoran  and
      Zhang, Yumeng  and
      Xu, Jiaqi  and
      Lu, Hongyuan  and
      Heng, Pheng-Ann  and
      Lam, Wai",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.51/",
    doi = "10.18653/v1/2024.naacl-long.51",
    pages = "884--899"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/