Fine-tuned Language Models are Continual Learners
Thomas ScialomTuhin ChakrabartySmaranda Muresan
Demonstrates that instruction-tuned language models can sequentially acquire new generation tasks with minimal rehearsal while preventing catastrophic forgetting, identifying self-supervised pre-training as the primary driver of this capability.
Large language models fine-tuned on natural language instructions demonstrate impressive general capabilities, but they struggle to handle new tasks outside their initial training data. Upgrading these models typically requires fine-tuning them on new data, which often triggers catastrophic forgetting—a major failure mode where a model loses previously learned skills and knowledge. Consequently, organizations face high computational expenses and operational friction because they must retrain models from scratch across all historical and new data whenever requirements expand.
The article demonstrates that instruction-tuned language models can act as effective continual learners, sequentially acquiring diverse new capabilities without forgetting prior skills.
To test this capability, the authors introduced an approach called Continual Learning via Rehearsal, creating the Continual-T0 model based on the pre-trained T0 architecture. Rather than retraining from scratch, the system replays a small external memory buffer containing a fraction of past task data alongside new training instances. The evaluation tracked progressive sequential training across 8 new language generation tasks (such as text simplification, haiku writing, constrained headline writing, and COVID-19 question answering) while measuring performance across 50 original training datasets and 12 completely unseen zero-shot evaluation datasets across more than 1,000 gradient steps.
The article establishes several key findings. First, allocating just 1% of previous task data to the memory buffer prevented catastrophic forgetting, enabling the model to retain 98.0% of peak performance on the 3-billion parameter version and 99.8% on the 11-billion parameter version. Second, the model maintained stable performance on zero-shot evaluation datasets despite including no rehearsal data for those tasks. Third, the system outperformed existing lifelong learning baselines and established a new state of the art on text simplification benchmarks. Fourth, the model demonstrated zero-shot instruction compositionality, successfully combining distinct, separately learned skills—such as applying emotional styles to haiku generation or adhering to multiple lexical constraints. Finally, comparative tests showed that continual learning ability is driven primarily by self-supervised pre-training rather than model scale or instruction tuning alone.
These findings suggest that artificial intelligence systems can be upgraded modularly like open-source software releases instead of requiring exhaustive retraining. This shift offers substantial reductions in computing costs, energy use, and training timelines for deploying specialized enterprise models. The results challenge the assumption that continual learning requires complex architectural adjustments or enormous parameter scales, showing that standard pre-trained architectures inherently possess strong adaptability when paired with light rehearsal.
Organizations developing or updating natural language processing systems should adopt memory-buffered continual training workflows for task expansion. When designing updates, teams should reserve a small representative replay buffer (around 1%) from historical datasets rather than rebuilding models from ground zero. Before broad deployment, practitioners should pilot test zero-shot instruction combinations to verify that novel composed prompts perform reliably.
Confidence in these findings is supported by consistent multi-task benchmarks across both 3-billion and 11-billion parameter scales, as well as tests showing task order invariance. However, decision-makers should note certain limitations: the experiments were restricted to English-language text, evaluated only 8 sequential tasks, and relied predominantly on automated evaluation metrics rather than comprehensive human reviews.
- Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Read this account of instruction tuning first to understand the instruction-following foundation that the source adapts for sequential learning.
- Paper: Experience Replay for Continual Learning, David Rolnick et al. (2018). Its experience-replay framework introduces the rehearsal principle that the source applies to preserve earlier task performance while learning new tasks.
- Paper: InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions, Yifan Wang et al. (2024). InsCL extends rehearsal-based continual instruction tuning by choosing replay examples according to task similarity and instruction complexity.
- Paper: An Empirical Investigation of the Role of Pre-training in Lifelong Learning, Sanket Vaibhav Mehta et al. (2023). This later empirical study tests and broadens the source’s finding that pretrained initialization helps models resist forgetting during sequential learning.
