Fine-tuned Language Models are Continual Learners

Thomas ScialomTuhin ChakrabartySmaranda Muresan

article2022EMNLP196 citations

Demonstrates that instruction-tuned language models can sequentially acquire new generation tasks with minimal rehearsal while preventing catastrophic forgetting, identifying self-supervised pre-training as the primary driver of this capability.

Listen

Large language models fine-tuned on natural language instructions demonstrate impressive general capabilities, but they struggle to handle new tasks outside their initial training data. Upgrading these models typically requires fine-tuning them on new data, which often triggers catastrophic forgetting—a major failure mode where a model loses previously learned skills and knowledge. Consequently, organizations face high computational expenses and operational friction because they must retrain models from scratch across all historical and new data whenever requirements expand.

The article demonstrates that instruction-tuned language models can act as effective continual learners, sequentially acquiring diverse new capabilities without forgetting prior skills.

To test this capability, the authors introduced an approach called Continual Learning via Rehearsal, creating the Continual-T0 model based on the pre-trained T0 architecture. Rather than retraining from scratch, the system replays a small external memory buffer containing a fraction of past task data alongside new training instances. The evaluation tracked progressive sequential training across 8 new language generation tasks (such as text simplification, haiku writing, constrained headline writing, and COVID-19 question answering) while measuring performance across 50 original training datasets and 12 completely unseen zero-shot evaluation datasets across more than 1,000 gradient steps.

The article establishes several key findings. First, allocating just 1% of previous task data to the memory buffer prevented catastrophic forgetting, enabling the model to retain 98.0% of peak performance on the 3-billion parameter version and 99.8% on the 11-billion parameter version. Second, the model maintained stable performance on zero-shot evaluation datasets despite including no rehearsal data for those tasks. Third, the system outperformed existing lifelong learning baselines and established a new state of the art on text simplification benchmarks. Fourth, the model demonstrated zero-shot instruction compositionality, successfully combining distinct, separately learned skills—such as applying emotional styles to haiku generation or adhering to multiple lexical constraints. Finally, comparative tests showed that continual learning ability is driven primarily by self-supervised pre-training rather than model scale or instruction tuning alone.

These findings suggest that artificial intelligence systems can be upgraded modularly like open-source software releases instead of requiring exhaustive retraining. This shift offers substantial reductions in computing costs, energy use, and training timelines for deploying specialized enterprise models. The results challenge the assumption that continual learning requires complex architectural adjustments or enormous parameter scales, showing that standard pre-trained architectures inherently possess strong adaptability when paired with light rehearsal.

Organizations developing or updating natural language processing systems should adopt memory-buffered continual training workflows for task expansion. When designing updates, teams should reserve a small representative replay buffer (around 1%) from historical datasets rather than rebuilding models from ground zero. Before broad deployment, practitioners should pilot test zero-shot instruction combinations to verify that novel composed prompts perform reliably.

Confidence in these findings is supported by consistent multi-task benchmarks across both 3-billion and 11-billion parameter scales, as well as tests showing task order invariance. However, decision-makers should note certain limitations: the experiments were restricted to English-language text, evaluated only 8 sequential tasks, and relied predominantly on automated evaluation metrics rather than comprehensive human reviews.

  • Paper: Finetuned Language Models Are Zero-Shot Learners, Jason Wei et al. (2022). Read this account of instruction tuning first to understand the instruction-following foundation that the source adapts for sequential learning.
  • Paper: Experience Replay for Continual Learning, David Rolnick et al. (2018). Its experience-replay framework introduces the rehearsal principle that the source applies to preserve earlier task performance while learning new tasks.
Cover for Fine-tuned Language Models are Continual Learners

Abstract

Recent work on large language models relies on the intuition that most natural language processing tasks can be described via natural language instructions and that models trained on these instructions show strong zero-shot performance on several standard datasets. However, these models even though impressive still perform poorly on a wide range of tasks outside of their respective training and evaluation sets. To address this limitation, we argue that a model should be able to keep extending its knowledge and abilities, without forgetting previous skills. In spite of the limited success of Continual Learning we show that Fine-tuned Language Models can be continual learners. We empirically investigate the reason for this success and conclude that Continual Learning emerges from self-supervision pre-training. Our resulting model Continual-T0 (CT0) is able to learn 8 new diverse language generation tasks, while still maintaining good performance on previous tasks, spanning in total 70 datasets. Finally, we show that CT0 is able to combine instructions in ways it was never trained for, demonstrating some level of instruction compositionality.^1

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Continual Learning for Fine-tuned Language Models
  • 3.1 Continual Learning via Rehearsal (CLR)
  • 3.2 Continual-T0 (CT0)
  • 3.3 Tasks
  • 3.4 Automatic Metrics
  • 4 Results
  • 4.1 Learning a New Task at a time
  • 4.2 Learning a Sequence of New Tasks
  • 4.3 Zero-shot Instruction Compositionality
  • 5 Discussion
  • 5.1 Why could LLMs be lifelong learners?
  • 5.2 Toward Concept Drift
  • 5.3 Data Efficiency
  • 6 Conclusion
  • References
  • 7 Appendix
  • 7.1 Tasks Order
  • 7.2 Tasks
  • 7.3 Automatic Metrics
  • 7.3.1 New Tasks
  • 7.4 Automatic Metrics
  • 7.5 Evaluation for T0 Train Set
  • 7.6 Additional Results
  • 7.7 Implementation Details

Knowls

  1. Knowl 1 — Continual Learning via Rehearsal Formulation for Language Models

    equation

    To train an instruction-tuned language model sequentially on a stream of tasks without suffering catastrophic forgetting, Continual Learning via Rehearsal (CLR) augments the training data of each incoming task with a sampled replay buffer from prior tasks. Given an ordered sequence of NN tasks T=(T1,T2,…,TN)\mathcal{T} = (T_1, T_2, \dots, T_N) with associated datasets DiD_i for task TiT_i, the replay-augmented dataset DirD_i^r used to train the model at step ii is defined as:

    Dir=Di∪(⋃j=1i−1rDj)D_i^r = D_i \cup \left( \bigcup_{j=1}^{i-1} r D_j \right)

    where r∈[0,1]r \in [0, 1] is the rehearsal hyperparameter governing the fraction of training examples sampled from each preceding dataset D1,…,Di−1D_1, \dots, D_{i-1}. When r=0r = 0, the setup reduces to standard sequential fine-tuning with no replay memory; when r=1r = 1, the setup corresponds to full multitask learning over all cumulative data. In the Continual-T0 (CT0) implementation, the initial task T1T_1 comprises the 50 multitask instruction datasets used to pre-tune T0, and each dataset is standardized to 100,000 examples (via upsampling or subsampling) such that r=1%r = 1\% samples exactly 1,000 replay examples per task. Held-out zero-shot evaluation datasets are never included in the rehearsal buffer.

  2. Knowl 2 — Performance Retention of Continual-T0 Across Sequentially Learned Tasks

    data/table

    Continual-T0 (CT0) fine-tunes T0_3B (3 billion parameters) and T0pp (11 billion parameters) over 8 sequential natural language generation tasks using 1%1\% rehearsal memory (r=1%r=1\%). The model maintains near-complete retention of both original training datasets (T0tr) and held-out zero-shot classification datasets (T0zs), while matching the performance Upper Bound (UB) achieved by fine-tuning separate models on individual tasks in isolation. In contrast, the baseline continual learning method LAMOL suffers catastrophic forgetting across generative tasks.

    Model T0tr T0zs ASSET Simp HGen Haiku CQA InqQG EmDg Exp TwSt
    R1 Acc B4/SARI B4/SARI R1/Cons HcustH_{cust} BS 1Tok/BS BS BS Clf/BS
    T0_3B 49.8 48.2 70.1/41.0 12.8/41.1 33.6/32.2 34.2 47.6 2.1/58.7 48.6 32.7 54.4/38.0
    T0pp 54.2 65.6 56.5/37.7 11.7/40.1 34.9/35.9 31.6 46.0 2.4/59.8 49.7 37.2 66.4/45.1
    UB_3B 49.8 48.2 79.9/45.2 13.8/44.6 39.7/81.0 62.6 90.0 5.3/63.3 55.7 71.8 74.8/56.5
    UB_pp 54.2 65.6 85.3/46.1 15.0/44.8 41.9/86.9 63.9 90.0 4.9/65.7 56.6 73.5 74.4/57.9
    LAMOL 32.6 33.6 37.3/12.6 8.4/21.4 22.9/33.5 25.8 46.6 1.8/47.9 45.1 27.6 50.1/35.2
    CT03B 47.9 46.6 78.0/44.5 14.6/43.7 37.3/77.5 60.4 86.8 5.2/61.9 55.3 72.4 74.8/56.5
    CT0pp 53.7 64.4 85.9/46.6 14.6/44.7 40.7/85.5 65.8 89.8 4.8/65.2 56.2 73.0 74.4/57.9
    revfinal 48.1 48.8 83.3/45.4 14.6/43.9 39.0/81.6 61.2 88.6 4.4/61.9 55.0 72.4 73.2/57.3

    CT0pp retains 99.8%99.8\% of its UB score across all tasks (with no task dropping more than 2%2\% relative to UB), while CT03B retains 98.0%98.0\%. CT0 also establishes a state-of-the-art score on the ASSET sentence simplification benchmark (85.9 BLEU-4 / 46.6 SARI).

  3. Knowl 3 — Emergence of Continual Learning Capability from Self-Supervised Pre-Training

    empirical result

    To identify whether continual learning capability in large language models arises from parameter scale, massive multitask instruction fine-tuning, or self-supervised pre-training, an ablation was conducted across four architectures evaluated sequentially on 8 generation tasks using 1%1\% rehearsal memory:

    Model ASSET Simp HGen Haiku CQA InqQG EmDg Exp TwSt
    B4/SARI B4/SARI R1/Cons HcustH_{cust} BS 1Tok/BS BS BS Clf/BS
    UB_rand 0.5/24.3 0.0/29.6 1.5/0.1 9.6 25.2 1.2/25.4 36.3 33.1 24.7
    UB_T5small 87.8/45.9 15.6/43.2 35.3/67.8 53.4 54.1 3.4/57.0 51.3 33.8 52.4/54.6
    UB_T53b 87.0/45.6 15.4/43.7 33.0/89.4 63.0 89.9 2.92/61.5 55.3 71.6 75.6/55.4
    UB_T0 79.9/45.2 13.8/44.6 39.7/81.0 62.6 90.0 5.3/63.3 55.7 71.8 74.8/56.5
    CTrand 0.0/22.9 0.0/28.5 0.2/0.0 9.6 25.2 1.2/27.9 28.1 30.7 24.7
    CT5small 85.5/45.8 15.0/42.8 34.6/64.8 51.8 49.5 3.3/56.0 51.2 32.3 52.4/54.6
    CT53B 84.6/45.8 14.8/44.0 38.3/88.3 62.3 85.8 4.64/62.1 55.5 73.1 75.6/55.4
    CT03B 78.0/44.5 14.6/43.7 37.3/77.5 60.4 86.8 5.2/61.9 55.3 72.4 74.8/56.5

    The randomly initialized 3B transformer (CTrand) fails entirely to acquire or maintain skills across the sequence, remaining far below its own upper bound (UB_rand). T5-small (CT5small) maintains performance close to its upper bound (UB_T5small) across tasks (e.g., ASSET B4 of 85.5 vs 87.8), indicating that parameter scale is not the core requirement for continual retention. Continual T5-3B (CT53B) performs comparably to Continual T0-3B (CT03B) across all tasks without having undergone T0's prior 50-dataset instruction tuning. These findings demonstrate that continual learning ability emerges primarily from the initial self-supervised pre-training phase.

  4. Knowl 4 — Zero-Shot Instruction Compositionality in Continual-T0

    empirical result

    Continual-T0 (CT0) demonstrates zero-shot instruction compositionality by adhering to novel combinations of constraints and styles never presented together during sequential training:

    1. Multi-Constraint Headline Generation: CT0 was trained exclusively on single-keyword constraints (n=1n = 1). When evaluated on unseen instructions containing 2 or 3 keyword constraints (e.g., containing "X" and "Y" and "Z"), CT0 satisfies multi-keyword constraints at rates far exceeding unconstrained models:
    HGen (% Respected) TwSt (% Respected)
    # Constraints 1 2 3 1
    CT0 77.0 56.4 39.5 46.4
    CT0_NoCons 33.6 15.4 8.1 10.7

    ROUGE-1 on headline generation correspondingly increases as constraint count rises: 30.2 (NoCons), 38.9 (1 constraint), 43.9 (2 constraints), and 47.4 (3 constraints).

    1. Cross-Task Composition: Composing headline keyword constraints with Twitter stylometry ("Write a tweet about X, in the style of Y, containing Z") yields a 46.4%46.4\% constraint satisfaction rate compared to 10.7%10.7\% for CT0 without constraints.

    2. Emotional Haiku Generation: Composing emotional directives (learned in Empathetic Dialogue) with Haiku generation prompts ("Generate a haiku about [topic]. The associated emotion is [emotion].") shifts the binary sentiment classification of generated haikus from under 40%40\% positive for negative emotions ("sad", "terrified", "angry") to over 75%75\% positive for positive emotions ("content", "faithful", "grateful").

  5. Knowl 5 — Rehearsal Proportion Sensitivity for Preventing Catastrophic Forgetting

    empirical result

    Varying the rehearsal ratio r∈{0.0%,0.25%,1.0%}r \in \{0.0\%, 0.25\%, 1.0\%\} during continual learning on Headline Generation with Constraint, Text Simplification, and Haiku Generation reveals distinct memory retention thresholds:

    1. At r=0.0%r = 0.0\% (zero rehearsal), the model learns the target new task effectively but suffers immediate and total catastrophic forgetting on the T0 zero-shot evaluation suite (T0zs), dropping by −100%-100\% in relative performance across training steps.
    2. At r=0.25%r = 0.25\% replay, task performance on T0 zero-shot benchmarks stabilizes almost completely, resisting forgetting throughout gradient updates.
    3. At r=1.0%r = 1.0\% replay, performance on T0zs remains fully stationary with zero degradation relative to the initial checkpoint, while target new tasks reach their maximum upper bound performance within 80 to 134 training steps.

    Target task acquisition trajectories are invariant to the rehearsal ratio across r∈{0.0%,0.25%,1.0%}r \in \{0.0\%, 0.25\%, 1.0\%\}.

  6. Knowl 6 — Task Order Invariance in Rehearsal-Based Continual Learning

    empirical result

    Training Continual-T0 (3B) on the sequence of 8 generation tasks in reverse order (revfinal: Twitter Stylometry →\to Explanation Generation →\to Empathetic Dialogue →\to Inquisitive Question Generation →\to Covid QA →\to Haiku →\to Headline Generation with Constraint →\to Text Simplification) results in performance virtually identical to forward sequential training:

    • T0 zero-shot accuracy: 48.8 (reverse) vs. 46.6 (forward)
    • ASSET: 83.3 BLEU-4 / 45.4 SARI (reverse) vs. 78.0 BLEU-4 / 44.5 SARI (forward)
    • WikiAuto Simplification: 14.6 BLEU-4 / 43.9 SARI (reverse) vs. 14.6 BLEU-4 / 43.7 SARI (forward)
    • Constrained Headline Generation: 39.0 ROUGE-1 / 81.6% Constraint (reverse) vs. 37.3 ROUGE-1 / 77.5% Constraint (forward)
    • Haiku Generation: 61.2 HcustH_{cust} (reverse) vs. 60.4 HcustH_{cust} (forward)
    • Covid QA: 88.6 BERTScore (reverse) vs. 86.8 BERTScore (forward)
    • Inquisitive QG: 4.4 1Tok / 61.9 BERTScore (reverse) vs. 5.2 1Tok / 61.9 BERTScore (forward)
    • Empathetic Dialogue: 55.0 BERTScore (reverse) vs. 55.3 BERTScore (forward)
    • Explanation Generation: 72.4 BERTScore (reverse) vs. 72.4 BERTScore (forward)
    • Twitter Stylometry: 73.2% Clf / 57.3 BERTScore (reverse) vs. 74.8% Clf / 56.5 BERTScore (forward)

    These results confirm that continual fine-tuning via rehearsal is robustly invariant to task presentation order.

  7. Knowl 7 — Eight-Task Generative Benchmark Suite for Continual Instruction Tuning

    experimental setup

    To evaluate continual learning beyond standard classification and summarization, eight diverse natural language generation tasks were introduced and converted to natural language instruction format:

    1. Text Simplification (Simpl): Paraphrasing text into simpler language trained on WikiAuto (400,000 sentence pairs; 4,000 test) and evaluated on WikiAuto and ASSET (2,000 multi-reference examples).
    2. Headline Generation with Constraint (HGen): Generating a news headline from an article given a keyword constraint (XX) required at the beginning, at the end, or anywhere in the headline. Derived from Gigaword (1,951 test instances per constraint type).
    3. Haiku Generation (Haiku): Generating a 5-7-5 syllable structured poem from a topic prompt, trained on 9,742 pairs and tested on 974 pairs from Reddit r/haiku.
    4. Covid QA (CQA): Context-free biomedical question answering on COVID-19, framed as direct question-to-answer retrieval without source context to evaluate closed-book memory acquisition and concept drift (2,019 QA pairs).
    5. Inquisitive Question Generation (InqQG): Generating open-ended inquisitive questions from text passages, trained on 61,710 pairs and tested on 1,681 pairs filtered from Reddit ELI5 threads.
    6. Empathetic Dialogue Generation (EmDg): Generating contextually grounded empathetic conversation responses given an emotion tag, situation, and dialogue history (58,770 train, 8,396 test examples).
    7. Explanation Generation (Exp): Generating natural language explanations justifying NLI premise-hypothesis labels (entailment, contradiction, neutral) using e-SNLI (100,000 train, 9,824 test examples).
    8. Twitter Stylometry (TwSt): Generating author-stylized tweets given a hashtag and target author handle, constructed from top 20 Twitter accounts (13,041 train, 250 test examples).
  8. Knowl 8 — Automatic and Customized Metrics for Open-Ended Continual Generation

    experimental setup

    Evaluating open-ended continual generation requires specific automated and custom metrics tailored to each task's structure:

    • Zero-shot Classification (T0zs): Evaluated using exact-match accuracy between decoded greedy text and ground truth, rather than candidate log-likelihood ranking, to directly assess decoder generation fidelity.
    • T0 Training Set (T0tr): Evaluated using ROUGE-1 across 1,000 sampled test instances spanning both NLU and NLG tasks.
    • Standard NLG Metrics: BLEU-4 and SARI for text simplification; ROUGE-1 for headline generation; DeBERTa-MNLI BERTScore for open-domain tasks (CQA, InqQG, EmDg, Exp, TwSt).
    • Haiku Metric (HcustH_{cust}): Defined as the unweighted average of four components: line-count error penalty, syllable-count error penalty (against the 17-syllable 3-line standard), BLEU score against reference, and constraint satisfiability (presence of the topic prompt).
    • First-Word Distribution (1Tok): For InqQG, computed as the inverse Jensen-Shannon Divergence between the distributions of the opening tokens (e.g., "Why", "How" vs. "What") in generated versus ground-truth questions.
    • Author Classification (ClfClf): For Twitter Stylometry, accuracy of a Ridge Classifier trained on the 20-author corpus (achieving 81%81\% standalone accuracy) in predicting the specified target author from the generated tweet text.
    • Constraint Satisfaction (ConsCons): Percentage of generated outputs strictly containing the specified keyword in the required location (beginning, end, or anywhere).
  9. Knowl 9 — Methodological and Structural Limitations of Continual-T0

    limitation

    The Continual-T0 framework exhibits several specific methodological limitations:

    1. Memory Buffer Scaling: Continual Learning via Rehearsal requires maintaining historical data buffers (O(N)O(N) memory scaling over task count NN), which restricts applicability in scenarios subject to data privacy constraints or strict data retention limits where past raw samples cannot be archived.
    2. Monolingual Scope: The base models (T5 and T0) and all 8 benchmark generation tasks are English-only; continual learning behavior on multilingual corpora was not evaluated.
    3. Task Stream Length: Empirical validation is conducted over a sequence of 8 tasks (70 datasets total including pre-training); dynamics over hundreds or thousands of successive tasks remain unverified.
    4. Reliance on Automatic Metrics: Validation relies on automated and custom proxy metrics (HcustH_{cust}, ClfClf, BERTScore, SARI); human evaluations of fluency and instruction adherence were not performed due to computational and benchmark scale.

Coverage note — None omitted; all core methods, empirical evaluations, ablation experiments on pre-training emergence, instruction compositionality tests, metric designs, and stated limitations are included.

References

  1. 1.Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, and Lucia Specia. 2020. ASSET: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4668–4679, Online. Association for Computational Linguistics.
  2. 2.Magdalena Biesialska, Katarzyna Biesialska, and Marta R. Costa-jussà. 2020. Continual lifelong learning in natural language processing: A survey. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6523–6541, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018. e-snli: Natural language inference with natural language explanations. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc.
  5. 5.Yue Cao, Hao-Ran Wei, Boxing Chen, and Xiaojun Wan. 2021. Continual learning for neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3964–3974, Online. Association for Computational Linguistics.
  6. 6.Andrea Cossu, Tinne Tuytelaars, Antonio Carta, Lucia Passaro, Vincenzo Lomonaco, and Davide Bacciu. 2022. Continual pre-training mitigates forgetting in language and vision. arXiv preprint arXiv:2205.09357.
  7. 7.Cyprien de Masson D’Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019. Episodic memory in lifelong language learning. Advances in Neural Information Processing Systems, 32.
  8. 8.Arthur Douillard, Alexandre Ramé, Guillaume Couairon, and Matthieu Cord. 2021. Dytox: Transformers for continual learning with dynamic token expansion. arXiv preprint arXiv:2111.11326.
  9. 9.Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long form question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3558–3567, Florence, Italy. Association for Computational Linguistics.
  10. 10.Robert M French. 1999. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135.
  11. 11.Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu. 2020. Neural CRF model for sentence alignment in text simplification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7943–7960. Association for Computational Linguistics.
  12. 12.Xisen Jin, Junyi Du, Arka Sadhu, Ram Nevatia, and Xiang Ren. 2020. Visually grounded continual learning of compositional phrases. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2018–2029, Online. Association for Computational Linguistics.
  13. 13.Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, and Xiang Ren. 2021. Lifelong pretraining: Continually adapting language models to emerging corpora. arXiv preprint arXiv:2110.08534.
  14. 14.Zixuan Ke, Hu Xu, and Bing Liu. 2021. Adapting BERT for continual learning of a sequence of aspect sentiment classification tasks. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4746–4755, Online. Association for Computational Linguistics.
  15. 15.Bill Yuchen Lin, Sida Wang, Xi Lin, Robin Jia, Lin Xiao, Xiang Ren, and Scott Yih. 2022. On continual model refinement in out-of-distribution data streams. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3128–3139, Dublin, Ireland. Association for Computational Linguistics.
  16. 16.Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81.
  17. 17.Vincenzo Lomonaco and Davide Maltoni. 2017. Core50: a new dataset and benchmark for continuous object recognition. In Conference on Robot Learning, pages 17–26. PMLR.
  18. 18.Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, and Zhiguang Wang. 2021. Continual learning in task-oriented dialogue systems. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7452–7467, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Louis Martin, Angela Fan, Éric de la Clergerie, Antoine Bordes, and Benoît Sagot. 2020. Muss: multilingual unsupervised sentence simplification by mining paraphrases. arXiv preprint arXiv:2005.00352.
  20. 20.Sanket Vaibhav Mehta, Darshan Patil, Sarath Chandar, and Emma Strubell. 2021. An empirical investigation of the role of pre-training in lifelong learning. arXiv preprint arXiv:2112.09153.
  21. 21.Fei Mi, Liangwei Chen, Mengjie Zhao, Minlie Huang, and Boi Faltings. 2020. Continual learning for natural language generation in task-oriented dialog systems. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3461–3474, Online. Association for Computational Linguistics.
  22. 22.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  23. 23.Timo Möller, Anthony Reina, Raghavan Jayakumar, and Malte Pietsch. 2020. COVID-QA: A question answering dataset for COVID-19. In Proceedings of the 1st Workshop on NLP for COVID-19 at ACL 2020, Online. Association for Computational Linguistics.
  24. 24.Nihal V Nayak, Peilin Yu, and Stephen H Bach. 2022. Learning to compose soft prompts for compositional zero-shot learning. arXiv preprint arXiv:2204.03574.
  25. 25.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL ’02, pages 311–318, Philadelphia, Pennsylvania. ACL.
  26. 26.German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. 2019. Continual lifelong learning with neural networks: A review. Neural Networks, 113:54–71.
  27. 27.Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al. 2011. Scikit-learn: Machine learning in python. the Journal of machine Learning research, 12:2825–2830.
  28. 28.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  30. 30.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
  31. 31.Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer. 2021. Effect of scale on catastrophic forgetting in neural networks. In International Conference on Learning Representations.
  32. 32.Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open-domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5370–5381, Florence, Italy. Association for Computational Linguistics.
  33. 33.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. 2017. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 2001–2010.
  34. 34.Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, and Jacopo Staiano. 2021. Synthetic data augmentation for zero-shot cross-lingual question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7016–7030.
  35. 35.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2022. Multitask prompted training enables zero-shot task generalization. In International Conference on Learning Representations.
  36. 36.Thomas Scialom and Jacopo Staiano. 2020. Ask to learn: A study on curiosity-driven question generation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2224–2235.
  37. 37.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017. Continual learning with deep generative replay. Advances in neural information processing systems, 30.
  38. 38.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  39. 39.Tejas Srinivasan, Ting-Yun Chang, Leticia Leonor Pinto Alva, Georgios Chochlakis, Mohammad Rostami, and Jesse Thomason. 2022. Climb: A continual learning benchmark for vision-and-language tasks. arXiv preprint arXiv:2206.09059.
  40. 40.Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee. 2019. Lamol: Language modeling for lifelong language learning. In International Conference on Learning Representations.
  41. 41.Bin Tareaf. 2017. R.: Tweets dataset-top 20 most followed users in twitter social platform. Harvard Dataverse, 2.
  42. 42.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  43. 43.Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016. Optimizing Statistical Machine Translation for Text Simplification. Transactions of the Association for Computational Linguistics, 4:401–415.
  44. 44.Wenpeng Yin, Jia Li, and Caiming Xiong. 2022. ConTinTin: Continual learning from task instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3062–3072, Dublin, Ireland. Association for Computational Linguistics.
  45. 45.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with bert. In International Conference on Learning Representations.

Citation

MLA
Scialom, T., et al. “Fine-tuned Language Models Are Continual Learners”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6107–22, https://doi.org/10.18653/v1/2022.emnlp-main.410.
APA
Scialom, T., Chakrabarty, T., & Muresan, S. (2022). Fine-tuned Language Models are Continual Learners. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6107–6122. https://doi.org/10.18653/v1/2022.emnlp-main.410
Chicago
Scialom, T., T. Chakrabarty, and S. Muresan. 2022. “Fine-tuned Language Models Are Continual Learners”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6107–22. https://doi.org/10.18653/v1/2022.emnlp-main.410.
Harvard
Scialom, T., Chakrabarty, T. and Muresan, S. (2022) “Fine-tuned Language Models are Continual Learners”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6107–6122. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.410.
Vancouver
1. Scialom T, Chakrabarty T, Muresan S (2022) Fine-tuned Language Models are Continual Learners. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6107–6122

BibTeX

@inproceedings{scialom-etal-2022-fine,
    title = "Fine-tuned Language Models are Continual Learners",
    author = "Scialom, Thomas  and
      Chakrabarty, Tuhin  and
      Muresan, Smaranda",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.410/",
    doi = "10.18653/v1/2022.emnlp-main.410",
    pages = "6107--6122"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/