Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora

Xisen JinDejiao ZhangHenghui ZhuWei XiaoShang-Wen LiXiaokai WeiAndrew O. ArnoldXiang Ren

article2022NAACL145 citations

Establishes a continual pretraining framework for language models across evolving domains and temporal data streams, demonstrating that distillation-based methods effectively prevent catastrophic forgetting on earlier tasks while improving adaptation to new data.

Listen

Organizations deploying natural language processing models face significant performance degradation over time as language, topics, and real-world conditions shift. Pretrained language models are traditionally trained on massive, static text corpora and then adapted to specific tasks. When new domains or temporal shifts emerge, maintaining these systems typically requires either expensive complete retraining from scratch or naive continuous updates that cause catastrophic forgetting, where the model loses its ability to handle earlier domains.

The article establishes a lifelong language model pretraining framework to evaluate how continuously updating a single model on incoming, unlabeled text streams impacts downstream performance. It specifically tests the model's ability to retain knowledge from past domains, adapt to the latest incoming data, and generalize across temporal gaps where downstream training data is outdated relative to evaluation data.

The researchers designed two real-world data streams using transformer models: a domain-incremental stream composed of millions of academic papers across four disciplines (biomedical, computer science, materials science, and physics) and a chronological stream comprising 100 million tweets across four separate years (2014, 2016, 2018, and 2020). Using these streams, they evaluated baseline approaches against multiple continual learning techniques, including parameter expansion (adapters), episodic memory replay, parameter regularization, and several knowledge distillation methods that penalize output or representation discrepancies between previous and updated model checkpoints.

The analysis revealed several critical findings. First, distillation-based continual learning methods—particularly output logit distillation—proved the most effective at preserving knowledge from earlier domains, improving downstream F1 scores by at least 1.0% to over 1.5% over sequential updating on earlier tasks and matching or exceeding offline multi-task retraining. Second, standard episodic memory replay largely failed to prevent forgetting and led to overfitting on cached examples, even when memory size was increased a hundredfold from 100,000 to 10 million instances. Third, continual pretraining improved performance on the latest temporal data and boosted temporal generalization across time gaps (such as applying models trained on 2016 data to 2020 test sets), outperforming models trained exclusively on current-year data. Finally, lifelong pretraining demonstrated its greatest relative performance gains in low-resource downstream fine-tuning scenarios where labeled target data is scarce.

These findings indicate that organizations do not need to choose between the prohibitive compute and storage costs of recurring full-corpus retraining and the performance degradation of naive continuous updating. Continuous pretraining using knowledge distillation delivers a single, robust model capable of maintaining legacy capabilities while adapting to new domains and temporal evolution, which significantly reduces compute overhead, storage requirements, and privacy risks associated with retaining historic data. However, distillation introduces a trade-off: strongly penalizing representational drift can slightly constrain the model's adaptability when absorbing radically distinct new domains.

Decision-makers should consider adopting distillation-based continual pretraining pipelines over simple replay buffers or parameter-regularization techniques when managing evolving text corpora. Teams should prioritize logit-distillation methods as the baseline continual learning strategy for production pipelines. Before large-scale deployment, organizations should conduct domain-similarity assessments; when domain shifts involve large vocabulary divergence (as in distinct scientific fields rather than social media), engineering teams should calibrate distillation regularization weights to avoid overly rigid model behavior on new tasks.

The primary limitations of the article involve task-dependent variations among advanced distillation methods, where no single contrastive or representation-based variant uniformly outperformed standard logit distillation across all settings. Furthermore, continuous pretraining risks compounding historical biases present in earlier streams unless targeted bias-mitigation techniques are applied. The core findings are supported by consistent evidence across distinct base model sizes (including base and large architectures) and diverse domain and temporal datasets, providing high confidence in distillation as a viable foundation for lifelong language model maintenance.

arXiv: 2110.08534
  • Paper: Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks, Suchin Gururangan et al. (2020). This paper establishes domain-adaptive and task-adaptive pretraining for language models, providing the baseline pretraining paradigms that lifelong pretraining directly adapts to continual data streams.
  • Paper: Learning without Forgetting, Zhizhong Li et al. (2016). This foundational work introduces distillation-based regularization for continual learning without past data, which serves as the core mechanism evaluated in the source for retaining downstream performance across domains.
  • Paper: Efficient Lifelong Learning with A-GEM, Arslan Chaudhry et al. (2018). This paper introduces gradient-based continual learning protocols and memory constraints that underpin standard evaluation setups in lifelong pretraining.
  • Paper: Dark Experience for General Continual Learning: a Strong, Simple Baseline, Pietro Buzzega et al. (2020). This paper proposes Dark Experience Replay to align output distributions over streaming data, representing a key distillation and replay baseline adapted in continuous pretraining setups.
  • Paper: A Continual Learning Survey: Defying Forgetting in Classification Tasks, Matthias De Lange et al. (2019). This comprehensive taxonomy outlines the stability-plasticity trade-offs and core algorithmic families that structure the lifelong pretraining methods evaluated in the source.
  • Paper: Continual Learning Through Synaptic Intelligence, Friedemann Zenke et al. (2017). This foundational work develops path-integral parameter regularization to prevent catastrophic forgetting, establishing the parameter-regularization baselines benchmarked against distillation.
Cover for Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora

Abstract

Pretrained language models (PTLMs) are typically learned over a large, static corpus and further fine-tuned for various downstream tasks. However, when deployed in the real world, a PTLM-based model must deal with data distributions that deviate from what the PTLM was initially trained on. In this paper, we study a lifelong language model pretraining challenge where a PTLM is continually updated so as to adapt to emerging data. Over a domain-incremental research paper stream and a chronologically-ordered tweet stream, we incrementally pretrain a PTLM with different continual learning algorithms, and keep track of the downstream task performance (after fine-tuning). We evaluate PTLM’s ability to adapt to new corpora while retaining learned knowledge in earlier corpora. Our experiments show distillation-based approaches to be most effective in retaining downstream performance in earlier domains. The algorithms also improve knowledge transfer, allowing models to achieve better downstream performance over the latest data, and improve temporal generalization when distribution gaps exist between training and evaluation because of time. We believe our problem formulation, methods, and analysis will inspire future studies towards continual pretraining of language models.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulation
  • 2.1 Lifelong Pretraining of PTLMs
  • 2.2 Data Streams & Downstream Datasets
  • 2.3 Evaluation Protocol
  • 3 Methods
  • 3.1 Simple Baselines
  • 3.2 Model-expansion and Regularization-based Methods
  • 3.3 Memory Replay Methods
  • 3.4 Distillation-based CL Methods
  • 4 Results
  • 4.1 Experiment Settings
  • 4.2 Domain Incremental Data Stream
  • 4.3 Temporal Data Stream
  • 5 Related Works
  • 6 Conclusion
  • References
  • A Detailed Experiment Settings
  • B Low-Resource Fine-Tuning
  • C Full Results over the Tweet Stream
  • D Dataset Details
  • E Details of Continual Learning Algorithms
  • E.1 Contrastive Distillation
  • E.2 SEED Distillation
  • F Analysis and Controlled Experiments of Computational Costs
  • G Experiments with RoBERTa-large
  • H Experiments with BERT on Tweet Stream After 2019
  • I Analysis of Data Streams
  • J Ethic Risks

Knowls

  1. Knowl 1 — Lifelong language model pretraining task

    definition

    Lifelong language model pretraining studies a language model that receives a sequence of unlabeled corpora D1,dots,DTD_1, dots,D_T, where the data distribution may change because of domain or time. After training on corpus DtD_t, the model becomes checkpoint ftf_t. The update on DtD_t cannot access the complete earlier corpora D1,dots,Dt−1D_1, dots,D_{t-1}, although it may use a bounded memory or information retained in an earlier model checkpoint. The initial model f0f_0 is RoBERTa-base in this study.

    The purpose of each checkpoint fif_i is to provide a useful initialization for downstream tasks associated with any domain already visited. During downstream fine-tuning, fif_i has no access to the pretraining corpora. The study evaluates three distinct utilities: knowledge retention on earlier domains, adaptation to tasks from the latest domain, and temporal generalization when a downstream model is trained on an earlier time period but tested on a later one.

  2. Knowl 2 — Two-stream benchmark for continual pretraining

    experimental setup

    The evaluation testbed contains two sequential unlabeled-text streams and associated downstream tasks.

    • Domain-incremental research-paper stream: four corpora are presented in the order biomedical science, computer science, materials science, and physics. The corpora contain approximately 6.6M, 12.1M, 7.8M, and 7.5M research papers, respectively. Biomedical evaluation uses ChemProt and RCT-Sample with micro-F1; computer-science evaluation uses ACL-ARC and SciERC with macro-F1; materials-science evaluation uses Synthesis with macro-F1 and MNER with micro-F1; physics evaluation uses Keyphrase and Hyponym with macro-F1. This stream primarily measures retention of downstream performance across domains.

    • Chronologically ordered tweet stream: four corpora contain tweets from 2014, 2016, 2018, and 2020, with 25M tweets per year. A separate 1M tweets from each year are used to create balanced downstream datasets: multi-label prediction of the 200 most frequent hashtags, scored by label-ranking average precision, and 20-way emoji prediction, scored by macro-F1. This stream measures adaptation to recent data and temporal generalization.

    The paper evaluates each continually pretrained checkpoint by fine-tuning it separately on the relevant downstream training data and measuring performance on the corresponding test data.

  3. Knowl 3 — Continual-learning baselines and model-update strategies

    model/method

    The study compares ordinary pretraining with several continual-learning strategies. RoBERTa-base without stream-specific pretraining is the no-adaptation baseline. Task-specific pretraining trains an independent model on each corpus. Sequential pretraining updates one model through D1,dots,DTD_1, dots,D_T without an explicit continual-learning mechanism, enabling transfer but risking catastrophic forgetting. Multi-task learning (MTL) shuffles all currently available corpora and repeatedly trains on them offline, requiring storage and reuse of earlier data.

    The continual-learning methods are grouped as follows:

    • Model expansion: Adapter training adds a separate trainable adapter layer at every transformer layer for each domain while keeping the main model parameters frozen. Layer Expansion instead learns domain-specific top transformer layers and a prediction head.
    • Parameter regularization: Online EWC penalizes changes to model parameters judged important to earlier domains.
    • Experience Replay (ER): a fixed memory MM stores a balanced sample from all domains seen so far. The memory has 100k examples by default, and one memory minibatch is replayed every 10 training steps while the full model is updated.
    • Distillation-based continual learning: one frozen copy of the previous checkpoint ft−1f_{t-1} supplies output or representation targets while the current checkpoint ftf_t learns from the current stream and, where available, replay-memory examples.
  4. Knowl 4 — Distillation objectives for continual pretraining

    algorithm

    For a current-domain or replay-memory minibatch xx, the student model ftf_t is trained with masked-language-modeling loss ell_{rm MLM} plus a distillation loss from the frozen teacher ft−1f_{t-1}:

    ell=ell_{rm MLM}+alphaell_{rm KD},

    where alpha controls the distillation strength. The paper evaluates four forms of distillation:

    • Logit distillation: let yty_t and yt−1y_{t-1} denote the student and teacher output distributions derived from their language-model logits. The loss is the Kullback Leibler divergence D_{rm KL}(y_t,y_{t-1}).

    • Representation distillation: if a sentence has NN tokens and htih_t^i and ht−1ih_{t-1}^i are the token representations before the prediction head, the loss is

      ell_{rm rep}=left t extstylesum_{i=1}^{N}left t extstyle|h_t^i-h_{t-1}^iright t extstyle|_2^2.

    • Contrastive distillation: sentence representations are trained with an unsupervised SimCSE objective. For a minibatch of nn examples, normalized sentence representations hih_i produce a similarity distribution

      B_{ij}=frac{exp(h_i^top h_j/tau)}{sum_{k=1}^{n}exp(h_i^top h_k/tau)},

      where tau is a temperature. The student matrix BtB_t is matched to the teacher matrix Bt−1B_{t-1} using

      ell_{rm contrast}=-frac{1}{n}sum_{i=1}^{n}sum_{j=1}^{n}B_{t-1,ij}log B_{t,ij}.

      The reported teacher and student temperatures are 0.050.05 and 0.010.01, respectively.

    • SEED distillation: the same type of similarity matching is performed between the current minibatch and a larger fixed-size queue QQ of examples from the current domain, allowing the teacher student relationship to be regularized against more examples than fit in one minibatch. SEED-Logit-KD combines this objective with logit distillation.

    The default distillation weight is alpha=1.0 for logit distillation and alpha=0.1 for the other distillation variants.

  5. Knowl 5 — Pretraining and evaluation configuration

    experimental setup

    All primary experiments initialize RoBERTa-base from its publicly pretrained weights, use maximum sequence length 128, and use an effective pretraining batch size of 2,048. On the research-paper stream, training lasts 8,000 steps on the first domain and 4,000 steps on each later domain. The main experimental description reports 4,000 steps per tweet domain, while the detailed settings report 8,000 steps for every tweet domain. Training examples are generally visited fewer than once because the corpora are large.

    The learning rate decreases linearly from 5×10−45\times10^{-4} on the research-paper stream and 3×10−43\times10^{-4} on the tweet stream. A held-out set of 128,000 sentences per corpus is used for masked-language-modeling validation. Replay and distillation normally use one memory minibatch every 10 training steps. Continual-learning hyperparameters are tuned using the first one or two domains, and downstream scores are averaged over repeated runs with the reported variability expressed as pm standard deviation.

  6. Knowl 6 — Distillation reduces forgetting in the research-paper stream

    empirical result

    On the final research-paper checkpoint f4f_4, sequential pretraining transfers knowledge to later domains but forgets the earliest biomedical tasks. For example, the final sequential model scores 82.09pm0.5 on ChemProt and 79.60pm0.5 on RCT-Sample, while logit distillation scores 83.39pm0.4 and 81.21pm0.1, respectively. On the computer-science tasks, sequential pretraining scores 72.73pm2.9 on ACL-ARC and 81.43pm0.8 on SciERC; logit distillation raises these to 73.70pm3.4 and 81.92pm0.8. Thus, distillation improves retention and transfer most clearly for domains early in the stream.

    The gains are not uniform over later domains. Sequential pretraining reaches 83.99pm0.3 on MNER and 92.10pm1.0 on Synthesis, compared with 83.96pm0.3 and 92.20pm1.0 for logit distillation. On the physics tasks, sequential pretraining scores 67.57pm1.0 and 74.68pm4.4, whereas logit distillation scores 64.75pm1.1 and 71.29pm3.6. SEED-Logit-KD is competitive with or better than ordinary logit distillation on some individual tasks, including SciERC at 83.03pm0.6, but it is not uniformly best.

    Sequential pretraining generally trails offline MTL on earlier domains, whereas logit-based continual learning can exceed MTL on tasks from the first and second domains. Online EWC gives little downstream improvement, adapters help mainly on biomedical tasks, and ER provides little benefit beyond sequential pretraining despite replaying earlier examples. The paper attributes the weak ER result partly to overfitting to the finite replay memory.

  7. Knowl 7 — Continual pretraining improves recent-data adaptation and temporal generalization

    empirical result

    On the tweet stream, continual pretraining improves both downstream performance on recent years and performance when training and testing occur in different years. Hashtag prediction uses label-ranking average precision, and emoji prediction uses macro-F1.

    For hashtag prediction on 2018 and 2020 data, the RoBERTa-base baseline scores 48.08pm1.0 and 56.42pm0.2, while sequential pretraining scores 56.79pm0.5 and 59.85pm0.4. Logit-KD further reaches 58.21pm0.5 and 60.52pm0.2, and SEED-Logit-KD reaches 57.75pm0.4 and 60.74pm0.6. For emoji prediction on the same years, sequential pretraining scores 29.30pm0.1 and 27.69pm0.1, while SEED-KD reaches 30.12pm0.1 and SEED-Logit-KD reaches 29.98pm0.1 and 27.84pm0.2.

    Temporal generalization is evaluated by fine-tuning on 2014 or 2016 and testing on 2020. Sequential pretraining improves hashtag scores from the RoBERTa-base values of 39.31pm2.7 and 42.23pm2.7 to 44.00pm1.1 and 49.87pm1.8. SEED-Logit-KD reaches 45.35pm0.6 and 51.56pm0.7. For cross-year emoji prediction, sequential pretraining reaches 14.20pm0.2 and 16.08pm1.4, while Contrast-KD reaches 14.42pm0.1 and 17.52pm0.1. These improvements show that continual pretraining can transfer information from earlier time periods rather than merely adapting to the final year.

  8. Knowl 8 — Lifelong pretraining is especially useful in low-resource fine-tuning

    empirical result

    The benefit of continual pretraining is larger when downstream fine-tuning has few labeled examples. With 100, 200, and 500 ChemProt training instances, RoBERTa-base obtains micro-F1 scores of 47.147.1, 55.055.0, and 70.770.7, whereas logit distillation obtains 53.453.4, 63.663.6, and 74.374.3. With the same numbers of SciERC training instances, RoBERTa-base obtains macro-F1 scores of 29.529.5, 40.140.1, and 63.963.9, while logit distillation obtains 42.042.0, 52.952.9, and 70.170.1.

    The comparison shows that continual pretraining supplies a more useful initialization when downstream supervision is scarce. The improvement narrows as the amount of downstream supervision increases, although the continually pretrained models remain competitive at 500 examples.

  9. Knowl 9 — Distillation requires additional computation that cannot be replaced by arbitrary extra training

    empirical result

    The paper measures computation by the number of forward and backward passes over the language model. For bb stream minibatches and replay every kk steps, sequential pretraining requires bb forward and bb backward passes, for total cost C=2bC=2b. ER requires (1+1/k)b(1+1/k)b of each type, giving C=(2+2/k)bC=(2+2/k)b. Logit or representation distillation adds a frozen-teacher forward pass, giving C=(3+3/k)bC=(3+3/k)b. SEED-Logit-KD additionally performs a contrastive-learning pass, giving C=(5+5/k)bC=(5+5/k)b.

    With k=10k=10, the reported wall times for 4,000 training steps after the first domain are approximately 4.0×1044.0\times10^4 seconds for sequential pretraining, 4.2×1044.2\times10^4 seconds for ER, 6.9×1046.9\times10^4 seconds for logit distillation, and 9.7×1049.7\times10^4 seconds for SEED-Logit-KD. Controlled-cost experiments show that simply training sequential pretraining for 1.2 times as many steps does not improve the latest-domain score and can worsen earlier-domain scores. Increasing replay frequency similarly increases overfitting to the replay memory. Reducing distillation frequency to match the cost of cheaper methods removes much of the distillation benefit, so performance cannot be exchanged for computation through a simple generic scaling rule.

  10. Knowl 10 — Transfer depends on stream similarity and distillation has task-dependent limits

    limitation

    The two streams have very different vocabulary-distribution shifts. Pairwise cosine distances between vocabulary distributions are on the order of 10−210^{-2} for the research-paper stream, with reported values ranging from 1.17×10−21.17\times10^{-2} to 6.58×10−26.58\times10^{-2}, but on the order of 10−510^{-5} for the tweet stream, with values ranging from 6.83×10−66.83\times10^{-6} to 2.54×10−52.54\times10^{-5}. The smaller and temporally ordered tweet shift makes transfer to recent data easier than transfer between unrelated research fields, helping explain why continual-learning methods improve latest-domain tweet performance but do not consistently improve the latest research-paper domain.

    Additional RoBERTa-large experiments show that the broad pattern is not limited to RoBERTa-base: logit distillation scores 86.18pm0.7 on ChemProt, 81.93pm0.7 on RCT-Sample, and 83.23pm0.6 on SciERC, outperforming the sequential baseline on those tasks. However, sequential pretraining is better on ACL-ARC than the tested continual-learning methods, with 73.44pm2.0 versus 72.10pm2.0 for logit distillation. The paper therefore concludes that no tested distillation variant consistently dominates logit distillation: the relative value of logit, representation, contrastive, and SEED distillation is task-dependent, and distillation can hinder acquisition of new knowledge when domains differ substantially.

Coverage note — The auxiliary BERT-on-post-2019-tweets experiment and the paper's ethical-risk discussion were omitted because they provide supporting analyses rather than core additions to the lifelong-pretraining formulation, methods, and primary evaluation.

References

  1. 1.Kristjan Arumae, Qing Sun, and Parminder Bhatia. 2020. An empirical investigation towards efficient multi-domain language model pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4854–4864, Online. Association for Computational Linguistics.
  2. 2.Isabelle Augenstein, Mrinal Das, Sebastian Riedel, Lakshmi Vikraman, and Andrew McCallum. 2017. SemEval 2017 task 10: ScienceIE - extracting keyphrases and relations from scientific publications. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 546–555, Vancouver, Canada. Association for Computational Linguistics.
  3. 3.Abdul Hameed Azeemi and Adeel Waheed. 2021. Covid-19 tweets analysis through transformer language models. ArXiv, abs/2103.00199.
  4. 4.Francesco Barbieri, Jose Camacho-Collados, Francesco Ronzano, Luis Espinosa-Anke, Miguel Ballesteros, Valerio Basile, Viviana Patti, and Horacio Saggion. 2018. SemEval 2018 task 2: Multilingual emoji prediction. In Proceedings of The 12th International Workshop on Semantic Evaluation, pages 24–33, New Orleans, Louisiana. Association for Computational Linguistics.
  5. 5.Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A pretrained language model for scientific text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3615–3620, Hong Kong, China. Association for Computational Linguistics.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  7. 7.Hyuntak Cha, Jaeho Lee, and Jinwoo Shin. 2021. Co2l: Contrastive continual learning. ArXiv, abs/2106.14413.
  8. 8.Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. 2019. On tiny episodic memories in continual learning. arXiv preprint arXiv:1902.10486.
  9. 9.Yung-Sung Chuang, Shang-Yu Su, and Yun-Nung Chen. 2020. Lifelong language knowledge distillation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2914–2924, Online. Association for Computational Linguistics.
  10. 10.Cyprien de Masson d’Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019. Episodic memory in lifelong language learning. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13122–13131.
  11. 11.Franck Dernoncourt and Ji Young Lee. 2017. PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 308–313, Taipei, Taiwan. Asian Federation of Natural Language Processing.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisen­schlos, D. Gillick, Jacob Eisenstein, and William W. Cohen. 2021. Time-aware language models as temporal knowledge bases. ArXiv, abs/2106.15110.
  14. 14.Zhiyuan Fang, Jianfeng Wang, Lijuan Wang, L. Zhang, Yezhou Yang, and Zicheng Liu. 2021. Seed: Self-supervised distillation for visual representation. ArXiv, abs/2101.04731.
  15. 15.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. ArXiv, abs/2104.08821.
  16. 16.Yuyun Gong and Qi Zhang. 2016. Hashtag recommendation using attention-based convolutional neural network. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, IJCAI 2016, New York, NY, USA, 9-15 July 2016, pages 2782–2788. IJCAI/AAAI Press.
  17. 17.Suchin Gururangan, Michael Lewis, Ari Holtzman, Noah A. Smith, and Luke Zettlemoyer. 2021. Demix layers: Disentangling domains for modular language modeling. ArXiv, abs/2108.05036.
  18. 18.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  19. 19.Geoffrey E. Hinton, Oriol Vinyals, and J. Dean. 2015. Distilling the knowledge in a neural network. ArXiv, abs/1503.02531.
  20. 20.Spurthi Amba Hombaiah, Tao Chen, Mingyang Zhang, Michael Bendersky, and Marc-Alexander Najork. 2021. Dynamic language models for continuously evolving content. Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining.
  21. 21.Saihui Hou, Xinyu Pan, Chen Change Loy, Zilei Wang, and Dahua Lin. 2018. Lifelong learning via progressive distillation and retrospection. In ECCV.
  22. 22.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 2790–2799. PMLR.
  23. 23.Xiaolei Huang and Michael J. Paul. 2018. Examining temporality in document classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 694–699, Melbourne, Australia. Association for Computational Linguistics.
  24. 24.Yufan Huang, Yanzhe Zhang, Jiaao Chen, Xuezhi Wang, and Diyi Yang. 2021. Continual learning for text classification with information disentanglement based regularization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2736–2746, Online. Association for Computational Linguistics.
  25. 25.Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2021. Towards continual knowledge learning of language models. arXiv preprint arXiv:2110.03215.
  26. 26.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174, Online. Association for Computational Linguistics.
  27. 27.David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky. 2018. Measuring the evolution of a scientific field through citation frames. Transactions of the Association for Computational Linguistics, 6:391–406.
  28. 28.Kasidis Kanwatchara, Thanapapas Horsuwan, Piyawat Lertvittayakumjorn, Boonserm Kijsirikul, and Peerapon Vateekul. 2021. Rational LAMOL: A rationale-based lifelong learning framework. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2942–2953, Online. Association for Computational Linguistics.
  29. 29.Angeliki Lazaridou, A. Kuncoro, E. Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Sebastian Ruder, Dani Yogatama, Kris Cao, Tomás Kociský, Susannah Young, and P. Blunsom. 2021. Assessing temporal generalization in neural language models. NeurIPS.
  30. 30.Zhizhong Li and Derek Hoiem. 2018. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40:2935–2947.
  31. 31.Tianlin Liu, Lyle Ungar, and João Sedoc. 2019a. Continual learning for sentence representations using conceptors. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3274–3279, Minneapolis, Minnesota. Association for Computational Linguistics.
  32. 32.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692.
  33. 33.Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. S2ORC: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4969–4983, Online. Association for Computational Linguistics.
  34. 34.Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018. Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3219–3232, Brussels, Belgium. Association for Computational Linguistics.
  35. 35.Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A. Smith. 2021. Time waits for no one! analysis and challenges of temporal misalignment.
  36. 36.Antonis Maronikolakis and Hinrich Schütze. 2021. Multidomain pretrained language models for green NLP. In Proceedings of the Second Workshop on Domain Adaptation for NLP, pages 1–8, Kyiv, Ukraine. Association for Computational Linguistics.
  37. 37.Michael Matena and Colin Raffel. 2021. Merging models with fisher-weighted averaging. ArXiv, abs/2111.09832.
  38. 38.Martin Müller, Marcel Salathé, and Per Egil Kummervold. 2020. Covid-twitter-bert: A natural language processing model to analyse covid-19 content on twitter. ArXiv, abs/2005.07503.
  39. 39.Sheshera Mysore, Zachary Jensen, Edward Kim, Kevin Huang, Haw-Shiuan Chang, Emma Strubell, Jeffrey Flanigan, Andrew McCallum, and Elsa Olivetti. 2019. The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures. In Proceedings of the 13th Linguistic Annotation Workshop, pages 56–64, Florence, Italy. Association for Computational Linguistics.
  40. 40.Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020. BERTweet: A pre-trained language model for English tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 9–14, Online. Association for Computational Linguistics.
  41. 41.Elsa A Olivetti, Jacqueline M Cole, Edward Kim, Olga Kononova, Gerbrand Ceder, Thomas Yong-Jin Han, and Anna M Hiszpanski. 2020. Data-driven materials research enabled by natural language processing and information extraction. Applied Physics Reviews, 7(4):041317.
  42. 42.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487–503, Online. Association for Computational Linguistics.
  43. 43.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. 2017. icarl: Incremental classifier and representation learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 5533–5542. IEEE Computer Society.
  44. 44.Shruti Rijhwani and Daniel Preotiuc-Pietro. 2020. Temporally-informed analysis of named entity recognition. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7605–7617, Online. Association for Computational Linguistics.
  45. 45.Anthony V. Robins. 1995. Catastrophic forgetting, rehearsal and pseudorehearsal. Connect. Sci., 7:123–146.
  46. 46.Paul Röttger and J. Pierrehumbert. 2021. Temporal adaptation of bert and performance on downstream document classification: Insights from social media. Findings of EMNLP.
  47. 47.Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. 2018. Progress & compress: A scalable framework for continual learning. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 4535–4544. PMLR.
  48. 48.Yangyang Shi, Martha Larson, and Catholijn M. Jonker. 2015. Recurrent neural network language model adaptation with curriculum learning. Comput. Speech Lang., 33:136–154.
  49. 49.Fan-Keng Sun, Cheng-Hao Ho, and Hung-Yi Lee. 2020. LAMOL: language modeling for lifelong language learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  50. 50.Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019. Patient knowledge distillation for BERT model compression. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4323–4332, Hong Kong, China. Association for Computational Linguistics.
  51. 51.Jens Vindahl. 2016. Chemprot-3.0: a global chemical biology diseases mapping.
  52. 52.Hong Wang, Wenhan Xiong, Mo Yu, Xiaoxiao Guo, Shiyu Chang, and William Yang Wang. 2019. Sentence embedding alignment for lifelong relation extraction. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 796–806, Minneapolis, Minnesota. Association for Computational Linguistics.
  53. 53.Zirui Wang, Sanket Vaibhav Mehta, Barnabas Poczos, and Jaime Carbonell. 2020. Efficient meta lifelong-learning with limited memory. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 535–548, Online. Association for Computational Linguistics.
  54. 54.Tongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li, Guilin Qi, and Gholamreza Haffari. 2022. Pre-trained language model in continual learning: A comparative study. In International Conference on Learning Representations.
  55. 55.Benfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang, Hongtao Xie, and Yongdong Zhang. 2020. Curriculum learning for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6095–6104, Online. Association for Computational Linguistics.
  56. 56.Yunzhi Yao, Shaohan Huang, Wenhui Wang, Li Dong, and Furu Wei. 2021. Adapt-and-distill: Developing small, fast and effective pretrained language models for domains. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 460–470, Online. Association for Computational Linguistics.

Citation

MLA
Jin, X., et al. “Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 4764–80, https://doi.org/10.18653/v1/2022.naacl-main.351.
APA
Jin, X., Zhang, D., Zhu, H., Xiao, W., Li, S.-W., Wei, X., Arnold, A., & Ren, X. (2022). Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4764–4780. https://doi.org/10.18653/v1/2022.naacl-main.351
Chicago
Jin, X., D. Zhang, H. Zhu, et al. 2022. “Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4764–80. https://doi.org/10.18653/v1/2022.naacl-main.351.
Harvard
Jin, X. et al. (2022) “Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 4764–4780. Available at: https://doi.org/10.18653/v1/2022.naacl-main.351.
Vancouver
1. Jin X, Zhang D, Zhu H, Xiao W, Li S-W, Wei X, Arnold A, Ren X (2022) Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 4764–4780

BibTeX

@inproceedings{jin-etal-2022-lifelong-pretraining,
    title = "Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora",
    author = "Jin, Xisen  and
      Zhang, Dejiao  and
      Zhu, Henghui  and
      Xiao, Wei  and
      Li, Shang-Wen  and
      Wei, Xiaokai  and
      Arnold, Andrew  and
      Ren, Xiang",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.351/",
    doi = "10.18653/v1/2022.naacl-main.351",
    pages = "4764--4780"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/