Lifelong Language Pretraining with Distribution-Specialized Experts

Wuyang ChenYanqi ZhouNan DuYanping HuangJames LaudonZhifeng ChenClaire Cui

article2023ICML85 citations

Proposes Lifelong-MoE, an extensible mixture-of-experts pretraining framework that continually adapts large language models to streaming data distributions and prevents catastrophic forgetting by progressively expanding and freezing specialized experts without increasing inference computation.

Listen

Large language models typically rely on static, high-quality pretraining datasets, but real-world textual data arrives in continuous, evolving streams such as new web pages, multilingual content, and social media discussions. Sequentially updating models on changing text distributions usually causes catastrophic forgetting, where the model overwrites previously acquired capabilities to fit new data. Completely retraining large models from scratch whenever new data appears is computationally prohibitive, while traditional continual learning approaches focus primarily on adapting models across specific downstream tasks rather than handling continuous, task-agnostic pretraining.

The article demonstrates an extensible framework, termed Lifelong-MoE, that enables large language models to continually pretrain on streaming data distributions without forgetting prior knowledge. It evaluates how dynamically expanding modular expert networks and applying targeted regularizations preserves past learning while maintaining constant computational costs during inference and training.

The authors designed an experimental evaluation using a sequential stream of diverse data distributions comprising hundreds of billions of tokens: English web text and Wikipedia, non-English multilingual text, and public conversational data. Building upon sparsely activated mixture-of-experts architectures—where only the top two expert modules are active per token—the approach progressively adds new specialized experts and gating units for new data distributions, freezes previously trained experts and their gating mechanisms, and applies output distillation regularization to shared dense layers. The evaluated models ranged up to 1.878 billion activated parameters and were benchmarked against standard dense architectures, memory replay strategies, and parameter regularization baselines across nineteen downstream language understanding and generation benchmarks.

The investigation produced several key findings. First, Lifelong-MoE significantly reduced catastrophic forgetting; after completing sequential training across all distributions, its performance dropped by only 39.9% on question answering and 15.3% on translation, compared to drops of 48.5% and 72.7% respectively for standard parameter regularization. Second, the proposed method achieved superior or competitive overall downstream performance, reaching a translation score of 19.16 compared to 11.14 for an oracle dense model trained on all data simultaneously. Third, ablation experiments revealed that explicit protection requires freezing both old experts and their corresponding gating dimensions together, as freezing either component alone degraded performance. Finally, a controlled partial expansion strategy (expanding experts from 4 to 7 to 10) outperformed naive duplication of experts while avoiding exponential growth in total parameter storage.

These findings indicate that modular, sparsely activated neural network architectures offer a practical path for maintaining up-to-date language models. Organizations can incorporate new domain data continuously rather than incurring the substantial financial, energy, and hardware costs of full retraining cycles. Because the architecture keeps the number of activated parameters per token constant, deployment inference latency and per-token compute costs remain unaffected even as total model capacity grows.

Engineering and research teams maintaining production language models should adopt modular expert expansion and selective parameter freezing strategies when integrating streaming, out-of-domain text corpora. When designing expert growth paths, teams should implement partial expert allocation rather than doubling modules to prevent excessive memory footprints, and calibrate output regularization scaling to maintain training stability. Future work should investigate automated criteria for deciding when a distribution shift warrants adding new experts, evaluate scaling behavior on models beyond the two-billion parameter regime, and test the architecture across longer sequences of data distributions.

The article provides strong empirical confidence through systematic ablations and extensive multi-task evaluations across diverse benchmarks. However, the evaluation assumes clear, distinct phase boundaries between training distributions rather than gradual, mixed distribution shifts. Practitioners should exercise caution when deploying the framework in settings with highly uncurated, continuous data mixtures where distinct distribution boundaries are difficult to identify.

arXiv: 2305.12281

No sufficiently relevant recommendations were found.

Cover for Lifelong Language Pretraining with Distribution-Specialized Experts

Abstract

Large-scale pretrained language models have achieved great success in various natural language processing tasks. However, they still suffer from catastrophic forgetting when adapted to new domains/tasks. To address this issue, we propose Lifelong Language Pretraining with Distribution-Specialized Experts (LLP-DSE), which continually adapts a pretrained model to multiple domains using specialized expert modules while avoiding interference between them.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Pretraining MoE without Forgetting
  • 3.1. Model Architecture
  • 3.2. Progressive Expert Expansion
  • 3.3. Expert/Gating Regularization
  • 4. Experiment Setup
  • 4.1. Training Datasets
  • 4.2. Architecture Setting
  • 4.3. Hyperparameters
  • 4.4. Pretraining Procedure
  • 4.5. Downstream Evaluations
  • 5. Experiments
  • 5.1. Lifelong Pretraining
  • 5.2. Ablation Study
  • 5.3. Lifelong-MoE Mitigates Forgetting Issues in Downstream Tasks
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Influence of different distributions on downstream decoding performance.

Knowls

  1. Knowl 1 — Lifelong-MoE adapts to shifting pretraining distributions by adding sparse capacity

    model/method

    Lifelong-MoE is a task-agnostic method for continuing language-model pretraining on a sequence of changing corpus distributions. At each distribution shift, it expands the mixture-of-experts (MoE) layers with additional experts and corresponding gating dimensions, while keeping the network’s depth and width unchanged. Previously learned experts and gates can retain distribution-specific knowledge, while the shared parts of the model continue to learn across distributions. Each token still activates only two experts per MoE layer, so adding experts increases model capacity without increasing the number of experts used to process an individual token.

  2. Knowl 2 — The underlying GLaM-style MoE uses token-wise top-two expert routing

    model/method

    The method is built on a sparsely activated Transformer in which every other Transformer layer replaces its feed-forward component with an MoE layer. Each MoE layer contains independent feed-forward expert networks and a learned softmax gating function whose output dimension matches the number of experts. For each input token, the gate selects the two highest-scoring experts; the same top-two routing is used during training and inference. Consequently, increasing the expert pool increases the model’s available capacity but does not increase the number of experts activated for each token.

  3. Knowl 3 — Expert expansion is partial and initialized from pretrained parameters

    model/method

    At a distribution shift, Lifelong-MoE adds only a subset of possible experts and gating dimensions rather than duplicating the whole expert pool, and it does not add dense layers. The authors report that random initialization of added experts and gates performed poorly, which they attribute to mismatched gradient directions and magnitudes relative to the pretrained dense and attention layers. Instead, new expert and gating parameters are initialized from pretrained ones, following a Net2WiderNet-style strategy. The reported expansion schedules include 4→7→10 experts for the smaller model and 16→28→32 for the large model across three successive distributions.

  4. Knowl 4 — Freezing old experts and gates preserves their learned parameters while shared layers adapt

    model/method

    When training on a new distribution, Lifelong-MoE can freeze previously trained experts and their gating dimensions, while optimizing the newly added experts and gates. Dense and attention layers are shared across distributions and remain trainable, allowing them to fit all distributions rather than being assigned to one distribution and frozen. In the reported ablation, freezing only experts or only gates was less effective than freezing both.

  5. Knowl 5 — Output-level distillation regularizes the expanded model

    equation

    Lifelong-MoE optionally combines next-token prediction loss with an output-level distillation loss. For a current training corpus XX, let xix_i be the model input at token position ii, let x0:ix_{0:i} be its preceding context, and let xi+1x_{i+1} be the next-token target. Let pold(xi)\mathbf p_{\mathrm{old}}(x_i) be the output probability vector from the model using the old experts and shared layers, and pnew(xi)\mathbf p_{\mathrm{new}}(x_i) the output vector from the expanded model; both vectors are distributions over the vocabulary. The logarithm of a probability vector is applied elementwise. The objective is

    L=Lnext+λLout,Lnext=−∑i:xi∈Xlog⁡pnew(xi+1∣x0:i),Lout=−∑i:xi∈Xpold(xi)Tlog⁡pnew(xi).\mathcal{L}=\mathcal{L}_{\mathrm{next}}+\lambda\mathcal{L}_{\mathrm{out}},\qquad \mathcal{L}_{\mathrm{next}}=-\sum_{i:x_i\in X}\log p_{\mathrm{new}}(x_{i+1}\mid x_{0:i}),\qquad \mathcal{L}_{\mathrm{out}}=-\sum_{i:x_i\in X}\mathbf p_{\mathrm{old}}(x_i)^{\mathsf T}\log\mathbf p_{\mathrm{new}}(x_i).

    Here, λ\lambda controls the distillation term’s weight. The authors describe this output loss as a KL-divergence regularizer that discourages the expanded model from departing too far from the previous model’s outputs. In the reported four-expert ablation, λ=1\lambda=1 improved TriviaQA F1 over λ=0\lambda=0, while values larger than 1 made pretraining unstable.

  6. Knowl 6 — Experiments use a deliberately shifting three-distribution pretraining stream

    experimental setup

    The lifelong-pretraining stream is ordered A→B→C. Distribution A contains Wikipedia (19% of the A mixture; 3 billion tokens) and filtered webpages (81%; 143 billion tokens). Distribution B is the i18n non-English corpus (366 billion tokens), and distribution C contains public-domain conversations (174 billion tokens). The distributions were chosen to create substantial shifts: A is associated with English question answering, B with translation, and C with conversational and other language-understanding tasks. Models are trained for 500,000 steps on each distribution, restoring the previous checkpoint before training on the next. Evaluation includes one-shot decoding, formed by concatenating one randomly selected training example and the evaluation example with two newlines; TriviaQA and WMT16 are used for generation, alongside 19 language-understanding tasks.

  7. Knowl 7 — Lifelong-MoE reduces forgetting during sequential pretraining

    empirical result

    During A→B→C pretraining, evaluation on distributions A and B showed sharp transition-related drops in next-token accuracy and increases in perplexity for the fixed-capacity baseline. Lifelong-MoE reduced these changes, indicating better retention of earlier-distribution performance. The comparison favored the baseline in capacity: it used 10 experts per expert layer throughout, whereas Lifelong-MoE expanded from 4 to 7 to 10 experts. In some phases, Lifelong-MoE performed better on distribution A despite having fewer experts than the baseline.

  8. Knowl 8 — Ablations identify the effective regularization and expansion choices

    empirical result

    The ablation evaluated TriviaQA F1 after pretraining on A→B→C. With four experts throughout and no freezing, F1 was 5.93 at output-regularization weight λ=0\lambda=0, 5.64 at λ=0.1\lambda=0.1, and 6.96 at λ=1\lambda=1. With expansion 4→8→16 and no output regularization, F1 was 6.90; freezing experts alone gave 6.39, freezing gates alone gave 6.82, and freezing both gave 6.98. With both experts and gates frozen and λ=1\lambda=1, expansion 4→5→6 yielded 5.82, while 4→7→10 yielded 7.06. Thus, the tested partial expansion 4→7→10 outperformed naive doubling 4→8→16 while using fewer experts, and the best reported ablation combined partial expansion, freezing both old experts and gates, and output regularization.

  9. Knowl 9 — Final Lifelong-MoE is competitive with dense, replay, and GLaM baselines

    empirical result

    After lifelong pretraining, the reported results compare TriviaQA F1, WMT16 BLEU, Ubuntu score, and the average score on 19 NLU tasks. Lifelong-MoE scored 20.22, 19.16, 27, and 50.26, respectively. Dense GShard with online L2 regularization scored 12.99, 5.66, 27, and 48.65; dense GShard with memory replay scored 14.18, 7.54, 26, and 48.65; the jointly pretrained dense oracle scored 21.25, 11.14, 26, and 49.03; and GLaM scored 21.76, 6.97, 26, and 50.9. Lifelong-MoE had the highest reported BLEU and Ubuntu scores, exceeded the dense oracle on the 19-task NLU average, and had a competitive but not highest TriviaQA F1. The authors note that GLaM’s higher TriviaQA result is associated with starting with more experts on distribution A.

  10. Knowl 10 — Lifelong-MoE retains more decoding performance across distribution shifts

    empirical result

    The sequential downstream results report TriviaQA F1 and WMT16 BLEU after training on A, A→B, and A→B→C. After A, online L2 regularization scored 25.23 F1 and 2.84 BLEU, memory replay scored 25.23 and 2.84, and Lifelong-MoE scored 33.66 and 4.41. After A→B, the respective scores were 17 (−32.6%) and 20.77 for online L2 regularization; 12.23 (−51.5%) and 12.34 for memory replay; and 26.81 (−20.4%) and 22.63 for Lifelong-MoE. After A→B→C, they were 12.99 (−48.5%) and 5.66 (−72.7%); 14.18 (−43.7%) and 7.54 (−38.8%); and 20.22 (−39.9%) and 19.16 (−15.3%), respectively. The parenthesized values are the reported performance drops. Lifelong-MoE achieved the strongest final scores on both tasks and the smallest reported final-stage drops.

Coverage note — No substantial contributed material was omitted; detailed optimizer, precision, and hardware settings are excluded as implementation-level experimental particulars.

References

  1. 1.Adiwardana, D., Luong, M., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., and Le, Q. V. Towards a human-like open-domain chatbot. CoRR, abs/2001.09977, 2020. URL https://arxiv.org/abs/2001.09977.
  2. 2.Ahrens, K., Abawi, F., and Wermter, S. Drill: Dynamic representations for imbalanced lifelong learning. In International Conference on Artificial Neural Networks, pp. 409–420. Springer, 2021.
  3. 3.Aljundi, R., Caccia, L., Belilovsky, E., Caccia, M., Lin, M., Charlin, L., and Tuytelaars, T. Online continual learning with maximally interfered retrieval. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Red Hook, NY, USA, 2019. Curran Associates Inc. URL https://dl.acm.org/doi/abs/10.5555/3454287.3455350.
  4. 4.Biesialska, M., Biesialska, K., and Costa-jussà, M. R. Continual lifelong learning in natural language processing: A survey. In Proceedings of the 28th International Conference on Computational Linguistics, pp. 6523–6541, Barcelona, Spain (Online), 2020. International Committee on Computational Linguistics. doi: 10.18653/v1/2020.coling-main.574. URL https://aclanthology.org/2020.coling-main.574.
  5. 5.Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huck, M., Yepes, A. J., Koehn, P., Logacheva, V., Monz, C., et al. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pp. 131–198, 2016.
  6. 6.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing Systems, pp. 1877–1901. Curran Associates, Inc.
  7. 7.Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M. Efficient lifelong learning with a-gem. arXiv preprint arXiv:1812.00420, 2018.
  8. 8.Chen, T., Goodfellow, I., and Shlens, J. Net2net: Accelerating learning via knowledge transfer. arXiv preprint arXiv:1511.05641, 2015.
  9. 9.Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020.
  10. 10.Dai, A. M. and Le, Q. V. Semi-supervised sequence learning. In Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., and Garnett, R. (eds.), Advances in Neural Information Processing Systems. Curran Associates, Inc.
  11. 11.d’Autume, C. d. M., Ruder, S., Kong, L., and Yogatama, D. Episodic memory in lifelong language learning. In Wallach, H., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32, pp. 13143–13152. Curran Associates, Inc., 2019. URL http://papers.nips.cc/paper/9471-episodic-memory-in-lifelong-language-learning.pdf.
  12. 12.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019.
  13. 13.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pp. 5547–5569. PMLR, 2022.
  14. 14.Fedus, W., Zoph, B., and Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. CoRR, abs/2101.03961, 2021. URL https://arxiv.org/abs/2101.03961.
  15. 15.Gordon, A., Kozareva, Z., and Roemmele, M. SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012), pp. 394–398, Montréal, Canada, 7-8 June 2012. Association for Computational Linguistics. URL https://aclanthology.org/S12-1052.
  16. 16.Greco, C., Plank, B., Fernández, R., and Bernardi, R. Psycholinguistics meets continual learning: Measuring catastrophic forgetting in visual question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 3601–3605, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1350. URL https://www.aclweb.org/anthology/P19-1350.
  17. 17.Gururangan, S., Lewis, M., Holtzman, A., Smith, N. A., and Zettlemoyer, L. Demix layers: Disentangling domains for modular language modeling. arXiv preprint arXiv:2108.05036, 2021.
  18. 18.Hestness, J., Narang, S., Ardalani, N., Diamos, G. F., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. Deep learning scaling is predictable, empirically. CoRR, abs/1712.00409, 2017. URL http://arxiv.org/abs/1712.00409.
  19. 19.Holla, N., Mishra, P., Yannakoudakis, H., and Shutova, E. Meta-learning with sparse experience replay for lifelong language learning. arXiv preprint arXiv:2009.04891, 2020.
  20. 20.Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. Parameter-efficient transfer learning for NLP. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp. 2790–2799. PMLR, 09–15 Jun 2019. URL https://proceedings.mlr.press/v97/houlsby19a.html.
  21. 21.Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M. X., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E. B., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 103–112, 2019.
  22. 22.Huang, Y., Zhang, Y., Chen, J., Wang, X., and Yang, D. Continual learning for text classification with information disentanglement based regularization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2736–2746, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.218. URL https://aclanthology.org/2021.naacl-main.218.
  23. 23.Hussain, A., Holla, N., Mishra, P., Yannakoudakis, H., and Shutova, E. Towards a robust experimental framework and benchmark for lifelong language learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2021.
  24. 24.Jin, X., Lin, B. Y., Rostami, M., and Ren, X. Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 714–729, Punta Cana, Dominican Republic, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.62. URL https://aclanthology.org/2021.findings-emnlp.62.
  25. 25.Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1601–1611, Vancouver, Canada, 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1147. URL https://aclanthology.org/P17-1147.
  26. 26.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  27. 27.Kirkpatrick, J., Pascanu, R., Rabinowitz, N. C., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114:3521 – 3526, 2017.
  28. 28.Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S. Skip-thought vectors. In Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., and Garnett, R. (eds.), Advances in Neural Information Processing Systems. Curran Associates, Inc.
  29. 29.Kudo, T. and Richardson, J. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In EMNLP, 2018.
  30. 30.Kudugunta, S., Huang, Y., Bapna, A., Krikun, M., Lepikhin, D., Luong, M.-T., and Firat, O. Beyond distillation: Task-level mixture-of-experts for efficient inference. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 3577–3599, 2021.
  31. 31.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. GShard: Scaling giant models with conditional computation and automatic sharding. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=qrwe7XHTmYb.
  32. 32.Li, Z. and Hoiem, D. Learning without forgetting. IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017.
  33. 33.Li, Z. and Hoiem, D. Learning without forgetting. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12):2935–2947, 2018.
  34. 34.Lin, B. Y., Wang, S., Lin, X. V., Jia, R., Xiao, L., Ren, X., and Yih, W.-t. On continual model refinement in out-of-distribution data streams. arXiv preprint arXiv:2205.02014, 2022.
  35. 35.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019.
  36. 36.Lopez-Paz, D. and Ranzato, M. Gradient episodic memory for continual learning. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 6467–6476.
  37. 37.Mallya, A. and Lazebnik, S. Packnet: Adding multiple tasks to a single network by iterative pruning. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7765–7773, 2018.
  38. 38.Mikolov, T., Karafiát, M., Burget, L., Cernocký, J. H., and Khudanpur, S. Recurrent neural network based language model. In INTERSPEECH, 2010.
  39. 39.Mikolov, T., Chen, K., Corrado, G., and Dean, J. Efficient estimation of word representations in vector space. In Bengio, Y. and LeCun, Y. (eds.), 1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings, 2013. URL http://arxiv.org/abs/1301.3781.
  40. 40.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I.
  41. 41.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
  42. 42.Rebuffi, S., Kolesnikov, A., Sperl, G., and Lampert, C. H. icarl: Incremental classifier and representation learning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pp. 5533–5542. IEEE Computer Society, 2017. doi: 10.1109/CVPR.2017.587. URL https://doi.org/10.1109/CVPR.2017.587.
  43. 43.Robins, A. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science, 7(2):123–146, 1995.
  44. 44.Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T. P., and Wayne, G. Experience replay for continual learning. In Wallach, H., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pp. 348–358.
  45. 45.Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R. Progressive neural networks. ArXiv preprint, abs/1606.04671, 2016. URL https://arxiv.org/abs/1606.04671.
  46. 46.Shazeer, N. and Stern, M. Adafactor: Adaptive learning rates with sublinear memory cost. ArXiv, abs/1804.04235, 2018.
  47. 47.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=B1ckMDqlg.
  48. 48.Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B. Mesh-tensorflow: Deep learning for supercomputers. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, pp. 10435–10444, Red Hook, NY, USA, 2018. Curran Associates Inc.
  49. 49.Shin, H., Lee, J. K., Kim, J., and Kim, J. Continual learning with deep generative replay. In Guyon, I., von Luxburg, U., Bengio, S., Wallach, H. M., Fergus, R., Vishwanathan, S. V. N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 2990–2999.
  50. 50.Sun, F., Ho, C., and Lee, H. LAMOL: language modeling for lifelong language learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020a. URL https://openreview.net/forum?id=Skgxcn4YDS.
  51. 51.Sun, F., Ho, C., and Lee, H. LAMOL: language modeling for lifelong language learning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020b. URL https://openreview.net/forum?id=Skgxcn4YDS.
  52. 52.Sun, F.-K., Ho, C.-H., and Lee, H.-Y. LAMOL: LAnguage MOdeling for Lifelong Language Learning. In International Conference on Learning Representations (ICLR), 2020c. URL https://openreview.net/forum?id=Skgxcn4YDS.
  53. 53.Sutskever, I., Martens, J., and Hinton, G. Generating text with recurrent neural networks. In Proceedings of the 28th International Conference on International Conference on Machine Learning, ICML’11, pp. 1017–1024, Madison, WI, USA, 2011. Omnipress. ISBN 9781450306195.
  54. 54.Sutskever, I., Vinyals, O., and Le, Q. V. Sequence to sequence learning with neural networks. In Advances in neural information processing systems, pp. 3104–3112, 2014.
  55. 55.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems. Curran Associates, Inc.
  56. 56.Wang, H., Xiong, W., Yu, M., Guo, X., Chang, S., and Wang, W. Y. Sentence embedding alignment for lifelong relation extraction. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 796–806, Minneapolis, Minnesota, June 2019a. Association for Computational Linguistics. doi: 10.18653/v1/N19-1086. URL https://www.aclweb.org/anthology/N19-1086.
  57. 57.Wang, H., Xiong, W., Yu, M., Guo, X., Chang, S., and Wang, W. Y. Sentence embedding alignment for lifelong relation extraction. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 796–806, Minneapolis, Minnesota, 2019b. Association for Computational Linguistics. doi: 10.18653/v1/N19-1086. URL https://aclanthology.org/N19-1086.
  58. 58.Wen, Y., Tran, D., and Ba, J. Batchensemble: an alternative approach to efficient ensemble and lifelong learning. In International Conference on Learning Representations (ICLR), 2020. URL https://openreview.net/forum?id=Sklf1yrYDr.
  59. 59.Xu, H., Liu, B., Shu, L., and Yu, P. BERT post-training for review reading comprehension and aspect-based sentiment analysis. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2324–2335, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1242. URL https://aclanthology.org/N19-1242.
  60. 60.Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R. R., and Le, Q. V. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32, 2019.
  61. 61.Yoon, J., Yang, E., Lee, J., and Hwang, S. J. Lifelong learning with dynamically expandable networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018. URL https://openreview.net/forum?id=Sk7KsfW0-.
  62. 62.Zhang, S., Liu, X., Liu, J., Gao, J., Duh, K., and Durme, B. V. Record: Bridging the gap between human and machine commonsense reading comprehension. CoRR, abs/1810.12885, 2018.
  63. 63.Zhou, G., Sohn, K., and Lee, H. Online incremental feature learning with denoising autoencoders. In Artificial intelligence and statistics, pp. 1453–1461. PMLR, 2012.

Citation

MLA
Chen, W., et al. “Lifelong Language Pretraining with Distribution-Specialized Experts”. International Conference on Machine Learning, vol. 202, 2023, pp. 5383–95, https://proceedings.mlr.press/v202/chen23aq.html.
APA
Chen, W., Zhou, Y., Du, N., Huang, Y., Laudon, J., Chen, Z., & Cui, C. (2023). Lifelong Language Pretraining with Distribution-Specialized Experts. International Conference on Machine Learning, 202, 5383–5395. https://proceedings.mlr.press/v202/chen23aq.html
Chicago
Chen, W., Y. Zhou, N. Du, et al. 2023. “Lifelong Language Pretraining with Distribution-Specialized Experts”. International Conference on Machine Learning 202: 5383–95. https://proceedings.mlr.press/v202/chen23aq.html.
Harvard
Chen, W. et al. (2023) “Lifelong Language Pretraining with Distribution-Specialized Experts”, International Conference on Machine Learning. PMLR, pp. 5383–5395. Available at: https://proceedings.mlr.press/v202/chen23aq.html.
Vancouver
1. Chen W, Zhou Y, Du N, Huang Y, Laudon J, Chen Z, Cui C (2023) Lifelong Language Pretraining with Distribution-Specialized Experts. In: International Conference on Machine Learning. PMLR, pp 5383–5395

BibTeX

@InProceedings{pmlr-v202-chen23aq,
  title = 	 {Lifelong Language Pretraining with Distribution-Specialized Experts},
  author =       {Chen, Wuyang and Zhou, Yanqi and Du, Nan and Huang, Yanping and Laudon, James and Chen, Zhifeng and Cui, Claire},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {5383--5395},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/chen23aq/chen23aq.pdf},
  url = 	 {https://proceedings.mlr.press/v202/chen23aq.html},
  abstract = 	 {Pretraining on a large-scale corpus has become a standard method to build general language models (LMs). Adapting a model to new data distributions targeting different downstream tasks poses significant challenges. Naive fine-tuning may incur catastrophic forgetting when the over-parameterized LMs overfit the new data but fail to preserve the pretrained features. Lifelong learning (LLL) aims to enable information systems to learn from a continuous data stream across time. However, most prior work modifies the training recipe assuming a static fixed network architecture. We find that additional model capacity and proper regularization are key elements to achieving strong LLL performance. Thus, we propose Lifelong-MoE, an extensible MoE (Mixture-of-Experts) architecture that dynamically adds model capacity via adding experts with regularized pretaining. Our results show that by only introducing a limited number of extra experts while keeping the computation cost constant, our model can steadily adapt to data distribution shifts while preserving the previous knowledge. Compared to existing lifelong learning approaches, Lifelong-MoE achieves better few-shot performance on NLP tasks. More impressively, Lifelong-MoE surpasses multi-task learning on 19 downstream NLU tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/