Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation

Dingcheng LiZheng ChenEunah ChoJie HaoXiaohu LiuFan XingChenlei GuoYang Liu

article2022NAACL67 citations

Proposes a framework combining adaptive parameter regularization with embedding-space domain drift estimation to prevent catastrophic forgetting in sequential sequence-to-sequence language generation without storing past task data.

Listen

Modern language generation systems, such as automated dialog assistants and paraphrasing tools, frequently struggle with catastrophic forgetting. When these sequence-to-sequence models are updated with data from a new domain or task, they tend to forget what they previously learned. Existing continuous learning solutions typically demand substantial extra memory storage to save past data samples for replay or fail to correct representation shifts between differing domains. As organizations deploy conversational artificial intelligence across expanding sets of features and operational environments, retaining previous knowledge without incurring massive storage overhead has become a critical challenge.

The article develops and evaluates a framework called RMR_DSE to mitigate catastrophic forgetting during sequential domain adaptation in text generation models. The approach demonstrates how models can acquire new domain skills while preserving prior capabilities without storing raw data from earlier training phases.

The proposed framework integrates two complementary techniques: a Regularized Memory Recall (RMR) mechanism and a Domain Shift Estimation (DSE) algorithm. RMR enhances training loss by selectively penalizing changes to critical parameters from prior models using gradient-based importance weights and vocabulary-ratio adjustments. DSE estimates semantic drift in internal representation spaces between consecutive models via unsupervised clustering and shifts, which is then subtracted during inference on past domains. The authors evaluated the system across two distinct generation benchmarks: sequential paraphrase generation (using over 270,000 sentence pairs across Quora, Twitter, and Wikipedia datasets) and task-oriented dialog response generation (using the multi-domain MultiWOZ-2.0 dataset across six domains and seven intent categories), testing standard architectures including BART, CVAE, and SCLSTM.

Experimental results show that the framework consistently outperforms standard sequential fine-tuning and state-of-the-art lifelong learning baselines on both current and historical tasks. In paraphrase generation, RMR_DSE achieved higher accuracy across all sequential test orderings, improving generation quality metrics by 3–9% over basic fine-tuning and significantly reducing performance drop when tested on earlier datasets. In dialog response generation without stored exemplars, the system dramatically reduced average slot error rates on previous domains from roughly 64–67% down to 48–52%, while maintaining higher text generation accuracy. When evaluated on initial task retention after learning all six dialog domains, the method cut error rates nearly in half compared to fine-tuning, reducing slot error from over 100% to 57–68%.

These findings indicate that artificial intelligence systems can expand into new domains without requiring burdensome archival databases of previous user interactions. This capability minimizes data storage costs, mitigates user privacy risks associated with long-term data retention, and improves the reliability of deployed natural language systems. Organizations can maintain stable conversational quality without having to retrain models from scratch across entire combined datasets every time a new business domain is added.

Engineering and product teams developing conversational systems should consider adopting adaptive regularization and domain shift estimation pipelines for sequential deployments. Teams can implement this approach as a lightweight drop-in optimizer and loss function adjustment without modifying underlying model architectures. For next steps, teams should pilot the framework across additional text generation tasks, such as automated summarization and machine translation, while experimenting with token-level drift estimation and deep contrastive learning objectives to further refine semantic alignment across domains.

While the empirical findings demonstrate strong confidence across two established text generation benchmarks, several operational considerations remain. The framework requires hyperparameter tuning—such as penalty weights and cluster numbers—that varies based on the structural complexity of the underlying neural network. In addition, semantic drift estimation showed modest improvements when embedding values between domains exhibited similar distributions. Stakeholders should validate cluster settings and parameter sensitivity on domain-specific pilot data before full-scale production rollouts.

Li et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation

Abstract

Seq2seq language generation models that are trained offline with multiple domains in a sequential fashion often suffer from catastrophic forgetting. Lifelong learning has been proposed to handle this problem. However, existing work such as experience replay or elastic weighted consolidation requires incremental memory space. In this work, we propose an innovative framework, RMR_DSE that leverages a recall optimization mechanism to selectively memorize important parameters of previous tasks via regularization, and uses a domain drift estimation algorithm to compensate for the drift between different domains in the embedding space. These designs enable the model to be trained on the current task while keeping the memory of previous tasks, and avoid much additional data storage. Furthermore, RMR_DSE can be combined with existing lifelong learning approaches. Our experiments on two seq2seq language generation tasks, paraphrase and dialog response generation, show that RMR_DSE outperforms state-of-the-art models by a considerable margin and greatly reduces forgetting.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Life Long Learning (LLL)
  • 2.2 LLL in Seq2seq Language Generation
  • 3 Proposed Method
  • 3.1 Adaptive Regularization of Memory Recall
  • 3.2 Deploying Domain Shift Estimation to Make Up Semantic Drift
  • 4 Experiments on Paraphrase Generation
  • 4.1 Experimental Setups
  • 4.2 Results
  • Evaluating on the Current Task
  • Evaluating on Previous Tasks
  • 4.3 Case Studies
  • 5 Experiments on Dialog Response Generation
  • 5.1 Task Definition
  • 5.2 Experimental Settings
  • 5.3 Baseline Methods
  • 5.4 Experimental Results
  • 6 Discussions
  • 7 Conclusion
  • References
  • A Appendix
  • A.1 Domain Order Permutation
  • A.2 Metrics Details
  • A.3 Bleu4 scores for MultiWoZ-2.0 dataset
  • A.4 Packages Used for Implementation
  • A.5 Parameter Update Analysis of RMR_DSE on Different Network Structures

Knowls

  1. Knowl 1 — Regularized Memory Recall for continual seq2seq training

    model/method

    Regularized Memory Recall (RMR) trains a seq2seq model on task tt while penalizing changes to parameters judged important, using the previous task’s parameter values as the reference:

    Lt(θ)=λ(τ)Lt(θ)+(1−λ(τ))γF∑i,jΠij(θij−θij∗)2.\mathcal{L}_t(\theta)=\lambda(\tau)L_t(\theta)+(1-\lambda(\tau))\gamma F\sum_{i,j}\Pi_{ij}(\theta_{ij}-\theta^*_{ij})^2.

    Here, θ\theta is the current model’s parameter vector, θ∗\theta^* is the parameter vector saved from the preceding task, and θij\theta_{ij} denotes a parameter associated with a connection between neurons in consecutive layers. LtL_t is the current-task loss, implemented as label-smoothed cross-entropy in the reported generation experiments. The weighting λ(τ)\lambda(\tau) is a sigmoid annealing function of training update step τ\tau, controlled by rate and midpoint hyperparameters kk and τ0\tau_0. The scalar FF is treated as a tunable hyperparameter rather than estimated from unavailable previous-task data. The scale γ\gamma is set using vocabulary sizes as γ=γbaseV1:t−1/Vt\gamma=\gamma_{\mathrm{base}}\sqrt{V_{1:t-1}/V_t}, where V1:t−1V_{1:t-1} and VtV_t are the vocabulary sizes for previous and current task data, respectively, and γbase\gamma_{\mathrm{base}} is a base tuning parameter.

    RMR weights each parameter penalty using current-task examples: Πij=1N∑n=1N∥gij(xn)∥22\Pi_{ij}=\frac{1}{N}\sum_{n=1}^{N}\|g_{ij}(x_n)\|_2^2, where NN is the number of current-task examples, xnx_n is example nn, and gij(xn)=∂G(xn;θ)/∂θijg_{ij}(x_n)=\partial G(x_n;\theta)/\partial\theta_{ij} is the gradient of the learned generative function GG with respect to parameter θij\theta_{ij}. Parameters with larger gradient-based importance receive stronger protection, while less important parameters can change to fit the new task. RMR therefore uses the previous model parameters but does not require old examples to compute these importance weights.

  2. Knowl 2 — Domain Shift Estimation compensates for embedding drift

    model/method

    Domain Shift Estimation (DSE) estimates how encoder representations move when a seq2seq model is updated from task t−1t-1 to task tt. For each example ii from the current task’s training data, it encodes the example with both models and computes the displacement δit−1→t=zit−zit−1\delta_i^{t-1\to t}=z_i^t-z_i^{t-1}, where zit−1z_i^{t-1} and zitz_i^t are the respective encoder outputs. DSE clusters the embeddings produced by the previous model, then summarizes the displacements within each cluster as a mean-shift-weighted average. For cluster kk, the estimate is Δdse,kt−1→t=∑i∈Ckmiδit−1→t/∑i∈Ckmi\Delta_{dse,k}^{t-1\to t}=\sum_{i\in C_k}m_i\delta_i^{t-1\to t}/\sum_{i\in C_k}m_i, where CkC_k is the set of examples assigned to that cluster and mim_i is the mean-shift weight computed for example ii.

    At inference on an earlier task, the current model encodes a test input and selects the stored shift associated with the most similar cluster center. It subtracts that shift from the current encoder representation before decoding. For a task several updates earlier, the method subtracts the sequence of estimated shifts for the intervening task transitions. DSE obtains the shifts from current-task examples encoded by the two available model checkpoints; it does not require storing previous-task training examples.

  3. Knowl 3 — DSE clustering and mean-shift procedure

    algorithm

    To estimate shifts for one transition from model t−1t-1 to model tt, DSE takes paired embeddings of current-task training examples from both models. It runs KK-means on embeddings from model t−1t-1; the reported best setting uses K=3K=3, with about five clustering runs to address sensitivity to center initialization. For each center, it uses FAISS to retrieve nearby examples, using 3,000 samples per center in the experiments. It forms each retrieved example’s displacement by subtracting its old-model embedding from its new-model embedding. It then iteratively applies Gaussian-kernel mean shift to the retrieved old-model embeddings and uses the resulting mean-shift values as weights for aggregating the paired displacements into a cluster shift vector. The paper describes the iteration as continuing until the mean-shift movement meets distance-based stopping conditions, but does not report numerical values for those thresholds.

    The output is one shift vector per cluster, together with the centers used to associate a later test embedding with a shift. At evaluation, DSE selects the relevant cluster shift by embedding similarity and subtracts it before generation. The paper does not state a computational-complexity bound.

  4. Knowl 4 — Paraphrase experiment uses sequential QTW training

    experimental setup

    The paraphrase experiment trains a BART seq2seq model in the order Quora, Twitter, then Wiki_data (QTW). The three corpora contain sentence-paraphrase pairs, with these reported split sizes:

    SplitQuoraTwitterWiki_dataTotal
    Train111,94785,97078,392276,309
    Validation8,0001,0008,15417,154
    Test37,3163,0009,32449,640

    Comparisons include independent models trained on one corpus, sequential fine-tuning, EWC, MR, RMR, and a Full model trained jointly on all three corpora. BLEU-4, ROUGE-L, and METEOR evaluate generation quality. To evaluate retention, models updated on later corpora are tested on earlier corpora. DSE is evaluated only on these previous-task tests, not on the current-task results.

  5. Knowl 5 — RMR improves paraphrase generation on current tasks

    empirical result

    In the QTW paraphrase experiment, the following table reports BLEU-4, ROUGE-L, and METEOR for models evaluated on each current corpus. DSE is not applied in this current-task comparison. MR and RMR improve on sequential fine-tuning and EWC across the later-task evaluations; the jointly trained Full model provides a reference that has seen all three corpora.

    ModelQuora BLEU-4Quora ROUGE-LQuora METEORTwitter BLEU-4Twitter ROUGE-LTwitter METEORWiki_data BLEU-4Wiki_data ROUGE-LWiki_data METEOR
    Quora-trained36.9858.1960.762.126.135.494.5111.2112.13
    Twitter-trained3.1811.469.0136.4747.4945.574.609.767.50
    Wiki_data-trained22.3843.4446.239.3217.9321.0348.0369.7067.43
    Finetune36.9858.1960.7635.7946.4645.9346.8768.9867.02
    EWC36.8958.1659.9835.5247.1446.1648.1569.5368.59
    MR37.9859.1961.1136.9849.3948.0253.9374.4974.53
    RMR38.4659.4861.1438.9451.2347.1254.1274.9875.13
    Full37.9959.3361.0439.5351.3347.6455.9376.5676.41
  6. Knowl 6 — DSE further reduces forgetting in paraphrase generation

    empirical result

    The QTW retention evaluation tests models trained on later corpora against earlier-corpus test sets. Each entry is BLEU-4, ROUGE-L, and METEOR in that order. RMR_DSE scores exceed fine-tuning on all three prior-task evaluations shown, and improve over RMR in each listed metric.

    Evaluation after later-task trainingMethodBLEU-4ROUGE-LMETEOR
    Quora test after Twitter trainingQuora-trained reference36.9858.1960.76
    Finetune20.7730.8041.75
    EWC21.6331.5342.03
    DSE21.5831.9542.98
    MR25.4735.8845.27
    RMR26.9736.3947.26
    RMR_DSE27.7436.9848.38
    Quora test after Wiki_data trainingQuora-trained reference36.9858.1960.76
    Finetune22.8342.1642.03
    EWC24.6344.3543.02
    DSE23.7943.4943.35
    MR28.4447.3755.43
    RMR29.7249.1557.15
    RMR_DSE30.7149.4357.99
    Twitter test after Wiki_data trainingTwitter-based reference36.4747.4945.57
    Finetune19.9937.2041.57
    EWC18.8438.6543.33
    DSE20.7840.0042.75
    MR21.9238.6944.36
    RMR24.1542.1145.59
    RMR_DSE26.7343.8546.23
  7. Knowl 7 — Dialog generation evaluation uses MultiWOZ continual learning

    experimental setup

    The dialog-response task maps a dialog act—an intent plus slot-value pairs—to a natural-language response. Experiments use MultiWOZ-2.0 with six domains (Attraction, Hotel, Restaurant, Booking, Taxi, and Train) and seven intents (Inform, Request, Select, Recommend, Book, Offer-Booked, and No-Offer), retaining the dataset’s original train, validation, and test splits. The generation backbones are a conditional variational autoencoder (CVAE) and a semantic-conditioned LSTM (SCLSTM). Comparisons include sequential fine-tuning, joint training on all domains (Full), ARPER, EWC, MR_DSE, and RMR_DSE, with evaluations both without exemplars and with 250 exemplars.

    The metrics are slot error rate (SER; lower is better) and BLEU-4 (higher is better). Continual-learning averages are Ωall=1T∑i=1TΩall,i\Omega_{all}=\frac{1}{T}\sum_{i=1}^{T}\Omega_{all,i} and Ωfirst=1T∑i=1TΩfirst,i\Omega_{first}=\frac{1}{T}\sum_{i=1}^{T}\Omega_{first,i}, where TT is the number of tasks, Ωall,i\Omega_{all,i} is average test performance across tasks learned through task ii, and Ωfirst,i\Omega_{first,i} is performance on the first task after learning task ii.

  8. Knowl 8 — RMR_DSE retains dialog performance without stored exemplars

    empirical result

    For continual learning across the six MultiWOZ-2.0 domains with zero exemplars, the table reports mean SER and BLEU-4 for all learned tasks (Ωall\Omega_{all}) and for the first task (Ωfirst\Omega_{first}). Without exemplars, RMR_DSE improves on fine-tuning and ARPER for both model backbones in both averages; ARPER without exemplars is the EWC-style comparison described by the paper.

    MethodΩall\Omega_{all} SERΩall\Omega_{all} BLEU-4Ωfirst\Omega_{first} SERΩfirst\Omega_{first} BLEU-4
    Finetune64.4636.1107.2725.3
    ARPER on CVAE63.5436.00102.8719.22
    ARPER on SCLSTM66.8735.64100.5621.09
    RMR_DSE on CVAE51.9239.4968.5625.67
    RMR_DSE on SCLSTM48.7939.8657.1830.32
  9. Knowl 9 — RMR_DSE remains competitive with 250 dialog exemplars

    empirical result

    With 250 exemplars in the six-domain MultiWOZ-2.0 continual-learning experiment, RMR_DSE is compared with ARPER and MR_DSE, as well as joint training (Full). The table reports SER and BLEU-4 for all-task and first-task averages. RMR_DSE has lower SER than ARPER for both backbones in both averages, and improves on MR_DSE in nearly all reported metrics; the CVAE Ωall\Omega_{all} BLEU-4 score is tied at 59.8. Full training has the lowest Ωall\Omega_{all} SER, while RMR_DSE’s first-task results show stronger retention than Full on the reported first-task SER values.

    MethodΩall\Omega_{all} SERΩall\Omega_{all} BLEU-4Ωfirst\Omega_{first} SERΩfirst\Omega_{first} BLEU-4
    ARPER on CVAE5.2458.32.9762.1
    ARPER on SCLSTM5.9756.73.5961.3
    MR_DSE on CVAE4.6859.82.8162.7
    MR_DSE on SCLSTM4.9559.92.6063.2
    RMR_DSE on CVAE4.5259.82.0263.5
    RMR_DSE on SCLSTM4.3860.32.1263.6
    Full4.2659.93.6061.6
  10. Knowl 10 — Reported RMR hyperparameters vary by generation backbone

    experimental setup

    The reported settings for four RMR hyperparameters differ across the BART, CVAE, and SCLSTM implementations. The update frequency is two for all three models; the other listed values are model-specific.

    HyperparameterBARTCVAESCLSTM
    updatefreq222
    regλ0.90.10.01
    annealw0.10.010.05
    pretraincof500050050

    The paper describes updatefreq as how often to update Π\Pi, regλ as the proportion of Π\Pi, annealw as the weight on parameter differences, and pretraincof as the coefficient for the quadratic penalty. The values are implementation settings reported for these experiments, not universal defaults.

  11. Knowl 11 — The paper identifies limits of sentence-level drift correction

    limitation

    The authors observe that embeddings from older and newer models can have different density patterns even when their value ranges are similar; they suggest that this similarity may partly explain why DSE contributes less to performance improvement than other components. Their DSE corrects sentence-level encoder representations, whereas the decoder generates with token-level beam search. They therefore identify token-level as well as sentence-level shift estimation, and adding deep contrastive learning to the label-smoothed cross-entropy training objective, as possible directions for improving the framework. These are proposed extensions, not demonstrated improvements.

Coverage note — Qualitative paraphrase examples, embedding visualizations, and the additional QWT and TQW domain orders are omitted because the quantitative results and DSE description capture the main contributed claims without duplicating evidence.

References

  1. 1.Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. 2018. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision (ECCV), pages 139–154.
  2. 2.Saket Anand, Sushil Mittal, Oncel Tuzel, and Peter Meer. 2013. Semi-supervised kernel mean shift clustering. IEEE transactions on pattern analysis and machine intelligence, 36(6):1201–1215.
  3. 3.Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018. Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018.
  4. 4.Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet Kumar Dokania, Philip H. S. Torr, and Marc’Aurelio Ranzato. 2019. Continual learning with tiny episodic memories. CoRR, abs/1902.10486.
  5. 5.Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu. 2020. Recall and learn: Fine-tuning deep pretrained language models with less forgetting. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020.
  6. 6.Cyprien de Masson d’Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019. Episodic memory in lifelong language learning. Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada.
  7. 7.Mihail Eric, Rahul Goel, Shachi Paul, Adarsh Kumar, Abhishek Sethi, Peter Ku, Anuj Kumar Goyal, Sanchit Agarwal, Shuyang Gao, and Dilek Hakkani-Tur. 2019. Multiwoz 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines. Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020.
  8. 8.Robert M French. 1999. Catastrophic forgetting in connectionist networks. Trends in cognitive sciences, 3(4):128–135.
  9. 9.Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James R. Glass, and Fuchun Peng. 2021. Analyzing the forgetting problem in pretrain-finetuning of open-domain dialogue response models. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021, pages 1121–1133. Association for Computational Linguistics.
  10. 10.Ferenc Huszár. 2018. Note on the quadratic penalties in elastic weight consolidation. Proceedings of the National Academy of Sciences, page 201717042.
  11. 11.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526.
  12. 12.Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li. 2017. Learning from noisy labels with distillation. In ICCV.
  13. 13.Zhizhong Li and Derek Hoiem. 2018. Learning without forgetting. IEEE Trans. Pattern Anal. Mach. Intell., 40(12):2935–2947.
  14. 14.David Lopez-Paz and Marc’Aurelio Ranzato. 2017. Gradient episodic memory for continual learning. Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA.
  15. 15.Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A. Crook, Bing Liu, Zhou Yu, Eunjoon Cho, and Zhiguang Wang. 2020. Continual learning in task-oriented dialogue systems. CoRR, abs/2012.15504.
  16. 16.Michael McCloskey and Neal J Cohen. 1989. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pages 109–165. Elsevier.
  17. 17.Fei Mi, Liangwei Chen, Mengjie Zhao, Minlie Huang, and Boi Faltings. 2020. Continual learning for natural language generation in task-oriented dialog systems. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, EMNLP 2020, Online Event, 16-20 November 2020.
  18. 18.Lorenzo Pellegrini, Gabriele Graffieti, Vincenzo Lomonaco, and Davide Maltoni. 2019. Latent replay for real-time continual learning. IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2020, Las Vegas, NV, USA, October 24, 2020 - January 24, 2021.
  19. 19.Mark Bishop Ring et al. 1994. Continual learning in reinforcement environments. Ph.D. thesis, University of Texas at Austin Austin, Texas 78712.
  20. 20.David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P Lillicrap, and Greg Wayne. 2019. Experience replay for continual learning. Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada.
  21. 21.Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim. 2017. Continual learning with deep generative replay. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 2990–2999.
  22. 22.Wenbo Wang, Yang Gao, He-Yan Huang, and Yuxiang Zhou. 2019. Concept pointer network for abstractive summarization. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3067–3076.
  23. 23.Yigong Wang, Zhuoyi Wang, Yu Lin, Latifur Khan, and Dingcheng Li. 2021a. Cifdm: continual and interactive feature distillation for multi-label stream learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2121–2125.
  24. 24.Zhuoyi Wang, Yuqiao Chen, Chen Zhao, Yu Lin, and Latifur Khan. 2021b. Clear: Contrastive-prototype learning with drift estimation for resource constrained stream mining. In Proceedings of The Web Conference.
  25. 25.Zirui Wang, Sanket Vaibhav Mehta, Barnabás Póczos, and Jaime Carbonell. 2020. Efficient meta lifelong-learning with limited memory. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020.
  26. 26.Lu Yu, Bartlomiej Twardowski, Xialei Liu, Luis Herranz, Kai Wang, Yongmei Cheng, Shangling Jui, and Joost van de Weijer. 2020. Semantic drift compensation for class-incremental learning. In CVPR, pages 6982–6991.
  27. 27.Friedemann Zenke, Ben Poole, and Surya Ganguli. 2017. Continual learning through synaptic intelligence. In Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017, volume 70 of Proceedings of Machine Learning Research, pages 3987–3995. PMLR.

Citation

MLA
Li, D., et al. “Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5441–54, https://doi.org/10.18653/v1/2022.naacl-main.398.
APA
Li, D., Chen, Z., Cho, E., Hao, J., Liu, X., Xing, F., Guo, C., & (刘扬), Y. L. (2022). Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5441–5454. https://doi.org/10.18653/v1/2022.naacl-main.398
Chicago
Li, D., Z. Chen, E. Cho, et al. 2022. “Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5441–54. https://doi.org/10.18653/v1/2022.naacl-main.398.
Harvard
Li, D. et al. (2022) “Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 5441–5454. Available at: https://doi.org/10.18653/v1/2022.naacl-main.398.
Vancouver
1. Li D, Chen Z, Cho E, Hao J, Liu X, Xing F, Guo C, (刘扬) YL (2022) Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 5441–5454

BibTeX

@inproceedings{li-etal-2022-overcoming,
    title = "Overcoming Catastrophic Forgetting During Domain Adaptation of Seq2seq Language Generation",
    author = "Li, Dingcheng  and
      Chen, Zheng  and
      Cho, Eunah  and
      Hao, Jie  and
      Liu, Xiaohu  and
      Xing, Fan  and
      Guo, Chenlei  and
      Liu, Yang",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.398/",
    doi = "10.18653/v1/2022.naacl-main.398",
    pages = "5441--5454"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/