Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation

Tu VuAditya BaruaBrian LesterDaniel CerMohit IyyerNoah Constant

article2022EMNLP79 citations

Proposes factorized prompt tuning and unlabeled data mixing to prevent multilingual language models from catastrophically forgetting non-English text generation capabilities when adapted only on English data.

Listen

Deploying language models across multiple languages is critical for modern applications, yet training data is often scarce for non-English languages. A key technical challenge is zero-shot cross-lingual generation—training a model on English generative tasks (such as summarization) and requiring it to perform the same task in unseen languages without relying on translation pipelines or target-language labeled examples. In this setting, models frequently suffer from catastrophic forgetting, inadvertently losing the capacity to generate coherent non-English text and instead defaulting to English.

The article evaluates parameter-efficient adaptation, particularly prompt tuning (which freezes the pre-trained model and tunes only a small set of input virtual tokens), against standard full-model tuning to determine how best to mitigate catastrophic forgetting across languages and model scales.

The researchers established a benchmark called WIKILINGUA-0 spanning 18 languages and evaluated models using a computationally efficient SentencePiece-based summarization metric, SP-ROUGE. They tested multilingual T5 models ranging from 300 million to 13 billion parameters across full-model fine-tuning and prompt tuning. To improve cross-lingual retention, they also evaluated mixing unlabeled multilingual data during training and developed a novel factorized prompts approach, which separates prompts into recombinable language and task components.

The experiments revealed three primary findings. First, standard full-model tuning rapidly overfits to English, generating substantial amounts of unwanted English text or code-switching when prompted in other languages. Second, scaling up overall model size while restricting tunable parameter capacity significantly reduces catastrophic forgetting; at the largest 13-billion-parameter scale, prompt tuning outperformed full-model tuning by 7.3 SP-ROUGE points on linguistically distant languages like Thai (37.4 versus 30.1). Third, adding a small fraction (1%) of unlabeled multilingual data during training consistently prevented forgetting across methods, whereas the novel factorized prompts technique yielded substantial gains primarily on smaller models facing severe language shifts.

These findings demonstrate that parameter-efficient prompt tuning is far more robust to cross-lingual domain shifts than full-model fine-tuning, while offering massive operational advantages by allowing a single frozen model to serve diverse language tasks at lower deployment and storage costs. Even so, fully supervised baselines trained directly on target-language data still surpass zero-shot prompt-tuned models by roughly 6 to 13 SP-ROUGE points, indicating that a performance gap remains.

Organizations aiming to build cross-lingual generative capabilities should prioritize using larger base models paired with parameter-efficient prompt tuning rather than fully fine-tuning model weights. Teams should also incorporate small amounts of unlabeled target-language data into the training pipeline to stabilize target-language output. If the target language is known in advance, mixing in data specifically for that language is recommended over broader mixtures.

Confidence in these findings is high for summarization across diverse script types, though decision-makers should note that the evaluation centered on a single instructional how-to dataset and a specific model family. Further evaluation is needed on broader generative tasks (such as open-ended generation or dialogue) and alternative parameter-efficient architectures before generalizing these outcomes across all language technologies.

Vu et al (2022).pdf

No sufficiently relevant recommendations were found.

No sufficiently relevant recommendations were found.

Cover for Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation

Abstract

Recent advances in speculative decoding have led to diverse perspectives among language models...

Table of Contents

  • 1 Introduction
  • 2 Challenge of zero-shot cross-lingual generation
  • 2.1 Problem formulation
  • 2.2 Experimental setup
  • 2.2.1 Baselines
  • 2.2.2 Training and implementation details
  • 2.3 Results and Discussion
  • 3 Mitigating catastrophic forgetting
  • 3.1 Methods
  • 3.2 Results and Discussion
  • 4 Qualitative Analysis
  • 5 Related Work
  • 6 Conclusion
  • 7 Limitations
  • Acknowledgements
  • References
  • Appendices
  • A Evaluation on zero-shot cross-lingual benchmarks
  • B Measuring the correlation between SP-RG and human judgments
  • C Zero-shot evaluation results on WIKILINGUA-0
  • D Language-Specific Prompt Clustering Analysis
  • E Mitigating catastrophic forgetting
  • F Intermediate tuning

Knowls

  1. Knowl 1 — WIKILINGUA-0 isolates zero-shot cross-lingual summarization

    experimental setup

    WIKILINGUA-0 adapts the multilingual WIKILINGUA summarization data to test generation in a language with no labeled training examples. It covers 18 languages. During ordinary training, the model receives labeled English article-summary pairs only; at evaluation, it receives articles in a non-English language and must generate summaries in that same language. Labeled data for other languages is excluded except in supervised ablations. The setting uses neither parallel data nor machine translation to provide target-language supervision.

  2. Knowl 2 — SP-ROUGE evaluates summaries using language-independent subword units

    definition

    SP-ROUGE applies ROUGE to SentencePiece-tokenized text, avoiding language-specific word segmentation assumptions that make standard ROUGE unsuitable for languages without spaces between words. The experiments report the summary-level ROUGE-L variant, SP-ROUGE-LSUM. On human-annotated multilingual summary evaluations, SP-ROUGE had mean Pearson correlations of 0.680.68 with human FOCUS judgments and 0.650.65 with COVERAGE judgments across eight languages. These averages were similar to BLEURT’s 0.680.68 and 0.700.70, respectively, while SP-ROUGE was reported to be substantially less computationally expensive.

  3. Knowl 3 — Adaptation compares full-model updates with a frozen-model soft prompt

    experimental setup

    The experiments use pretrained mT5 models with 300 million, 580 million, 1.2 billion, 3.7 billion, or 13 billion parameters. In MODELTUNING, all pretrained model weights are updated on the English summarization task. In PROMPTTUNING, pretrained weights are frozen and only a sequence of trainable virtual-token embeddings prepended to the input is updated; the standard prompt length is 100 tokens, initialized by sampling vocabulary embeddings. Prompt tuning uses mT5 checkpoints further trained for 100,000 steps on multilingual mC4 with a prefix-language-model objective. Both adaptation methods train for 100,000 steps; inputs and targets are clipped to 1,024 and 512 SentencePiece tokens, respectively. Checkpoints are evaluated on 250 validation examples per language, with the best target-language validation checkpoint reported. Generation uses beam search with beam size 4 and length penalty 0.60.6.

  4. Knowl 4 — Prompt tuning often transfers better than full-model tuning under large language shifts

    empirical result

    On WIKILINGUA-0, both MODELTUNING and PROMPTTUNING lose substantial summarization quality when trained on English and evaluated on non-English articles, and both can generate unwanted English. At the 13-billion-parameter scale, PROMPTTUNING nevertheless exceeds MODELTUNING on each of the four highlighted target languages: SP-ROUGE is 37.4 versus 35.5 for French, 38.0 versus 34.0 for Vietnamese, 29.2 versus 27.2 for Russian, and 37.4 versus 30.1 for Thai. The reversal is especially large for Thai, despite PROMPTTUNING scoring lower on English summarization at this scale (43.4 versus 46.7). The advantage is not universal across all model sizes or tasks, but is most evident under larger language shifts.

  5. Knowl 5 — Model scale mitigates forgetting, while excessive prompt capacity can hurt transfer

    empirical result

    Larger mT5 models are less prone to losing the ability to generate the target language after English-only summarization training. For Russian, target-language identification confidence rises from 0.0% with SMALL MODELTUNING and 10.1% with SMALL PROMPTTUNING to 57.5% and 84.4%, respectively, at XXL scale. Prompt length has a non-monotonic effect: additional tokens can improve task learning initially, but greater capacity can later increase forgetting. For XXL PROMPTTUNING on Thai, increasing prompt length from 1 to 10 tokens improves SP-ROUGE by 4.8 points while reducing Thai language-identification confidence by 8.0 points; increasing the prompt further to 100 tokens reduces both summarization quality and target-language confidence. Training trajectories also differ by method: prompt tuning increasingly introduces English as English-only training continues, whereas full-model tuning already produces more than 60% ASCII characters in Russian and Thai outputs by its first saved checkpoint.

  6. Knowl 6 — Mixing unlabeled multilingual text counters forgetting, especially for full-model tuning

    model/method

    The mixing method adds a span-corruption objective on unlabeled mC4 text to the English supervised summarization training mixture. Its mixing rate is κ=1\kappa=1, meaning 1% of training examples are unsupervised and 99% are WIKILINGUA-0 examples. The unsupervised data can cover only the intended target language or all WIKILINGUA-0 languages. For MODELTUNING, mixing generally improves both target-language generation and SP-ROUGE: at BASE scale, for example, the Russian and Thai SP-ROUGE scores rise from 15.3 to 24.1 and from 17.9 to 25.2, respectively. For PROMPTTUNING, effects are less consistent: at BASE scale, mixing improves Russian and Thai SP-ROUGE (14.3 to 16.1; 17.3 to 20.9) but lowers French and Vietnamese SP-ROUGE (23.8 to 20.3; 24.4 to 23.1). Mixing all languages gives broadly similar results, with a marginal performance drop relative to mixing only the known target language.

  7. Knowl 7 — Factorized prompts separate language and task sub-prompts for recombination

    model/method

    Factorized prompts (FP) represent a prompt as a language-specific sub-prompt concatenated with a task-specific sub-prompt. The method jointly trains these components using unlabeled mC4 examples from all 18 WIKILINGUA-0 languages and seven unsupervised tasks per language: prefix language modeling, span corruption, IID denoising, causal language modeling with an empty encoder input, missing-prefix prediction, prediction of the first nn input tokens, and prediction of a missing nn-token prefix. Each randomly initialized language and task sub-prompt has 50 tokens, giving a 100-token concatenated prompt; this multilingual multitask stage runs for 200,000 steps. For English summarization, a new task sub-prompt is trained while the learned English language sub-prompt is frozen, initialized from the span-corruption task sub-prompt. At inference in a target language, the English language component is replaced by that language’s learned component while retaining the summarization component. The authors report that FP improves target-language accuracy across conditions and can improve SP-ROUGE where ordinary prompt tuning forgets most severely; gains are limited when language accuracy is already high, including at XXL scale. At BASE scale, Russian and Thai SP-ROUGE rise from 14.3 to 17.8 and from 17.3 to 21.1, respectively, with FP.

  8. Knowl 8 — Fully supervised target-language training leaves substantial headroom

    empirical result

    Even the best zero-shot methods remain below training with labeled target-language data. At XXL scale, the all-language supervised MODELTUNING baseline exceeds the best zero-shot result by 5.8 SP-ROUGE points for Vietnamese and 12.8 points for Thai. For Thai, the supervised result also substantially exceeds the approaches that use machine-translated training or evaluation data, indicating that those translation-based conditions do not close the observed gap.

  9. Knowl 9 — Human inspection finds different code-switching patterns in generated summaries

    empirical result

    Two native-speaker annotators inspected 50 XXL-model predictions per method for Hindi and Vietnamese. MODELTUNING more often alternated between English and the target language within a summary; PROMPTTUNING more often stayed consistently in the target language, although some prompt-tuned outputs were entirely English. The inspected code-switched summaries were generally understandable to bilingual readers and conveyed reasonable summaries. In Hindi, PROMPTTUNING had mean SP-ROUGE 17.9 with standard deviation 5.1 across runs, compared with MODELTUNING’s 23.1 and 0.7; low-scoring prompt-tuned runs often produced entirely English output. In Vietnamese, PROMPTTUNING scored 38.0 (standard deviation at most 0.5) versus MODELTUNING’s 34.0 (standard deviation at most 0.5); prompt-tuned summaries were often entirely Vietnamese, and the authors judged them comparable to or better than the references.

  10. Knowl 10 — The evaluation covers one specialized summarization domain and one PEFT method

    limitation

    The experiments evaluate only WIKILINGUA-0 summarization, whose inputs are how-to guides and whose reference summaries are step headings rather than conventional news summaries. The authors note minor data-quality issues, including HTML in some targets and apparent machine-translation artifacts in some non-English documents. The parameter-efficient comparison focuses on prompt tuning; whether the findings extend to other parameter-efficient adaptation methods or to other generation tasks and domains remains untested.

Coverage note — The supplementary language-prompt clustering analysis and intermediate-tuning experiments are omitted: they are diagnostic or secondary comparisons rather than core methods for the paper’s zero-shot summarization contribution.

References

  1. 1.Shengnan An, Yifei Li, Zeqi Lin, Qian Liu, Bei Chen, Qiang Fu, Weizhu Chen, Nanning Zheng, and Jian-Guang Lou. 2022. Input-tuning: Adapting unfamiliar inputs to frozen pretrained models. arXiv preprint arXiv:2203.03131.
  2. 2.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 4623–4637.
  3. 3.Zewen Chi, Li Dong, Furu Wei, Wenhui Wang, Xian-Ling Mao, and Heyan Huang. 2020. Cross-lingual natural language generation via pre-training. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2020), 34(05):7570–7577.
  4. 4.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics (TACL 2020), 8:454–470.
  5. 5.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 8440–8451.
  6. 6.Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS 2019), volume 32.
  7. 7.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP 2018), pages 2475–2485.
  8. 8.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2019), pages 4171–4186.
  9. 9.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021. The GEM benchmark: Natural language generation, its evaluation and metrics. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021), pages 96–120.
  10. 10.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzman, and Angela Fan. 2021. The flores-101 evaluation benchmark for low-resource and multilingual machine translation. arXiv preprint arXiv:2106.03193.
  11. 11.David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003. English gigaword. Linguistic Data Consortium, Philadelphia, 4(1):34.
  12. 12.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2022. Towards a unified view of parameter-efficient transfer learning. Proceedings of the 10th International Conference on Learning Representations (ICLR 2022).
  13. 13.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (PMLR 2019), volume 97, pages 2790–2799.
  14. 14.Jeremy Howard and Sebastian Ruder. 2018. Universal language model fine-tuning for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL 2018), pages 328–339.
  15. 15.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. Proceedings of the 10th International Conference on Learning Representations (ICLR 2022).
  16. 16.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In Proceedings of the 37th International Conference on Machine Learning (ICML 2020), volume 119 of Proceedings of Machine Learning Research, pages 4411–4421.
  17. 17.Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021. Compacter: Efficient low-rank hypercomplex adapter layers. In Proceedings of the 35th International Conference on Neural Information Processing Systems (NeurIPS 2021), volume 34, pages 1022–1035.
  18. 18.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2017. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences (PNAS 2017), 114(13):3521–3526.
  19. 19.Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021. Evaluating the efficacy of summarization evaluation across languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 (Findings of ACL-IJCNLP 2021), pages 801–812.
  20. 20.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics (TACL 2022), 10:50–72.
  21. 21.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP 2018 System Demonstrations), pages 66–71.
  22. 22.Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020 (Findings of EMNLP 2020)), pages 4034–4048.
  23. 23.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 3045–3059.
  24. 24.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 7315–7330.
  25. 25.Junyi Jessy Li, Marine Carpuat, and Ani Nenkova. 2014. Assessing the discourse factors that influence the quality of machine translation. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (ACL 2014), pages 283–288.
  26. 26.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL 2021), pages 4582–4597.
  27. 27.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Workshop of Text Summarization Branches Out (WAS 2004), pages 74–81.
  28. 28.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. arXiv preprint arXiv:2205.05638.
  29. 29.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  30. 30.Kaushal Kumar Maurya, Maunendra Sankar Desarkar, Yoshinobu Kano, and Kumari Deepshikha. 2021. ZmBART: An unsupervised cross-lingual transfer framework for language generation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 (Findings of ACL-IJCNLP 2021), pages 2804–2818.
  31. 31.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2022. MetaICL: Learning to Learn In Context. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2022).
  32. 32.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), pages 1946–1958.
  33. 33.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL 2002), pages 311–318.
  34. 34.Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2020. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 7654–7673.
  35. 35.Jason Phang, Iacer Calixto, Phu Mon Htut, Yada Pruksachatkun, Haokun Liu, Clara Vania, Katharina Kann, and Samuel R. Bowman. 2020. English intermediate-task training improves zero-shot cross-lingual transfer too. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (AACL 2020), pages 557–575.
  36. 36.Jason Phang, Thibault Févry, and Samuel R Bowman. 2019. Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks. arXiv preprint arXiv:1811.01088.
  37. 37.Chengwei Qin and Shafiq Joty. 2022. Lfpt5: A unified framework for lifelong few-shot language learning based on prompt tuning of t5. Proceedings of the 10th International Conference on Learning Representations (ICLR 2022).
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR 2020), 21(140):1–67.
  39. 39.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017. Learning multiple visual domains with residual adapters. In Proceedings of the 31th Conference on Neural Information Processing Systems (NeurIPS 2017), volume 30.
  40. 40.Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasminn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo. 2022. Scaling up models and data with t5x and seqio. arXiv preprint arXiv:2203.17189.
  41. 41.Anthony Robins. 1995. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science, 7(2):123–146.
  42. 42.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 10215–10245.
  43. 43.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush. 2022. Multitask Prompted Training Enables Zero-Shot Task Generalization. In Proceedings of the 10th International Conference on Learning Representations (ICLR 2022).
  44. 44.Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), pages 1073–1083.
  45. 45.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020), pages 7881–7892.
  46. 46.Siamak Shakeri, Noah Constant, Mihir Kale, and Linting Xue. 2021. Towards zero-shot multilingual synthetic question and answer generation for cross-lingual reading comprehension. In Proceedings of the 14th International Conference on Natural Language Generation (INLG 2021), pages 35–45.
  47. 47.Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer. 2022. SPoT: Better frozen model adaptation through soft prompt transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022), pages 5039–5059.
  48. 48.Tu Vu, Minh-Thang Luong, Quoc Le, Grady Simon, and Mohit Iyyer. 2021. STraTA: Self-training with task augmentation for better few-shot learning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 5715–5731.
  49. 49.Tu Vu, Tong Wang, Tsendsuren Munkhdalai, Alessandro Sordoni, Adam Trischler, Andrew Mattarella-Micke, Subhransu Maji, and Mohit Iyyer. 2020. Exploring and predicting transferability across NLP tasks. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), pages 7882–7926.
  50. 50.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned Language Models Are Zero-Shot Learners. In Proceedings of the 10th International Conference on Learning Representations (ICLR 2022).
  51. 51.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2021), pages 483–498.
  52. 52.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019), pages 3687–3692.
  53. 53.Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. arXiv preprint arXiv:2106.10199.
  54. 54.Mengjie Zhao and Hinrich Schütze. 2021. Discrete and soft prompting for multilingual models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), pages 8547–8555.

Citation

MLA
Vu, T., et al. “Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 9279–300, https://doi.org/10.18653/v1/2022.emnlp-main.630.
APA
Vu, T., Barua, A., Lester, B., Cer, D., Iyyer, M., & Constant, N. (2022). Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9279–9300. https://doi.org/10.18653/v1/2022.emnlp-main.630
Chicago
Vu, T., A. Barua, B. Lester, D. Cer, M. Iyyer, and N. Constant. 2022. “Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 9279–9300. https://doi.org/10.18653/v1/2022.emnlp-main.630.
Harvard
Vu, T. et al. (2022) “Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9279–9300. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.630.
Vancouver
1. Vu T, Barua A, Lester B, Cer D, Iyyer M, Constant N (2022) Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9279–9300

BibTeX

@inproceedings{vu-etal-2022-overcoming,
    title = "Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation",
    author = "Vu, Tu  and
      Barua, Aditya  and
      Lester, Brian  and
      Cer, Daniel  and
      Iyyer, Mohit  and
      Constant, Noah",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.630/",
    doi = "10.18653/v1/2022.emnlp-main.630",
    pages = "9279--9300"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/