Overcoming Catastrophic Forgetting in Zero-Shot Cross-Lingual Generation
Tu VuAditya BaruaBrian LesterDaniel CerMohit IyyerNoah Constant
Proposes factorized prompt tuning and unlabeled data mixing to prevent multilingual language models from catastrophically forgetting non-English text generation capabilities when adapted only on English data.
Deploying language models across multiple languages is critical for modern applications, yet training data is often scarce for non-English languages. A key technical challenge is zero-shot cross-lingual generation—training a model on English generative tasks (such as summarization) and requiring it to perform the same task in unseen languages without relying on translation pipelines or target-language labeled examples. In this setting, models frequently suffer from catastrophic forgetting, inadvertently losing the capacity to generate coherent non-English text and instead defaulting to English.
The article evaluates parameter-efficient adaptation, particularly prompt tuning (which freezes the pre-trained model and tunes only a small set of input virtual tokens), against standard full-model tuning to determine how best to mitigate catastrophic forgetting across languages and model scales.
The researchers established a benchmark called WIKILINGUA-0 spanning 18 languages and evaluated models using a computationally efficient SentencePiece-based summarization metric, SP-ROUGE. They tested multilingual T5 models ranging from 300 million to 13 billion parameters across full-model fine-tuning and prompt tuning. To improve cross-lingual retention, they also evaluated mixing unlabeled multilingual data during training and developed a novel factorized prompts approach, which separates prompts into recombinable language and task components.
The experiments revealed three primary findings. First, standard full-model tuning rapidly overfits to English, generating substantial amounts of unwanted English text or code-switching when prompted in other languages. Second, scaling up overall model size while restricting tunable parameter capacity significantly reduces catastrophic forgetting; at the largest 13-billion-parameter scale, prompt tuning outperformed full-model tuning by 7.3 SP-ROUGE points on linguistically distant languages like Thai (37.4 versus 30.1). Third, adding a small fraction (1%) of unlabeled multilingual data during training consistently prevented forgetting across methods, whereas the novel factorized prompts technique yielded substantial gains primarily on smaller models facing severe language shifts.
These findings demonstrate that parameter-efficient prompt tuning is far more robust to cross-lingual domain shifts than full-model fine-tuning, while offering massive operational advantages by allowing a single frozen model to serve diverse language tasks at lower deployment and storage costs. Even so, fully supervised baselines trained directly on target-language data still surpass zero-shot prompt-tuned models by roughly 6 to 13 SP-ROUGE points, indicating that a performance gap remains.
Organizations aiming to build cross-lingual generative capabilities should prioritize using larger base models paired with parameter-efficient prompt tuning rather than fully fine-tuning model weights. Teams should also incorporate small amounts of unlabeled target-language data into the training pipeline to stabilize target-language output. If the target language is known in advance, mixing in data specifically for that language is recommended over broader mixtures.
Confidence in these findings is high for summarization across diverse script types, though decision-makers should note that the evaluation centered on a single instructional how-to dataset and a specific model family. Further evaluation is needed on broader generative tasks (such as open-ended generation or dialogue) and alternative parameter-efficient architectures before generalizing these outcomes across all language technologies.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
