MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators
Zhixing TanXiangwen ZhangShuo WangYang Liu
Proposes a multi-stage continuous prompting framework that splits the translation process in decoder-only language models into separate encoding, re-encoding, and decoding stages, substantially improving translation quality over standard prompt-tuning approaches while requiring minimal parameter updates.
Deploying dedicated machine translation models for multiple language pairs requires massive storage, high computational overhead, and complex maintenance. While large pre-trained language models can perform diverse tasks within a single architecture, guiding them to translate effectively remains challenging. Standard prompting strategies—which guide a model using short text instructions or continuous numerical vectors—struggle with unidirectional decoder architectures and the fundamental objective mismatch between pre-training text completion and strict bilingual translation.
The article evaluates Multi-Stage Prompting, a lightweight prompting framework designed to turn pre-trained generative language models into high-performing translators. The core objective is to demonstrate that decomposing the translation pipeline into distinct, prompted stages significantly enhances translation quality while keeping the underlying model fixed.
The approach introduces a three-stage sequence: an encoding stage that maps the source sentence into initial internal representations, an intermediate re-encoding stage that refines these representations into richer context, and a decoding stage that generates the final target translation. Each stage uses a dedicated set of continuous prompts optimized via back-propagation. To evaluate this approach, the authors trained prompts on a 560-million-parameter multilingual model across Romanian-English, English-German, and English-Chinese benchmarks, alongside a multilingual corpus covering five additional language pairs.
The key findings show substantial performance and efficiency improvements. First, Multi-Stage Prompting achieved an average translation score of 28.0 BLEU across the three benchmark tasks, outperforming standard embedding prompt tuning by 18.6 points and prefix-tuning by 4.1 points. Second, on the English-Chinese task, the 560-million-parameter model achieved 28.1 BLEU, surpassing much larger 10-to-13-billion-parameter models trained with fine-tuning or alternative prompts. Third, across five language pairs, the method exceeded a strong dedicated Transformer translation baseline by 3.4 to 3.9 BLEU points on average. Fourth, internal analysis confirmed that core translation knowledge already exists within the pre-trained language model, while the multi-stage prompts act as effective operational guides to transition smoothly between language spaces.
These findings have direct operational and cost implications. By freezing the underlying base model and only learning compact prompt vectors, supporting a new translation direction requires only 19 megabytes of storage—compared to over 60 to 450 megabytes for a dedicated translation model—and reduces training time from 72 hours down to 21 hours on standard hardware. Crucially, the base model retains its versatility for other tasks, enabling unified system deployments without performance degradation.
Organizations seeking to streamline translation infrastructure should consider multi-stage continuous prompting over training standalone translation models, especially when supporting multiple language pairs under strict storage and computational budgets. Before wide production rollout, teams should conduct pilot evaluations on high-resource language pairs where dedicated translation models may still retain an edge, such as English-German. Future work should focus on exploring dynamic, sentence-specific prompts and testing the approach across larger base models.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Its soft-prompt method establishes the frozen-backbone, trainable-continuous-vector approach that MSP adapts into a staged translation pipeline.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). Read this early account of gradient-trained continuous prompts to understand the prompt-tuning foundation MSP extends beyond language-understanding tasks.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). mT5’s multilingual text-to-text model provides the relevant pretrained-model context for MSP’s multilingual translation experiments.
No sufficiently relevant recommendations were found.
