MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators

Zhixing TanXiangwen ZhangShuo WangYang Liu

article2022ACL60 citations

Proposes a multi-stage continuous prompting framework that splits the translation process in decoder-only language models into separate encoding, re-encoding, and decoding stages, substantially improving translation quality over standard prompt-tuning approaches while requiring minimal parameter updates.

Listen

Deploying dedicated machine translation models for multiple language pairs requires massive storage, high computational overhead, and complex maintenance. While large pre-trained language models can perform diverse tasks within a single architecture, guiding them to translate effectively remains challenging. Standard prompting strategies—which guide a model using short text instructions or continuous numerical vectors—struggle with unidirectional decoder architectures and the fundamental objective mismatch between pre-training text completion and strict bilingual translation.

The article evaluates Multi-Stage Prompting, a lightweight prompting framework designed to turn pre-trained generative language models into high-performing translators. The core objective is to demonstrate that decomposing the translation pipeline into distinct, prompted stages significantly enhances translation quality while keeping the underlying model fixed.

The approach introduces a three-stage sequence: an encoding stage that maps the source sentence into initial internal representations, an intermediate re-encoding stage that refines these representations into richer context, and a decoding stage that generates the final target translation. Each stage uses a dedicated set of continuous prompts optimized via back-propagation. To evaluate this approach, the authors trained prompts on a 560-million-parameter multilingual model across Romanian-English, English-German, and English-Chinese benchmarks, alongside a multilingual corpus covering five additional language pairs.

The key findings show substantial performance and efficiency improvements. First, Multi-Stage Prompting achieved an average translation score of 28.0 BLEU across the three benchmark tasks, outperforming standard embedding prompt tuning by 18.6 points and prefix-tuning by 4.1 points. Second, on the English-Chinese task, the 560-million-parameter model achieved 28.1 BLEU, surpassing much larger 10-to-13-billion-parameter models trained with fine-tuning or alternative prompts. Third, across five language pairs, the method exceeded a strong dedicated Transformer translation baseline by 3.4 to 3.9 BLEU points on average. Fourth, internal analysis confirmed that core translation knowledge already exists within the pre-trained language model, while the multi-stage prompts act as effective operational guides to transition smoothly between language spaces.

These findings have direct operational and cost implications. By freezing the underlying base model and only learning compact prompt vectors, supporting a new translation direction requires only 19 megabytes of storage—compared to over 60 to 450 megabytes for a dedicated translation model—and reduces training time from 72 hours down to 21 hours on standard hardware. Crucially, the base model retains its versatility for other tasks, enabling unified system deployments without performance degradation.

Organizations seeking to streamline translation infrastructure should consider multi-stage continuous prompting over training standalone translation models, especially when supporting multiple language pairs under strict storage and computational budgets. Before wide production rollout, teams should conduct pilot evaluations on high-resource language pairs where dedicated translation models may still retain an edge, such as English-German. Future work should focus on exploring dynamic, sentence-specific prompts and testing the approach across larger base models.

No sufficiently relevant recommendations were found.

Cover for MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators

Abstract

Prompting has recently been shown as a promising approach for applying pre-trained language models to perform downstream tasks. We present Multi-Stage Prompting, a simple and automatic approach for leveraging pre-trained language models to translation tasks. To better mitigate the discrepancy between pre-training and translation, MSP divides the translation process via pre-trained language models into multiple separate stages: the encoding stage, the re-encoding stage, and the decoding stage. During each stage, we independently apply different continuous prompts for allowing pre-trained language models better shift to translation tasks. We conduct extensive experiments on three translation tasks. Experiments show that our method can significantly improve the translation performance of pre-trained language models.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Prompting
  • 2.2 mGPT
  • 3 Multi-Stage Prompting
  • 3.1 Deep Continuous Prompts
  • 3.2 Stages
  • 3.3 Training Objective
  • 3.4 Reparameterization
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Main Results
  • 4.3 Comparison with Other LMs
  • 4.4 Comparison with Transformer
  • 4.5 Effect of Prompt Length
  • 4.6 Effect of Stages
  • 4.7 Effect of Reparameterization
  • 4.8 Analysis
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Details of Multilingual GPT
  • A.2 Preprocessing and Postprocessing
  • A.3 Alignment Examples

Knowls

  1. Knowl 1 — Multi-stage prompting separates translation into three passes through one causal language model

    model/method

    For translation from a source sentence xx to a target sentence yy, MSP uses the same fixed multilingual GPT (mGPT) in three consecutive stages, with a separate continuous prompt for each stage. First, an encoding pass processes xx with an encoding prompt and retains its activations. Second, a re-encoding pass processes xx again, conditioned on the retained activations and a re-encoding prompt; this is intended to produce refined source representations in which each representation can use information from the full source sentence. Third, a decoding pass generates yy autoregressively, conditioned on those refined source representations and a decoding prompt. The stages distinguish source encoding, source refinement, and target-language generation rather than asking one prompt to steer all three behaviors at once. The authors also report that additional re-encoding stages can be used; their default configuration has one.

  2. Knowl 2 — MSP uses deep continuous prompts that affect every Transformer attention layer

    definition

    A deep continuous prompt is a sequence of LL trainable continuous vectors. For an mGPT with NN Transformer layers and hidden size dd, each prompt vector concatenates the key-value prompt components for all layers and therefore has dimension 2Nd2Nd. Unlike a prompt that only adds virtual tokens at the embedding layer, a deep prompt directly affects the computation of every attention layer. MSP assigns such prompts to its encoding, re-encoding, and decoding stages, allowing each pass to be steered independently.

  3. Knowl 3 — MSP learns prompts with target-token cross-entropy while freezing mGPT

    equation

    MSP trains its prompts to minimize the average next-token cross-entropy for the reference target sentence. For target tokens y1,…,yTy_1,\ldots,y_T, let gt(d)g_t^{(d)} be the decoder-stage output vector used to predict target position tt, let eve_v be the mGPT embedding of vocabulary token vv, and let V\mathcal{V} be the vocabulary. The loss is

    L=−1T∑t=1Tlog⁡exp⁡(eyt⊤gt(d))∑v∈Vexp⁡(ev⊤gt(d)).\mathcal{L}=-\frac{1}{T}\sum_{t=1}^{T}\log\frac{\exp(e_{y_t}^{\top}g_t^{(d)})}{\sum_{v\in\mathcal{V}}\exp(e_v^{\top}g_t^{(d)})}.

    The source sentence and preceding target tokens condition each prediction through the MSP stages. During training, the parameters of the pretrained mGPT network remain fixed; only the continuous-prompt parameters are optimized.

  4. Knowl 4 — A scalar-scaled prompt parameterization reduces the number of learned parameters

    model/method

    MSP uses a scaled reparameterization for the encoding, re-encoding, and decoding prompts. For stage s∈{e,r,d}s\in\{e,r,d\}, the prompt is parameterized as

    Ps=max⁡(αs,1.0) ϕs,P_s=\max(\alpha_s,1.0)\,\phi_s,

    where αs\alpha_s is a learned scalar and ϕs\phi_s is the stage-specific learned embedding, with ϕs∈R2N×d\phi_s\in\mathbb{R}^{2N\times d} for an mGPT with NN layers and hidden size dd. Each scalar is initialized to 1.01.0. Thus the trainable set is {αe,αr,αd,ϕe,ϕr,ϕd}\{\alpha_e,\alpha_r,\alpha_d,\phi_e,\phi_r,\phi_d\}. The authors describe this as a simpler, lower-parameter alternative to the MLP reparameterization used in prefix-tuning.

  5. Knowl 5 — Experiments use a frozen 560-million-parameter multilingual GPT on three WMT tasks

    experimental setup

    The backbone is mGPT, a 560-million-parameter decoder-only Transformer pretrained on mC4, which covers 101 languages. It has 24 Transformer layers, hidden size 1,024, and a 250,100-token vocabulary. Translation experiments use WMT16 Romanian–English (0.6 million parallel sentence pairs plus 2 million back-translated pairs; newsdev2016 for development and newstest2016 for testing), WMT14 English–German (4.5 million pairs; newstest2013 and newstest2014), and WMT20 English–Chinese (28 million pairs; newstest2019 and newstest2020). Evaluation uses case-sensitive BLEU from SACREBLEU.

    The default prompt length is 128. Prompt tuning and prefix-tuning are trained for 80,000 steps, while MSP is trained for 40,000 steps. Training uses Adam with β1=0.9\beta_1=0.9, β2=0.98\beta_2=0.98, and ϵ=10−9\epsilon=10^{-9}, batches of roughly 32K tokens, 4,000 warmup steps, and a maximum learning rate of 0.02 for prompt tuning and MSP or 7×10−47\times10^{-4} for prefix-tuning. Inference uses beam search with beam size 4; the length penalty is 1.0 for English–Chinese and 0.0 for the other two tasks.

  6. Knowl 6 — MSP substantially outperforms single-prompt baselines on all three translation tasks

    empirical result

    With the same mGPT backbone, MSP obtains the highest BLEU score on each of the three tested translation tasks. The scores, listed in Romanian–English, English–German, English–Chinese order, are 17.7, 5.9, and 4.5 for prompt tuning; 32.5, 17.5, and 21.9 for prefix-tuning; and 34.7, 21.2, and 28.1 for MSP. Their respective three-task averages are 9.4, 23.9, and 28.0 BLEU. The corresponding numbers of trainable parameters per translation task are 131K, 26M, and 19M. Thus MSP improves the average over prompt tuning by 18.6 BLEU and over prefix-tuning by 4.1 BLEU, while using fewer trainable parameters than prefix-tuning.

  7. Knowl 7 — Comparisons show strong English–Chinese and TedTalks results, but a gap to Transformer on WMT14 English–German

    empirical result

    On WMT20 English–Chinese, mGPT with MSP scores 28.1 BLEU using a 560M-parameter decoder-only backbone. The reported encoder–decoder comparison systems score 24.0 for 13B-parameter mT5-XXL with fine-tuning; 24.1 for 11B CPM-2 with prompt tuning and 26.2 with fine-tuning; and 26.8 for 10B Ernie 3.0 with fine-tuning.

    On TedTalks, the multilingual Transformer baseline is trained on parallel data covering 59 languages, whereas MSP is trained separately for each reported language pair. For Bulgarian, Estonian, Italian, Russian, and Turkish, respectively, X→En BLEU is 35.2, 38.0, 34.2, 22.6, and 21.0 for the 437M-parameter Transformer (average 30.2), versus 38.9, 42.1, 37.8, 24.4, and 24.9 for MSP (average 33.6; 19M tunable parameters per direction). In En→X, the corresponding Transformer scores are 29.2, 34.0, 29.2, 16.7, and 11.6 (average 24.1), while MSP scores 34.1, 38.4, 32.8, 19.2, and 15.6 (average 28.0). The paper reports MSP advantages of 3.4 BLEU for X→En and 3.9 for En→X on average.

    On WMT14 English–German, however, the 450M-parameter Transformer-big baseline scores 27.9 BLEU compared with 21.2 for MSP, which uses 19M tunable parameters. The paper reports 21 hours to train MSP prompts for this task versus 72 hours to train the Transformer.

  8. Knowl 8 — Stage ablations show gains from separating stages and adding re-encoding

    empirical result

    On WMT14 English–German and WMT20 English–Chinese, respectively, single-stage prompting scores 17.9 and 22.8 BLEU, using 6.3M trainable parameters and 14 hours of training. Separating encoding and decoding raises the scores to 20.2 and 25.2, with 12.6M parameters and the same reported training time. Adding one re-encoding stage (the default MSP configuration) yields 21.2 and 28.1 with 19.0M parameters and 21 hours of training. Adding a second re-encoding stage yields 21.8 and 28.4 with 25.1M parameters and 29 hours. Sharing one prompt across the encoding, re-encoding, and decoding stages yields 19.8 and 24.5 with 6.3M parameters and 21 hours.

    Reported inference speeds are 0.10 seconds per sentence for the first two settings and 0.11 seconds per sentence for the other three, measured on the test set using 8 GPUs. The results support improvements from stage separation and re-encoding; prompt sharing also improves over single-stage prompting without increasing the parameter count. The authors report no notable inference slowdown from added re-encoding stages because those stages are computed once in parallel during inference.

  9. Knowl 9 — Longer prompts help, while scaled reparameterization speeds convergence

    empirical result

    On WMT14 English–German, increasing prompt length from 64 to 128, 192, and 256 raises MSP BLEU from 19.0 to 21.2, 22.2, and 22.4. Prefix-tuning at the same lengths scores 14.8, 17.5, 18.2, and 18.2. Thus the gains from longer prompts diminish, and MSP with length 64 still exceeds prefix-tuning with length 256 (19.0 versus 18.2 BLEU). The paper reports that longer prompts do not significantly reduce GPU decoding speed.

    In a separate WMT14 English–German training comparison, MSP with scaled reparameterization converges faster than MSP without reparameterization, while the two approaches reach nearly the same BLEU performance once training has converged. The authors consequently use scaled reparameterization to reduce training time rather than to claim a higher final score.

  10. Knowl 10 — Analyses suggest prompts steer generation while the backbone contains translation associations

    limitation

    As evidence about where translation information resides, the authors remove prompts, feed mGPT concatenated parallel sentences, and compare source and target hidden activations. They report that nearest source–target token pairs frequently coincide with bilingual alignments, and infer cautiously that translation knowledge resides mainly in the pretrained LM while prompts guide its use during generation.

    A separate free-generation probe on 100 samples from the English–Chinese task shows different language outputs under different prompts. The top two language identifications and their proportions are: no prompt, English 16% and Russian 10%; prefix-tuning prompt, Chinese 80% and Japanese 12%; MSP encoding prompt, English 51% and Latin 14%; MSP re-encoding prompt, English 24% and Latin 17%; and MSP decoding prompt, Chinese 91% and Japanese 9%. The authors interpret this pattern as a transition from source-language generation toward target-language generation across MSP stages. They did not find the learned prompts interpretable.

    The paper also identifies a performance limitation: a separate Transformer encoder plus adapter network that maps source sentences into deep prompts, with mGPT used as decoder, obtains 25.9 BLEU on WMT14 English–German using 378M trainable parameters, compared with MSP’s 21.2 BLEU. The authors take this as evidence of room to improve prompting, while noting that performance may also be bottlenecked by the backbone LM.

Coverage note — The detailed tokenization and punctuation-handling procedures and the individual bilingual alignment examples are omitted because they are ancillary implementation details and illustrations rather than load-bearing contributions.

References

  1. 1.Dzmitry Bahdanau, KyungHyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of ICLR.
  2. 2.Ankur Bapna and Orhan Firat. 2019. Simple, scalable adaptation for neural machine translation. In Proceedings of EMNLP, pages 1538–1548.
  3. 3.Graeme Blackwood, Miguel Ballesteros, and Todd Ward. 2018. Multilingual neural machine translation with task-specific attention. In Proceedings of COLING, pages 3112–3122.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL, pages 4171–4186.
  6. 6.Tianyu Gao, Adam Fisch, and Danqi Chen. 2020. Making pre-trained language models better few-shot learners. In Proceedings of ACL, pages 3816–3830.
  7. 7.Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, volume 9, pages 249–256.
  8. 8.Junliang Guo, Zhirui Zhang, Linli Xu, Hao-Ran Wei, Boxing Chen, and Enhong Chen. 2020. Incorporating BERT into Parallel Sequence Decoding with Adapters. In Advances in Neural Information Processing Systems, volume 33, pages 10843–10854.
  9. 9.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of ICLR.
  10. 10.Philipp Koehn and Rebecca Knowles. 2017. Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, pages 28–39.
  11. 11.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of EMNLP, pages 3045–3059.
  12. 12.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of ACL, pages 4582–4597.
  13. 13.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  14. 14.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of ACL, pages 311–318.
  15. 15.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation, pages 186–191.
  16. 16.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. OpenAI blog.
  17. 17.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  18. 18.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  19. 19.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of EMNLP, pages 4222–4235.
  20. 20.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053.
  21. 21.Asa Cooper Stickland, Xian Li, and Marjan Ghazvininejad. 2021. Recipes for adapting pre-trained monolingual and multilingual models to machine translation. In Proceedings of EACL, pages 3440–3453.
  22. 22.Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, et al. 2021a. Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation. arXiv preprint arXiv:2107.02137.
  23. 23.Zewei Sun, Mingxuan Wang, and Lei Li. 2021b. Multilingual translation via grafting pre-trained language models. In Findings of EMNLP, pages 2735–2747.
  24. 24.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. Advances in Neural Information Processing Systems, 27:3104–3112.
  25. 25.Zhixing Tan, Jiacheng Zhang, Xuancheng Huang, Gang Chen, Shuo Wang, Maosong Sun, Huanbo Luan, and Yang Liu. 2020. THUMT: An Open-Source Toolkit for Neural Machine Translation. In Proceedings of AMTA, pages 116–122.
  26. 26.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008.
  27. 27.Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP: System Demonstrations, pages 38–45.
  28. 28.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of NAACL, pages 483–498.
  29. 29.Zhengyan Zhang, Yuxian Gu, Xu Han, Shengqi Chen, Chaojun Xiao, Zhenbo Sun, Yuan Yao, Fanchao Qi, Jian Guan, Pei Ke, et al. 2021. CPM-2: Large-scale Cost-effective Pre-trained Language Models. AIOpen, 2:216–224.

Citation

MLA
Tan, Z., et al. “MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators”. arXiv, 2021, http://arxiv.org/abs/2110.06609v2.
APA
Tan, Z., Zhang, X., Wang, S., & Liu, Y. (2021). MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators. arXiv. http://arxiv.org/abs/2110.06609v2
Chicago
Tan, Z., X. Zhang, S. Wang, and Y. Liu. 2021. “MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators”. arXiv. http://arxiv.org/abs/2110.06609v2.
Harvard
Tan, Z. et al. (2021) “MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.06609v2.
Vancouver
1. Tan Z, Zhang X, Wang S, Liu Y (2021) MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators. arXiv

BibTeX

@article{tan2021msp,
  title = {MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators},
  author = {Tan, Zhixing and Zhang, Xiangwen and Wang, Shuo and Liu, Yang},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.06609v2},
  eprint = {2110.06609}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/