Prompting Large Language Model for Machine Translation: A Case Study
Biao ZhangBarry HaddowAlexandra Birch
Presents a systematic evaluation of prompting strategies for machine translation, identifying how demonstration quality, pseudo-parallel data from monolingual text, and cross-domain transfer govern large language model performance.
Large language models have shown remarkable capabilities across diverse tasks without task-specific training, yet their application to machine translation remains underexplored. The article systematically examines strategies for translation prompting, focusing on prompt templates, demonstration example selection, the role of monolingual data, and transfer learning across languages and domains.
The article evaluates these prompting techniques using the raw, 130-billion-parameter GLM-130B model across English, German, and Chinese language pairs. The empirical evaluation covers standard multilingual benchmarks across general, news, and specialized domains, testing zero-shot and few-shot configurations alongside automated translation quality metrics.
The evaluation reveals four primary findings. First, simple English-language templates specifying source and target tags outperform complex task instructions or templates in other languages. Second, few-shot prompting generally improves translation quality over zero-shot baselines as the number of demonstration examples increases, though performance variance remains high. Third, directly feeding monolingual or mismatched data as demonstrations degrades translation quality; however, synthesizing pseudo-parallel examples via back-translation effectively enhances performance. Fourth, demonstration features such as sequence length, semantic similarity, and model likelihood correlate with output quality, but these correlations are weak, meaning top-performing examples in one domain or language pair rarely transfer their superiority to another.
These findings indicate that large language models require explicit source-to-target mapping signals rather than generic context to translate accurately. The model's cross-lingual capabilities remain heavily English-centric, struggling with direct non-English translations unless routed through English pivoting. Additionally, translation prompting exhibits vulnerability to hallucinations, entity errors, and prompt traps, where prompt instructions are mistakenly copied into the output.
Organizations implementing large language models for translation should utilize concise English prompt templates and construct balanced, pseudo-parallel demonstrations using back-translation when parallel data is scarce. Prompt engineering should be tailored per language pair and domain, and direct non-English translations should use English as an intermediate pivot. Because the evaluation relies on a single quantized model across three languages, further validation across other architectures and language families is recommended before large-scale deployment.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). Establishes semantic-similarity-based demonstration retrieval for in-context learning, providing the direct conceptual foundation for the source paper's exploration of prompt example selection.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Investigates how different components and label qualities of in-context demonstrations affect language model outputs, directly informing the source's empirical study of prompt example quality and formatting.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). Provides a comprehensive taxonomy and formal framework for prompting paradigms across natural language processing tasks, framing the foundational concepts evaluated in this case study.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Demonstrates early findings on zero-shot versus few-shot prompt formulation in translation tasks using large language models, which the source builds upon with broader empirical investigations.
- Paper: Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity, Yao Lu et al. (2021). Analyzes prompt sensitivity and demonstration ordering effects in few-shot learning, establishing the instability challenges that the source addresses during prompt strategy evaluation.
- Paper: Calibrate Before Use: Improving Few-Shot Performance of Language Models, Tony Z. Zhao et al. (2021). Explores inherent model biases and volatility driven by prompt example selection, motivating the systematic analysis of demonstration features conducted in the source.
- Paper: Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering, Zhiyong Wu et al. (2023). Extends the study of demonstration selection by proposing an information-theoretic framework to dynamically optimize both in-context example selection and ordering for individual test queries.
- Paper: UPRISE: Universal Prompt Retrieval for Improving Zero-Shot Evaluation, Daixuan Cheng et al. (2023). Generalizes prompt demonstration selection into a universal cross-task retrieval framework that improves zero-shot and few-shot evaluation across varied model families.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). Broadens the evaluation of prompt-based multilingual and cross-lingual capabilities across 54 typologically diverse languages to assess the limits of few-shot cross-lingual transfer.
- Paper: A Survey on In-context Learning, Qingxiu Dong et al. (2024). Synthesizes empirical findings on demonstration selection, retrieval strategies, and scoring stability into a comprehensive broader survey on in-context learning mechanisms.
