Prompting PaLM for Translation: Assessing Strategies and Performance
David VilarMarkus FreitagColin CherryJiaming LuoViresh RatnakarGeorge F. Foster
Demonstrates that example quality outweighs domain match and semantic proximity in few-shot prompt selection for translation with PaLM, while establishing that even optimized large language models still lag behind dedicated supervised translation systems across modern benchmarks and human evaluation.
Recent advances in large language models (LLMs) have demonstrated an unexpected capability to perform multilingual translation despite being trained without explicit parallel text. This article addresses whether general-purpose LLMs can realistically match or replace dedicated, state-of-the-art translation engines in enterprise applications. The main objective of the article is to systematically assess prompting strategies for Google’s 540-billion-parameter Pathways Language Model (PaLM) and rigorously evaluate its translation performance against specialized supervised systems.
To evaluate performance credibility, the researchers conducted extensive sentence-level translation experiments across three high-resource language pairs paired with English (German, Chinese, and French). The study evaluated various few-shot example selection strategies, comparing standard random selection against customized nearest-neighbor retrieval. To prevent data contamination, the analysis utilized recent standard benchmark test sets and assessed quality using both modern neural automated metrics and comprehensive, expert human evaluations based on standardized error-weighting metrics.
Four primary findings emerge from the study. First, prompt example quality is the most critical determinant of output quality; selecting examples from clean, high-quality reference pools consistently improved results, whereas semantic matching via nearest-neighbor search introduced vulnerability to data noise and alignment errors. Second, while PaLM demonstrates impressive few-shot translation ability, its performance consistently lags behind specialized state-of-the-art systems by 1 to 3 metric points and underperforms commercial off-the-shelf tools, showing better relative capability when translating into English rather than out of it. Third, human error analysis reveals that PaLM achieves natural fluency comparable to dedicated systems but suffers from significant accuracy deficits, notably omitting important source information and occasionally hallucinating details. Fourth, earlier claims suggesting LLMs rivaled supervised systems were partly influenced by training-data overlap on older benchmarks, which inflated historical scores by up to 0.7 automated metric points.
These findings indicate that while LLMs produce natural and stylistically sound translations, their tendency to omit content and hallucinate poses compliance, safety, and brand risks in high-stakes operational environments. Furthermore, because LLM translation requires orders of magnitude more computational time and cost than dedicated engines, direct substitution is currently neither cost-effective nor risk-free. Organizations should maintain dedicated machine translation engines for accurate, production-level workflows, using LLMs primarily where fluency and stylistic adaptation take precedence over strict fidelity.
Future development should focus on document-level translation to leverage the extended context windows of LLMs, as well as soft-prompt tuning to reduce factual errors without compromising fluency. Readers should note that these findings are limited to high-resource languages translated to and from English in isolated, sentence-level formats; performance on lower-resource languages or full-context documents may yield different quality characteristics.
- Paper: PaLM: Scaling Language Modeling with Pathways, Aakanksha Chowdhery et al. (2023). Introduces the Pathways Language Model (PaLM) architecture, training details, and baseline few-shot capabilities evaluated and revisited in this study.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Establishes foundational insights into what makes in-context demonstration examples effective in large language models, directly informing the prompt-selection strategies analyzed in the paper.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Provides fundamental principles of prompt programming and zero- versus few-shot task framing in large language models that underpin the paper's prompting investigations.
- Paper: Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation, Melvin Johnson et al. (2016). Lays the groundwork for multilingual neural translation architectures and zero-shot cross-lingual transfer that multilingual language models build upon.
- Paper: Improving Neural Machine Translation Models with Monolingual Data, Rico Sennrich et al. (2016). Pioneers the use of monolingual data and back-translation in neural machine translation, contextualizing how models translate without explicit parallel training data.
- Paper: PaLM 2 Technical Report, Rohan Anil et al. (2023). Introduces PaLM 2, directly scaling and refining the multilingual pre-training mixture and translation capabilities analyzed in the PaLM prompting study.
- Paper: Prompting Large Language Model for Machine Translation: A Case Study, Biao Zhang et al. (2023). Expands on machine translation prompting strategies by systematically investigating demonstration selection, templates, and pseudo-parallel data on another massive language model.
- Paper: Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation, Haoran Xu et al. (2024). Builds on LLM translation limitations identified in prompt-based studies by introducing preference optimization to surpass supervised translation benchmarks.
- Paper: Document-Level Machine Translation with Large Language Models, Longyue Wang et al. (2023). Extends the evaluation of few-shot LLM translation from sentence-level prompting to full document-level translation and discourse-level properties.
- Paper: BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer, Akari Asai et al. (2024). Provides a broader, standardized multi-task evaluation comparing few-shot in-context learning with fine-tuning across 54 typologically diverse languages.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Generalizes the evaluation of multilingual prompting across 70 languages to benchmark the broader non-English capabilities of generative language models.
- Paper: Do Llamas Work in English? On the Latent Language of Multilingual Transformers, Chris Wendler et al. (2024). Investigates the latent internal mechanisms and representation spaces that enable multilingual transformers to translate and process non-English text.
