Prompting PaLM for Translation: Assessing Strategies and Performance

David VilarMarkus FreitagColin CherryJiaming LuoViresh RatnakarGeorge F. Foster

article2023ACL245 citations

Demonstrates that example quality outweighs domain match and semantic proximity in few-shot prompt selection for translation with PaLM, while establishing that even optimized large language models still lag behind dedicated supervised translation systems across modern benchmarks and human evaluation.

Listen

Recent advances in large language models (LLMs) have demonstrated an unexpected capability to perform multilingual translation despite being trained without explicit parallel text. This article addresses whether general-purpose LLMs can realistically match or replace dedicated, state-of-the-art translation engines in enterprise applications. The main objective of the article is to systematically assess prompting strategies for Google’s 540-billion-parameter Pathways Language Model (PaLM) and rigorously evaluate its translation performance against specialized supervised systems.

To evaluate performance credibility, the researchers conducted extensive sentence-level translation experiments across three high-resource language pairs paired with English (German, Chinese, and French). The study evaluated various few-shot example selection strategies, comparing standard random selection against customized nearest-neighbor retrieval. To prevent data contamination, the analysis utilized recent standard benchmark test sets and assessed quality using both modern neural automated metrics and comprehensive, expert human evaluations based on standardized error-weighting metrics.

Four primary findings emerge from the study. First, prompt example quality is the most critical determinant of output quality; selecting examples from clean, high-quality reference pools consistently improved results, whereas semantic matching via nearest-neighbor search introduced vulnerability to data noise and alignment errors. Second, while PaLM demonstrates impressive few-shot translation ability, its performance consistently lags behind specialized state-of-the-art systems by 1 to 3 metric points and underperforms commercial off-the-shelf tools, showing better relative capability when translating into English rather than out of it. Third, human error analysis reveals that PaLM achieves natural fluency comparable to dedicated systems but suffers from significant accuracy deficits, notably omitting important source information and occasionally hallucinating details. Fourth, earlier claims suggesting LLMs rivaled supervised systems were partly influenced by training-data overlap on older benchmarks, which inflated historical scores by up to 0.7 automated metric points.

These findings indicate that while LLMs produce natural and stylistically sound translations, their tendency to omit content and hallucinate poses compliance, safety, and brand risks in high-stakes operational environments. Furthermore, because LLM translation requires orders of magnitude more computational time and cost than dedicated engines, direct substitution is currently neither cost-effective nor risk-free. Organizations should maintain dedicated machine translation engines for accurate, production-level workflows, using LLMs primarily where fluency and stylistic adaptation take precedence over strict fidelity.

Future development should focus on document-level translation to leverage the extended context windows of LLMs, as well as soft-prompt tuning to reduce factual errors without compromising fluency. Readers should note that these findings are limited to high-resource languages translated to and from English in isolated, sentence-level formats; performance on lower-resource languages or full-context documents may yield different quality characteristics.

arXiv: 2211.09102
Cover for Prompting PaLM for Translation: Assessing Strategies and Performance

Abstract

Large language models (LLMs) that have been trained on multilingual but not parallel text exhibit a remarkable ability to translate between languages. We probe this ability in an in-depth study of the pathways language model (PaLM), which has demonstrated the strongest machine translation (MT) performance among similarly-trained LLMs to date. We investigate various strategies for choosing translation examples for few-shot prompting, concluding that example quality is the most important factor. Using optimized prompts, we revisit previous assessments of PaLM’s MT capabilities with more recent test sets, modern MT metrics, and human evaluation, and find that its performance, while impressive, still lags that of state-of-the-art supervised systems. We conclude by providing an analysis of PaLM’s MT output which reveals some interesting properties and prospects for future work.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • Après nous, le déluge
  • 3 Prompting for Machine Translation
  • 4 Data
  • 5 Experiments
  • 5.1 Selection strategies and pools
  • 5.2 Results on all language pairs
  • 5.3 Comparison to previous results
  • 6 Analysis
  • 6.1 kNN versus random prompts
  • 6.2 Example Translations
  • 6.3 Overlap of test and training data
  • 7 Conclusion
  • Limitations
  • Ethical Considerations
  • References
  • Appendices
  • A Prompt Exploration
  • B High-end pool
  • C Variability of Random Runs
  • D Detailed MQM Scores
  • E Significance numbers
  • F Example Prompts
  • G Example Translations
  • H Overlap Analysis
  • I Fixed versus random prompts

Knowls

  1. Knowl 1 — In-Context Machine Translation Prompting with Decoder-Only LLMs

    algorithm

    Few-shot machine translation with a decoder-only large language model such as PaLM (Pathways Language Model, 540B parameters) formats parallel demonstration sentence pairs alongside the test sentence into an in-context prompt. Demonstration examples are prepended with the language name in English, followed by the source input and a target prefix. Inference is performed via greedy search (temperature T=0T = 0) until a newline character is generated.

    Input: Source text xx, source language name SnameS_{\text{name}}, target language name TnameT_{\text{name}}, parallel pool P\mathcal{P}, shot count nn
    Output: Target translation yy
    Select nn translation example pairs (x1,y1),…,(xn,yn)∈P(x_1, y_1), \dots, (x_n, y_n) \in \mathcal{P}
    Initialize prompt string P←""P \leftarrow \text{""}
    for i←1i \leftarrow 1 to nn do
        $P \leftarrow P + S_{\text{name}} + \text{": "} + x_i + \text{"\n"}
        $P \leftarrow P + T_{\text{name}} + \text{": "} + y_i + \text{"\n"}
    end for
    $P \leftarrow P + S_{\text{name}} + \text{": "} + x + \text{"\n"}
    $P \leftarrow P + T_{\text{name}} + \text{": "}
    Initialize generated tokens y←""y \leftarrow \text{""}
    while true do
        Sample next token t←arg⁡max⁡vPLM(v∣P+y)t \leftarrow \arg\max_v P_{\text{LM}}(v \mid P + y)
        if t="\n"t = \text{"\n"} or t=EOSt = \text{EOS} then
            break
        end if
        y←y+ty \leftarrow y + t
    end while
    return yy
  2. Knowl 2 — Impact of Prompt Template Formatting and Demonstration Count on Translation

    data/table

    In few-shot machine translation with large language models, prompt formatting has a substantial effect when the number of demonstrations is very small (0 to 1 shot), but performance converges across templates as the number of shots reaches 5 or more. Beyond 5 shots, returns diminish.

    The table below presents median sentence-level BLEURT scores across 5 runs on English→\toGerman translation evaluated on WMT news data with randomly selected demonstrations under different prompt templates and shot counts:

    Prompt Template 0 shots 1 shot 2 shots 5 shots 10 shots
    Language 63.9 69.1 71.7 73.6 74.4
    Codes 59.0 68.5 71.2 73.4 74.1
    Header 72.4 69.1 70.7 73.4 74.1
    Textual 36.9 67.5 71.8 73.0 73.7
    Deutsch 72.6 70.8 71.9 73.5 74.1
    None 3.2 38.5 59.6 73.0 74.1

    Template descriptions:

    • Language: Prepends source and target with full English language names (e.g., English: ... \n German: ...).
    • Codes: Uses two-letter ISO language codes (e.g., en: ... \n de: ...).
    • Header: Adds an introductory instruction Translate following sentences: above the Language format.
    • Textual: Explicit instruction sentence Translate X from English into German: Y.
    • Deutsch: Uses target language names in German (Englisch: ... \n Deutsch: ...).
    • None: Concatenates raw source and target sentence pairs with no identifying prefixes.

    Without explicit demonstration pairs (0 shots), templates lacking clear task descriptions (None, Textual, Codes) fail or degrade significantly, whereas localized and instructional headers achieve higher zero-shot scores. At 5 shots, all templates achieve between 73.0 and 73.6 BLEURT.

  3. Knowl 3 — Source-Side kNN Example Retrieval for Machine Translation Prompts

    model/method

    To dynamically condition few-shot prompts on a test input xx, kk-nearest neighbor (kkNN) retrieval selects the kk parallel sentence pairs (xi,yi)(x_i, y_i) from a pool P\mathcal{P} whose source sentences xix_i have the lowest distance to xx. Two representations are used:

    1. Bag-of-Words (BOW): Sentences are represented by sparse word count vectors v(x)∈R∣V∣\mathbf{v}(x) \in \mathbb{R}^{|V|}. Proximity is measured by cosine distance: distBOW(x,x′)=1−v(x)⋅v(x′)∥v(x)∥2∥v(x′)∥2\text{dist}_{\text{BOW}}(x, x') = 1 - \frac{\mathbf{v}(x) \cdot \mathbf{v}(x')}{\|\mathbf{v}(x)\|_2 \|\mathbf{v}(x')\|_2}
    2. Multilingual Dense Embeddings (RoBERTa): Sentences are mapped to continuous dense embeddings e(x)∈Rd\mathbf{e}(x) \in \mathbb{R}^d via a pretrained multilingual RoBERTa model. Proximity is measured using Euclidean distance: distRoBERTa(x,x′)=∥e(x)−e(x′)∥2\text{dist}_{\text{RoBERTa}}(x, x') = \|\mathbf{e}(x) - \mathbf{e}(x')\|_2

    Dense vector indexing and retrieval across large parallel pools (such as the tens of millions of sentence pairs in full WMT training sets) are computed using anisotropic vector quantization (ScaNN).

  4. Knowl 4 — Demonstration Quality Dominates Lexico-Semantic Proximity in Few-Shot MT

    empirical result

    In few-shot machine translation with PaLM 540B, the average quality and domain reliability of the parallel demonstration candidate pool matters substantially more than the lexical or semantic proximity between prompt examples and the input sentence.

    Key observations comparing prompt pools and selection methods on WMT benchmarks:

    1. Randomly selecting demonstrations from a clean, curated development pool (WMT-dev) consistently outperforms kkNN RoBERTa retrieval from the full WMT training corpus (WMT-full), even though WMT-full allows kkNN to find examples with substantially closer semantic matches to the source sentence.
    2. In English↔\leftrightarrowGerman experiments, random selection from WMT-dev yields 74.8 BLEURT (En→\toDe) and 75.9 BLEURT (De→\toEn), compared to 73.0 BLEURT (En→\toDe) and 73.8 BLEURT (De→\toEn) for kkNN RoBERTa on WMT-full.
    3. The inferiority of kkNN over large web-crawled pools arises because kkNN matches strictly on the source side. Web-crawled training corpora contain alignment noise and misaligned target sentences; when kkNN retrieves a source sentence from an incorrectly aligned document, adjacent or semantically similar retrieved demonstrations in the same prompt are also prone to alignment errors, inducing severe hallucinations in the LLM output. Independent random selection avoids correlated alignment noise.
  5. Knowl 5 — Comparative Performance of Few-Shot PaLM Against Supervised SOTA and Commercial MT

    data/table

    Evaluating 5-shot PaLM (540B parameters) against competition-grade supervised systems and commercial translation engines on modern WMT test sets demonstrates that PaLM achieves strong translation ability without parallel pretraining, but significantly lags specialized state-of-the-art (SOTA) MT systems.

    Evaluation is performed on WMT21 news test sets (newstest2021) for German↔\leftrightarrowEnglish and Chinese↔\leftrightarrowEnglish, and WMT14 (newstest2014) for French↔\leftrightarrowEnglish. Quality is measured by BLEURT (RemBERT-based cased), SacreBLEU, and document-context Multidimensional Quality Metrics (MQM, where lower is better, evaluating the first 12 segments per document):

    Language Pair System / Setting MQM ↓\downarrow BLEURT ↑\uparrow BLEU ↑\uparrow
    en →\to de WMT21 Facebook Submission 1.18 76.9 42.0
    Google Translate 1.59 75.7 39.8
    PaLM (WMT-dev, random) 1.58 74.8 32.8
    PaLM (high-end, random) 1.67 74.7 32.9
    PaLM (WMT-full, random) 1.90 73.7 32.9
    PaLM (WMT-full, kkNN) 1.93 73.0 32.5
    de →\to en WMT21 Facebook Submission 1.31 76.9 41.9
    Google Translate 1.71 76.4 40.9
    PaLM (WMT-dev, random) 1.92 75.9 38.0
    PaLM (high-end, random) 1.89 75.8 38.8
    PaLM (WMT-full, random) 2.38 74.7 38.3
    PaLM (WMT-full, kkNN) 3.03 73.8 35.4
    en →\to zh WMT21 WeChat Submission 2.47 66.6 36.9
    Google Translate 3.23 65.0 36.2
    PaLM (WMT-dev, random) 3.24 64.1 29.2
    PaLM (high-end, random) 3.70 63.9 29.6
    PaLM (WMT-full, random) 4.35 62.2 28.6
    PaLM (WMT-full, kkNN) 5.06 60.7 28.5
    zh →\to en WMT21 Borderline Submission 3.11 70.0 33.4
    Google Translate 3.12 69.5 32.2
    PaLM (WMT-dev, random) 3.60 67.5 25.3
    PaLM (high-end, random) 3.89 67.7 25.1
    PaLM (WMT-full, random) 3.95 67.2 25.8
    PaLM (WMT-full, kkNN) 4.06 65.8 23.8
    en →\to fr Google Translate – 76.5 45.7
    PaLM (WMT-full, random) – 75.9 42.3
    PaLM (WMT-dev, random) – 75.4 41.9
    fr →\to en Google Translate – 77.7 43.2
    PaLM (WMT-full, random) – 77.7 42.7
    PaLM (high-end, random) – 77.6 40.4

    Supervised SOTA systems maintain a 1.0 to 3.0 BLEURT point advantage over the best PaLM configurations. PaLM consistently performs closer to supervised systems when translating into English than in the reverse translation direction.

  6. Knowl 6 — Fluency vs Accuracy Asymmetry in Large Language Model Translation Errors

    empirical result

    Detailed Multidimensional Quality Metrics (MQM) expert human evaluation reveals that the quality gap between few-shot PaLM (WMT-dev random) and supervised SOTA systems is primarily driven by accuracy errors rather than fluency errors.

    1. Fluency Parity: PaLM achieves fluency error rates comparable to or better than supervised SOTA models. For instance, in style assessment, PaLM generates fewer Minor Style/Awkward errors than SOTA in De→\toEn (73 for PaLM vs 81 for SOTA) and Zh→\toEn (205 for PaLM vs 284 for SOTA).
    2. Accuracy Deficit: PaLM exhibits a substantially higher rate of Major Accuracy/Omission errors across all evaluated language pairs:
      • De→\toEn: 51 major omission errors for PaLM vs 19 for SOTA.
      • En→\toDe: 26 major omission errors for PaLM vs 7 for SOTA.
      • Zh→\toEn: 109 major omission errors for PaLM vs 42 for SOTA.
      • En→\toZh: 80 major omission errors for PaLM vs 46 for SOTA.
    3. Translation Style: Supervised neural machine translation generates literal, faithful sentence-level translations, preventing omissions but occasionally yielding unidiomatic expressions (such as translating street names literally or copying source date formats). PaLM generates less literal, more fluent target phrasing, but frequently drops source clauses or hallucinates unsupported facts.
  7. Knowl 7 — Impact of Pretraining Target Overlap on Machine Translation Evaluation

    empirical result

    Evaluating LLM translation on older test sets introduces performance inflation due to target-side training data contamination. Measuring 15-gram target matches via the mBERT tokenizer against PaLM's 780B token training corpus demonstrates substantial contamination in older WMT benchmarks and negligible overlap in newer benchmarks.

    Clean test set proportions (fraction of sentences lacking 15-gram target overlap with PaLM training data):

    • WMT14 French →\to English: 69.2%69.2\% clean (30.8%30.8\% overlap)
    • WMT14 English →\to French: 93.6%93.6\% clean (6.4%6.4\% overlap)
    • WMT16 German →\to English: 80.3%80.3\% clean (19.7%19.7\% overlap)
    • WMT16 English →\to German: 97.3%97.3\% clean (2.7%2.7\% overlap)
    • WMT21 English →\to German: 99.6%99.6\% clean
    • WMT21 German →\to English: 97.9%97.9\% clean
    • WMT21 English →\to Chinese: 99.7%99.7\% clean
    • WMT21 Chinese →\to English: 98.1%98.1\% clean

    Comparing 5-shot PaLM to contamination-free Google Translate (GT) on clean vs overlapping (eg egclean) subsets shows that target-side overlap inflates PaLM's relative performance by 0.50.5 to 0.80.8 BLEU points on older benchmarks (e.g., De→\toEn 2016 original delta: 1.51.5 BLEU vs clean delta: 2.02.0 BLEU; Fr→\toEn 2014 original delta: 0.10.1 BLEU vs clean delta: 0.80.8 BLEU). Neural metrics like BLEURT exhibit lower sensitivity to overlap, shifting relative deltas by ≤0.3\le 0.3 points.

  8. Knowl 8 — Maximum-Likelihood Prompt Selection for Fixed Few-Shot In-Context MT

    empirical result

    Instead of dynamically sampling random demonstrations for each input sentence, a single static high-quality demonstration prompt can be selected for all inputs by maximizing the language model probability over a held-out set of parallel examples.

    For a candidate 1-shot paragraph prompt PP from a curated high-end pool, its score is computed as the total log-likelihood log⁡PPaLM(Dheld-out∣P)\log P_{\text{PaLM}}(\mathcal{D}_{\text{held-out}} \mid P) of a held-out set Dheld-out\mathcal{D}_{\text{held-out}} conditioned on PP. The top-ranked prompt is then fixed for all evaluation inputs.

    Comparison of BLEURT scores on standard WMT test sets between the fixed maximum-likelihood prompt and random prompt selection (showing min, average, and max over 5 runs):

    Language Pair Fixed Prompt Random (min) Random (avg) Random (max)
    en →\to de 74.7 74.5 74.7 75.0
    de →\to en 76.3 75.6 75.8 75.9
    en →\to zh 64.7 63.7 63.9 64.0
    zh →\to en 67.0 67.3 67.5 67.7
    en →\to fr 75.5 75.2 75.2 75.3
    fr →\to en 77.9 77.4 77.6 77.6

    Using a single high-likelihood prompt meets or exceeds the average performance of per-input random demonstration sampling across almost all language directions while eliminating runtime variance.

Coverage note — None was omitted; all key contributed prompting strategies, empirical evaluations against SOTA across language pairs, MQM error analyses, contamination evaluations, and prompt selection methodologies are fully documented.

References

  1. 1.Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2022. In-context examples selection for machine translation. arXiv preprint arXiv:2212.02437.
  2. 2.Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. 2021. Findings of the 2021 conference on machine translation (WMT21). In Proceedings of the Sixth Conference on Machine Translation, pages 1–88, Online. Association for Computational Linguistics.
  3. 3.Anonymous. 2023. Does gpt-3 produces less literal translations? Anonymous preprint under review.
  4. 4.Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al. 2022. Promptsource: An integrated development environment and repository for natural language prompts. arXiv preprint arXiv:2202.01279.
  5. 5.Rachel Bawden and François Yvon. 2023. Investigating the translation performance of a large multilingual language model: the case of bloom. arXiv preprint arXiv:2303.01911.
  6. 6.Eleftheria Briakou, Colin Cherry, and George Foster. 2023. Searching for needles in a haystack: On the role of incidental bilingualism in palm’s translation capability. arXiv preprint arXiv:2305.10266.
  7. 7.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder. 2020. Rethinking embedding coupling in pre-trained language models. arXiv preprint:2010.12821.
  10. 10.Marta R. Costa-jussà, Eric Smith, Christophe Ropers, Daniel Licht, Javier Ferrando, and Carlos Escolano. 2022. Toxicity in multilingual machine translation at scale. arXiv preprint arXiv:2210.03070.
  11. 11.Daniel Deutsch, Rotem Dror, and Dan Roth. 2021. A statistical analysis of summarization evaluation metrics using resampling methods. Transactions of the Association for Computational Linguistics, 9:1132–1146.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Markus Freitag, Isaac Caswell, and Scott Roy. 2019. APE at scale and its implications on MT evaluation biases. In Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers), pages 34–44, Florence, Italy. Association for Computational Linguistics.
  14. 14.Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a. Experts, errors, and context: A large-scale study of human evaluation for machine translation. Transactions of the Association for Computational Linguistics, 9:1460–1474.
  15. 15.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online. Association for Computational Linguistics.
  16. 16.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830.
  17. 17.Xavier Garcia, Yamini Bansal, Colin Cherry, George Foster, Maxim Krikun, Fangxiaoyu Feng, Melvin Johnson, and Orhan Firat. 2023. The unreasonable effectiveness of few-shot learning for machine translation. arXiv preprint arXiv:2302.01398.
  18. 18.Xavier Garcia and Orhan Firat. 2022. Using natural language prompts for machine translation. arXiv preprint arXiv:2202.11822.
  19. 19.Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant. 2022. What can transformers learn in-context? a case study of simple function classes. arXiv preprint arXiv:2208.01066.
  20. 20.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  21. 21.Marjan Ghazvininejad, Hila Gonen, and Luke Zettlemoyer. 2023. Dictionary-based phrase-level prompting of large language models for machine translation. arXiv preprint arXiv:2302.07856.
  22. 22.Nuno M Guerreiro, Duarte Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and André FT Martins. 2023. Hallucinations in large multilingual translation models. arXiv preprint arXiv:2303.16104.
  23. 23.Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020. Accelerating large-scale inference with anisotropic vector quantization. In International Conference on Machine Learning.
  24. 24.Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Zeyu Wang, Luke Zettlemoyer, and Noah A. Smith. 2022. Whose language counts as high quality? measuring language ideologies in text data selection. arXiv preprint arXiv:2201.10474.
  25. 25.Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021. WARP: Word-level Adversarial ReProgramming. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4921–4933, Online. Association for Computational Linguistics.
  26. 26.Zhiwei He, Tian Liang, Wenxiang Jiao, Zhuosheng Zhang, Yujiu Yang, Rui Wang, Zhaopeng Tu, Shuming Shi, and Xing Wang. 2023. Exploring human-like translation strategy with large language models. arXiv preprint arXiv:2305.04118.
  27. 27.Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210.
  28. 28.Yutai Hou, Hongyuan Dong, Xinghao Wang, Bohan Li, and Wanxiang Che. 2022. Metaprompting: Learning to learn better prompts. arXiv preprint arXiv:2209.11486.
  29. 29.Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745.
  30. 30.Alex Jones, Isaac Caswell, Ishank Saxena, and Orhan Firat. 2023. Bilex rx: Lexical data augmentation for massively multilingual machine translation. arXiv preprint arXiv:2303.15265.
  31. 31.Marzena Karpinska and Mohit Iyyer. 2023. Large language models effectively leverage document-level context for literary translation, but critical errors persist. arXiv preprint arXiv:2304.03245.
  32. 32.Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2021. Nearest neighbor machine translation. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  33. 33.Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation. In Proceedings of the Sixth Conference on Machine Translation, pages 478–494, Online. Association for Computational Linguistics.
  34. 34.Philipp Koehn. 2004. Statistical significance tests for machine translation evaluation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pages 388–395, Barcelona, Spain. Association for Computational Linguistics.
  35. 35.Sawan Kumar and Partha Talukdar. 2021. Reordering examples helps during priming-based few-shot learning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4507–4518.
  36. 36.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  37. 37.Jiaoda Li, Ryan Cotterell, and Mrinmaya Sachan. 2022a. Probing via prompting. arXiv preprint arXiv:2207.01736.
  38. 38.Junyi Li, Tianyi Tang, Jian-Yun Nie, Ji-Rong Wen, and Xin Zhao. 2022b. Learning to transfer prompts for text generation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3506–3518, Seattle, United States. Association for Computational Linguistics.
  39. 39.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  40. 40.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022a. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. arXiv preprint arXiv:2205.05638.
  41. 41.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022b. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  42. 42.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  43. 43.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  44. 44.Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014. Multidimensional Quality Metrics (MQM) : A Framework for Declaring and Describing Translation Quality Metrics. Tradumàtica, pages 0455–463.
  45. 45.Hongyuan Lu, Haoyang Huang, Dongdong Zhang, Haoran Yang, Wai Lam, and Furu Wei. 2023. Chain-of-dictionary prompting elicits translation in large language models. arXiv preprint arXiv:2305.06575.
  46. 46.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  47. 47.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  48. 48.Yasmin Moslem, Rejwanul Haque, and Andy Way. 2023. Adaptive machine translation with large language models. arXiv preprint arXiv:2301.13294.
  49. 49.Ajay Patel, Bryan Li, Mohammad Sadegh Rasooli, Noah Constant, Colin Raffel, and Chris Callison-Burch. 2022. Bidirectional language models are also few-shot learners. arXiv preprint arXiv:2209.14500.
  50. 50.Jonathan Pilault, Xavier Garcia, Arthur Bražinskas, and Orhan Firat. 2023. Interactive-chain-prompting: Ambiguity resolution for crosslingual conditional generation with interaction. arXiv preprint arXiv:2301.10309.
  51. 51.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium. Association for Computational Linguistics.
  52. 52.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  53. 53.Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–7.
  54. 54.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269.
  55. 55.Andrea Schioppa, Xavier Garcia, and Orhan Firat. 2023. Cross-lingual supervision improves large language models pre-training. arXiv preprint arXiv:2305.11778.
  56. 56.Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, Online. Association for Computational Linguistics.
  57. 57.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235.
  58. 58.Chandan Singh, John X Morris, Jyoti Aneja, Alexander M Rush, and Jianfeng Gao. 2022. Explaining patterns in data with language models via interpretable autoprompting. arXiv preprint arXiv:2210.01848.
  59. 59.Hendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover, Johanna Beyer, Hanspeter Pfister, and Alexander M Rush. 2022. Interactive and visual prompt engineering for ad-hoc task adaptation with large language models. IEEE transactions on visualization and computer graphics.
  60. 60.Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021. Facebook AI’s WMT21 news translation task submission. In Proceedings of the Sixth Conference on Machine Translation, pages 205–215, Online. Association for Computational Linguistics.
  61. 61.Josef Valvoda, Yimai Fang, and David Vandyke. 2022. Prompting for a conversation: How to control a dialog model? arXiv preprint arXiv:2209.11068.
  62. 62.Longyue Wang, Mu Li, Fangxu Liu, Shuming Shi, Zhaopeng Tu, Xing Wang, Shuangzhi Wu, Jiali Zeng, and Wen Zhang. 2021. Tencent translation system for the WMT21 news translation task. In Proceedings of the Sixth Conference on Machine Translation, pages 216–224, Online. Association for Computational Linguistics.
  63. 63.Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, and Zhaopeng Tu. 2023. Document-level machine translation with large language models. arXiv preprint arXiv:2304.02210.
  64. 64.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  65. 65.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  66. 66.Xianfeng Zeng, Yijin Liu, Ernan Li, Qiu Ran, Fandong Meng, Peng Li, Jinan Xu, and Jie Zhou. 2021. WeChat neural machine translation systems for WMT21. In Proceedings of the Sixth Conference on Machine Translation, pages 243–254, Online. Association for Computational Linguistics.
  67. 67.Biao Zhang, Barry Haddow, and Alexandra Birch. 2023. Prompting large language model for machine translation: A case study. arXiv preprint arXiv:2301.07069.
  68. 68.Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Lingpeng Kong, Jiajun Chen, Lei Li, and Shujian Huang. 2023. Multilingual machine translation with large language models: Empirical results and analysis. arXiv preprint arXiv:2304.04675.

Citation

MLA
Vilar, D., et al. “Prompting PaLM for Translation: Assessing Strategies and Performance”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15406–27, https://doi.org/10.18653/v1/2023.acl-long.859.
APA
Vilar, D., Freitag, M., Cherry, C., Luo, J., Ratnakar, V., & Foster, G. (2023). Prompting PaLM for Translation: Assessing Strategies and Performance. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15406–15427. https://doi.org/10.18653/v1/2023.acl-long.859
Chicago
Vilar, D., M. Freitag, C. Cherry, J. Luo, V. Ratnakar, and G. Foster. 2023. “Prompting PaLM for Translation: Assessing Strategies and Performance”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15406–27. https://doi.org/10.18653/v1/2023.acl-long.859.
Harvard
Vilar, D. et al. (2023) “Prompting PaLM for Translation: Assessing Strategies and Performance”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15406–15427. Available at: https://doi.org/10.18653/v1/2023.acl-long.859.
Vancouver
1. Vilar D, Freitag M, Cherry C, Luo J, Ratnakar V, Foster G (2023) Prompting PaLM for Translation: Assessing Strategies and Performance. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15406–15427

BibTeX

@inproceedings{vilar-etal-2023-prompting,
    title = "Prompting {P}a{LM} for Translation: Assessing Strategies and Performance",
    author = "Vilar, David  and
      Freitag, Markus  and
      Cherry, Colin  and
      Luo, Jiaming  and
      Ratnakar, Viresh  and
      Foster, George",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.859/",
    doi = "10.18653/v1/2023.acl-long.859",
    pages = "15406--15427"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/