Prompting Large Language Model for Machine Translation: A Case Study

Biao ZhangBarry HaddowAlexandra Birch

article2023ICML398 citations

Presents a systematic evaluation of prompting strategies for machine translation, identifying how demonstration quality, pseudo-parallel data from monolingual text, and cross-domain transfer govern large language model performance.

Listen

Large language models have shown remarkable capabilities across diverse tasks without task-specific training, yet their application to machine translation remains underexplored. The article systematically examines strategies for translation prompting, focusing on prompt templates, demonstration example selection, the role of monolingual data, and transfer learning across languages and domains.

The article evaluates these prompting techniques using the raw, 130-billion-parameter GLM-130B model across English, German, and Chinese language pairs. The empirical evaluation covers standard multilingual benchmarks across general, news, and specialized domains, testing zero-shot and few-shot configurations alongside automated translation quality metrics.

The evaluation reveals four primary findings. First, simple English-language templates specifying source and target tags outperform complex task instructions or templates in other languages. Second, few-shot prompting generally improves translation quality over zero-shot baselines as the number of demonstration examples increases, though performance variance remains high. Third, directly feeding monolingual or mismatched data as demonstrations degrades translation quality; however, synthesizing pseudo-parallel examples via back-translation effectively enhances performance. Fourth, demonstration features such as sequence length, semantic similarity, and model likelihood correlate with output quality, but these correlations are weak, meaning top-performing examples in one domain or language pair rarely transfer their superiority to another.

These findings indicate that large language models require explicit source-to-target mapping signals rather than generic context to translate accurately. The model's cross-lingual capabilities remain heavily English-centric, struggling with direct non-English translations unless routed through English pivoting. Additionally, translation prompting exhibits vulnerability to hallucinations, entity errors, and prompt traps, where prompt instructions are mistakenly copied into the output.

Organizations implementing large language models for translation should utilize concise English prompt templates and construct balanced, pseudo-parallel demonstrations using back-translation when parallel data is scarce. Prompt engineering should be tailored per language pair and domain, and direct non-English translations should use English as an intermediate pivot. Because the evaluation relies on a single quantized model across three languages, further validation across other architectures and language families is recommended before large-scale deployment.

arXiv: 2301.07069
Cover for Prompting Large Language Model for Machine Translation: A Case Study

Abstract

Research on prompting has shown excellent performance with little or even no supervised training across many tasks. However, prompting for machine translation is still under-explored in the literature. We fill this gap by offering a systematic study on prompting strategies for translation, examining various factors for prompt template and demonstration example selection. We further explore the use of monolingual data and the feasibility of cross-lingual, cross-domain, and sentence-to-document transfer learning in prompting. Extensive experiments with GLM-130B (Zeng et al., 2022) as the testbed show that 1) the number and the quality of prompt examples matter, where using suboptimal examples degenerates translation; 2) several features of prompt examples, such as semantic similarity, show significant Spearman correlation with their prompting performance; yet, none of the correlations are strong enough; 3) using pseudo parallel prompt examples constructed from monolingual data via zero-shot prompting could improve translation; and 4) improved performance is achievable by transferring knowledge from prompt examples selected in other settings. We finally provide an analysis on the model outputs and discuss several problems that prompting still suffers from.

Table of Contents

  • 1 Introduction
  • 2 Setup
  • 3 Prompting Strategy for MT
  • 4 Monolingual Data for Prompting
  • 5 Transfer Learning for Prompting
  • 6 Discussion
  • 7 Related Work
  • 8 Conclusion and Future Work
  • References
  • A Appendix

Knowls

  1. Knowl 1 — Prompt Template Formulation and Language Choice for Machine Translation Prompting

    empirical result

    In prompting large language models (LLMs) for machine translation (MT), a test source sentence XX is converted into a prompt via template T\mathcal{T}, and an LLM L\mathcal{L} generates target translation YY.

    The zero-shot MT prompt template format is: [src]: X [tgt]:\text{[src]: } X \text{ [tgt]:} where [src]\text{[src]} and [tgt]\text{[tgt]} denote the natural language names of the source and target languages, respectively.

    For KK-shot prompting with demonstration set DP={Xi′,Yi′}i=1K\mathcal{D}^P = \{X'_i, Y'_i\}_{i=1}^K, prompt examples are concatenated: [psrc]: X1′ [ptgt]: Y1′…[psrc]: XK′ [ptgt]: YK′ [src]: X [tgt]:\text{[psrc]: } X'_1 \text{ [ptgt]: } Y'_1 \ldots \text{[psrc]: } X'_K \text{ [ptgt]: } Y'_K \text{ [src]: } X \text{ [tgt]:} where [psrc]\text{[psrc]} and [ptgt]\text{[ptgt]} denote the source and target language names of the prompt demonstration examples.

    Evaluating 6 prompt templates across English, German, and Chinese template languages on the FLORES Wiki Ablation benchmark across 6 translation directions (En↔De\text{En}\leftrightarrow\text{De}, En↔Zh\text{En}\leftrightarrow\text{Zh}, De↔Zh\text{De}\leftrightarrow\text{Zh}) using GLM-130B (INT4 quantized) demonstrates:

    1. A simple English template without line breaks specifying only language names (Template A: [src]: [input] [tgt]:) achieves the best zero-shot translation performance, averaging 25.9825.98 BLEU and 38.7838.78 COMET (wmt20-comet-da).
    2. Inserting line breaks into the template consistently degrades quality (e.g., Template A drops to 24.7824.78 BLEU and 31.1731.17 COMET).
    3. Expressing template instructions in non-English languages (German or Chinese) severely degrades translation quality on average (German Template A achieves COMET −26.15-26.15; Chinese Template A achieves COMET 14.8214.82), though a Chinese template provides modest gains when translating specifically into Chinese.
  2. Knowl 2 — Correlation between Demonstration Features and MT In-Context Learning Performance

    empirical result

    To analyze which demonstration properties predict in-context translation quality, seven demonstration features were evaluated for 1-shot MT prompting using GLM-130B on the FLORES Wiki Ablation benchmark:

    • SLength\text{SLength} and TLength\text{TLength}: Number of tokens in the demonstration source and target sentence, respectively.
    • LMScore\text{LMScore}: Length-normalized log-likelihood of the demonstration sentence pair under GLM-130B.
    • MTScore\text{MTScore}: Demonstration translation quality scored by the COMET QE model (wmt20-comet-qe-da).
    • SemScore\text{SemScore}: Semantic similarity computed as the cosine similarity between LASER2 sentence embeddings of the demonstration source and target sentences.
    • CaseSemScore-Src\text{CaseSemScore-Src}: Cosine similarity between the test input sentence embedding and the demonstration source sentence embedding.
    • CaseSemScore-Tgt\text{CaseSemScore-Tgt}: Cosine similarity between the test input sentence embedding and the demonstration target sentence embedding.

    Key findings based on Spearman rank correlation ρ\rho with COMET and BLEU:

    1. When demonstrations are drawn exclusively from a high-quality candidate pool, correlations with translation performance are weak and frequently statistically insignificant (ρ∈[−0.01,0.23]\rho \in [-0.01, 0.23] for COMET; ρ∈[0.04,0.23]\rho \in [0.04, 0.23] for BLEU), indicating that diverse high-quality examples yield similar translation quality.
    2. When combining high-quality and low-quality (WikiMatrix.v1) pools, correlations strengthen. Average COMET correlations rank as follows: LMScore\text{LMScore} (ρ=0.31\rho = 0.31), CaseSemScore-Tgt\text{CaseSemScore-Tgt} (ρ=0.31\rho = 0.31), SemScore\text{SemScore} (ρ=0.30\rho = 0.30), TLength\text{TLength} (ρ=0.29\rho = 0.29), CaseSemScore-Src\text{CaseSemScore-Src} (ρ=0.28\rho = 0.28), SLength\text{SLength} (ρ=0.26\rho = 0.26), and MTScore\text{MTScore} (ρ=0.19\rho = 0.19).
    3. All Spearman correlations remain below 0.50.5, showing that no single feature strongly guarantees optimal demonstration performance, and simple length features (S/TLength\text{S/TLength}) correlate almost as strongly as semantic or LLM-based scores.
  3. Knowl 3 — Demonstration Selection Strategies for High- and Low-Quality MT Candidate Pools

    algorithm

    Demonstration selection strategies for 1-shot and 5-shot machine translation prompting differ depending on whether the candidate example pool is clean (high-quality) or noisy (low-quality).

    Input: Candidate demonstration pool DD, test source sentence XX, desired shot count KK, pool quality flag Q∈{High,Low}Q \in \{\text{High}, \text{Low}\}
    Output: Ordered demonstration set DP={(X1′,Y1′),…,(XK′,YK′)}D^P = \{(X'_1, Y'_1), \ldots, (X'_K, Y'_K)\}
    if Q==HighQ == \text{High} then
        Filter Dfiltered={(X′,Y′)∈D:10≤Length(X′)≤100 and 10≤Length(Y′)≤100}D_{filtered} = \{(X', Y') \in D : 10 \le \text{Length}(X') \le 100 \text{ and } 10 \le \text{Length}(Y') \le 100\}
        Compute score s(X′,Y′)=SemScore(X′,Y′)s(X', Y') = \text{SemScore}(X', Y') for each (X′,Y′)∈Dfiltered(X', Y') \in D_{filtered}
        Select the top-KK examples with highest s(X′,Y′)s(X', Y')
        Sort these KK examples in ascending order of their score ss
        return Sorted top-KK examples as DPD^P
    else
        Compute SemScore(X′,Y′)=cos⁡(LASER2(X′),LASER2(Y′))\text{SemScore}(X', Y') = \cos(\text{LASER2}(X'), \text{LASER2}(Y')) for all (X′,Y′)∈D(X', Y') \in D
        Select top 11,000 examples with highest SemScore\text{SemScore}
        Drop the top 1,000 examples to remove uninformative verbatim or near-identical matches
        Retain the remaining 10,000 examples as DcleanD_{clean}
        Compute length-normalized log-likelihood LMScore(X′,Y′)\text{LMScore}(X', Y') under GLM-130B for all (X′,Y′)∈Dclean(X', Y') \in D_{clean}
        Select top 1,000 examples with highest LMScore\text{LMScore} as DlmD_{lm}
        Rank DlmD_{lm} by target sequence token length TLength(Y′)\text{TLength}(Y')
        Select top-KK examples with longest TLength\text{TLength}
        Sort selected KK examples in ascending order of length
        return Sorted top-KK examples as DPD^P
    end if

    On the FLORES Wiki Full set, feature-based selection from the high-quality pool achieves 26.7326.73 BLEU / 49.3449.34 COMET in 1-shot (vs. 26.3126.31 / 48.2948.29 for random selection) and 27.3627.36 BLEU / 51.6651.66 COMET in 5-shot (vs. 27.4627.46 / 51.1151.11 for random). On the low-quality pool, the combined strategy achieves 24.9424.94 BLEU / 39.8839.88 COMET (vs. 24.7524.75 / 38.8638.86 for random selection).

  4. Knowl 4 — Impact of Demonstration Example Quantity and Instability in MT In-Context Learning

    empirical result

    In-context learning for MT was evaluated by varying the number of prompt demonstration examples K∈{1,5,10,20}K \in \{1, 5, 10, 20\} using GLM-130B across language pairs (En-De, En-Zh, De-Zh) on the FLORES Wiki Ablation set.

    1. Scaling KK: On average, increasing the number of demonstration examples from K=1K=1 to K=20K=20 yields progressive improvements in BLEU and COMET. For instance, on Wiki De→En\text{De}\to\text{En}, COMET increases from ∼70.8\sim 70.8 at K=1K=1 to ∼72.8\sim 72.8 at K=20K=20; on En→Zh\text{En}\to\text{Zh}, COMET increases from ∼54\sim 54 to ∼63\sim 63.
    2. Computational Cost: Inference latency per generated token scales linearly with KK. On four A100-40GB GPUs, generating En↔De\text{En}\leftrightarrow\text{De} translations requires ∼0.19\sim 0.19 seconds per token at K=1K=1, increasing to ∼0.29\sim 0.29 seconds per token at K=20K=20.
    3. Performance Instability: Demonstrations exhibit high performance variance across different random samples at the same KK. A suboptimal 1-shot demonstration frequently underperforms the 0-shot baseline on average. In addition, an effective 5-shot demonstration can outperform a 10- or 20-shot demonstration.
    4. Language Disparity: Few-shot prompting provides an exceptionally large boost when translating into Chinese (e.g., COMET rises from ∼35\sim 35 in 0-shot to >60>60 in few-shot) because zero-shot GLM-130B defaults to generating Traditional Chinese with messy corruptions, whereas demonstration examples steer output generation to Simplified Chinese.
  5. Knowl 5 — Failure of Unpaired Monolingual Demonstrations in MT Prompting

    empirical result

    In text classification, prior work found that in-context demonstrations primarily convey the input distribution, label space, and format, such that replacing demonstration labels randomly does not harm performance. However, evaluating unpaired monolingual data for MT prompting with GLM-130B on the FLORES Wiki Ablation benchmark disproves this behavior for translation.

    Three unpaired monolingual prompting settings were evaluated across K∈{1,5,10,20}K \in \{1, 5, 10, 20\}:

    1. Random Example: Randomly pairing unrelated monolingual source sentences with target sentences.
    2. Source Example Only: Using only monolingual source sentences formatted in the prompt.
    3. Target Example Only: Using only monolingual target sentences formatted in the prompt.

    Experimental results:

    • All three unpaired monolingual demonstration methods cause catastrophic performance drops compared to the zero-shot baseline, with COMET scores plunging from positive values (>30>30) to negative values between −50-50 and −150-150.
    • The degradation intensifies monotonically as the number of examples KK increases from 1 to 20.
    • Random pairing misleads the LLM most severely, performing worst overall. Source-only examples perform slightly better than target-only examples, except when translating into Chinese.

    Conclusion: Machine translation in-context learning strictly requires maintaining genuine semantic correspondence between source and target pairs in demonstrations.

  6. Knowl 6 — Pseudo-Parallel Demonstration Construction via Zero-Shot Back-Translation

    model/method

    When parallel bilingual data is unavailable for prompt demonstration selection, pseudo-parallel demonstrations can be constructed from monolingual corpora using zero-shot LLM translation.

    Two data augmentation strategies for prompt demonstration construction were formulated using GLM-130B:

    1. Forward-Translation Augmentation: Monolingual source sentences are translated into the target language using zero-shot prompting with the LLM; the resulting pairs are used as demonstrations.
    2. Back-Translation Augmentation: Monolingual target sentences are translated into the source language using zero-shot prompting with the LLM; the resulting pairs are used as demonstrations.

    Empirical evaluation on the FLORES Wiki Ablation sets (K∈{1,5,10,20}K \in \{1, 5, 10, 20\}) demonstrates:

    • Despite containing zero-shot translation errors, pseudo-parallel demonstrations significantly improve translation over the zero-shot baseline, with translation quality improving as KK increases.
    • Back-translation augmentation consistently outperforms forward-translation augmentation across all language pairs and exhibits greater robustness, closely matching the translation performance achieved with authentic human-parallel demonstrations (e.g., on Zh→En\text{Zh}\to\text{En} at K=20K=20, target back-translation COMET reaches ∼65.8\sim 65.8, approaching the parallel COMET score of ∼66.2\sim 66.2).
  7. Knowl 7 — Transferability of In-Context Demonstrations Across Languages, Domains, and Document Granularities

    empirical result

    Demonstrations selected in one translation setting S1S_1 transfer imperfectly to another setting S2S_2 under 1-shot prompting with GLM-130B.

    1. Cross-Lingual Transfer:
    • Demonstration rankings do not generalize across language pairs; Spearman rank correlations of demonstration performance between different language pairs are weak (average ρ=0.08\rho = 0.08 for source-shared, ρ=0.20\rho = 0.20 for target-shared, and ρ=0.15\rho = 0.15 for reversed translation directions).
    • However, using out-of-setting demonstrations still improves average translation over the zero-shot baseline: target-shared transfer yields ΔCOMET=+9.67\Delta\text{COMET} = +9.67, source-shared yields ΔCOMET=+7.03\Delta\text{COMET} = +7.03, and reversed direction yields ΔCOMET=+11.56\Delta\text{COMET} = +11.56.
    1. Cross-Domain Transfer:
    • Transferring demonstrations from Wiki to Multi-Domain (IT, Medical, Law) shows weak Spearman rank correlations (ρ∈[0.07,0.27]\rho \in [0.07, 0.27] in COMET).
    • Absolute gains over zero-shot are substantial when the target domain is distant and in-domain parallel data is noisy: transferring Wiki demonstrations to IT and Medical yields COMET improvements of +19.52+19.52 and +7.80+7.80 on En→De\text{En}\to\text{De}, and +19.46+19.46 and +1.24+1.24 on De→En\text{De}\to\text{En}.
    1. Sentence-to-Document Transfer:
    • Applying sentence-level demonstrations selected via SemScore\text{SemScore} or LMScore\text{LMScore} to document-level translation on the PDC Zh→En\text{Zh}\to\text{En} dataset improves document-level BLEU (d-BLEU from 30.230.2 to 30.530.5) and discourse-specific metrics (Tense Consistency TC from 47.547.5 to 53.053.0; Pronoun Translation PT from 41.641.6 to 43.243.2; Discourse Connective TCP from 42.442.4 to 42.942.9).
  8. Knowl 8 — English Pivoting for Direct Non-English Translation in Prompted LLMs

    empirical result

    Pretrained bilingual LLMs such as GLM-130B (pretrained primarily on English and Chinese monolingual corpora) exhibit severe quality deficits when prompted directly on non-English translation pairs such as German-Chinese (De↔Zh\text{De}\leftrightarrow\text{Zh}).

    Setting 0-shot COMET 1-shot COMET
    De→Zh\text{De}\to\text{Zh} Zh→De\text{Zh}\to\text{De} De→Zh\text{De}\to\text{Zh} Zh→De\text{Zh}\to\text{De}
    Direct Prompting 2.80 10.05 47.23 11.75
    Pivoting via English 19.23 19.53 48.25 25.31

    Pivoting translation decomposes the non-English translation path into two prompted stages: source→English→target\text{source} \to \text{English} \to \text{target}. On the FLORES Wiki Full set:

    • In zero-shot translation, English pivoting improves De→Zh\text{De}\to\text{Zh} COMET from 2.802.80 to 19.2319.23 (+16.43+16.43) and Zh→De\text{Zh}\to\text{De} COMET from 10.0510.05 to 19.5319.53 (+9.48+9.48).
    • In 1-shot translation, English pivoting improves Zh→De\text{Zh}\to\text{De} COMET from 11.7511.75 to 25.3125.31 (+13.56+13.56).
    • These results demonstrate that cross-lingual capabilities in current LLMs center heavily on English, and direct non-English translation without English pivoting remains poor without explicit multilingual parallel data in pretraining or finetuning.
  9. Knowl 9 — Prompt Trap Vulnerability and Translation Pathologies in MT Prompting

    definition

    A prompt trap is a failure mode in prompt-based machine translation where the input text to be translated contains phrases, formatting, or delimiters that resemble the surrounding prompt template instructions (e.g., Translate from English to Chinese: or English: ... Chinese:). Instead of translating the entire source content including the meta-phrases, the language model misidentifies the prompt phrase inside the input as a task instruction or control token, resulting in copying the untranslated prompt phrase or truncating translation.

    In addition to prompt traps, prompted raw large language models (such as GLM-130B) display several characteristic generation errors during MT:

    1. Source copying and code-switching: Verbatim reproduction of source-language spans or named entities into the target hypothesis without translation.
    2. Entity and numerical distortion: Hallucinatory alteration of numbers, years, and dates (e.g., altering the year 2019 to 2109 or 2018 to 2808).
    3. Rejection and empty generation: Emitting empty outputs or repeating formatting tokens instead of translating.
    4. Off-target language generation: Generating in an unintended script or dialect (e.g., emitting Traditional Chinese characters with byte/character corruption when Simplified Chinese is requested).

Coverage note — All core contributions from the paper—including template selection, demonstration feature analysis, selection algorithms, example scaling/variance, monolingual data ablation, pseudo-parallel generation via back-translation, transfer learning, English pivoting, and prompt traps—are fully represented. Minor appendix per-language breakdown tables were omitted in favor of the summarized findings and full set evaluations.

References

  1. 1.Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2022. In-context examples selection for machine translation. arXiv preprint arXiv:2212.02437.
  2. 2.Roee Aharoni and Yoav Goldberg. 2020. Unsupervised domain clusters in pretrained language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7747–7763, Online. Association for Computational Linguistics.
  3. 3.Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. 2021. Findings of the 2021 conference on machine translation (wmt21). In Proceedings of the Sixth Conference on Machine Translation, pages 1–88, Online. Association for Computational Linguistics.
  4. 4.Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019. Massively multilingual neural machine translation in the wild: Findings and challenges.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  6. 6.Isaac Caswell, Ciprian Chelba, and David Grangier. 2019. Tagged back-translation. In Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers), pages 53–63, Florence, Italy. Association for Computational Linguistics.
  7. 7.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  8. 8.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  10. 10.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  11. 11.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  12. 12.Xavier Garcia and Orhan Firat. 2022. Using natural language prompts for machine translation. arXiv preprint arXiv:2202.11822.
  13. 13.Tanya Goyal, Junyi Jessy Li, and Greg Durrett. 2022. News summarization and evaluation in the era of gpt-3. arXiv preprint arXiv:2209.12356.
  14. 14.Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor O.K. Li. 2018. Search engine guided neural machine translation. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’18/IAAI’18/EAAI’18. AAAI Press.
  15. 15.Kevin Heffernan, Onur Çelebi, and Holger Schwenk. 2022. Bitext mining using distilled sentence representations for low-resource languages. arXiv preprint arXiv:2205.12654.
  16. 16.Melvin Johnson, Mike Schuster, Quoc Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017. Google’s multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5(0):339–351.
  17. 17.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  18. 18.Yafu Li, Yongjing Yin, Jing Li, and Yue Zhang. 2022. Prompt-driven neural machine translation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2579–2590, Dublin, Ireland. Association for Computational Linguistics.
  19. 19.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668.
  20. 20.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  21. 21.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  22. 22.Nikita Moghe, Tom Sherborne, Mark Steedman, and Alexandra Birch. 2022. Extrinsic evaluation of machine translation metrics. arXiv preprint arXiv:2212.10297.
  23. 23.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  24. 24.Matt Post. 2018. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Belgium, Brussels. Association for Computational Linguistics.
  25. 25.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  26. 26.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021. WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1351–1361, Online. Association for Computational Linguistics.
  27. 27.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a. Controlling politeness in neural machine translation via side constraints. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 35–40, San Diego, California. Association for Computational Linguistics.
  28. 28.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016b. Improving neural machine translation models with monolingual data. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 86–96, Berlin, Germany. Association for Computational Linguistics.
  29. 29.Raphael Shu, Hideki Nakayama, and Kyunghyun Cho. 2019. Generating diverse translations with sentence codes. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1823–1827, Florence, Italy. Association for Computational Linguistics.
  30. 30.Taylor Sorensen, Joshua Robinson, Christopher Rytting, Alexander Shaw, Kyle Rogers, Alexia Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate. 2022. An information-theoretic approach to prompt engineering without ground truth labels. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 819–862, Dublin, Ireland. Association for Computational Linguistics.
  31. 31.Zewei Sun, Mingxuan Wang, Hao Zhou, Chengqi Zhao, Shujian Huang, Jiajun Chen, and Lei Li. 2022. Rethinking document-level neural machine translation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3537–3548, Dublin, Ireland. Association for Computational Linguistics.
  32. 32.Zhixing Tan, Xiangwen Zhang, Shuo Wang, and Yang Liu. 2021. Msp: Multi-stage prompting for making pre-trained language models better translators. arXiv preprint arXiv:2110.06609.
  33. 33.David Vilar, Markus Freitag, Colin Cherry, Jiaming Luo, Viresh Ratnakar, and George Foster. 2022. Prompting palm for translation: Assessing strategies and performance. arXiv preprint arXiv:2211.09102.
  34. 34.Chengyu Wang, Jianing Wang, Minghui Qiu, Jun Huang, and Ming Gao. 2021. TransPrompt: Towards an automatic transferable prompting framework for few-shot text classification. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2792–2802, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  35. 35.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022a. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  36. 36.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022b. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  37. 37.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022c. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  38. 38.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  39. 39.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  40. 40.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020. Improving massively multilingual neural machine translation and zero-shot translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1628–1639, Online. Association for Computational Linguistics.
  41. 41.Jiajun Zhang and Chengqing Zong. 2016. Exploiting source-side monolingual data in neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1535–1545.
  42. 42.Jingyi Zhang, Masao Utiyama, Eiichro Sumita, Graham Neubig, and Satoshi Nakamura. 2018. Guiding neural machine translation with retrieved translation pieces. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1325–1335, New Orleans, Louisiana. Association for Computational Linguistics.
  43. 43.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022a. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  44. 44.Yiming Zhang, Shi Feng, and Chenhao Tan. 2022b. Active example selection for in-context learning. arXiv preprint arXiv:2211.04486.
  45. 45.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.
  46. 46.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.

Citation

MLA
Zhang, B., et al. “Prompting Large Language Model for Machine Translation: A Case Study”. arXiv, 2023, http://arxiv.org/abs/2301.07069v2.
APA
Zhang, B., Haddow, B., & Birch, A. (2023). Prompting Large Language Model for Machine Translation: A Case Study. arXiv. http://arxiv.org/abs/2301.07069v2
Chicago
Zhang, B., B. Haddow, and A. Birch. 2023. “Prompting Large Language Model for Machine Translation: A Case Study”. arXiv. http://arxiv.org/abs/2301.07069v2.
Harvard
Zhang, B., Haddow, B. and Birch, A. (2023) “Prompting Large Language Model for Machine Translation: A Case Study”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.07069v2.
Vancouver
1. Zhang B, Haddow B, Birch A (2023) Prompting Large Language Model for Machine Translation: A Case Study. arXiv

BibTeX

@article{zhang2023prompting,
  title = {Prompting Large Language Model for Machine Translation: A Case Study},
  author = {Zhang, Biao and Haddow, Barry and Birch, Alexandra},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.07069v2},
  eprint = {2301.07069}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/