Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Haoran XuAmr SharafYunmo ChenWeiting TanLingfeng ShenBenjamin Van DurmeKenton MurrayYoung Jin Kim

article2024ICML436 citations

Proposes Contrastive Preference Optimization, a training method that prevents moderate-sized language models from mimicking imperfect reference translations, enabling a 13B model trained on only 22K sentences to match or exceed the translation performance of GPT-4 and WMT competition winners.

Listen

Moderate-sized large language models (featuring 7B to 13B parameters) have shown substantial potential for machine translation, but they consistently trail behind massive models like GPT-4 and specialized competition-winning translation engines. Traditional supervised fine-tuning trains models by forcing them to replicate gold-standard human references; however, these human-written references frequently contain errors, omissions, or stylistic flaws. Consequently, conventional fine-tuning caps model capabilities at the quality of the reference data and fails to teach systems how to reject subtle, near-perfect translation errors.

The article introduces Contrastive Preference Optimization (CPO), a novel, resource-efficient training method designed to guide models to generate superior translations while explicitly rejecting imperfect candidates. The study evaluates whether training moderate-sized models on preference triplets—combining automated outputs, reference translations, and neural quality ratings—can push language models beyond the limitations of standard supervised fine-tuning.

The authors constructed a compact preference dataset across 10 translation directions using 22,000 paired sentences derived from the FLORES-200 benchmark. Candidate translations generated by GPT-4 and an existing translation model (ALMA-13B-LoRA) were paired with human references and ranked using state-of-the-art reference-free evaluation models (KIWI-XXL and XCOMET). Using these ranked triplets, the authors applied CPO by training only 12 million low-rank adaptation parameters (0.1% of the model's weights) on top of the 13B base model for a single training epoch. The resulting model, ALMA-13B-R, was evaluated across test sets from WMT’21, WMT’22, and WMT’23 alongside human evaluation.

The investigation produced several key findings. First, advanced translation models often produce translations superior to human gold references; reference-free evaluations showed model translations surpassed human references in up to 73% to 79% of English-target instances. Second, ALMA-13B-R matched or exceeded the performance of GPT-4 and specialized WMT competition winners across all evaluated benchmarks, achieving top-tier scores such as 85.74 on KIWI-XXL for English-to-target translations compared to GPT-4's 83.83. Third, CPO fundamentally outperformed traditional fine-tuning and standard Direct Preference Optimization (DPO), which both failed to reliably improve model quality on the same preference data. Fourth, ablation analyses demonstrated that the high quality of rejected examples is critical: using realistic, near-perfect translations as negative examples yielded substantially higher translation quality than using artificially corrupted negative data. Finally, human evaluators confirmed these improvements in blind reviews, preferring ALMA-13B-R over the baseline model in 77.8% of test comparisons.

These findings indicate that organizations do not necessarily need to deploy massive, expensive models or rely on massive datasets to achieve state-of-the-art translation. By updating only a tiny fraction of parameters using contrastive preference learning, organizations can drastically lower computing costs, inference latency, and memory footprints while attaining enterprise-grade quality. Furthermore, the findings demonstrate that traditional reference-based metrics like BLEU are becoming increasingly unreliable for assessing advanced translation systems, as they penalize high-quality, diverse outputs that differ from imperfect gold references.

Organizations developing translation solutions should adopt contrastive preference optimization methods and incorporate high-performing negative examples rather than relying purely on imitation-based fine-tuning. Teams should also transition toward validated reference-free neural evaluation frameworks rather than relying solely on strict lexical overlap metrics like BLEU. While these results demonstrate high statistical confidence across multiple metrics and human evaluations, the current study is limited to 10 language directions centered around English and relied heavily on automated preference labeling. Before deploying this approach across specialized enterprise domains or rare languages, teams should conduct targeted pilot tests to assess performance under domain-specific terminology and low-resource constraints.

arXiv: 2401.08417
Cover for Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Abstract

Moderate-sized large language models (LLMs) -- those with 7B or 13B parameters -- exhibit promising machine translation (MT) performance. However, even the top-performing 13B LLM-based translation models, like ALMA, does not match the performance of state-of-the-art conventional encoder-decoder translation models or larger-scale LLMs such as GPT-4. In this study, we bridge this performance gap. We first assess the shortcomings of supervised fine-tuning for LLMs in the MT task, emphasizing the quality issues present in the reference data, despite being human-generated. Then, in contrast to SFT which mimics reference translations, we introduce Contrastive Preference Optimization (CPO), a novel approach that trains models to avoid generating adequate but not perfect translations. Applying CPO to ALMA models with only 22K parallel sentences and 12M parameters yields significant improvements. The resulting model, called ALMA-R, can match or exceed the performance of the WMT competition winners and GPT-4 on WMT'21, WMT'22 and WMT'23 test datasets.

Table of Contents

  • 1 Introduction
  • 2 Gold or Gilded? Scrutinizing Gold Reference Quality
  • 3 Contrastive Preference Optimization
  • 3.1 Triplet Preference Data
  • 3.2 Deriving the CPO Objective
  • 4 Experiments
  • 4.1 Data
  • 4.2 Training Setup
  • 4.3 Baselines
  • 4.4 WMT’21 and WMT’22 Results
  • 4.5 WMT’23 Results
  • 5 Analyses
  • 5.1 Are Translations Really Better or Just Metric-Preferred?
  • 5.2 Human Evaluation
  • 5.3 Ablation Study
  • 5.4 Does The Quality of Dis-preferred Data Matter?
  • 6 Conclusion
  • References
  • A Comprehensive Results of WMT’21 and WMT’22
  • B Prompts for Translations
  • C Theory
  • C.1 Proof of The Upper Boundary
  • C.2 BC Regularizer Simplification
  • D Details And Influence of Human-Labeled Preference Data
  • D.1 Data Construction Details
  • D.2 Influence on Performance
  • E WMT Winner Systems
  • E.1 Systems For WMT’21 And WMT’22
  • E.2 Systems For WMT’23
  • F Estimated Accuracy with Human Agreements
  • G Full Results of WMT’23
  • H Evaluation on ALMA-R with Non-Comet Metric
  • I The Effectiveness of The BC Regularizer for DPO

Knowls

  1. Knowl 1 — Contrastive Preference Optimization Objective

    equation

    Contrastive Preference Optimization (CPO) optimizes a parameterized language model policy πθ\pi_\theta using labeled preference data D={(x(i),yw(i),yl(i))}i=1N\mathcal{D} = \{(x^{(i)}, y_w^{(i)}, y_l^{(i)})\}_{i=1}^N, where xx represents a source sentence, ywy_w denotes a preferred target translation, and yly_l denotes a dis-preferred target translation. The CPO training objective minimizes the sum of a contrastive preference loss Lprefer(πθ,U)\mathcal{L}_{\text{prefer}}(\pi_\theta, U) based on a uniform prior UU and a negative log-likelihood (NLL) behavior cloning regularizer LNLL(πθ)\mathcal{L}_{\text{NLL}}(\pi_\theta) evaluated on the preferred translations:

    min⁡θLCPO(θ)=min⁡θ(Lprefer(πθ,U)+LNLL(πθ))\min_\theta \mathcal{L}_{\text{CPO}}(\theta) = \min_\theta \left( \mathcal{L}_{\text{prefer}}(\pi_\theta, U) + \mathcal{L}_{\text{NLL}}(\pi_\theta) \right)

    where the individual loss components are defined as:

    Lprefer(πθ,U)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)−βlog⁡πθ(yl∣x))]\mathcal{L}_{\text{prefer}}(\pi_\theta, U) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \pi_\theta(y_w|x) - \beta \log \pi_\theta(y_l|x) \right) \right]

    LNLL(πθ)=−E(x,yw)∼D[log⁡πθ(yw∣x)]\mathcal{L}_{\text{NLL}}(\pi_\theta) = -\mathbb{E}_{(x, y_w) \sim \mathcal{D}} \left[ \log \pi_\theta(y_w|x) \right]

    Here, σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}} is the standard logistic sigmoid function, β>0\beta > 0 is a scalar temperature hyperparameter controlling the preference scale (set by default to 0.10.1), πθ(y∣x)\pi_\theta(y|x) is the sequence-level conditional likelihood of generating translation yy given source xx, and LNLL\mathcal{L}_{\text{NLL}} acts as a behavior cloning constraint derived from a Kullback-Leibler (KL) divergence penalty E(x,yw)∼D[KL(πw(yw∣x)∥πθ(yw∣x))]<ϵ\mathbb{E}_{(x, y_w) \sim \mathcal{D}}[\text{KL}(\pi_w(y_w|x) \parallel \pi_\theta(y_w|x))] < \epsilon with respect to the ideal distribution πw\pi_w of preferred data.

  2. Knowl 2 — Upper Bound on DPO Loss via Uniform Reference Distribution

    theoretical result

    Let πw\pi_w be an ideal reference policy perfectly aligned with the true preferred translation distribution such that for any preference triplet (x,yw,yl)∼D(x, y_w, y_l) \sim \mathcal{D}, the conditions πw(yw∣x)=1\pi_w(y_w|x) = 1 and 0≤πw(yl∣x)≤10 \le \pi_w(y_l|x) \le 1 hold. Under Direct Preference Optimization (DPO), the loss parameterized with reference policy πw\pi_w satisfies:

    L(πθ;πw)+C≤L(πθ;U)\mathcal{L}(\pi_\theta; \pi_w) + C \le \mathcal{L}(\pi_\theta; U)

    where L(πθ;πw)\mathcal{L}(\pi_\theta; \pi_w) is defined as:

    L(πθ;πw)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πw(yw∣x)−βlog⁡πθ(yl∣x)πw(yl∣x))]\mathcal{L}(\pi_\theta; \pi_w) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w|x)}{\pi_w(y_w|x)} - \beta \log \frac{\pi_\theta(y_l|x)}{\pi_w(y_l|x)} \right) \right]

    UU denotes a uniform reference model where probabilities cancel out, L(πθ;U)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)−βlog⁡πθ(yl∣x))]\mathcal{L}(\pi_\theta; U) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \pi_\theta(y_w|x) - \beta \log \pi_\theta(y_l|x) \right) \right], and C=E(x,yl)∼D[log⁡πw(yl∣x)β]C = \mathbb{E}_{(x, y_l) \sim \mathcal{D}} [\log \pi_w(y_l|x)^\beta] is a constant independent of the policy parameters θ\theta. Consequently, minimizing L(πθ;U)\mathcal{L}(\pi_\theta; U) minimizes an upper bound of the ideal DPO loss without requiring a secondary reference model during training.

  3. Knowl 3 — Triplet-Based Preference Dataset Construction for Machine Translation

    model/method

    To construct preference training pairs (x,yw,yl)(x, y_w, y_l) for machine translation without relying solely on human annotations:

    1. For each source sentence xx in parallel training data, three translation candidates are compiled to form a triplet y=(yref,ygpt-4,yalma)y = (y_{\text{ref}}, y_{\text{gpt-4}}, y_{\text{alma}}), where yrefy_{\text{ref}} is the dataset's human gold reference, ygpt-4y_{\text{gpt-4}} is generated zero-shot by GPT-4 (gpt-4-1106-preview), and yalmay_{\text{alma}} is generated by ALMA-13B-LoRA.
    2. Candidate translations are evaluated using two 10B-parameter reference-free metric models: KIWI-XXL (wmt23-cometkiwi-da-xxl) and XCOMET (Unbabel/XCOMET-XXL).
    3. The average score for each candidate i∈{ref,gpt-4,alma}i \in \{\text{ref}, \text{gpt-4}, \text{alma}\} is computed as si=sKIWI-XXL(x,yi)+sXCOMET(x,yi)2s_i = \frac{s_{\text{KIWI-XXL}}(x, y_i) + s_{\text{XCOMET}}(x, y_i)}{2}.
    4. The highest-scoring translation is assigned as the preferred target yw=arg⁡max⁡i(s)y_w = \arg\max_i (s), and the lowest-scoring translation is assigned as the dis-preferred target yl=arg⁡min⁡i(s)y_l = \arg\min_i (s). The candidate with the median score is discarded.

    Applying this protocol to the FLORES-200 development and test splits across 10 translation directions (Czech, German, Icelandic, Chinese, and Russian to/from English) yields a dataset of 2,0092{,}009 paired instances per direction, totaling approximately 20,00020{,}000 training triplets.

  4. Knowl 4 — Quality Discrepancy Between Human Gold References and LLM Translations

    empirical result

    An evaluation of human-authored FLORES-200 gold references against translations produced by GPT-4 and ALMA-13B-LoRA across 5 language pairs (cs, de, is, zh, ru to and from English) using 10B reference-free quality estimation models reveals that system-generated outputs frequently surpass human gold references.

    Translation KIWI-XXL Win Ratio (%) XCOMET Win Ratio (%)
    Translating to English (xx→\rightarrowen)
    Reference 85.31 – 88.82 –
    ALMA-13B-LoRA 88.33 73.24 92.68 60.17
    GPT-4 89.21 79.43 94.66 54.25
    Translating from English (en→\rightarrowxx)
    Reference 87.85 – 94.42 –
    ALMA-13B-LoRA 85.62 42.15 93.07 35.46
    GPT-4 87.30 49.13 94.21 38.09

    In translations into English (xx→\rightarrowen), ALMA-13B-LoRA and GPT-4 outperform the gold references in 73.24%73.24\% and 79.43%79.43\% of sentences according to KIWI-XXL (and 60.17%60.17\% and 54.25%54.25\% under XCOMET), producing average score advantages of 3.02--3.90 points on KIWI-XXL and 3.86--5.84 points on XCOMET. In the en→\rightarrowxx direction, system translations surpass gold references in approximately 35%–49%35\%\text{--}49\% of test cases. Gold references often omit full named entity expansions or context present in system translations.

  5. Knowl 5 — ALMA-R Fine-Tuning Setup and Parameter Efficiency

    experimental setup

    ALMA-R models are developed by fine-tuning ALMA-LoRA checkpoints (ALMA-7B-LoRA and ALMA-13B-LoRA) with Contrastive Preference Optimization (CPO). The fine-tuning procedure specifies:

    • Trainable Parameters: Only newly added Low-Rank Adaptation (LoRA) parameters are updated, using rank r=16r = 16. For the 13B model, this introduces 12M trainable parameters (0.1% of the total model size).
    • Preference Dataset: 20K sentence triplets covering 10 language directions derived from FLORES-200 (cs↔\leftrightarrowen, de↔\leftrightarrowen, is↔\leftrightarrowen, zh↔\leftrightarrowen, ru↔\leftrightarrowen), supplemented with 1K human-annotated pairs for en→\rightarrowde and en→\rightarrowzh.
    • Hyperparameters: Temperature parameter β=0.1\beta = 0.1, batch size of 128 sequences, warm-up ratio of 0.01, trained for 1 single epoch with a maximum sequence length of 512 tokens using DeepSpeed.
    • Loss Masking: Loss is computed exclusively over target translation tokens, omitting the prompt tokens.
  6. Knowl 6 — ALMA-13B-R Performance on WMT Benchmarks

    empirical result

    Fine-tuning ALMA-13B-LoRA with CPO yields ALMA-13B-R, which matches or surpasses GPT-4 (gpt-4-1106-preview) and WMT competition winning systems on WMT'21 (is), WMT'22 (de, cs, zh, ru), and WMT'23 test sets across reference-free metrics KIWI-22, KIWI-XXL, and XCOMET:

    WMT'21/22 en→\rightarrowxx (Avg.) WMT'21/22 xx→\rightarrowen (Avg.)
    Model KIWI-22 KIWI-XXL XCOMET KIWI-22 KIWI-XXL XCOMET
    Gold Reference 82.05 83.47 92.85 79.91 80.10 85.77
    WMT Winners 83.41 84.81 93.78 80.92 81.19 87.13
    GPT-4 82.94 83.83 93.23 81.28 82.60 89.41
    ALMA-13B-LoRA 82.48 82.66 92.76 80.53 81.50 86.74
    + SFT (preferred) 82.57 82.42 92.54 80.96 81.99 88.40
    + DPO 82.27 82.07 92.25 80.51 81.36 86.58
    + CPO (ALMA-13B-R) 83.34 85.74 94.05 81.33 82.43 89.11

    On the WMT'23 benchmark across 6 directions (de↔\leftrightarrowen, zh↔\leftrightarrowen, ru↔\leftrightarrowen), ALMA-13B-R attains average scores of 80.55 (KIWI-22), 78.97 (KIWI-XXL), and 89.74 (XCOMET), surpassing TowerInstruct (80.31 / 77.18 / 88.11), ALMA-13B-LoRA (79.48 / 76.00 / 87.16), and WMT Winners (80.57 / 77.72 / 88.24).

  7. Knowl 7 — Impact of Dis-Preferred Data Quality: Natural vs. Noised Negatives

    empirical result

    Training CPO using naturally generated, high-quality dis-preferred translations (yly_l) outperforms training on synthetic negative examples generated via rule-based corruptions (random word deletion with p=0.15p=0.15 and adjacent word swaps with p=0.3p=0.3 applied to ywy_w):

    Dis-Preferred Data Type KIWI-22 KIWI-XXL XCOMET
    Translating to English (xx→\rightarrowen)
    Manually Noised 81.01 82.18 88.23
    Natural (CPO Triplet) 81.33 82.43 89.11
    Translating from English (en→\rightarrowxx)
    Manually Noised 82.71 83.13 92.80
    Natural (CPO Triplet) 83.34 85.74 94.05

    Using natural dis-preferred translations from strong models that contain minor omissions or subtle flaws teaches the model to discriminate fine-grained translation errors, whereas artificial noisy negatives lead to lower translation quality.

  8. Knowl 8 — Computational and Memory Efficiency of CPO vs. DPO

    empirical result

    Direct Preference Optimization (DPO) requires maintaining two separate model instances in memory (the active policy πθ\pi_\theta and the frozen reference policy πref\pi_{\text{ref}}) and performing sequential forward passes across both models. In contrast, CPO sets πref\pi_{\text{ref}} to a uniform prior UU, eliminating reference model storage and inference entirely:

    Objective KIWI-22 KIWI-XXL XCOMET Memory Cost FLOPs/Token
    Translating to English (xx→\rightarrowen)
    LDPO\mathcal{L}_{\text{DPO}} 80.51 81.36 86.58 2×2\times 2×2\times
    LDPO+LNLL\mathcal{L}_{\text{DPO}} + \mathcal{L}_{\text{NLL}} 81.28 82.42 89.05 2×2\times 2×2\times
    Lprefer+LNLL\mathcal{L}_{\text{prefer}} + \mathcal{L}_{\text{NLL}} (CPO) 81.33 82.43 89.11 1×\mathbf{1\times} 1×\mathbf{1\times}
    Translating from English (en→\rightarrowxx)
    LDPO\mathcal{L}_{\text{DPO}} 82.27 82.07 92.25 2×2\times 2×2\times
    LDPO+LNLL\mathcal{L}_{\text{DPO}} + \mathcal{L}_{\text{NLL}} 83.13 84.74 93.53 2×2\times 2×2\times
    Lprefer+LNLL\mathcal{L}_{\text{prefer}} + \mathcal{L}_{\text{NLL}} (CPO) 83.34 85.74 94.05 1×\mathbf{1\times} 1×\mathbf{1\times}

    CPO halves GPU memory overhead and per-token FLOPs compared to DPO while achieving equal or superior translation performance.

  9. Knowl 9 — Ablation of CPO Loss Terms and Candidate Sources

    empirical result

    Ablation studies on the CPO loss formulation and triplet source composition on WMT'21/22 show that both loss terms and candidate generators are necessary for optimal performance:

    1. Loss Components: Optimizing solely with Lprefer\mathcal{L}_{\text{prefer}} yields average reference-free scores of 82.8182.81 (xx→\rightarrowen) and 85.5085.50 (en→\rightarrowxx). Optimizing solely with LNLL\mathcal{L}_{\text{NLL}} (SFT on preferred targets) yields 83.7883.78 and 85.8485.84. Combining both in CPO (Lprefer+LNLL\mathcal{L}_{\text{prefer}} + \mathcal{L}_{\text{NLL}}) achieves the highest scores: 84.2984.29 (xx→\rightarrowen) and 87.7187.71 (en→\rightarrowxx).
    2. Candidate Sources: Excluding ALMA-13B-LoRA outputs from the triplet selection degrades average en→\rightarrowxx reference-free evaluation from 87.7187.71 to 86.6686.66. Excluding GPT-4 outputs degrades xx→\rightarrowen performance from 84.2984.29 to 83.7083.70, confirming that both model generators contribute complementary preference signals.
  10. Knowl 10 — Human Evaluation and Non-COMET Metric Validation of ALMA-13B-R

    empirical result

    Human evaluation and independent non-COMET automated evaluation confirm that ALMA-13B-R improvements reflect genuine translation quality rather than metric overfitting:

    • Human Evaluation: A blind evaluation by four bilingual evaluators on 400 sampled WMT'22 Chinese→\rightarrowEnglish (zh→\rightarrowen) translations rated on a 0--6 scale demonstrated that ALMA-13B-R scored an average of 5.16 with an average rank of 1.40 and a win ratio of 77.80%, outperforming ALMA-13B-LoRA (average score 4.86, rank 1.60, win ratio 62.50%, with 40.30% ties).
    • BLEURT-20 Evaluation: Across all 5 language pairs evaluated with reference-based BLEURT-20, ALMA-13B-R improved average performance over ALMA-13B-LoRA from 73.96 to 74.79 in xx→\rightarrowen and from 75.02 to 76.04 in en→\rightarrowxx, demonstrating consistent gains on metrics structurally independent from COMET.

Coverage note — Omitted the list of specific WMT winning system names per direction (Appendix E), prompt templates (Appendix B), and intermediate mathematical derivation steps of the BC regularizer simplification.

References

  1. 1.Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G. Falcon-40B: an open large language model with state-of-the-art performance. 2023.
  2. 2.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  3. 3.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pp. 1597–1607. PMLR, 2020.
  4. 4.Chen, Y., Liu, Y., Meng, F., Chen, Y., Xu, J., and Zhou, J. Improving translation faithfulness of large language models via augmenting instructions. arXiv preprint arXiv:2308.12674, 2023.
  5. 5.Fan, A., Bhosale, S., Schwenk, H., Ma, Z., El-Kishky, A., Goyal, S., Baines, M., Celebi, O., Wenzek, G., Chaudhary, V., et al. Beyond english-centric multilingual machine translation. Journal of Machine Learning Research, 22 (107):1–48, 2021.
  6. 6.Freitag, M., Mathur, N., Lo, C.-k., Avramidis, E., Rei, R., Thompson, B., Kocmi, T., Blain, F., Deutsch, D., Stewart, C., Zerva, C., Castilho, S., Lavie, A., and Foster, G. Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent. In Koehn, P., Haddow, B., Kocmi, T., and Monz, C. (eds.), Proceedings of the Eighth Conference on Machine Translation, pp. 578–628, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.wmt-1.51. URL https://aclanthology.org/2023.wmt-1.51.
  7. 7.Guerreiro, N. M., Rei, R., van Stigt, D., Coheur, L., Colombo, P., and Martins, A. F. xcomet: Transparent machine translation evaluation through fine-grained error detection. arXiv preprint arXiv:2310.10482, 2023.
  8. 8.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729–9738, 2020.
  9. 9.Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D. Contrastive prefence learning: Learning from human feedback without rl. arXiv preprint arXiv:2310.13639, 2023.
  10. 10.Hendy, A., Abdelrehim, M., Sharaf, A., Raunak, V., Gabr, M., Matsushita, H., Kim, Y. J., Afify, M., and Awadalla, H. H. How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210, 2023.
  11. 11.Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  12. 12.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  13. 13.Jiao, W., Huang, J.-t., Wang, W., He, Z., Liang, T., Wang, X., Shi, S., and Tu, Z. ParroT: Translating during chat using large language models tuned with human translation and feedback. In Bouamor, H., Pino, J., and Bali, K. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 15009–15020, Singapore, December 2023a. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.1001. URL https://aclanthology.org/2023.findings-emnlp.1001.
  14. 14.Jiao, W., Wang, W., Huang, J.-t., Wang, X., and Tu, Z. Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745, 2023b.
  15. 15.Kocmi, T., Bawden, R., Bojar, O., Dvorkovich, A., Federmann, C., Fishel, M., Gowda, T., Graham, Y., Grundkiewicz, R., Haddow, B., Knowles, R., Koehn, P., Monz, C., Morishita, M., Nagata, M., Nakazawa, T., Novák, M., Popel, M., and Popović, M. Findings of the 2022 conference on machine translation (WMT22). In Koehn, P., Barrault, L., Bojar, O., Bougares, F., Chatterjee, R., Costajussà, M. R., Federmann, C., Fishel, M., Fraser, A., Freitag, M., Graham, Y., Grundkiewicz, R., Guzman, P., Haddow, B., Huck, M., Jimeno Yepes, A., Kocmi, T., Martins, A., Morishita, M., Monz, C., Nagata, M., Nakazawa, T., Negri, M., Nevéol, A., Neves, M., Popel, M., Turchi, M., and Zampieri, M. (eds.), Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 1–45, Abu Dhabi, United Arab Emirates (Hybrid), December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.wmt-1.1.
  16. 16.Kocmi, T., Avramidis, E., Bawden, R., Bojar, O., Dvorkovich, A., Federmann, C., Fishel, M., Freitag, M., Gowda, T., Grundkiewicz, R., Haddow, B., Koehn, P., Marie, B., Monz, C., Morishita, M., Murray, K., Nagata, M., Nakazawa, T., Popel, M., Popović, M., and Shmatova, M. Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet. In Koehn, P., Haddow, B., Kocmi, T., and Monz, C. (eds.), Proceedings of the Eighth Conference on Machine Translation, pp. 1–42, Singapore, December 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.wmt-1.1.
  17. 17.Kocmi, T., Zouhar, V., Federmann, C., and Post, M. Navigating the metrics maze: Reconciling score magnitudes and accuracies. arXiv preprint arXiv:2401.06760, 2024.
  18. 18.Kudugunta, S., Caswell, I., Zhang, B., Garcia, X., Choquette-Choo, C. A., Lee, K., Xin, D., Kusupati, A., Stella, R., Bapna, A., and Firat, O. Madlad-400: A multilingual and document-level large audited dataset, 2023.
  19. 19.Li, J., Zhou, H., Huang, S., Chen, S., and Chen, J. Eliciting the translation ability of large language models via multilingual finetuning with translation instructions. arXiv preprint arXiv:2305.15083, 2023.
  20. 20.Maillard, J., Gao, C., Kalbassi, E., Sadagopan, K. R., Goswami, V., Koehn, P., Fan, A., and Guzman, F. Small data, big impact: Leveraging minimal data for effective machine translation. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2740–2756, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.154. URL https://aclanthology.org/2023.acl-long.154.
  21. 21.NLLB TEAM, Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al. No language left behind: Scaling human-centered machine translation. arXiv preprint arXiv:2207.04672, 2022.
  22. 22.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  23. 23.OpenAI. Gpt-4 technical report, 2023.
  24. 24.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  25. 25.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Isabelle, P., Charniak, E., and Lin, D. (eds.), Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311–318, Philadelphia, Pennsylvania, USA, July 2002. Association for Computational Linguistics. doi: 10.3115/1073083.1073135. URL https://aclanthology.org/P02-1040.
  26. 26.Post, M. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pp. 186–191, Brussels, Belgium, October 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-6319. URL https://aclanthology.org/W18-6319.
  27. 27.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290, 2023.
  28. 28.Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pp. 3505–3506, 2020.
  29. 29.Rei, R., C. de Souza, J. G., Alves, D., Zerva, C., Farinha, A. C., Glushkova, T., Lavie, A., Coheur, L., and Martins, A. F. T. COMET-22: Unbabel-IST 2022 submission for the metrics shared task. In Proceedings of the Seventh Conference on Machine Translation (WMT), pp. 578–585, Abu Dhabi, United Arab Emirates (Hybrid), December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.wmt-1.52.
  30. 30.Rei, R., Guerreiro, N. M., Pombal, J., van Stigt, D., Treviso, M., Coheur, L., de Souza, J. G., and Martins, A. F. Scaling up cometkiwi: Unbabel-ist 2023 submission for the quality estimation shared task. arXiv preprint arXiv:2309.11925, 2023.
  31. 31.Robinson, J. D., Chuang, C.-Y., Sra, S., and Jegelka, S. Contrastive learning with hard negative samples. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=CR1XOQ0UTh-.
  32. 32.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  33. 33.Sellam, T., Das, D., and Parikh, A. BLEURT: Learning robust metrics for text generation. In Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J. (eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 7881–7892, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.704. URL https://aclanthology.org/2020.acl-main.704.
  34. 34.Tan, W., Heffernan, K., Schwenk, H., and Koehn, P. Multilingual representation distillation with contrastive learning. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 1469–1482, 2023.
  35. 35.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023a.
  36. 36.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and finetuned chat models. arXiv preprint arXiv:2307.09288, 2023b.
  37. 37.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  38. 38.Wu, Y. and Hu, G. Exploring prompt engineering with GPT language models for document-level machine translation: Insights and findings. In Koehn, P., Haddow, B., Kocmi, T., and Monz, C. (eds.), Proceedings of the Eighth Conference on Machine Translation, pp. 166–169, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.wmt-1.15. URL https://aclanthology.org/2023.wmt-1.15.
  39. 39.Xu, H., Van Durme, B., and Murray, K. BERT, mBERT, or BiBERT? a study on contextualized embeddings for neural machine translation. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6663–6675, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.534. URL https://aclanthology.org/2021.emnlp-main.534.
  40. 40.Xu, H., Kim, Y. J., Sharaf, A., and Awadalla, H. H. A paradigm shift in machine translation: Boosting translation performance of large language models, 2023.
  41. 41.Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C. mT5: A massively multilingual pre-trained text-to-text transformer. In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 483–498, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.41. URL https://aclanthology.org/2021.naacl-main.41.
  42. 42.Yang, W., Li, C., Zhang, J., and Zong, C. Bigtrans: Augmenting large language models with multilingual translation capability over 100 languages. arXiv preprint arXiv:2305.18098, 2023.
  43. 43.Zeng, J., Meng, F., Yin, Y., and Zhou, J. Tim: Teaching large language models to translate with comparison. arXiv preprint arXiv:2307.04408, 2023.
  44. 44.Zhang, S., Fang, Q., Zhang, Z., Ma, Z., Zhou, Y., Huang, L., Bu, M., Gui, S., Chen, Y., Chen, X., et al. Bayling: Bridging cross-lingual alignment and instruction following through interactive translation for large language models. arXiv preprint arXiv:2306.10968, 2023.
  45. 45.Zhu, W., Liu, H., Dong, Q., Xu, J., Kong, L., Chen, J., Li, L., and Huang, S. Multilingual machine translation with large language models: Empirical results and analysis. arXiv preprint arXiv:2304.04675, 2023a.
  46. 46.Zhu, W., Lv, Y., Dong, Q., Yuan, F., Xu, J., Huang, S., Kong, L., Chen, J., and Li, L. Extrapolating large language models to non-english by aligning languages. arXiv preprint arXiv:2308.04948, 2023b.
  47. 47.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.

Citation

MLA
Xu, H., et al. “Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation”. arXiv, 2024, http://arxiv.org/abs/2401.08417v4.
APA
Xu, H., Sharaf, A., Chen, Y., Tan, W., Shen, L., Durme, B. V., Murray, K., & Kim, Y. J. (2024). Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation. arXiv. http://arxiv.org/abs/2401.08417v4
Chicago
Xu, H., A. Sharaf, Y. Chen, et al. 2024. “Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation”. arXiv. http://arxiv.org/abs/2401.08417v4.
Harvard
Xu, H. et al. (2024) “Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.08417v4.
Vancouver
1. Xu H, Sharaf A, Chen Y, Tan W, Shen L, Durme BV, Murray K, Kim YJ (2024) Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation. arXiv

BibTeX

@article{xu2024contrastive,
  title = {Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation},
  author = {Xu, Haoran and Sharaf, Amr and Chen, Yunmo and Tan, Weiting and Shen, Lingfeng and Durme, Benjamin Van and Murray, Kenton and Kim, Young Jin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.08417v4},
  eprint = {2401.08417}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/