Contrastive Decoding: Open-ended Text Generation as Optimization

Xiang Lisa LiAri HoltzmanDaniel FriedPercy LiangJason EisnerTatsunori HashimotoLuke ZettlemoyerMike Lewis

article2023ACL759 citations

Proposes a training-free decoding strategy that searches for text maximizing the log-likelihood difference between a large expert and a small amateur language model under an adaptive plausibility constraint, outperforming standard sampling methods in coherence and fluency.

Listen

Generating high-quality, open-ended text with large language models remains a core technical challenge. Traditional methods force an unsatisfactory compromise: searching for the most probable text produces dull and repetitive loops, while standard random sampling introduces incoherence, stylistic shifts, and topic drift over long sequences. The article presents contrastive decoding, an optimization-based text generation method designed to reliably produce coherent, fluent, and diverse text without requiring any model retraining or fine-tuning.

The core objective of the article is to demonstrate that contrasting probability predictions between a large language model (termed the expert) and a small language model (termed the amateur) dramatically suppresses undesirable text generation patterns. The approach pairs frozen, pre-trained models—such as contrasting OPT-13B against OPT-125M or GPT-2 XL against GPT-2 small—and uses beam search to choose text that maximizes the difference in their likelihoods. To prevent selecting implausible text, the method applies an adaptive plausibility filter that restricts choices to tokens where the expert model maintains high confidence. The researchers validated the technique across three distinct domains (news, Wikipedia, and stories) using both automated quality metrics and structured human assessments.

The findings show that contrastive decoding significantly outperforms leading decoding techniques, including nucleus sampling and typical decoding. In human evaluations, evaluators preferred the coherence of contrastive decoding text 2.6 times more often than nucleus sampling and 6.4 times more often than typical decoding, while also preferring its fluency 1.4 to 3.5 times more. Automated benchmarks confirmed substantial gains in text coherence and overall distributional quality without sacrificing vocabulary diversity. Ablation analyses revealed that performance peaks when using the widest capacity gap between expert and amateur models, confirming that smaller models act as effective, domain-agnostic proxies for the failure modes present in larger systems.

These results offer significant operational advantages for deploying language models in open-ended settings like creative writing and content generation. Because contrastive decoding operates purely at inference time using frozen models, organizations can achieve substantially higher output reliability and eliminate topic drift without the heavy computational costs of fine-tuning or specialized training. Furthermore, because the amateur model is exceptionally small, it introduces minimal computational overhead during generation.

Decision-makers should consider adopting contrastive decoding for open-ended text applications to reduce post-generation editing and improve factual continuity. When implementing the system, practitioners should pair their primary large models with the smallest available baseline model from the same family. However, stakeholders should note that the article’s empirical successes are limited to open-ended generation; contrastive decoding is currently not suitable for targeted tasks like summarization or machine translation, where smaller models generate high-quality text and penalizing them degrades overall performance. Further research and piloting are required before extending the method beyond open-ended tasks.

Cover for Contrastive Decoding: Open-ended Text Generation as Optimization

Abstract

Given a language model (LM), maximum probability is a poor decoding objective for open-ended generation, because it produces short and repetitive text. On the other hand, sampling can often produce incoherent text that drifts from the original topics. We propose contrastive decoding (CD), a reliable decoding approach that optimizes a contrastive objective subject to a plausibility constraint. The contrastive objective returns the difference between the likelihood under a large LM (called the expert, e.g. OPT-13B) and a small LM (called the amateur, e.g. OPT-125M), and the constraint ensures that the outputs are plausible. CD is inspired by the fact that the failures of larger LMs (e.g., repetition, incoherence) are even more prevalent in smaller LMs, and that this difference signals which texts should be preferred. CD requires zero additional training, and produces higher quality text than decoding from the larger LM alone. It also works across model scales (OPT-13B and GPT2-1.5B) and significantly outperforms four strong decoding algorithms (e.g., nucleus, top-k) in automatic and human evaluations across wikipedia, news and story domains.

Table of Contents

  • 1 Introduction
  • 2 Problem Statement
  • 3 Contrastive Decoding
  • 3.1 Contrastive Objective
  • 3.2 V head: Adaptive Plausibility Constraint
  • 3.3 Full Method
  • 3.4 Choice of Amateur
  • 4 CD as Pragmatic Communication
  • 4.1 Special Cases of Contrastive Decoding
  • 5 Experimental Setup
  • 5.1 Datasets and Metrics
  • 5.2 Baselines
  • 5.3 Models and Hyperparameters
  • 6 Main Results
  • 6.1 Automatic Evaluation
  • 6.2 Human Evaluation
  • 6.3 Qualitative Examples
  • 7 Ablation Studies
  • 7.1 Size of Amateur and Expert LMs
  • 7.2 The Impact of Amateur Temperature
  • 7.3 Sampling v.s. Search
  • 7.4 Plausibility Constraints
  • 7.5 Prompt Inclusion
  • 8 Related Work
  • 9 Conclusion and Future Work
  • Limitations
  • References
  • A CD-Score Analysis
  • B Quantitative Analysis of LM decoding
  • C CD as Distinguishability objective
  • D Additional Related Work
  • E Potential Ethics Risks and Societal Impact
  • F Compute Resources
  • G Human Evaluation Details
  • H Expert and Amateurs from Different model Families
  • I Full Automatic Evaluation Results
  • J Additional Ablation Results
  • K Additional Ablation Results for Sample v.s. Search
  • L More Qualitative Examples
  • M Variant of CD: Training the Amateur LM

Knowls

  1. Knowl 1 — Contrastive Decoding Objective and Token-Level Formulation

    model/method

    Given an input prompt prefix xpre=x1…xnx_{pre} = x_1 \dots x_n of tokens from vocabulary V\mathcal{V} and an autoregressive continuation sequence xcont=xn+1…xn+mx_{cont} = x_{n+1} \dots x_{n+m}, Contrastive Decoding (CD) selects a continuation by maximizing the log-likelihood difference between a large pre-trained expert language model pEXPp_{EXP} and a smaller amateur language model pAMAp_{AMA}:

    max⁡xcontLCD(xcont,xpre)=log⁡pEXP(xcont∣xpre)−log⁡pAMA(xcont∣xpre)\max_{x_{cont}} \mathcal{L}_{CD}(x_{cont}, x_{pre}) = \log p_{EXP}(x_{cont} \mid x_{pre}) - \log p_{AMA}(x_{cont} \mid x_{pre})

    subject to token-level adaptive plausibility constraints xi∈Vhead(x<i)x_i \in V_{head}(x_{<i}) for all xi∈xcontx_i \in x_{cont}, where x<i=x1…xi−1x_{<i} = x_1 \dots x_{i-1}.

    To make sequence generation tractable, the sequence-level objective is factored into step-wise token scores:

    CD-score(xi;x<i)={log⁡pEXP(xi∣x<i)−log⁡pAMA(xi∣x<i),if xi∈Vhead(x<i)−∞,otherwise\text{CD-score}(x_i; x_{<i}) = \begin{cases} \log p_{EXP}(x_i \mid x_{<i}) - \log p_{AMA}(x_i \mid x_{<i}), & \text{if } x_i \in V_{head}(x_{<i}) \\ -\infty, & \text{otherwise} \end{cases}

    This objective penalizes undesirable failure modes that are disproportionately prevalent in smaller amateur language models (such as repetitive loops, vague boilerplate, and semantic drift) while rewarding informative and factual tokens prioritized by the larger expert model.

  2. Knowl 2 — Adaptive Plausibility Constraint for Truncating Implausible Tokens

    definition

    The adaptive plausibility constraint Vhead(x<i)V_{head}(x_{<i}) is a dynamic candidate set that restricts the allowable next token xi∈Vx_i \in \mathcal{V} based on the probability distribution of the expert language model pEXPp_{EXP}:

    Vhead(x<i)={xi∈V:pEXP(xi∣x<i)≥αmax⁡w∈VpEXP(w∣x<i)}V_{head}(x_{<i}) = \left\{ x_i \in \mathcal{V} : p_{EXP}(x_i \mid x_{<i}) \ge \alpha \max_{w \in \mathcal{V}} p_{EXP}(w \mid x_{<i}) \right\}

    where α∈[0,1]\alpha \in [0, 1] is a truncation hyperparameter (set to α=0.1\alpha = 0.1 by default). The constraint resolves two failure modes of an unconstrained contrastive objective:

    1. False positives: Highly implausible tokens with negligible probability under the expert model can achieve an artificially high contrastive score if the amateur model assigns an even smaller probability (for instance, pEXP=3×10−9p_{EXP} = 3 \times 10^{-9} versus pAMA=8×10−14p_{AMA} = 8 \times 10^{-14} yields a contrast score of 10.610.6). VheadV_{head} filters out tokens whose probability is less than α\alpha times the mode probability under pEXPp_{EXP}.

    2. False negatives: On trivial or deterministic decisions (such as subwords completing a single lexical token, e.g., 'unic' →\to 'orn'), both expert and amateur assign probabilities close to 1.01.0, resulting in a contrast near zero. When the expert is highly confident in the top candidate, VheadV_{head} collapses to that single token, bypassing the contrastive penalty and forcing the correct token to be chosen.

  3. Knowl 3 — Contrastive Decoding Beam Search Algorithm

    algorithm

    Contrastive Decoding uses beam search to greedily optimize the cumulative token-level contrastive score subject to the adaptive plausibility constraint at each step.

    Input: Prompt sequence xpre=x1…xnx_{pre} = x_1 \dots x_n, continuation length mm, beam width BB, expert LM pEXPp_{EXP}, amateur LM pAMAp_{AMA}, threshold parameter α\alpha, amateur temperature τ\tau
    Output: Generated continuation xcont=xn+1…xn+mx_{cont} = x_{n+1} \dots x_{n+m}
    Initialize beam pool Beams←{(xpre,0.0)}\text{Beams} \leftarrow \{ (x_{pre}, 0.0) \}
    for step t=1t = 1 to mm do
        Candidates←∅\text{Candidates} \leftarrow \emptyset
        for each (x<n+t,score)(x_{<n+t}, \text{score}) in Beams\text{Beams} do
            Compute expert next-token distribution pEXP(⋅∣x<n+t)p_{EXP}(\cdot \mid x_{<n+t})
            Compute amateur next-token distribution pAMA(⋅∣ctxAMA)p_{AMA}(\cdot \mid \text{ctx}_{AMA}) with temperature τ\tau
            w∗←arg⁡max⁡w∈VpEXP(w∣x<n+t)w^* \leftarrow \arg\max_{w \in \mathcal{V}} p_{EXP}(w \mid x_{<n+t})
            Vhead←{w∈V:pEXP(w∣x<n+t)≥α⋅pEXP(w∗∣x<n+t)}V_{head} \leftarrow \{ w \in \mathcal{V} : p_{EXP}(w \mid x_{<n+t}) \ge \alpha \cdot p_{EXP}(w^* \mid x_{<n+t}) \}
            for each w∈Vheadw \in V_{head} do
                step_score←log⁡pEXP(w∣x<n+t)−log⁡pAMA(w∣ctxAMA)\text{step\_score} \leftarrow \log p_{EXP}(w \mid x_{<n+t}) - \log p_{AMA}(w \mid \text{ctx}_{AMA})
                cand_seq←x<n+t∘w\text{cand\_seq} \leftarrow x_{<n+t} \circ w
                cand_score←score+step_score\text{cand\_score} \leftarrow \text{score} + \text{step\_score}
                Candidates←Candidates∪{(cand_seq,cand_score)}\text{Candidates} \leftarrow \text{Candidates} \cup \{ (\text{cand\_seq}, \text{cand\_score}) \}
            end for
        end for
        Beams←the B sequences in Candidates with highest cumulative scores\text{Beams} \leftarrow \text{the } B \text{ sequences in Candidates with highest cumulative scores}
    end for
    (x∗,final_score)←arg⁡max⁡(x,s)∈Beamss(x^*, \text{final\_score}) \leftarrow \arg\max_{(x, s) \in \text{Beams}} s
    return xn+1:n+m∗x_{n+1:n+m}^*
  4. Knowl 4 — Theoretical Interpretations: Maximizing Distinguishability and Pragmatic Communication

    theoretical result

    Contrastive Decoding can be understood through two theoretical frameworks:

    1. Distinguishability via Pointwise Mutual Information (PMI): Let I∈{0,1}I \in \{0, 1\} be a binary indicator denoting whether a sequence was generated by the expert LM (I=1I = 1, distribution pEXPp_{EXP}) or the amateur LM (I=0I = 0, distribution pAMAp_{AMA}), assuming equal prior probabilities p(I=1)=p(I=0)=0.5p(I=1) = p(I=0) = 0.5. The contrastive decoding objective is mathematically equivalent to maximizing the PMI between the continuation xcontx_{cont} and the indicator variable I=1I = 1:

    PMI(xcont,I=1)=log⁡p(xcont∣I=1)p(xcont)=log⁡pEXP(xcont)0.5pEXP(xcont)+0.5pAMA(xcont)=−log⁡(0.5+0.5pAMA(xcont)pEXP(xcont))\text{PMI}(x_{cont}, I = 1) = \log \frac{p(x_{cont} \mid I = 1)}{p(x_{cont})} = \log \frac{p_{EXP}(x_{cont})}{0.5 p_{EXP}(x_{cont}) + 0.5 p_{AMA}(x_{cont})} = -\log \left(0.5 + 0.5 \frac{p_{AMA}(x_{cont})}{p_{EXP}(x_{cont})}\right)

    Maximizing this quantity selects the text sequence that is most distinguishable as originating from the expert LM rather than the amateur LM.

    1. Pragmatic Speaker-Listener Communication: Under linguistic pragmatics (Gricean maxims, Horn, and Levinson), effective communication balances speaker truthfulness/relevance with listener informativeness. In Contrastive Decoding, pEXPp_{EXP} represents a knowledgeable speaker ensuring relevance and grammatical fluency, while pAMAp_{AMA} represents a less-informed listener; penalizing tokens with high pAMAp_{AMA} suppresses language that is trivially predictable to the listener, ensuring the continuation is non-redundant and informative.
  5. Knowl 5 — Reductions of Contrastive Decoding to Canonical Decoding Strategies

    theoretical result

    Under specific configurations of the amateur distribution pAMAp_{AMA} and amateur context constraints, Contrastive Decoding reduces to several existing decoding algorithms:

    1. Maximum Probability / Beam Search: Setting pAMAp_{AMA} to a uniform distribution over the vocabulary (pAMA(xi)=1∣V∣p_{AMA}(x_i) = \frac{1}{|\mathcal{V}|}) makes log⁡pAMA\log p_{AMA} a constant, reducing the objective strictly to standard log-probability maximization under the expert language model pEXPp_{EXP}.

    2. Soft N-gram Blocking: Setting pAMAp_{AMA} as an nn-gram language model whose nn-gram counts are dynamically updated from the generated prefix yields a decoding objective with soft nn-gram blocking. As the amateur temperature τ→0\tau \to 0, it approaches the canonical rule of forbidding repeated nn-grams.

    3. Maximum Mutual Information (MMI) Decoding: Setting the amateur to be the exact same model as the expert (pAMA=pEXP=pLMp_{AMA} = p_{EXP} = p_{LM}) while restricting the context window of pAMAp_{AMA} to condition only on the final prompt token xnx_n (or no context) equates the objective to log⁡pLM(xcont∣xpre)pLM(xcont)\log \frac{p_{LM}(x_{cont} \mid x_{pre})}{p_{LM}(x_{cont})}, which is the MMI objective used to promote diversity in dialogue systems.

  6. Knowl 6 — Open-Ended Generation Evaluation Benchmark and Metrics

    experimental setup

    Evaluation of open-ended generation is performed across three domains:

    • Wikipedia: WikiText-103 dataset.
    • News: Wikinews dataset.
    • Stories: BookCorpus (Project Gutenberg split).

    Models condition on a 32-word prompt xprex_{pre} and decode a 256-token continuation xcontx_{cont}. Evaluations assess quality along several dimensions:

    • MAUVE: Divergence score in [0,1][0, 1] measuring the distributional alignment between the generated text continuations and gold human references (higher is better).
    • Diversity (DIV): Lexical variation score defined as the product of unique nn-gram ratios: DIV=∏n=24∣unique n-grams(xcont)∣∣total n-grams(xcont)∣\text{DIV} = \prod_{n=2}^4 \frac{|\text{unique } n\text{-grams}(x_{cont})|}{|\text{total } n\text{-grams}(x_{cont})|}
    • Coherence (COH): Cosine similarity between sentence embeddings of the prompt xprex_{pre} and continuation xcontx_{cont} using pre-trained SimCSE embeddings EMB(⋅)\text{EMB}(\cdot): COH(xcont,xpre)=EMB(xpre)⋅EMB(xcont)∥EMB(xpre)∥∥EMB(xcont)∥\text{COH}(x_{cont}, x_{pre}) = \frac{\text{EMB}(x_{pre}) \cdot \text{EMB}(x_{cont})}{\|\text{EMB}(x_{pre})\| \|\text{EMB}(x_{cont})\|}
    • Human Evaluation: Amazon Mechanical Turk annotators perform blinded pairwise comparisons evaluating which continuation is more fluent (grammatical, natural flow) and more coherent (stays on topic with the prompt without drift).
  7. Knowl 7 — Automatic Evaluation Comparison of Contrastive Decoding and Baselines

    data/table

    Automatic evaluation results comparing Contrastive Decoding (CD, beam size 5, α=0.1\alpha=0.1) against standard baselines—greedy search (max prob), top-kk sampling (k=50k=50), nucleus sampling (p=0.95p=0.95), typical decoding (τ=0.95\tau=0.95), and contrastive search (CS)—on continuations of length 256 tokens from OPT-13B (paired with OPT-125M amateur) and GPT-2 XL (1.5B, paired with GPT-2 small amateur).

    Wikinews WikiText-103 Story
    Expert LM Method DIV MAUVE COH DIV MAUVE COH DIV MAUVE COH
    OPT-13B max prob 0.08 0.30 0.65 0.03 0.08 0.63 0.02 0.05 0.51
    OPT-13B k=50k=50 0.91 0.92 0.64 0.72 0.77 0.64 0.91 0.90 0.51
    OPT-13B p=0.95p=0.95 0.92 0.92 0.62 0.92 0.89 0.55 0.93 0.91 0.48
    OPT-13B typical=0.95 0.94 0.90 0.59 0.89 0.86 0.58 0.95 0.91 0.46
    OPT-13B CS 0.92 0.87 0.59 0.87 0.77 0.52 0.81 0.78 0.47
    OPT-13B CD 0.94 0.94 0.69 0.91 0.91 0.69 0.89 0.94 0.62
    GPT-2 XL max prob 0.04 0.14 0.65 0.02 0.05 0.62 0.01 0.03 0.49
    GPT-2 XL k=50k=50 0.92 0.88 0.64 0.87 0.79 0.61 0.91 0.87 0.51
    GPT-2 XL p=0.95p=0.95 0.94 0.90 0.60 0.92 0.87 0.57 0.94 0.91 0.46
    GPT-2 XL typical=0.95 0.95 0.91 0.56 0.95 0.84 0.53 0.96 0.88 0.43
    GPT-2 XL CS 0.93 0.82 0.62 0.86 0.75 0.59 0.88 0.78 0.48
    GPT-2 XL CD 0.92 0.94 0.69 0.89 0.92 0.69 0.83 0.94 0.64

    CD attains the highest MAUVE score and Coherence (COH) across all evaluated domains and model sizes, outperforming all sampling and search algorithms while maintaining lexical diversity substantially higher than greedy search and comparable to stochastic sampling.

  8. Knowl 8 — Human Evaluation Preferences for Coherence and Fluency

    data/table

    Pairwise human evaluation comparing Contrastive Decoding (CD) against Nucleus Sampling (p=0.95p=0.95) and Typical Decoding (τ=0.95\tau=0.95) across three domains (WikiText-103, Wikinews, BookCorpus stories) using GPT-2 XL and OPT-13B.

    Coherence Fluency
    Domain CD Setup Baseline Setup CD Win Same Base Win CD Win Same Base Win
    WikiText CD (GPT-2 XL) Nucleus 0.714 0.083 0.202 0.548 0.083 0.369
    WikiText CD (GPT-2 XL) Typical 0.887 0.046 0.067 0.703 0.082 0.215
    WikiText CD (OPT-13B) Nucleus 0.556 0.202 0.242 0.419 0.197 0.384
    WikiText CD (OPT-13B) Typical 0.773 0.106 0.121 0.687 0.152 0.162
    Wikinews CD (GPT-2 XL) Nucleus 0.708 0.042 0.250 0.583 0.120 0.297
    Wikinews CD (GPT-2 XL) Typical 0.771 0.151 0.078 0.755 0.151 0.094
    Wikinews CD (OPT-13B) Nucleus 0.585 0.221 0.195 0.518 0.123 0.359
    Wikinews CD (OPT-13B) Typical 0.693 0.099 0.208 0.490 0.297 0.214
    Story CD (GPT-2 XL) Nucleus 0.636 0.045 0.318 0.404 0.106 0.490
    Story CD (GPT-2 XL) Typical 0.506 0.256 0.238 0.387 0.363 0.250
    Story CD (OPT-13B) Nucleus 0.616 0.101 0.283 0.449 0.293 0.258
    Story CD (OPT-13B) Typical 0.626 0.202 0.172 0.520 0.212 0.268

    Averaged across domains and models, evaluators preferred Contrastive Decoding over Nucleus sampling by 2.6×2.6\times on coherence and 1.4×1.4\times on fluency. Compared to Typical decoding, CD was preferred by 6.4×6.4\times on coherence and 3.5×3.5\times on fluency.

  9. Knowl 9 — Impact of Amateur-Expert Scale Discrepancy and Model Capacity

    empirical result

    Systematic evaluations pairing expert and amateur language models across different parameter scales (within GPT-2: small, medium, large, XL; within OPT: 125M, 350M, 1.3B, 2.7B, 6.7B, 13B) demonstrate that:

    1. Monotonic Benefit from Scale Gap: Text quality (measured by MAUVE and lexical diversity) improves monotonically as the parameter gap between the expert model and the amateur model widens. The optimal configuration pairs the largest available expert with the smallest amateur in the model family (e.g., OPT-13B with OPT-125M, or GPT-2 XL with GPT-2 small).

    2. Degeneracy on Self-Contrasting and Inverted Scales: Pairing identical models as both expert and amateur yields repetitive output (diversity dropping to 0.41−0.500.41 - 0.50), as no meaningful contrast exists. Using an expert smaller than the amateur produces severely degraded generation quality.

    3. Failure of Low-Capacity N-Gram Amateurs: Using an nn-gram model (e.g., trigram) as the amateur yields poor quality (MAUVE of 0.730.73). Effective contrast requires the amateur to share the inductive biases and failure modes of neural language models, which nn-gram models fail to represent.

    4. Cross-Family Compatibility: CD remains effective when expert and amateur originate from different model families that share a tokenizer. Contrasting GPT-J (6B) as expert against GPT-2 small (117M) as amateur yields MAUVE=0.93\text{MAUVE} = 0.93 and DIV=0.91\text{DIV} = 0.91.

  10. Knowl 10 — Ablations on Plausibility Truncation, Search versus Sampling, and Amateur Temperature

    empirical result

    Ablation experiments on GPT-2 XL (1.5B) identify the individual contributions of the core components of Contrastive Decoding:

    1. Plausibility Truncation (VheadV_{head}): Ablating the adaptive constraint VheadV_{head} leads to severe text collapse. On WikiText-103, omitting VheadV_{head} causes MAUVE to drop from 0.920.92 to 0.010.01, coherence to fall from 0.690.69 to 0.230.23, and model perplexity to explode from 17.7717.77 to 2×1052 \times 10^5, demonstrating that the constraint is essential to eliminate false-positive tokens.

    2. Beam Search vs. Softmax Sampling: Sampling tokens from the softmax-normalized distribution of CD-score(xi;x<i)\text{CD-score}(x_i; x_{<i}) results in lower generation quality than beam search. On WikiText-103, sampling drops MAUVE from 0.920.92 to 0.850.85 and lexical diversity from 0.890.89 to 0.810.81. In human evaluations, beam search is preferred over sampling for coherence (53.5% vs. 42.4% on 1.5B; 46.5% vs. 37.4% on 13B) and fluency (43.4% vs. 23.2% on 1.5B; 47.5% vs. 39.4% on 13B).

    3. Amateur Temperature (τ\tau): Adjusting amateur temperature modifies the penalization of the amateur mode. Setting τ∈[0.5,1.0]\tau \in [0.5, 1.0] robustly yields optimal quality. Increasing τ>1.5\tau > 1.5 flattens the amateur distribution towards uniform and reintroduces repetition, while overly aggressive sharpening (τ=0.1\tau = 0.1) reduces fluency (human preference drops to 34% vs. 44%).

  11. Knowl 11 — Inapplicability of Scale-Contrastive Decoding to Task-Oriented Generation

    limitation

    Contrastive decoding by contrasting smaller and larger models fails to improve performance in task-oriented generation settings such as abstractive summarization and machine translation. In task-oriented tasks, the high-probability modes of both smaller and larger fine-tuned models already represent fluent and relevant outputs, rather than repetitive degeneration. Empirically, using a smaller fine-tuned summarization model (BART-small fine-tuned on summarization data) as the amateur against a larger fine-tuned BART model yields lower ROUGE scores than using a uniform amateur distribution (which corresponds to standard beam search on expert log-probabilities).

Coverage note — None was omitted; all primary contributed concepts, equations, algorithms, empirical results across domains/scales, human evaluation tables, ablations, and stated limitations are fully represented.

References

  1. 1.Chen An, Jiangtao Feng, Kai Lv, Lingpeng Kong, Xipeng Qiu, and Xuanjing Huang. 2022. Cont: Contrastive neural text generation. ArXiv, abs/2205.14690.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  3. 3.Alexandra DeLucia, Aaron Mueller, Xiang Lisa Li, and João Sedoc. 2020. Decoding methods for neural narrative generation. CoRR, abs/2010.07375.
  4. 4.Bryan Eikema and Wilker Aziz. 2020. Is map decoding all you need? the inadequacy of the mode in neural machine translation. In COLING, pages 4506–4520.
  5. 5.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, Melbourne, Australia. Association for Computational Linguistics.
  6. 6.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple contrastive learning of sentence embeddings. In Empirical Methods in Natural Language Processing (EMNLP).
  7. 7.H. Paul Grice. 1975. Logic and conversation. In Peter Cole and Jerry L. Morgan, editors, Speech Acts, volume 3 of Syntax and Semantics.
  8. 8.He He, Nanyun Peng, and Percy Liang. 2019. Pun generation with surprise. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1734–1744, Minneapolis, Minnesota. Association for Computational Linguistics.
  9. 9.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  10. 10.Laurence Horn. 1984. Toward a new taxonomy for pragmatic inference: Q-based and r-based implicature. Meaning, form, and use in context: Linguistic applications, 11:42.
  11. 11.Alex M Lamb, Anirudh Goyal ALIAS PARTH GOYAL, Ying Zhang, Saizheng Zhang, Aaron C Courville, and Yoshua Bengio. 2016. Professor forcing: A new algorithm for training recurrent networks. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc.
  12. 12.Stephen C Levinson. 2000. Presumptive Meanings: The Theory of Generalized Conversational Implicature. MIT Press.
  13. 13.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119, San Diego, California. Association for Computational Linguistics.
  14. 14.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  15. 15.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. DExperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706, Online. Association for Computational Linguistics.
  16. 16.Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1906–1919, Online. Association for Computational Linguistics.
  17. 17.Clara Meister, Tiago Pimentel, Gian Wiher, and Ryan Cotterell. 2022. Typical decoding for natural language generation. CoRR, abs/2202.00666.
  18. 18.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In International Conference on Learning Representations.
  19. 19.Romain Paulus, Caiming Xiong, and Richard Socher. 2018. A deep reinforced model for abstractive summarization. In International Conference on Learning Representations.
  20. 20.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021. MAUVE: Measuring the gap between neural text and human text using divergence frontiers. In Advances in Neural Information Processing Systems.
  21. 21.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. https://openai.com/blog/better-language-models/.
  22. 22.Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016. Sequence level training with recurrent neural networks. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
  23. 23.Yixuan Su and Nigel Collier. 2022. Contrastive search is what you need for neural text generation. arXiv preprint arXiv:2210.14140.
  24. 24.Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022. A contrastive framework for neural text generation. Neurips, abs/2202.06417.
  25. 25.Arun Venkatraman, Martial Hebert, and J.. Bagnell. 2015. Improving multi-step prediction of learned time series models. Proceedings of the AAAI Conference on Artificial Intelligence, 29(1).
  26. 26.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020. Neural text generation with unlikelihood training. In International Conference on Learning Representations.
  27. 27.Sam Wiseman and Alexander M. Rush. 2016. Sequence-to-sequence learning as beam-search optimization. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1296–1306, Austin, Texas. Association for Computational Linguistics.
  28. 28.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016. Google’s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144.
  29. 29.Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In arXiv preprint arXiv:1506.06724.

Citation

MLA
Li, X. L., et al. “Contrastive Decoding: Open-ended Text Generation as Optimization”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12286–312, https://doi.org/10.18653/v1/2023.acl-long.687.
APA
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T. B., Zettlemoyer, L., & Lewis, M. (2023). Contrastive Decoding: Open-ended Text Generation as Optimization. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12286–12312. https://doi.org/10.18653/v1/2023.acl-long.687
Chicago
Li, X. L., A. Holtzman, D. Fried, et al. 2023. “Contrastive Decoding: Open-ended Text Generation as Optimization”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12286–312. https://doi.org/10.18653/v1/2023.acl-long.687.
Harvard
Li, X.L. et al. (2023) “Contrastive Decoding: Open-ended Text Generation as Optimization”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12286–12312. Available at: https://doi.org/10.18653/v1/2023.acl-long.687.
Vancouver
1. Li XL, Holtzman A, Fried D, Liang P, Eisner J, Hashimoto TB, Zettlemoyer L, Lewis M (2023) Contrastive Decoding: Open-ended Text Generation as Optimization. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12286–12312

BibTeX

@inproceedings{li-etal-2023-contrastive,
    title = "Contrastive Decoding: Open-ended Text Generation as Optimization",
    author = "Li, Xiang Lisa  and
      Holtzman, Ari  and
      Fried, Daniel  and
      Liang, Percy  and
      Eisner, Jason  and
      Hashimoto, Tatsunori  and
      Zettlemoyer, Luke  and
      Lewis, Mike",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.687/",
    doi = "10.18653/v1/2023.acl-long.687",
    pages = "12286--12312"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/