A Contrastive Framework for Neural Text Generation

Yixuan SuTian LanYan WangDani YogatamaLingpeng KongNigel Collier

article2022NeurIPS308 citations

Proposes SimCTG and contrastive search to eliminate repetitive text degeneration by calibrating anisotropic token representations, enabling neural language models to generate coherent and diverse text that surpasses standard decoding methods.

Listen

Neural text generation models are critical for applications such as dialogue systems, contextual text completion, and automated story generation. However, standard training and decoding techniques frequently produce degenerate text plagued by unnatural, repetitive phrasing. Existing remedies introduce stochastic sampling or adjust token probabilities, but these often compromise semantic coherence and cause the generated text to diverge from the original context. The article addresses this fundamental limitation by tracing model degeneration to the narrow, highly similar nature of internal token representations and proposing a contrastive framework to solve it.

The article demonstrates that learning a more discriminative and isotropic representation space combined with an adaptive decoding strategy produces superior text quality. To achieve this, the authors introduce SimCTG, a contrastive training objective that separates distinct token representations, paired with contrastive search, a decoding method that balances model prediction confidence against a degeneration penalty based on context similarity. The authors evaluate this framework using GPT-2 architectures across document generation on the Wikitext-103 benchmark and multi-turn open-domain dialogue generation in both English (DailyDialog) and Chinese (LCCC), supporting their findings with automated metrics and extensive human evaluations comprising 9,000 annotated samples.

The key findings show that the proposed contrastive approach significantly improves generation performance across multiple criteria. In document generation, SimCTG with contrastive search achieved the highest diversity scores and the lowest repetitive phrasing, reducing two-gram repetitions from around 70% seen in deterministic searches to less than 4%, while attaining a superior semantic coherence score above 0.60. Human evaluations demonstrated that contrastive search combined with SimCTG statistically outperformed state-of-the-art nucleus sampling in both coherence and fluency across benchmarks, with the larger model configuration performing comparably to human-written text in fluency. Additionally, the decoding latency of contrastive search remained comparable to, and at higher beam widths faster than, traditional beam search, demonstrating practical computational efficiency.

These results indicate that organizations deploying language models can achieve higher semantic consistency, lower repetition, and human-level fluency without incurring prohibitive computational latency. The framework eliminates the traditional trade-off between repetitive loops and out-of-context hallucinations. Notably, the findings also revealed that for certain languages like Chinese, standard models already exhibit sufficient representation diversity, allowing contrastive search decoding to be applied directly off-the-shelf without additional model retraining.

Based on these findings, teams maintaining text generation systems should consider integrating contrastive search as a drop-in replacement for conventional sampling and beam search methods. For English and similar languages, adopting contrastive training objectives is recommended to fully unlock the benefits of the decoding strategy. Future initiatives should pilot this approach on very large-scale foundation models exceeding billions of parameters. While the findings provide high confidence across the tested architectures, users should exercise caution given that evaluations were limited to models up to the size of GPT-2-large and experiments were not repeated across multiple random seeds due to compute constraints.

Cover for A Contrastive Framework for Neural Text Generation

Abstract

Text generation is of great importance to many natural language processing applications. However, maximization-based decoding methods (e.g. beam search) of neural language models often lead to degenerate solutions -- the generated text is unnatural and contains undesirable repetitions. Existing approaches introduce stochasticity via sampling or modify training objectives to decrease probabilities of certain tokens (e.g., unlikelihood training). However, they often lead to solutions that lack coherence. In this work, we show that an underlying reason for model degeneration is the anisotropic distribution of token representations. We present a contrastive solution: (i) SimCTG, a contrastive training objective to calibrate the model's representation space, and (ii) a decoding method -- contrastive search -- to encourage diversity while maintaining coherence in the generated text. Extensive experiments and analyses on three benchmarks from two languages demonstrate that our proposed approach significantly outperforms current state-of-the-art text generation methods as evaluated by both human and automatic metrics.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Language Modelling
  • 2.2 Open-ended Text Generation
  • 3 Methodology
  • 3.1 Contrastive Training
  • 3.2 Contrastive Search
  • 4 Document Generation
  • 4.1 Evaluation Metrics
  • 4.1.1 Language Modelling Quality
  • 4.1.2 Generation Quality
  • 4.2 Results
  • 4.3 Human Evaluation
  • 5 Open-domain Dialogue Generation
  • 6 Further Analysis
  • 6.1 Token Representation Self-similarity
  • 6.2 The Effect of Contrastive Loss Margin
  • 6.3 Contrastive Search versus Nucleus Sampling
  • 6.4 Decoding Latency Comparison
  • 6.5 Case Study
  • 6.6 Comparison of Token Similarity Matrix
  • 7 Conclusion
  • References
  • Checklist
  • A Future Work
  • B Related Work
  • C Software Package
  • D Experiments on Different Language Models
  • E Ablation Study on the Hyperparameters of Contrastive Search
  • F Gen-ppl Results Measured by Different Models
  • G Human Evaluation Guidelines
  • G.1 Coherence
  • G.2 Fluency
  • G.3 Informativeness
  • H Self-similarity of Chinese Language Models
  • I Training Efficiency Comparison
  • J Generated Examples on Open-domain Dialogue Generation
  • K More Generated Examples of SimCTG + Contrastive Search
  • L Diverse Contrastive Search

Knowls

  1. Knowl 1 — SimCTG Contrastive Training Objective

    model/method

    SimCTG (Simple Contrastive Framework for Neural Text Generation) is a training framework designed to mitigate representation anisotropy in autoregressive neural language models by enforcing token-level discriminability. Given an input token sequence x={x1,x2,…,x∣x∣}x = \{x_1, x_2, \dots, x_{|x|}\} and language model parameters θ\theta, the SimCTG training objective LSimCTG\mathcal{L}_{\text{SimCTG}} augments the standard Maximum Likelihood Estimation (MLE) loss LMLE\mathcal{L}_{\text{MLE}} with a token-level contrastive loss LCL\mathcal{L}_{\text{CL}}:

    LSimCTG=LMLE+LCL\mathcal{L}_{\text{SimCTG}} = \mathcal{L}_{\text{MLE}} + \mathcal{L}_{\text{CL}}

    The components are defined as:

    LMLE=−1∣x∣∑i=1∣x∣log⁡pθ(xi∣x<i)\mathcal{L}_{\text{MLE}} = - \frac{1}{|x|} \sum_{i=1}^{|x|} \log p_\theta(x_i \mid x_{<i})

    LCL=1∣x∣(∣x∣−1)∑i=1∣x∣∑j=1,j≠i∣x∣max⁡{0,ρ−s(hxi,hxi)+s(hxi,hxj)}\mathcal{L}_{\text{CL}} = \frac{1}{|x|(|x|-1)} \sum_{i=1}^{|x|} \sum_{j=1, j \neq i}^{|x|} \max\{0, \rho - s(h_{x_i}, h_{x_i}) + s(h_{x_i}, h_{x_j})\}

    where hxi∈Rdh_{x_i} \in \mathbb{R}^d is the hidden representation of token xix_i taken from the output layer of the Transformer model, ρ∈[−1,1]\rho \in [-1, 1] is a predefined similarity margin hyperparameter, and s(u,v)s(u, v) is the cosine similarity between two vectors:

    s(u,v)=u⊤v∥u∥∥v∥s(u, v) = \frac{u^\top v}{\|u\| \|v\|}

    Because s(hxi,hxi)=1.0s(h_{x_i}, h_{x_i}) = 1.0, the contrastive penalty for token pair (xi,xj)(x_i, x_j) activates whenever the cosine similarity of distinct tokens within the sequence exceeds 1−ρ1 - \rho. When ρ=0\rho = 0, LCL\mathcal{L}_{\text{CL}} reduces to zero, recovering standard MLE training.

  2. Knowl 2 — Contrastive Search Decoding Algorithm

    algorithm

    Contrastive search is an autoregressive decoding algorithm that selects tokens by balancing model predictive probability with a degeneration penalty computed from previous contextual representations.

    At decoding step tt, given the generated context x<t={x1,…,xt−1}x_{<t} = \{x_1, \dots, x_{t-1}\}, the algorithm identifies the top-kk most probable next-token candidates V(k)V^{(k)} from the language model's predictive distribution pθ(⋅∣x<t)p_\theta(\cdot \mid x_{<t}). For each candidate token v∈V(k)v \in V^{(k)}, the model computes its contextualized representation hvh_v by passing the sequence concatenation [x<t,v][x_{<t}, v] through the model. The output token xtx_t is selected as:

    xt=arg⁡max⁡v∈V(k){(1−α)⋅pθ(v∣x<t)−α⋅max⁡1≤j≤t−1s(hv,hxj)}x_t = \arg\max_{v \in V^{(k)}} \left\{ (1 - \alpha) \cdot p_\theta(v \mid x_{<t}) - \alpha \cdot \max_{1 \le j \le t-1} s(h_v, h_{x_j}) \right\}

    where s(hv,hxj)=hv⊤hxj∥hv∥∥hxj∥s(h_v, h_{x_j}) = \frac{h_v^\top h_{x_j}}{\|h_v\| \|h_{x_j}\|} is the cosine similarity between candidate vv and context token xjx_j, and α∈[0,1]\alpha \in [0, 1] is a hyperparameter governing the weight of the degeneration penalty. When α=0\alpha = 0, the algorithm reduces to greedy search.

    Input: Language model pθp_\theta, context x<tx_{<t}, candidate pool size kk, penalty weight α\alpha, continuation length TT
    Output: Continuation sequence xt:t+T−1x_{t:t+T-1}
    for step = tt to t+T−1t+T-1 do
        V(k)←V^{(k)} \leftarrow top-kk candidate tokens from distribution pθ(⋅∣x<step)p_\theta(\cdot \mid x_{<\text{step}})
        for each candidate token v∈V(k)v \in V^{(k)} do
            hv←h_v \leftarrow output layer representation of candidate vv given sequence [x<step,v][x_{<\text{step}}, v]
            penalty(v)←max⁡1≤j<stephv⊤hxj∥hv∥∥hxj∥\text{penalty}(v) \leftarrow \max_{1 \le j < \text{step}} \frac{h_v^\top h_{x_j}}{\|h_v\| \|h_{x_j}\|}
            score(v)←(1−α)⋅pθ(v∣x<step)−α⋅penalty(v)\text{score}(v) \leftarrow (1 - \alpha) \cdot p_\theta(v \mid x_{<\text{step}}) - \alpha \cdot \text{penalty}(v)
        end for
        xstep←arg⁡max⁡v∈V(k)score(v)x_{\text{step}} \leftarrow \arg\max_{v \in V^{(k)}} \text{score}(v)
        Append xstepx_{\text{step}} to context x<stepx_{<\text{step}}
    end for
    return xt:t+T−1x_{t:t+T-1}
  3. Knowl 3 — Token Representation Self-Similarity Metric

    definition

    Token representation self-similarity measures the degree of anisotropy within a sequence of contextualized token representations produced by a neural language model. For a text sequence x={x1,…,x∣x∣}x = \{x_1, \dots, x_{|x|}\} with corresponding hidden representations {hx1,…,hx∣x∣}\{h_{x_1}, \dots, h_{x_{|x|}}\} produced at a given model layer, the sequence self-similarity is defined as:

    self-similarity(x)=1∣x∣(∣x∣−1)∑i=1∣x∣∑j=1,j≠i∣x∣hxi⊤hxj∥hxi∥∥hxj∥\text{self-similarity}(x) = \frac{1}{|x|(|x|-1)} \sum_{i=1}^{|x|} \sum_{j=1, j \neq i}^{|x|} \frac{h_{x_i}^\top h_{x_j}}{\|h_{x_i}\| \|h_{x_j}\|}

    A higher self-similarity indicates that token representations are clustered in a narrow cone in embedding space (anisotropic distribution), reducing token discriminability and predisposing the model to degenerate repetitive generation. A lower self-similarity indicates an isotropic representation space where distinct tokens remain well-separated.

  4. Knowl 4 — Evaluation of SimCTG and Contrastive Search on Wikitext-103

    data/table

    The table compares language modelling quality and open-ended document generation quality on the Wikitext-103 test set (prefix length 32, continuation length 128) across GPT-2 base (117M parameters) models trained with standard MLE, unlikelihood training, and SimCTG (margin ρ=0.5\rho=0.5). Decoding strategies include greedy search, beam search (size 10), nucleus sampling (p=0.95p=0.95), and contrastive search (k=8,α=0.6k=8, \alpha=0.6). Language modelling metrics include test perplexity (ppl ↓\downarrow), next-token prediction accuracy (acc ↑\uparrow), prefix repetition rate (rep ↓\downarrow), and wrong-token repetition rate (wrep ↓\downarrow). Generation quality metrics include nn-gram repetition (rep-2/3/4 ↓\downarrow), generation diversity (diversity ↑=∏n=24(1−rep-n/100)\uparrow = \prod_{n=2}^4 (1 - \text{rep-n}/100)), distribution closeness to human text (MAUVE ↑\uparrow), semantic coherence via SimCSE (coherence ↑\uparrow), and generation perplexity (gen-ppl).

    Model ppl↓\downarrow acc↑\uparrow rep↓\downarrow wrep↓\downarrow Method rep-2↓\downarrow rep-3↓\downarrow rep-4↓\downarrow diversity↑\uparrow MAUVE↑\uparrow coherence↑\uparrow gen-ppl
    MLE 24.32 39.63 52.82 29.97 greedy 69.21 65.18 62.05 0.04 0.03 0.587 7.32
    beam 71.94 68.97 66.62 0.03 0.03 0.585 6.42
    nucleus 4.45 0.81 0.43 0.94 0.90 0.577 49.71
    contrastive 44.20 37.07 32.44 0.24 0.18 0.599 9.90
    Unlike. 28.57 38.41 51.23 28.57 greedy 24.12 13.35 8.04 0.61 0.69 0.568 37.82
    beam 11.83 5.11 2.86 0.81 0.75 0.524 34.73
    nucleus 4.01 0.80 0.42 0.95 0.87 0.563 72.03
    contrastive 7.48 3.23 1.40 0.88 0.83 0.574 43.61
    SimCTG 23.82 40.91 51.66 28.65 greedy 67.36 63.33 60.17 0.05 0.05 0.596 7.16
    beam 70.32 67.17 64.64 0.04 0.06 0.591 6.36
    nucleus 4.05 0.79 0.37 0.94 0.92 0.584 47.19
    contrastive 3.93 0.78 0.31 0.95 0.94 0.610 18.26
    Human - - 36.19 - - 3.92 0.88 0.28 0.95 1.00 0.644 24.01

    SimCTG + contrastive search achieves the lowest repetition rates, highest generation diversity (0.95), highest MAUVE score (0.94), highest coherence score (0.610), and a generation perplexity (18.26) closest to human text (24.01) among all automated generation methods.

  5. Knowl 5 — Human Evaluation of Open-Ended Document Generation

    data/table

    Human evaluation was conducted on 200 prefixes (length 32) from the Wikitext-103 test set, generating continuations of length 128 across different models (MLE, Unlikelihood, SimCTG base with 117M parameters, and SimCTG-large with 36 layers/20 heads based on GPT-2 large) paired with nucleus sampling (p=0.95p=0.95) and contrastive search (k=8,α=0.6k=8, \alpha=0.6). Five annotators scored 9,000 blind, randomized continuations on a 5-point Likert scale across Coherence, Fluency, and Informativeness. Inter-annotator agreement was evaluated using Fleiss' kappa coefficient.

    Model Decoding Method Coherence Fluency Informativeness
    Agreement (Fleiss' κ\kappa) - 0.51 0.64 0.70
    MLE nucleus 2.92 3.32 3.91
    contrastive 2.78 2.29 2.56
    Unlikelihood nucleus 2.59 3.02 3.58
    contrastive 2.76 2.90 3.35
    SimCTG nucleus 2.96 3.34 3.96
    contrastive 3.25* 3.57* 3.96
    SimCTG-large nucleus 3.01 3.37 3.98
    contrastive 3.33* 3.66* 3.98
    Human - 3.70 3.71 4.21

    Asterisks (*) indicate statistically significant superiority over nucleus sampling with p<0.05p < 0.05 (Sign Test). SimCTG-large + contrastive search achieved fluency (3.66) comparable to human text (3.71, p>0.4p > 0.4).

  6. Knowl 6 — Human Evaluation on Open-Domain Dialogue Benchmarks

    data/table

    The performance of SimCTG and contrastive search (k=5,α=0.6k=5, \alpha=0.6) was evaluated on multi-turn dialogue generation across two languages: Chinese (LCCC dataset, fine-tuned for 40k steps) and English (DailyDialog dataset, fine-tuned for 5k steps). Five annotators evaluated 200 test dialogue contexts on a 5-point Likert scale across Coherence, Fluency, and Informativeness.

    LCCC (Chinese) DailyDialog (English)
    Model Method Coherence Fluency Informativeness Coherence Fluency Informativeness
    Agreement (Fleiss' κ\kappa) - 0.73 0.61 0.57 0.64 0.60 0.55
    MLE greedy 3.01 3.27 1.97 3.28 3.51 2.92
    beam 2.60 2.90 1.55 3.16 3.43 2.78
    nucleus 2.78 3.55 2.64 2.67 3.58 3.42
    contrastive 3.28* 3.84* 3.06* 3.27 3.41 2.82
    SimCTG greedy 3.04 3.32 2.01 3.31 3.50 2.94
    beam 2.57 2.93 1.59 3.19 3.45 2.79
    nucleus 2.84 3.58 2.72 2.75 3.59 3.39
    contrastive 3.32* 3.96* 3.13* 3.73* 3.85* 3.46
    Human - 3.42 3.76 3.20 4.11 3.98 3.74

    Asterisks (*) indicate statistically significant superiority over greedy, beam, and nucleus search (p<0.05p < 0.05 by Sign Test). On the Chinese benchmark (LCCC), SimCTG with contrastive search scored higher in fluency (3.96) than the human reference responses (3.76).

  7. Knowl 7 — Layer-Wise Token Representation Isotropy and Margin Hyperparameter Effects

    empirical result

    Analysis of hidden representations of GPT-2 on the Wikitext-103 test set reveals the following properties of representation isotropy and contrastive margin tuning:

    1. Layer-wise self-similarity: In layers 1 through 9, token representation self-similarity remains roughly stable across models between 0.40 and 0.60. At the final output layer (layer 12), the self-similarity of standard MLE explodes to over 0.95, reflecting severe representation anisotropy. Unlikelihood training reduces output self-similarity to ~0.50 but degrades token accuracy (acc drops from 39.63 to 38.41). SimCTG achieves an output self-similarity of ~0.42 while simultaneously improving next-token prediction accuracy to 40.91.

    2. Effect of contrastive margin ρ\rho: Varying ρ∈[0.1,1.0]\rho \in [0.1, 1.0] in LCL\mathcal{L}_{\text{CL}} shows that contrastive training consistently outperforms vanilla MLE (ppl 24.32) across all values. Perplexity improves as ρ\rho increases from 0.1 to 0.5, reaching optimal performance at ρ=0.5\rho = 0.5 (ppl 23.82). Higher values (ho>0.5 ho > 0.5) force excessive isotropy, leading to sub-optimal perplexity (rising to ~24.10 at ρ=1.0\rho = 1.0).

  8. Knowl 8 — Diversity-Perplexity Trade-Off in Contrastive Search vs Nucleus Sampling

    empirical result

    Evaluating open-ended generation on Wikitext-103 using SimCTG demonstrates a fundamental trade-off disparity between nucleus sampling and contrastive search:

    • Nucleus sampling: Sweeping sampling threshold p∈[0.5,1.0]p \in [0.5, 1.0] shows that at low thresholds (p≤0.7p \le 0.7), generation perplexity matches human level (≈20–25\approx 20\text{--}25), but generation diversity drops severely (<0.80< 0.80), causing repetitive loops. At high thresholds (p≥0.95p \ge 0.95), generation diversity reaches human level (0.94–0.950.94\text{--}0.95), but generation perplexity surges (>45–70> 45\text{--}70), reflecting incoherent and highly improbable token sequences.

    • Contrastive search: Sweeping degeneration penalty weight α∈[0.4,1.0]\alpha \in [0.4, 1.0] with candidate pool size k=8k=8 demonstrates that for α∈[0.5,0.8]\alpha \in [0.5, 0.8], both generation diversity (0.90–0.950.90\text{--}0.95) and generation perplexity (15–2515\text{--}25) remain aligned with human performance (extdiversity=0.95 ext{diversity} = 0.95, gen-ppl=24.01\text{gen-ppl} = 24.01), effectively avoiding both repetitive degeneration and semantic incoherence.

  9. Knowl 9 — Decoding Latency of Contrastive Search

    empirical result

    A comparative evaluation of inference latency across decoding methods using SimCTG on Wikitext-103 (generating fixed-length continuations with batch size 1) demonstrates that:

    1. Greedy search exhibits the lowest latency among all evaluated approaches.
    2. When candidate pool size kk and beam width bb are small (k,b≤6k, b \le 6), the relative decoding latency of contrastive search and beam search is nearly identical.
    3. When k,b>6k, b > 6, contrastive search becomes faster than beam search because contrastive search only computes representations for kk top candidate tokens at each autoregressive step rather than maintaining and expanding multiple parallel beam candidate hypotheses.

Coverage note — None was omitted; all key contributions, algorithms, theoretical models, empirical evaluations, and analyses from the paper are represented.

References

  1. 1.Sourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, and Lav R. Varshney. Mirostat: a neural text decoding algorithm that directly controls perplexity. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  3. 3.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 2978–2988. Association for Computational Linguistics, 2019.
  4. 4.Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander H. Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, Shrimai Prabhumoye, Alan W. Black, Alexander I. Rudnicky, Jason Williams, Joelle Pineau, Mikhail S. Burtsev, and Jason Weston. The second conversational intelligence challenge (convai2). CoRR, abs/1902.00098, 2019.
  5. 5.Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas. Attention is not all you need: pure attention loses rank doubly exponentially with depth. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 2793–2803. PMLR, 2021.
  6. 6.Kawin Ethayarajh. How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT-2 embeddings. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 55–65. Association for Computational Linguistics, 2019.
  7. 7.Angela Fan, Mike Lewis, and Yann N. Dauphin. Hierarchical neural story generation. In Iryna Gurevych and Yusuke Miyao, editors, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 889–898. Association for Computational Linguistics, 2018.
  8. 8.J.L. Fleiss et al. Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5):378–382, 1971.
  9. 9.Tianyu Gao, Xingcheng Yao, and Danqi Chen. Simcse: Simple contrastive learning of sentence embeddings. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6894–6910. Association for Computational Linguistics, 2021.
  10. 10.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020.
  11. 11.Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. Generalization through memorization: Nearest neighbor language models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020.
  12. 12.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015.
  13. 13.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. A diversity-promoting objective function for neural conversation models. In Kevin Knight, Ani Nenkova, and Owen Rambow, editors, NAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12-17, 2016, pages 110–119. The Association for Computational Linguistics, 2016.
  14. 14.Piji Li, Haisong Zhang, Xiaojiang Liu, and Shuming Shi. Rigid formats controlled text generation. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 742–751. Association for Computational Linguistics, 2020.
  15. 15.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. Dailydialog: A manually labelled multi-turn dialogue dataset. In Greg Kondrak and Taro Watanabe, editors, Proceedings of the Eighth International Joint Conference on Natural Language Processing, IJCNLP 2017, Taipei, Taiwan, November 27 - December 1, 2017 - Volume 1: Long Papers, pages 986–995. Asian Federation of Natural Language Processing, 2017.
  16. 16.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017.
  17. 17.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. MAUVE: Measuring the gap between neural text and human text using divergence frontiers. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021.
  18. 18.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  19. 19.Yixuan Su and Nigel Collier. Contrastive search is what you need for neural text generation. arXiv preprint arXiv:2210.14140, 2022.
  20. 20.Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier. Language models can see: Plugging visual controls in text generation. arXiv preprint arXiv:2205.02655, 2022.
  21. 21.Yixuan Su, Fangyu Liu, Zaiqiao Meng, Tian Lan, Lei Shu, Ehsan Shareghi, and Nigel Collier. Tacl: Improving BERT pre-training with token-aware contrastive learning. CoRR, abs/2111.04198, 2021.
  22. 22.Yixuan Su, Yan Wang, Deng Cai, Simon Baker, Anna Korhonen, and Nigel Collier. PROTOTYPE-TO-STYLE: dialogue generation with style-aware editing on retrieval memory. IEEE ACM Trans. Audio Speech Lang. Process., 29:2152–2161, 2021.
  23. 23.Yixuan Su and Jialu Xu. An empirical study on contrastive search and contrastive decoding for open-ended text generation. arXiv preprint arXiv:2211.10797, 2022.
  24. 24.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Kathleen S. Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. Lamda: Language models for dialog applications. CoRR, abs/2201.08239, 2022.
  25. 25.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017.
  26. 26.Yida Wang, Pei Ke, Yinhe Zheng, Kaili Huang, Yong Jiang, Xiaoyan Zhu, and Minlie Huang. A large-scale chinese short-text conversation dataset. In Xiaodan Zhu, Min Zhang, Yu Hong, and Ruifang He, editors, Natural Language Processing and Chinese Computing - 9th CCF International Conference, NLPCC 2020, Zhengzhou, China, October 14-18, 2020, Proceedings, Part I, volume 12430 of Lecture Notes in Computer Science, pages 91–103. Springer, 2020.
  27. 27.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. Neural text generation with unlikelihood training. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020.
  28. 28.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771, 2019.
  29. 29.Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. Adaptive semiparametric language models. Trans. Assoc. Comput. Linguistics, 9:362–373, 2021.
  30. 30.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Asli Celikyilmaz and Tsung-Hsien Wen, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, ACL 2020, Online, July 5-10, 2020, pages 270–278. Association for Computational Linguistics, 2020.
  31. 31.Zhe Zhao, Hui Chen, Jinbin Zhang, Xin Zhao, Tao Liu, Wei Lu, Xi Chen, Haotang Deng, Qi Ju, and Xiaoyong Du. UER: an open-source toolkit for pre-training models. In Sebastian Padó and Ruihong Huang, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019 - System Demonstrations, pages 241–246. Association for Computational Linguistics, 2019.

Citation

MLA
Su, Y., et al. “A Contrastive Framework for Neural Text Generation”. arXiv, 2022, http://arxiv.org/abs/2202.06417v3.
APA
Su, Y., Lan, T., Wang, Y., Yogatama, D., Kong, L., & Collier, N. (2022). A Contrastive Framework for Neural Text Generation. arXiv. http://arxiv.org/abs/2202.06417v3
Chicago
Su, Y., T. Lan, Y. Wang, D. Yogatama, L. Kong, and N. Collier. 2022. “A Contrastive Framework for Neural Text Generation”. arXiv. http://arxiv.org/abs/2202.06417v3.
Harvard
Su, Y. et al. (2022) “A Contrastive Framework for Neural Text Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2202.06417v3.
Vancouver
1. Su Y, Lan T, Wang Y, Yogatama D, Kong L, Collier N (2022) A Contrastive Framework for Neural Text Generation. arXiv

BibTeX

@article{su2022contrastive,
  title = {A Contrastive Framework for Neural Text Generation},
  author = {Su, Yixuan and Lan, Tian and Wang, Yan and Yogatama, Dani and Kong, Lingpeng and Collier, Nigel},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2202.06417v3},
  eprint = {2202.06417}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission