A Contrastive Framework for Neural Text Generation
Yixuan SuTian LanYan WangDani YogatamaLingpeng KongNigel Collier
Proposes SimCTG and contrastive search to eliminate repetitive text degeneration by calibrating anisotropic token representations, enabling neural language models to generate coherent and diverse text that surpasses standard decoding methods.
Neural text generation models are critical for applications such as dialogue systems, contextual text completion, and automated story generation. However, standard training and decoding techniques frequently produce degenerate text plagued by unnatural, repetitive phrasing. Existing remedies introduce stochastic sampling or adjust token probabilities, but these often compromise semantic coherence and cause the generated text to diverge from the original context. The article addresses this fundamental limitation by tracing model degeneration to the narrow, highly similar nature of internal token representations and proposing a contrastive framework to solve it.
The article demonstrates that learning a more discriminative and isotropic representation space combined with an adaptive decoding strategy produces superior text quality. To achieve this, the authors introduce SimCTG, a contrastive training objective that separates distinct token representations, paired with contrastive search, a decoding method that balances model prediction confidence against a degeneration penalty based on context similarity. The authors evaluate this framework using GPT-2 architectures across document generation on the Wikitext-103 benchmark and multi-turn open-domain dialogue generation in both English (DailyDialog) and Chinese (LCCC), supporting their findings with automated metrics and extensive human evaluations comprising 9,000 annotated samples.
The key findings show that the proposed contrastive approach significantly improves generation performance across multiple criteria. In document generation, SimCTG with contrastive search achieved the highest diversity scores and the lowest repetitive phrasing, reducing two-gram repetitions from around 70% seen in deterministic searches to less than 4%, while attaining a superior semantic coherence score above 0.60. Human evaluations demonstrated that contrastive search combined with SimCTG statistically outperformed state-of-the-art nucleus sampling in both coherence and fluency across benchmarks, with the larger model configuration performing comparably to human-written text in fluency. Additionally, the decoding latency of contrastive search remained comparable to, and at higher beam widths faster than, traditional beam search, demonstrating practical computational efficiency.
These results indicate that organizations deploying language models can achieve higher semantic consistency, lower repetition, and human-level fluency without incurring prohibitive computational latency. The framework eliminates the traditional trade-off between repetitive loops and out-of-context hallucinations. Notably, the findings also revealed that for certain languages like Chinese, standard models already exhibit sufficient representation diversity, allowing contrastive search decoding to be applied directly off-the-shelf without additional model retraining.
Based on these findings, teams maintaining text generation systems should consider integrating contrastive search as a drop-in replacement for conventional sampling and beam search methods. For English and similar languages, adopting contrastive training objectives is recommended to fully unlock the benefits of the decoding strategy. Future initiatives should pilot this approach on very large-scale foundation models exceeding billions of parameters. While the findings provide high confidence across the tested architectures, users should exercise caution given that evaluations were limited to models up to the size of GPT-2-large and experiments were not repeated across multiple random seeds due to compute constraints.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). This paper first identifies and analyzes the phenomenon of neural text degeneration and establishes nucleus sampling as the standard stochastic decoding baseline evaluated against in the contrastive framework.
- Paper: How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, Kawin Ethayarajh (2019). This study demonstrates that contextualized representations in models like GPT-2 are anisotropic and occupy narrow cones, providing the foundational representation problem that SimCTG aims to rectify.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). This work introduces a foundational contrastive learning objective for NLP representation spaces that directly inspires the token-level contrastive training objective used in SimCTG.
- Paper: DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation, Yizhe Zhang et al. (2019). This research provides the foundational autoregressive conversational modeling paradigm and benchmarks that motivate the open-domain dialogue evaluations in the target work.
- Paper: Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation, Jin Xu et al. (2022). This work directly continues the investigation into sentence-level repetition loops in neural text generation by analyzing the self-reinforcement copying mechanism and proposing training-time penalization.
- Paper: Controlled Text Generation with Natural Language Instructions, Wangchunshu Zhou et al. (2023). This paper extends controlled and non-degenerate text generation by moving from search-time penalization to natural language instruction fine-tuning.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). This paper extends decoding-time search strategies for generative language models by integrating heuristic lookaheads and multi-faceted preference scoring.
- Paper: Aligning Large Language Models with Representation Editing: A Control Perspective, Lingkai Kong et al. (2024). This study builds on representation-level modifications during text generation by introducing dynamical control theory adjustments to intermediate states.
- Paper: Stay on Topic with Classifier-Free Guidance, Guillaume Sanchez et al. (2024). This research explores an alternative inference-time contrastive mechanism by applying classifier-free guidance to autoregressive language model distributions.
