Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation
Jin XuXiaojiang LiuJianhao YanDeng CaiHuayang LiJian Li
Reveals the self-reinforcement mechanism behind repetitive neural text generation loops and introduces DITTO, a training method that penalizes repetition probabilities using pseudo-data without sacrificing text quality or perplexity.
Large-scale neural language models often get trapped in undesirable loops of repeating entire sentences consecutively when generating text, especially under standard greedy or maximum-probability decoding strategies. This behavior contradicts natural human language, where full consecutive sentence repetitions are virtually nonexistent (around 0.02% in benchmark corpora). As natural language generation becomes integral to automated summarization, customer interactions, and content creation, repetitive loops severely degrade text quality, user trust, and overall system utility.
The article investigates the fundamental mechanisms driving these sentence-level repetitive loops and evaluates a novel training framework designed to suppress redundant repetitions while preserving overall language generation quality.
To analyze this phenomenon, the researchers evaluated token probability behavior across 1,000 synthetic test sequences from multiple data sources, including Wikipedia, BookCorpus, and random word sets. They systematically measured probability changes as sentences were repeated up to 100 times. Using these diagnostic findings, the authors designed Pseudo-Repetition Penalization (DITTO), a lightweight fine-tuning strategy that feeds artificially repeated sentences into the model and penalizes repetition probabilities using an exponential decay factor. DITTO was evaluated against existing baselines across open-ended generation tasks on Wikitext-103 using a 750-million-parameter Transformer and abstractive summarization on CNN/DailyMail using a BART-large model, utilizing automated metrics (such as MAUVE and ROUGE) alongside human evaluations.
The investigation produced several key findings regarding neural text generation. First, language models exhibit an inherent copying shortcut: introducing just one sentence-level repetition causes the probability of repeating the next token to rise in over 90% of evaluated cases. Second, repetitions suffer from a strong self-reinforcement effect, where token probabilities increase almost monotonically with each subsequent repetition until reaching a high ceiling value. Third, sentences with higher initial probabilities—such as those chosen by greedy decoding—experience a significantly faster escalation into repetitive loops. Fourth, fine-tuning with DITTO effectively dismantles this reinforcement dynamic; on Wikitext-103 greedy decoding, DITTO reduced sentence repetition rates to 2.85% (down from 14.50% in standard models) and improved human similarity MAUVE scores from 0.34 to 0.77 while maintaining lower perplexity (24.33 vs. 25.68). In human evaluation matchups, DITTO achieved win rates between 62% and 84% over competing methods.
These findings demonstrate that repetition is an intrinsic reinforcement artifact in probability assignment rather than a simple vocabulary limitation. Unlike earlier training interventions that penalize all context tokens indiscriminately and degrade fluency, DITTO selectively targets over-repetition without harming necessary natural repetitions such as proper nouns. This yields direct operational benefits by producing higher-fidelity text generation without incurring additional computational overhead during inference.
For practical implementation, engineering teams deploying text generation pipelines should adopt pseudo-repetition fine-tuning alongside standard maximum likelihood training. Implementations should balance real and pseudo data equally (a 50-50 mix) and calibrate the decay factor based on task freedom—using a stronger penalty (such as 0.5) for open-ended generation and a milder penalty (such as 0.9) for constrained tasks like summarization. While the empirical results demonstrate robust improvements across diverse model families and decoding setups, the authors note that the study focused primarily on sequence-level probability dynamics without exploring deeper underlying token-embedding or neural architecture causes, warranting continued exploration into model architectural safeguards.
- Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). This foundational work establishes how standard maximization decoding causes repetitive degeneration in language models, providing the core problem context that the source paper's training penalty seeks to analyze and mitigate.
- Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). Understanding exposure bias and the drawbacks of step-by-step cross-entropy training under greedy decoding clarifies why neural text generation models fall into self-reinforcing repetitive loops.
- Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). This paper introduces architectural coverage mechanisms to track generation history and reduce repetition in sequence-to-sequence generation, serving as an important predecessor to fine-tuning strategies for repetition mitigation.
- Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). Reading this work provides critical background on how intra-attention and objective modifications are used to overcome exposure bias and prevent repetitive phrasing in multi-sentence generation.
- Paper: A Contrastive Framework for Neural Text Generation, Yixuan Su et al. (2022). This work directly extends the investigation of repetition in maximization-based decoding by proposing a contrastive training and decoding framework to prevent text degeneration.
- Paper: Controlled Text Generation with Natural Language Instructions, Wangchunshu Zhou et al. (2023). Building on training-time interventions for controlled text generation, this paper explores training language models with natural language instructions to satisfy structured constraints without degenerative decoding.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). This work advances post-training optimization by introducing direct preference optimization to eliminate degenerate or dispreferred model behaviors without explicit reinforcement learning loops.
- Paper: Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions, John Joon Young Chung et al. (2023). This article applies logit suppression and diversity techniques to address repetitive generation issues when synthesizing training data from large language models.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). This paper explores an alternative post-generation paradigm, using iterative self-feedback to detect and refine degenerate or suboptimal text outputs without retraining.
