Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation

Jin XuXiaojiang LiuJianhao YanDeng CaiHuayang LiJian Li

article2022NeurIPS104 citations

Reveals the self-reinforcement mechanism behind repetitive neural text generation loops and introduces DITTO, a training method that penalizes repetition probabilities using pseudo-data without sacrificing text quality or perplexity.

Listen

Large-scale neural language models often get trapped in undesirable loops of repeating entire sentences consecutively when generating text, especially under standard greedy or maximum-probability decoding strategies. This behavior contradicts natural human language, where full consecutive sentence repetitions are virtually nonexistent (around 0.02% in benchmark corpora). As natural language generation becomes integral to automated summarization, customer interactions, and content creation, repetitive loops severely degrade text quality, user trust, and overall system utility.

The article investigates the fundamental mechanisms driving these sentence-level repetitive loops and evaluates a novel training framework designed to suppress redundant repetitions while preserving overall language generation quality.

To analyze this phenomenon, the researchers evaluated token probability behavior across 1,000 synthetic test sequences from multiple data sources, including Wikipedia, BookCorpus, and random word sets. They systematically measured probability changes as sentences were repeated up to 100 times. Using these diagnostic findings, the authors designed Pseudo-Repetition Penalization (DITTO), a lightweight fine-tuning strategy that feeds artificially repeated sentences into the model and penalizes repetition probabilities using an exponential decay factor. DITTO was evaluated against existing baselines across open-ended generation tasks on Wikitext-103 using a 750-million-parameter Transformer and abstractive summarization on CNN/DailyMail using a BART-large model, utilizing automated metrics (such as MAUVE and ROUGE) alongside human evaluations.

The investigation produced several key findings regarding neural text generation. First, language models exhibit an inherent copying shortcut: introducing just one sentence-level repetition causes the probability of repeating the next token to rise in over 90% of evaluated cases. Second, repetitions suffer from a strong self-reinforcement effect, where token probabilities increase almost monotonically with each subsequent repetition until reaching a high ceiling value. Third, sentences with higher initial probabilities—such as those chosen by greedy decoding—experience a significantly faster escalation into repetitive loops. Fourth, fine-tuning with DITTO effectively dismantles this reinforcement dynamic; on Wikitext-103 greedy decoding, DITTO reduced sentence repetition rates to 2.85% (down from 14.50% in standard models) and improved human similarity MAUVE scores from 0.34 to 0.77 while maintaining lower perplexity (24.33 vs. 25.68). In human evaluation matchups, DITTO achieved win rates between 62% and 84% over competing methods.

These findings demonstrate that repetition is an intrinsic reinforcement artifact in probability assignment rather than a simple vocabulary limitation. Unlike earlier training interventions that penalize all context tokens indiscriminately and degrade fluency, DITTO selectively targets over-repetition without harming necessary natural repetitions such as proper nouns. This yields direct operational benefits by producing higher-fidelity text generation without incurring additional computational overhead during inference.

For practical implementation, engineering teams deploying text generation pipelines should adopt pseudo-repetition fine-tuning alongside standard maximum likelihood training. Implementations should balance real and pseudo data equally (a 50-50 mix) and calibrate the decay factor based on task freedom—using a stronger penalty (such as 0.5) for open-ended generation and a milder penalty (such as 0.9) for constrained tasks like summarization. While the empirical results demonstrate robust improvements across diverse model families and decoding setups, the authors note that the study focused primarily on sequence-level probability dynamics without exploring deeper underlying token-embedding or neural architecture causes, warranting continued exploration into model architectural safeguards.

arXiv: 2206.02369
  • Paper: The Curious Case of Neural Text Degeneration, Ari Holtzman et al. (2020). This foundational work establishes how standard maximization decoding causes repetitive degeneration in language models, providing the core problem context that the source paper's training penalty seeks to analyze and mitigate.
  • Paper: Sequence Level Training with Recurrent Neural Networks, Marc'Aurelio Ranzato et al. (2015). Understanding exposure bias and the drawbacks of step-by-step cross-entropy training under greedy decoding clarifies why neural text generation models fall into self-reinforcing repetitive loops.
  • Paper: Get To The Point: Summarization with Pointer-Generator Networks, Abigail See et al. (2017). This paper introduces architectural coverage mechanisms to track generation history and reduce repetition in sequence-to-sequence generation, serving as an important predecessor to fine-tuning strategies for repetition mitigation.
  • Paper: A Deep Reinforced Model for Abstractive Summarization, Romain Paulus et al. (2017). Reading this work provides critical background on how intra-attention and objective modifications are used to overcome exposure bias and prevent repetitive phrasing in multi-sentence generation.
Cover for Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation

Abstract

While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (e.g., greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in human corpora (e.g., 0.02% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probabilities of the repetitive tokens and their previous repetitions in the context. Through our quantitative experiments, we find that 1) Language models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a self-reinforcement effect: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method DITTO (PseuDo-RepeTition PenalizaTion), where the model learns to penalize probabilities of sentence-level repetitions from pseudo repetitive data. Although our method is motivated by mitigating repetitions, experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method. Code is released at https://github.com/Jxu-Thu/DITTO.

Table of Contents

  • 1 Introduction
  • 2 Analyzing Repetition
  • 2.1 Experiment Design
  • 2.2 Results and Analyses
  • 3 Pseudo-repetition Penalization Training
  • 4 Experiments
  • 4.1 Setup
  • 4.2 Results of Open-ended Generation
  • 4.3 Results of Directed Generation
  • 4.4 Analyses
  • 5 Related Work
  • 6 Conclusion and Future Work
  • 7 Acknowledgements
  • References
  • Checklist

Knowls

  1. Knowl 1 — DITTO: Pseudo-Repetition Penalization Training

    model/method

    Pseudo-Repetition Penalization (DITTO) is a training method designed to mitigate sentence-level loops in neural text generation by training the model to decay the probability of repetitive sentences.

    Given a sentence ss sampled from a training corpus, a pseudo repetitive sequence x=(s0,s1,…,sN)x = (s^0, s^1, \dots, s^N) is constructed by repeating ss for N+1N+1 iterations until a maximum sequence length is reached, where sn=(xn,1,…,xn,Ls)s^n = (x_{n,1}, \dots, x_{n,L_s}) represents the nn-th occurrence of the sentence and LsL_s is its token length. The previous context of sentence ss can optionally be appended as a prefix to xx.

    For the ll-th token xn,lx_{n,l} in the nn-th repetition (n≥1n \ge 1), the per-step DITTO penalization loss is defined as:

    LDITTOn,l(Pθ(xn,l∣x<n,l))=−log⁡(1−∣Pθ(xn,l∣x<n,l)−λ⋅Pθ∗(xn−1,l∣x<n−1,l)∣)\mathcal{L}^{n,l}_{\text{DITTO}}\left(P_\theta(x_{n,l} \mid x_{<n,l})\right) = -\log \left( 1 - \left| P_\theta(x_{n,l} \mid x_{<n,l}) - \lambda \cdot P^*_\theta(x_{n-1,l} \mid x_{<n-1,l}) \right| \right)

    where:

    • x<n,l=(s0,…,sn−1,xn,1,…,xn,l−1)x_{<n,l} = (s^0, \dots, s^{n-1}, x_{n,1}, \dots, x_{n,l-1}) is the full preceding context across repetitions and within the current sentence.
    • Pθ(⋅∣⋅)P_\theta(\cdot \mid \cdot) is the conditional probability predicted by the model parameterized by heta heta.
    • Pθ∗(xn−1,l∣x<n−1,l)P^*_\theta(x_{n-1,l} \mid x_{<n-1,l}) denotes the probability assigned to the corresponding token in the previous repetition, detached from gradient computation.
    • λ∈(0,1]\lambda \in (0, 1] is a scalar penalization factor. Setting λ=1\lambda = 1 stabilizes repetition probability across iterations, while λ<1\lambda < 1 enforces an exponential probability decay with rate λ\lambda across repetitions.

    Training is performed by fine-tuning an MLE-pretrained language model using an equal mixture of standard MLE loss updates on genuine data and DITTO loss updates on pseudo repetitive sequences (mix ratio γ=0.5\gamma = 0.5).

  2. Knowl 2 — Metrics for Quantifying Contextual Token Probabilities in Repetitive Sentences

    definition

    To quantitatively evaluate how historical repetitions influence model probabilities, a test sequence x=(s0,s1,…,sN)x = (s^0, s^1, \dots, s^N) is constructed by repeating a sentence ss of length LsL_s tokens NN times, where sn=(xn,1,…,xn,Ls)s^n = (x_{n,1}, \dots, x_{n,L_s}). A token xn,lx_{n,l} is said to have a sentence-level context repetition if an identical preceding token xi,l=xn,lx_{i,l} = x_{n,l} exists with matching intra-sentence context xi,<l=xn,<lx_{i,<l} = x_{n,<l} for i<ni < n.

    For a model PθP_\theta, probability dynamics across the nn-th repetition sns^n are evaluated using three metrics:

    1. Average Token Probability (TP\text{TP}): The mean conditional probability across all tokens in sns^n:

    TP(sn)=1Ls∑l=1LsPθ(xn,l∣x<n,l)\text{TP}(s^n) = \frac{1}{L_s} \sum_{l=1}^{L_s} P_\theta(x_{n,l} \mid x_{<n,l})

    where TP(s0)\text{TP}(s^0) denotes the initial token probability in the unrepeated sentence.

    1. Rate of Increased Token Probability (IP\text{IP}): The proportion of tokens in sns^n whose conditional probability strictly exceeds their probability in the initial sentence s0s^0:

    IP(sn)=1Ls∑l=1Ls1(Pθ(xn,l∣x<n,l)>Pθ(x0,l∣x<0,l))\text{IP}(s^n) = \frac{1}{L_s} \sum_{l=1}^{L_s} \mathbf{1}\left(P_\theta(x_{n,l} \mid x_{<n,l}) > P_\theta(x_{0,l} \mid x_{<0,l})\right)

    where 1(⋅)\mathbf{1}(\cdot) is the indicator function.

    1. Winner Rate (WR\text{WR}): The proportion of tokens in sns^n that both increased in probability relative to s0s^0 and constitute the argmax (top-1) token prediction of the model:

    WR(sn)=1Ls∑l=1Ls1(Pθ(xn,l∣x<n,l)>Pθ(x0,l∣x<0,l)  ∧  xn,l=arg⁡max⁡Pθ(⋅∣x<n,l))\text{WR}(s^n) = \frac{1}{L_s} \sum_{l=1}^{L_s} \mathbf{1}\left(P_\theta(x_{n,l} \mid x_{<n,l}) > P_\theta(x_{0,l} \mid x_{<0,l}) \;\land\; x_{n,l} = \arg\max P_\theta(\cdot \mid x_{<n,l})\right)

    Averaged across a corpus D\mathcal{D}, the aggregate metrics are TPn=1∣D∣∑s∈DTP(sn)\text{TP}_n = \frac{1}{|\mathcal{D}|}\sum_{s \in \mathcal{D}} \text{TP}(s^n), IPn=1∣D∣∑s∈DIP(sn)\text{IP}_n = \frac{1}{|\mathcal{D}|}\sum_{s \in \mathcal{D}} \text{IP}(s^n), and WRn=1∣D∣∑s∈DWR(sn)\text{WR}_n = \frac{1}{|\mathcal{D}|}\sum_{s \in \mathcal{D}} \text{WR}(s^n).

  3. Knowl 3 — Self-Reinforcement Effect and Sentence Repetition Dynamics in Language Models

    empirical result

    Quantitative experiments feeding repeated sentence sequences into standard Transformer decoders (such as a 16-layer, 750M-parameter model) across random vocabulary sequences (DrandomD_{\text{random}}), BookCorpus sentences (DbookD_{\text{book}}), and Wikitext-103 sentences (DwikiD_{\text{wiki}}) demonstrate three key phenomena:

    1. Immediate Probability Elevation: Even with only a single previous context repetition (n=1n=1), the Rate of Increased Token Probability (IP1\text{IP}_1) exceeds 90%90\% across all corpora. When encountering a previously seen sentence prefix, the model disproportionately boosts the probability of copying the succeeding token rather than generating novel continuation.
    2. Self-Reinforcement Effect: As the repetition count nn grows from 1 to 100, Average Token Probability (TPn\text{TP}_n), Rate of Increased Token Probability (IPn\text{IP}_n), and Winner Rate (WRn\text{WR}_n) increase almost monotonically and converge toward ceiling values. The more times a sentence is present in the context, the higher the probability of continuing to repeat it under greedy and maximization-based decoding.
    3. Dependence on Initial Probability: Natural sentences with higher baseline likelihood TP0\text{TP}_0 (from DwikiD_{\text{wiki}} and DbookD_{\text{book}}) exhibit a steeper self-reinforcement rate, reaching ceiling values in fewer repetitions than low-likelihood random sequences (DrandomD_{\text{random}}). Maximization-based decoding algorithms naturally generate high-probability sentences, accelerating the onset of self-reinforcing repetition loops.
  4. Knowl 4 — Wikitext-103 Open-Ended Generation Performance Under Greedy Decoding

    data/table

    On the Wikitext-103 test set for open-ended generation (50-token prefix, 100 greedily decoded tokens), fine-tuning a 750M-parameter Transformer with DITTO (λ=0.5\lambda = 0.5) substantially reduces phrase-level and sentence-level repetitions while improving next-token prediction accuracy, perplexity, and MAUVE score relative to Maximum Likelihood Estimation (MLE), unlikelihood training (UL-token, UL-token+seq), and straight-to-gradient (SG).

    Model MAUVE ↑\uparrow Perplexity ↓\downarrow Accuracy ↑\uparrow Repetition-4 ↓\downarrow Repetition-Sen ↓\downarrow
    MLE 0.34±0.020.34 \pm 0.02 25.68±0.0425.68 \pm 0.04 0.39±0.000.39 \pm 0.00 44.20±1.43%44.20 \pm 1.43\% 14.50±1.59%14.50 \pm 1.59\%
    UL-token 0.57±0.010.57 \pm 0.01 26.98±0.1226.98 \pm 0.12 0.39±0.000.39 \pm 0.00 28.30±0.78%28.30 \pm 0.78\% 7.40±0.83%7.40 \pm 0.83\%
    UL-token+seq 0.48±0.030.48 \pm 0.03 25.95±0.0825.95 \pm 0.08 0.40±0.000.40 \pm 0.00 7.60±0.46%7.60 \pm 0.46\% 0.05±0.03%0.05 \pm 0.03\%
    SG 0.74±0.010.74 \pm 0.01 25.84±0.0625.84 \pm 0.06 0.40±0.000.40 \pm 0.00 23.00±0.28%23.00 \pm 0.28\% 5.24±0.75%5.24 \pm 0.75\%
    DITTO 0.77±0.01\mathbf{0.77 \pm 0.01} 24.33±0.04\mathbf{24.33 \pm 0.04} 0.42±0.00\mathbf{0.42 \pm 0.00} 22.00±0.31%22.00 \pm 0.31\% 2.85±0.74%2.85 \pm 0.74\%
    Human — — — 1.10%1.10\% 0.01%0.01\%

    Repetition-4 is defined as 1.0−∣unique 4-grams∣/∣4-grams∣1.0 - |\text{unique 4-grams}| / |\text{4-grams}|, and Repetition-Sen is defined as 1.0−∣unique sentences∣/∣sentences∣1.0 - |\text{unique sentences}| / |\text{sentences}|. While unlikelihood baselines achieve repetition reduction at the cost of worsening perplexity (e.g., 26.9826.98 for UL-token vs 25.6825.68 for MLE), DITTO lowers perplexity to 24.3324.33 and raises accuracy to 0.420.42 while achieving the highest MAUVE score (0.770.77).

  5. Knowl 5 — Compatibility of DITTO with Stochastic Decoding Strategies

    data/table

    When paired with stochastic decoding methods—top-kk sampling (k=50k=50) and nucleus top-pp sampling (p=0.9p=0.9)—on Wikitext-103 open-ended generation (50 prefix tokens, 100 generated tokens), DITTO achieves repetition rates closely aligned with human text and produces the highest MAUVE scores (0.960.96).

    Search Model MAUVE ↑\uparrow Repetition-4 Repetition-Sen
    Top-kk (k=50k=50) MLE 0.94±0.000.94 \pm 0.00 1.60±0.09%1.60 \pm 0.09\% 0.25±0.06 \textperthousand0.25 \pm 0.06\text{ \textperthousand}
    UL-token 0.95±0.000.95 \pm 0.00 0.70±0.13%0.70 \pm 0.13\% 0.00±0.00 \textperthousand0.00 \pm 0.00\text{ \textperthousand}
    UL-token+seq 0.93±0.010.93 \pm 0.01 0.09±0.11%0.09 \pm 0.11\% 0.06±0.02 \textperthousand0.06 \pm 0.02\text{ \textperthousand}
    SG 0.93±0.010.93 \pm 0.01 0.50±0.19%0.50 \pm 0.19\% 0.00±0.00 \textperthousand0.00 \pm 0.00\text{ \textperthousand}
    DITTO 0.96±0.00\mathbf{0.96 \pm 0.00} 1.00±0.10%\mathbf{1.00 \pm 0.10\%} 0.09±0.01 \textperthousand\mathbf{0.09 \pm 0.01\text{ \textperthousand}}
    Nucleus (p=0.9p=0.9) MLE 0.94±0.000.94 \pm 0.00 1.40±0.08%1.40 \pm 0.08\% 0.08±0.01 \textperthousand0.08 \pm 0.01\text{ \textperthousand}
    UL-token 0.94±0.000.94 \pm 0.00 0.47±0.08%0.47 \pm 0.08\% 0.00±0.00 \textperthousand0.00 \pm 0.00\text{ \textperthousand}
    UL-token+seq 0.94±0.010.94 \pm 0.01 0.08±0.05%0.08 \pm 0.05\% 0.02±0.02 \textperthousand0.02 \pm 0.02\text{ \textperthousand}
    SG 0.93±0.010.93 \pm 0.01 0.40±0.19%0.40 \pm 0.19\% 0.06±0.01 \textperthousand0.06 \pm 0.01\text{ \textperthousand}
    DITTO 0.96±0.00\mathbf{0.96 \pm 0.00} 0.98±0.09%\mathbf{0.98 \pm 0.09\%} 0.08±0.01 \textperthousand\mathbf{0.08 \pm 0.01\text{ \textperthousand}}
    Human — — 1.10%1.10\% 0.10 \textperthousand0.10\text{ \textperthousand}

    Under top-kk sampling, DITTO yields 1.00%1.00\% Repetition-4 and 0.09\textperthousand0.09\text{\textperthousand} Repetition-Sen compared to human scores of 1.10%1.10\% and 0.10\textperthousand0.10\text{\textperthousand}. Under nucleus sampling, DITTO yields 0.98%0.98\% Repetition-4 and 0.08\textperthousand0.08\text{\textperthousand} Repetition-Sen, avoiding both under-repetition (over-penalization seen in UL-token+seq) and over-repetition (MLE).

  6. Knowl 6 — Abstractive Text Summarization Performance on CNN/DailyMail

    data/table

    On the CNN/DailyMail abstractive summarization benchmark, fine-tuning BART-large with DITTO (λ=0.9\lambda = 0.9) using beam search (beam size =5= 5) and tri-gram blocking achieves superior ROUGE-1, ROUGE-2, and ROUGE-L scores compared to baseline BART-large models trained with MLE, UL, or SG, as well as other dedicated architectures.

    Model ROUGE-1 ↑\uparrow ROUGE-2 ↑\uparrow ROUGE-L ↑\uparrow
    Pointer-generator + Coverage 39.53 17.28 36.38
    Mask Attention Network 40.98 18.29 37.88
    BertSum 42.13 19.60 39.18
    UniLM 43.08 20.43 40.34
    UniLM V2 43.16 20.42 40.14
    ERNIE-GEN-large 44.02 21.17 41.26
    PEGASUS 44.17 21.47 41.11
    ProphetNet 44.20 21.17 41.30
    PALM 44.30 21.12 41.14
    BART-large w.t. MLE 44.11±0.0344.11 \pm 0.03 21.21±0.0121.21 \pm 0.01 40.83±0.0240.83 \pm 0.02
    BART-large w.t. UL-token 44.17±0.0444.17 \pm 0.04 21.20±0.0221.20 \pm 0.02 40.83±0.0340.83 \pm 0.03
    BART-large w.t. UL-token+seq 44.13±0.0744.13 \pm 0.07 21.15±0.1121.15 \pm 0.11 40.71±0.0940.71 \pm 0.09
    BART-large w.t. SG 44.18±0.0644.18 \pm 0.06 21.17±0.0721.17 \pm 0.07 40.89±0.0540.89 \pm 0.05
    BART-large w.t. DITTO 44.41±0.03\mathbf{44.41 \pm 0.03} 21.45±0.01\mathbf{21.45 \pm 0.01} 41.16±0.02\mathbf{41.16 \pm 0.02}

    For directed summarization, pseudo-data is created by sampling a sentence from the reference summary and repeating it in the target sequence while leaving the input article unchanged. The results demonstrate that DITTO is effective for directed generation and compatible with inference-time decoding heuristics such as nn-gram blocking.

  7. Knowl 7 — Hyperparameter Effects and Decoding Length Robustness in DITTO

    empirical result

    Systematic evaluations of DITTO across hyperparameter choices and decoding constraints show:

    1. Mix Ratio (γ\gamma): Performance measured by MAUVE and ROUGE displays an inverted U-shape across pseudo-to-actual data mixing ratios γ∈[0,1]\gamma \in [0, 1], reaching maximum performance at γ=0.5\gamma = 0.5 on both Wikitext-103 and CNN/DailyMail. Balancing DITTO penalty updates equally with standard MLE updates is optimal.
    2. Penalization Factor (λ\lambda): Optimal values of λ\lambda depend on task constraints. Open-ended generation (Wikitext-103) benefits from stronger penalization (λ=0.5\lambda = 0.5), whereas directed summarization (CNN/DailyMail) achieves highest performance with milder penalization (λ=0.9\lambda = 0.9).
    3. Decoding Length Robustness: Across output sequence lengths ranging from 60 to 200 tokens during greedy auto-completion, DITTO consistently maintains higher MAUVE scores and lower sentence repetition rates than MLE baselines.
  8. Knowl 8 — Architectural and Representation Scope Limitations of DITTO

    limitation

    Although DITTO effectively characterizes and mitigates the self-reinforcement effect via pseudo-data penalty training, it operates at the input-output sequence and loss level without explaining or addressing the internal representation mechanisms (such as token embedding dynamics, self-attention routing, or intrinsic linguistic structure) that cause neural architectures to develop the shortcut of assigning higher probabilities to copied context tokens in the first place.

Coverage note — None was omitted; the extracted knowls comprehensively represent the empirical analysis of repetition dynamics, the DITTO training framework, main open-ended and directed generation results, hyperparameter analyses, and stated limitations.

References

  1. 1.Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. Towards a human-like open-domain chatbot. arXiv preprint arXiv:2001.09977, 2020.
  2. 2.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, et al. Unilmv2: Pseudo-masked language models for unified language model pre-training. In International Conference on Machine Learning, pages 642–652. PMLR, 2020.
  3. 3.Sourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, and Lav R Varshney. Mirostat: A neural text decoding algorithm that directly controls perplexity. In International Conference on Learning Representations, 2020.
  4. 4.Bin Bi, Chenliang Li, Chen Wu, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. Palm: Pre-training an autoencoding&autoregressive language model for context-conditioned generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8681–8691, 2020.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  6. 6.Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language understanding and generation. Advances in Neural Information Processing Systems, 32, 2019.
  7. 7.Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 889–898, 2018.
  8. 8.Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, and Xuan-Jing Huang. Mask attention networks: Rethinking and strengthen transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1692–1701, 2021.
  9. 9.Zihao Fu, Wai Lam, Anthony Man-Cho So, and Bei Shi. A theoretical analysis of the repetition problem in text generation. arXiv preprint arXiv:2012.14660, 2020.
  10. 10.Jian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu, Wenbiao Ding, and Minlie Huang. Long text generation by modeling sentence-level and discourse-level coherence. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6379–6393, 2021.
  11. 11.Tianxing He, Jingzhao Zhang, Zhiming Zhou, and James Glass. Exposure bias versus self-recovery: Are distortions really incremental for autoregressive text generation? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5087–5102, 2021.
  12. 12.Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and comprehend. Advances in neural information processing systems, 28, 2015.
  13. 13.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, 2019.
  14. 14.Ari Holtzman, Jan Buys, Maxwell Forbes, Antoine Bosselut, David Golub, and Yejin Choi. Learning to write with cooperative discriminators. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1638–1649, 2018.
  15. 15.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3128–3137, 2015.
  16. 16.Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M Rush. Opennmt: Open-source toolkit for neural machine translation. In Proceedings of ACL 2017, System Demonstrations, pages 67–72, 2017.
  17. 17.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, 2020.
  18. 18.Jiwei Li, Will Monroe, and Dan Jurafsky. A simple, fast diverse decoding algorithm for neural generation. arXiv preprint arXiv:1611.08562, 2016.
  19. 19.Xiang Lin, Simeng Han, and Shafiq Joty. Straight to the gradient: Learning to use novel tokens for neural text generation. In International Conference on Machine Learning, pages 6642–6653. PMLR, 2021.
  20. 20.Yang Liu and Mirella Lapata. Text summarization with pretrained encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3730–3740, 2019.
  21. 21.Clara Meister, Gian Wiher, Tiago Pimentel, and Ryan Cotterell. On the probability-quality paradox in language generation. arXiv preprint arXiv:2203.17217, 2022.
  22. 22.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016.
  23. 23.Ramesh Nallapati, Bowen Zhou, Cıcero Nogueira dos Santos, Caglar Gulcehre, and Bing Xiang. Abstractive text summarization using sequence-to-sequence rnns and beyond. In Yoav Goldberg and Stefan Riezler, editors, Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, CoNLL 2016, Berlin, Germany, August 11-12, 2016, pages 280–290. ACL, 2016.
  24. 24.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations, 2019.
  25. 25.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc., 2019.
  26. 26.Romain Paulus, Caiming Xiong, and Richard Socher. A deep reinforced model for abstractive summarization. In International Conference on Learning Representations, 2018.
  27. 27.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. Mauve: Measuring the gap between neural text and human text using divergence frontiers. Advances in Neural Information Processing Systems, 34, 2021.
  28. 28.Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou. Prophetnet: Predicting future n-gram for sequence-to-sequencepre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2401–2410, 2020.
  29. 29.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019.
  30. 30.Laria Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–7, 2021.
  31. 31.Lin CY ROUGE. A package for automatic evaluation of summaries. In Proceedings of Workshop on Text Summarization of ACL, Spain, 2004.
  32. 32.Abigail See, Peter J Liu, and Christopher D Manning. Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1073–1083, 2017.
  33. 33.Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. A contrastive framework for neural text generation. arXiv preprint arXiv:2202.06417, 2022.
  34. 34.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  35. 35.Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, and Kyunghyun Cho. Consistency of a recurrent language model with respect to incomplete decoding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5553–5568, 2020.
  36. 36.Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. Neural text generation with unlikelihood training. In International Conference on Learning Representations, 2019.
  37. 37.Dongling Xiao, Han Zhang, Yukun Li, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-gen: an enhanced multi-flow pre-training and fine-tuning framework for natural language generation. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, pages 3997–4003, 2021.
  38. 38.Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pages 11328–11339. PMLR, 2020.
  39. 39.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068, 2022.
  40. 40.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In The IEEE International Conference on Computer Vision (ICCV), December 2015.

Citation

MLA
Xu, J., et al. “Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 3082–95, https://proceedings.neurips.cc/paper_files/paper/2022/file/148c0aeea1c5da82f4fa86a09d4190da-Paper-Conference.pdf.
APA
Xu, J., Liu, X., Yan, J., Cai, D., Li, H., & Li, J. (2022). Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation. Advances in Neural Information Processing Systems, 35, 3082–3095. https://proceedings.neurips.cc/paper_files/paper/2022/file/148c0aeea1c5da82f4fa86a09d4190da-Paper-Conference.pdf
Chicago
Xu, J., X. Liu, J. Yan, D. Cai, H. Li, and J. Li. 2022. “Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation”. Advances in Neural Information Processing Systems 35: 3082–95. https://proceedings.neurips.cc/paper_files/paper/2022/file/148c0aeea1c5da82f4fa86a09d4190da-Paper-Conference.pdf.
Harvard
Xu, J. et al. (2022) “Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 3082–3095. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/148c0aeea1c5da82f4fa86a09d4190da-Paper-Conference.pdf.
Vancouver
1. Xu J, Liu X, Yan J, Cai D, Li H, Li J (2022) Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 3082–3095

BibTeX

@inproceedings{xu2022learning,
  title = {Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation},
  author = {Xu, Jin and Liu, Xiaojiang and Yan, Jianhao and Cai, Deng and Li, Huayang and Li, Jian},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {3082-3095},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/148c0aeea1c5da82f4fa86a09d4190da-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors