Latent Diffusion for Language Generation

Justin LovelaceVarsha KishoreChao WanEliot ShekhtmanKilian Q. Weinberger

article2023NeurIPS173 citations

Presents a framework that applies continuous diffusion models within the compact latent space of pretrained encoder-decoder language models, outperforming existing text diffusion methods across conditional and sequence-to-sequence generation tasks with significantly fewer sampling steps.

Listen

Diffusion models have driven major breakthroughs in generating continuous media such as images and audio, yet applying them to discrete text has proven difficult. Previous efforts attempted to replace existing language models by learning diffusion directly on individual word embeddings, but these models often suffered from training instability, degraded generation quality, and high computational costs. The article evaluates a new framework, Latent Diffusion for Language Generation (LD4LG), which demonstrates that continuous diffusion processes and pretrained language models can work complementarily rather than in competition.

The approach uses a pretrained encoder-decoder model (such as BART or FLAN-T5) paired with lightweight compression and reconstruction networks. The compression network reduces variable-length, high-dimensional text representations into a compact, fixed-size continuous latent space (reducing dimensions by a factor of 24×). A continuous diffusion model is then trained purely within this low-dimensional semantic space. During generation, the diffusion model produces a latent representation from random noise, which the pretrained decoder then translates back into natural text. The method was evaluated across multiple generation tasks using benchmark datasets, including ROCStories, AG News, Quora Question Pairs, XSum summarization, and WMT 2014 English-German translation.

The findings show that this latent diffusion framework substantially outperforms previous text diffusion methods across all evaluated tasks while requiring significantly fewer generation steps. On the ROCStories benchmark, LD4LG achieved a MAUVE quality score of 0.716 using 250 sampling steps, compared to 0.043 across 2,000 steps for Diffusion-LM. For sequence-to-sequence summarization on XSum, LD4LG scored a ROUGE-L of 31.9, more than doubling the 14.1 score of the DiffuSeq baseline. Furthermore, compared to a fine-tuned GPT-2 autoregressive model, LD4LG demonstrated substantially lower rates of training data memorization (for instance, 0.293 versus 0.829 on AG News) and showed superior steering ability in class-conditional topic generation.

These results demonstrate that separating discrete text decoding from continuous semantic planning makes diffusion viable and efficient for natural language generation. By operating in a low-dimensional latent space, the framework achieves a nearly fourfold training speedup over uncompressed diffusion baselines and reduces inference steps by roughly eightfold compared to previous text diffusion systems. This capability reduces compliance and privacy risks associated with models memorizing sensitive training data, while maintaining competitive text generation quality.

Future efforts should focus on adapting rapid sampling and model distillation techniques from the computer vision domain to reduce the required inference steps down toward single-step generation. Organizations adopting text diffusion should also explore improved re-ranking and candidate selection techniques, as optimal sample selection was shown to consistently outperform standard fine-tuned baselines. The evidence provides high confidence in the quality and memorization advantages of the latent diffusion approach, though practical deployment is currently bounded by the higher latency of iterative sampling relative to standard single-pass autoregressive decoders.

  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). Its discrete denoising framework establishes how diffusion can model categorical data, a foundation for understanding why this paper instead moves language generation into a continuous autoencoder latent space.
  • Paper: Generating Sentences from a Continuous Space, Samuel R. Bowman et al. (2016). Its sentence-level variational autoencoder introduces continuous latent representations for text, clarifying the autoencoding step that this paper uses before applying diffusion.
  • Paper: Self-conditioned Embedding Diffusion for Text Generation, Robin Strudel et al. (2022). Its continuous embedding diffusion applies denoising directly to text representations, providing a close point of comparison for this paper’s move to diffusion over learned language latents.
  • Paper: Continuous diffusion for categorical data, Sander Dieleman et al. (2022). Its method adapts continuous diffusion to categorical language through token embeddings, providing a useful contrast to this paper’s pretrained autoencoder and latent-space approach.
  • Paper: The Diffusion Duality, Subham Sekhar Sahoo et al. (2025). It maps continuous Gaussian diffusion to discrete text generation, extending the continuous-to-discrete bridge that this paper explores through language autoencoder latents.
  • Paper: ELF: Embedded Language Flows, Keya Hu et al. (2026). Its continuous language flows denoise in embedding space and defer token conversion, carrying forward this paper’s strategy of using continuous generative modeling for language.
Cover for Latent Diffusion for Language Generation

Abstract

Diffusion models have achieved great success in modeling continuous data modalities such as images, audio, and video, but have seen limited use in discrete domains such as language. Recent attempts to adapt diffusion to language have presented diffusion as an alternative to existing pretrained language models. We view diffusion and existing language models as complementary. We demonstrate that encoder-decoder language models can be utilized to efficiently learn high-quality language autoencoders. We then demonstrate that continuous diffusion models can be learned in the latent space of the language autoencoder, enabling us to sample continuous latent representations that can be decoded into natural language with the pretrained decoder. We validate the effectiveness of our approach for unconditional, class-conditional, and sequence-to-sequence language generation. We demonstrate across multiple diverse data sets that our latent language diffusion models are significantly more effective than previous diffusion language models. Our code is available at https://github.com/justinlovelace/latent-diffusion-for-language.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Background
  • 3 Latent Diffusion For Language
  • 3.1 Language Autoencoder
  • 3.2 Latent Language Diffusion
  • 4 Datasets
  • 4.1 Evaluation Metrics.
  • 5 Experiments
  • 5.1 Language Autoencoder
  • 5.2 Unconditional Language Generation
  • 5.3 Conditional Language Generation
  • 5.4 Sequence-to-Sequence Language Generation
  • 6 Future Work
  • 7 Related Work
  • 8 Conclusion
  • Acknowledgements
  • References
  • Appendix A: Diffusion Models
  • Additional Language Autoencoder Results

Knowls

  1. Knowl 1 — LD4LG generates text by diffusing in a compact language-model latent space

    model/method

    Latent Diffusion for Language Generation (LD4LG) combines a pretrained encoder–decoder language model with continuous diffusion. For an input token sequence ww, the encoder produces contextual features, a learned compression network maps those features to a short, fixed-length continuous representation, and a diffusion model learns the distribution of these representations. To generate text, LD4LG samples and denoises a latent, maps it back to decoder-compatible features with a learned reconstruction network, and uses the pretrained autoregressive decoder to produce a variable-length token sequence. The encoder–decoder handles the mapping between continuous representations and discrete text, while diffusion models the compact continuous latent distribution. The method is evaluated for unconditional, class-conditional, and source-conditioned sequence-to-sequence generation.

  2. Knowl 2 — A Perceiver-based autoencoder compresses and reconstructs language features

    model/method

    LD4LG’s language autoencoder uses a pretrained encoder EE to map a token sequence ww of length LL to contextual features E(w)∈RL×dLME(w)\in\mathbb{R}^{L\times d_{\rm LM}}. A Perceiver Resampler compression network uses ℓ\ell learnable queries to cross-attend to the encoder features and to the queries themselves, producing a fixed-length sequence. A linear projection reduces its feature dimension to daed_{\rm ae}, giving x=fϕ(E(w))∈Rℓ×daex=f_\phi(E(w))\in\mathbb{R}^{\ell\times d_{\rm ae}}. A reconstruction network projects xx back to dimension dLMd_{\rm LM}, adds learned absolute positional embeddings, and processes the result with a transformer; the language decoder cross-attends to these reconstructed features and is trained to reproduce ww using cross-entropy loss.

    For monolingual datasets, the experiments use ℓ=32\ell=32, dae=64d_{\rm ae}=64, and three layers in each learned autoencoding network. The pretrained encoder and decoder are normally frozen. Latent vectors are norm-constrained so each latent position xix_i has squared feature norm ∥xi∥22=dae\|x_i\|_2^2=d_{\rm ae}, except with FLAN-T5, where this constraint reduced performance. For machine translation, the autoencoder uses mT5-base, jointly fine-tunes the pretrained language model and autoencoding modules, and uses one layer in each learned module.

  3. Knowl 3 — Latent denoising uses a transformer, velocity prediction, and self-conditioning

    model/method

    For a clean autoencoder latent x∈Rℓ×daex\in\mathbb{R}^{\ell\times d_{\rm ae}}, LD4LG constructs a noisy latent zt=αtx+1−αtϵz_t=\sqrt{\alpha_t}x+\sqrt{1-\alpha_t}\epsilon, where t∈[0,1]t\in[0,1] is the diffusion time, ϵ\epsilon is independent standard Gaussian noise of the same shape as xx, and the default schedule is αt=cos⁡2(πt/2)\alpha_t=\cos^2(\pi t/2). The denoiser uses velocity prediction, with target v=αtϵ−1−αtxv=\sqrt{\alpha_t}\epsilon-\sqrt{1-\alpha_t}x. It is a 12-layer, 768-dimensional pre-LayerNorm transformer with learned positional encodings, GeGLU activations, dense connections, and time conditioning. The default sampler is DDPM with 250 denoising steps; final text is produced by the reconstruction network and decoder using beam search with four beams.

    The denoiser uses self-conditioning: it can receive its previous estimate of the clean latent concatenated with the noisy latent. During training, with probability 0.50.5 it predicts without a previous estimate; otherwise it first makes an estimate without self-conditioning, then makes a second prediction conditioned on the stop-gradient version of that estimate and computes the loss from the second prediction. When no estimate is supplied, a learned embedding is concatenated instead. At inference, the initial estimate is made without self-conditioning and subsequent denoising steps can use the preceding estimate.

  4. Knowl 4 — Class labels guide generation across AG News topics

    model/method

    For class-conditional generation, LD4LG adds a learned embedding for the class label to the denoiser’s time embedding. Training replaces the ground-truth label with a learned null label with probability 0.10.1, retaining the ability to generate unconditionally; at generation time, a chosen label conditions the sampled text. On AG News, where the classes are World, Sports, Business, and Sci/Tech, the highest MAUVE scores occur when the requested class matches the evaluation class. The aligned MAUVE scores are 0.8420.842, 0.8450.845, 0.7520.752, and 0.8130.813 for LD4LG with BART-base, and 0.8090.809, 0.8360.836, 0.7650.765, and 0.8430.843 with FLAN-T5-base, respectively. LD4LG is more consistent than the conditional GPT-2 baseline on the similar Business and Sci/Tech classes: their aligned scores are 0.6290.629 and 0.6970.697 for GPT-2, compared with 0.7520.752 and 0.8130.813 for BART-based LD4LG and 0.7650.765 and 0.8430.843 for FLAN-T5-based LD4LG. The reported memorization measure is the fraction of generated four-grams found in training data; GPT-2 has higher memorization than either LD4LG model for all four classes.

  5. Knowl 5 — Source-conditioned diffusion supports summarization, paraphrasing, and translation

    model/method

    For a source–target pair (wsrc,wtrg)(w_{\rm src},w_{\rm trg}), LD4LG encodes and compresses the target into a latent and trains the denoiser to predict that latent while conditioning on features of the source. The denoiser adds a cross-attention layer after each self-attention layer; the source features come from a frozen language encoder. The experiments use the autoencoder’s pretrained encoder by default, and use a frozen mT5-XL encoder for machine-translation conditioning. Classifier-free guidance is trained by dropping source conditioning with probability 0.10.1 and attending to a learned embedding in its place. At sampling, the conditional and unconditional predictions are combined with guidance weight 2.02.0 for the sequence-to-sequence tasks.

    LD4LG can generate several candidates for each source using independent Gaussian starting latents. Minimum Bayes Risk decoding selects, from five candidates, the one with the lowest average loss against the other candidates; the paper reports this as MBR-5. It also reports Oracle-5, which selects the candidate with the lowest loss against the known reference and is therefore an upper bound rather than a deployable selection procedure. For the sequence-to-sequence experiments, the diffusion regression loss is L1L_1; the authors report that this choice improved fidelity metrics at the cost of some diversity.

  6. Knowl 6 — The compact autoencoder reconstructs text with near-perfect fidelity

    empirical result

    On held-out ROCStories and AG News examples, LD4LG’s language autoencoders preserve reconstruction quality while reducing the encoder feature representation from up to 49,15249{,}152 hidden units to 2,0482{,}048—a 24×24\times compression. Rouge values are Rouge-1/2/L, and BLEU is reported as a percentage.

    Method Latent dimensions Hidden units ROCStories Rouge-1/2/L; BLEU AG News Rouge-1/2/L; BLEU
    BART-base L×768L\times768 ≤49,152\leq49{,}152 98.9/98.2/98.8; 97.5 99.6/99.4/99.6; 98.6
    BART-base autoencoder 32×6432\times64 2,048 99.2/98.5/99.2; 97.6 99.7/99.4/99.7; 98.8
    FLAN-T5-base L×768L\times768 ≤49,152\leq49{,}152 21.5/11.8/19.4; 0.7 63.6/53.0/59.6; 42.3
    FLAN-T5-base autoencoder 32×6432\times64 2,048 98.4/96.9/98.4; 95.8 99.1/98.3/99.1; 96.8

    The BART-base autoencoder slightly improves on BART’s already strong copying behavior, while the learned modules turn FLAN-T5-base—which does not ordinarily copy the input—into an effective autoencoder. The paper reports similarly near-perfect reconstructions on XSum, QQP, and WMT14 English and German.

  7. Knowl 7 — Unconditional latent diffusion improves distribution matching over prior text diffusion

    empirical result

    Unconditional generation was evaluated on ROCStories and AG News. MAUVE compares generated and reference text distributions; perplexity (Ppl) is measured with GPT-2-Large; diversity (Div) is the paper’s n-gram diversity measure; memorization (Mem) is the proportion of generated four-grams present in training data. Entries below are reported mean ±\pm standard deviation.

    Method and metric ROCStories AG News
    Reference: MAUVE / Ppl / Div / Mem .951±\pm.007 / 21.1±\pm.3 / .414±\pm.003 / .362±\pm.003 .951±\pm.014 / 43.6±\pm1.2 / .658±\pm.002 / .385±\pm.005
    Diffusion-LM, 2,000 steps .043±\pm.006 / 47.3±\pm.6 / .128±\pm.002 / .434±\pm.002 .012±\pm.001 / 67.1±\pm1.2 / .043±\pm.002 / .086±\pm.006
    LD4LG BART-base, 250 steps .716±\pm.019 / 30.6±\pm.5 / .331±\pm.005 / .441±\pm.004 .866±\pm.016 / 100.6±\pm2.9 / .540±\pm.006 / .293±\pm.001
    LD4LG FLAN-T5-base, 250 steps .481±\pm.007 / 37.5±\pm.4 / .389±\pm.002 / .387±\pm.002 .859±\pm.020 / 122.0±\pm3.9 / .624±\pm.008 / .221±\pm.003
    GPT-2-medium .788±\pm.025 / 20.0±\pm.2 / .372±\pm.002 / .688±\pm.006 .820±\pm.012 / 37.3±\pm1.1 / .532±\pm.017 / .829±\pm.005

    Both LD4LG variants substantially exceed Diffusion-LM in MAUVE with 250 rather than 2,000 sampling steps. GPT-2 has lower perplexity, but the authors note that evaluating with a GPT-2 language model may favor the fine-tuned GPT-2 baseline; GPT-2 also has much greater memorization than LD4LG, especially on AG News. On ROCStories, compressing BART features improves MAUVE from 0.605±0.0240.605\pm0.024 for BART-Diffusion to 0.716±0.0190.716\pm0.019 and reaches BART-Diffusion’s peak validation MAUVE in 3.86×3.86\times less time. Removing self-conditioning lowers LD4LG’s ROCStories MAUVE to 0.480±0.0180.480\pm0.018 from 0.716±0.0190.716\pm0.019 and raises perplexity to 79.3±1.079.3\pm1.0 from 30.6±0.530.6\pm0.5, while increasing diversity and lowering memorization.

  8. Knowl 8 — LD4LG is competitive with fine-tuning and prior diffusion on QQP and XSum

    empirical result

    On QQP paraphrasing and XSum summarization, LD4LG was evaluated with random sampling, MBR selection from five candidates, and Oracle selection from five candidates. The table gives Rouge-1/2/L and BERTScore. The main comparison is that LD4LG substantially exceeds DiffuSeq, particularly on the challenging XSum task; LD4LG with MBR-5 is competitive with fine-tuned encoder–decoder models, while Oracle-5 demonstrates that its sampled candidate sets can contain stronger outputs than ordinary selection retrieves.

    Dataset Method and selection Rouge-1/2/L BERTScore
    QQP DiffuSeq, random 55.2/29.2/52.7 82.4
    QQP LD4LG BART-base, random 62.6/39.0/60.3 85.8
    QQP LD4LG FLAN-T5-base, random 62.1/38.4/59.7 85.8
    QQP LD4LG BART-base, MBR-5 63.3/40.3/61.1 86.2
    QQP LD4LG FLAN-T5-base, MBR-5 63.0/39.7/60.7 86.1
    QQP LD4LG BART-base, Oracle-5 68.0/46.6/66.0 87.2
    QQP LD4LG FLAN-T5-base, Oracle-5 67.8/46.0/65.7 87.2
    QQP BART-base, fine-tuned beam search 61.9/39.0/59.5 85.5
    QQP FLAN-T5-base, fine-tuned beam search 63.0/40.1/60.5 86.2
    XSum DiffuSeq, random 18.9/1.3/13.6 46.8
    XSum LD4LG BART-base, random 37.6/15.5/30.8 74.1
    XSum LD4LG FLAN-T5-base, random 38.1/15.9/31.2 74.8
    XSum LD4LG BART-base, MBR-5 38.2/16.2/31.5 74.5
    XSum LD4LG FLAN-T5-base, MBR-5 38.7/16.6/31.9 75.2
    XSum LD4LG BART-base, Oracle-5 42.4/19.4/36.4 75.3
    XSum LD4LG FLAN-T5-base, Oracle-5 43.0/20.0/37.2 76.1
    XSum Fine-tuned BART-base, beam search 39.9/18.0/32.6 75.6
    XSum Fine-tuned FLAN-T5-base, beam search 39.7/17.7/32.3 75.3

    LD4LG’s MBR-5 results narrowly exceed fine-tuned models on QQP, whereas fine-tuned models are slightly stronger on XSum. On XSum, the LD4LG Oracle-5 scores also exceed the reported GENIE-with-pretraining Oracle-5 Rouge scores of 41.2/19.1/33.4.

  9. Knowl 9 — Latent diffusion improves over some machine-translation diffusion baselines

    empirical result

    On WMT14 English–German, LD4LG uses an mT5-base autoencoder and reports SacreBLEU for both translation directions. The comparison below includes random generation where reported and candidate selection by MBR. LD4LG exceeds Diffusion-LM and CDCD, but does not reach DINOISER’s MBR-5 scores.

    Method Selection English→\toGerman SacreBLEU German→\toEnglish SacreBLEU
    CDCD Random 19.3 24.9
    LD4LG mT5-base Random 21.4 26.2
    Diffusion-LM MBR-5 15.3 17.3
    CDCD MBR-10 19.7 25.4
    DINOISER MBR-5 24.3 28.8
    LD4LG mT5-base MBR-5 22.4 27.0

    These results demonstrate that LD4LG can use pretrained multilingual language models for machine translation, while leaving a gap to the strongest reported diffusion baseline in this comparison.

  10. Knowl 10 — Iterative sampling and candidate selection remain limitations

    limitation

    LD4LG requires iterative diffusion sampling and uses 250 denoising steps in its default setup, so generation remains slow relative to one-pass decoding; the paper identifies faster diffusion sampling as future work. In sequence-to-sequence generation, Oracle-5 substantially outperforms MBR-5, showing that strong candidates can be present in the sampled set even when the paper’s MBR selection procedure does not identify the best one. The authors therefore identify improved sampling and candidate reranking as open needs, particularly for summarization and machine translation.

Coverage note — The paper’s qualitative sample galleries and detailed training hyperparameter tables are omitted because they illustrate behavior or document implementation settings without adding a separate load-bearing contribution beyond the methods and benchmark results captured here.

References

  1. 1.Chi-square Distribution, pages 70–72. Springer New York, New York, NY, 2008. ISBN 978-0-387-32833-1. doi: 10.1007/978-0-387-32833-1_54. URL https://doi.org/10.1007/978-0-387-32833-1_54.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35: 23716–23736, 2022.
  3. 3.Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. Structured denoising diffusion models in discrete state-spaces. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021. URL https://openreview.net/forum?id=h7-XixPCAL.
  4. 4.Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. All are worth words: A vit backbone for diffusion models. In CVPR, 2023.
  5. 5.Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Ale s Tamchyna. Findings of the 2014 workshop on statistical machine translation. In Proceedings of the Ninth Workshop on Statistical Machine Translation, pages 12–58, Baltimore, Maryland, USA, June 2014. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W/W14/W14-3302.
  6. 6.Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan. Wavegrad: Estimating gradients for waveform generation. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=NsMLjcFaO8O.
  7. 7.Ting Chen. On the importance of noise scheduling for diffusion models. arXiv preprint arXiv:2301.10972, 2023.
  8. 8.Ting Chen, Ruixiang Zhang, and Geoffrey Hinton. Analog bits: Generating discrete data using diffusion models with self-conditioning. arXiv preprint arXiv:2208.04202, 2022.
  9. 9.Zihang Chen, Hongbo Zhang, Xiaoji Zhang, and Leqi Zhao. Quora question pairs. 2017.
  10. 10.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  11. 11.Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. arXiv preprint arXiv:2302.05442, 2023.
  12. 12.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 8780–8794. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper/2021/file/49ad23d1ec9fa4bd8d77d02681df5cfa-Paper.pdf.
  13. 13.Sander Dieleman, Laurent Sartran, Arman Roshannai, Nikolay Savinov, Yaroslav Ganin, Pierre H Richemond, Arnaud Doucet, Robin Strudel, Chris Dyer, Conor Durkan, et al. Continuous diffusion for categorical data. arXiv preprint arXiv:2211.15089, 2022.
  14. 14.Jessica Ficler and Yoav Goldberg. Controlling linguistic style aspects in neural language generation. EMNLP 2017, page 94, 2017.
  15. 15.Vaibhava Goel and William J Byrne. Minimum bayes-risk automatic speech recognition. Computer Speech & Language, 14(2):115–135, 2000. ISSN 0885-2308. doi: https://doi.org/10.1006/csla.2000.0138. URL https://www.sciencedirect.com/science/article/pii/S0885230800901384.
  16. 16.Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and Lingpeng Kong. Diffuseq: Sequence to sequence text generation with diffusion models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=jQj-_rLVXsj.
  17. 17.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger, editors, Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014. URL https://proceedings.neurips.cc/paper/2014/file/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf.
  18. 18.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=XPZIaotutsD.
  19. 19.Zhengfu He, Tianxiang Sun, Qiong Tang, Kuanning Wang, Xuanjing Huang, and Xipeng Qiu. DiffusionBERT: Improving generative masked language models with diffusion models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4521–4534, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.248. URL https://aclanthology.org/2023.acl-long.248.
  20. 20.Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  21. 21.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  22. 22.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. CoRR, abs/2006.11239, 2020. URL https://arxiv.org/abs/2006.11239.
  23. 23.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  24. 24.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rygGQyrFvH.
  25. 25.Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré, and Max Welling. Argmax flows and multinomial diffusion: Towards non-autoregressive language models. CoRR, abs/2102.05379, 2021. URL https://arxiv.org/abs/2102.05379.
  26. 26.Emiel Hoogeboom, Alexey A. Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans. Autoregressive diffusion models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=Lm8T39vLDTE.
  27. 27.Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. arXiv preprint arXiv:2301.11093, 2023.
  28. 28.Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.
  29. 29.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858, 2019.
  30. 30.Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. Advances in neural information processing systems, 34:21696–21707, 2021.
  31. 31.Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. On density estimation with diffusion models. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, 2021. URL https://openreview.net/forum?id=2LdBqxc1Yv.
  32. 32.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=a-xFK8Ymz5J.
  33. 33.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L Buckley, Jason Phang, Samuel R Bowman, and Ethan Perez. Pretraining language models with human preferences. arXiv preprint arXiv:2302.08582, 2023.
  34. 34.Shankar Kumar and William Byrne. Minimum Bayes-risk decoding for statistical machine translation. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004, pages 169–176, Boston, Massachusetts, USA, May 2 - May 7 2004. Association for Computational Linguistics. URL https://aclanthology.org/N04-1022.
  35. 35.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, 2019. URL https://arxiv.org/abs/1910.13461.
  36. 36.Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto. Diffusion-lm improves controllable text generation, 2022. URL https://arxiv.org/abs/2205.14217.
  37. 37.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004.
  38. 38.Zhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu, Zhihao Fan, Chen Lin, Weizhu Chen, and Nan Duan. Text generation with diffusion language models: A pre-training approach with continuous paragraph denoise. 2023.
  39. 39.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7.
  40. 40.Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Amm anabrolu, and Yejin Choi. QUARK: Controllable text generation with reinforced unlearning. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=5HaIds3ux5O.
  41. 41.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=aBsCjcPu_tE.
  42. 42.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, 2016.
  43. 43.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1206. URL https://aclanthology.org/D18-1206.
  44. 44.Shashi Narayan, Shay B Cohen, and Mirella Lapata. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. arXiv preprint arXiv:1808.08745, 2018.
  45. 45.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8162–8171. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/nichol21a.html.
  46. 46.William Peebles and Saining Xie. Scalable diffusion models with transformers. arXiv preprint arXiv:2212.09748, 2022.
  47. 47.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. Mauve: Measuring the gap between neural text and human text using divergence frontiers. Advances in Neural Information Processing Systems, 34:4816–4828, 2021.
  48. 48.Matt Post. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium, October 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-6319. URL https://aclanthology.org/W18-6319.
  49. 49.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  50. 50.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  51. 51.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021. URL https://arxiv.org/abs/2112.10752.
  52. 52.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022.
  53. 53.Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. Palette: Image-to-image diffusion models. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–10, 2022.
  54. 54.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo-Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=08Yk-n5l2Al.
  55. 55.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 36479–36494. Curran Associates, Inc., 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/ec795aeadae0b7d230fa35cbaf04c041-Paper-Conference.pdf.
  56. 56.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  57. 57.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=TIdIXIpzhoI.
  58. 58.Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf. Mo\ˆ usai: Text-to-music generation with long-context latent diffusion. arXiv preprint arXiv:2301.11757, 2023.
  59. 59.Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  60. 60.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642, 2013.
  61. 61.Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics, 2015. URL https://arxiv.org/abs/1503.03585.
  62. 62.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. CoRR, abs/2010.02502, 2020. URL https://arxiv.org/abs/2010.02502.
  63. 63.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 32, 2019.
  64. 64.Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon. Solving inverse problems in medical imaging with score-based generative models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=vaRCHVj0uGI.
  65. 65.Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models, 2023.
  66. 66.Robin Strudel, Corentin Tallec, Florent Altché, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, et al. Self-conditioned embedding diffusion for text generation. arXiv preprint arXiv:2211.04236, 2022.
  67. 67.Robin Strudel, Corentin Tallec, Florent Altché, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, and Rémi Leblond. Self-conditioned embedding diffusion for text generation, 2022. URL https://arxiv.org/abs/2211.04236.
  68. 68.Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. A contrastive framework for neural text generation. arXiv preprint arXiv:2202.06417, 2022.
  69. 69.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6000–6010, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN 9781510860964.
  70. 70.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In Hal Daumé III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 10524–10533. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/v119/xiong20b.html.
  71. 71.Minkai Xu, Alexander Powers, Ron Dror, Stefano Ermon, and Jure Leskovec. Geometric latent diffusion models for 3d molecule generation. arXiv preprint arXiv:2305.01140, 2023.
  72. 72.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.41. URL https://aclanthology.org/2021.naacl-main.41.
  73. 73.Jiasheng Ye, Zaixiang Zheng, Yu Bao, Lihua Qian, and Mingxuan Wang. Dinoiser: Diffused conditional sequence learning by manipulating noises. arXiv preprint arXiv:2302.10025, 2023.
  74. 74.Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  75. 75.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019.
  76. 76.Lin Zheng, Jianbo Yuan, Lei Yu, and Lingpeng Kong. A reparameterized discrete diffusion model for text generation. arXiv preprint arXiv:2302.05737, 2023.

Citation

MLA
Lovelace, J., et al. “Latent Diffusion for Language Generation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 56998–7025, https://proceedings.neurips.cc/paper_files/paper/2023/file/b2a2bd5d5051ff6af52e1ef60aefd255-Paper-Conference.pdf.
APA
Lovelace, J., Kishore, V., Wan, C., Shekhtman, E., & Weinberger, K. (2023). Latent Diffusion for Language Generation. Advances in Neural Information Processing Systems, 36, 56998–57025. https://proceedings.neurips.cc/paper_files/paper/2023/file/b2a2bd5d5051ff6af52e1ef60aefd255-Paper-Conference.pdf
Chicago
Lovelace, J., V. Kishore, C. Wan, E. Shekhtman, and K. Weinberger. 2023. “Latent Diffusion for Language Generation”. Advances in Neural Information Processing Systems 36: 56998–57025. https://proceedings.neurips.cc/paper_files/paper/2023/file/b2a2bd5d5051ff6af52e1ef60aefd255-Paper-Conference.pdf.
Harvard
Lovelace, J. et al. (2023) “Latent Diffusion for Language Generation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 56998–57025. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/b2a2bd5d5051ff6af52e1ef60aefd255-Paper-Conference.pdf.
Vancouver
1. Lovelace J, Kishore V, Wan C, Shekhtman E, Weinberger K (2023) Latent Diffusion for Language Generation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 56998–57025

BibTeX

@inproceedings{lovelace2023latent,
  title = {Latent Diffusion for Language Generation},
  author = {Lovelace, Justin and Kishore, Varsha and Wan, Chao and Shekhtman, Eliot and Weinberger, Kilian},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {56998-57025},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/b2a2bd5d5051ff6af52e1ef60aefd255-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission