The Diffusion Duality

Subham Sekhar SahooJustin DeschenauxAaron GokaslanGuanghan WangJustin T. ChiuVolodymyr Kuleshov

article2025ICML123 citations

Establishes a theoretical link proving uniform-state discrete diffusion emerges from continuous Gaussian processes, enabling curriculum training and discrete consistency distillation that doubles training speed and accelerates text generation by two orders of magnitude.

Listen

Modern generative language models typically generate text sequentially from left to right, which limits processing speed and prevents models from revising earlier errors. Discrete diffusion models provide an alternative approach that can update and refine entire text sequences in parallel, offering inherent self-correction capabilities. However, discrete diffusion models have historically lagged behind standard sequential models and masked diffusion models in terms of output quality and training efficiency, while also requiring hundreds of iterative steps during inference.

The article demonstrates that discrete diffusion processes naturally emerge from continuous Gaussian diffusion processes through mathematical mapping. Building on this theoretical insight, the authors introduce Duo, a framework designed to transfer proven continuous diffusion techniques to discrete text generation to significantly accelerate both model training and text generation.

To evaluate this framework, the authors conducted extensive experiments using 170-million-parameter transformer architectures trained on major standard benchmarks, including the One Billion Word Benchmark and OpenWebText. The approach incorporates two main technical innovations: a curriculum learning strategy that gradually transitions training inputs from smooth continuous approximations to hard discrete tokens, and a discrete consistency distillation algorithm that creates deterministic trajectories to compress multi-step generation into very few steps.

The experimental findings show significant improvements across training, quality, and generation speed. First, the curriculum learning strategy halved the training time needed to reach target performance by substantially reducing gradient variance. Second, the resulting models outperformed standard autoregressive models on zero-shot perplexity across 3 out of 7 standard benchmark datasets. Third, the consistency distillation technique, combined with a greedy sampling refinement, reduced the required generation steps from 1,024 down to just 8 to 16 steps—an acceleration of up to two orders of magnitude—while outperforming competing distilled diffusion models in the few-step regime.

These results demonstrate that bridging continuous and discrete diffusion principles removes major computational and quality barriers that previously hindered non-autoregressive text models. In practical terms, accelerating inference by up to two orders of magnitude dramatically lowers computing costs, reduces latency, and unlocks viable deployment for high-throughput language generation applications where real-time speed and error correction are paramount.

Organizations developing or deploying high-speed language generation systems should consider piloting discrete diffusion frameworks with consistency distillation as a high-throughput alternative to traditional sequential models. For broader adoption, practitioners should evaluate the trade-off between sampling steps and generation diversity depending on whether the downstream task prioritizes peak speed or rich creative variation.

The study has certain limitations, as evaluations were focused on 170-million-parameter models and established text benchmarks rather than modern web-scale foundation models with tens of billions of parameters. Nonetheless, the mathematical proofs and empirical gains provide high confidence that continuous diffusion techniques can reliably enhance discrete sequence generation.

arXiv: 2506.10892
Cover for The Diffusion Duality

Abstract

Uniform-state discrete diffusion models hold the promise of fast text generation due to their inherent ability to self-correct. However, they are typically outperformed by autoregressive models and masked diffusion models. In this work, we narrow this performance gap by leveraging a key insight: Uniform-state diffusion processes naturally emerge from an underlying Gaussian diffusion. Our method, Duo, transfers powerful techniques from Gaussian diffusion to improve both training and sampling. First, we introduce a curriculum learning strategy guided by the Gaussian process, doubling training speed by reducing variance. Models trained with curriculum learning surpass autoregressive models in zero-shot perplexity on 3 of 7 benchmarks. Second, we present Discrete Consistency Distillation, which adapts consistency distillation from the continuous to the discrete setting. This algorithm unlocks few-step generation in diffusion language models by accelerating sampling by two orders of magnitude. We provide the code and model checkpoints on the project page: https://s-sahoo.com/duo

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Discrete Diffusion Models
  • 2.2. Gaussian Diffusion Models
  • 2.3. Consistency Distillation
  • 3. The Diffusion Duality
  • 4. Applications
  • 4.1. Faster Training using Curriculum Learning
  • 4.1.1. DISCRETE NELBO WITH GAUSSIAN LATENTS
  • 4.1.2. LOW-VARIANCE TRAINING LOSS
  • 4.2. Discrete Consistency Distillation
  • 5. Experiments
  • 5.1. Improved Training
  • 5.2. Improved Sampling
  • 6. Related Work
  • 7. Conclusion
  • Impact Statement
  • References
  • Appendices
  • A. The Diffusion Duality
  • A.1. Discrete Marginals
  • A.2. Time Evolution of Probability Densities of Discrete Marginals
  • A.3. Relationship between Gaussian and Discrete Diffusion Trajectories
  • A.4. Gaussian ELBO vs Discrete ELBO
  • A.5. Negative Evidence Lower Bound
  • A.6. Reverse Process Visualizations
  • B. Additional Proofs
  • B.1. ELBO Equivalence
  • B.2. Discrete Consistency Distillation
  • B.2.1. OPTIMAL GAUSSIAN PF-ODES
  • B.2.2. DISCRETE CONSISTENCY DISTILLATION ABLATION
  • C. Experimental details
  • C.1. Plaid Baseline
  • C.2. Denoising Model
  • C.3. Low Discrepancy Sampler
  • C.4. Likelihood Evaluation
  • C.5. Language Modeling
  • C.6. Zeroshot Likelihood
  • C.7. Curriculum Learning
  • C.8. Distillation Experiments
  • D. Additional Experiments
  • D.1. Gradient Variance and Loss Variance
  • D.2. Tau Ablations
  • D.3. Sample Quality Base Model
  • D.4. Duo Ablations
  • D.5. Distillation
  • D.6. Samples
  • D.6.1. DUO
  • D.6.2. DISTILLED DUO VIA DDT
  • D.6.4. DISTILLED MDLM VIA SDTT

Knowls

  1. Knowl 1 — Diffusion Duality via the Argmax Pushforward

    theoretical result

    A continuous Gaussian diffusion process on categorical data representations naturally transforms into a discrete Uniform-state Diffusion Model (USDM) under the arg⁡max⁡\arg\max operator.

    Let clean data x∈V={v∈{0,1}K:∑i=1Kvi=1}x \in \mathcal{V} = \{v \in \{0, 1\}^K : \sum_{i=1}^K v_i = 1\} be a one-hot column vector of dimension KK. In Gaussian diffusion, the forward marginal distribution of continuous latents wt∈RKw_t \in \mathbb{R}^K at diffusion timestep t∈[0,1]t \in [0, 1] is given by:

    q~t(wt∣x;α~t)=N(α~tx,(1−α~t2)IK)\tilde{q}_t(w_t \mid x; \tilde{\alpha}_t) = \mathcal{N}\left(\tilde{\alpha}_t x, (1 - \tilde{\alpha}_t^2) I_K\right)

    where α~t∈[0,1]\tilde{\alpha}_t \in [0, 1] is a monotonically decreasing function of tt with α~0=1\tilde{\alpha}_0 = 1 and α~1=0\tilde{\alpha}_1 = 0, and IKI_K is the K×KK \times K identity matrix.

    Applying the coordinate-wise arg⁡max⁡\arg\max mapping arg⁡max⁡(w)=arg⁡max⁡z∈Vz⊤w\arg\max(w) = \arg\max_{z \in \mathcal{V}} z^\top w defines discrete latents zt=arg⁡max⁡(wt)∈Vz_t = \arg\max(w_t) \in \mathcal{V}. The marginal distribution of ztz_t conditioned on xx under the pushforward operation [arg⁡max⁡]∗q~t[ \arg\max]_* \tilde{q}_t is identically a categorical distribution of a uniform-state discrete diffusion process:

    qt(zt∣x;αt)=[arg⁡max⁡]∗q~t(wt∣x;α~t)=Cat(zt;αtx+(1−αt)1K1)q_t(z_t \mid x; \alpha_t) = [\arg\max]_* \tilde{q}_t(w_t \mid x; \tilde{\alpha}_t) = \text{Cat}\left(z_t; \alpha_t x + (1 - \alpha_t)\frac{1}{K}\mathbf{1}\right)

    where 1={1}K\mathbf{1} = \{1\}^K, and the discrete diffusion parameter αt∈[0,1]\alpha_t \in [0, 1] is related to the Gaussian parameter α~t\tilde{\alpha}_t via the Diffusion Transformation operator T:[0,1]→[0,1]\mathcal{T} : [0, 1] \to [0, 1]:

    αt=T(α~t)=KK−1[∫−∞∞ϕ(z−α~t1−α~t2)ΦK−1(z)dz−1K]\alpha_t = \mathcal{T}(\tilde{\alpha}_t) = \frac{K}{K - 1} \left[ \int_{-\infty}^\infty \phi\left(z - \frac{\tilde{\alpha}_t}{\sqrt{1 - \tilde{\alpha}_t^2}}\right) \Phi^{K-1}(z) dz - \frac{1}{K} \right]

    where ϕ(z)=12πexp⁡(−z2/2)\phi(z) = \frac{1}{\sqrt{2\pi}}\exp(-z^2/2) is the standard normal probability density function and Φ(z)=∫−∞zϕ(u)du\Phi(z) = \int_{-\infty}^z \phi(u)du is its cumulative distribution function.

  2. Knowl 2 — Time Evolution ODE of Discretized Gaussian Marginals

    theoretical result

    When a continuous vector wt∈RKw_t \in \mathbb{R}^K undergoes Gaussian forward diffusion, the marginal probability mass function Pt(⋅∣x)=Cat(⋅;T(α~t)x+(1−T(α~t))1K1)P_t(\cdot \mid x) = \text{Cat}(\cdot; \mathcal{T}(\tilde{\alpha}_t)x + (1 - \mathcal{T}(\tilde{\alpha}_t))\frac{1}{K}\mathbf{1}) of the discretized latent zt=arg⁡max⁡(wt)z_t = \arg\max(w_t) satisfies the characteristic linear ordinary differential equation (ODE) of a continuous-time Markovian Uniform-State Discrete Diffusion Model:

    ddtPt=QtPt\frac{d}{dt} P_t = Q_t P_t

    where the state transition matrix Qt∈RK×KQ_t \in \mathbb{R}^{K \times K} is given by:

    Qt=−T′(α~t)KT(α~t)[11⊤−KIK]Q_t = -\frac{\mathcal{T}'(\tilde{\alpha}_t)}{K \mathcal{T}(\tilde{\alpha}_t)} \left[ \mathbf{1}\mathbf{1}^\top - K I_K \right]

    Here, T′(α~t)\mathcal{T}'(\tilde{\alpha}_t) denotes the time derivative ddtT(α~t)\frac{d}{dt}\mathcal{T}(\tilde{\alpha}_t), 1\mathbf{1} is the all-ones vector in RK\mathbb{R}^K, and IKI_K is the K×KK \times K identity matrix.

    While the marginal distributions of the discretized latents evolve identically to a Markovian uniform-state discrete diffusion process, an individual continuous Gaussian diffusion trajectory projected through arg⁡max⁡\arg\max does not generally constitute a discrete diffusion trajectory because the pushforward transition kernel [arg⁡max⁡]∗q~t∣s(⋅∣ws)[\arg\max]_* \tilde{q}_{t \mid s}(\cdot \mid w_s) does not equal the discrete transition kernel qt∣s(⋅∣zs)q_{t \mid s}(\cdot \mid z_s) for intermediate steps 0≤s<t≤10 \le s < t \le 1.

  3. Knowl 3 — Relative Tightness of Discrete vs Gaussian Evidence Lower Bounds

    theoretical result

    For any categorical data distribution qdataq_{\text{data}} over one-hot vectors x∈Vx \in \mathcal{V}, the variational Evidence Lower Bound (ELBO) defined directly on the discrete uniform-state diffusion process is strictly tighter than the ELBO defined on the underlying continuous Gaussian diffusion process.

    Let pθp_\theta denote the reverse denoiser in the Gaussian diffusion space, and let pˉθ:=[arg⁡max⁡]∗pθ\bar{p}_\theta := [\arg\max]_* p_\theta denote the induced discrete denoiser. By the data processing inequality for Kullback-Leibler (KL) divergence under the Markov kernel k=arg⁡max⁡k = \arg\max, for any distributions pp and qq, DKL([k]∗p,[k]∗q)≤DKL(p,q)D_{\text{KL}}([k]_* p, [k]_* q) \le D_{\text{KL}}(p, q). Consequently:

    DKL(q,pˉθ)=DKL([arg⁡max⁡]∗q~,[arg⁡max⁡]∗pθ)≤DKL(q~,pθ)D_{\text{KL}}(q, \bar{p}_\theta) = D_{\text{KL}}\left([\arg\max]_* \tilde{q}, [\arg\max]_* p_\theta\right) \le D_{\text{KL}}(\tilde{q}, p_\theta)

    which establishes the hierarchy of lower bounds on the true log-likelihood log⁡pθ(x)\log p_\theta(x):

    log⁡pθ(x)≥ELBO(q,pˉθ;x)≥ELBO(q~,pθ;x)\log p_\theta(x) \ge \text{ELBO}(q, \bar{p}_\theta; x) \ge \text{ELBO}(\tilde{q}, p_\theta; x)

    Because the discrete ELBO provides a tighter lower bound, training denoising models on discrete categorical latents is theoretically preferred over training purely in Gaussian continuous space.

  4. Knowl 4 — Rao-Blackwellized Negative ELBO Objective for USDMs

    equation

    For sequence modeling over discrete tokens x1:L=(x1,…,xL)∈VLx^{1:L} = (x^1, \dots, x^L) \in \mathcal{V}^L, the Negative Evidence Lower Bound (NELBO) for Uniform-State Discrete Diffusion Models (USDMs) decomposes into a sum of token-level losses:

    NELBO(q,pθ;x1:L)=Et∼U[0,1],zt1:L∼qt(zt1:L∣x1:L;αt)[∑ℓ=1LfDuo(ztℓ,xθℓ(zt1:L,t),αt;xℓ)]\text{NELBO}(q, p_\theta; x^{1:L}) = \mathbb{E}_{t \sim U[0, 1], z_t^{1:L} \sim q_t(z_t^{1:L} \mid x^{1:L}; \alpha_t)} \left[ \sum_{\ell=1}^L f_{\text{Duo}}\left(z_t^\ell, x_\theta^\ell(z_t^{1:L}, t), \alpha_t; x^\ell\right) \right]

    where xθ:VL×[0,1]→ΔLx_\theta : \mathcal{V}^L \times [0, 1] \to \Delta^L is the denoising network predicting probabilities on the simplex ΔK\Delta^K. The per-token Rao-Blackwellized loss fDuof_{\text{Duo}} analytically computes expectations to reduce estimator variance without materializing KK-dimensional one-hot vectors:

    fDuo(zt,xθ(zt,t),αt;x)=−αt′Kαt[Kxˉi−K(xˉθ)i−(κt1zt=x+1zt≠x)∑j=1Klog⁡(xˉθ)i(xˉθ)j−Kαt1−αtlog⁡((xˉθ)i(xˉθ)m)1zt≠x−((K−1)κt1zt=x−1κt1zt≠x)log⁡κt]f_{\text{Duo}}(z_t, x_\theta(z_t, t), \alpha_t; x) = -\frac{\alpha_t'}{K\alpha_t} \left[ \frac{K}{\bar{x}_i} - \frac{K}{(\bar{x}_\theta)_i} - (\kappa_t \mathbf{1}_{z_t = x} + \mathbf{1}_{z_t \neq x}) \sum_{j=1}^K \log \frac{(\bar{x}_\theta)_i}{(\bar{x}_\theta)_j} - \frac{K\alpha_t}{1 - \alpha_t} \log\left(\frac{(\bar{x}_\theta)_i}{(\bar{x}_\theta)_m}\right) \mathbf{1}_{z_t \neq x} - \left((K - 1)\kappa_t \mathbf{1}_{z_t = x} - \frac{1}{\kappa_t}\mathbf{1}_{z_t \neq x}\right) \log \kappa_t \right]

    where i=arg⁡max⁡j∈[K](zt)ji = \arg\max_{j \in [K]} (z_t)_j is the active index of ztz_t, mm is the active index of clean token xx (xm=1x_m = 1), αt′=ddtαt\alpha_t' = \frac{d}{dt}\alpha_t, and the scalar terms are defined by:

    κt=1−αtKαt+1−αt,xˉi=Kαtxi+(1−αt),(xˉθ)i=Kαt(xθ(zt,t))i+(1−αt)\kappa_t = \frac{1 - \alpha_t}{K\alpha_t + 1 - \alpha_t}, \quad \bar{x}_i = K\alpha_t x_i + (1 - \alpha_t), \quad (\bar{x}_\theta)_i = K\alpha_t (x_\theta(z_t, t))_i + (1 - \alpha_t)

  5. Knowl 5 — Gaussian-Guided Low-Variance Curriculum Learning

    model/method

    By exploiting the diffusion duality equivalence Eqt[fDuo]=Eq~t[fDuo]\mathbb{E}_{q_t}[f_{\text{Duo}}] = \mathbb{E}_{\tilde{q}_t}[f_{\text{Duo}}], USDMs can be trained with continuous Gaussian latents relaxed via tempered softmax to smooth optimization and drastically reduce gradient variance.

    During training, continuous noisy latents wt1:L∼q~t(wt1:L∣x1:L;α~t)w_t^{1:L} \sim \tilde{q}_t(w_t^{1:L} \mid x^{1:L}; \tilde{\alpha}_t) are sampled. The hard arg⁡max⁡\arg\max in the denoiser input is replaced with a temperature-relaxed softmax softmax(wtℓ/τ)\text{softmax}(w_t^\ell / \tau):

    Ltrain=Ex1:L∼qdata,t∼U[β,γ],wt1:L∼q~t[∑ℓ=1LfDuo(ztℓ:=arg⁡max⁡(wtℓ),xθ([softmax(wtℓ′/τ)]ℓ′=1L,t),αt:=T(α~t);xℓ)]\mathcal{L}_{\text{train}} = \mathbb{E}_{x^{1:L} \sim q_{\text{data}}, t \sim U[\beta, \gamma], w_t^{1:L} \sim \tilde{q}_t} \left[ \sum_{\ell=1}^L f_{\text{Duo}}\left( z_t^\ell := \arg\max(w_t^\ell), x_\theta\left( \left[ \text{softmax}(w_t^{\ell'} / \tau) \right]_{\ell'=1}^L, t \right), \alpha_t := \mathcal{T}(\tilde{\alpha}_t); x^\ell \right) \right]

    where τ>0\tau > 0 controls task difficulty by preserving soft continuous signal. Timestep tt is sampled uniformly from a sub-interval [β,γ][ \beta, \gamma] corresponding to discrete diffusion parameters αt=T(α~t)∈[0.05,0.95]\alpha_t = \mathcal{T}(\tilde{\alpha}_t) \in [0.05, 0.95] ([β,γ]=[0.03,0.15][\beta, \gamma] = [0.03, 0.15]), which excludes boundary regions where learning signal derivative αt′≈0\alpha_t' \approx 0.

    In training schedules, τ=0.001\tau = 0.001 is maintained for the initial phase (e.g., first 500k500\text{k} iterations), after which τ=0\tau = 0 is used for the remaining training up to 1M1\text{M} steps to transition the model from soft continuous representations to hard discrete tokens. This curriculum reduces gradient variance by an order of magnitude and accelerates convergence by 2×2\times.

  6. Knowl 6 — Deterministic Discrete Trajectories (PDDT)

    definition

    A Deterministic Discrete Trajectory PDDT(x1:L,ϵ1:L)P_{\text{DDT}}(x^{1:L}, \epsilon^{1:L}) is a deterministic sequence of discrete states connecting clean data x1:L∼qdatax^{1:L} \sim q_{\text{data}} to uniform categorical prior noise across continuous time t∈[0,1]t \in [0, 1].

    Given Gaussian standard normal noise vectors ϵ1:L={ϵℓ∼N(0,IK)}ℓ=1L\epsilon^{1:L} = \{\epsilon^\ell \sim \mathcal{N}(0, I_K)\}_{\ell=1}^L, the continuous Probability Flow ODE (PF-ODE) under an optimal denoiser xθ(wt,t)=xx_\theta(w_t, t) = x defines the continuous trajectory:

    PODE(x1:L,ϵ1:L)={[α~txℓ+1−α~t2ϵℓ]ℓ=1L}t∈[0,1]P_{\text{ODE}}(x^{1:L}, \epsilon^{1:L}) = \left\{ \left[ \tilde{\alpha}_t x^\ell + \sqrt{1 - \tilde{\alpha}_t^2} \epsilon^\ell \right]_{\ell=1}^L \right\}_{t \in [0, 1]}

    Projecting this continuous trajectory into the discrete space using the arg⁡max⁡\arg\max operator yields the deterministic discrete trajectory:

    PDDT(x1:L,ϵ1:L)={[arg⁡max⁡(α~txℓ+1−α~t2ϵℓ)]ℓ=1L}t∈[0,1]P_{\text{DDT}}(x^{1:L}, \epsilon^{1:L}) = \left\{ \left[ \arg\max\left( \tilde{\alpha}_t x^\ell + \sqrt{1 - \tilde{\alpha}_t^2} \epsilon^\ell \right) \right]_{\ell=1}^L \right\}_{t \in [0, 1]}

    Along PDDTP_{\text{DDT}}, each token position ℓ\ell at timestep tt takes one of two values: the clean token xℓx^\ell (when tt is small) or the prior noise token arg⁡max⁡(ϵℓ)\arg\max(\epsilon^\ell) (when tt is near 1). Once a token transitions during reverse time, it remains fixed, serving as a deterministic discrete path suitable for consistency distillation.

  7. Knowl 7 — Discrete Consistency Distillation (DCD)

    algorithm

    Discrete Consistency Distillation (DCD) distills a multi-step uniform-state discrete diffusion teacher model into a few-step student model by aligning their clean data predictions along Deterministic Discrete Trajectories (PDDTP_{\text{DDT}}).

    Input: Dataset DD, learning rate η\eta, distillation rounds NN, iterations per round MM, EMA parameter μ\mu, initial model weights θ\theta, EMA weights θema\theta_{\text{ema}}, initial step size δ\delta
    Output: Distilled model weights θema\theta_{\text{ema}}
    for i=1i = 1 to NN do
        θ−←stopgrad(θ)\theta^- \leftarrow \text{stopgrad}(\theta)
        for j=1j = 1 to MM do
            Sample x1:L∼Dx^{1:L} \sim D, t∼U[0,1]t \sim U[0, 1], and ϵℓ∼N(0,IK)\epsilon^\ell \sim \mathcal{N}(0, I_K) for all ℓ∈{1,…,L}\ell \in \{1, \dots, L\}
            s←max⁡(t−δ,0)s \leftarrow \max(t - \delta, 0)
            zs1:L←[arg⁡max⁡(α~sxℓ+1−α~s2ϵℓ)]ℓ=1Lz_s^{1:L} \leftarrow [\arg\max(\tilde{\alpha}_s x^\ell + \sqrt{1 - \tilde{\alpha}_s^2} \epsilon^\ell)]_{\ell=1}^L
            zt1:L←[arg⁡max⁡(α~txℓ+1−α~t2ϵℓ)]ℓ=1Lz_t^{1:L} \leftarrow [\arg\max(\tilde{\alpha}_t x^\ell + \sqrt{1 - \tilde{\alpha}_t^2} \epsilon^\ell)]_{\ell=1}^L
            LDCD(θ;θ−)←∑ℓ=1LDKL(xθℓ(zt1:L,t),xθ−ℓ(zs1:L,s))\mathcal{L}_{\text{DCD}}(\theta; \theta^-) \leftarrow \sum_{\ell=1}^L D_{\text{KL}}(x_\theta^\ell(z_t^{1:L}, t), x_{\theta^-}^\ell(z_s^{1:L}, s))
            θ←θ−η∇θLDCD(θ;θ−)\theta \leftarrow \theta - \eta \nabla_\theta \mathcal{L}_{\text{DCD}}(\theta; \theta^-)
            θema←stopgrad(μθema+(1−μ)θ)\theta_{\text{ema}} \leftarrow \text{stopgrad}(\mu \theta_{\text{ema}} + (1 - \mu)\theta)
        end for
        δ←2⋅δ\delta \leftarrow 2 \cdot \delta
    end for
    return θema\theta_{\text{ema}}

    In standard implementations, N=5N = 5 rounds are trained with M=10kM = 10\text{k} steps per round, starting with δ=1/512\delta = 1/512 and doubling δ\delta at the end of each round. Teacher targets are generated directly using the student weights θ\theta from the previous round (θ− \theta^-) rather than pre-trained EMA weights, which significantly improves distillation quality.

  8. Knowl 8 — Greedy-Tail Sampling for Discrete Diffusion

    model/method

    Greedy-Tail Sampler is an inference method designed for Uniform-State Discrete Diffusion Models that reduces sample entropy and improves text generation quality.

    In standard ancestral sampling, the reverse process starts from uniform prior noise z1∼Cat(1K1)z_1 \sim \text{Cat}(\frac{1}{K}\mathbf{1}) and iteratively samples zs∼ps∣tθ(⋅∣zt)z_s \sim p_{s \mid t}^\theta(\cdot \mid z_t) for discrete timesteps descending from t=1t = 1 to t=δt = \delta, where δ\delta is the final discretization step size. At the final transition step to clean data t=0t = 0, instead of drawing stochastically via x~∼p0∣δθ(⋅∣zδ)\tilde{x} \sim p_{0 \mid \delta}^\theta(\cdot \mid z_\delta), Greedy-Tail sampling executes deterministic greedy decoding:

    x~=arg⁡max⁡(p0∣δθ(⋅∣zδ))\tilde{x} = \arg\max\left(p_{0 \mid \delta}^\theta(\cdot \mid z_\delta)\right)

    This deterministic final step curbs excess distribution entropy at the end of sampling (analogous to nucleus sampling in autoregressive models), boosting generation perplexity across all sampling step budgets and enabling high sample quality down to T=8T = 8 function evaluations.

  9. Knowl 9 — Language Modeling Perplexity Benchmarks

    data/table

    Models were trained for 1M1\text{M} steps with batch size 512 using a 170M-parameter Diffusion Transformer (DiT) architecture. Duo was trained using Gaussian-guided curriculum learning (with τ=0.001\tau = 0.001 for 500k500\text{k} steps, then τ=0\tau = 0) and evaluated on test perplexity (PPL, ↓\downarrow) against autoregressive, absorbing-state (masked), uniform-state, and Gaussian diffusion models.

    Model LM1B (unpacked) LM1B (packed) OWT
    Autoregressive Transformer 22.3 22.8 17.5
    Diffusion (absorbing state)
    BERT-Mouth - 142.9 -
    D3PM Absorb - 76.9 -
    DiffusionBert - 63.8 -
    SEDD Absorb 32.7 - 24.1
    MDLM 27.0 31.8 23.2
    Diffusion (Uniform-state / Gaussian)
    D3PM Uniform - 137.9 -
    Diffusion-LM - 118.6 -
    SEDD Uniform 40.3 - 29.7
    UDLM 31.3 36.7 27.4
    Duo (Ours) 29.9 33.7 25.2

    Duo achieves state-of-the-art test likelihood among uniform-state and Gaussian diffusion language models, narrowing the performance gap with absorbing-state diffusion (MDLM) to within 2 PPL points on OpenWebText.

  10. Knowl 10 — Zero-Shot Likelihood Generalization across 7 Benchmarks

    data/table

    Models pre-trained on OpenWebText (OWT) for 1M1\text{M} steps were evaluated for zero-shot generalization by measuring test perplexity (PPL, ↓\downarrow) across 7 downstream datasets: Penn Tree Bank (PTB), WikiText, LM1B, LAMBADA, AG News, PubMed, and arXiv.

    Model PTB WikiText LM1B LAMBADA AG News PubMed arXiv
    Autoregressive Transformer 82.05 25.75 51.25 51.28 52.09 49.01 41.73
    Diffusion (absorbing state)
    SEDD Absorb 100.09 34.28 68.20 49.86 62.09 44.53 38.48
    D3PM Absorb 200.82 50.86 138.92 93.47 - - -
    MDLM 95.26 32.83 67.01 47.52 61.15 41.89 37.37
    Diffusion (Uniform / Gaussian)
    SEDD Uniform 105.51 41.10 82.62 57.29 82.64 55.89 50.86
    PLAID 142.60 50.86 91.12 57.28 - - -
    UDLM 112.82 39.42 77.59 53.57 80.96 50.98 44.08
    Duo (Ours) 89.35 33.57 73.86 49.78 67.81 44.48 40.39

    Duo outperforms all uniform-state and continuous Gaussian diffusion baselines across all 7 benchmarks, surpasses SEDD Absorbing on 4 of 7 datasets, and outperforms the Autoregressive Transformer on 3 of 7 zero-shot benchmarks (LAMBADA: 49.78 vs 51.28; PubMed: 44.48 vs 49.01; arXiv: 40.39 vs 41.73).

  11. Knowl 11 — Few-Step Generation Sample Quality and Sampling Acceleration

    empirical result

    Discrete Consistency Distillation (DCD) combined with Greedy-Tail sampling accelerates inference by two orders of magnitude while matching or exceeding the generative quality of the 1024-step base diffusion model on OpenWebText (evaluated via GPT-2 Large Generative Perplexity, Gen PPL ↓\downarrow, and sequence entropy ↑\uparrow).

    1. Base Model Performance: Undistilled Duo achieves better Gen PPL than SEDD Uniform and MDLM across sampling step budgets T∈{8,16,32,64,128,256,512,1024}T \in \{8, 16, 32, 64, 128, 256, 512, 1024\} (e.g., at T=1024T = 1024, Duo Gen PPL is 77.69 vs 99.90 for SEDD Uniform and 104.85 for MDLM).
    2. Distillation Acceleration: After 5 rounds of DCD, distilled Duo evaluated with ancestral sampling matches the 1024-step base model quality in 16 steps (Gen PPL 75.24 vs 77.69, a 64×64\times speedup). Evaluated with the Greedy-Tail sampler, distilled Duo achieves Gen PPL 69.58 (entropy 5.30) in only T=8T = 8 steps (a 128×128\times speedup).
    3. Low-NFE Dominance over Masked Diffusion: In the few-step regime (T≤32T \le 32), distilled Duo outperforms MDLM distilled via SDTT (at T=8T = 8, distilled Duo achieves Gen PPL 111.88 with ancestral and 69.58 with Greedy-Tail, compared to 193.05 for SDTT-distilled MDLM and 830.82 for undistilled MDLM). This advantage stems from the self-correcting property of uniform-state diffusion, which allows tokens to be revised across steps unlike masked diffusion.
  12. Knowl 12 — Ablation of Training Loss Components in Duo

    empirical result

    An ablation study on the LM1B dataset (with sequence packing) evaluates the individual contributions of Duo's two primary training innovations over the baseline Uniform Discrete Language Model (UDLM):

    Method PPL (↓\downarrow)
    Duo 33.7
    w/o Curriculum Learning (CL) 35.0
    w/o Improved Training Loss (fDuof_{\text{Duo}}) 36.7

    The total 3.0-point PPL reduction over UDLM is divided approximately equally between the components:

    • Replacing the original loss formulation with the Rao-Blackwellized loss fDuof_{\text{Duo}} accounts for approximately 1.7 PPL points of improvement by reducing estimation variance and memory overhead.
    • Adding the Gaussian-guided curriculum learning strategy contributes the remaining 1.3 PPL points by easing optimization via continuous relaxations.

Coverage note — No substantial contributed material was omitted; all theoretical foundations (Diffusion Duality, ODEs, ELBO tightness), algorithmic techniques (Rao-Blackwellized NELBO, Curriculum Learning, DCD, Greedy-Tail sampling), and main empirical results (likelihood, zero-shot benchmarks, few-step distillation, and ablations) are fully covered.

References

  1. 1.Anderson, W. J. Continuous-time Markov chains: An applications-oriented approach. Springer Science & Business Media, 2012.
  2. 2.Arriola, M., Sahoo, S. S., Gokaslan, A., Yang, Z., Qi, Z., Han, J., Chiu, J. T., and Kuleshov, V. Block diffusion: Interpolating between autoregressive and diffusion language models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=tyEyYT267x.
  3. 3.Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and Van Den Berg, R. Structured denoising diffusion models in discrete state-spaces. Advances in Neural Information Processing Systems, 34:17981–17993, 2021.
  4. 4.Bengio, Y., Louradour, J., Collobert, R., and Weston, J. Curriculum learning. In International Conference on Machine Learning, 2009. URL https://api.semanticscholar.org/CorpusID:873046.
  5. 5.Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S. W., Fidler, S., and Kreis, K. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22563–22575, 2023.
  6. 6.Campbell, A., Benton, J., De Bortoli, V., Rainforth, T., Deligiannidis, G., and Doucet, A. A continuous time framework for discrete denoising models. Advances in Neural Information Processing Systems, 35:28266–28279, 2022.
  7. 7.Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T. One billion word benchmark for measuring progress in statistical language modeling, 2014.
  8. 8.Chen, T., ZHANG, R., and Hinton, G. Analog bits: Generating discrete data using diffusion models with self-conditioning. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=3itjR9QxFw.
  9. 9.Cohan, A., Dernoncourt, F., Kim, D. S., Bui, T., Kim, S., Chang, W., and Goharian, N. A discourse-aware attention model for abstractive summarization of long documents. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), 2018. doi: 10.18653/v1/n18-2097. URL http://dx.doi.org/10.18653/v1/n18-2097.
  10. 10.Cover, T. and Thomas, J. Elements of Information Theory. Wiley, 2012. ISBN 9781118585771. URL https://books.google.com/books?id=VWq5GG6ycxMC.
  11. 11.CSISZAR, I. Information-type measures of difference of probability distributions and indirect observation. Studia Scientiarum Mathematicarum Hungarica, 2:229–318, 1967. URL https://cir.nii.ac.jp/crid/1571417125811646464.
  12. 12.Deschenaux, J. and Gulcehre, C. Beyond autoregression: Fast llms via self-distillation through time. arXiv preprint arXiv:2410.21035, 2024.
  13. 13.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Dieleman, S., Sartran, L., Roshannai, A., Savinov, N., Ganin, Y., Richemond, P. H., Doucet, A., Strudel, R., Dyer, C., Durkan, C., et al. Continuous diffusion for categorical data. arXiv preprint arXiv:2211.15089, 2022.
  15. 15.Esser, P., Chiu, J., Atighehchian, P., Granskog, J., and Germanidis, A. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7346–7356, 2023.
  16. 16.Frans, K., Hafner, D., Levine, S., and Abbeel, P. One step diffusion via shortcut models. arXiv preprint arXiv:2410.12557, 2024.
  17. 17.Gat, I., Remez, T., Shaul, N., Kreuk, F., Chen, R. T. Q., Synnaeve, G., Adi, Y., and Lipman, Y. Discrete flow matching, 2024. URL https://arxiv.org/abs/2407.15595.
  18. 18.Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019.
  19. 19.Gulrajani, I. and Hashimoto, T. B. Likelihood-based diffusion language models. Advances in Neural Information Processing Systems, 36, 2024.
  20. 20.Han, X., Kumar, S., and Tsvetkov, Y. Ssd-lm: Semi-autoregressive simplex-based diffusion language model for text generation and modular control. arXiv preprint arXiv:2210.17432, 2022.
  21. 21.He, Z., Sun, T., Wang, K., Huang, X., and Qiu, X. Diffusionbert: Improving generative masked language models with diffusion models. arXiv preprint arXiv:2211.15029, 2022.
  22. 22.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  23. 23.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. arXiv:2204.03458, 2022.
  24. 24.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=rygGQyrFvH.
  25. 25.Jang, E., Gu, S., and Poole, B. Categorical reparameterization with gumbel-softmax. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=rkE3y85ee.
  26. 26.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. Advances in neural information processing systems, 35:26565–26577, 2022.
  27. 27.Kingma, D., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. Advances in neural information processing systems, 34:21696–21707, 2021.
  28. 28.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=a-xFK8Ymz5J.
  29. 29.Lee, S., Kreis, K., Veccham, S. P., Liu, M., Reidenbach, D., Peng, Y., Paliwal, S., Nie, W., and Vahdat, A. Genmol: A drug discovery generalist with discrete diffusion. arXiv preprint arXiv:2501.06158, 2025.
  30. 30.Li, X., Thickstun, J., Gulrajani, I., Liang, P. S., and Hashimoto, T. B. Diffusion-lm improves controllable text generation. Advances in Neural Information Processing Systems, 35:4328–4343, 2022.
  31. 31.Liu, C., Fan, W., Liu, Y., Li, J., Li, H., Liu, H., Tang, J., and Li, Q. Generative diffusion models on graphs: Methods and applications. arXiv preprint arXiv:2302.02591, 2023a.
  32. 32.Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., and Plumbley, M. D. Audioldm: text-to-audio generation with latent diffusion models. In Proceedings of the 40th International Conference on Machine Learning, pp. 21450–21474, 2023b.
  33. 33.Lou, A., Meng, C., and Ermon, S. Discrete diffusion language modeling by estimating the ratios of the data distribution. arXiv preprint arXiv:2310.16834, 2023.
  34. 34.Luhman, E. and Luhman, T. Knowledge distillation in iterative generative models for improved sampling speed, 2021. URL https://arxiv.org/abs/2101.02388.
  35. 35.Maddison, C. J., Mnih, A., and Teh, Y. W. The concrete distribution: A continuous relaxation of discrete random variables. In International Conference on Learning Representations, 2017. URL https://openreview.net/forum?id=S1jE5L5gl.
  36. 36.Marcus, M., Santorini, B., and Marcinkiewicz, M. A. Building a large annotated corpus of english: The penn treebank. Computational linguistics, 19(2):313–330, 1993.
  37. 37.Mena, G., Belanger, D., Linderman, S., and Snoek, J. Learning latent permutations with gumbel-sinkhorn networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=Byt3oJ-0W.
  38. 38.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models, 2016.
  39. 39.Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernandez, R. The LAMBADA dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1525–1534, Berlin, Germany, August 2016. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/P16-1144.
  40. 40.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023.
  41. 41.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  42. 42.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020.
  43. 43.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  44. 44.Sahoo, S. S., Arriola, M., Gokaslan, A., Marroquin, E. M., Rush, A. M., Schiff, Y., Chiu, J. T., and Kuleshov, V. Simple and effective masked diffusion language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024a. URL https://openreview.net/forum?id=L4uaAR4ArM.
  45. 45.Sahoo, S. S., Gokaslan, A., Sa, C. D., and Kuleshov, V. Diffusion models with learned adaptive noise. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024b. URL https://openreview.net/forum?id=loMa99A4p8.
  46. 46.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models, 2022. URL https://arxiv.org/abs/2202.00512.
  47. 47.Schiff, Y., Sahoo, S. S., Phung, H., Wang, G., Boshar, S., Dalla-torre, H., de Almeida, B. P., Rush, A. M., PIERROT, T., and Kuleshov, V. Simple guidance mechanisms for discrete diffusion models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=i5MrJ6g5G1.
  48. 48.Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K. Simplified and generalized masked diffusion for discrete data, 2025. URL https://arxiv.org/abs/2406.04329.
  49. 49.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pp. 2256–2265. PMLR, 2015.
  50. 50.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=St1giarCHLP.
  51. 51.Song, Y. and Dhariwal, P. Improved techniques for training consistency models, 2023. URL https://arxiv.org/abs/2310.14189.
  52. 52.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  53. 53.Song, Y., Dhariwal, P., Chen, M., and Sutskever, I. Consistency models, 2023. URL https://arxiv.org/abs/2303.01469.
  54. 54.Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding, 2023.
  55. 55.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  56. 56.Wang, A. and Cho, K. Bert has a mouth, and it must speak: Bert as a markov random field language model. arXiv preprint arXiv:1902.04094, 2019.
  57. 57.Wang, G., Schiff, Y., Sahoo, S., and Kuleshov, V. Remasking discrete diffusion models with inference-time scaling. arXiv preprint arXiv:2503.00307, 2025.
  58. 58.Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7623–7633, 2023.
  59. 59.Yin, T., Gharbi, M., Zhang, R., Shechtman, E., Durand, F., Freeman, W. T., and Park, T. One-step diffusion with distribution matching distillation. In CVPR, 2024.
  60. 60.Zhang, X., Zhao, J. J., and LeCun, Y. Character-level convolutional networks for text classification. In NIPS, 2015.
  61. 61.Zhao, Y., Shi, J., Mackey, L., and Linderman, S. Informed correctors for discrete diffusion models. arXiv preprint arXiv:2407.21243, 2024.
  62. 62.Zheng, K., Lu, C., Chen, J., and Zhu, J. Improved techniques for maximum likelihood estimation for diffusion odes. In International Conference on Machine Learning, pp. 42363–42389. PMLR, 2023.
  63. 63.Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J., and Zhang, Q. Masked diffusion models are secretly time-agnostic masked models and exploit inaccurate categorical sampling. arXiv preprint arXiv:2409.02908, 2024.

Citation

MLA
Sahoo, S. S., et al. “The Diffusion Duality”. arXiv, 2025, http://arxiv.org/abs/2506.10892v3.
APA
Sahoo, S. S., Deschenaux, J., Gokaslan, A., Wang, G., Chiu, J., & Kuleshov, V. (2025). The Diffusion Duality. arXiv. http://arxiv.org/abs/2506.10892v3
Chicago
Sahoo, S. S., J. Deschenaux, A. Gokaslan, G. Wang, J. Chiu, and V. Kuleshov. 2025. “The Diffusion Duality”. arXiv. http://arxiv.org/abs/2506.10892v3.
Harvard
Sahoo, S.S. et al. (2025) “The Diffusion Duality”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2506.10892v3.
Vancouver
1. Sahoo SS, Deschenaux J, Gokaslan A, Wang G, Chiu J, Kuleshov V (2025) The Diffusion Duality. arXiv

BibTeX

@article{sahoo2025the,
  title = {The Diffusion Duality},
  author = {Sahoo, Subham Sekhar and Deschenaux, Justin and Gokaslan, Aaron and Wang, Guanghan and Chiu, Justin and Kuleshov, Volodymyr},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2506.10892v3},
  eprint = {2506.10892}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/