One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

Fan BaoShen NieKaiwen XueChongxuan LiShi PuYaole WangGang YueYue CaoHang SuJun Zhu

article2023ICML228 citations

Proposes UniDiffuser, a single transformer-based framework that captures marginal, conditional, and joint multi-modal distributions to handle diverse generation tasks—including text-to-image, image-to-text, and paired generation—without requiring task-specific models or extra computational overhead.

Listen

Generative artificial intelligence has traditionally relied on specialized, single-task systems that separately handle text-to-image creation, image captioning, or unconditional content generation. While these dedicated frameworks produce high-quality outputs, maintaining distinct architectures for individual tasks creates significant operational redundancy, escalates infrastructure costs, and prevents systems from sharing cross-modal knowledge. A unified framework capable of learning all joint, conditional, and marginal distributions simultaneously across multiple modalities has remained a key challenge due to computational complexity and architectural limits.

The article demonstrates UniDiffuser, a single transformer-based diffusion model designed to capture marginal, conditional, and joint distributions across image and text modalities within a unified framework. It evaluates whether a single general-purpose system can match or exceed the performance of specialized generative models without incurring additional training or inference overhead.

The research team formulated the training process as a unified noise prediction task where perturbation levels vary independently across modalities. By assigning zero noise to a condition modality and maximum noise to ignore an unconditioned modality, the system dynamically switches among generative modes. The implementation employs a 952-million-parameter transformer operating in continuous latent spaces, leveraging pre-trained vision encoders and language decoders. The framework was evaluated on subsets of the open-access LAION-5B dataset comprising hundreds of millions of image-text pairs and benchmarked against standard datasets like MS-COCO.

The evaluation yielded several key findings. First, the unified model achieved a zero-shot Fréchet Inception Distance score of 9.71 on MS-COCO text-to-image synthesis, outperforming existing general-purpose models like Versatile Diffusion (10.09) and specialized systems like DALL·E 2 (10.39) while remaining competitive with Stable Diffusion (8.59). Second, the model consistently surpassed multi-task baselines in cross-modal alignment across both text-to-image and image-to-text generation tasks. Third, it enabled classifier-free guidance during sampling at zero additional training cost by directly incorporating marginal score estimates. Finally, it demonstrated high operational efficiency, generating batches in 19.77 seconds with 48.30 gigabytes of memory, outperforming comparable systems in both processing speed and memory footprint.

These findings indicate that generative workflows can be consolidated into single, scalable foundation models without sacrificing output fidelity. Adopting unified architectures lowers the risks and costs associated with managing multiple task-specific systems, simplifies model maintenance, and expands functionality into complex capabilities such as continuous image-text translation, variation generation, and latent interpolation.

Decision-makers should explore unified diffusion architectures when designing multi-modal infrastructure rather than deploying disparate point solutions. Prior to production deployment, organizations should conduct targeted pilot tests and integrate safety safeguards, including automated watermarking and detection protocols, to mitigate misuse risks such as synthetic misinformation.

The framework presents certain operational limitations. The quality and fluency of generated text remain constrained by noise in the underlying training datasets. Additionally, validating the architecture across more diverse modalities—such as audio, video, and 3D data—requires further empirical study. Overall confidence in the performance and efficiency gains of the unified model remains high for two-modality vision and language deployments.

  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work establishes the foundational denoising diffusion probabilistic framework that UniDiffuser modifies and unifies across multiple modalities.
  • Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). This work introduces Diffusion Transformers (DiTs), providing the core transformer-based diffusion backbone that UniDiffuser adapts to process multi-modal inputs.
  • Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). This paper establishes classifier-free guidance, a fundamental sampling technique used in conditional and joint multi-modal diffusion generation.
  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). This paper investigates diffusion formulations across discrete data spaces, offering essential background for extending diffusion models beyond continuous images to text.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). This work introduces non-Markovian implicit sampling (DDIM), which provides the mathematical mechanics for fast and deterministic diffusion sampling across distributions.
  • Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This paper formalizes the design space, noise schedules, and preconditioning of diffusion models, clarifying the noise prediction parameterization UniDiffuser builds upon.
  • Paper: Multimodal learning with deep Boltzmann machines, Nitish Srivastava et al. (2012). This foundational paper introduces joint, marginal, and conditional generative modeling across paired image and text channels in a unified probabilistic architecture.
Cover for One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

Abstract

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoke models (e.g., Stable Diffusion and DALL·E 2) in representative tasks (e.g., text-to-image generation). Our code is available at https://github.com/thu-ml/unidiffuser.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Method
  • 3.1. UniDiffuser: One Diffusion Fits All Distributions
  • 3.2. Classifier-Free Guidance for Free
  • 4. UniDiffuser on Images and Texts
  • 4.1. Encoding Images and Texts into Latent Space
  • 4.2. Transformer as Joint Noise Prediction Network
  • 5. Related Work
  • 6. Experiments
  • 6.1. Setup
  • 6.2. Main Results
  • 6.3. Data Variation and Gibbs Sampling
  • 6.4. Interpolation between Two Images in the Wild
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. More Examples
  • B. The Training and Sampling Algorithms
  • C. Summary of Classifier-Free Guidance Models
  • D. Details of the U-ViT
  • E. Details of the GPT-2 Text Decoders
  • F. The Interpolation Algorithm
  • G. Comparison of Examples
  • H. Efficiency Comparison
  • I. Licences

Knowls

  1. Knowl 1 — Unified Multi-Modal Diffusion Framework

    model/method

    UniDiffuser unifies the modeling of all distributions over paired multi-modal data (x0,y0)∼q(x0,y0)(x_0, y_0) \sim q(x_0, y_0)—including marginals q(x0)q(x_0) and q(y0)q(y_0), conditionals q(x0∣y0)q(x_0|y_0) and q(y0∣x0)q(y_0|x_0), and the joint distribution q(x0,y0)q(x_0, y_0)—into a single diffusion network ϵθ(xtx,yty,tx,ty)\epsilon_\theta(x_{t_x}, y_{t_y}, t_x, t_y).

    The fundamental insight is that diffusion models learn conditional expectations over injected noise E[ϵx,ϵy∣xtx,yty]\mathbb{E}[\epsilon^x, \epsilon^y \mid x_{t_x}, y_{t_y}], where tx,ty∈{0,1,…,T}t_x, t_y \in \{0, 1, \dots, T\} denote modality-specific perturbation timesteps and xtx,ytyx_{t_x}, y_{t_y} denote the perturbed data states. By manipulating txt_x and tyt_y, different distribution objectives are realized:

    • Marginal generation of x0x_0 (q(x0)q(x_0)): Setting ty=Tt_y = T marginalizes yy, yielding E[ϵx∣xtx,yT]≈E[ϵx∣xtx]\mathbb{E}[\epsilon^x \mid x_{t_x}, y_T] \approx \mathbb{E}[\epsilon^x \mid x_{t_x}] because yTy_T approaches standard Gaussian noise ϵy\epsilon^y.
    • Conditional generation of x0x_0 given y0y_0 (q(x0∣y0)q(x_0|y_0)): Setting ty=0t_y = 0 conditions on clean data y0y_0, yielding E[ϵx∣xtx,y0]\mathbb{E}[\epsilon^x \mid x_{t_x}, y_0]. Symmetrically, tx=0t_x = 0 yields conditional generation of y0y_0 given x0x_0 (q(y0∣x0)q(y_0|x_0)).
    • Joint generation of (x0,y0)(x_0, y_0) (q(x0,y0)q(x_0, y_0)): Setting tx=ty=tt_x = t_y = t ties the timesteps, estimating E[ϵx,ϵy∣xt,yt]\mathbb{E}[\epsilon^x, \epsilon^y \mid x_t, y_t].
  2. Knowl 2 — Training Objective for Unified Multi-Modal Diffusion

    equation

    UniDiffuser trains a joint noise prediction network ϵθ(xtx,yty,tx,ty)\epsilon_\theta(x_{t_x}, y_{t_y}, t_x, t_y) to simultaneously predict the noises injected into both modalities using a single regression loss:

    min⁡θE(x0,y0),ϵx,ϵy,tx,ty∥ϵθ(xtx,yty,tx,ty)−[ϵx,ϵy]∥22\min_\theta \mathbb{E}_{(x_0, y_0), \epsilon^x, \epsilon^y, t_x, t_y} \left\| \epsilon_\theta(x_{t_x}, y_{t_y}, t_x, t_y) - [\epsilon^x, \epsilon^y] \right\|_2^2

    where (x0,y0)∼q(x0,y0)(x_0, y_0) \sim q(x_0, y_0) is a data sample from the paired multi-modal distribution, [⋅,⋅][\cdot, \cdot] denotes vector concatenation, ϵx,ϵy∼N(0,I)\epsilon^x, \epsilon^y \sim \mathcal{N}(0, I) are independent standard Gaussian noise vectors, and the discrete timesteps tx,ty∼Uniform({1,2,…,T})t_x, t_y \sim \text{Uniform}(\{1, 2, \dots, T\}) are sampled independently. The perturbed representations are given by:

    xtx=αˉtxx0+1−αˉtxϵx,yty=αˉtyy0+1−αˉtyϵyx_{t_x} = \sqrt{\bar{\alpha}_{t_x}} x_0 + \sqrt{1 - \bar{\alpha}_{t_x}} \epsilon^x, \quad y_{t_y} = \sqrt{\bar{\alpha}_{t_y}} y_0 + \sqrt{1 - \bar{\alpha}_{t_y}} \epsilon^y

    where αˉt=∏i=1tαi=∏i=1t(1−βi)\bar{\alpha}_t = \prod_{i=1}^t \alpha_i = \prod_{i=1}^t (1 - \beta_i), with βi\beta_i representing the variance schedule at step ii. A single parameter update optimizes all marginal, conditional, and joint generative tasks via a single forward-backward computation.

  3. Knowl 3 — Classifier-Free Guidance for Free in UniDiffuser

    model/method

    Because UniDiffuser natively models both conditional distributions (t=0t=0) and marginal distributions (t=Tt=T), classifier-free guidance (CFG) can be applied during inference for conditional and joint sampling without modifying the training process or reserving special null tokens ∅\emptyset.

    Let ϵθ=[ϵθx,ϵθy]\epsilon_\theta = [\epsilon^x_\theta, \epsilon^y_\theta], ss be the guidance scale, and ϵx,ϵy∼N(0,I)\epsilon^x, \epsilon^y \sim \mathcal{N}(0, I) be standard Gaussian noise vectors. The guidance formulations are:

    • Conditional text-to-image sampling (x0x_0 conditioned on y0y_0): ϵ^θx(xt,y0,t)=(1+s)ϵθx(xt,y0,t,0)−sϵθx(xt,ϵy,t,T)\hat{\epsilon}^x_\theta(x_t, y_0, t) = (1 + s)\epsilon^x_\theta(x_t, y_0, t, 0) - s\epsilon^x_\theta(x_t, \epsilon^y, t, T)

    • Conditional image-to-text sampling (y0y_0 conditioned on x0x_0): ϵ^θy(yt,x0,t)=(1+s)ϵθy(x0,yt,0,t)−sϵθy(ϵx,yt,T,t)\hat{\epsilon}^y_\theta(y_t, x_0, t) = (1 + s)\epsilon^y_\theta(x_0, y_t, 0, t) - s\epsilon^y_\theta(\epsilon^x, y_t, T, t)

    • Joint sampling of (x0,y0)(x_0, y_0): ϵ^θ(xt,yt,t)=(1+s)ϵθ(xt,yt,t,t)−s[ϵθx(xt,ϵy,t,T),ϵθy(ϵx,yt,T,t)]\hat{\epsilon}_\theta(x_t, y_t, t) = (1 + s)\epsilon_\theta(x_t, y_t, t, t) - s\left[ \epsilon^x_\theta(x_t, \epsilon^y, t, T), \epsilon^y_\theta(\epsilon^x, y_t, T, t) \right]

    • Unconditional sampling of x0x_0 and y0y_0: ϵ^θx(xt,t)=ϵθx(xt,ϵy,t,T),ϵ^θy(yt,t)=ϵθy(ϵx,yt,T,t)\hat{\epsilon}^x_\theta(x_t, t) = \epsilon^x_\theta(x_t, \epsilon^y, t, T), \quad \hat{\epsilon}^y_\theta(y_t, t) = \epsilon^y_\theta(\epsilon^x, y_t, T, t)

  4. Knowl 4 — Latent Space Multi-Modal Encoders and Decoders

    model/method

    UniDiffuser operates in continuous latent spaces constructed via modality-specific encoders and decoders:

    • Image Latent Space: The image representation x0=[x0AE,x0CLIP]x_0 = [x_0^{\text{AE}}, x_0^{\text{CLIP}}] concatenates an autoencoder latent x0AEx_0^{\text{AE}} (extracted using the Stable Diffusion VAE encoder EAE\mathcal{E}^{\text{AE}}) with a 512-dimensional semantic embedding x0CLIPx_0^{\text{CLIP}} from a CLIP ViT-B/32 encoder. Image reconstruction utilizes only x0AEx_0^{\text{AE}} via the decoder DAE\mathcal{D}^{\text{AE}}, while x0CLIPx_0^{\text{CLIP}} provides high-level semantic information needed for image-to-text tasks.
    • Text Latent Space: The text encoder extracts 77 token vectors (each 768-dimensional) from CLIP text encoder, which are projected down to 64 dimensions per token via a linear layer to produce text latent y0∈R77×64y_0 \in \mathbb{R}^{77 \times 64}. A 124M-parameter GPT-2 decoder Dtext\mathcal{D}^{\text{text}} takes y0y_0 as a prefix embedding and is fine-tuned to reconstruct text autoregressively via cross-entropy loss: min⁡ϕET[log⁡p(T∣y0)]=ET[∑i=1Nlog⁡p(Ti∣T1:i−1,y0)]\min_\phi \mathbb{E}_T [\log p(T|y_0)] = \mathbb{E}_T \left[ \sum_{i=1}^N \log p(T_i \mid T_{1:i-1}, y_0) \right]

    Both latent spaces naturally occupy the range [−2,2][-2, 2] with near-standard normal distributions (image: mean 0.02690.0269, std 0.79190.7919; text: mean 0.01270.0127, std 0.59570.5957), eliminating the need for post-hoc normalization.

  5. Knowl 5 — U-ViT Architecture Modifications for UniDiffuser

    model/method

    UniDiffuser implements its joint noise prediction network using the U-ViT transformer architecture, modified to support multi-modal diffusion at scale:

    1. Multi-Modal Tokenization: The network treats the perturbed image latent tokens, text latent tokens, and their respective individual timestep embeddings (txt_x and tyt_y) as unified sequence tokens.
    2. Training Stability with Post-Layer Normalization: To prevent numerical overflow when training with mixed precision (fp16), the pre-layer normalization of standard U-ViT is replaced with post-layer normalization, and an additional layer normalization is inserted immediately following the concatenation in long skip connections between shallow and deep layers.
    3. Network Configuration: The model comprises 31 transformer layers with a patch size of 2, hidden size of 1536, MLP intermediate dimension of 6144, 24 attention heads, totaling 952 million parameters.
  6. Knowl 6 — UniDiffuser Multi-Modal Training and Sampling Algorithms

    algorithm

    The training and general sampling procedures of UniDiffuser handle joint, conditional, and unconditional tasks within one unified framework:

    Algorithm: UniDiffuser Training
    Input: Paired dataset distribution q(x0,y0)q(x_0, y_0), total timesteps TT, noise schedule coefficients αˉ1,…,αˉT\bar{\alpha}_1, \dots, \bar{\alpha}_T
    Output: Trained joint noise prediction network ϵθ\epsilon_\theta
    repeat
        Sample (x0,y0)∼q(x0,y0)(x_0, y_0) \sim q(x_0, y_0)
        Sample tx,ty∼Uniform({1,2,…,T})t_x, t_y \sim \text{Uniform}(\{1, 2, \dots, T\}) independently
        Sample ϵx,ϵy∼N(0,I)\epsilon^x, \epsilon^y \sim \mathcal{N}(0, I)
        Compute xtx=αˉtxx0+1−αˉtxϵxx_{t_x} = \sqrt{\bar{\alpha}_{t_x}} x_0 + \sqrt{1 - \bar{\alpha}_{t_x}} \epsilon^x
        Compute yty=αˉtyy0+1−αˉtyϵyy_{t_y} = \sqrt{\bar{\alpha}_{t_y}} y_0 + \sqrt{1 - \bar{\alpha}_{t_y}} \epsilon^y
        Compute gradient step ∇θ∥ϵθ(xtx,yty,tx,ty)−[ϵx,ϵy]∥22\nabla_\theta \|\epsilon_\theta(x_{t_x}, y_{t_y}, t_x, t_y) - [\epsilon^x, \epsilon^y]\|_2^2
        Update parameters θ\theta
    until convergence
    return ϵθ\epsilon_\theta
    Algorithm: UniDiffuser Joint Sampling of (x0,y0)(x_0, y_0)
    Input: Timestep schedule T,…,1T, \dots, 1, noise coefficients αt,βt,αˉt,σt\alpha_t, \beta_t, \bar{\alpha}_t, \sigma_t
    Output: Generated pair (x0,y0)(x_0, y_0)
    Initialize xT,yT∼N(0,I)x_T, y_T \sim \mathcal{N}(0, I)
    for t=T,…,1t = T, \dots, 1 do
        if t>1t > 1 then
            Sample zx,zy∼N(0,I)z^x, z^y \sim \mathcal{N}(0, I)
        else
            Set zx=0,zy=0z^x = 0, z^y = 0
        end if
        xt−1=1αt(xt−βt1−αˉtϵθx(xt,yt,t,t))+σtzxx_{t-1} = \frac{1}{\sqrt{\alpha_t}} \left( x_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon^x_\theta(x_t, y_t, t, t) \right) + \sigma_t z^x
        yt−1=1αt(yt−βt1−αˉtϵθy(xt,yt,t,t))+σtzyy_{t-1} = \frac{1}{\sqrt{\alpha_t}} \left( y_t - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon^y_\theta(x_t, y_t, t, t) \right) + \sigma_t z^y
    end for
    return x0,y0x_0, y_0
  7. Knowl 7 — In-the-Wild Image Interpolation via Dual-Modality Diffusion Inversion

    algorithm

    UniDiffuser performs semantic interpolation between two arbitrary images IaI^a and IbI^b by leveraging deterministic DPM-Solver inversion across both text and image latent spaces:

    Input: Source images IaI^a and IbI^b, guidance scale ss, total timestep TT, spherical interpolation parameter θ∈[0,1]\theta \in [0, 1]
    Output: Interpolated image IθI^\theta
    Define conditional image-to-text model ϵ^θy(yt,x0,t)=(1+s)ϵθy(x0,yt,0,t)−sϵθy(ϵx,yt,T,t)\hat{\epsilon}^y_\theta(y_t, x_0, t) = (1 + s)\epsilon^y_\theta(x_0, y_t, 0, t) - s\epsilon^y_\theta(\epsilon^x, y_t, T, t)
    Define conditional text-to-image model ϵ^θx(xt,y0,t)=(1+s)ϵθx(xt,y0,t,0)−sϵθx(xt,ϵy,t,T)\hat{\epsilon}^x_\theta(x_t, y_0, t) = (1 + s)\epsilon^x_\theta(x_t, y_0, t, 0) - s\epsilon^x_\theta(x_t, \epsilon^y, t, T)
    Encode Ia,IbI^a, I^b into latent embeddings x0a,x0bx^a_0, x^b_0
    Sample initial text noise yT∼N(0,I)y_T \sim \mathcal{N}(0, I)
    Generate text latents:
        y0a=DPM-Solver(init=yT,start=T,end=0,model=ϵ^θy(⋅,x0a,⋅))y^a_0 = \text{DPM-Solver}(\text{init}=y_T, \text{start}=T, \text{end}=0, \text{model}=\hat{\epsilon}^y_\theta(\cdot, x^a_0, \cdot))
        y0b=DPM-Solver(init=yT,start=T,end=0,model=ϵ^θy(⋅,x0b,⋅))y^b_0 = \text{DPM-Solver}(\text{init}=y_T, \text{start}=T, \text{end}=0, \text{model}=\hat{\epsilon}^y_\theta(\cdot, x^b_0, \cdot))
    Invert image latents into noise states:
        xTa=DPM-Solver(init=x0a,start=0,end=T,model=ϵ^θx(⋅,y0a,⋅))x^a_T = \text{DPM-Solver}(\text{init}=x^a_0, \text{start}=0, \text{end}=T, \text{model}=\hat{\epsilon}^x_\theta(\cdot, y^a_0, \cdot))
        xTb=DPM-Solver(init=x0b,start=0,end=T,model=ϵ^θx(⋅,y0b,⋅))x^b_T = \text{DPM-Solver}(\text{init}=x^b_0, \text{start}=0, \text{end}=T, \text{model}=\hat{\epsilon}^x_\theta(\cdot, y^b_0, \cdot))
    Compute spherical linear interpolations:
        y0θ=slerp(y0a,y0b,θ)y^\theta_0 = \text{slerp}(y^a_0, y^b_0, \theta)
        xTθ=slerp(xTa,xTb,θ)x^\theta_T = \text{slerp}(x^a_T, x^b_T, \theta)
    Generate interpolated image latent:
        x0θ=DPM-Solver(init=xTθ,start=T,end=0,model=ϵ^θx(⋅,y0θ,⋅))x^\theta_0 = \text{DPM-Solver}(\text{init}=x^\theta_T, \text{start}=T, \text{end}=0, \text{model}=\hat{\epsilon}^x_\theta(\cdot, y^\theta_0, \cdot))
    Decode x0θx^\theta_0 to output image IθI^\theta
    return IθI^\theta
  8. Knowl 8 — Cross-Modal Variation and Blocked Gibbs Sampling

    model/method

    Because UniDiffuser natively fits both conditional distributions q(x0∣y0)q(x_0|y_0) and q(y0∣x0)q(y_0|x_0), it enables iterative translation and variation applications between image and text modalities without additional training:

    • Image Variation: An input image x0x_0 is mapped to text latent y0∼q(y0∣x0)y_0 \sim q(y_0|x_0), and a variation image x0′∼q(x0∣y0)x'_0 \sim q(x_0|y_0) is subsequently sampled, generating semantically consistent visual variations.
    • Text Variation: An input text y0y_0 is mapped to image latent x0∼q(x0∣y0)x_0 \sim q(x_0|y_0), followed by caption generation y0′∼q(y0∣x0)y'_0 \sim q(y_0|x_0) to obtain semantic textual paraphrases.
    • Blocked Gibbs Sampling: Modalities are sampled in an alternating Markov chain: y0(0)∼q(y0∣x0(0))→x0(1)∼q(x0∣y0(0))→y0(1)∼q(y0∣x0(1))→…y_0^{(0)} \sim q(y_0|x_0^{(0)}) \to x_0^{(1)} \sim q(x_0|y_0^{(0)}) \to y_0^{(1)} \sim q(y_0|x_0^{(1)}) \to \dots, exploring the joint distribution space through cyclic cross-modal transitions.
  9. Knowl 9 — Zero-Shot Text-to-Image Generation Performance on MS-COCO

    data/table

    The zero-shot Fréchet Inception Distance (FID) on the MS-COCO validation set demonstrates that UniDiffuser outperforms general-purpose multi-modal baselines and performs competitively against bespoke single-task text-to-image models:

    Model FID ↓\downarrow
    Bespoken models
    GLIDE 12.24
    Make-A-Scene 11.84
    DALL·E 2 10.39
    Stable Diffusion 8.59
    Imagen 7.27
    Parti 7.23
    General-purpose models
    Versatile Diffusion 10.09
    UniDiffuser (ours) 9.71

    UniDiffuser achieves an FID of 9.71 using a classifier-free guidance scale of s=3s=3, outperforming the direct general-purpose baseline Versatile Diffusion (FID 10.09) and the bespoke model DALL·E 2 (FID 10.39).

  10. Knowl 10 — Computational Efficiency and Scaling Comparison

    data/table

    UniDiffuser provides lower inference latency, reduced GPU memory consumption, and lower training overhead compared to multi-task and single-task diffusion architectures:

    Model Model size Inference time Inference memory Training cost
    Bespoken models
    DALL·E 2 4.5B – – –
    Imagen 2B – – –
    Parti 20B – – –
    Stable Diffusion 860M 25.43s 67.83GB 150K (A100 40GB GPU hours)
    General-purpose models
    Versatile Diffusion 2566M 23.89s 76.53GB –
    UniDiffuser (ours) 952M 19.77s 48.30GB 59K (A100 80GB GPU hours)

    Inference benchmarks measure the generation of 10 samples across 25 denoising steps on a single A100 80GB GPU. Compared to Stable Diffusion (860M parameters), UniDiffuser requires only ∼10%\sim 10\% more parameters (952M) to simultaneously support 5 tasks (marginal image, marginal text, text-to-image, image-to-text, and joint image-text generation), while achieving lower inference time (19.77s vs 25.43s) and lower peak GPU memory (48.30GB vs 67.83GB).

  11. Knowl 11 — Text Generation Smoothness Limitation in UniDiffuser

    limitation

    Although the GPT-2 text decoder achieves high reconstruction accuracy on clean benchmarks (BLEU-1 score of 0.969 and BLEU-4 score of 0.894 on the MS-COCO test set), text generated unconditionally or conditionally from UniDiffuser exhibits limited linguistic smoothness. This limitation is primarily attributed to web-scraped training captions in the LAION-5B dataset containing significant noise, ungrammatical fragments, and artifacts, rather than an architectural deficiency of the unified diffusion framework.

Coverage note — No substantial contributed material was omitted. Specific training dataset pre-processing heuristics (such as regex cleaning rules for LAION-5B) and visual sample galleries were summarized into their respective functional knowls.

References

  1. 1.Bao, F., Li, C., Sun, J., Zhu, J., and Zhang, B. Estimating the optimal covariance with imperfect mean in diffusion probabilistic models. In ICML, 2022a.
  2. 2.Bao, F., Li, C., Zhu, J., and Zhang, B. Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models. In ICLR, 2022b.
  3. 3.Bao, F., Li, C., Cao, Y., and Zhu, J. All are worth words: a vit backbone for score-based diffusion models. In CVPR, 2023a.
  4. 4.Bao, F., Zhao, M., Hao, Z., Li, P., Li, C., and Zhu, J. Equivariant energy-guided sde for inverse molecular design. In ICLR, 2023b.
  5. 5.Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., and Wei, F. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. In NeurIPS, 2022c.
  6. 6.Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W. Wavegrad: Estimating gradients for waveform generation. In ICLR, 2021.
  7. 7.Chen, T., Zhang, R., and Hinton, G. Analog bits: Generating discrete data using diffusion models with self-conditioning. ArXiv preprint, 2022.
  8. 8.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. ArXiv preprint, 2021.
  9. 9.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al. Cogview: Mastering text-to-image generation via transformers. In NeurIPS, 2021.
  10. 10.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  11. 11.Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y. Make-a-scene: Scene-based text-to-image generation with human priors. In ECCV, 2022.
  12. 12.Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., and Guo, B. Vector quantized diffusion model for text-to-image synthesis. In CVPR, 2022.
  13. 13.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017.
  14. 14.Ho, J. and Salimans, T. Classifier-free diffusion guidance. In NeurIPS, 2021.
  15. 15.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In NeurIPS, 2020.
  16. 16.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. ArXiv preprint, 2022a.
  17. 17.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. In ICLR Workshop on Deep Generative Models for Highly Structured Data, 2022b.
  18. 18.Hoogeboom, E., Satorras, V. G., Vignac, C., and Welling, M. Equivariant diffusion for molecule generation in 3d. In ICML, 2022.
  19. 19.Hu, M., Zheng, C., Zheng, H., Cham, T.-J., Wang, C., Yang, Z., Tao, D., and Suganthan, P. N. Unified discrete diffusion for simultaneous vision-language generation. arXiv preprint arXiv:2211.14842, 2022.
  20. 20.Karpathy, A. and Fei-Fei, L. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015.
  21. 21.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. In NeurIPS, 2022.
  22. 22.Kim, W., Son, B., and Kim, I. Vilt: Vision-and-language transformer without convolution or region supervision. In ICML, 2021.
  23. 23.Kingma, D. P., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. In NeurIPS, 2021.
  24. 24.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. In ICLR, 2021.
  25. 25.Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022.
  26. 26.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In ECCV, 2014.
  27. 27.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In ICLR, 2019.
  28. 28.Lu, C., Zheng, K., Bao, F., Chen, J., Li, C., and Zhu, J. Maximum likelihood training for score-based diffusion odes by high order denoising score matching. In ICML, 2022a.
  29. 29.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. In NeurIPS, 2022b.
  30. 30.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022c.
  31. 31.Luo, S. and Hu, W. Diffusion probabilistic models for 3d point cloud generation. In CVPR, 2021.
  32. 32.Mokady, R., Hertz, A., and Bermano, A. H. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734, 2021.
  33. 33.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In ICML, 2022.
  34. 34.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In ICML, 2021.
  35. 35.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  36. 36.Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M. A. Grad-tts: A diffusion probabilistic model for text-to-speech. In ICML, 2021.
  37. 37.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  38. 38.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  39. 39.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In ICML, 2021.
  40. 40.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. ArXiv preprint, 2022.
  41. 41.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  42. 42.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022.
  43. 43.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. In ICLR, 2022.
  44. 44.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. In NeurIPS, 2022.
  45. 45.Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In ICLR, 2015.
  46. 46.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In ICLR, 2021a.
  47. 47.Song, Y., Durkan, C., Murray, I., and Ermon, S. Maximum likelihood training of score-based diffusion models. In NeurIPS, 2021b.
  48. 48.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In ICLR, 2021c.
  49. 49.Srivastava, N. and Salakhutdinov, R. R. Multimodal learning with deep boltzmann machines. In NeurIPS, 2012.
  50. 50.Vahdat, A., Kreis, K., and Kautz, J. Score-based generative modeling in latent space. ArXiv preprint, 2021.
  51. 51.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In NeurIPS, 2017.
  52. 52.Vincent, P. A connection between score matching and denoising autoencoders. Neural computation, 2011.
  53. 53.Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. ArXiv preprint, 2022.
  54. 54.Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T. On layer normalization in the transformer architecture. In ICML, 2020.
  55. 55.Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., and He, X. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In CVPR, 2018.
  56. 56.Xu, X., Wang, Z., Zhang, E., Wang, K., and Shi, H. Versatile diffusion: Text, images and variations all in one diffusion model. arXiv preprint arXiv:2211.08332, 2022.
  57. 57.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. ArXiv preprint, 2022.
  58. 58.Zhao, M., Bao, F., Li, C., and Zhu, J. Egsde: Unpaired image-to-image translation via energy-guided stochastic differential equations. In NeurIPS, 2022.

Citation

MLA
Bao, F., et al. “One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale”. International Conference on Machine Learning, vol. 202, 2023, pp. 1692–717, https://proceedings.mlr.press/v202/bao23a.html.
APA
Bao, F., Nie, S., Xue, K., Li, C., Pu, S., Wang, Y., Yue, G., Cao, Y., Su, H., & Zhu, J. (2023). One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale. International Conference on Machine Learning, 202, 1692–1717. https://proceedings.mlr.press/v202/bao23a.html
Chicago
Bao, F., S. Nie, K. Xue, et al. 2023. “One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale”. International Conference on Machine Learning 202: 1692–1717. https://proceedings.mlr.press/v202/bao23a.html.
Harvard
Bao, F. et al. (2023) “One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale”, International Conference on Machine Learning. PMLR, pp. 1692–1717. Available at: https://proceedings.mlr.press/v202/bao23a.html.
Vancouver
1. Bao F, Nie S, Xue K, Li C, Pu S, Wang Y, Yue G, Cao Y, Su H, Zhu J (2023) One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale. In: International Conference on Machine Learning. PMLR, pp 1692–1717

BibTeX

@InProceedings{pmlr-v202-bao23a,
  title = 	 {One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale},
  author =       {Bao, Fan and Nie, Shen and Xue, Kaiwen and Li, Chongxuan and Pu, Shi and Wang, Yaole and Yue, Gang and Cao, Yue and Su, Hang and Zhu, Jun},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {1692--1717},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/bao23a/bao23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/bao23a.html},
  abstract = 	 {This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoken models (e.g., Stable Diffusion and DALL-E 2) in representative tasks (e.g., text-to-image generation).}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/