One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale
Fan BaoShen NieKaiwen XueChongxuan LiShi PuYaole WangGang YueYue CaoHang SuJun Zhu
Proposes UniDiffuser, a single transformer-based framework that captures marginal, conditional, and joint multi-modal distributions to handle diverse generation tasks—including text-to-image, image-to-text, and paired generation—without requiring task-specific models or extra computational overhead.
Generative artificial intelligence has traditionally relied on specialized, single-task systems that separately handle text-to-image creation, image captioning, or unconditional content generation. While these dedicated frameworks produce high-quality outputs, maintaining distinct architectures for individual tasks creates significant operational redundancy, escalates infrastructure costs, and prevents systems from sharing cross-modal knowledge. A unified framework capable of learning all joint, conditional, and marginal distributions simultaneously across multiple modalities has remained a key challenge due to computational complexity and architectural limits.
The article demonstrates UniDiffuser, a single transformer-based diffusion model designed to capture marginal, conditional, and joint distributions across image and text modalities within a unified framework. It evaluates whether a single general-purpose system can match or exceed the performance of specialized generative models without incurring additional training or inference overhead.
The research team formulated the training process as a unified noise prediction task where perturbation levels vary independently across modalities. By assigning zero noise to a condition modality and maximum noise to ignore an unconditioned modality, the system dynamically switches among generative modes. The implementation employs a 952-million-parameter transformer operating in continuous latent spaces, leveraging pre-trained vision encoders and language decoders. The framework was evaluated on subsets of the open-access LAION-5B dataset comprising hundreds of millions of image-text pairs and benchmarked against standard datasets like MS-COCO.
The evaluation yielded several key findings. First, the unified model achieved a zero-shot Fréchet Inception Distance score of 9.71 on MS-COCO text-to-image synthesis, outperforming existing general-purpose models like Versatile Diffusion (10.09) and specialized systems like DALL·E 2 (10.39) while remaining competitive with Stable Diffusion (8.59). Second, the model consistently surpassed multi-task baselines in cross-modal alignment across both text-to-image and image-to-text generation tasks. Third, it enabled classifier-free guidance during sampling at zero additional training cost by directly incorporating marginal score estimates. Finally, it demonstrated high operational efficiency, generating batches in 19.77 seconds with 48.30 gigabytes of memory, outperforming comparable systems in both processing speed and memory footprint.
These findings indicate that generative workflows can be consolidated into single, scalable foundation models without sacrificing output fidelity. Adopting unified architectures lowers the risks and costs associated with managing multiple task-specific systems, simplifies model maintenance, and expands functionality into complex capabilities such as continuous image-text translation, variation generation, and latent interpolation.
Decision-makers should explore unified diffusion architectures when designing multi-modal infrastructure rather than deploying disparate point solutions. Prior to production deployment, organizations should conduct targeted pilot tests and integrate safety safeguards, including automated watermarking and detection protocols, to mitigate misuse risks such as synthetic misinformation.
The framework presents certain operational limitations. The quality and fluency of generated text remain constrained by noise in the underlying training datasets. Additionally, validating the architecture across more diverse modalities—such as audio, video, and 3D data—requires further empirical study. Overall confidence in the performance and efficiency gains of the unified model remains high for two-modality vision and language deployments.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work establishes the foundational denoising diffusion probabilistic framework that UniDiffuser modifies and unifies across multiple modalities.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). This work introduces Diffusion Transformers (DiTs), providing the core transformer-based diffusion backbone that UniDiffuser adapts to process multi-modal inputs.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). This paper establishes classifier-free guidance, a fundamental sampling technique used in conditional and joint multi-modal diffusion generation.
- Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). This paper investigates diffusion formulations across discrete data spaces, offering essential background for extending diffusion models beyond continuous images to text.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). This work introduces non-Markovian implicit sampling (DDIM), which provides the mathematical mechanics for fast and deterministic diffusion sampling across distributions.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This paper formalizes the design space, noise schedules, and preconditioning of diffusion models, clarifying the noise prediction parameterization UniDiffuser builds upon.
- Paper: Multimodal learning with deep Boltzmann machines, Nitish Srivastava et al. (2012). This foundational paper introduces joint, marginal, and conditional generative modeling across paired image and text channels in a unified probabilistic architecture.
- Paper: Generative Multimodal Models are In-Context Learners, Quan Sun et al. (2024). This work scales unified generative multimodal architectures to foundation-scale in-context learners that combine language backbones with diffusion decoders.
- Paper: There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation, Gabe Guo et al. (2026). This volume extends bidirectional multimodality translation from noise-based multi-modal diffusion to data-to-data diffusion bridges.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This research advances unified multimodal transformer architectures by developing Multimodal Diffusion Transformers (MM-DiT) parameterized via rectified flows.
- Paper: UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild, Can Qin et al. (2023). This paper builds on unified diffusion modeling to consolidate diverse visual conditionings and text prompts into a single multi-task diffusion framework.
- Paper: Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners, Yazhou Xing et al. (2024). This paper expands multimodal diffusion generation into cross-modal visual-audio alignment and joint audiovisual synthesis.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). This paper unifies multifaceted conditional and unconditional generation within a single large diffusion transformer by learning cross-frame dynamics.
