One Diffusion to Generate Them All
Duong H. LeTuan PhamSangho LeeChristopher ClarkAniruddha KembhaviStephan MandtRanjay KrishnaJiasen Lu
Presents OneDiffusion, a unified diffusion model that integrates bidirectional image synthesis and visual understanding into a single architecture by formulating diverse vision tasks as multi-frame sequences with variable noise scales.
Computer vision models currently rely on separate, specialized architectures and external add-on modules to perform distinct tasks such as text-to-image synthesis, multi-view generation, and image understanding. This fragmented approach increases deployment complexity, limits scalability, and prevents models from generalizing across varied visual domains like large language models do.
The article demonstrates OneDiffusion, a unified 2.8-billion parameter diffusion model trained from scratch to execute both generative and predictive computer vision tasks within a single architecture. The primary objective is to evaluate whether casting diverse visual inputs and tasks into a unified sequence framework can eliminate the need for specialized modules or task-specific losses.
The researchers developed a sequence-based flow-matching framework that treats all task conditions and targets as variable sequences of image frames or views with independent noise levels. Training occurred in three stages across public and synthetic datasets totaling tens of millions of samples, covering text-to-image synthesis, image-to-image translation, identity customization, and multi-view generation. During inference, specific views serve as conditions while others are generated from noise, allowing the model to perform bidirectional generation and prediction across arbitrary resolutions.
Evaluation shows that OneDiffusion achieves competitive performance across all tested domains despite using a relatively compact dataset. On text-to-image alignment, the model achieved a 0.65 GenEval score, outperforming larger baselines like SD3-medium and matching 12-billion parameter models. In multi-view tasks, it matched or surpassed dedicated baselines on standard reconstruction metrics while handling novel setups with unknown camera poses. For identity customization, it successfully modified viewpoints, gaze directions, and non-human subjects where specialized face-embedding models failed. Additionally, it delivered competitive monocular depth estimation while exhibiting strong robustness on unconventional, open-world imagery.
These findings indicate that unified diffusion architectures can streamline computer vision pipelines, substantially reducing engineering overhead, deployment costs, and architectural fragmentation. Removing task-specific adapters allows organizations to maintain a single foundation model that dynamically switches between generating content and predicting underlying visual properties without compromising output quality.
Organizations developing or deploying visual AI systems should consider adopting unified sequential diffusion frameworks instead of maintaining disparate, task-specific pipelines. To prepare for practical deployment, teams should run targeted pilot evaluations to assess latency and throughput under production workloads and test task performance against specialized proprietary models on domain-specific data.
Confidence in the core methodology and results is high based on the standardized benchmarks provided across diverse vision tasks. However, users should note that on standard identity benchmarks focused strictly on facial replication, the unified attention mechanism achieved lower identity similarity scores than models with dedicated face-recognition networks, indicating a trade-off between flexible manipulation and strict identity preservation.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Read this foundational account of the forward-noising and learned denoising process that OneDiffusion adapts into its unified framework.
- Paper: One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale, Fan Bao et al. (2023). UniDiffuser establishes the closely related idea of one diffusion transformer covering multiple conditional and joint generation modes through independently varied noise levels.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). Its rectified-flow transformer formulation provides essential context for OneDiffusion’s flow-matching approach to sequence-based visual generation.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). DiT introduces the transformer backbone for diffusion that helps make sense of OneDiffusion’s large sequence-based architecture.
- Paper: Show-o: One Single Transformer to Unify Multimodal Understanding and Generation, Jinheng Xie et al. (2025). Show-o carries unified generation-and-understanding research into a multimodal language-model setting, extending the architectural ambition beyond OneDiffusion’s vision-task framework.
