One Diffusion to Generate Them All

Duong H. LeTuan PhamSangho LeeChristopher ClarkAniruddha KembhaviStephan MandtRanjay KrishnaJiasen Lu

article2025CVPR52 citations

Presents OneDiffusion, a unified diffusion model that integrates bidirectional image synthesis and visual understanding into a single architecture by formulating diverse vision tasks as multi-frame sequences with variable noise scales.

Listen

Computer vision models currently rely on separate, specialized architectures and external add-on modules to perform distinct tasks such as text-to-image synthesis, multi-view generation, and image understanding. This fragmented approach increases deployment complexity, limits scalability, and prevents models from generalizing across varied visual domains like large language models do.

The article demonstrates OneDiffusion, a unified 2.8-billion parameter diffusion model trained from scratch to execute both generative and predictive computer vision tasks within a single architecture. The primary objective is to evaluate whether casting diverse visual inputs and tasks into a unified sequence framework can eliminate the need for specialized modules or task-specific losses.

The researchers developed a sequence-based flow-matching framework that treats all task conditions and targets as variable sequences of image frames or views with independent noise levels. Training occurred in three stages across public and synthetic datasets totaling tens of millions of samples, covering text-to-image synthesis, image-to-image translation, identity customization, and multi-view generation. During inference, specific views serve as conditions while others are generated from noise, allowing the model to perform bidirectional generation and prediction across arbitrary resolutions.

Evaluation shows that OneDiffusion achieves competitive performance across all tested domains despite using a relatively compact dataset. On text-to-image alignment, the model achieved a 0.65 GenEval score, outperforming larger baselines like SD3-medium and matching 12-billion parameter models. In multi-view tasks, it matched or surpassed dedicated baselines on standard reconstruction metrics while handling novel setups with unknown camera poses. For identity customization, it successfully modified viewpoints, gaze directions, and non-human subjects where specialized face-embedding models failed. Additionally, it delivered competitive monocular depth estimation while exhibiting strong robustness on unconventional, open-world imagery.

These findings indicate that unified diffusion architectures can streamline computer vision pipelines, substantially reducing engineering overhead, deployment costs, and architectural fragmentation. Removing task-specific adapters allows organizations to maintain a single foundation model that dynamically switches between generating content and predicting underlying visual properties without compromising output quality.

Organizations developing or deploying visual AI systems should consider adopting unified sequential diffusion frameworks instead of maintaining disparate, task-specific pipelines. To prepare for practical deployment, teams should run targeted pilot evaluations to assess latency and throughput under production workloads and test task performance against specialized proprietary models on domain-specific data.

Confidence in the core methodology and results is high based on the standardized benchmarks provided across diverse vision tasks. However, users should note that on standard identity benchmarks focused strictly on facial replication, the unified attention mechanism achieved lower identity similarity scores than models with dedicated face-recognition networks, indicating a trade-off between flexible manipulation and strict identity preservation.

Cover for One Diffusion to Generate Them All

Abstract

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables conditional generation from inputs such as text, depth, pose, layout, and semantic maps, while also handling tasks like image deblurring, upscaling, and reverse processes such as depth estimation and segmentation. Additionally, OneDiffusion allows for multi-view generation, camera pose estimation, and instant personalization using sequential image inputs. Our model takes a straightforward yet effective approach by treating all tasks as frame sequences with varying noise scales during training, allowing any frame to act as a conditioning image at inference time. Our unified training framework removes the need for specialized architectures, supports scalable multi-task training, and adapts smoothly to any resolution, enhancing both generalization and scalability. Experimental results demonstrate competitive performance across tasks in both generation and prediction such as text-to-image, multiview generation, ID preservation, depth estimation and camera pose estimation despite relatively small training dataset. Our code and checkpoint are freely available at this https URL

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Methodology
  • 3.1 Flow matching for generative modeling
  • 3.2 Proposed Approach
  • 3.3 Implementation Details
  • 4 One-Gen Datasets
  • 5 Experiments
  • 5.1 Text-to-Image
  • 5.2 Controllable Image generation
  • 5.3 Multiview Generation
  • 5.4 ID Customization
  • 5.5 Depth Estimation
  • 6 Conclusion
  • 7 Acknowledgements
  • References
  • 8 Additional quantitative results
  • 8.1 Camera Pose Estimation
  • 8.2 Image Editing and Subject-driven generation
  • 9 Additional qualitative results
  • 10 Summary Datasets

Knowls

  1. Knowl 1 — A single sequence model covers image synthesis and understanding

    model/method

    OneDiffusion represents a task as a sequence of NN views, where each view is an image or image-like condition and text specifies the task and any additional instructions. Text-to-image uses one view; image-to-image tasks such as depth-to-image or pose-to-image use two; and identity customization or multiview generation can use more than two. The same learned model can generate any selected subset of views while treating the others as conditions, so reversing the direction of an image translation does not require a separate task-specific model. The framework is applied to text-to-image synthesis, image editing and inpainting, upscaling, generation conditioned on edges, depth, pose, semantic maps or boxes, and prediction of outputs such as depth, pose, segmentation, edges, and camera poses.

  2. Knowl 2 — Joint flow-matching objective for independently noised views

    equation

    For a training example with NN views, let xix_i be the clean tensor for view ii, and let ϵi\epsilon_i be independent standard Gaussian noise with the same shape as xix_i. The dimensionless time ti∈[0,1]t_i\in[0,1] determines the noisy view xi(ti)x_i(t_i); ti=0t_i=0 is pure noise and ti=1t_i=1 is the clean view. The target velocity for each view is ui=xi−ϵiu_i=x_i-\epsilon_i. OneDiffusion predicts all view velocities jointly with vθv_\theta, where θ\theta denotes the model parameters, and trains by minimizing

    xi(ti)=tixi+(1−ti)ϵi,ui=xi−ϵi,L(θ)=E[∥vθ(t1,…,tN,x1(t1),…,xN(tN))−(u1,…,uN)∥22].x_i(t_i)=t_i x_i+(1-t_i)\epsilon_i,\qquad u_i=x_i-\epsilon_i,\qquad \mathcal{L}(\theta)=\mathbb{E}\left[\left\|v_\theta(t_1,\ldots,t_N,x_1(t_1),\ldots,x_N(t_N))-(u_1,\ldots,u_N)\right\|_2^2\right].

    The expectation is over training examples, times, and Gaussian noises; the squared Euclidean norm covers the concatenated per-view outputs. Training samples a separate time and noise for each view, allowing conditioned and target views to have different noise levels in the same example.

  3. Knowl 3 — Conditional generation by fixing selected views during flow integration

    model/method

    To generate a chosen set of target views K⊆{1,…,N}K\subseteq\{1,\ldots,N\}, OneDiffusion initializes those views with Gaussian noise and holds the complementary views at their known values xˉ∖K\bar{x}_{\setminus K}. At integration time t∈[0,1]t\in[0,1], it sets every target-view time to tt and every conditioning-view time to zero, then evaluates the corresponding components of the joint model vector field:

    vθK(t,xK∣xˉ∖K)=[vθ(tK=t, t∖K=0, xK, x∖K=xˉ∖K)]K.v^K_\theta(t,x_K\mid\bar{x}_{\setminus K})= \left[v_\theta(t_K=t,\ t_{\setminus K}=0,\ x_K,\ x_{\setminus K}=\bar{x}_{\setminus K})\right]_K.

    Here xKx_K denotes the current target-view tensors, xˉ∖K\bar{x}_{\setminus K} the fixed conditioning tensors, and [⋅]K[\cdot]_K selects the predicted velocities for target views. Integrating dxK/dt=vθK\mathrm{d}x_K/\mathrm{d}t=v^K_\theta from t=0t=0 to t=1t=1 produces the conditional samples. This setup permits arbitrary choices of input and output views, including generation conditioned on multiple images or on camera information.

  4. Knowl 4 — Transformer and latent representation support variable view counts and resolutions

    model/method

    OneDiffusion uses the Next-DiT full-transformer architecture with a VAE tokenizer. Each image or image-like condition is encoded separately as a latent tensor, and the latents are concatenated along the view dimension, allowing the sequence length in views to vary with the task. The model uses 3D rotary positional embeddings (3D RoPE) and task-label tokens in the text input to identify the requested operation. The paper reports that this design accommodates different resolutions and aspect ratios and enables zero-shot high-resolution generation at resolutions not encountered during training.

  5. Knowl 5 — Camera poses are represented as independently denoised Plücker-ray views

    model/method

    For multiview generation and camera-pose prediction, OneDiffusion encodes each camera ray by its Plücker coordinates r=(o×d,d)r=(o\times d,d), where oo is the ray origin and dd its direction. The six-coordinate representation is computed at the latent patch resolution H/8×W/8H/8\times W/8, then replicated across channels to form a 16-channel embedding; it is scaled to unit variance. Rather than concatenating camera information into image channels, the model treats each ray-embedding tensor as its own view in the sequence. Consequently, camera rays can be fixed as conditioning views to generate images, or treated as target views to estimate poses from image conditions.

  6. Knowl 6 — One-Gen combines public and synthetic data across task families

    data/table

    One-Gen supplies paired or grouped examples for the unified tasks. For text-to-image training, the paper uses PixelProse, Unsplash, Coyo, and JourneyDB, together with 10 million internally generated images recaptioned using LLaVA-NeXT and Molmo; the text-to-image benchmark comparison reports 75 million training examples for OneDiffusion. For simpler image-to-image tasks—including deblurring, inpainting, Canny-edge conditioning, and upscaling—the authors use a one-million-sample synthetic-data subset and apply task-specific preprocessing. Other synthetic image-to-image data include 350,000 image/semantic-map/box triplets, depth estimates generated for 500,000 images plus 40,000 captioned HyperSim images, and 50,000 human-pose examples. Identity customization uses approximately 1.3 million captioned images of about 60,000 subjects, filtered to retain subjects with at least four images. Multiview training draws on DL3DV-10K, an 80,000-object filtered Objaverse split, and CO3D. This mix supplies conditions and targets in the sequence format used by the shared model.

  7. Knowl 7 — Three-stage training recipe trains the model from scratch

    experimental setup

    The 2.8-billion-parameter model is trained from scratch in three stages. First, text-to-image pretraining runs for 500,000 steps at 256×256256\times256 resolution and another 500,000 steps at 512×512512\times512. Second, mixed-task training runs for 1 million steps, using 512×512512\times512 for text-to-image and 256×256256\times256 for other tasks. Third, text-to-image fine-tuning uses 1024×10241024\times1024 resolution. In each stage, text-to-image, image-to-image, identity customization, and multiview tasks are sampled with equal probability. The reported setup uses AdamW with learning rate 0.00050.0005, a noise-scheduler shift of 3, and global batch size 256 in the first two stages; the final stage uses 64 H100 GPUs with the same configuration. Identity-customization fine-tuning uses 2–5 views: 512×512512\times512 for 2–3 views and 256×256256\times256 for larger view counts.

  8. Knowl 8 — Text-to-image performance on GenEval

    empirical result

    On GenEval at 1024×10241024\times1024 resolution, OneDiffusion scores 0.65 with 2.8 billion parameters and 75 million reported training examples. For evaluation, the authors generate four images per prompt using a 100-step Euler solver and guidance scale 5. The score is above the reported values for similarly sized LUMINA-Next (0.46; 2.0B parameters, 14M data), PixArt-Σ\Sigma (0.54; 0.6B, 33M), SDXL (0.55; 2.6B), SD3-medium (0.62; 2.0B, 1,000M), and Hunyuan-DiT (0.63; 1.5B). It is below FLUX-schnell (0.71; 12.0B) and the reported 0.67 scores for DALL·E 3 and FLUX-dev (12.0B). The comparison supports the paper's claim of competitive text-to-image performance alongside multitask training, rather than a best-in-table result.

  9. Knowl 9 — Multiview reconstruction quality varies with number and pose knowledge of inputs

    empirical result

    On the Google Scanned Objects multiview evaluation, OneDiffusion reports PSNR 19.01 with one conditioning view; 19.83 with two views and unknown poses; 20.22 with two views and known poses; 20.64 with three views and unknown poses; and 21.79 with three views and known poses. For comparison, Zero123 scores 18.51 and Zero123-XL 18.93 with one view, while EscherNet scores 20.24 with one view, 22.91 with two, and 24.09 with three. Thus OneDiffusion exceeds the two Zero123 baselines in the one-view setting and produces useful results with unknown poses, but its reported PSNR is below EscherNet's in the listed view-count comparisons. The authors also demonstrate consistent views generated from one input image and text-to-multiview generation conditioned on camera poses.

  10. Knowl 10 — Identity customization supports varied references and manipulations

    empirical result

    The identity-customization evaluation includes expression and gaze changes, viewpoint changes, and non-human identities; the paper reports examples generated from one or multiple reference images and states that competing personalization methods often fail on the more complex manipulations or non-human inputs. On Unsplash-50, the reported ID and CLIP-T scores (both marked higher-is-better) are: PhotoMaker 0.193 and 27.38; InstantID 0.648 and 26.41; PuLID 0.654 and 31.23; OneDiffusion 0.283 and 26.80. OneDiffusion therefore does not lead these quantitative metrics, despite the paper's qualitative demonstrations of broader manipulation and input coverage.

  11. Knowl 11 — Monocular depth results are competitive but not best on either benchmark

    empirical result

    On NYUv2 and DIODE, OneDiffusion reports AbsRel values of 6.8 and 29.4, respectively, and δ1\delta_1 values of 95.2 and 75.2, respectively. Lower AbsRel and higher δ1\delta_1 indicate better results. For context, Marigold reports NYUv2 values of 6.0 and 95.9 and DIODE values of 31.0 and 77.2; Depth Anything-2 reports 4.6 and 97.7 on NYUv2 and 27.1 and 74.8 on DIODE. OneDiffusion is therefore worse than these methods on NYUv2 for both metrics; on DIODE it has lower AbsRel than Marigold but lower δ1\delta_1, and it has higher δ1\delta_1 but higher AbsRel than Depth Anything-2. The paper additionally reports qualitative robustness on open-world images such as paintings, hazy scenes, and images with unconventional textures.

Coverage note — Detailed qualitative examples for individual edge, pose, segmentation, and bounding-box tasks are not separate knowls because the paper reports them without task-specific quantitative analyses; their scope is captured in the unified task formulation.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Michael S Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022.
  3. 3.Jason Baldridge, Jakob Bauer, Mukul Bhutani, Nicole Brichtova, Andrew Bunner, Kelvin Chan, Yichang Chen, Sander Dieleman, Yuqing Du, Zach Eaton-Rosen, et al. Imagen 3. arXiv preprint arXiv:2408.07009, 2024.
  4. 4.James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn.openai.com/papers/dall-e-3.pdf, 2(3):8, 2023.
  5. 5.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023.
  6. 6.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset, 2022.
  7. 7.Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion. arXiv preprint arXiv:2407.01392, 2024.
  8. 8.Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023.
  9. 9.Junsong Chen, Chongjian Ge, Enze Xie, Yue Wu, Lewei Yao, Xiaozhe Ren, Zhongdao Wang, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation. arXiv preprint arXiv:2403.04692, 2024.
  10. 10.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. 2023 ieee. In CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13142–13153, 2022.
  11. 11.Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Christopher Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz, Aaron Sarnat, Byron Bischoff, Pete Walsh, Christopher Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Jennifer Dumas, Crystal Nam, Sophie Lebrecht, Caitlin Marie Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hanna Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi. Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models. arXiv preprint arXiv:2409.17146, 2024.
  12. 12.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. Advances in Neural Information Processing Systems, 36, 2024.
  13. 13.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automation (ICRA), pages 2553–2560. IEEE, 2022.
  14. 14.Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10786–10796, 2021.
  15. 15.Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning, 2024.
  16. 16.Rinon Gal, Or Lichter, Elad Richardson, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Lcm-lookahead for encoder-based text-to-image personalization. arXiv preprint arXiv:2404.03620, 2024.
  17. 17.Yulu Gan, Sungwoo Park, Alexander Schubert, Anthony Philippakis, and Ahmed M Alaa. Instructcv: Instruction-tuned text-to-image diffusion models as vision generalists. arXiv preprint arXiv:2310.00390, 2023.
  18. 18.Ruiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T Barron, and Ben Poole. Cat3d: Create anything in 3d with multi-view diffusion models. arXiv preprint arXiv:2405.10314, 2024.
  19. 19.Zigang Geng, Binxin Yang, Tiankai Hang, Chen Li, Shuyang Gu, Ting Zhang, Jianmin Bao, Zheng Zhang, Houqiang Li, Han Hu, et al. Instructdiffusion: A generalist modeling interface for vision tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12709–12720, 2024.
  20. 20.Dhruba Ghosh, Hannaneh Hajishirzi, and Ludwig Schmidt. Geneval: An object-focused framework for evaluating text-to-image alignment. Advances in Neural Information Processing Systems, 36, 2024.
  21. 21.Zinan Guo, Yanze Wu, Zhuowei Chen, Lang Chen, and Qian He. Pulid: Pure and lightning id customization via contrastive alignment. arXiv preprint arXiv:2404.16022, 2024.
  22. 22.Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, and Qiang Xu. Direct inversion: Boosting diffusion-based editing with 3 lines of code. arXiv preprint arXiv:2310.01506, 2023.
  23. 23.Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler. Repurposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9492–9502, 2024.
  24. 24.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023.
  25. 25.Xin Kong, Shikun Liu, Xiaoyang Lyu, Marwan Taher, Xiaojuan Qi, and Andrew J Davison. Eschernet: A generative model for scalable view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9503–9513, 2024.
  26. 26.Daiqing Li, Aleks Kamko, Ehsan Akhgari, Ali Sabet, Linmiao Xu, and Suhail Doshi. Playground v2. 5: Three insights towards enhancing aesthetic quality in text-to-image generation. arXiv preprint arXiv:2402.17245, 2024.
  27. 27.Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Open-vocabulary object segmentation with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7667–7676, 2023.
  28. 28.Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming-Ming Cheng, and Ying Shan. Photomaker: Customizing realistic human photos via stacked id embedding. In CVPR, 2024.
  29. 29.Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, et al. Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding. arXiv preprint arXiv:2405.08748, 2024.
  30. 30.Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22160–22169, 2024.
  31. 31.Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023.
  32. 32.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, 2024.
  33. 33.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9298–9309, 2023.
  34. 34.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022.
  35. 35.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023.
  36. 36.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. In The Eleventh International Conference on Learning Representations, 2022.
  37. 37.Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi. Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action. arXiv preprint arXiv:2312.17172, 2023.
  38. 38.Ao Luo, Xin Li, Fan Yang, Jiangyu Liu, Haoqiang Fan, and Shuaicheng Liu. Flowdiffuser: Advancing optical flow estimation with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19167–19176, 2024.
  39. 39.Tiange Luo, Chris Rockwell, Honglak Lee, and Justin Johnson. Scalable 3d captioning with pretrained models. Advances in Neural Information Processing Systems, 36, 2024.
  40. 40.Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 4296–4304, 2024.
  41. 41.Junting Pan, Keqiang Sun, Yuying Ge, Hao Li, Haodong Duan, Xiaoshi Wu, Renrui Zhang, Aojun Zhou, Zipeng Qin, Yi Wang, Jifeng Dai, Yu Qiao, and Hongsheng Li. Journeydb: A benchmark for generative image understanding, 2023.
  42. 42.Kushagra Pandey, Ruihan Yang, and Stephan Mandt. Fast samplers for inverse problems in iterative refinement models. Advances in Neural Information Processing Systems, 37:26872–26914, 2024.
  43. 43.Kushagra Pandey, Farrin Marouf Sofian, Felix Draxler, Theofanis Karaletsos, and Stephan Mandt. Variational control for guidance in diffusion models. arXiv preprint arXiv:2502.03686, 2025.
  44. 44.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  45. 45.Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, et al. Unicontrol: A unified diffusion model for controllable visual generation in the wild. arXiv preprint arXiv:2305.11147, 2023.
  46. 46.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821–8831. Pmlr, 2021.
  47. 47.Rene Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637, 2020.
  48. 48.Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10901–10911, 2021.
  49. 49.Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10912–10922, 2021.
  50. 50.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  51. 51.David Ruhe, Jonathan Heek, Tim Salimans, and Emiel Hoogeboom. Rolling diffusion models. arXiv preprint arXiv:2402.09470, 2024.
  52. 52.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22500–22510, 2023.
  53. 53.Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J Fleet. The surprising effectiveness of diffusion models for optical flow and monocular depth estimation. Advances in Neural Information Processing Systems, 36, 2024.
  54. 54.Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023.
  55. 55.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12, pages 746–760. Springer, 2012.
  56. 56.Vasu Singla, Kaiyu Yue, Sukriti Paul, Reza Shirkavand, Mayuka Jayawardhana, Alireza Ganjdanesh, Heng Huang, Abhinav Bhatele, Gowthami Somepalli, and Tom Goldstein. From pixels to prose: A large dataset of dense image captions. arXiv preprint arXiv:2406.10328, 2024.
  57. 57.Jiaming Song, Arash Vahdat, Morteza Mardani, and Jan Kautz. Pseudoinverse-guided diffusion models for inverse problems. In ICLR, 2023.
  58. 58.Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  59. 59.Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen, Tengfei Wang, Gang Zeng, and Ziwei Liu. Lgm: Large multi-view gaussian model for high-resolution 3d content creation. arXiv preprint arXiv:2402.05054, 2024.
  60. 60.Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
  61. 61.Igor Vasiljevic, Nick Kolkin, Shanyi Zhang, Ruotian Luo, Haochen Wang, Falcon Z Dai, Andrea F Daniele, Mohammadreza Mostajabi, Steven Basart, Matthew R Walter, et al. Diode: A dense indoor and outdoor depth dataset. arXiv preprint arXiv:1908.00463, 2019.
  62. 62.Peng Wang and Yichun Shi. Imagedream: Image-prompt multi-view diffusion for 3d generation. arXiv preprint arXiv:2312.02201, 2023.
  63. 63.Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, and Anthony Chen. Instantid: Zero-shot identity-preserving generation in seconds. arXiv preprint arXiv:2401.07519, 2024.
  64. 64.Shitao Xiao, Yueze Wang, Junjie Zhou, Huaying Yuan, Xingrun Xing, Ruiran Yan, Shuting Wang, Tiejun Huang, and Zheng Liu. Omnigen: Unified image generation. arXiv preprint arXiv:2409.11340, 2024.
  65. 65.Xingqian Xu, Zhangyang Wang, Gong Zhang, Kai Wang, and Humphrey Shi. Versatile diffusion: Text, images and variations all in one diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7754–7765, 2023.
  66. 66.Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vit-pose: Simple vision transformer baselines for human pose estimation. Advances in Neural Information Processing Systems, 35:38571–38584, 2022.
  67. 67.Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10371–10381, 2024.
  68. 68.Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv preprint arXiv:2308.06721, 2023.
  69. 69.Wei Yin, Xinlong Wang, Chunhua Shen, Yifan Liu, Zhi Tian, Songcen Xu, Changming Sun, and Dou Renyin. Diversedepth: Affine-invariant depth prediction using diverse data. arXiv preprint arXiv:2002.00569, 2020.
  70. 70.Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 204–213, 2021.
  71. 71.Chi Zhang, Wei Yin, Billzb Wang, Gang Yu, Bin Fu, and Chunhua Shen. Hierarchical normalization for robust monocular depth estimation. Advances in Neural Information Processing Systems, 35:14128–14139, 2022.
  72. 72.Jason Y Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, and Shubham Tulsiani. Cameras as rays: Pose estimation via ray diffusion. arXiv preprint arXiv:2402.14817, 2024.
  73. 73.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023.
  74. 74.Zihan Zhang, Richard Liu, Rana Hanocka, and Kfir Aberman. Tedi: Temporally-entangled diffusion for long-term motion synthesis. In ACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024.
  75. 75.Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models.(may 2023), 2023.
  76. 76.Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024.
  77. 77.Le Zhuo, Ruoyi Du, Han Xiao, Yangguang Li, Dongyang Liu, Rongjie Huang, Wenze Liu, Lirui Zhao, Fu-Yun Wang, Zhanyu Ma, et al. Lumina-next: Making lumina-t2x stronger and faster with next-dit. arXiv preprint arXiv:2406.18583, 2024.

Citation

MLA
Le, D. H., et al. “One Diffusion to Generate Them All”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 2671–82, https://doi.org/10.1109/CVPR52734.2025.00255.
APA
Le, D. H., Pham, T., Lee, S., Clark, C., Kembhavi, A., Mandt, S., Krishna, R., & Lu, J. (2025). One Diffusion to Generate Them All. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2671–2682. https://doi.org/10.1109/CVPR52734.2025.00255
Chicago
Le, D. H., T. Pham, S. Lee, et al. 2025. “One Diffusion to Generate Them All”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2671–82. https://doi.org/10.1109/CVPR52734.2025.00255.
Harvard
Le, D.H. et al. (2025) “One Diffusion to Generate Them All”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 2671–2682. Available at: https://doi.org/10.1109/CVPR52734.2025.00255.
Vancouver
1. Le DH, Pham T, Lee S, Clark C, Kembhavi A, Mandt S, Krishna R, Lu J (2025) One Diffusion to Generate Them All. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2671–2682

BibTeX

@inproceedings{Le_2025, title={One Diffusion to Generate Them All}, url={http://dx.doi.org/10.1109/CVPR52734.2025.00255}, DOI={10.1109/cvpr52734.2025.00255}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Le, Duong H. and Pham, Tuan and Lee, Sangho and Clark, Christopher and Kembhavi, Aniruddha and Mandt, Stephan and Krishna, Ranjay and Lu, Jiasen}, year={2025}, month=June, pages={2671–2682} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/