Diffusion Models in Vision: A Survey
Florinel-Alin CroitoruVlad HondruRadu Tudor IonescuMubarak Shah
Systematizes the theoretical foundations of diffusion models across probabilistic, score-based, and stochastic differential equation frameworks while analyzing their vision applications, generative trade-offs, and computational bottlenecks.
Deep generative models have seen rapid growth in artificial intelligence, with diffusion models emerging as a leading approach for synthesizing high-fidelity visual content. Understanding the theoretical foundations, operational trade-offs, and practical capabilities of these models is critical for leaders evaluating next-generation visual AI deployment. The article provides a comprehensive review and multi-perspective taxonomy of denoising diffusion models in computer vision, analyzing their mathematical formulations, structural connections to other generative paradigms, and applications across diverse vision tasks.
The article conducts a systematic literature review and comparative analysis of the diffusion modeling landscape. It formalizes three generic frameworks: Denoising Diffusion Probabilistic Models (inspired by non-equilibrium thermodynamics), Noise Conditioned Score Networks (trained via score matching to estimate data gradients), and continuous Stochastic Differential Equations, which generalize the first two formulations. The authors evaluate how these frameworks map to a wide spectrum of tasks—including unconditional image generation, conditional and text-to-image synthesis, super-resolution, inpainting, segmentation, medical imaging, and video generation—while analyzing the structural trade-offs against conventional generative models such as generative adversarial networks, variational auto-encoders, and normalizing flows.
The review establishes several core findings. First, diffusion models have matched and in many domains surpassed generative adversarial networks in sample quality, fine detail, and output diversity, notably avoiding common training instabilities like mode collapse. Second, the fundamental operating principle across all frameworks consists of a two-phase process: a forward stage that gradually degrades data into standard Gaussian noise, and a parameterized reverse stage where neural networks learn to iteratively remove the noise. Third, text-to-image models (such as Latent Diffusion Models and Imagen) exhibit high generalization capabilities, synthesizing complex, out-of-distribution concepts with minimal visual artifacts. Fourth, the primary operational bottleneck of diffusion models is poor inference speed, historically requiring hundreds to thousands of iterative network evaluations to produce a single image.
These findings have direct operational and commercial implications. The high visual quality and training stability make diffusion models dependable for critical applications, including medical anomaly detection, image super-resolution, and artistic generation. However, the high computational cost and latency at inference time present a significant barrier for real-time applications and low-latency production pipelines compared to single-step generative methods. Furthermore, text-conditioned models inherit specific failure modes from external components, such as spelling errors caused by image-text embedding encoders lacking granular character representations.
To overcome these limitations, organizations and researchers should pursue algorithmic and architectural optimizations. Key recommended strategies include adopting accelerated differential equation solvers, distilling knowledge into student networks with significantly fewer sampling steps (e.g., reducing steps by factors of 20 to 40 or distilling down to single-digit steps), and executing diffusion processes in lower-dimensional latent spaces rather than raw pixel spaces. Future development should also focus on extending diffusion representations to discriminative tasks, long-term video synthesis, and unified multi-task architectures. The findings are backed by an extensive body of recent empirical research, though practitioners should remain cautious regarding inference latency constraints and domain-specific conditioning limits.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Reading this foundational paper first is essential because the survey extensively reviews the Denoising Diffusion Probabilistic Models framework introduced here.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). This paper provides the crucial continuous-time stochastic differential equation formulation that the survey covers as one of its three core generic diffusion frameworks.
- Paper: Generative Modeling by Estimating Gradients of the Data Distribution, Yang Song et al. (2019). Understanding this noise-conditioned score network work is necessary since the survey builds directly upon score-based generative modeling as a foundational perspective.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). This paper builds directly upon the survey's foundation by unifying and streamlining the broader design space of diffusion-based generative models.
- Paper: Diffusion policy: Visuomotor policy learning via action diffusion, Cheng Chi et al. (2023). This work extends the survey's computer vision principles into the domain of robotic visuomotor policy learning via action diffusion.
- Paper: A Continuous Time Framework for Discrete Denoising Models, Andrew Campbell et al. (2022). This paper continues the survey's trajectory by generalizing diffusion models from continuous vision domains to a complete continuous-time framework for discrete data.
