keyword
conditional diffusion
Conditional diffusion is a generative machine learning framework that synthesizes new data samples conforming to specific conditioning inputs, such as class labels, text prompts, or reference images, by steering a progressive denoising process. Unlike unconditional diffusion models that generate data from an unconstrained distribution, conditional diffusion models learn the conditional probability distribution of target data given an external guiding signal. During generation, the model iteratively removes noise from an initially random representation while incorporating the conditioning information into the neural network architecture or through sampling guidance mechanisms, such as classifier guidance and classifier-free guidance. This enables precise control over the generated content, making conditional diffusion widely applicable to multimodal synthesis tasks, including text-to-image generation, image-to-image translation, and solving inverse problems.
10 items

Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance
Xinyu Peng, Ziyang Zheng, Wenrui Dai, Nuoqian Xiao, Chenglin Li, Junni Zou, Hongkai Xiong
Why you should read this
Unifies zero-shot diffusion solvers for inverse problems under a posterior covariance framework and derives maximum-likelihood-optimized covariance estimators that boost image reconstruction quality without manual hyperparameter tuning.
Recent diffusion models provide a promising zero-shot solution to noisy linear inverse problems without retraining for specific inverse problems. In this paper, we reveal that recent methods can be uniformly interpreted as employing a Gaussian approximation with hand-crafted isotropic covariance for the intractable denoising posterior to approximate the conditional posterior mean. Inspired by this finding, we propose to improve recent methods by using more principled covariance determined by maximum likelihood estimation. To achieve posterior covariance optimization without retraining, we provide general plug-and-play solutions based on two approaches specifically designed for leveraging pre-trained models with and without reverse covariance. We further propose a scalable method for learning posterior covariance prediction based on representation with orthonormal basis. Experimental results demonstrate that the proposed methods significantly enhance reconstruction performance without requiring hyperparameter tuning.
Added
2026-10-04

Guidance with Spherical Gaussian Constraint for Conditional Diffusion
Lingxiao Yang, Shutong Ding, Yifan Cai, Jingyi Yu, Jingya Wang, Ye Shi
Why you should read this
Proposes a training-free guidance method using spherical Gaussian constraints to resolve intermediate trajectory drift in conditional diffusion models, providing a closed-form solution that allows larger step sizes and faster sampling with minimal code changes.
Added
2026-10-02

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale
Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, Jun Zhu
Why you should read this
Proposes UniDiffuser, a single transformer-based framework that captures marginal, conditional, and joint multi-modal distributions to handle diverse generation tasks—including text-to-image, image-to-text, and paired generation—without requiring task-specific models or extra computational overhead.
This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoke models (e.g., Stable Diffusion and DALL·E 2) in representative tasks (e.g., text-to-image generation). Our code is available at https://github.com/thu-ml/unidiffuser.
Added
2026-09-28

Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models
Seyedmorteza Sadat, Otmar Hilliges, Romann M. Weber
Why you should read this
Proposes Adaptive Projected Guidance, a plug-and-play modification to classifier-free guidance that decomposes diffusion updates to eliminate oversaturation and visual artifacts at high guidance scales with negligible computational overhead.
Classifier-free guidance (CFG) is crucial for improving both generation quality and alignment between the input condition and final output in diffusion models. While a high guidance scale is generally required to enhance these aspects, it also causes oversaturation and unrealistic artifacts. In this paper, we revisit the CFG update rule and introduce modifications to address this issue. We first decompose the update term in CFG into parallel and orthogonal components with respect to the conditional model prediction and observe that the parallel component primarily causes oversaturation, while the orthogonal component enhances image quality. Accordingly, we propose down-weighting the parallel component to achieve high-quality generations without oversaturation. Additionally, we draw a connection between CFG and gradient ascent and introduce a new rescaling and momentum method for the CFG update rule based on this insight. Our approach, termed adaptive projected guidance (APG), retains the quality-boosting advantages of CFG while enabling the use of higher guidance scales without oversaturation. APG is easy to implement and introduces practically no additional computational overhead to the sampling process. Through extensive experiments, we demonstrate that APG is compatible with various conditional diffusion models and samplers, leading to improved FID, recall, and saturation scores while maintaining precision comparable to CFG, making our method a superior plug-and-play alternative to standard classifier-free guidance.
Added
2026-09-26

Diffusion Models for Black-Box Optimization
Siddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya Grover
Why you should read this
Proposes Denoising Diffusion Optimization Models, an inverse approach that pairs conditional diffusion with objective reweighting and classifier-free guidance to generate candidate optima that surpass the best observations in offline datasets across diverse continuous and discrete benchmarks.
Added
2026-09-26

Conditional Text Image Generation with Diffusion Models
Yuanzhi Zhu, Zhaohai Li, Tianwei Wang, Mengchao He, Cong Yao
Why you should read this
Proposes a conditional diffusion model that controls text, style, and visual attributes across four generation modes to synthesize realistic scene and handwritten text images that improve downstream recognition accuracy and handle out-of-vocabulary words.
Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore the problem of text image generation, by taking advantage of the powerful abilities of Diffusion Models in generating photo-realistic and diverse image samples with given conditions, and propose a method called Conditional Text Image Generation with Diffusion Models (CTIG-DM for short). To conform to the characteristics of text images, we devise three conditions: image condition, text condition, and style condition, which can be used to control the attributes, contents, and styles of the samples in the image generation process. Specifically, four text image generation modes, namely: (1) synthesis mode, (2) augmentation mode, (3) recovery mode, and (4) imitation mode, can be derived by combining and configuring these three conditions. Extensive experiments on both handwritten and scene text demonstrate that the proposed CTIG-DM is able to produce image samples that simulate real-world complexity and diversity, and thus can boost the performance of existing text recognizers. Besides, CTIG-DM shows its appealing potential in domain adaptation and generating images containing Out-Of-Vocabulary (OOV) words.
Added
2026-09-26

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models
Guanhua Zhang, Jiabao Ji, Yang Zhang, Mo Yu, Tommi S. Jaakkola, Shiyu Chang
Why you should read this
Proposes CoPaint, a Bayesian framework for diffusion-based image inpainting that jointly modifies revealed and unrevealed regions to eliminate incoherence while driving approximation errors to zero to strictly match reference constraints.
Image inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions, but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics. The codes are available at https://github.com/UCSB-NLP-Chang/CoPaint/.
Added
2026-09-26

Palette: Image-to-Image Diffusion Models
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David J. Fleet, Mohammad Norouzi
Why you should read this
Develops a unified conditional diffusion framework that outperforms task-specific GAN baselines across colorization, inpainting, uncropping, and restoration without specialized architectures, auxiliary losses, or hyperparameter tuning.
This paper develops a unified framework for image-to-image translation based on conditional diffusion models and evaluates this framework on four challenging image-to-image translation tasks, namely colorization, inpainting, uncropping, and JPEG restoration. Our simple implementation of image-to-image diffusion models outperforms strong GAN and regression baselines on all tasks, without task-specific hyper-parameter tuning, architecture customization, or any auxiliary loss or sophisticated new techniques needed. We uncover the impact of an L2 vs. L1 loss in the denoising diffusion objective on sample diversity, and demonstrate the importance of self-attention in the neural architecture through empirical studies. Importantly, we advocate a unified evaluation protocol based on ImageNet, with human evaluation and sample quality scores (FID, Inception Score, Classification Accuracy of a pre-trained ResNet-50, and Perceptual Distance against original images). We expect this standardized evaluation protocol to play a role in advancing image-to-image translation research. Finally, we show that a generalist, multi-task diffusion model performs as well or better than task-specific specialist counterparts. Check out this https URL for an overview of the results.
Added
2026-09-15

Diffusion Models Beat GANs on Image Synthesis
Prafulla Dhariwal, Alex Nichol
Why you should read this
Demonstrates that diffusion models can achieve superior image sample quality compared to state-of-the-art GANs while maintaining better distribution coverage and offering practical advancements like classifier guidance.
We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. We achieve this on unconditional image synthesis by finding a better architecture through a series of ablations. For conditional image synthesis, we further improve sample quality with classifier guidance: a simple, compute-efficient method for trading off diversity for fidelity using gradients from a classifier. We achieve an FID of 2.97 on ImageNet 128128, 4.59 on ImageNet 256256, and 7.72 on ImageNet 512512, and we match BigGAN-deep even with as few as 25 forward passes per sample, all while maintaining better coverage of the distribution. Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.94 on ImageNet 256256 and 3.85 on ImageNet 512512. We release our code at this https URL
Added
2026-03-26
License
Published with permission

Classifier-Free Diffusion Guidance
Jonathan Ho, Tim Salimans
Why you should read this
Introduces a training technique to condition generation by balancing unconditional and conditional score estimates, eliminating the need for external classifiers.
Classifier guidance is a recently introduced method to trade off mode coverage and sample fidelity in conditional diffusion models post training, in the same spirit as low temperature sampling or truncation in other types of generative models. Classifier guidance combines the score estimate of a diffusion model with the gradient of an image classifier and thereby requires training an image classifier separate from the diffusion model. It also raises the question of whether guidance can be performed without a classifier. We show that guidance can be indeed performed by a pure generative model without such a classifier: in what we call classifier-free guidance, we jointly train a conditional and an unconditional diffusion model, and we combine the resulting conditional and unconditional score estimates to attain a trade-off between sample quality and diversity similar to that obtained using classifier guidance.
Added
2026-03-26
License
Published with permission
