Built independently by an author, for readers. Read the story and support ChapterPal

keyword

conditional diffusion

Conditional diffusion is a generative machine learning framework that synthesizes new data samples conforming to specific conditioning inputs, such as class labels, text prompts, or reference images, by steering a progressive denoising process. Unlike unconditional diffusion models that generate data from an unconstrained distribution, conditional diffusion models learn the conditional probability distribution of target data given an external guiding signal. During generation, the model iteratively removes noise from an initially random representation while incorporating the conditioning information into the neural network architecture or through sampling guidance mechanisms, such as classifier guidance and classifier-free guidance. This enables precise control over the generated content, making conditional diffusion widely applicable to multimodal synthesis tasks, including text-to-image generation, image-to-image translation, and solving inverse problems.

10 items

Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Xinyu Peng, Ziyang Zheng, Wenrui Dai, Nuoqian Xiao, Chenglin Li, Junni Zou, Hongkai Xiong

OrganizationsSchool of Electronic Information and Electrical EngineeringShanghai Jiao Tong University

Why you should read this

Unifies zero-shot diffusion solvers for inverse problems under a posterior covariance framework and derives maximum-likelihood-optimized covariance estimators that boost image reconstruction quality without manual hyperparameter tuning.

Recent diffusion models provide a promising zero-shot solution to noisy linear inverse problems without retraining for specific inverse problems. In this paper, we reveal that recent methods can be uniformly interpreted as employing a Gaussian approximation with hand-crafted isotropic covariance for the intractable denoising posterior to approximate the conditional posterior mean. Inspired by this finding, we propose to improve recent methods by using more principled covariance determined by maximum likelihood estimation. To achieve posterior covariance optimization without retraining, we provide general plug-and-play solutions based on two approaches specifically designed for leveraging pre-trained models with and without reverse covariance. We further propose a scalable method for learning posterior covariance prediction based on representation with orthonormal basis. Experimental results demonstrate that the proposed methods significantly enhance reconstruction performance without requiring hyperparameter tuning.

Added

2026-10-04

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, Jun Zhu

OrganizationsBeijing Academy of Artificial IntelligencePazhou Laboratory (Huangpu)Renmin University of ChinaShengshu TechnologyTsinghua University

Why you should read this

Proposes UniDiffuser, a single transformer-based framework that captures marginal, conditional, and joint multi-modal distributions to handle diverse generation tasks—including text-to-image, image-to-text, and paired generation—without requiring task-specific models or extra computational overhead.

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i.e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e.g., the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoke models (e.g., Stable Diffusion and DALL·E 2) in representative tasks (e.g., text-to-image generation). Our code is available at https://github.com/thu-ml/unidiffuser.

Added

2026-09-28

Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models

Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models

Seyedmorteza Sadat, Otmar Hilliges, Romann M. Weber

OrganizationsDisney ResearchETH Zurich

Why you should read this

Proposes Adaptive Projected Guidance, a plug-and-play modification to classifier-free guidance that decomposes diffusion updates to eliminate oversaturation and visual artifacts at high guidance scales with negligible computational overhead.

Classifier-free guidance (CFG) is crucial for improving both generation quality and alignment between the input condition and final output in diffusion models. While a high guidance scale is generally required to enhance these aspects, it also causes oversaturation and unrealistic artifacts. In this paper, we revisit the CFG update rule and introduce modifications to address this issue. We first decompose the update term in CFG into parallel and orthogonal components with respect to the conditional model prediction and observe that the parallel component primarily causes oversaturation, while the orthogonal component enhances image quality. Accordingly, we propose down-weighting the parallel component to achieve high-quality generations without oversaturation. Additionally, we draw a connection between CFG and gradient ascent and introduce a new rescaling and momentum method for the CFG update rule based on this insight. Our approach, termed adaptive projected guidance (APG), retains the quality-boosting advantages of CFG while enabling the use of higher guidance scales without oversaturation. APG is easy to implement and introduces practically no additional computational overhead to the sampling process. Through extensive experiments, we demonstrate that APG is compatible with various conditional diffusion models and samplers, leading to improved FID, recall, and saturation scores while maintaining precision comparable to CFG, making our method a superior plug-and-play alternative to standard classifier-free guidance.

Added

2026-09-26

Conditional Text Image Generation with Diffusion Models

Conditional Text Image Generation with Diffusion Models

Yuanzhi Zhu, Zhaohai Li, Tianwei Wang, Mengchao He, Cong Yao

OrganizationsAlibaba Group

Why you should read this

Proposes a conditional diffusion model that controls text, style, and visual attributes across four generation modes to synthesize realistic scene and handwritten text images that improve downstream recognition accuracy and handle out-of-vocabulary words.

Current text recognition systems, including those for handwritten scripts and scene text, have relied heavily on image synthesis and augmentation, since it is difficult to realize real-world complexity and diversity through collecting and annotating enough real text images. In this paper, we explore the problem of text image generation, by taking advantage of the powerful abilities of Diffusion Models in generating photo-realistic and diverse image samples with given conditions, and propose a method called Conditional Text Image Generation with Diffusion Models (CTIG-DM for short). To conform to the characteristics of text images, we devise three conditions: image condition, text condition, and style condition, which can be used to control the attributes, contents, and styles of the samples in the image generation process. Specifically, four text image generation modes, namely: (1) synthesis mode, (2) augmentation mode, (3) recovery mode, and (4) imitation mode, can be derived by combining and configuring these three conditions. Extensive experiments on both handwritten and scene text demonstrate that the proposed CTIG-DM is able to produce image samples that simulate real-world complexity and diversity, and thus can boost the performance of existing text recognizers. Besides, CTIG-DM shows its appealing potential in domain adaptation and generating images containing Out-Of-Vocabulary (OOV) words.

Added

2026-09-26

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Guanhua Zhang, Jiabao Ji, Yang Zhang, Mo Yu, Tommi S. Jaakkola, Shiyu Chang

OrganizationsIBMMassachusetts Institute of TechnologyMIT-IBM Watson AI LabUniversity of California, Santa Barbara

Why you should read this

Proposes CoPaint, a Bayesian framework for diffusion-based image inpainting that jointly modifies revealed and unrevealed regions to eliminate incoherence while driving approximation errors to zero to strictly match reference constraints.

Image inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions, but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics. The codes are available at https://github.com/UCSB-NLP-Chang/CoPaint/.

Added

2026-09-26

Palette: Image-to-Image Diffusion Models

Palette: Image-to-Image Diffusion Models

Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David J. Fleet, Mohammad Norouzi

OrganizationsGoogle

Why you should read this

Develops a unified conditional diffusion framework that outperforms task-specific GAN baselines across colorization, inpainting, uncropping, and restoration without specialized architectures, auxiliary losses, or hyperparameter tuning.

This paper develops a unified framework for image-to-image translation based on conditional diffusion models and evaluates this framework on four challenging image-to-image translation tasks, namely colorization, inpainting, uncropping, and JPEG restoration. Our simple implementation of image-to-image diffusion models outperforms strong GAN and regression baselines on all tasks, without task-specific hyper-parameter tuning, architecture customization, or any auxiliary loss or sophisticated new techniques needed. We uncover the impact of an L2 vs. L1 loss in the denoising diffusion objective on sample diversity, and demonstrate the importance of self-attention in the neural architecture through empirical studies. Importantly, we advocate a unified evaluation protocol based on ImageNet, with human evaluation and sample quality scores (FID, Inception Score, Classification Accuracy of a pre-trained ResNet-50, and Perceptual Distance against original images). We expect this standardized evaluation protocol to play a role in advancing image-to-image translation research. Finally, we show that a generalist, multi-task diffusion model performs as well or better than task-specific specialist counterparts. Check out this https URL for an overview of the results.

Added

2026-09-15

Diffusion Models Beat GANs on Image Synthesis

Diffusion Models Beat GANs on Image Synthesis

Prafulla Dhariwal, Alex Nichol

OrganizationsOpenAI

Why you should read this

Demonstrates that diffusion models can achieve superior image sample quality compared to state-of-the-art GANs while maintaining better distribution coverage and offering practical advancements like classifier guidance.

We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. We achieve this on unconditional image synthesis by finding a better architecture through a series of ablations. For conditional image synthesis, we further improve sample quality with classifier guidance: a simple, compute-efficient method for trading off diversity for fidelity using gradients from a classifier. We achieve an FID of 2.97 on ImageNet 128×\times128, 4.59 on ImageNet 256×\times256, and 7.72 on ImageNet 512×\times512, and we match BigGAN-deep even with as few as 25 forward passes per sample, all while maintaining better coverage of the distribution. Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.94 on ImageNet 256×\times256 and 3.85 on ImageNet 512×\times512. We release our code at this https URL

Added

2026-03-26

License

Published with permission

Classifier-Free Diffusion Guidance

Classifier-Free Diffusion Guidance

Jonathan Ho, Tim Salimans

OrganizationsGoogle

Why you should read this

Introduces a training technique to condition generation by balancing unconditional and conditional score estimates, eliminating the need for external classifiers.

Classifier guidance is a recently introduced method to trade off mode coverage and sample fidelity in conditional diffusion models post training, in the same spirit as low temperature sampling or truncation in other types of generative models. Classifier guidance combines the score estimate of a diffusion model with the gradient of an image classifier and thereby requires training an image classifier separate from the diffusion model. It also raises the question of whether guidance can be performed without a classifier. We show that guidance can be indeed performed by a pure generative model without such a classifier: in what we call classifier-free guidance, we jointly train a conditional and an unconditional diffusion model, and we combine the resulting conditional and unconditional score estimates to attain a trade-off between sample quality and diversity similar to that obtained using classifier guidance.

Added

2026-03-26

License

Published with permission