keyword
Diffusion model
A diffusion model is a class of generative machine learning models that learns to produce new data by reversing a gradual corruption process. During the forward process, random noise is incrementally added to training data over a sequence of steps until the data becomes indistinguishable from pure noise. A neural network is trained to estimate and invert this corruption at each step. During generation, the model begins with pure noise and iteratively removes the predicted noise to synthesize realistic, high-fidelity data samples across diverse domains, including images, audio, and physical structures.
5 items

AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model
Teng Hu, Jiangning Zhang, Ran Yi, Yuzhen Du, Xu Chen, Liang Liu, Yabiao Wang, Chengjie Wang
Why you should read this
Proposes a few-shot diffusion framework that disentangles defect appearance from spatial masks and applies an adaptive attention re-weighting mechanism to generate realistic, precisely aligned industrial anomaly images for downstream inspection tasks.
Anomaly inspection plays an important role in industrial manufacture. Existing anomaly inspection methods are limited in their performance due to insufficient anomaly data. Although anomaly generation methods have been proposed to augment the anomaly data, they either suffer from poor generation authenticity or inaccurate alignment between the generated anomalies and masks. To address the above problems, we propose AnomalyDiffusion, a novel diffusion-based few-shot anomaly generation model, which utilizes the strong prior information of latent diffusion model learned from large-scale dataset to enhance the generation authenticity under few-shot training data. Firstly, we propose Spatial Anomaly Embedding, which consists of a learnable anomaly embedding and a spatial embedding encoded from an anomaly mask, disentangling the anomaly information into anomaly appearance and location information. Moreover, to improve the alignment between the generated anomalies and the anomaly masks, we introduce a novel Adaptive Attention Re-weighting Mechanism. Based on the disparities between the generated anomaly image and normal sample, it dynamically guides the model to focus more on the areas with less noticeable generated anomalies, enabling generation of accurately-matched anomalous image-mask pairs. Extensive experiments demonstrate that our model significantly outperforms the state-of-the-art methods in generation authenticity and diversity, and effectively improves the performance of downstream anomaly inspection tasks. The code and data are available in https://github.com/sjtuplayer/anomalydiffusion.
Added
2026-10-06

DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking
Gabriele Corso, Hannes Stärk, Bowen Jing, Regina Barzilay, Tommi Jaakkola
Why you should read this
Develops DiffDock, a diffusion generative model that treats molecular docking over translational, rotational, and torsional degrees of freedom, substantially outperforming traditional and deep learning baselines on both crystal and computationally folded protein structures.
Predicting the binding structure of a small molecule ligand to a protein -- a task known as molecular docking -- is critical to drug design. Recent deep learning methods that treat docking as a regression problem have decreased runtime compared to traditional search-based methods but have yet to offer substantial improvements in accuracy. We instead frame molecular docking as a generative modeling problem and develop DiffDock, a diffusion generative model over the non-Euclidean manifold of ligand poses. To do so, we map this manifold to the product space of the degrees of freedom (translational, rotational, and torsional) involved in docking and develop an efficient diffusion process on this space. Empirically, DiffDock obtains a 38% top-1 success rate (RMSD<2A) on PDBBind, significantly outperforming the previous state-of-the-art of traditional docking (23%) and deep learning (20%) methods. Moreover, while previous methods are not able to dock on computationally folded structures (maximum accuracy 10.4%), DiffDock maintains significantly higher precision (21.7%). Finally, DiffDock has fast inference times and provides confidence estimates with high selective accuracy.
Added
2026-10-05

Robust Classification via a Single Diffusion Model
Huanran Chen, Yinpeng Dong, Zhengyi Wang, Xiao Yang, Chengqi Duan, Hang Su, Jun Zhu
Why you should read this
Presents a generative classification framework that converts a single pre-trained diffusion model into an adversarial defender by maximizing input data likelihood and predicting class probabilities via Bayes' theorem, achieving superior accuracy against adaptive attacks and unseen threats on CIFAR-10 without adversarial training.
Despite the remarkable progress in deep learning, achieving robust classification remains challenging because models trained under empirical risk minimization tend to fit spurious features instead of intrinsic ones. In contrast, human vision exhibits robustness by adapting perception conditioned solely on target class information while effectively ignoring irrelevant contexts. Inspired by such characteristics, we propose Robust Classification via a Single Diffusion Model (~RCSD~), which improves robust accuracy against various distribution shifts using just *one* diffusion model pre-trained on training domain data; unlike prior approaches requiring multiple generative classifiers matched to each possible corrupted label distributions.The key insight lies in how ~RCSD~ guides reverse sampling toward desired labels through Bayes' rule – substituting missing likelihood estimates typically obtained separately across all candidate categories–thereby reducing computational burdens substantially.A theoretical analysis proves our method generates samples belonging closer together around true decision boundaries than standard multi-classifier alternatives given enough time steps .Experiments demonstrate state-of-the art performance improvements over existing algorithms benchmark datasets CIFAR10-C CIFAR100C ImageNetC alongside five real-world applications further validating effectiveness proposed approach
Added
2026-10-02

FiT: Flexible Vision Transformer for Diffusion Model
Zeyu Lu, Zidong Wang, Di Huang, Chengyue Wu, Xihui Liu, Wanli Ouyang, Lei Bai
Why you should read this
Proposes a flexible vision transformer architecture that treats images as variable-length token sequences using 2D rotary positional embeddings, enabling diffusion models to generate high-fidelity images across arbitrary resolutions and aspect ratios without cropping.
Nature is infinitely resolution-free. In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically designed for generating images with unrestricted resolutions and aspect ratios. Unlike traditional methods that perceive images as static-resolution grids, FiT conceptualizes images as sequences of dynamically-sized tokens. This perspective enables a flexible training strategy that effortlessly adapts to diverse aspect ratios during both training and inference phases, thus promoting resolution generalization and eliminating biases induced by image cropping. Enhanced by a meticulously adjusted network structure and the integration of training-free extrapolation techniques, FiT exhibits remarkable flexibility in resolution extrapolation generation. Comprehensive experiments demonstrate the exceptional performance of FiT across a broad range of resolutions. Repository available at https://github.com/whlzy/FiT.
Added
2026-09-26

DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, Bryan Catanzaro
Why you should read this
Applies bidirectional diffusion to raw audio waveforms to match the quality of autoregressive vocoders with parallel sampling.
In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured waveform through a Markov chain with a constant number of steps at synthesis. It is efficiently trained by optimizing a variant of variational bound on the data likelihood. DiffWave produces high-fidelity audios in different waveform generation tasks, including neural vocoding conditioned on mel spectrogram, class-conditional generation, and unconditional generation. We demonstrate that DiffWave matches a strong WaveNet vocoder in terms of speech quality (MOS: 4.44 versus 4.43), while synthesizing orders of magnitude faster. In particular, it significantly outperforms autoregressive and GAN-based waveform models in the challenging unconditional generation task in terms of audio quality and sample diversity from various automatic and human evaluations.
Added
2026-02-25
