Robust Classification via a Single Diffusion Model

Huanran ChenYinpeng DongZhengyi WangXiao YangChengqi DuanHang SuJun Zhu

article2024ICML108 citations

Presents a generative classification framework that converts a single pre-trained diffusion model into an adversarial defender by maximizing input data likelihood and predicting class probabilities via Bayes' theorem, achieving superior accuracy against adaptive attacks and unseen threats on CIFAR-10 without adversarial training.

Listen

Deep learning models are widely vulnerable to adversarial attacks—subtle, malicious modifications to inputs that trigger severe classification errors. This vulnerability poses critical safety and security risks across high-stakes domains such as automated driving, biometric facial recognition, and medical diagnostics. Prevailing defense strategies, including adversarial training and diffusion-based noise purification, suffer from fundamental trade-offs: adversarial training struggles to defend against threat types not encountered during training, while purification techniques remain vulnerable to strong adaptive attacks that exploit the downstream classifier.

The article introduces and evaluates the Robust Diffusion Classifier (RDC), a framework that converts a single pre-trained diffusion model into an inherently robust generative classifier. Rather than relying on standard discriminative prediction or fragile preprocessing pipelines, the article demonstrates how to directly estimate class probabilities from data density models while optimizing computational efficiency.

The evaluated approach uses a two-stage mechanism. First, an input image undergoes Likelihood Maximization, an optimization step that shifts the data into regions of higher estimated likelihood within a constrained distance. Next, the model computes the class probabilities of the optimized input via Bayes' theorem by evaluating the conditional likelihood across classes. To make this computationally feasible, the authors introduce a multi-head diffusion architecture that outputs noise predictions for all target classes in a single forward pass, substantially reducing computational evaluations.

Key experimental evaluations on standard image benchmarks demonstrate substantial performance gains. On the CIFAR-10 benchmark under standard adaptive attacks, RDC achieved a 75.67% robust accuracy, outperforming the previous state-of-the-art adversarial training model by 4.77 percentage points and dynamic defenses by 3.01 percentage points. Under unseen threat categories, RDC maintained superior generalizability, improving average robust accuracy across multiple attack types by more than 30 percentage points over baseline models. Additionally, comprehensive gradient analyses confirmed that these robustness gains stem from true defensive capability rather than flawed evaluations or masked gradients, while experiments on CIFAR-100 and Restricted ImageNet showed consistent performance advantages over existing defenses.

These findings suggest that generative classifiers built directly from diffusion models offer a fundamentally stronger defense paradigm than traditional discriminative classifiers. In operational settings, RDC mitigates critical security risks against unexpected or adaptive attacks without requiring retraining for every specific threat type. The primary operational trade-off is computational overhead: generating predictions with RDC requires multiple network evaluations per image, making inference slower than conventional classifiers.

Decision-makers and engineering teams seeking resilient computer vision systems should consider generative diffusion architectures as viable defenses for security-critical pipelines, especially where threat vectors are diverse and unpredictable. Before broad deployment, technical teams should conduct pilot testing to measure the latency impact within target applications and explore advanced acceleration techniques—such as consistency models or distilled sampling—to minimize real-time processing costs.

Confidence in these findings is high regarding core image classification benchmarks, supported by rigorous adaptive evaluations. However, readers should note limitations regarding sample scope: high computational costs restricted certain adversarial attack evaluations to representative subsets of test sets. Further validation on full-scale, complex enterprise datasets is warranted as efficiency optimizations mature.

Cover for Robust Classification via a Single Diffusion Model

Abstract

Despite the remarkable progress in deep learning, achieving robust classification remains challenging because models trained under empirical risk minimization tend to fit spurious features instead of intrinsic ones. In contrast, human vision exhibits robustness by adapting perception conditioned solely on target class information while effectively ignoring irrelevant contexts. Inspired by such characteristics, we propose Robust Classification via a Single Diffusion Model (RCSD), which improves robust accuracy against various distribution shifts using just one diffusion model pre-trained on training domain data; unlike prior approaches requiring multiple generative classifiers matched to each possible corrupted label distributions.The key insight lies in how RCSD guides reverse sampling toward desired labels through Bayes' rule – substituting missing likelihood estimates typically obtained separately across all candidate categories–thereby reducing computational burdens substantially.A theoretical analysis proves our method generates samples belonging closer together around true decision boundaries than standard multi-classifier alternatives given enough time steps .Experiments demonstrate state-of-the art performance improvements over existing algorithms benchmark datasets CIFAR10-C CIFAR100C ImageNetC alongside five real-world applications further validating effectiveness proposed approach

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methodology
  • 3.1. Preliminary: diffusion models
  • 3.2. Diffusion model for classification
  • 3.3. Robustness analysis under the optimal setting
  • 3.4. Likelihood maximization
  • 3.5. Time complexity reduction
  • 4. Experiments
  • 4.1. Experimental settings
  • 4.2. Comparison with the state-of-the-art
  • 4.3. Defense against unseen threats
  • 4.4. Evaluation of gradient obfuscation
  • 4.5. Ablation studies
  • 5. Conclusion
  • Acknowledgements
  • References
  • A. Proofs and Derivations
  • A.1. Proof of Theorem 3.1
  • A.2. Derivation of the optimal diffusion classifier
  • A.2.1. OPTIMAL DIFFUSION MODEL: PROOF OF THEOREM 3.2
  • A.2.2. OPTIMAL DIFFUSION CLASSIFIER: PROOF OF THEOREM 3.3
  • A.3. Derivation of conditional elbos in Eq. (5)
  • A.4. Connection between Energy-Based Models (EBMs)
  • A.5. Computing gradient without computing UNet jacobi
  • B. More experimental results
  • B.1. Training details
  • B.2. More Analysis and Discussion
  • B.3. Experiment on Restricted ImageNet
  • B.4. Experiment on CIFAR-100
  • B.5. Discussions.
  • C. Limitations

Knowls

  1. Knowl 1 — Diffusion likelihoods yield class probabilities

    model/method

    Robust Diffusion Classifier (RDC) turns a class-conditional diffusion model into a generative classifier. For an input image xx, class y∈{1,…,K}y\in\{1,\ldots,K\}, and class-conditional noise predictor ϵθ(xt,t,y)\epsilon_\theta(x_t,t,y), define the diffusion loss Ly(x)=Et,ϵ[wt∥ϵθ(xt,t,y)−ϵ∥22]L_y(x)=\mathbb{E}_{t,\epsilon}[w_t\|\epsilon_\theta(x_t,t,y)-\epsilon\|_2^2]. Here tt is a diffusion timestep, ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I), xt=αtx+σtϵx_t=\sqrt{\alpha_t}x+\sigma_t\epsilon, and wtw_t is the diffusion-loss weight. The loss approximates the negative conditional log-likelihood −log⁡pθ(x∣y)-\log p_\theta(x\mid y) up to an additive constant and the gap between the likelihood and its variational lower bound. With a uniform class prior and a negligible gap for every class, Bayes’ rule gives

    pθ(y∣x)=exp⁡[−Ly(x)]∑j=1Kexp⁡[−Lj(x)].p_\theta(y\mid x)=\frac{\exp[-L_y(x)]}{\sum_{j=1}^{K}\exp[-L_j(x)]}.

    Thus, RDC predicts the class with the lowest conditional diffusion loss. For a nonuniform prior, the class logit additionally includes log⁡p(y)\log p(y). In practice, the likelihood-gap condition is an approximation rather than a guarantee.

  2. Knowl 2 — Likelihood maximization pre-optimizes the input

    model/method

    Before classifying an input, RDC optionally moves it toward a region assigned higher likelihood by the unconditional diffusion model. Given input image xx, the method finds an optimized image x^\hat{x} by minimizing unconditional diffusion loss subject to an ℓ∞\ell_\infty displacement budget η\eta:

    x^∈arg⁡min⁡z: ∥z−x∥∞≤η  Et,ϵ ⁣[wt∥ϵθ(zt,t)−ϵ∥22],zt=αtz+σtϵ.\hat{x}\in\arg\min_{z:\,\|z-x\|_\infty\leq\eta}\;\mathbb{E}_{t,\epsilon}\!\left[w_t\|\epsilon_\theta(z_t,t)-\epsilon\|_2^2\right],\qquad z_t=\sqrt{\alpha_t}z+\sigma_t\epsilon.

    The objective maximizes a variational lower bound on the data log-likelihood without requiring the input’s label. The budget is intended to prevent the optimization from moving the image into another class’s region. The paper’s rationale is that an adversarial image usually remains near its ground-truth-class example, so moving toward higher unconditional likelihood may also improve its conditional likelihood for that class. The optimization is performed with a finite number of gradient-based steps; the paper does not claim that this necessarily finds a global likelihood maximum.

  3. Knowl 3 — RDC inference procedure and default optimization settings

    algorithm

    RDC takes an image xx and a diffusion noise-prediction model, uses likelihood maximization to obtain x^\hat{x}, then predicts the class from conditional diffusion losses. In the reported CIFAR-10 setup, likelihood maximization uses N=5N=5 steps, momentum decay μ=1\mu=1, budget η=8/255\eta=8/255, and step size γ=0.1\gamma=0.1. Each optimization step estimates its gradient using one randomly sampled timestep and one Gaussian noise sample. Classification estimates the loss across all TT timesteps using one noise sample per timestep, with all class predictions evaluated simultaneously by the multi-head model.

    Input: Image x; diffusion model; optimization budget η; step size γ; steps N; momentum decay μ
    Initialize x_hat = x and m = 0
    For n = 1 to N:
        Sample timestep t uniformly from {1, ..., T} and noise ε from N(0, I)
        Form x_hat_t = sqrt(α_t) x_hat + σ_t ε
        Estimate g = gradient with respect to x_hat of w_t ||ε_θ(x_hat_t, t) - ε||_2^2
        Set m = μ m - g / ||g||_1
        Project x_hat + γ m onto the ℓ∞ ball of radius η centered at x
    For each class y, estimate L_y(x_hat) by averaging its weighted noise-prediction error over all T timesteps
    Set p(y | x_hat) = softmax_y(-L_y(x_hat))
    Return the class with greatest p(y | x_hat)

    The projection in the update enforces ∥x^−x∥∞≤η\|\hat{x}-x\|_\infty\leq\eta. The final class-loss calculation uses the uniform-prior form of the diffusion classifier; a nonuniform prior can be added to the logits.

  4. Knowl 4 — Multi-head diffusion reduces class-wise inference cost

    model/method

    A conventional conditional diffusion network predicts noise for one class at a time, so evaluating all KK class losses over TT timesteps requires K×TK\times T network function evaluations (NFEs). RDC’s multi-head diffusion modifies the final convolution of a UNet to output noise predictions for all classes in parallel; for RGB images, the output has K×3K\times3 channels. This reduces the classifier’s cost to TT NFEs, independent of the number of classes. The paper distills this multi-head network from a pretrained conditional diffusion model by matching its noise predictions for sampled noisy images, timesteps, and labels; this avoids adversarial-example-specific training. With NN likelihood-maximization iterations, the resulting RDC uses N+TN+T NFEs per image. In the reported CIFAR-10 timing on one NVIDIA 3090 GPU, the N+TN+T configuration took 1.43 seconds per image, compared with 9.86 seconds for the N+TKN+T K configuration.

  5. Knowl 5 — The optimal diffusion classifier has class-prototype distance logits

    theoretical result

    Let DyD_y be the finite set of data examples in class yy, and suppose the examples within each class are equally weighted. For noise level tt, the optimal noise predictor for minimizing squared noise-prediction error is a posterior-weighted average over class examples:

    ϵD∗(xt,t,y)=∑x(i)∈Dys(xt,x(i))xt−αtx(i)σt,s(xt,x(i))=exp⁡ ⁣[−∥xt−αtx(i)∥222σt2]∑x(j)∈Dyexp⁡ ⁣[−∥xt−αtx(j)∥222σt2].\epsilon_D^*(x_t,t,y)=\sum_{x^{(i)}\in D_y}s(x_t,x^{(i)})\frac{x_t-\sqrt{\alpha_t}x^{(i)}}{\sigma_t},\qquad s(x_t,x^{(i)})=\frac{\exp\!\left[-\frac{\|x_t-\sqrt{\alpha_t}x^{(i)}\|_2^2}{2\sigma_t^2}\right]}{\sum_{x^{(j)}\in D_y}\exp\!\left[-\frac{\|x_t-\sqrt{\alpha_t}x^{(j)}\|_2^2}{2\sigma_t^2}\right]}.

    Here xtx_t is a noisy image, x(i)x^{(i)} and x(j)x^{(j)} are examples in DyD_y, and αt,σt\alpha_t,\sigma_t are the diffusion signal and noise scales. Substituting this optimal predictor into the diffusion classifier yields a softmax over class logits

    fD∗(x)y=−Et,ϵ ⁣[αtσt2∥∑x(i)∈Dys(x,x(i),ϵ,t)(x−x(i))∥22],f_D^*(x)_y=-\mathbb{E}_{t,\epsilon}\!\left[\frac{\alpha_t}{\sigma_t^2}\left\|\sum_{x^{(i)}\in D_y}s(x,x^{(i)},\epsilon,t)(x-x^{(i)})\right\|_2^2\right],

    where xx is the input, ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I), xt=αtx+σtϵx_t=\sqrt{\alpha_t}x+\sigma_t\epsilon, and s(x,x(i),ϵ,t)s(x,x^{(i)},\epsilon,t) is the same normalized exponential weight as above with noisy point xtx_t. The optimal classifier therefore scores a class by the negative expected squared norm of a noise-dependent weighted displacement from the input to examples of that class; the expectation weights diffusion levels through αt/σt2\alpha_t/\sigma_t^2.

  6. Knowl 6 — The idealized classifier is robust, but trained density estimates fall short

    empirical result

    On the paper’s CIFAR-10 robustness evaluation, the classifier formed from the optimal diffusion predictor achieved 100% robust accuracy under AutoAttack for both an ℓ∞\ell_\infty threat with ϵ∞=8/255\epsilon_\infty=8/255 and an ℓ2\ell_2 threat with ϵ2=0.5\epsilon_2=0.5. The trained diffusion classifier without likelihood maximization achieved 35.94% and 76.95% robust accuracy under those respective threats. The paper attributes this gap to practical models’ inaccurate conditional-density estimates and/or a non-negligible gap between log-likelihood and its variational lower bound. The 100% figure is the reported evaluation of the idealized classifier, not a general guarantee for trained diffusion networks.

  7. Knowl 7 — CIFAR-10 results show robustness across norm and semantic threats

    data/table

    The paper evaluates 512 randomly selected CIFAR-10 test images. Robust accuracy is measured against AutoAttack under ℓ∞\ell_\infty perturbations with ϵ∞=8/255\epsilon_\infty=8/255 and ℓ2\ell_2 perturbations with ϵ2=0.5\epsilon_2=0.5, and against StAdv spatial transformations using 100 steps and bound 0.05. “Avg” is the mean of the three robust-accuracy columns. The results show that RDC improves over the strongest listed norm-specific adversarial-training baseline on both norm threats and retains high accuracy against the unseen spatial threat.

    Method Clean Acc. ℓ∞\ell_\infty ℓ2\ell_2 StAdv Avg.
    AT-EDM-ℓ∞\ell_\infty 93.36 70.90 69.73 2.93 47.85
    AT-EDM-ℓ2\ell_2 95.90 53.32 84.77 5.08 47.72
    DiffPure (t∗=0.1t^*=0.1) 90.97 44.53 72.65 12.89 43.35
    JEM 92.90 8.20 26.37 0.05 11.54
    LM (ours) 87.89 71.68 75.00 87.50 78.06
    DC (ours) 93.55 35.94 76.95 93.55 68.81
    RDC (ours) 89.85 75.67 82.03 89.45 82.38

    All entries are percentages. RDC’s 75.67%75.67\% ℓ∞\ell_\infty robust accuracy is 4.77 percentage points above AT-EDM-ℓ∞\ell_\infty; its 82.03%82.03\% ℓ2\ell_2 accuracy is also higher than the listed ℓ2\ell_2-trained baseline’s 84.77%84.77\%? No: AT-EDM-ℓ2\ell_2 is higher at 84.77%84.77\%. RDC’s principal cross-threat advantage is its much higher StAdv accuracy and average across the three threats.

  8. Knowl 8 — Adaptive-attack checks support the robustness evaluation

    empirical result

    Because attacking likelihood maximization requires differentiating through an optimization procedure, the paper checks whether its reported robustness is an artifact of gradient obfuscation. With five likelihood-maximization steps, RDC obtained 75.67% robust accuracy under BPDA and 77.54% under a Lagrange adaptive attack. With one optimization step, exact-gradient and BPDA attacks gave 69.53% and 69.92%, respectively; exact gradients were only evaluated at one step because of memory costs. The iterative updates in these evaluations were conducted by AutoAttack. The near agreement between exact-gradient and BPDA results supports the use of BPDA for the reported setting. The authors also report low gradient randomness for RDC relative to DiffPure and an average absolute gradient magnitude of 8.2×10−68.2\times10^{-6}, comparable in scale to the listed adversarial-training models; these checks were used to argue against gradient obfuscation and vanishing gradients.

  9. Knowl 9 — RDC transfers to CIFAR-100 and Restricted ImageNet

    empirical result

    Additional evaluations test whether RDC’s performance extends beyond CIFAR-10. For CIFAR-100, the authors randomly sampled 128 test images and evaluated an ℓ∞\ell_\infty threat with ϵ∞=8/255\epsilon_\infty=8/255. For Restricted ImageNet, they randomly sampled 256 test images from its nine superclasses and evaluated ℓ∞\ell_\infty perturbations with ϵ∞=4/255\epsilon_\infty=4/255. The reported clean and robust accuracies (%) are:

    Dataset Method Clean Acc. Robust Acc.
    CIFAR-100 Rebuffi et al. (2021) 63.56 34.64
    CIFAR-100 Wang et al. (2023b) 75.22 42.67
    CIFAR-100 DiffPure 39.06 7.81
    CIFAR-100 DC (ours) 79.69 39.06
    CIFAR-100 RDC (ours) 80.47 53.12
    Restricted ImageNet Engstrom et al. (2019) 87.11 53.12
    Restricted ImageNet Wong et al. (2020) 83.98 46.88
    Restricted ImageNet Salman et al. (2020) 86.72 56.64
    Restricted ImageNet Debenedetti et al. (2022) 80.08 38.67
    Restricted ImageNet DiffPure 81.25 29.30
    Restricted ImageNet RDC (ours) 87.50 58.40

    RDC exceeds the best listed CIFAR-100 robust accuracy by 10.45 percentage points and the best listed Restricted ImageNet result by 1.76 points. These results are reported on small randomly sampled test subsets and under different threat budgets across the two datasets.

  10. Knowl 10 — Compute and out-of-distribution performance remain limitations

    limitation

    The reported RDC configuration still requires N+TN+T network function evaluations per image: NN likelihood-maximization updates plus TT timestep evaluations for classification. The authors identify more efficient diffusion models as a possible way to reduce this cost and note that diffusion architectures designed specifically for classification might improve performance beyond the off-the-shelf models used. The paper also reports that unconditional ELBO and likelihood scores distinguish CIFAR-10 images from some CIFAR-10-C corruptions but have difficulty separating in-distribution images from corruptions such as fog and frost, limiting their usefulness as general out-of-distribution detectors.

Coverage note — The paper’s auxiliary energy-based-model interpretation and detailed multi-head training diagnostics were omitted because they add less to reconstruction of RDC’s main method, theoretical result, and evaluated robustness than the included knowls.

References

  1. 1.Athalye, A., Carlini, N., and Wagner, D. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International Conference on Machine Learning, pp. 274–283, 2018.
  2. 2.Blau, T., Ganz, R., Kawar, B., Bronstein, A., and Elad, M. Threat model-agnostic adversarial defense using diffusion models. arXiv preprint arXiv:2207.08089, 2022.
  3. 3.Blau, T., Ganz, R., Baskin, C., Elad, M., and Bronstein, A. Classifier robustness enhancement via test-time transformation. arXiv preprint arXiv:2303.15409, 2023.
  4. 4.Cao, Y., Wang, N., Xiao, C., Yang, D., Fang, J., Yang, R., Chen, Q. A., Liu, M., and Li, B. Invisible for both camera and lidar: Security of multi-sensor fusion based perception in autonomous driving under physical-world attacks. In 2021 IEEE Symposium on Security and Privacy, pp. 176–194, 2021.
  5. 5.Carlini, N. and Wagner, D. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy, pp. 39–57, 2017.
  6. 6.Carlini, N., Tramer, F., Dvijotham, K. D., Rice, L., Sun, M., and Kolter, J. Z. (certified!!) adversarial robustness for free! In International Conference on Learning Representations, 2023.
  7. 7.Chen, H., Zhang, Y., Dong, Y., Yang, X., Su, H., and Zhu, J. Rethinking model ensemble in transfer-based adversarial attacks. In The Twelfth International Conference on Learning Representations, 2023.
  8. 8.Chen, H., Dong, Y., Shao, S., Hao, Z., Yang, X., Su, H., and Zhu, J. Your diffusion model is secretly a certifiably robust classifier. arXiv preprint arXiv:2402.02316, 2024.
  9. 9.Clark, K. and Jaini, P. Text-to-image diffusion models are zero-shot classifiers. arXiv preprint arXiv:2303.15233, 2023.
  10. 10.Cohen, J., Rosenfeld, E., and Kolter, Z. Certified adversarial robustness via randomized smoothing. In International Conference on Machine Learning, pp. 1310–1320, 2019.
  11. 11.Croce, F. and Hein, M. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International Conference on Machine Learning, pp. 2206–2216, 2020.
  12. 12.Croce, F., Andriushchenko, M., Sehwag, V., Debenedetti, E., Flammarion, N., Chiang, M., Mittal, P., and Hein, M. Robustbench: a standardized adversarial robustness benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2020.
  13. 13.Debenedetti, E., Sehwag, V., and Mittal, P. A light recipe to train robust vision transformers. arXiv preprint arXiv:2209.07399, 2022.
  14. 14.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, pp. 8780–8794, 2021.
  15. 15.Dong, M., Chen, X., Wang, Y., and Xu, C. Random normalization aggregation for adversarial defense. Advances in Neural Information Processing Systems, pp. 33676–33688, 2022.
  16. 16.Dong, Y., Liao, F., Pang, T., Su, H., Zhu, J., Hu, X., and Li, J. Boosting adversarial attacks with momentum. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9185–9193, 2018.
  17. 17.Dong, Y., Su, H., Wu, B., Li, Z., Liu, W., Zhang, T., and Zhu, J. Efficient decision-based black-box adversarial attacks on face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7714–7722, 2019.
  18. 18.Dong, Y., Chen, H., Chen, J., Fang, Z., Yang, X., Zhang, Y., Tian, Y., Su, H., and Zhu, J. How robust is google’s bard to adversarial image attacks? arXiv preprint arXiv:2309.11751, 2023.
  19. 19.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020.
  20. 20.Du, Y. and Mordatch, I. Implicit generation and modeling with energy based models. Advances in Neural Information Processing Systems, 2019.
  21. 21.Engstrom, L., Ilyas, A., Salman, H., Santurkar, S., and Tsipras, D. Robustness (python library), 2019. URL https://github.com/MadryLab/robustness.
  22. 22.Finlayson, S. G., Bowers, J. D., Ito, J., Zittrain, J. L., Beam, A. L., and Kohane, I. S. Adversarial attacks on medical machine learning. Science, pp. 1287–1289, 2019.
  23. 23.Fu, Y., Yu, Q., Li, M., Chandra, V., and Lin, Y. Double-win quant: Aggressively winning robustness of quantized deep neural networks via random precision training and inference. In International Conference on Machine Learning, pp. 3492–3504, 2021.
  24. 24.Goodfellow, I. J., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. In International Conference on Learning Representations, 2015.
  25. 25.Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K. Your classifier is secretly an energy based model and you should treat it like one. In International Conference on Learning Representations, 2019.
  26. 26.Han, X., Zheng, H., and Zhou, M. Card: Classification and regression diffusion models. Advances in Neural Information Processing Systems, pp. 18100–18115, 2022.
  27. 27.Hao, Z., Ying, C., Dong, Y., Su, H., Song, J., and Zhu, J. Gsmooth: Certified robustness against semantic transformations via generalized randomized smoothing. In International Conference on Machine Learning, pp. 8465–8483, 2022.
  28. 28.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 770–778, 2016.
  29. 29.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, pp. 6840–6851, 2020.
  30. 30.Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., and Welling, M. Argmax flows and multinomial diffusion: Learning categorical distributions. Advances in Neural Information Processing Systems, pp. 12454–12465, 2021.
  31. 31.Jing, P., Tang, Q., Du, Y., Xue, L., Luo, X., Wang, T., Nie, S., and Wu, S. Too good to be safe: Tricking lane detection in autonomous driving with crafted perturbations. In Proceedings of USENIX Security Symposium, pp. 3237–3254, 2021.
  32. 32.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, pp. 26565–26577, 2022.
  33. 33.Kingma, D., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. Advances in Neural Information Processing Systems, pp. 21696–21707, 2021.
  34. 34.Krizhevsky, A. and Hinton, G. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  35. 35.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, 2012.
  36. 36.Laidlaw, C., Singla, S., and Feizi, S. Perceptual adversarial robustness: Defense against unseen threat models. In International Conference on Learning Representations, 2021.
  37. 37.LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., and Huang, F. A tutorial on energy-based learning. Predicting structured data, 1(0), 2006.
  38. 38.Li, A. C., Prabhudesai, M., Duggal, S., Brown, E., and Pathak, D. Your diffusion model is secretly a zero-shot classifier. arXiv preprint arXiv:2303.16203, 2023.
  39. 39.Li, Y., Bradshaw, J., and Sharma, Y. Are generative classifiers more robust to adversarial attacks? In International Conference on Machine Learning, pp. 3804–3814, 2019.
  40. 40.Liao, F., Liang, M., Dong, Y., Pang, T., Hu, X., and Zhu, J. Defense against adversarial attacks using high-level representation guided denoiser. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1778–1787, 2018.
  41. 41.Liu, X., Zhang, X., Ma, J., Peng, J., and Liu, Q. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. arXiv preprint arXiv:2309.06380, 2023.
  42. 42.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022, 2021.
  43. 43.Mackowiak, R., Ardizzone, L., Kothe, U., and Rother, C. Generative classifiers as a basis for trustworthy image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2971–2981, 2021.
  44. 44.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations, 2018.
  45. 45.Ng, A. Y. and Jordan, M. I. On discriminative vs. generative classifiers: a comparison of logistic regression and naive bayes. In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic, pp. 841–848, 2001.
  46. 46.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pp. 8162–8171, 2021.
  47. 47.Nie, W., Guo, B., Huang, Y., Xiao, C., Vahdat, A., and Anandkumar, A. Diffusion models for adversarial purification. In International Conference on Machine Learning, pp. 16805–16827, 2022.
  48. 48.Papamakarios, G., Pavlakou, T., and Murray, I. Masked autoregressive flow for density estimation. Advances in neural information processing systems, 30, 2017.
  49. 49.Pérez, J. C., Alfarra, M., Jeanneret, G., Rueda, L., Thabet, A., Ghanem, B., and Arbeláez, P. Enhancing adversarial robustness via test-time transformation ensembling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 81–91, 2021.
  50. 50.Poole, B., Jain, A., Barron, J. T., and Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representations, 2022.
  51. 51.Raghunathan, A., Steinhardt, J., and Liang, P. Certified defenses against adversarial examples. In International Conference on Learning Representations, 2018.
  52. 52.Raina, R., Shen, Y., Ng, A. Y., and McCallum, A. Classification with hybrid generative/discriminative models. In Proceedings of the 16th International Conference on Neural Information Processing Systems, pp. 545–552, 2003.
  53. 53.Rebuffi, S.-A., Gowal, S., Calian, D. A., Stimberg, F., Wiles, O., and Mann, T. Fixing data augmentation to improve adversarial robustness. arXiv preprint arXiv:2103.01946, 2021.
  54. 54.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  55. 55.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3): 211–252, 2015.
  56. 56.Sabour, S., Cao, Y., Faghri, F., and Fleet, D. J. Adversarial manipulation of deep representations. arXiv preprint arXiv:1511.05122, 2015.
  57. 57.Salman, H., Ilyas, A., Engstrom, L., Kapoor, A., and Madry, A. Do adversarially robust imagenet models transfer better? Advances in Neural Information Processing Systems, 33:3533–3545, 2020.
  58. 58.Samangouei, P., Kabkab, M., and Chellappa, R. Defense-gan: Protecting classifiers against adversarial attacks using generative models. In International Conference on Learning Representations, 2018.
  59. 59.Schott, L., Rauber, J., Bethge, M., and Brendel, W. Towards the first adversarially robust neural network model on mnist. In International Conference on Learning Representations, 2019.
  60. 60.Schwinn, L., Bungert, L., Nguyen, A., Raab, R., Pulsmeyer, F., Precup, D., Eskofier, B., and Zanca, D. Improving robustness against real-world and worst-case distribution shifts through decision region quantification. In International Conference on Machine Learning, pp. 19434–19449, 2022.
  61. 61.Shao, S., Dai, X., Yin, S., Li, L., Chen, H., and Hu, Y. Catch-up distillation: You only need to train once for accelerating sampling. arXiv preprint arXiv:2305.10769, 2023.
  62. 62.Sharif, M., Bhagavatula, S., Bauer, L., and Reiter, M. K. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In ACM Sigsac Conference on Computer and Communications Security, pp. 1528–1540, 2016.
  63. 63.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265, 2015.
  64. 64.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations, 2020.
  65. 65.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pp. 11918–11930, 2019.
  66. 66.Song, Y., Kim, T., Nowozin, S., Ermon, S., and Kushman, N. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. In International Conference on Learning Representations, 2018.
  67. 67.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  68. 68.Song, Y., Dhariwal, P., Chen, M., and Sutskever, I. Consistency models. arXiv preprint arXiv:2303.01469, 2023.
  69. 69.Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. Intriguing properties of neural networks. International Conference on Learning Representations, 2014.
  70. 70.Tramèr, F. and Boneh, D. Adversarial training and robustness for multiple perturbations. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pp. 5866–5876, 2019.
  71. 71.Tramer, F., Carlini, N., Brendel, W., and Madry, A. On adaptive attacks to adversarial example defenses. In Advances in Neural Information Processing Systems, pp. 1633–1645, 2020.
  72. 72.Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A. Robustness may be at odds with accuracy. In International Conference on Learning Representations, 2019.
  73. 73.Wang, J., Lyu, Z., Lin, D., Dai, B., and Fu, H. Guided diffusion model for adversarial purification. arXiv preprint arXiv:2205.14969, 2022.
  74. 74.Wang, Z., Lu, C., Wang, Y., Bao, F., Li, C., Su, H., and Zhu, J. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. arXiv preprint arXiv:2305.16213, 2023a.
  75. 75.Wang, Z., Pang, T., Du, C., Lin, M., Liu, W., and Yan, S. Better diffusion models further improve adversarial training. arXiv preprint arXiv:2302.04638, 2023b.
  76. 76.Wei, Z., Wang, Y., Guo, Y., and Wang, Y. Cfa: Class-wise calibrated fair adversarial training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8193–8201, 2023a.
  77. 77.Wei, Z., Wang, Y., and Wang, Y. Jailbreak and guard aligned language models with only few in-context demonstrations. arXiv preprint arXiv:2310.06387, 2023b.
  78. 78.Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., and Schmidhuber, J. Natural evolution strategies. The Journal of Machine Learning Research, pp. 949–980, 2014.
  79. 79.Wong, E. and Kolter, Z. Provable defenses against adversarial examples via the convex outer adversarial polytope. In International Conference on Machine Learning, pp. 5286–5295, 2018.
  80. 80.Wong, E., Rice, L., and Kolter, J. Z. Fast is better than free: Revisiting adversarial training. arXiv preprint arXiv:2001.03994, 2020.
  81. 81.Xiao, C., Zhu, J.-Y., Li, B., He, W., Liu, M., and Song, D. Spatially transformed adversarial examples. In International Conference on Learning Representations, 2018.
  82. 82.Xiao, C., Chen, Z., Jin, K., Wang, J., Nie, W., Liu, M., Anandkumar, A., Li, B., and Song, D. Densepure: Understanding diffusion models for adversarial robustness. In International Conference on Learning Representations, 2023.
  83. 83.Yang, X., Shih, S.-M., Fu, Y., Zhao, X., and Ji, S. Your vit is secretly a hybrid discriminative-generative diffusion model. arXiv preprint arXiv:2208.07791, 2022.
  84. 84.Zhang, H., Yu, Y., Jiao, J., Xing, E., El Ghaoui, L., and Jordan, M. Theoretically principled trade-off between robustness and accuracy. In International Conference on Machine Learning, pp. 7472–7482, 2019.
  85. 85.Zhang, J., Chen, Z., Zhang, H., Xiao, C., and Li, B. {DiffSmooth}: Certifiably robust learning via diffusion models and local smoothing. In 32nd USENIX Security Symposium, pp. 4787–4804, 2023.
  86. 86.Zhang, Y., He, H., Zhu, J., Chen, H., Wang, Y., and Wei, Z. On the duality between sharpness-aware minimization and adversarial training. In International Conference on Machine Learning, 2024.
  87. 87.Zimmermann, R. S., Schott, L., Song, Y., Dunn, B. A., and Klindt, D. A. Score-based generative classifiers. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/