WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

Changhoon KimKyle MinMaitreya PatelSheng ChengYezhou Yang

article2024CVPR62 citations

Proposes a distributor-oriented framework that embeds user-specific identifiers directly into the weights of text-to-image diffusion models via weight modulation, enabling reliable user attribution and resistance against post-processing manipulations without sacrificing image quality.

Listen

Rapid advances in artificial intelligence now enable the creation of highly realistic synthetic images from simple text descriptions. While these text-to-image models unlock immense creative potential, they also pose severe risks regarding misinformation, deepfakes, and political manipulation. Existing safeguards, such as standalone watermarking modules, are easily bypassed in open-source environments by simply disabling a line of code. The article addresses this accountability challenge by evaluating a distributor-focused framework called WOUAF (Weight Modulation for User Attribution and Fingerprinting), which embeds unique digital identifiers directly into model weights so that synthetic images can be traced back to the specific user who generated them.

To achieve this, the authors developed a fine-tuning technique applied specifically to the image decoder component of the popular Stable Diffusion model. When an authorized user requests a model from an open-source hub or distributor, the system applies a distinct mathematical modulation to the decoder weights and registers the corresponding identifier in a central database. When an image is generated, it inherently carries an invisible identifier that a dedicated decoding network can extract. The authors evaluated this approach against standard image datasets (MS-COCO and LAION-Aesthetics) across diverse image modifications and adversarial attempts to remove the fingerprint.

The findings show that this weight-modulation technique achieves near-perfect user attribution accuracy of 98% to 99% while causing negligible degradation to image visual quality or text alignment. Using a 32-bit identifier, the system can distinguish over four billion unique users without interference. Generating a customized, fingerprinted model requires less than one second—a major computational advantage over prior fine-tuning baselines that require minutes or hours per user. Furthermore, the embedded identifiers proved robust against common transformations such as cropping, rotation, blurring, and compression, outperforming alternative approaches by an average of 11% in post-processing resilience. When tested against deliberate removal attempts via image compression or model fine-tuning, the identifier could only be obscured by severely degrading the visual quality of the output.

These results demonstrate that model distributors can effectively enforce accountability without sacrificing model performance or incurring heavy computational expenses. By embedding user attribution directly into the model weights rather than in optional modular code, distributors can reliably trace malicious content and deter harmful misuse. Organizations distributing open-source generative models should consider integrating weight-modulation fingerprinting as a scalable compliance and safety measure. Decision-makers should note that while the current evaluation provides high confidence for image-generation frameworks, further validation is necessary before applying these techniques across other media modalities such as audio, text, and video.

arXiv: 2306.04744
Cover for WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

Abstract

The rapid advancement of generative models, facilitating the creation of hyper-realistic images from textual descriptions, has concurrently escalated critical societal concerns such as misinformation. Although providing some mitigation, traditional fingerprinting mechanisms fall short in attributing responsibility for the malicious use of synthetic images. This paper introduces a novel approach to model fingerprinting that assigns responsibility for the generated images, thereby serving as a potential countermeasure to model misuse. Our method modifies generative models based on each user's unique digital fingerprint, imprinting a unique identifier onto the resultant content that can be traced back to the user. This approach, incorporating fine-tuning into Text-to-Image (T2I) tasks using the Stable Diffusion Model, demonstrates near-perfect attribution accuracy with a minimal impact on output quality. Through extensive evaluation, we show that our method outperforms baseline methods with an average improvement of 11% in handling image post-processes. Our method presents a promising and novel avenue for accountable model distribution and responsible use. Our code is available in https://github.com/kylemin/WOUAF.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methods
  • 3.1. Preliminaries
  • 3.2. User-specific Weight Modulation
  • 3.3. Training Objectives
  • 4. Experiments
  • 4.1. Experiment Settings
  • 4.2. Fingerprint Capacity
  • 4.3. Attribution Accuracy and Image Quality
  • 4.4. Attribution Analysis for Diverse Image Sources
  • 4.5. Benefits of Finetuning only Decoder
  • 4.6. Robust User Attribution against Image Post-processes
  • 5. Deliberate Fingerprint Manipulations
  • 5.1. Resilience Against Deep Classifier
  • 5.2. Resilience Against Auto-Encoders
  • 5.3. Resilience Against Model Purification
  • 6. Conclusion
  • 7. Acknowledgment
  • References

Knowls

  1. Knowl 1 — Distributor-controlled fingerprinting through model weights

    model/method

    WOUAF is a distributor-oriented fingerprinting method for Stable Diffusion text-to-image models. A model distributor first maintains a generic Stable Diffusion model and a database of user identifiers. For each download request, the distributor samples a user-specific binary fingerprint, uses it to modify the decoder weights, registers the fingerprint in the database, and distributes the resulting model. Images generated by the distributed model can later be passed to a fingerprint decoder to recover the identifier and attribute the image to the corresponding user.

    The fingerprint is embedded directly into the model parameters rather than added by an independent image-watermarking module. WOUAF does not change Stable Diffusion's architecture, and producing a new user-specific model requires only a forward pass through the fingerprint-mapping network. This design is intended to prevent users from bypassing fingerprinting by deleting a separate embedding operation from the generation code.

  2. Knowl 2 — User-specific weight modulation of the Stable Diffusion decoder

    model/method

    Let ϕ∈{0,1}dϕ\phi\in\{0,1\}^{d_\phi} be a binary fingerprint sampled from Bernoulli⁡(0.5)dϕ\operatorname{Bernoulli}(0.5)^{d_\phi}. A mapping network M:Rdϕ→RdMM:\mathbb{R}^{d_\phi}\rightarrow\mathbb{R}^{d_M} converts ϕ\phi into an intermediate representation. For each modulated decoder layer ll, an affine map Al:RdM→RdjA_l:\mathbb{R}^{d_M}\rightarrow\mathbb{R}^{d_j} produces one scale value for each of the djd_j output channels. If W∈Ri×j×kW\in\mathbb{R}^{i\times j\times k} is a pretrained convolutional weight tensor, where ii, jj, and kk index input channels, output channels, and kernel positions, respectively, WOUAF forms the fingerprinted weights as

    Wi,j,kϕ=ujWi,j,k,uj=Al(M(ϕ))j.W^{\phi}_{i,j,k}=u_jW_{i,j,k},\qquad u_j=A_l(M(\phi))_j.

    The modulation is applied only to the Stable Diffusion image decoder DD, not to the latent diffusion denoising network ϵθ\epsilon_\theta. Given a latent code zz, the resulting user-specific decoder is denoted D(ϕ,z)D(\phi,z). The released user model therefore differs from the generic model through its decoder weights while retaining the same model architecture.

  3. Knowl 3 — Joint fingerprint recovery and image-quality training objective

    equation

    WOUAF trains a fingerprint-decoding network F:Rdx→RdϕF:\mathbb{R}^{d_x}\rightarrow\mathbb{R}^{d_\phi}, implemented as a ResNet-50, to recover the dϕd_\phi fingerprint bits from generated images. For an image x∼pdatax\sim p_{\mathrm{data}}, its Stable Diffusion latent representation is z=E(x)z=E(x), and the fingerprinted reconstruction is xϕ=D(ϕ,z)x_\phi=D(\phi,z). The decoder is trained with a bitwise binary cross-entropy loss between the true bits ϕi\phi_i and the sigmoid outputs σ(F(xϕ)i)\sigma(F(x_\phi)_i):

    Lϕ=−Ex∼pdata, ϕ∼Bernoulli⁡(0.5)dϕ[∑i=1dϕ(ϕilog⁡σ(F(D(ϕ,E(x)))i)+(1−ϕi)log⁡(1−σ(F(D(ϕ,E(x)))i)))].\mathcal{L}_{\phi}=-\mathbb{E}_{x\sim p_{\mathrm{data}},\,\phi\sim\operatorname{Bernoulli}(0.5)^{d_\phi}} \left[\sum_{i=1}^{d_\phi}\left(\phi_i\log\sigma(F(D(\phi,E(x)))_i)+(1-\phi_i)\log\left(1-\sigma(F(D(\phi,E(x)))_i)\right)\right)\right].

    To preserve image quality, WOUAF also minimizes a perceptual distance ℓ\ell between the original image and its fingerprinted reconstruction:

    Lquality=Ex∼pdata, ϕ∼Bernoulli⁡(0.5)dϕ[ℓ(x,D(ϕ,E(x)))].\mathcal{L}_{\mathrm{quality}}=\mathbb{E}_{x\sim p_{\mathrm{data}},\,\phi\sim\operatorname{Bernoulli}(0.5)^{d_\phi}}\left[\ell\bigl(x,D(\phi,E(x))\bigr)\right].

    The trainable parameters are the mapping network MM, the affine modulation maps AlA_l, the decoder DD, and the fingerprint decoder FF. They are optimized jointly using

    min⁡A,M,D,F  λ1Lϕ+λ2Lquality,λ1=λ2=1.0.\min_{A,M,D,F}\;\lambda_1\mathcal{L}_{\phi}+\lambda_2\mathcal{L}_{\mathrm{quality}}, \qquad \lambda_1=\lambda_2=1.0.

    After training, the distributor samples a new fingerprint once per user, modulates the decoder, and gives the user the resulting decoder that accepts only latent inputs.

  4. Knowl 4 — Differentiable robust training against image post-processing

    algorithm

    WOUAF can be trained to decode fingerprints after images have undergone a differentiable transformation T:Rdx→RdxT:\mathbb{R}^{d_x}\rightarrow\mathbb{R}^{d_x}. Instead of decoding directly from D(ϕ,E(x))D(\phi,E(x)), the fingerprint decoder receives T(D(ϕ,E(x)))T(D(\phi,E(x))). The robust fingerprint loss is

    Lrobust=−Ex∼pdata, ϕ∼Bernoulli⁡(0.5)dϕ[∑i=1dϕ(ϕilog⁡σ(F(T(D(ϕ,E(x)))i)+(1−ϕi)log⁡(1−σ(F(T(D(ϕ,E(x)))i)))].\mathcal{L}_{\mathrm{robust}}=-\mathbb{E}_{x\sim p_{\mathrm{data}},\,\phi\sim\operatorname{Bernoulli}(0.5)^{d_\phi}} \left[\sum_{i=1}^{d_\phi}\left(\phi_i\log\sigma(F(T(D(\phi,E(x)))_i)+(1-\phi_i)\log\left(1-\sigma(F(T(D(\phi,E(x)))_i)\right)\right)\right].

    The robust model is optimized with the same quality term and weights as the ordinary model, replacing Lϕ\mathcal{L}_{\phi} with Lrobust\mathcal{L}_{\mathrm{robust}}. Training uses random intensities of eight transformations: random erasing with erase ratios in {5%,10%,15%,20%}\{5\%,10\%,15\%,20\%\}; rotation with an angle sampled from (−30,30)(-30,30) degrees; Gaussian blur with kernel size in {3,5,7}\{3,5,7\}; random cropping with crop-out ratios in {5%,10%,15%,20%}\{5\%,10\%,15\%,20\%\}; brightness jitter with a factor in (−0.3,0.3)(-0.3,0.3); Gaussian noise with standard deviation sampled from U[0,0.2]U[0,0.2]; JPEG compression with quality in {90,80,70,60,50}\{90,80,70,60,50\}; and a combined transformation that independently selects a subset of the seven individual operations with probability 0.50.5.

  5. Knowl 5 — Evaluation protocol and attribution metric

    experimental setup

    The method is fine-tuned on the MS-COCO training data using the Karpathy split and evaluated on the held-out MS-COCO test set and randomly sampled LAION-Aesthetics images. Unless otherwise stated, generation uses the Euler scheduler with T=20T=20 denoising steps and classifier-free guidance scale 7.57.5. The mapping network has two fully connected layers, and the default fingerprint length is dϕ=32d_\phi=32 bits. WOUAF is evaluated as WOUAF-conv, which modulates only decoder convolutional layers, and WOUAF-all, which modulates all decoder layers, including self-attention and convolutional layers. The baselines are DAG and Stable Signature, with Stable Signature retrained under the same experimental conditions.

    For a true fingerprint ϕ∈{0,1}dϕ\phi\in\{0,1\}^{d_\phi} and an image xϕx_\phi, the estimated fingerprint is ϕ^i=1[σ(F(xϕ)i)>0.5]\hat{\phi}_i=\mathbf{1}[\sigma(F(x_\phi)_i)>0.5]. Attribution accuracy is the fraction of correctly recovered bits:

    Acc⁡(ϕ,ϕ^)=1dϕ∑i=1dϕ1(ϕi=ϕ^i).\operatorname{Acc}(\phi,\hat{\phi})=\frac{1}{d_\phi}\sum_{i=1}^{d_\phi}\mathbf{1}(\phi_i=\hat{\phi}_i).

    Image quality is measured with Fréchet Inception Distance (FID), where lower is better, and CLIP score, where higher is better. Fingerprinting time measures the time required to produce a user-specific model.

  6. Knowl 6 — Near-perfect attribution with negligible quality degradation

    data/table

    On held-out MS-COCO and LAION-Aesthetics data, both WOUAF variants recover fingerprints substantially better than DAG and comparably to or better than Stable Signature, while retaining image-quality scores close to the original Stable Diffusion model. The benchmark values are:

    Model Fingerprinting time MS-COCO LAION
    Attribution Acc. CLIP-score FID Attribution Acc. CLIP-score FID
    Original SD - - 0.73 24.48 - 0.50 19.67
    DAG 8.4 hr 0.70 0.73 26.54 0.71 0.49 23.13
    Stable Signature <1 min 0.99 0.73 24.55 0.98 0.50 20.02
    WOUAF-conv <1 sec 0.99 0.73 24.43 0.98 0.51 20.46
    WOUAF-all <1 sec 0.99 0.73 24.42 0.99 0.51 19.91

    Thus, with one training process and less than one second of per-user fingerprinting time, WOUAF reaches 0.990.99 attribution accuracy on MS-COCO for both variants. WOUAF-all also achieves the best reported FID on MS-COCO and LAION among the fingerprinted models, while qualitative comparisons show only small pixel-wise differences from the original Stable Diffusion outputs.

  7. Knowl 7 — Fingerprint capacity and the length–decodability trade-off

    empirical result

    The number of distinguishable user identifiers is controlled primarily by the fingerprint length dϕd_\phi. The measured attribution accuracies for different lengths are:

    Fingerprint dimension dϕd_\phi 16 32 64 128
    WOUAF-conv 0.99 0.99 0.98 0.94
    WOUAF-all 0.99 0.99 0.99 0.97

    Attribution accuracy decreases as more bits must be decoded, although both variants remain at 0.990.99 for dimensions 1616, 3232, and 6464. At dϕ=128d_\phi=128, WOUAF-all is more accurate than WOUAF-conv. The selected default, dϕ=32d_\phi=32, provides approximately 2322^{32}, or more than four billion, possible fingerprints while retaining near-perfect attribution accuracy.

  8. Knowl 8 — No false attribution for images from unrelated sources

    empirical result

    WOUAF was tested for whether non-fingerprinted or externally generated images could accidentally match a registered user. The evaluation contained 1,000 images from each of five sources: authentic MS-COCO test images, non-fingerprinted Stable Diffusion images, ProGAN images, StyleGAN images, and StyleGAN2 images. These images were compared against a database containing one million registered fingerprints.

    For the default 32-bit fingerprints, an image was considered a database match only when at least 32×0.95≈3032\times0.95\approx30 bits agreed with a stored fingerprint. No image from any of the five source categories was incorrectly attributed to a fingerprint in the one-million-entry database. This result indicates that the attribution procedure can reject images whose origin is external to the distributed WOUAF models under the tested threshold.

  9. Knowl 9 — Modulating only the decoder preserves the attribution–quality balance

    empirical result

    The authors compared the proposed decoder-only modulation with a variant that modulates both the diffusion denoising network ϵθ\epsilon_\theta and the image decoder DD. The latter resembles weight-modulation strategies used for GANs but performs poorly when combined with Stable Diffusion fine-tuning: its highest observed attribution accuracy was 89%89\%, accompanied by a CLIP score of 0.680.68 and an FID of 63.4863.48.

    The decoder-only design avoids this severe quality degradation while achieving approximately 0.990.99 attribution accuracy and FID values near the original Stable Diffusion model. The experiments therefore support restricting fingerprint-dependent weight changes to DD rather than applying them to the latent diffusion network as well.

  10. Knowl 10 — Robust attribution under common image transformations

    empirical result

    When trained with randomized differentiable post-processing, WOUAF remains more accurate than robust versions of DAG and Stable Signature across random erasing, rotation, Gaussian blurring, random cropping, brightness jittering, Gaussian noise, JPEG compression, and combined attacks. Averaged across the tested post-processing conditions, WOUAF improves attribution accuracy by 11%11\% over Stable Signature and by 29%29\% over DAG.

    For every method, stronger transformations generally reduce attribution accuracy. The decline is especially visible for rotation, blur, noise, JPEG compression, and the combined attack, whereas WOUAF remains comparatively stable over the tested intensity ranges. These results show that incorporating the transformations into training improves resilience to post-processing rather than merely relying on the unmodified fingerprint.

  11. Knowl 11 — Resistance to fingerprint detection and deliberate removal

    limitation

    The paper evaluates three deliberate attacks. First, a ResNet-50 binary classifier trained on 10,000 images—5,000 original Stable Diffusion images and 5,000 fingerprinted images—achieves 98%98\% training accuracy, but only 0.660.66 accuracy on a separate 5,000-image WOUAF-conv evaluation set and 0.560.56 on WOUAF-all, close to random guessing. This suggests that the embedded fingerprint is difficult to detect directly, especially for WOUAF-all.

    Second, when an attacker passes outputs through learned image-compression auto-encoders, attribution accuracy decreases toward approximately 50%50\% as compression increases. The decrease is accompanied by worsening image quality, so removing the fingerprint requires sacrificing visual fidelity in the tested setting.

    Third, an attacker can fine-tune the downloaded fingerprinted decoder on a proprietary image set to perform model purification. Increasing purification reduces attribution accuracy, but it also worsens image quality, as reflected by increasing FID. The paper therefore finds an empirical trade-off rather than proving that fingerprint removal is impossible: attacks can weaken attribution, but the tested attacks incur substantial quality degradation.

Coverage note — No substantial contribution from the main paper was omitted; routine implementation details and appendix-only extensions were excluded because they are not load-bearing for reconstructing the method or results.

References

  1. 1.Johannes Ballé, David Minnen, Saurabh Singh, Sung Jin Hwang, and Nick Johnston. Variational image compression with a scale hyperprior. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018.
  2. 2.Ali Breland. The bizarre and terrifying case of the “deepfake” video that helped bring an african nation to the brink. motherjones, 2019.
  3. 3.Zhengxue Cheng, Heming Sun, Masaru Takeuchi, and Jiro Katto. Learned image compression with discretized gaussian mixture likelihoods and attention modules. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020.
  4. 4.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021.
  5. 5.Pierre Fernandez, Guillaume Couairon, Herve Jégou, Matthijs Douze, and Teddy Furon. The stable signature: Rooting watermarks in latent diffusion models. arXiv preprint arXiv:2303.15435, 2023.
  6. 6.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  7. 7.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718, 2021.
  8. 8.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, pages 6626–6637, 2017.
  9. 9.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  10. 10.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017.
  11. 11.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019.
  12. 12.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. arXiv preprint arXiv:1912.04958, 2019.
  13. 13.Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative adversarial networks with limited data. Advances in neural information processing systems, 33:12104–12114, 2020.
  14. 14.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020.
  15. 15.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. arXiv preprint arXiv:2206.00364, 2022.
  16. 16.Changhoon Kim, Yi Ren, and Yezhou Yang. Decentralized attribution of generative models. In International Conference on Learning Representations, 2021.
  17. 17.Arijeta Lajka. New ai voice-cloning tools ‘add fuel’ to misinformation fire. AP News, 2023.
  18. 18.Yue Li, Hongxia Wang, and Mauro Barni. A survey of deep neural network watermarking techniques. ArXiv, abs/2103.09274, 2021.
  19. 19.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014.
  20. 20.David Minnen, Johannes Balle, and George Toderici. Joint autoregressive and hierarchical priors for learned image compression. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, 3-8 December 2018, Montreal, Canada, pages 10794–10803, 2018.
  21. 21.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, pages 16784–16804. PMLR, 2022.
  22. 22.Guangyu Nie, Changhoon Kim, Yezhou Yang, and Yi Ren. Attributing image generative models using latent fingerprints. arXiv preprint arXiv:2304.09752, 2023.
  23. 23.Matt Novak. Ai image creator midjourney halts free trials but it has nothing to do with the pope’s jacket. forbes, 2023.
  24. 24.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  25. 25.E. Riba, D. Mishkin, D. Ponsa, E. Rublee, and G. Bradski. Kornia: an open source differentiable computer vision library for pytorch. In Winter Conference on Applications of Computer Vision, 2020.
  26. 26.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022.
  27. 27.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015.
  28. 28.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022.
  29. 29.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022.
  30. 30.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  31. 31.Yuxin Wen, John Kirchenbauer, Jonas Geiping, and Tom Goldstein. Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust. arXiv preprint arXiv:2305.20030, 2023.
  32. 32.Ning Yu, Vladislav Skripniuk, Dingfan Chen, Larry Davis, and Mario Fritz. Responsible disclosure of generative models using scalable fingerprinting. arXiv preprint arXiv:2012.08726, 2020.
  33. 33.Ning Yu, Vladislav Skripniuk, Sahar Abdelnabi, and Mario Fritz. Artificial fingerprinting for generative models: Rooting deepfake attribution in training data. In Proceedings of the IEEE/CVF International conference on computer vision, pages 14448–14457, 2021.
  34. 34.Kevin Alex Zhang, Lei Xu, Alfredo Cuesta-Infante, and Kalyan Veeramachaneni. Robust invisible video watermarking with attention. 2019.
  35. 35.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018.
  36. 36.Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Ngai-Man Cheung, and Min Lin. A recipe for watermarking diffusion models. ArXiv, abs/2303.10137, 2023.
  37. 37.Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. Hidden: Hiding data with deep networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 657–672, 2018.

Citation

MLA
Kim, C., et al. “WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2306.04744v3.
APA
Kim, C., Min, K., Patel, M., Cheng, S., & Yang, Y. (2023). WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models. arXiv. http://arxiv.org/abs/2306.04744v3
Chicago
Kim, C., K. Min, M. Patel, S. Cheng, and Y. Yang. 2023. “WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models”. arXiv. http://arxiv.org/abs/2306.04744v3.
Harvard
Kim, C. et al. (2023) “WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.04744v3.
Vancouver
1. Kim C, Min K, Patel M, Cheng S, Yang Y (2023) WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models. arXiv

BibTeX

@article{kim2023wouaf,
  title = {WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models},
  author = {Kim, Changhoon and Min, Kyle and Patel, Maitreya and Cheng, Sheng and Yang, Yezhou},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.04744v3},
  eprint = {2306.04744}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE