CLIP2Protect: Protecting Facial Privacy Using Text-Guided Makeup via Adversarial Latent Search

Fahad ShamshadMuzammal NaseerKarthik Nandakumar

article2023CVPR60 citations

Proposes a text-guided generative framework that applies natural-looking adversarial makeup to face images, effectively deceiving commercial black-box face recognition systems while preserving human-perceived identity.

Listen

Widespread deployment of automated face recognition systems by commercial and governmental entities poses growing privacy and mass surveillance risks, particularly as images are routinely scraped from public social media platforms. Existing privacy-preserving methods typically add noticeable pixel noise, degrade visual appeal, or alter human-perceived identity, which severely undermines the user experience. The article introduces and evaluates a privacy-protection framework called CLIP2Protect, designed to generate naturalistic, artifact-free facial modifications guided by plain-text makeup descriptions to reliably deceive unknown, proprietary face recognition systems.

The framework operates in two main stages using a pretrained generative image model (StyleGAN) and a vision-language alignment model (CLIP). First, the input face is inverted into the model's internal latent representation, accompanied by fine-tuning of the generator to guarantee near-perfect initial visual reconstruction. Second, the system explores the generative model's latent space to create targeted adversarial modifications guided by user-selected text prompts (such as specific makeup styles) alongside an identity-preserving regularizer that restricts alterations to fine-grained makeup attributes while locking foundational facial geometry. The approach was evaluated against multiple leading black-box face recognition architectures across standard benchmark datasets (CelebA-HQ, LADN, and LFW) as well as commercial face-matching application programming interfaces.

The experimental findings show substantial improvements over existing privacy protection methods. In black-box face verification tests, the framework achieved an average protection success rate of 64.90%, yielding an absolute gain of roughly 12% over the leading unrestricted makeup-transfer technique and about 14% over leading noise-based masking methods. In face identification evaluations, it consistently outperformed baselines across both top-1 and top-5 retrieval settings for impersonation and dodging tasks. Real-world validation against commercial engines, including the Face++ platform, demonstrated superior matching deception scores (averaging between 52.8 and 73.8) compared to prior baselines. Furthermore, the approach exhibited strong image quality with lower visual distortion scores and proved robust across diverse makeup prompts.

These results demonstrate that online facial privacy can be effectively secured without sacrificing image aesthetics or relying on complex, image-paired training datasets. Organizations and platform providers can deploy semantic, text-driven latent adjustments to protect users against non-consensual biometric tracking while maintaining a visually appealing social media experience. As a next step, developing automated mechanisms to recommend optimal text prompts and target identities tailored to individual user images will further streamline usability.

The primary limitation of the method is its computational overhead during the generation phase, which requires iterative generator fine-tuning and latent optimization per image. Nevertheless, the reported evidence provides high confidence in the technique's capability to generalize across diverse black-box models, though stakeholders should account for processing latency in large-scale, real-time deployment environments.

Cover for CLIP2Protect: Protecting Facial Privacy Using Text-Guided Makeup via Adversarial Latent Search

Abstract

The success of deep learning based face recognition systems has given rise to serious privacy concerns due to their ability to enable unauthorized tracking of users in the digital world. Existing methods for enhancing privacy fail to generate “naturalistic” images that can protect facial privacy without compromising user experience. We propose a novel two-step approach for facial privacy protection that relies on finding adversarial latent codes in the low-dimensional manifold of a pretrained generative model. The first step inverts the given face image into the latent space and finetunes the generative model to achieve an accurate reconstruction of the given image from its latent code. This step produces a good initialization, aiding the generation of high-quality faces that resemble the given identity. Subsequently, user-defined makeup text prompts and identity-preserving regularization are used to guide the search for adversarial codes in the latent space. Extensive experiments demonstrate that faces generated by our approach have stronger black-box transferability with an absolute gain of 12.06% over the state-of-the-art facial privacy protection approach under the face verification task. Finally, we demonstrate the effectiveness of the proposed approach for commercial face recognition systems. Our code is available at https://github.com/fahadshamshad/Clip2Protect.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Proposed Approach for Facial Privacy
  • 3.1. Preliminaries
  • 3.2. Makeup Text-Guided Adversarial Faces
  • 3.2.1 Latent Code Initialization
  • 3.2.2 Text-guided adversarial optimization
  • 4. Experiments
  • 4.1. Experimental Results
  • 4.2. Effectiveness in Real-World Applications
  • 4.3. Ablation Studies
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Privacy protection as a naturalistic adversarial-face objective

    model/method

    Given an original face image xx, the method seeks a protected image xpx^p that remains natural and visually related to xx while deceiving an unknown face-recognition system. Let D(a,b)D(a,b) measure dissimilarity between two face images in the recognition system's feature space, let H(xp,x)H(x^p,x) measure unnaturalness, let xtx^t be a target face with a different true identity from xx, and let u>0 u>0 bound the allowed unnaturalness. The targeted impersonation objective is

    min⁡xp  D(xp,xt)−D(xp,x)subject toH(xp,x)≤ν.\min_{x^p}\; D(x^p,x^t)-D(x^p,x) \quad\text{subject to}\quad H(x^p,x)\leq \nu.

    Minimizing this objective makes the protected image close to the target identity and far from the original identity. For dodging, the desired behavior is instead a large recognition distance between xpx^p and another image of the same person. Because the malicious recognition model is black-box, the proposed method optimizes surrogate recognition models and relies on transferability to the unknown system.

  2. Knowl 2 — Two-stage adversarial latent search with text-guided makeup

    model/method

    CLIP2Protect searches for a protected face on the natural-image manifold of a pretrained StyleGAN rather than directly adding a bounded pixel perturbation. The method has two stages. First, an encoder maps the input face xx to an initial latent code, and the StyleGAN generator is fine-tuned so that this code reconstructs xx accurately. Second, only selected components of the latent code are optimized to generate a face that simultaneously satisfies three goals: it fools an ensemble of white-box surrogate face-recognition models, follows a user-provided makeup description through CLIP, and remains close to the identity-preserving part of the initial latent code. The resulting generated image is the protected face. The pipeline diagram on page 4 shows this sequence: inversion and generator tuning, followed by latent optimization using adversarial, textual, and latent-preservation losses.

  3. Knowl 3 — Generator fine-tuning for identity-preserving inversion

    model/method

    Let x∈Rnx\in\mathbb{R}^n be an input face, IϕI_\phi be a pretrained e4e encoder, winv=Iϕ(x)w_{\mathrm{inv}}=I_\phi(x) its latent code, and GθG_\theta a pretrained StyleGAN generator. Instead of optimizing the encoder, CLIP2Protect keeps winvw_{\mathrm{inv}} fixed and fine-tunes the generator weights using a perceptual reconstruction loss and a pixelwise loss:

    θ∗=arg⁡min⁡θ  LLPIPS(x,Gθ(winv))+λ2L2(x,Gθ(winv)),\theta^* = \arg\min_{\theta}\; \mathcal{L}_{\mathrm{LPIPS}}\bigl(x,G_\theta(w_{\mathrm{inv}})\bigr)+\lambda_2\mathcal{L}_2\bigl(x,G_\theta(w_{\mathrm{inv}})\bigr),

    where LLPIPS\mathcal{L}_{\mathrm{LPIPS}} is the LPIPS perceptual distance, L2\mathcal{L}_2 is pixelwise squared-error similarity, and λ2\lambda_2 is a scalar weight. The reconstruction used for subsequent optimization is xinv∗=Gθ∗(winv)x^*_{\mathrm{inv}}=G_{\theta^*}(w_{\mathrm{inv}}). In the paper's LFW example, generator fine-tuning increased the Face++ matching score from 62.3862.38 for ordinary encoder inversion to 98.9698.96, producing a near-perfect reconstruction and a better starting point for natural adversarial editing.

  4. Knowl 4 — Directional CLIP loss for localized text-guided makeup

    equation

    Let tmakeupt_{\mathrm{makeup}} be a user-selected makeup prompt, let tsrct_{\mathrm{src}} be the source-image text description (set to face in the experiments), and let ETE_T and EIE_I be the text and image encoders of a pretrained CLIP model. For a candidate latent code ww, the generated image is Gθ∗(w)G_{\theta^*}(w). CLIP2Protect forms the text and image changes

    ΔT=ET(tmakeup)−ET(tsrc),ΔI=EI(Gθ∗(w))−EI(x),\Delta T=E_T(t_{\mathrm{makeup}})-E_T(t_{\mathrm{src}}), \qquad \Delta I=E_I\bigl(G_{\theta^*}(w)\bigr)-E_I(x),

    and minimizes the directional loss

    Lclip=1−ΔI⋅ΔT∥ΔI∥ ∥ΔT∥.\mathcal{L}_{\mathrm{clip}}=1-\frac{\Delta I\cdot\Delta T}{\lVert\Delta I\rVert\,\lVert\Delta T\rVert}.

    Here ⋅\cdot is the Euclidean inner product and the norms are Euclidean norms. Aligning the change from the original image with the change specified by the makeup text encourages modifications such as lipstick or eyeshadow without forcing the entire face to become a generic image matching the prompt. The paper argues that this directional formulation improves style diversity and hides adversarial changes inside a semantically meaningful makeup effect.

  5. Knowl 5 — Transferable adversarial loss against black-box recognition

    equation

    Let f(⋅)f(\cdot) be a normalized face-recognition feature extractor and define the cosine distance between images aa and bb as

    D(a,b)=1−f(a)Tf(b).D(a,b)=1-f(a)^{\mathsf T}f(b).

    For a candidate latent code ww, a target face xtx^t, the original face xx, and the fine-tuned generator Gθ∗G_{\theta^*}, CLIP2Protect uses

    Ladv(w)=D(Gθ∗(w),xt)−D(Gθ∗(w),x).\mathcal{L}_{\mathrm{adv}}(w)=D\bigl(G_{\theta^*}(w),x^t\bigr)-D\bigl(G_{\theta^*}(w),x\bigr).

    Minimizing this loss makes the generated face closer to the target and farther from the original in surrogate recognition-feature space. Since the malicious face-recognition model and its distance function are unavailable, the loss is optimized using an ensemble of three white-box surrogate recognition models while the fourth model is held out as the black-box target. The ensemble is intended to approximate a transferable decision boundary rather than directly query or access the protected recognition system.

  6. Knowl 6 — Identity-preserving latent regularization and combined objective

    equation

    StyleGAN's early latent layers control high-level properties such as pose, hairstyle, and face shape, whereas deeper layers more strongly control fine appearance attributes such as makeup. Let w∈R18×512w\in\mathbb{R}^{18\times512} be the candidate StyleGAN latent code and winvw_{\mathrm{inv}} the inversion code. CLIP2Protect uses an identity mask mid∈{0,1}18×512m_{\mathrm{id}}\in\{0,1\}^{18\times512} whose entries are 00 for StyleGAN layers 1--7 and 11 for layers 8--18. The latent-preservation term is

    Llatent(w)=∥(w⊙mid)−(winv⊙mid)∥2,\mathcal{L}_{\mathrm{latent}}(w)=\left\|(w\odot m_{\mathrm{id}})-(w_{\mathrm{inv}}\odot m_{\mathrm{id}})\right\|_2,

    where ⊙\odot denotes elementwise multiplication. During optimization, the identity-related early layers are not changed, while the deeper layers can encode the requested makeup and adversarial changes but are discouraged from drifting far from the inversion. The complete optimization objective is

    Ltotal(w)=λadvLadv(w)+λclipLclip(w)+λlatentLlatent(w),\mathcal{L}_{\mathrm{total}}(w)=\lambda_{\mathrm{adv}}\mathcal{L}_{\mathrm{adv}}(w)+\lambda_{\mathrm{clip}}\mathcal{L}_{\mathrm{clip}}(w)+\lambda_{\mathrm{latent}}\mathcal{L}_{\mathrm{latent}}(w),

    where the three λ\lambda values are positive loss weights. The latent-manifold restriction and the text and latent regularizers serve as the method's implicit naturalness constraint.

  7. Knowl 7 — Evaluation protocol and implementation configuration

    experimental setup

    The experiments use StyleGAN2 pretrained on FFHQ and a vision-transformer CLIP model. Generator inversion uses 450450 fine-tuning iterations with λ2=0.5\lambda_2=0.5. Protected-face optimization uses Adam with learning rate 0.010.01, β1=0.9\beta_1=0.9, β2=0.999\beta_2=0.999, and 5050 iterations; the loss weights are λadv=1\lambda_{\mathrm{adv}}=1, λclip=0.5\lambda_{\mathrm{clip}}=0.5, and λlatent=0.01\lambda_{\mathrm{latent}}=0.01. The method uses 4040 collected makeup prompts, while reported aggregate results are averaged over five styles.

    Face-verification impersonation experiments use 1,0001{,}000 CelebA-HQ images and all 332332 LADN images divided among four target identities. Verification dodging uses 500500 randomly selected subjects from CelebA-HQ and LFW, with two images per subject. Identification experiments also use 500500 subjects from CelebA-HQ and LFW, assigning one image to a gallery and the other to a probe set; four target identities are added to the gallery for impersonation.

    The black-box target models are IRSE50, IR152, FaceNet, and MobileFace. For each target, the other three models generate the adversarial images. Protection success rate (PSR) is the fraction of protected faces misclassified by the recognition system. Identification additionally uses Rank-NN untargeted success, where the top-NN list contains no image of the original identity, and Rank-NN targeted success, where the target identity appears in the top-NN list. PSNR and SSIM measure similarity to the original image, FID measures generated-image realism, and commercial APIs are evaluated using their returned confidence scores.

  8. Knowl 8 — Black-box verification impersonation performance

    data/table

    The black-box verification experiment measures PSR for targeted impersonation: a higher percentage means more protected faces fail to be recognized as the intended original identity while pursuing the target. Each target-model column is evaluated with the other three recognition models used as surrogates. The proposed method has the highest average PSR on both datasets, reaching 64.90%64.90\% and exceeding AMT-GAN's 52.84%52.84\% by 12.0612.06 percentage points.

    Could not parse LaTeX table

    On CelebA-HQ the proposed method obtains PSR values of 81.10%81.10\%, 48.42%48.42\%, 41.72%41.72\%, and 75.26%75.26\% against IRSE50, IR152, FaceNet, and MobileFace, respectively. On LADN the corresponding values are 91.57%91.57\%, 53.31%53.31\%, 47.91%47.91\%, and 79.94%79.94\%.

  9. Knowl 9 — Black-box identification dodging and impersonation performance

    data/table

    The identification experiment evaluates both untargeted dodging and targeted impersonation on LFW. For untargeted attacks, higher Rank-11 or Rank-55 success means that the original identity is absent from the corresponding top-NN gallery candidates. For targeted attacks, higher Rank-NN success means that the chosen target identity appears in the top-NN candidates. The proposed method gives the strongest average untargeted performance at both ranks and the strongest average targeted performance at Rank-11 and Rank-55 among the compared methods.

    Could not parse LaTeX table

    The method's average untargeted results are 82.2%82.2\% at Rank-11 and 58.5%58.5\% at Rank-55, compared with 77.8%77.8\% and 54.9%54.9\% for TIP-IM. Its average targeted results are 23.4%23.4\% and 51.05%51.05\%, compared with 19.7%19.7\% and 41.9%41.9\% for TIP-IM. The paper highlights the untargeted identification result as evidence that the generative approach can support dodging without requiring a specific impersonation target.

  10. Knowl 10 — Naturalness and effectiveness against a commercial API

    data/table

    The method improves the privacy-naturalness tradeoff relative to prior approaches. On CelebA-HQ and LADN, FID is lower for the proposed method than for TIP-IM and AMT-GAN, while its PSR gain relative to Adv-Makeup is the highest among the listed methods.

    Could not parse LaTeX table

    Adv-Makeup has the lowest FID because it changes mainly the eye region, but that restriction also gives it poor protection success. The proposed method changes more of the face while remaining more realistic by FID than TIP-IM and AMT-GAN. In the commercial Face++ verification experiment, 100 randomly selected CelebA-HQ faces were protected for targeted impersonation. The bar chart on page 8 reports average confidence scores of 52.852.8 on CelebA-HQ and 73.8273.82 on LADN for the proposed method, the highest values among the compared methods; a higher confidence score indicates stronger similarity to the target returned by the API.

  11. Knowl 11 — Ablation tradeoffs and robustness to makeup wording

    empirical result

    The ablations show that the text and identity components control different properties of the protected image. Removing text guidance allows the adversarial loss to induce visible artifacts, whereas directional makeup guidance hides the changes in a natural-looking style. Optimizing all latent layers increases attack freedom and PSR but can alter the user's visual identity; restricting optimization to the deeper layers with the identity-preserving regularizer retains identity while still applying the requested makeup.

    The latent-loss weight produces an explicit privacy-realism tradeoff. The reported values are:

    Could not parse LaTeX table

    Reducing λlatent\lambda_{\mathrm{latent}} permits greater latent deviation, increasing PSR but worsening FID and therefore naturalness. Finally, when 1,000 CelebA-HQ images are protected using five makeup descriptions—tanned, pale, pink eyeshadows, red lipstick, and Matte—the PSR values are 74.1%74.1\%, 77.3%77.3\%, 78.4%78.4\%, 78.7%78.7\%, and 79.2%79.2\%, respectively, with standard deviation 1.241.24. This small variation supports robustness to the selected textual makeup style.

Coverage note — The paper's supplementary PSNR/SSIM results, detailed Tencent Yunshentu results, and the brief statement that generation is computationally expensive were omitted because they add evaluation detail without introducing a separate load-bearing method or result.

References

  1. 1.Shane Ahern, Dean Eckles, Nathaniel S Good, Simon King, Mor Naaman, and Rahul Nair. Over-exposed? privacy patterns and considerations in online and mobile photo sharing. In Proceedings of the SIGCHI conference on Human factors in computing systems, pages 357–366, 2007.
  2. 2.Amit H Bermano, Rinon Gal, Yuval Alaluf, Ron Mokady, Yotam Nitzan, Omer Tov, Oren Patashnik, and Daniel Cohen-Or. State-of-the-art in the architecture, methods and applications of stylegan. In Computer Graphics Forum, volume 41, pages 591–611. Wiley Online Library, 2022.
  3. 3.Anand Bhattad, Min Jin Chong, Kaizhao Liang, Bo Li, and David A Forsyth. Unrestricted adversarial examples via semantic manipulation. arXiv preprint arXiv:1904.06347, 2019.
  4. 4.Jia-Wei Chen, Li-Ju Chen, Chia-Mu Yu, and Chun-Shien Lu. Perceptual indistinguishability-net (pi-net): Facial image obfuscation with manipulable semantics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6478–6487, 2021.
  5. 5.Sheng Chen, Yang Liu, Xiang Gao, and Zhen Han. Mobilefacenets: Efficient cnns for accurate real-time face verification on mobile devices. In Chinese Conference on Biometric Recognition, pages 428–438. Springer, 2018.
  6. 6.Valeriia Cherepanova, Micah Goldblum, Harrison Foley, Shiyuan Duan, John P Dickerson, Gavin Taylor, and Tom Goldstein. Lowkey: Leveraging adversarial attacks to protect social media users from facial recognition. In International Conference on Learning Representations, 2020.
  7. 7.William L Croft, Jorg-R¨udiger Sack, and Wei Shi. Differentially private facial obfuscation via generative adversarial networks. Future Generation Computer Systems, 129:358–379, 2022.
  8. 8.Ali Dabouei, Sobhan Soleymani, Jeremy Dawson, and Nasser Nasrabadi. Fast geometrically-perturbed adversarial faces. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1979–1988. IEEE, 2019.
  9. 9.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4690–4699, 2019.
  10. 10.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR’18), pages 9185–9193, 2018.
  11. 11.Yinpeng Dong, Tianyu Pang, Hang Su, and Jun Zhu. Evading defenses to transferable adversarial examples by translation-invariant attacks. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR’19), pages 4312–4321, 2019.
  12. 12.Hang Du, Hailin Shi, Dan Zeng, Xiao-Ping Zhang, and Tao Mei. The elements of end-to-end deep face recognition: A survey of recent advances. ACM Computing Surveys (CSUR), 54(10s):1–42, 2022.
  13. 13.Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao. A survey of vision-language pre-trained models. arXiv preprint arXiv:2202.10936, 2022.
  14. 14.Kevin Frans, Lisa B Soros, and Olaf Witkowski. Clipdraw: Exploring text-to-drawing synthesis through language-image encoders. arXiv preprint arXiv:2106.14843, 2021.
  15. 15.Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial networks. CoRR, abs/1406.2661, 2014.
  16. 16.Qiao Gu, Guanzhi Wang, Mang Tik Chiu, Yu-Wing Tai, and Chi-Keung Tang. Ladn: Local adversarial disentangling network for facial makeup and de-makeup. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10481–10490, 2019.
  17. 17.Nitzan Guetta, Asaf Shabtai, Inderjeet Singh, Satoru Momiyama, and Yuval Elovici. Dodging attack using carefully crafted natural makeup. arXiv preprint arXiv:2109.06467, 2021.
  18. 18.Rebecca Heilweil. The world’s scariest facial recognition company explained. Vox, May, 8, 2020.
  19. 19.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  20. 20.Kashmir Hill. The secretive company that might end privacy as we know it. The New York Times, 18:2020, 2020.
  21. 21.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018.
  22. 22.Shengshan Hu, Xiaogeng Liu, Yechao Zhang, Minghui Li, Leo Yu Zhang, Hai Jin, and Libing Wu. Protecting facial privacy: Generating adversarial identity masks via style-robust makeup transfer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15014–15023, 2022.
  23. 23.Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. Labeled faces in the wild: A database forstudying face recognition in unconstrained environments. In Workshop on faces in ’Real-Life’ Images: detection, alignment, and recognition, 2008.
  24. 24.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017.
  25. 25.Kazuya Kakizaki and Kosuke Yoshida. Adversarial image translation: Unrestricted adversarial examples in face recognition systems. arXiv preprint arXiv:1905.03421, 2019.
  26. 26.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In International Conference on Learning Representations, 2018.
  27. 27.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020.
  28. 28.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2426–2435, 2022.
  29. 29.Stepan Komkov and Aleksandr Petiushko. Advhat: Real-world adversarial attack on arcface face id system. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 819–826. IEEE, 2021.
  30. 30.Zhenzhong Kuang, Zhiqiang Guo, Jinglong Fang, Jun Yu, Noboru Babaguchi, and Jianping Fan. Unnoticeable synthetic face replacement for image privacy protection. Neurocomputing, 457:322–333, 2021.
  31. 31.Zhenzhong Kuang, Huigui Liu, Jun Yu, Aikui Tian, Lei Wang, Jianping Fan, and Noboru Babaguchi. Effective de-identification generative adversarial network for face anonymization. In Proceedings of the 29th ACM International Conference on Multimedia, pages 3182–3191, 2021.
  32. 32.Gihyun Kwon and Jong Chul Ye. Clipstyler: Image style transfer with a single text condition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18062–18071, 2022.
  33. 33.Tao Li and Min Soo Choi. Deepblur: A simple and effective method for natural image obfuscation. arXiv preprint arXiv:2104.02655, 1, 2021.
  34. 34.Tao Li and Lei Lin. Anonymousnet: Natural face de-identification with measurable privacy. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 0–0, 2019.
  35. 35.Fangcheng Liu, Chao Zhang, and Hongyang Zhang. Towards transferable unrestricted adversarial examples with minimum changes. arXiv preprint arXiv:2201.01102, 2022.
  36. 36.Suolan Liu, Lizhi Kong, and Hongyuan Wang. Face detection and encryption for privacy preserving in surveillance video. In Chinese Conference on Pattern Recognition and Computer Vision (PRCV), pages 162–172. Springer, 2018.
  37. 37.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In Proceedings of the 6th International Conference on Learning Representations (ICLR’18), 2018.
  38. 38.Blaz Meden, Peter Rot, Philipp Terh¨orst, Naser Damer, Arjan Kuijper, Walter J Scheirer, Arun Ross, Peter Peer, and Vitomir Struc. Privacy–enhancing face biometrics: A comprehensive survey. IEEE Transactions on Information Forensics and Security, 2021.
  39. 39.Dongbin Na, Sangwoo Ji, and Jong Kim. Unrestricted black-box adversarial attack using gan with limited queries. arXiv preprint arXiv:2208.11613, 2022.
  40. 40.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  41. 41.Seong Joon Oh, Rodrigo Benenson, Mario Fritz, and Bernt Schiele. Faceless person recognition: Privacy implications in social media. In European Conference on Computer Vision, pages 19–35. Springer, 2016.
  42. 42.Seong Joon Oh, Mario Fritz, and Bernt Schiele. Adversarial image perturbation for privacy protection a game theory perspective. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 1491–1500. IEEE, 2017.
  43. 43.Omkar M Parkhi, Andrea Vedaldi, and Andrew Zisserman. Deep face recognition. 2015.
  44. 44.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2085–2094, 2021.
  45. 45.P Jonathon Phillips, Amy N Yates, Ying Hu, Carina A Hahn, Eilidh Noyes, Kelsey Jackson, Jacqueline G Cavazos, Geraldine Jeckeln, Rajeev Ranjan, Swami Sankaranarayanan, et al. Face recognition accuracy of forensic examiners, superrecognizers, and face recognition algorithms. Proceedings of the National Academy of Sciences, 115(24):6171–6176, 2018.
  46. 46.Omid Poursaeed, Tianxing Jiang, Harry Yang, Serge Belongie, and Ser-Nam Lim. Robustness and generalization via generative adversarial training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15711–15720, 2021.
  47. 47.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  48. 48.Rajeev Ranjan, Swami Sankaranarayanan, Ankan Bansal, Navaneeth Bodla, Jun-Cheng Chen, Vishal M Patel, Carlos D Castillo, and Rama Chellappa. Deep learning for understanding faces: Machines may be just as good, or better, than humans. IEEE Signal Processing Magazine, 35(1):66–83, 2018.
  49. 49.Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real images. ACM Transactions on Graphics (TOG), 42(1):1–13, 2022.
  50. 50.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. Advances in neural information processing systems, 29, 2016.
  51. 51.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 815–823, 2015.
  52. 52.Sachith Seneviratne, Nuran Kasthuriarachchi, Sanka Rasnayaka, Danula Hettiachchi, and Ridwan Shariffdeen. Does a face mask protect my privacy?: Deep learning to predict protected attributes from masked face images. In Australasian Joint Conference on Artificial Intelligence, pages 91–102. Springer, 2022.
  53. 53.Shawn Shan, Emily Wenger, Jiayun Zhang, Huiying Li, Haitao Zheng, and Ben Y Zhao. Fawkes: Protecting privacy against unauthorized deep learning models. In 29th USENIX security symposium (USENIX Security 20), pages 1589–1604, 2020.
  54. 54.Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K Reiter. A general framework for adversarial examples with objectives. ACM Transactions on Privacy and Security (TOPS), 22(3):1–30, 2019.
  55. 55.Yang Song, Rui Shu, Nate Kushman, and Stefano Ermon. Constructing unrestricted adversarial examples with generative models. Advances in Neural Information Processing Systems, 31, 2018.
  56. 56.Qianru Sun, Liqian Ma, Seong Joon Oh, Luc Van Gool, Bernt Schiele, and Mario Fritz. Natural and effective obfuscation by head inpainting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5050–5059, 2018.
  57. 57.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations, 2014.
  58. 58.Huan Tian, Tianqing Zhu, and Wanlei Zhou. Fairness and privacy preservation for facial images: Gan-based methods. Computers & Security, 122:102902, 2022.
  59. 59.Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. Designing an encoder for stylegan image manipulation. ACM Transactions on Graphics (TOG), 40(4):1–14, 2021.
  60. 60.Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3835–3844, 2022.
  61. 61.Mei Wang and Weihong Deng. Deep face recognition: A survey. Neurocomputing, 429:215–244, 2021.
  62. 62.Shunxin Wang, Una M Kelly, and Raymond NJ Veldhuis. Gender obfuscation through face morphing. In 2021 IEEE International Workshop on Biometrics and Forensics (IWBF), pages 1–6. IEEE, 2021.
  63. 63.Ya Wang, Tianlong Bao, Chunhui Ding, and Ming Zhu. Face recognition in real-world surveillance videos with deep learning method. In 2017 2nd international conference on image, vision and computing (icivc), pages 239–243. IEEE, 2017.
  64. 64.Yinggui Wang, Jian Liu, Man Luo, Le Yang, and Li Wang. Privacy-preserving face recognition in the frequency domain. 2022.
  65. 65.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004.
  66. 66.Tianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao, Zhentao Tan, Lu Yuan, Weiming Zhang, and Nenghai Yu. Hairclip: Design your hair by text and reference image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18072–18081, 2022.
  67. 67.Emily Wenger, Shawn Shan, Haitao Zheng, and Ben Y Zhao. Sok: Anti-facial recognition technology. arXiv preprint arXiv:2112.04558, 2021.
  68. 68.Chaowei Xiao, Jun-Yan Zhu, Bo Li, Warren He, Mingyan Liu, and Dawn Song. Spatially transformed adversarial examples. arXiv preprint arXiv:1801.02612, 2018.
  69. 69.Zihao Xiao, Xianfeng Gao, Chilin Fu, Yinpeng Dong, Wei Gao, Xiaolu Zhang, Jun Zhou, and Jun Zhu. Improving transferability of adversarial patches on face recognition with generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11845–11854, 2021.
  70. 70.Xiao Yang, Yinpeng Dong, Tianyu Pang, Hang Su, Jun Zhu, Yuefeng Chen, and Hui Xue. Towards face encryption by generating adversarial identity masks. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV’21), pages 3897–3907, 2021.
  71. 71.Bangjie Yin, Wenxuan Wang, Taiping Yao, Junfeng Guo, Zelun Kong, Shouhong Ding, Jilin Li, and Cong Liu. Adv-makeup: A new imperceptible and transferable attack on face recognition. In Proceedings of the 30th International Joint Conference on Artificial Intelligence (IJCAI’21), pages 1252–1258, 2021.
  72. 72.Bangjie Yin, Wenxuan Wang, Taiping Yao, Junfeng Guo, Zelun Kong, Shouhong Ding, Jilin Li, and Cong Liu. Adv-makeup: A new imperceptible and transferable attack on face recognition. arXiv preprint arXiv:2105.03162, 2021.
  73. 73.Shengming Yuan, Qilong Zhang, Lianli Gao, Yaya Cheng, and Jingkuan Song. Natural color fool: Towards boosting black-box unrestricted attacks. arXiv preprint arXiv:2210.02041, 2022.
  74. 74.Jiaming Zhang, Jitao Sang, Xian Zhao, Xiaowen Huang, Yanfeng Sun, and Yongli Hu. Adversarial privacy-preserving filter. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1423–1431, 2020.
  75. 75.Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE signal processing letters, 23(10):1499–1503, 2016.
  76. 76.Zhengyu Zhao, Zhuoran Liu, and Martha Larson. Towards large yet imperceptible adversarial image perturbations with perceptual color distance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1039–1048, 2020.
  77. 77.Yaoyao Zhong and Weihong Deng. Opom: Customized invisible cloak towards face privacy protection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  78. 78.Jizhe Zhou and Chi-Man Pun. Personal privacy protection via irrelevant faces tracking and pixelation in video live streaming. IEEE Transactions on Information Forensics and Security, 16:1088–1103, 2020.
  79. 79.Zheng-An Zhu, Yun-Zhong Lu, and Chen-Kuo Chiang. Generating adversarial examples by makeup attacks on face recognition. In 2019 IEEE International Conference on Image Processing (ICIP), pages 2516–2520. IEEE, 2019.

Citation

MLA
Shamshad, F., et al. “CLIP2Protect: Protecting Facial Privacy Using Text-Guided Makeup via Adversarial Latent Search”. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, Pp. 20595-20605, 2023, http://arxiv.org/abs/2306.10008v2.
APA
Shamshad, F., Naseer, M., & Nandakumar, K. (2023). CLIP2Protect: Protecting Facial Privacy using Text-Guided Makeup via Adversarial Latent Search. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, Pp. 20595-20605. http://arxiv.org/abs/2306.10008v2
Chicago
Shamshad, F., M. Naseer, and K. Nandakumar. 2023. “CLIP2Protect: Protecting Facial Privacy Using Text-Guided Makeup via Adversarial Latent Search”. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, Pp. 20595-20605. http://arxiv.org/abs/2306.10008v2.
Harvard
Shamshad, F., Naseer, M. and Nandakumar, K. (2023) “CLIP2Protect: Protecting Facial Privacy using Text-Guided Makeup via Adversarial Latent Search”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 20595-20605 [Preprint]. Available at: http://arxiv.org/abs/2306.10008v2.
Vancouver
1. Shamshad F, Naseer M, Nandakumar K (2023) CLIP2Protect: Protecting Facial Privacy using Text-Guided Makeup via Adversarial Latent Search. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 20595-20605

BibTeX

@article{shamshad2023clip2protect,
  title = {CLIP2Protect: Protecting Facial Privacy using Text-Guided Makeup via Adversarial Latent Search},
  author = {Shamshad, Fahad and Naseer, Muzammal and Nandakumar, Karthik},
  year = {2023},
  journal = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 20595-20605},
  url = {http://arxiv.org/abs/2306.10008v2},
  eprint = {2306.10008}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE