Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition

Zexin LiBangjie YinTaiping YaoJunfeng GuoShouhong DingSimin ChenCong Liu

article2023CVPR58 citations

Proposes a multi-task adversarial framework that leverages gradient information from face attribute recognition to substantially improve black-box attack transferability against commercial face recognition systems.

Listen

Commercial face recognition systems are widely deployed across security and identity-verification applications, yet they remain susceptible to adversarial attacks where carefully modified images fool the underlying software. In practical security contexts, attackers lack internal access to proprietary target systems (a black-box setting) and must rely on transferable attacks created on a local model. Existing methods rely strictly on single-task face recognition information and struggle against modern enterprise systems, typically achieving attack success rates below 50% against commercial platforms.

The article develops and evaluates Sibling-Attack, an adversarial attack method designed to significantly increase transferability against black-box face recognition systems by incorporating information from a complementary auxiliary task—specifically Attribute Recognition. The researchers designed a multi-task optimization framework using a shared-parameter architecture alongside two core algorithms: an alternating Joint-Task Meta Optimization method to align cross-task gradient directions and a Cross-Task Gradient Stabilization technique to prevent optimization instability. Experiments evaluated targeted impersonation attacks across 1,000 face pairs from two benchmark datasets (CelebA-HQ and LFW), testing against multiple offline face recognition models as well as two widely used commercial cloud platforms, Face++ and Microsoft Azure Face API.

The findings show that Sibling-Attack outperforms existing state-of-the-art attack approaches by substantial margins. Across offline target models, the proposed method increased the average attack success rate by 12.61 percentage points compared to existing techniques. Against commercial online systems, it increased the average success rate by 55.77 percentage points. Notably, when attacking Face++, Sibling-Attack attained success rates of 86.50% on CelebA-HQ and 96.10% on LFW, whereas prior state-of-the-art methods achieved at most 58.10% and 64.30%, respectively. It also maintained visual indistinguishability comparable to baseline methods while achieving higher cross-system transferability.

These results demonstrate a critical security vulnerability: proprietary, black-box facial recognition systems are significantly more exposed to transferable digital attacks than previously assumed. The findings confirm that multi-task learning representations capture generalizable facial features that make adversarial examples substantially more potent across different model architectures and commercial deployments. To mitigate these risks, organizations deploying face recognition systems should strengthen their defenses by implementing targeted adversarial training and specialized image de-noising filters, while avoiding reliance on architectural obscurity alone. Confidence in the empirical results is high across digital evaluation benchmarks, though the scope remains bounded by digital image inputs and limited to the evaluated commercial APIs.

arXiv: 2303.12512

No sufficiently relevant recommendations were found.

Cover for Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition

Abstract

A hard challenge in developing practical face recognition (FR) attacks is due to the black-box nature of the target FR model, i.e., inaccessible gradient and parameter information to attackers. While recent research took an important step towards attacking black-box FR models through leveraging transferability, their performance is still limited, especially against online commercial FR systems that can be pessimistic (e.g., a less than 50% ASR–attack success rate on average). Motivated by this, we present Sibling-Attack, a new FR attack technique for the first time explores a novel multi-task perspective (i.e., leveraging extra information from multi-correlated tasks to boost attacking transferability). Intuitively, Sibling-Attack selects a set of tasks correlated with FR and picks the Attribute Recognition (AR) task as the task used in Sibling-Attack based on theoretical and quantitative analysis. Sibling-Attack then develops an optimization framework that fuses adversarial gradient information through (1) constraining the cross-task features to be under the same space, (2) a joint-task meta optimization framework that enhances the gradient compatibility among tasks, and (3) a cross-task gradient stabilization method which mitigates the oscillation effect during attacking. Extensive experiments demonstrate that Sibling-Attack outperforms state-of-the-art FR attack techniques by a non-trivial margin, boosting ASR by 12.61% and 55.77% on average on state-of-the-art pre-trained FR models and two well-known, widely used commercial FR systems.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Adversarial Attacks
  • 2.2. Multi-task Learning
  • 3. Methodology
  • 3.1. Overview
  • 3.2. Sibling-Attack Framework
  • 3.3. Joint-Task Meta Optimization
  • 3.4. Cross-Task Gradient Stabilization
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Why Select the AR Task?
  • 4.3. Experimental Results
  • 4.4. Ablation Study
  • 4.5. Visualization and Analysis
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Sibling-task attack architecture

    model/method

    Sibling-Attack generates transferable targeted impersonation examples for a black-box face-recognition target without accessing the target model’s parameters or gradients. It trains a white-box surrogate with hard parameter sharing, denoted S(P;F;A)S(P;F;A), where PP is a shared encoder, FF is a face-recognition branch, and AA is a facial-attribute-recognition branch. For an attacker image and a victim image, both branches produce high-level feature embeddings and contribute adversarial gradients. Sharing the encoder constrains the two tasks to use a common feature space, reducing the feature-space variance that would arise from independently optimized face-recognition and attribute-recognition models.

  2. Knowl 2 — Joint impersonation objective

    equation

    Let xa,xv∈[0,1]H×W×Cx_a,x_v\in[0,1]^{H\times W\times C} be the attacker and victim face images, respectively, and let ϵa∈RH×W×C\epsilon_a\in\mathbb{R}^{H\times W\times C} be an additive perturbation constrained by ∥ϵa∥p≤ξ\|\epsilon_a\|_p\leq\xi, where p∈{0,2,∞}p\in\{0,2,\infty\} and ξ>0\xi>0 is the perturbation budget. For branch ∗∈{F,A}*\in\{F,A\}, let fa∗f_a^* and fv∗f_v^* be the branch embeddings of xa+ϵax_a+\epsilon_a and xvx_v. The branch impersonation loss is one minus cosine similarity:

    Ladv∗=1−cos⁡(fa∗,fv∗),∗∈{F,A}.L_{\mathrm{adv}}^*=1-\operatorname{cos}(f_a^*,f_v^*),\qquad *\in\{F,A\}.

    Sibling-Attack minimizes the weighted losses of both branches:

    min⁡ϵa  λ1LadvF+λ2LadvAsubject to∥ϵa∥p≤ξ,\min_{\epsilon_a}\;\lambda_1L_{\mathrm{adv}}^F+\lambda_2L_{\mathrm{adv}}^A \quad\text{subject to}\quad \|\epsilon_a\|_p\leq\xi,

    where λ1,λ2≥0\lambda_1,\lambda_2\geq0 are task trade-off weights. The targeted objective makes the perturbed attacker image resemble the victim image in both the face-recognition and attribute-recognition surrogate feature spaces.

  3. Knowl 3 — Joint-Task Meta Optimization

    model/method

    Joint-Task Meta Optimization (JTMO) improves compatibility between adversarial gradients from the face-recognition and attribute-recognition branches by alternating the branches within every optimization iteration rather than directly averaging their losses. Let ∗∈{F,A}*\in\{F,A\} be the selected branch and let ∗ˉ\bar * denote the other branch. With current perturbation ϵa\epsilon_a, learning rate α>0\alpha>0, branch-update weights γ1,γ2\gamma_1,\gamma_2, and projection Π\Pi onto the feasible ℓ∞\ell_\infty perturbation set, JTMO first computes a provisional update using the selected branch:

    ϵ~a=Π ⁣(ϵa−α sign⁡ ⁣(γ1∇ϵaLadv∗(xa+ϵa,xv))).\widetilde{\epsilon}_a=\Pi\!\left(\epsilon_a-\alpha\,\operatorname{sign}\!\left(\gamma_1\nabla_{\epsilon_a}L_{\mathrm{adv}}^*(x_a+\epsilon_a,x_v)\right)\right).

    It then evaluates the other branch after this provisional update and makes a second update:

    ϵanew=Π ⁣(ϵ~a−α sign⁡ ⁣(γ2∇ϵ~aLadv∗ˉ(xa+ϵ~a,xv))).\epsilon_a^{\mathrm{new}}=\Pi\!\left(\widetilde{\epsilon}_a-\alpha\,\operatorname{sign}\!\left(\gamma_2\nabla_{\widetilde{\epsilon}_a}L_{\mathrm{adv}}^{\bar *}(x_a+\widetilde{\epsilon}_a,x_v)\right)\right).

    Here ∇\nabla denotes the pixel gradient, sign⁡\operatorname{sign} is applied elementwise, and Π\Pi enforces the perturbation constraint. The branch order can be reversed without changing the reported final performance.

  4. Knowl 4 — Cross-Task Gradient Stabilization

    model/method

    Cross-Task Gradient Stabilization (CTGS) reduces oscillation caused by conflicting update directions from the two tasks. Suppose the face-recognition branch FF is selected for an iteration, and NN consecutive provisional perturbations ϵa1′F,…,ϵaN′F\epsilon_{a1}^{\prime F},\ldots,\epsilon_{aN}^{\prime F} are generated using face-recognition gradients. The attribute-recognition branch AA evaluates each provisional example and supplies gradients giA=∇ϵai′FLadvA(xa+ϵai′F,xv)g_i^A=\nabla_{\epsilon_{ai}^{\prime F}}L_{\mathrm{adv}}^A(x_a+\epsilon_{ai}^{\prime F},x_v). The final update on branch AA uses the current gradient together with the preceding N−1N-1 historical gradients:

    ϵa′′A=Π ⁣(ϵaN′F−α sign⁡ ⁣[γ2(gNA+γ3∑i=1N−1giA)]).\epsilon_a^{\prime\prime A}=\Pi\!\left(\epsilon_{aN}^{\prime F}-\alpha\,\operatorname{sign}\!\left[\gamma_2\left(g_N^A+\gamma_3\sum_{i=1}^{N-1}g_i^A\right)\right]\right).

    The same construction applies with the tasks exchanged. Here α\alpha is the pixel step size, γ2\gamma_2 weights the current cross-task gradient, γ3\gamma_3 controls the contribution of historical gradients, and Π\Pi projects onto the allowed perturbation set. The historical gradients provide auxiliary information while remaining weaker than the current update.

  5. Knowl 5 — End-to-end Sibling-Attack optimization procedure

    algorithm

    The attack takes attacker images, victim images, a pretrained hard-sharing surrogate S(P;F;A)S(P;F;A), an iteration count TT, a within-iteration update count NN, perturbation budget ξ\xi, step size α\alpha, and weights γ1,γ2,γ3\gamma_1,\gamma_2,\gamma_3. It returns the final perturbation.

    Input: attacker image xax_a, victim image xvx_v, surrogate S(P;F;A)S(P;F;A), iterations TT, update count NN, budget ξ\xi, step size α\alpha, weights γ1,γ2,γ3\gamma_1,\gamma_2,\gamma_3
    Output: optimized perturbation ϵaopt\epsilon_a^{opt}
    Initialize the perturbation and set the adversarial image to xax_a
    For each outer iteration from 1 to TT
        Select one branch, either FF or AA
        Starting from the current perturbation, perform NN consecutive projected sign updates on the selected branch
        Store all NN provisional perturbations
        Evaluate the other branch on every stored provisional adversarial image
        Aggregate the current gradient and the previous N−1N-1 cross-task gradients using CTGS
        Perform the projected sign update on the other branch
        Update the adversarial image and retain the resulting perturbation
    Return the final perturbation

    For the reported experiments, the surrogate used an IR152 backbone for both task branches, ξ=40/255\xi=40/255 under an ℓ∞\ell_\infty constraint, α=2/255\alpha=2/255, T=200T=200, N=4N=4, and (γ1,γ2,γ3)=(0.1,0.9,0.01)(\gamma_1,\gamma_2,\gamma_3)=(0.1,0.9,0.01). The attack is trained only on the white-box surrogate; the black-box face-recognition system is used for transfer evaluation rather than gradient computation.

  6. Knowl 6 — Attribute recognition is the selected sibling task

    data/table

    The paper selects facial attribute recognition (AR) as the sibling task because face-recognition representations implicitly encode facial attributes, while AR models can learn robust identity-relevant features. The authors compare AR with face landmark detection (FLD) and face parsing (FP) using the same hard-parameter-sharing architecture and the same cosine-based branch losses. Attack success rate (ASR, in percent) is measured against offline IR50 and ResNet101 targets on CelebA-HQ and LFW.

    Could not parse LaTeX table

    AR is the only tested auxiliary task that is consistently best across all four dataset-target combinations. The weaker FR+FLD and FR+FP results show that merely adding a face-related task is insufficient; the auxiliary task must provide complementary information that is strongly coupled to identity.

  7. Knowl 7 — Experimental evaluation protocol

    experimental setup

    The evaluation uses 1,000 randomly sampled pairs of different-identity faces from each of CelebA-HQ and LFW. CelebA-HQ contains 30,000 high-quality face images, while LFW contains 13,233 images of 5,749 subjects. Attack success rate is defined as the fraction of comparisons for which the black-box system assigns similarity at least its acceptance threshold τ\tau:

    ASR⁡=number of comparisons with similarity≥τtotal number of comparisons.\operatorname{ASR}=\frac{\text{number of comparisons with similarity}\geq\tau}{\text{total number of comparisons}}.

    Offline target models are IR50 and ResNet101; IR152, IRSE50, and FaceNet are used as white-box face-recognition source models. Online targets are the Face++ and Microsoft face-verification systems. For offline evaluation, thresholds are 0.2770.277 for IR50 and 0.2000.200 for ResNet101. Face++ uses its cosine-similarity threshold at a false-positive rate of 0.0010.001; Microsoft exposes only a score and decision for each query, so its reported successful decisions are counted directly. The baselines include Adv-Hat, Adv-Glasses, Adv-Face, Adv-Makeup, GenAP, PGD, TAP, MI-FGSM, and VMI-FGSM. Sibling-Attack uses an IR152-based FR/AR surrogate, while competing transfer attacks use two face-recognition source models for comparison fairness.

  8. Knowl 8 — Black-box impersonation performance

    data/table

    Sibling-Attack substantially improves transferability on both offline and commercial face-recognition targets. The entries below are ASR percentages; each parenthetical value gives the strongest competing baseline under the same source-model and target condition.

    Could not parse LaTeX table

    The method dominates the strongest transfer-based or face-based competitor in every listed condition. Averaged over the reported offline pretrained targets, it improves ASR by 12.6112.61 percentage points; averaged over the online commercial targets, it improves ASR by 55.7755.77 percentage points. In particular, its Face++ ASR reaches 86.50%86.50\% on CelebA-HQ and 96.10%96.10\% on LFW, compared with the strongest competing values of 58.10%58.10\% and 64.30%64.30\% reported for the corresponding best overall comparisons.

  9. Knowl 9 — Component ablation

    data/table

    The ablation isolates the contributions of AR information, hard parameter sharing, JTMO, and CTGS on LFW. ASR values are percentages against IR50, ResNet101, Face++, and Microsoft. Single-model and ensemble rows are conventional face-recognition attacks; the Sibling-Attack rows add the indicated component progressively.

    Could not parse LaTeX table

    The basic FR+AR framework already improves over most face-recognition-only alternatives, supporting the value of cross-task information. Hard sharing raises the two online ASRs from 69.80%69.80\% to 77.40%77.40\% and from 37.20%37.20\% to 45.40%45.40\%. JTMO adds 18.1018.10 and 5.805.80 percentage points over hard sharing on Face++ and Microsoft, respectively. CTGS adds a further 0.600.60 and 8.108.10 percentage points, showing that all three proposed components contribute to the final transfer performance.

  10. Knowl 10 — Transferability and perceptual analyses

    empirical result

    Additional analyses support both the mechanism and the scope of Sibling-Attack. Grad-CAM responses on an offline FR model show that conventional black-box ensemble attacks often emphasize background regions or overfit to local facial areas, whereas Sibling-Attack and the white-box target response focus on more similar key facial regions. This alignment is used as qualitative evidence for stronger transferability.

    On LFW, the reported perceptual metrics are:

    Could not parse LaTeX table

    Sibling-Attack has competitive structural similarity and mean squared error while achieving much higher black-box ASR. For the 18 facial attributes used to train the AR model, attributes are grouped into eye, nose, mouth, and other regions; mouth-region attributes produce weaker FR transferability than the other groups. The method also transfers information in the reverse direction against a black-box AR target. For 1,000 image pairs, the summed attribute-score difference is computed as

    D=∑s=11000∥AB(xadvs)−AB(xs)∥1,D=\sum_{s=1}^{1000}\left\|A_B(x_{\mathrm{adv}}^s)-A_B(x^s)\right\|_1,

    where AB(⋅)A_B(\cdot) is the black-box AR model’s vector of attribute scores and xsx^s is the original image. For eye, nose, mouth, and other attribute groups, a white-box AR baseline obtains 148.85148.85, 223.97223.97, 184.77184.77, and 201.79201.79, whereas Sibling-Attack obtains 162.41162.41, 241.29241.29, 195.46195.46, and 214.02214.02, respectively. Thus, the cross-task attack improves transfer not only toward FR targets but also toward AR targets.

Coverage note — The conclusion-level discussion that the method is currently focused on digital attacks, along with mitigation suggestions and future extension to other tasks, was omitted because it adds no separate methodological or quantitative result.

References

  1. 1.Jonathan Baxter. A bayesian/information theoretic model of learning to learn via multiple task sampling. Machine learning, 28(1):7–39, 1997.
  2. 2.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In 2017 ieee symposium on security and privacy (sp), pages 39–57. IEEE, 2017.
  3. 3.Rich Caruana. Multitask learning. Machine learning, 28(1):41–75, 1997.
  4. 4.Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In Proceedings of the 10th ACM Workshop On Artificial Intelligence and Security, pages 15–26, 2017.
  5. 5.Simin Chen, Soroush Bateni, Sampath Grandhi, Xiaodi Li, Cong Liu, and Wei Yang. Denas: automated rule generation by knowledge extraction from neural networks. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 813–825, 2020.
  6. 6.Simin Chen, Mirazul Haque, Cong Liu, and Wei Yang. Deepperform: An efficient approach for performance testing of resource-constrained neural networks. In 37th IEEE/ACM International Conference on Automated Software Engineering, pages 1–13, 2022.
  7. 7.Simin Chen, Hamed Khanpour, Cong Liu, and Wei Yang. Learn to reverse dnns from AI programs automatically. In Luc De Raedt, editor, Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, Vienna, Austria, 23-29 July 2022, pages 666–672. ijcai.org, 2022.
  8. 8.Simin Chen, Cong Liu, Mirazul Haque, Zihe Song, and Wei Yang. Nmtsloth: understanding and testing efficiency degradation of neural machine translation systems. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pages 1148–1160, 2022.
  9. 9.Sheng Chen, Yang Liu, Xiang Gao, and Zhen Han. Mobilefacenets: Efficient cnns for accurate real-time face verification on mobile devices. In Proceedings of the Chinese Conference on Biometric Recognition (CCBR), pages 428–438, 2018.
  10. 10.Simin Chen, Zihe Song, Mirazul Haque, Cong Liu, and Wei Yang. Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15365–15374, 2022.
  11. 11.Yiming Chen, Yan Zhang, Bin Wang, Zuozhu Liu, and Haizhou Li. Generate, discriminate and contrast: A semi-supervised sentence representation learning framework. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8150–8161, Abu Dhabi, United Arab Emirates, Dec. 2022. Association for Computational Linguistics.
  12. 12.Yiming Chen, Yan Zhang, Chen Zhang, Grandee Lee, Ran Cheng, and Haizhou Li. Revisiting self-training for few-shot learning of language model. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9125–9135, Online and Punta Cana, Dominican Republic, Nov. 2021. Association for Computational Linguistics.
  13. 13.Debayan Deb, Jianbang Zhang, and Anil K Jain. Advfaces: Adversarial face synthesis. In 2020 IEEE International Joint Conference on Biometrics (IJCB), pages 1–10. IEEE, 2020.
  14. 14.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4690–4699, 2019.
  15. 15.Matheus Alves Diniz and William Robson Schwartz. Face attributes as cues for deep face recognition understanding. In 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020), pages 307–313. IEEE, 2020.
  16. 16.Xin Dong, Junfeng Guo, Ang Li, Wei-Te Ting, Cong Liu, and HT Kung. Neural mean discrepancy for efficient out-of-distribution detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19217–19227, 2022.
  17. 17.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9185–9193, 2018.
  18. 18.Yinpeng Dong, Tianyu Pang, Hang Su, and Jun Zhu. Evading defenses to transferable adversarial examples by translation-invariant attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4312–4321, 2019.
  19. 19.Yinpeng Dong, Hang Su, Baoyuan Wu, Zhifeng Li, Wei Liu, Tong Zhang, and Jun Zhu. Efficient decision-based black-box adversarial attacks on face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7714–7722, 2019.
  20. 20.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the International Conference on Machine Learning (ICML), pages 1126–1135. PMLR, 2017.
  21. 21.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the International Conference on Machine Learning (ICML), pages 1126–1135. PMLR, 2017.
  22. 22.Salah Ghamizi, Maxime Cordy, Mike Papadakis, and Yves Le Traon. Adversarial robustness in multi-task learning: Promises and illusions. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 697–705, 2022.
  23. 23.Ian Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR), 2015.
  24. 24.Junfeng Guo, Ang Li, and Cong Liu. AEVA: Black-box backdoor detection using adversarial extreme value analysis. In International Conference on Learning Representations, 2022.
  25. 25.Junfeng Guo, Yiming Li, Xun Chen, Hanqing Guo, Lichao Sun, and Cong Liu. SCALE-UP: An efficient black-box input-level backdoor detection via analyzing scaled prediction consistency. In The Eleventh International Conference on Learning Representations, 2023.
  26. 26.Junfeng Guo and Cong Liu. Practical poisoning attacks on neural networks. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVII 16, pages 142–158. Springer, 2020.
  27. 27.Pengxin Guo, Yuancheng Xu, Baijiong Lin, and Yu Zhang. Multi-task adversarial attack. arXiv preprint arXiv:2011.09824, 2020.
  28. 28.Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao. Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. In Proceedings of the European Conference on Computer Vision (ECCV), pages 87–102. Springer, 2016.
  29. 29.Naresh Kumar Gurulingan, Elahe Arani, and Bahram Zonooz. Uninet: A unified scene understanding network and exploring multi-task relationships through the lens of adversarial attacks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 2239–2248, 2021.
  30. 30.Bangyan He, Jian Liu, Yiming Li, Siyuan Liang, Jingzhi Li, Xiaojun Jia, and Xiaochun Cao. Generating transferable 3d adversarial point cloud via random perturbation factorization. In AAAI, 2023.
  31. 31.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016.
  32. 32.Guosheng Hu, Yang Hua, Yang Yuan, Zhihong Zhang, Zheng Lu, Sankha S Mukherjee, Timothy M Hospedales, Neil M Robertson, and Yongxin Yang. Attribute-enhanced face recognition with neural tensor fusion networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 3744–3753, 2017.
  33. 33.Gary B Huang, Marwan Mattar, Tamara Berg, and Eric Learned-Miller. Labeled faces in the wild: A database forstudying face recognition in unconstrained environments. In Workshop on faces in’Real-Life’Images: detection, alignment, and recognition, 2008.
  34. 34.Yi Huang and Adams Wai-Kin Kong. Transferable adversarial attack based on integrated gradients. In The Tenth International Conference on Learning Representations (ICLR). OpenReview.net, 2022.
  35. 35.Shuai Jia, Bangjie Yin, Taiping Yao, Shouhong Ding, Chunhua Shen, Xiaokang Yang, and Chao Ma. Adv-attribute: Inconspicuous and transferable adversarial attack on face recognition. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems (NIPS), 2022.
  36. 36.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations (ICLR), 2018.
  37. 37.S. Komkov and A. Petiushko. Advhat: Real-world adversarial attack on arcface face id system. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 819–826, Los Alamitos, CA, USA, jan 2021. IEEE Computer Society.
  38. 38.Zelun Kong, Junfeng Guo, Ang Li, and Cong Liu. Physgan: Generating physical-world-resilient adversarial examples for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14254–14263, 2020.
  39. 39.Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. Adversarial examples in the physical world. In Artificial intelligence safety and security, pages 99–112. Chapman and Hall/CRC, 2018.
  40. 40.Yiming Li, Baoyuan Wu, Yan Feng, Yanbo Fan, Yong Jiang, Zhifeng Li, and Shu-Tao Xia. Semi-supervised robust training with generalized perturbed neighborhood. Pattern Recognition, 124:108472, 2022.
  41. 41.Jiadong Lin, Chuanbiao Song, Kun He, Liwei Wang, and John E. Hopcroft. Nesterov accelerated gradient and scale invariance for adversarial attacks. In International Conference on Learning Representations (ICLR), 2020.
  42. 42.Shikun Liu, Edward Johns, and Andrew J Davison. End-to-end multi-task learning with attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1871–1880, 2019.
  43. 43.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black-box attacks. In Proceedings of 5th International Conference on Learning Representations (ICLR), 2017.
  44. 44.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE/CVF International Conference On Computer Vision (ICCV), pages 3730–3738, 2015.
  45. 45.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. International Conference on Learning Representations (ICLR), 2018.
  46. 46.Chengzhi Mao, Amogh Gupta, Vikram Nitin, Baishakhi Ray, Shuran Song, Junfeng Yang, and Carl Vondrick. Multitask learning strengthens adversarial robustness. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16 (ECCV), pages 158–174. Springer, 2020.
  47. 47.Hans Marmolin. Subjective mse measures. IEEE transactions on systems, man, and cybernetics, 16(3):486–489, 1986.
  48. 48.MEGVII. Online face verification. https://www.faceplusplus.com.cn/, 2021.
  49. 49.Fei Miao, Sihong He, Lynn Pepin, Shuo Han, Abdeltawab Hendawi, Mohamed E Khalefa, John A Stankovic, and George Pappas. Data-driven distributionally robust optimization for vehicle balancing of mobility-on-demand systems. ACM Transactions on Cyber-Physical Systems, 5(2):1–27, 2021.
  50. 50.Microsoft. Online face verification. https://azure.microsoft.com/, 2021.
  51. 51.Alex Nichol and John Schulman. Reptile: a scalable metalearning algorithm. arXiv preprint arXiv:1803.02999, 2(3):4, 2018.
  52. 52.Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami. Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security (ASIA CCS), pages 506–519, 2017.
  53. 53.Florian Schroff, Kalenichenko Dmitry, and Philbin James. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  54. 54.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision (ICCV), pages 618–626, 2017.
  55. 55.Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. Advances in neural information processing systems, 31, 2018.
  56. 56.Rui Shao, Xiangyuan Lan, and Pong C Yuen. Regularized fine-grained meta face anti-spoofing. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 34, pages 11974–11981, 2020.
  57. 57.Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K Reiter. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of the 2016 acm sigsac conference on computer and communications security (CCS), pages 1528–1540, 2016.
  58. 58.Trevor Standley, Amir Zamir, Dawn Chen, Leonidas Guibas, Jitendra Malik, and Silvio Savarese. Which tasks should be learned together in multi-task learning? In International Conference on Machine Learning (ICML), pages 9120–9132. PMLR, 2020.
  59. 59.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.
  60. 60.Fariborz Taherkhani, Nasser M Nasrabadi, and Jeremy Dawson. A deep face identification network enhanced by facial attributes prediction. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops (CVPR), pages 553–560, 2018.
  61. 61.Wenxuan Wang, Bangjie Yin, Taiping Yao, Li Zhang, Yanwei Fu, Shouhong Ding, Jilin Li, Feiyue Huang, and Xiangyang Xue. Delving into data: Effectively substitute training for black-box attack. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4761–4770, 2021.
  62. 62.Xiaosen Wang and Kun He. Enhancing the transferability of adversarial attacks through variance tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1924–1933, 2021.
  63. 63.Zhou Wang and Alan C Bovik. A universal image quality index. IEEE signal processing letters, 9(3):81–84, 2002.
  64. 64.Zhanxiong Wang, Keke He, Yanwei Fu, Rui Feng, Yu-Gang Jiang, and Xiangyang Xue. Multi-task deep neural network for joint face recognition and facial attribute prediction. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval (ICMR), pages 365–374, 2017.
  65. 65.Weibin Wu, Yuxin Su, Michael R Lyu, and Irwin King. Improving the transferability of adversarial samples with adversarial transformations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9024–9033, 2021.
  66. 66.Zihao Xiao, Xianfeng Gao, Chilin Fu, Yinpeng Dong, Wei Gao, Xiaolu Zhang, Jun Zhou, and Jun Zhu. Improving transferability of adversarial patches on face recognition with generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11845–11854, 2021.
  67. 67.Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L Yuille. Improving transferability of adversarial examples with input diversity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2730–2739, 2019.
  68. 68.Yifeng Xiong, Jiadong Lin, Min Zhang, John E Hopcroft, and Kun He. Stochastic variance reduced ensemble adversarial attack for boosting the adversarial transferability. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14983–14992, 2022.
  69. 69.Bangjie Yin, Wenxuan Wang, Taiping Yao, Junfeng Guo, Zelun Kong, Shouhong Ding, Jilin Li, and Cong Liu. Adv-makeup: A new imperceptible and transferable attack on face recognition. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI), pages 1252–1258. International Joint Conferences on Artificial Intelligence Organization, 8 2021. Main Track.
  70. 70.Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan. Theoretically principled trade-off between robustness and accuracy. In International Conference on Machine Learning (ICML), pages 7472–7482. PMLR, 2019.
  71. 71.Jianping Zhang, Weibin Wu, Jen-tse Huang, Yizhan Huang, Wenxuan Wang, Yuxin Su, and Michael R Lyu. Improving adversarial transferability via neuron attribution-based attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14993–15002, 2022.
  72. 72.Yaoyao Zhong and Weihong Deng. Towards transferable adversarial attack against deep face recognition. IEEE Transactions on Information Forensics and Security (TIFS), 16:1452–1466, 2020.
  73. 73.Husheng Zhou, Wei Li, Zelun Kong, Junfeng Guo, Yuqun Zhang, Bei Yu, Lingming Zhang, and Cong Liu. Deep-billboard: Systematic physical-world testing of autonomous driving systems. In Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering, pages 347–358, 2020.
  74. 74.Ling Zhou, Zhen Cui, Chunyan Xu, Zhenyu Zhang, Chaoqun Wang, Tong Zhang, and Jian Yang. Pattern-structure diffusion for multi-task learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4514–4523, 2020.
  75. 75.Wen Zhou, Xin Hou, Yongjun Chen, Mengyun Tang, Xiangqi Huang, Xiang Gan, and Yong Yang. Transferable adversarial perturbations. In Proceedings of the European Conference on Computer Vision (ECCV), pages 452–467, 2018.

Citation

MLA
Li, Z., et al. “Sibling-Attack: Rethinking Transferable Adversarial Attacks Against Face Recognition”. arXiv, 2023, http://arxiv.org/abs/2303.12512v1.
APA
Li, Z., Yin, B., Yao, T., Guo, J., Ding, S., Chen, S., & Liu, C. (2023). Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition. arXiv. http://arxiv.org/abs/2303.12512v1
Chicago
Li, Z., B. Yin, T. Yao, et al. 2023. “Sibling-Attack: Rethinking Transferable Adversarial Attacks Against Face Recognition”. arXiv. http://arxiv.org/abs/2303.12512v1.
Harvard
Li, Z. et al. (2023) “Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.12512v1.
Vancouver
1. Li Z, Yin B, Yao T, Guo J, Ding S, Chen S, Liu C (2023) Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition. arXiv

BibTeX

@article{li2023sibling,
  title = {Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition},
  author = {Li, Zexin and Yin, Bangjie and Yao, Taiping and Guo, Juefeng and Ding, Shouhong and Chen, Simin and Liu, Cong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.12512v1},
  eprint = {2303.12512}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE