TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text

Muchao YeChenglin MiaoTing WangFenglong Ma

article2022AAAI64 citations

Proposes TextHoaxer, a gradient-based framework that formulates hard-label text adversarial attacks in continuous word embedding space to generate high-similarity adversarial examples under strict query budgets without relying on query-expensive genetic algorithms.

Listen

Deep neural networks are widely deployed across natural language processing systems, yet their vulnerability to adversarial manipulation poses significant security risks. In commercial environments, these models are typically accessible only through restricted application programming interfaces that provide final discrete decisions rather than detailed confidence probabilities, while also enforcing strict query rate limits to prevent exploitation. Existing evaluation techniques rely on genetic algorithms that require thousands of repeated queries across large candidate pools, causing them to fail when query budgets are limited. The article addresses this operational gap by introducing TextHoaxer, a framework designed to generate high-quality adversarial text samples efficiently under tight query limits in decision-only settings.

The researchers formulated the attack process as a gradient-based optimization task within a continuous word embedding space, bypassing the inefficiency of combinatorial discrete search. Instead of managing a large population of candidates, TextHoaxer refines a single candidate text using an objective function that balances overall semantic meaning, individual word-level distances, and the total number of altered words. The authors evaluated the approach across eight benchmark datasets covering text classification and natural language inference against three widely used model architectures (BERT, WordCNN, and WordLSTM) under a tight limit of 1,000 queries, supported by human evaluation panels.

The findings show that TextHoaxer consistently outperformed existing methods by generating adversarial examples with higher semantic preservation and fewer modified words. On benchmark sentiment classification tasks, TextHoaxer achieved semantic similarity scores 3.7 to 4.8 percentage points higher than the strongest competing hard-label methods, while decreasing the required word perturbation rate by roughly 2.0 to 2.6 percentage points. Testing across varying budgets from 100 to 1,000 queries demonstrated that TextHoaxer reliably produced superior text quality at all constraint levels. Independent human reviewers also ranked its output among the most semantically faithful to original text.

These results demonstrate that standard API query throttling and decision-only output restrictions do not provide sufficient defense against sophisticated, low-query adversarial attacks. Organizations deploying language models should not rely solely on query rate limiting for protection, but should instead incorporate budgeted adversarial evaluation into model robustness auditing and defense pipelines. While the study provides strong confidence across standard benchmarks, it focused primarily on word-level synonym replacements within classification tasks; future work should explore broader multi-word perturbations, defensive hardening strategies, and defenses across larger modern generative architectures.

Cover for TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text

Abstract

This paper focuses on a newly challenging setting in hard-label adversarial attacks on text data by taking the budget information into account. Although existing approaches can successfully generate adversarial examples in the hard-label setting, they follow an ideal assumption that the victim model does not restrict the number of queries. However, in real-world applications the query budget is usually tight or limited. Moreover, existing hard-label adversarial attack techniques use the genetic algorithm to optimize discrete text data by maintaining a number of adversarial candidates during optimization, which can lead to the problem of generating low-quality adversarial examples in the tight-budget setting. To solve this problem, in this paper, we propose a new method named TextHoaxer by formulating the budgeted hard-label adversarial attack task on text data as a gradient-based optimization problem of perturbation matrix in the continuous word embedding space. Compared with the genetic algorithm-based optimization, our solution only uses a single initialized adversarial example as the adversarial candidate for optimization, which significantly reduces the number of queries. The optimization is guided by a new objective function consisting of three terms, i.e., semantic similarity term, pair-wise perturbation constraint, and sparsity constraint. Semantic similarity term and pair-wise perturbation constraint can ensure the high semantic similarity of adversarial examples from both comprehensive text-level and individual word-level, while the sparsity constraint explicitly restricts the number of perturbed words, which is also helpful for enhancing the quality of generated text. We conduct extensive experiments on eight text datasets against three representative natural language models, and experimental results show that TextHoaxer can generate high-quality adversarial examples with higher semantic similarity and lower perturbation rate under the tight-budget setting.

Table of Contents

  • Introduction
  • Methodology
  • Problem Formulation
  • The Proposed TextHoaxer
  • Experiments
  • Experimental Settings
  • Experimental Results
  • Related Work
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — TextHoaxer’s single-candidate continuous attack framework

    model/method

    TextHoaxer attacks a text classifier in the hard-label black-box setting while respecting a limited query budget. It maintains only one initialized adversarial text candidate rather than a population of candidates. The candidate is represented by a perturbation matrix in a continuous word-embedding space, optimized using zeroth-order gradient estimates, and mapped back to discrete words through synonym selection. The victim model is queried only for its predicted class label; its probabilities, parameters, and gradients are not required. This design is intended to avoid the repeated queries required by genetic-algorithm attacks that maintain and evaluate many adversarial candidates.

  2. Knowl 2 — Hard-label text-adversarial attack formulation

    definition

    Let $x=[w_1,[... ELLIPSIZATION ...]

Coverage note — No substantial contributed material was deliberately omitted; the extraction covers the attack formulation, objective, optimization procedure, experimental protocol, quantitative results, budget analysis, human evaluation, and qualitative examples.

References

  1. 1.Bowman, S. R.; Angeli, G.; Potts, C.; and Manning, C. D. 2015. A large annotated corpus for learning natural language inference. In EMNLP, 632–642. The Association for Computational Linguistics.
  2. 2.Brendel, W.; Rauber, J.; and Bethge, M. 2018. Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models. In ICLR. OpenReview.net.
  3. 3.Carlini, N.; and Wagner, D. A. 2017. Towards Evaluating the Robustness of Neural Networks. In S&P, 39–57. IEEE Computer Society.
  4. 4.Cer, D.; Yang, Y.; Kong, S.-y.; Hua, N.; Limtiaco, N.; John, R. S.; Constant, N.; Guajardo-Cespedes, M.; Yuan, S.; Tar, ´ C.; et al. 2018. Universal sentence encoder.
  5. 5.Chen, J.; Jordan, M. I.; and Wainwright, M. J. 2020. HopSkipJumpAttack: A Query-Efficient Decision-Based Attack. In S&P, 1277–1294. IEEE.
  6. 6.Chen, P.; Zhang, H.; Sharma, Y.; Yi, J.; and Hsieh, C. 2017. ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models. In AISec@CCS, 15–26. ACM.
  7. 7.Cheng, M.; Le, T.; Chen, P.; Zhang, H.; Yi, J.; and Hsieh, C. 2019. Query-Efficient Hard-label Black-box Attack: An Optimization-based Approach. In ICLR. OpenReview.net.
  8. 8.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT, 4171–4186. Association for Computational Linguistics.
  9. 9.Ebrahimi, J.; Rao, A.; Lowd, D.; and Dou, D. 2018. HotFlip: White-Box Adversarial Examples for Text Classification. In ACL, 31–36. Association for Computational Linguistics.
  10. 10.Gao, J.; Lanchantin, J.; Soffa, M. L.; and Qi, Y. 2018. Black-Box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers. In S&P Workshops, 50–56. IEEE Computer Society.
  11. 11.Goodfellow, I. J.; Shlens, J.; and Szegedy, C. 2015. Explaining and Harnessing Adversarial Examples. In ICLR.
  12. 12.Gorban, A. N.; and Tyukin, I. Y. 2018. Blessing of dimensionality: mathematical foundations of the statistical physics of data. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 376(2118): 20170237.
  13. 13.Hochreiter, S.; and Schmidhuber, J. 1997. Long Short-Term Memory. Neural Comput., 9(8): 1735–1780.
  14. 14.Jia, R.; and Liang, P. 2017. Adversarial Examples for Evaluating Reading Comprehension Systems. In EMNLP, 2021–2031. Association for Computational Linguistics.
  15. 15.Jin, D.; Jin, Z.; Zhou, J. T.; and Szolovits, P. 2020. Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment. In AAAI, 8018–8025. AAAI Press.
  16. 16.Kim, Y. 2014. Convolutional Neural Networks for Sentence Classification. In EMNLP, 1746–1751. ACL.
  17. 17.Kurakin, A.; Goodfellow, I. J.; and Bengio, S. 2017. Adversarial examples in the physical world. In ICLR. OpenReview.net.
  18. 18.Li, J.; Ji, S.; Du, T.; Li, B.; and Wang, T. 2019. TextBugger: Generating Adversarial Text Against Real-world Applications. In NDSS. The Internet Society.
  19. 19.Maas, A. L.; Daly, R. E.; Pham, P. T.; Huang, D.; Ng, A. Y.; and Potts, C. 2011. Learning Word Vectors for Sentiment Analysis. In ACL, 142–150. The Association for Computational Linguistics.
  20. 20.Maheshwary, R.; Maheshwary, S.; and Pudi, V. 2021. Generating Natural Language Attacks in a Hard Label Black Box Setting. In AAAI, 13525–13533. AAAI Press.
  21. 21.Mrksic, N.; Seaghdha, D. ´ O.; Thomson, B.; Gasic, M.; ´ Rojas-Barahona, L. M.; Su, P.; Vandyke, D.; Wen, T.; and Young, S. J. 2016. Counter-fitting Word Vectors to Linguistic Constraints. In Knight, K.; Nenkova, A.; and Rambow, O., eds., NAACL-HLT, 142–148. The Association for Computational Linguistics.
  22. 22.Pang, B.; and Lee, L. 2005. Seeing Stars: Exploiting Class Relationships for Sentiment Categorization with Respect to Rating Scales. In ACL, 115–124. The Association for Computational Linguistics.
  23. 23.Pennington, J.; Socher, R.; and Manning, C. D. 2014. Glove: Global Vectors for Word Representation. In EMNLP, 1532–1543. ACL.
  24. 24.Ren, S.; Deng, Y.; He, K.; and Che, W. 2019. Generating Natural Language Adversarial Examples through Probability Weighted Word Saliency. In ACL, 1085–1097. Association for Computational Linguistics.
  25. 25.Williams, A.; Nangia, N.; and Bowman, S. R. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In NAACL-HLT, 1112–1122. Association for Computational Linguistics.
  26. 26.Zhang, X.; Zhao, J. J.; and LeCun, Y. 2015. Character-level Convolutional Networks for Text Classification. In NeurIPS, 649–657.
  27. 27.Zhou, Y.; Wang, H.; He, J.; and Wang, H. 2021a. From Intrinsic to Counterfactual: On the Explainability of Contextualized Recommender Systems. arXiv:2110.14844.
  28. 28.Zhou, Y.; Wu, J.; Wang, H.; and He, J. 2021b. Adversarial Robustness through Bias Variance Decomposition: A New Perspective for Federated Learning. arXiv:2009.09026.
  29. 29.Zhou, Y.; Xu, J.; Wu, J.; Nasrabadi, Z. T.; Korpeoglu, E.; ¨ Achan, K.; and He, J. 2021c. PURE: Positive-Unlabeled Recommendation with Generative Adversarial Network. In KDD, 2409–2419. ACM.
  30. 30.Zhu, C.; Cheng, Y.; Gan, Z.; Sun, S.; Goldstein, T.; and Liu, J. 2020. FreeLB: Enhanced Adversarial Training for Natural Language Understanding. In ICLR. OpenReview.net.

Citation

MLA
Ye, M., et al. “TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 4, 2022, pp. 3877–84, https://doi.org/10.1609/AAAI.V36I4.20303.
APA
Ye, M., Miao, C., Wang, T., & Ma, F. (2022). TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text. Proceedings of the AAAI Conference on Artificial Intelligence, 36(4), 3877–3884. https://doi.org/10.1609/AAAI.V36I4.20303
Chicago
Ye, M., C. Miao, T. Wang, and F. Ma. 2022. “TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (4): 3877–84. https://doi.org/10.1609/AAAI.V36I4.20303.
Harvard
Ye, M. et al. (2022) “TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(4), pp. 3877–3884. Available at: https://doi.org/10.1609/AAAI.V36I4.20303.
Vancouver
1. Ye M, Miao C, Wang T, Ma F (2022) TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text. Proceedings of the AAAI Conference on Artificial Intelligence 36:3877–3884

BibTeX

@article{Ye_2022, title={TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I4.20303}, DOI={10.1609/aaai.v36i4.20303}, number={4}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Ye, Muchao and Miao, Chenglin and Wang, Ting and Ma, Fenglong}, year={2022}, month=June, pages={3877–3884} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF