TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text
Muchao YeChenglin MiaoTing WangFenglong Ma
Proposes TextHoaxer, a gradient-based framework that formulates hard-label text adversarial attacks in continuous word embedding space to generate high-similarity adversarial examples under strict query budgets without relying on query-expensive genetic algorithms.
Deep neural networks are widely deployed across natural language processing systems, yet their vulnerability to adversarial manipulation poses significant security risks. In commercial environments, these models are typically accessible only through restricted application programming interfaces that provide final discrete decisions rather than detailed confidence probabilities, while also enforcing strict query rate limits to prevent exploitation. Existing evaluation techniques rely on genetic algorithms that require thousands of repeated queries across large candidate pools, causing them to fail when query budgets are limited. The article addresses this operational gap by introducing TextHoaxer, a framework designed to generate high-quality adversarial text samples efficiently under tight query limits in decision-only settings.
The researchers formulated the attack process as a gradient-based optimization task within a continuous word embedding space, bypassing the inefficiency of combinatorial discrete search. Instead of managing a large population of candidates, TextHoaxer refines a single candidate text using an objective function that balances overall semantic meaning, individual word-level distances, and the total number of altered words. The authors evaluated the approach across eight benchmark datasets covering text classification and natural language inference against three widely used model architectures (BERT, WordCNN, and WordLSTM) under a tight limit of 1,000 queries, supported by human evaluation panels.
The findings show that TextHoaxer consistently outperformed existing methods by generating adversarial examples with higher semantic preservation and fewer modified words. On benchmark sentiment classification tasks, TextHoaxer achieved semantic similarity scores 3.7 to 4.8 percentage points higher than the strongest competing hard-label methods, while decreasing the required word perturbation rate by roughly 2.0 to 2.6 percentage points. Testing across varying budgets from 100 to 1,000 queries demonstrated that TextHoaxer reliably produced superior text quality at all constraint levels. Independent human reviewers also ranked its output among the most semantically faithful to original text.
These results demonstrate that standard API query throttling and decision-only output restrictions do not provide sufficient defense against sophisticated, low-query adversarial attacks. Organizations deploying language models should not rely solely on query rate limiting for protection, but should instead incorporate budgeted adversarial evaluation into model robustness auditing and defense pipelines. While the study provides strong confidence across standard benchmarks, it focused primarily on word-level synonym replacements within classification tasks; future work should explore broader multi-word perturbations, defensive hardening strategies, and defenses across larger modern generative architectures.
- Paper: Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment, Di Jin et al. (2019). Introduces TextFooler, a seminal black-box text adversarial attack framework that establishes the core concepts of synonym substitution and semantic preservation used in TextHoaxer.
- Paper: Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models, Wieland Brendel et al. (2017). Establishes the fundamental decision-based (hard-label) attack paradigm that TextHoaxer adapts and optimizes for discrete natural language under tight query budgets.
- Paper: Black-box Adversarial Attacks with Limited Queries and Information, Andrew Ilyas et al. (2018). Formulates the query-budgeted and label-only threat model for adversarial attacks that motivates TextHoaxer's focus on query efficiency.
- Paper: Practical Black-Box Attacks against Machine Learning, Nicolas Papernot et al. (2017). Provides foundational methodology for black-box adversarial generation against machine learning classifiers with restricted API query access.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). Advances query-efficient black-box adversarial attacks to modern large language models, achieving successful jailbreaks under strict query constraints.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Extends gradient-guided discrete text optimization to generate universal adversarial sequences that transfer across language model architectures.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). Improves the optimization efficiency and search dynamics of token-level gradient-based adversarial attacks on language models.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). Explores how genetic algorithms can be refined with language model fluency objectives to generate stealthy, semantically coherent adversarial text.
