Boundary Point Jailbreaking of Black-Box LLMs
Xander DaviesGiorgi GiglemianiEdmund LauEric WinsorGeoffrey IrvingYarin Gal
Develops an automated black-box jailbreak method that circumvents frontier language model defenses using only single-bit binary feedback per query, presenting the first successful automated attack against Constitutional Classifiers without relying on human seeds.
Frontier artificial intelligence models increasingly rely on auxiliary safety classifiers to block harmful prompts and prevent policy violations. While these classifier-based safeguards have resisted extensive red-teaming and standard automated jailbreaks, evaluating and ensuring their true resilience remains a critical challenge. Attackers face an environment with minimal feedback, often receiving only a binary decision indicating whether a prompt was flagged or permitted. Without internal probability scores or gradient access, finding vulnerabilities in these systems has historically required specialized human expertise, potentially creating a false sense of security regarding classifier robustness.
The article introduces and evaluates Boundary Point Jailbreaking (BPJ), a fully automated, black-box optimization method designed to bypass leading safety classifiers. The primary objective is to demonstrate that an automated algorithm using only single-bit feedback can discover universal adversarial prefixes capable of evading industry-grade safeguards.
To overcome the absence of informative feedback from robust classifiers, the approach combines curriculum learning with active boundary point selection. The algorithm constructs an intermediate ladder of targets by corrupting harmful prompts with varying levels of noise, gradually progressing toward the uncorrupted text. Within each curriculum stage, an evolutionary algorithm mutates candidate text strings and evaluates them against selected boundary points—evaluation inputs where candidate attacks disagree, maximizing the signal per query. The authors evaluated BPJ against Anthropic's Constitutional Classifiers guarding Claude Sonnet 4.5, OpenAI's GPT-5 input classifier, and a benchmark dataset using a prompted GPT-4.1-nano baseline.
The study demonstrates several key findings:
- Evading Production Safeguards: BPJ represents the first automated black-box attack to successfully discover universal jailbreaks against Constitutional Classifiers and GPT-5's input classifier without relying on human-curated seeds.
- Substantial Elicitation Increases: On unseen, severe biological misuse benchmarks, BPJ increased the average rubric score of elicited responses from 0% to 25.5% (68.0% with basic elicitation) on Constitutional Classifiers and from 0% to 75.6% on GPT-5's input monitor.
- Extreme Query Efficiency Over Baselines: BPJ converged approximately 5 times faster than curriculum learning alone and succeeded where random baseline methods failed completely, generating successful attacks in roughly 660,000 to 800,000 queries (330 in API costs).
- Strong Cross-Domain Transferability: Adversarial prefixes optimized on a single harmful query generalized effectively to broad sets of unseen questions across diverse policy categories.
These findings indicate that single-interaction, text-based classification safeguards are fundamentally vulnerable to automated, decision-based optimization even under minimal feedback constraints. Because optimized adversarial prefixes transfer across arbitrary queries, relying solely on input-output filters at the individual request level presents significant risk to model safety, compliance, and deployment pipelines.
To mitigate these risks, organizations should transition from isolated prompt-level defenses to layered, batch-level telemetry and monitoring. Since BPJ generates thousands of flagged queries during optimization, monitoring request patterns, flagging frequencies, and behavioral anomalies over time can detect attacks prior to successful evasion. Additionally, defenders should explore probe-based internal classifiers, ensemble model randomization, and adversarial training using automated jailbreak artifacts.
The analysis is subject to certain boundary conditions. The attack was conducted in an environment without automated account bans, whereas real-world rate limiting and policy enforcement could disrupt optimization budgets. Furthermore, BPJ was evaluated primarily against deterministic text classifiers rather than stochastic defenses and was paired with human-found jailbreaks to bypass internal model refusals. Nonetheless, the high consistency and cross-domain transfer demonstrate that single-interaction classifiers alone cannot guarantee model safety against automated black-box search.
- Paper: Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models, Wieland Brendel et al. (2017). The Boundary Attack establishes how adversarial search can operate using only final binary decisions, the central black-box constraint BPJ adapts to text safety classifiers.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Its universal adversarial suffixes provide the direct precedent for BPJ’s goal of optimizing a reusable jailbreak prefix that transfers across harmful requests.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). PAIR shows how automated black-box feedback can iteratively refine jailbreak prompts, making BPJ’s alternative search strategy and efficiency claims easier to assess.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). AutoDAN supplies an earlier evolutionary prompt-optimization approach, clarifying BPJ’s shift from seeded attacks on models to seed-free search against classifiers.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This analysis of jailbreak failure modes gives context for BPJ’s evaluation of how adversarial prompts exploit weaknesses in LLM safety mechanisms.
No sufficiently relevant recommendations were found.
