Boundary Point Jailbreaking of Black-Box LLMs

Xander DaviesGiorgi GiglemianiEdmund LauEric WinsorGeoffrey IrvingYarin Gal

article2026arXiv9 citations

Develops an automated black-box jailbreak method that circumvents frontier language model defenses using only single-bit binary feedback per query, presenting the first successful automated attack against Constitutional Classifiers without relying on human seeds.

Listen

Frontier artificial intelligence models increasingly rely on auxiliary safety classifiers to block harmful prompts and prevent policy violations. While these classifier-based safeguards have resisted extensive red-teaming and standard automated jailbreaks, evaluating and ensuring their true resilience remains a critical challenge. Attackers face an environment with minimal feedback, often receiving only a binary decision indicating whether a prompt was flagged or permitted. Without internal probability scores or gradient access, finding vulnerabilities in these systems has historically required specialized human expertise, potentially creating a false sense of security regarding classifier robustness.

The article introduces and evaluates Boundary Point Jailbreaking (BPJ), a fully automated, black-box optimization method designed to bypass leading safety classifiers. The primary objective is to demonstrate that an automated algorithm using only single-bit feedback can discover universal adversarial prefixes capable of evading industry-grade safeguards.

To overcome the absence of informative feedback from robust classifiers, the approach combines curriculum learning with active boundary point selection. The algorithm constructs an intermediate ladder of targets by corrupting harmful prompts with varying levels of noise, gradually progressing toward the uncorrupted text. Within each curriculum stage, an evolutionary algorithm mutates candidate text strings and evaluates them against selected boundary points—evaluation inputs where candidate attacks disagree, maximizing the signal per query. The authors evaluated BPJ against Anthropic's Constitutional Classifiers guarding Claude Sonnet 4.5, OpenAI's GPT-5 input classifier, and a benchmark dataset using a prompted GPT-4.1-nano baseline.

The study demonstrates several key findings:

  1. Evading Production Safeguards: BPJ represents the first automated black-box attack to successfully discover universal jailbreaks against Constitutional Classifiers and GPT-5's input classifier without relying on human-curated seeds.
  2. Substantial Elicitation Increases: On unseen, severe biological misuse benchmarks, BPJ increased the average rubric score of elicited responses from 0% to 25.5% (68.0% with basic elicitation) on Constitutional Classifiers and from 0% to 75.6% on GPT-5's input monitor.
  3. Extreme Query Efficiency Over Baselines: BPJ converged approximately 5 times faster than curriculum learning alone and succeeded where random baseline methods failed completely, generating successful attacks in roughly 660,000 to 800,000 queries (210to210 to 330 in API costs).
  4. Strong Cross-Domain Transferability: Adversarial prefixes optimized on a single harmful query generalized effectively to broad sets of unseen questions across diverse policy categories.

These findings indicate that single-interaction, text-based classification safeguards are fundamentally vulnerable to automated, decision-based optimization even under minimal feedback constraints. Because optimized adversarial prefixes transfer across arbitrary queries, relying solely on input-output filters at the individual request level presents significant risk to model safety, compliance, and deployment pipelines.

To mitigate these risks, organizations should transition from isolated prompt-level defenses to layered, batch-level telemetry and monitoring. Since BPJ generates thousands of flagged queries during optimization, monitoring request patterns, flagging frequencies, and behavioral anomalies over time can detect attacks prior to successful evasion. Additionally, defenders should explore probe-based internal classifiers, ensemble model randomization, and adversarial training using automated jailbreak artifacts.

The analysis is subject to certain boundary conditions. The attack was conducted in an environment without automated account bans, whereas real-world rate limiting and policy enforcement could disrupt optimization budgets. Furthermore, BPJ was evaluated primarily against deterministic text classifiers rather than stochastic defenses and was paired with human-found jailbreaks to bypass internal model refusals. Nonetheless, the high consistency and cross-domain transfer demonstrate that single-interaction classifiers alone cannot guarantee model safety against automated black-box search.

arXiv: 2602.15001

No sufficiently relevant recommendations were found.

Cover for Boundary Point Jailbreaking of Black-Box LLMs

Abstract

Frontier LLMs are safeguarded against attempts to extract harmful information via adversarial prompts known as "jailbreaks". Recently, defenders have developed classifier-based systems that have survived thousands of hours of human red teaming. We introduce Boundary Point Jailbreaking (BPJ), a new class of automated jailbreak attacks that evade the strongest industry-deployed safeguards. Unlike previous attacks that rely on white/grey-box assumptions (such as classifier scores or gradients) or libraries of existing jailbreaks, BPJ is fully black-box and uses only a single bit of information per query: whether or not the classifier flags the interaction. To achieve this, BPJ addresses the core difficulty in optimising attacks against robust real-world defences: evaluating whether a proposed modification to an attack is an improvement. Instead of directly trying to learn an attack for a target harmful string, BPJ converts the string into a curriculum of intermediate attack targets and then actively selects evaluation points that best detect small changes in attack strength ("boundary points"). We believe BPJ is the first fully automated attack algorithm that succeeds in developing universal jailbreaks against Constitutional Classifiers, as well as the first automated attack algorithm that succeeds against GPT-5's input classifier without relying on human attack seeds. BPJ is difficult to defend against in individual interactions but incurs many flags during optimisation, suggesting that effective defence requires supplementing single-interaction methods with batch-level monitoring.

Table of Contents

  • 1 Introduction
  • 2 Boundary Point Jailbreaking
  • 2.1 Generating Within-Level Boundary Points
  • 2.2 Improving Attack Strings
  • 2.3 Replacing Solved Boundary Points and Iterating
  • 2.4 Pseudocode
  • 3 Experiments
  • 3.1 Attack Settings
  • 3.2 Results
  • 4 Theoretical Formulation
  • 5 Related Work
  • 6 Discussion
  • References
  • A Theoretical Formulation and Analysis of BPJ
  • A.1 Problem Formulation
  • A.2 Formalisation of evolution dynamics
  • A.3 What drives progress?
  • A.4 Lifting via continuation method
  • A.4.1 Within-level Convergence
  • A.4.2 Continuation
  • A.4.3 Conditions for success: alignment between fqf_{q} and the base objective
  • A.5 Summary
  • A.6 Surrogate Fitness Function and Boundary Points
  • A.6.1 BP Sampling Induces a Biased but Rank-preserving Surrogate Fitness Function
  • A.6.2 Advantage of BP over i.i.d. sampling
  • A.7 Theory limitations
  • B Target Questions for GPT-4.1-nano Classifier
  • C GPT-4.1-Nano Prompt
  • D Additional Plots
  • E Single-Interaction Defences
  • F Pseudocode

Knowls

  1. Knowl 1 — Black-box threat model and universal-prefix objective

    definition

    Boundary Point Jailbreaking (BPJ) assumes an attacker can query a deterministic classifier-guarded language model and observe only a binary result: flagged or not flagged. Write C(s)=0C(s)=0 when input string ss is flagged and C(s)=1C(s)=1 when it passes. For a harmful target string xx, the attacker seeks a prefix aa such that C(ax)=1C(ax)=1; a prefix optimized once and then prepended to different harmful targets is a universal adversarial prefix. The experiments attack the safeguards’ classifiers, rather than independently jailbreaking the underlying language model.

  2. Knowl 2 — BPJ combines a difficulty curriculum with disagreement-based evaluation

    model/method

    BPJ addresses sparse binary feedback by turning a difficult target into progressively harder intermediate targets and evaluating candidate prefixes on inputs where the current candidates disagree about whether the classifier will flag. A population-based search uses those evaluations to favor better-performing prefixes. When an evaluation input becomes easy for every current candidate, it is replaced; as performance improves at the current difficulty, optimization moves toward less-corrupted targets. The central design is to make progress measurable before the final harmful target itself can be passed.

  3. Knowl 3 — Noise interpolation defines the relaxed optimization objective

    equation

    Let xx be the harmful target string, aa a candidate prefix, CC a binary classifier whose output 11 means the input passes, and Nq,xN_{q,x} a distribution over corrupted versions of xx, indexed by a noise level q∈[0,1]q\in[0,1]. BPJ scores a prefix by its expected pass rate on that distribution: fq(a)=Ex′∼Nq,x[C(ax′)]f_q(a)=\mathbb{E}_{x'\sim N_{q,x}}[C(ax')]. At q=0q=0, the distribution is concentrated on xx, so f0(a)=C(ax)f_0(a)=C(ax) recovers the original binary objective. Larger noise levels are intended to produce easier intermediate targets; in the paper’s experiments, character-level noise interpolation gradually changes the probability that a target is flagged rather than imposing a sharp easy-versus-hard cutoff.

  4. Knowl 4 — Boundary points isolate classifier inputs that distinguish candidates

    definition

    For a finite set AA of candidate prefixes, a boundary point is a target variant bb for which at least one candidate passes and at least one candidate is flagged: there are a,a′∈Aa,a'\in A such that C(ab)=1C(ab)=1 and C(a′b)=0C(a'b)=0. Inputs on which every candidate receives the same binary outcome do not distinguish the candidates. Under hard rank-based selection with a fixed candidate set, such non-discriminating inputs cannot change their fitness ranking; evaluating on boundary points therefore concentrates queries on inputs that can affect selection.

  5. Knowl 5 — BPJ succeeds against two deployed safeguard systems

    empirical result

    On harmful biological-information rubrics for unseen questions, BPJ increased the mean score on non-empty responses from 0% to 25.5% against Anthropic’s Constitutional Classifiers (CC), with a best score of 39.9% across 50 attempts. With basic elicitation, the CC mean was 68% and the best score was 80.4%. Against GPT-5’s input classifier, the mean score rose from 0% to 75.6%, with a best score of 94.3% across 50 attempts. The CC attack development used approximately 660,000 queries and 330;theGPT−5input−classifierattackusedapproximately800,000queriesand330; the GPT-5 input-classifier attack used approximately 800,000 queries and 210. In both settings, a simple human-found jailbreak was used to get past the underlying model’s refusal behavior, so the reported automation concerns classifier evasion rather than the entire end-to-end jailbreak.

  6. Knowl 6 — Boundary-point selection improves optimization over two baselines

    empirical result

    In experiments against a prompted GPT-4.1-nano classifier, BPJ was compared with random-prefix Best-of-N search and a curriculum-only variant that used randomly selected evaluation inputs rather than boundary points. BPJ converged on average about five times faster than the curriculum-only variant; the paper reports speedups of 3–8 times across different target questions. Both curriculum-based methods substantially outperformed Best-of-N, which did not converge within the reported experiment budget. Individual BPJ runs for the illustrated target finished after roughly 25,000–30,000 queries.

  7. Knowl 7 — Prefixes optimized on one question transfer to unseen questions

    empirical result

    When BPJ optimized a prefix against a single HarmBench question and periodically evaluated it on the wider HarmBench set, transfer success generally increased during optimization for both pass@1 and pass@10. The transfer included questions from topic areas beyond the optimization question’s original category. The authors report similar generalization to unseen biological-misuse questions for CC and GPT-5’s input classifier. Across attacks trained on different target questions, they also observed a positive association between the number of queries needed to solve the training target and transfer to unseen questions; this is an empirical correlation, not evidence that greater target difficulty causes transfer.

  8. Knowl 8 — Evolutionary selection needs variation in candidate fitness

    theoretical result

    For a candidate-prefix population distributed as rr, let f(a)f(a) be candidate fitness and let w(a)≥0w(a)\geq 0 be a selection weight with 0<Er[w]<∞0<\mathbb{E}_{r}[w]<\infty. Reweighting the population gives p′(a)=r(a)w(a)/Er[w]p'(a)=r(a)w(a)/\mathbb{E}_{r}[w]. The change in mean fitness due to selection is Ep′[f]−Er[f]=Cov⁡r(f,w)/Er[w]\mathbb{E}_{p'}[f]-\mathbb{E}_{r}[f]=\operatorname{Cov}_{r}(f,w)/\mathbb{E}_{r}[w]. For quantile selection, where the mean selection weight is α\alpha, this gain is Cov⁡r(f,w)/α\operatorname{Cov}_{r}(f,w)/\alpha. Thus selection can make little progress when candidate fitness has little variation. The paper uses this result to explain why direct optimization of the sparse, noiseless binary objective can behave much like random drift, while intermediate noise levels can provide more graded fitness differences.

  9. Knowl 9 — Successful continuation requires alignment and trackable equilibria

    theoretical result

    The paper’s stylized analysis gives sufficient conditions—not an end-to-end guarantee—for a noise curriculum to retain probability mass on prefixes that pass the original target. At a given noise level, selection by the relaxed fitness must favor truly successful prefixes over unsuccessful ones: under the relaxed-fitness ranking, the true-positive rate for membership in the base success set must exceed the false-positive rate. Mutation must also avoid removing successful-prefix mass too quickly. Across noise levels, the population must be able to track a locally stable, smoothly varying equilibrium, which requires sufficiently small noise-reduction steps and no intervening singularity or abrupt change in the equilibrium branch. The authors note that relaxed objectives can instead favor candidates exploiting artifacts of corrupted inputs, so the alignment condition is not guaranteed for a real classifier.

  10. Knowl 10 — Theoretical query-efficiency claims have important scope limits

    theoretical result

    For a fixed candidate population and hard quantile selection, the paper shows that evaluation inputs on which all candidates agree add the same contribution to every candidate’s estimated fitness and therefore do not affect the ranking used for selection. Conditioning evaluation on boundary points can consequently avoid queries that provide no ranking signal, under those assumptions. A separate variance result concerns comparing two candidates on a shared set of evaluation inputs: if their classifier outcomes on an input are nonnegatively correlated, the shared-set estimate of their fitness difference has lower variance than estimates formed from independent sets. These claims do not establish query-complexity guarantees for BPJ’s adaptive, finite-population implementation.

  11. Knowl 11 — Operational and theoretical limitations, and the need for batch monitoring

    limitation

    The reported real-world attacks required hundreds of thousands of classifier queries and generated many flags; the authors warn that realistic systems may ban accounts and argue that effective defense should add batch-level monitoring rather than rely only on single-interaction safeguards. BPJ was not used to find the underlying model jailbreaks, and the experiments did not cover highly stochastic classifiers. The theoretical analysis is diagnostic: it uses an infinite-population approximation and often smooth selection, whereas the implemented search uses finite populations, hard selection, and finite or adaptive evaluations. The paper therefore does not prove that the stated alignment or continuation conditions hold in deployed systems. Preliminary defensive observations reported by the authors suggest that probe-based classifiers may resist better than text-based classifiers, and that training on BPJ examples or randomly serving one of two similarly trained classifiers may help; these observations are not presented as established guarantees.

Coverage note — Detailed executable pseudocode, attack strings, and implementation refinements are deliberately omitted because they would make the safeguard-evasion procedure more directly reproducible; supplementary per-question plots are also omitted because they add detail beyond the main reported comparisons and findings.

References

  1. 1.Allgower, E. L. and Georg, K. Numerical Continuation Methods: An Introduction. Springer Series in Computational Mathematics. Springer, Berlin, Germany, October 2011.
  2. 2.Andriushchenko, M., Croce, F., and Flammarion, N. Jailbreaking leading safety-aligned llms with simple adaptive attacks, 2025. URL https://arxiv.org/abs/2404.02151.
  3. 3.Anthropic. Expanding our model safety bug bounty program. https://www.anthropic.com/news/model-safety-bug-bounty, August 2024. Accessed: 2026-01-27.
  4. 4.Anthropic. Strengthening our safeguards through collaboration with us caisi and uk aisi, September 2025. URL https://www.anthropic.com/news/strengthening-our-safeguards-through-collaboration-with-us-caisi-and-uk-aisi. News post, Sep 12, 2025. Accessed 2025-10-11.
  5. 5.Anthropic Safeguards Research Team. Constitutional classifiers: Defending against universal jailbreaks. https://www.anthropic.com/research/constitutional-classifiers, February 2025. Accessed: 2025-11-13.
  6. 6.Ben-Tov, M., Geva, M., and Sharif, M. Universal jailbreak suffixes are strong attention hijackers, 2025. URL https://arxiv.org/abs/2506.12880.
  7. 7.Bengio, Y., Louradour, J., Collobert, R., and Weston, J. Curriculum learning. In Proceedings of the 26th Annual International Conference on Machine Learning, New York, NY, USA, June 2009. ACM.
  8. 8.Brendel, W., Rauber, J., and Bethge, M. Decision-based adversarial attacks: Reliable attacks against black-box machine learning models, 2018. URL https://arxiv.org/abs/1712.04248.
  9. 9.Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. Jailbreaking black box large language models in twenty queries, 2024. URL https://arxiv.org/abs/2310.08419.
  10. 10.Chen, J., Jordan, M. I., and Wainwright, M. J. Hop-skipjumpattack: A query-efficient decision-based attack, 2020. URL https://arxiv.org/abs/1904.02144.
  11. 11.Cheng, M., Le, T., Chen, P.-Y., Yi, J., Zhang, H., and Hsieh, C.-J. Query-efficient hard-label black-box attack:an optimization-based approach, 2018. URL https://arxiv.org/abs/1807.04457.
  12. 12.Chowdhury, N., Schwettmann, S., and Steinhardt, J. Automatically jailbreaking frontier language models with investigator agents. https://transluce.org/jailbreaking-frontier-models, September 2025.
  13. 13.Cunningham, H., Wei, J., Wang, Z., Persic, A., Peng, A., Abderrachid, J., Agarwal, R., Chen, B., Cohen, A., Dau, A., Dimitriev, A., Gilson, R., Howard, L., Hua, Y., Kaplan, J., Leike, J., Lin, M., Liu, C., Mikulik, V., Mittapalli, A., O’Hara, C., Pan, J., Saxena, N., Silverstein, A., Song, Y., Yu, X., Zhou, G., Perez, E., and Sharma, M. Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks, 2026. URL https://arxiv.org/abs/2601.04603.
  14. 14.David C. Prompt injection is not sql injection (it may be worse), 2025. URL https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection. NCSC Technical Director for Platforms Research. Accessed: 2026-01-28.
  15. 15.Glasserman, P. and Yao, D. D. Some guidelines and guarantees for common random numbers. Manage. Sci., 38(6): 884–908, June 1992.
  16. 16.Hayase, J., Borevkovic, E., Carlini, N., Tramèr, F., and Nasr, M. Query-based adversarial prompt generation, 2024. URL https://arxiv.org/abs/2402.12329.
  17. 17.Holland, J. H. Adaptation in natural and artificial systems: An introductory analysis with applications to biology, control, and artificial intelligence. Complex Adaptive Systems. Bradford Books, Cambridge, MA, June 2019.
  18. 18.Huang, D., Shah, A., Araujo, A., Wagner, D., and Sitawarin, C. Stronger universal and transferable attacks by suppressing refusals. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 5850–5876, 2025.
  19. 19.Hughes, J., Price, S., Lynch, A., Schaeffer, R., Barez, F., Koyejo, S., Sleight, H., Jones, E., Perez, E., and Sharma, M. Best-of-n jailbreaking, 2024. URL https://arxiv.org/abs/2412.03556.
  20. 20.Krauskopf, B., Osinga, H. M., and Galan-Vioque, J. (eds.). Numerical Continuation Methods for Dynamical Systems: Path following and boundary value problems. Understanding Complex Systems. Springer, New York, NY, July 2007.
  21. 21.Liu, H., Xu, Z., Zhang, X., Zhang, F., Ma, F., Chen, H., Yu, H., and Zhang, X. Hqa-attack: Toward high quality black-box hard-label adversarial attack on text, 2024. URL https://arxiv.org/abs/2402.01806.
  22. 22.Liu, Y., Chen, X., Liu, C., and Song, D. Delving into transferable adversarial examples and black-box attacks, 2017. URL https://arxiv.org/abs/1611.02770.
  23. 23.Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024. URL https://arxiv.org/abs/2402.04249.
  24. 24.McKenzie, I. R., Hollinsworth, O. J., Tseng, T., Davies, X., Casper, S., Tucker, A. D., Kirk, R., and Gleave, A. Stack: Adversarial attacks on llm safeguard pipelines, 2025. URL https://arxiv.org/abs/2506.24068.
  25. 25.Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A. Tree of attacks: Jailbreaking black-box llms automatically, 2024. URL https://arxiv.org/abs/2312.02119.
  26. 26.Nasr, M., Carlini, N., Sitawarin, C., Schulhoff, S. V., Hayes, J., Ilie, M., Pluto, J., Song, S., Chaudhari, H., Shumailov, I., Thakurta, A., Xiao, K. Y., Terzis, A., and Tramèr, F. The attacker moves second: Stronger adaptive attacks bypass defenses against llm jailbreaks and prompt injections, 2025. URL https://arxiv.org/abs/2510.09023.
  27. 27.National Cyber Security Centre (NCSC). From bugs to bypasses: Adapting vulnerability disclosure for ai safeguards. Blog post (with the UK AI Safety Institute), September 2025. URL https://www.ncsc.gov.uk/blog-post/from-bugs-to-bypasses-adapting-vulnerability-disclosure-for-ai-safeguards. Accessed: 2026-01-23.
  28. 28.OpenAI. Working with us caisi and uk aisi to build more secure ai systems. https://openai.com/index/us-caisi-uk-aisi-ai-update/, September 2025. Accessed: YYYY-MM-DD.
  29. 29.OpenAI. Introducing gpt-4.1 in the api. https://openai.com/index/gpt-4-1/, April 2025a. Accessed: 2026-01-23.
  30. 30.OpenAI. Gpt-5 system card. System card, OpenAI, August 2025b. URL https://cdn.openai.com/gpt-5-system-card.pdf. Accessed: 2026-01-23.
  31. 31.Price, G. R. Selection and covariance. Nature, 227(5257): 520–521, August 1970.
  32. 32.Sadasivan, V. S., Saha, S., Sriramanan, G., Kattakinda, P., Chegini, A., and Feizi, S. Fast adversarial attacks on language models in one gpu minute, 2024. URL https://arxiv.org/abs/2402.15570.
  33. 33.Settles, B. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison, 2009.
  34. 34.Seung, H. S., Opper, M., and Sompolinsky, H. Query by committee. In Proceedings of the fifth annual workshop on Computational learning theory, New York, NY, USA, July 1992. ACM.
  35. 35.Sharma, M., Tong, M., Mu, J., Wei, J., Kruthoff, J., Goodfriend, S., Ong, E., Peng, A., Agarwal, R., Anil, C., Askell, A., Bailey, N., Benton, J., Bluemke, E., Bowman, S. R., Christiansen, E., Cunningham, H., Dau, A., Gopal, A., Gilson, R., Graham, L., Howard, L., Kalra, N., Lee, T., Lin, K., Lofgren, P., Mosconi, F., O’Hara, C., Olsson, C., Petrini, L., Rajani, S., Saxena, N., Silverstein, A., Singh, T., Sumers, T., Tang, L., Troy, K. K., Weisser, C., Zhong, R., Zhou, G., Leike, J., Kaplan, J., and Perez, E. Constitutional classifiers: Defending against universal jailbreaks across thousands of hours of red teaming, 2025. URL https://arxiv.org/abs/2501.18837.
  36. 36.Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2024. URL https://arxiv.org/abs/2308.03825.
  37. 37.Ye, M., Chen, J., Miao, C., Wang, T., and Ma, F. Leapattack: Hard-label adversarial attack on text via gradient-based optimization. In KDD 2022 - Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 2307–2315. Association for Computing Machinery, August 2022. doi: 10.1145/3534678.3539357. Publisher Copyright: © 2022 ACM.; 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2022 ; Conference date: 14-08-2022 Through 18-08-2022.
  38. 38.Zhang, J., Ding, M., Liu, Y., Hong, J., and Tramèr, F. Black-box optimization of llm outputs by asking for directions, 2025. URL https://arxiv.org/abs/2510.16794.
  39. 39.Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models, 2023. URL https://arxiv.org/abs/2307.15043.
  40. 40.Zou, A., Lin, M., Jones, E., Nowak, M., Dziemian, M., Winter, N., Grattan, A., Nathanael, V., Croft, A., Davies, X., Patel, J., Kirk, R., Burnikell, N., Gal, Y., Hendrycks, D., Kolter, J. Z., and Fredrikson, M. Security challenges in ai agent deployment: Insights from a large scale public competition, 2025. URL https://arxiv.org/abs/2507.20526.

Citation

MLA
Davies, X., et al. “Boundary Point Jailbreaking of Black-Box LLMs”. arXiv, 2026, http://arxiv.org/abs/2602.15001v2.
APA
Davies, X., Giglemiani, G., Lau, E., Winsor, E., Irving, G., & Gal, Y. (2026). Boundary Point Jailbreaking of Black-Box LLMs. arXiv. http://arxiv.org/abs/2602.15001v2
Chicago
Davies, X., G. Giglemiani, E. Lau, E. Winsor, G. Irving, and Y. Gal. 2026. “Boundary Point Jailbreaking of Black-Box LLMs”. arXiv. http://arxiv.org/abs/2602.15001v2.
Harvard
Davies, X. et al. (2026) “Boundary Point Jailbreaking of Black-Box LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.15001v2.
Vancouver
1. Davies X, Giglemiani G, Lau E, Winsor E, Irving G, Gal Y (2026) Boundary Point Jailbreaking of Black-Box LLMs. arXiv

BibTeX

@article{davies2026boundary,
  title = {Boundary Point Jailbreaking of Black-Box LLMs},
  author = {Davies, Xander and Giglemiani, Giorgi and Lau, Edmund and Winsor, Eric and Irving, Geoffrey and Gal, Yarin},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.15001v2},
  eprint = {2602.15001}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/