FlipAttack: Jailbreak LLMs via Flipping

Yue Liu 0008Xiaoxin HeMiao XiongJinlan FuShumin DengYingwei MaJiaheng ZhangBryan Hooi

article2025ICML55 citations

Presents FlipAttack, a stealthy single-query black-box jailbreak method that bypasses LLM guardrails with an average 78.97% success rate across eight leading models by perturbing input text from the left and prompting models to flip and execute the disguised request.

Listen

As large language models become increasingly integral to security-critical applications such as finance and medicine, ensuring their safety alignment and resistance to adversarial manipulation is paramount. Existing jailbreak attack methods designed to probe these vulnerabilities often suffer from significant practical limitations: white-box methods require internal model weights and substantial computing power, while conventional black-box techniques demand expensive multi-round interactions or rely on overly complex auxiliary tasks like coding or ciphering that frequently fail. Addressing these challenges is necessary to accurately assess safety boundaries and guardrail efficacy across commercial systems.

The article evaluates model vulnerability by introducing FlipAttack, a black-box jailbreak method that exploits the left-to-right processing nature of autoregressive language models. The primary objective is to demonstrate how simple text-flipping transformations can bypass content moderation guardrails and manipulate models into fulfilling prohibited requests within a single query.

To evaluate this vulnerability, the researchers conducted extensive empirical testing across eight state-of-the-art language models, including major commercial systems like GPT-4 and Claude 3.5 Sonnet, as well as open-source architectures like LLaMA 3.1 405B and Mixtral 8x22B. The attack disguise mechanism reverses text (across word order, character order, whole sentences, or mixed modes) to disrupt standard left-to-right token comprehension and inflate guardrail perplexity, while a structured guidance prompt instructs the model to decode and execute the underlying request.

The key findings show that FlipAttack achieved an average attack success rate of 78.97% across eight models, outperforming the closest black-box baseline by roughly 22 percentage points. Specifically, it achieved success rates of 94.04% on GPT-4 Turbo and 88.08% on Claude 3.5 Sonnet. Furthermore, the disguised prompts bypassed five leading automated guardrail systems with an average bypass rate of 98.08%, including a 100% bypass rate on OpenAI Moderation. Standard heuristic defenses, such as simple system-prompt instructions and basic perplexity-based filtering, failed to mitigate the attack effectively.

These findings indicate that existing guardrail filters and alignment strategies rely heavily on surface-level, standard-order text patterns, leaving substantial blind spots against simple, non-standard structural perturbations. Because this approach succeeds in a single query without requiring model weight access, it represents a cost-effective and highly transferable vulnerability. Standard safety filters that scan for direct keywords or expected text structures will require significant updates to handle disguised transformations.

To mitigate these risks, the article suggests that safety teams integrate advanced red-teaming and alignment procedures that train models to detect structural obfuscation directly. The authors noted that simple countermeasures, like adjusting perplexity thresholds, offer minimal protection while inadvertently raising false-positive rejection rates on benign inputs. Organizations deploying language models should evaluate reasoning-based safeguards rather than relying solely on surface-level guardrails.

The scope of this evaluation presents several limitations. FlipAttack showed lower success rates against certain complex or highly sensitive categories, such as physical harm, compared to categories like digital fraud and software vulnerabilities. Additionally, the approach showed lower effectiveness against specialized reasoning-dense models, and few-shot demonstrations can occasionally expose plain harmful words that trigger standard filters. While confidence in the vulnerability of current standard systems is high, readers should view these findings as a baseline for updating moderation architectures rather than an exhaustive evaluation across all emerging reasoning models.

Cover for FlipAttack: Jailbreak LLMs via Flipping

Abstract

This paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we reveal that LLMs tend to understand the text from left to right and find that they struggle to comprehend the text when the perturbation is added to the left side. Motivated by these insights, we propose to disguise the harmful prompt by constructing a left-side perturbation merely based on the prompt itself, then generalize this idea to 4 flipping modes. Second, we verify the strong ability of LLMs to perform the text-flipping task and then develop 4 variants to guide LLMs to understand and execute harmful behaviors accurately. These designs keep FlipAttack universal, stealthy, and simple, allowing it to jailbreak black-box LLMs within only 1 query. Experiments on 8 LLMs demonstrate the superiority of FlipAttack. Remarkably, it achieves ~78.97% attack success rate across 8 LLMs on average and ~98% bypass rate against 5 guard models on average. The codes are available^1.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. FlipAttack
  • 3.1.1. ATTACK DISGUISE MODULE
  • 3.1.2. FLIPPING GUIDANCE MODULE
  • 3.2. Defense Strategy
  • 4. Experiment
  • 4.1. Attack Performance
  • 4.2. Ablation Study
  • 4.3. Exploring Why FlipAttack Succeeds
  • 5. Conclusion
  • Impact Statement
  • References
  • A. Appendix
  • A.1. Experimental Setup
  • A.1.1. EXPERIMENTAL ENVIRONMENT
  • A.1.2. BENCHMARK
  • A.1.3. BASELINE
  • A.1.4. TARGET LLM
  • A.1.5. EVALUATION
  • A.1.6. IMPLEMENTATION
  • A.2. Additional Experiment
  • A.3. Testing of Evaluation Metric
  • A.4. Testing of Defense Strategy
  • A.4.1. TESTING OF SYSTEM PROMPT DEFENSE
  • A.4.2. TESTING OF PERPLEXITY-BASED GUARDRAIL FILTER
  • A.5. Testing of Stealthiness
  • A.6. Case Study
  • A.7. Limitation
  • A.8. Prompt Design
  • A.9. Ethical Consideration

Knowls

  1. Knowl 1 — Left-side perturbations disrupt language-model understanding

    empirical result

    The paper argues that autoregressive LLMs tend to process text from left to right, so perturbations placed before a target sentence interfere more strongly with understanding than equally long perturbations placed after it. For an input sentence XX and a random string NN of the same length, the paper compares XX, X+NX+N, and N+XN+X using perplexity (PPL), where lower PPL indicates better model understanding. Across three LLMs and four guard models, the mean PPL values are 74.7574.75 for XX, 477.09477.09 for X+NX+N, and 815.93815.93 for N+XN+X.

    Could not parse LaTeX table

    The authors hypothesize that errors introduced early in the sequence propagate through later positions because subsequent autoregressive predictions depend on earlier context.

  2. Knowl 2 — Attack disguise by self-derived flipping

    model/method

    FlipAttack disguises a harmful request using only information already present in that request, rather than adding an externally optimized suffix or unrelated perturbation. The attack repeatedly moves the unprocessed portion of a request to the left of the currently processed portion, disrupting the model's left-to-right interpretation while preserving all original information.

    The method instantiates this idea through four modes:

    1. Flip Word Order: reverse the order of words while preserving the characters inside each word.
    2. Flip Characters in Word: preserve word order but reverse the characters within every word.
    3. Flip Characters in Sentence: reverse the character sequence of the entire sentence.
    4. Fool Model Mode: reverse the characters of the sentence but instruct the victim LLM to recover the request by reversing word order, creating a mismatch between the actual transformation and the stated recovery rule.

    The resulting prompt remains a transformed version of the original request, but its unusual word order or character structure substantially increases the difficulty of direct recognition by safety classifiers and aligned victim models.

  3. Knowl 3 — Flipping guidance module for recovery and task execution

    model/method

    After disguising a request, FlipAttack adds guidance intended to make the victim LLM recover the underlying text and complete the recovered task. The guidance has four nested variants:

    1. Vanilla: ask the model to read the disguised request, reverse the relevant flipping operation, and complete the recovered task while not explicitly discussing the harmful behavior or changing the task.
    2. Vanilla+CoT: add step-by-step reasoning instructions for completing the recovery task.
    3. Vanilla+CoT+LangGPT: organize the instructions as a role, profile, rules, target, and initialization structure so that the model has a more explicit task specification.
    4. Vanilla+CoT+LangGPT+Few-shot: add task-oriented demonstrations constructed from the original request itself. The request is split at its midpoint, and transformed halves are used as examples, avoiding the introduction of unrelated information.

    The ablations show that Vanilla already performs strongly on some capable models, reaching 98.08%98.08\% ASR on GPT-4 Turbo, 88.85%88.85\% on GPT-4, and 86.35%86.35\% on GPT-4o, but reaches only 30.58%30.58\% on GPT-3.5 Turbo. CoT improves most cases, including a reported 16.9216.92 percentage-point gain on Claude 3.5 Sonnet, although it can reduce performance on GPT-4o mini by making the harmful nature of the task more apparent. LangGPT raises one reported GPT-3.5 Turbo result from 39.04%39.04\% to 70.38%70.38\%, and task-oriented few-shot learning gives a reported 16.1616.16 percentage-point improvement on GPT-3.5 Turbo. The few-shot design can also expose harmful words and increase detectability.

  4. Knowl 4 — One-query FlipAttack procedure

    algorithm

    FlipAttack takes as input a harmful request, a black-box victim LLM, one of the four flipping modes, and one of the four guidance variants. It returns the victim LLM's response to the disguised request.

    1. Transform the harmful request using a selected flipping mode. The transformation is deterministic and uses no model weights, gradients, external suffix, cipher, or iterative search.
    2. Construct a guidance prompt that tells the victim LLM how to recover the transformed request and complete it. Select the guidance variant according to the target model's ability to perform the flipping task.
    3. Submit the constructed system/user prompt to the victim LLM once.
    4. Treat the single returned response as the attack result; FlipAttack performs no iterative refinement or additional victim queries.

    The configurations used in the paper are: Vanilla with Flip Word Order for GPT-3.5 Turbo; Vanilla+CoT with Flip Characters in Word for GPT-4 and GPT-4 Turbo; Vanilla+CoT+LangGPT with Flip Characters in Sentence for GPT-4o mini; Vanilla+CoT+LangGPT+Few-shot for GPT-4o and Mixtral 8x22B; and Vanilla+CoT with Fool Model Mode for Claude 3.5 Sonnet and LLaMA 3.1 405B. The default Vanilla configuration uses Flip Characters in Sentence mode. The method's computational interaction cost with the victim is therefore one query, although the transformed prompt can contain more tokens than the original.

  5. Knowl 5 — Evaluation protocol for black-box jailbreak performance

    experimental setup

    The main experiments use the 520 harmful-behavior prompts in AdvBench, with an additional 50-prompt AdvBench subset for comparison. FlipAttack is compared with four white-box methods—GCG, AutoDAN, MAC, and COLD-Attack—and eleven black-box baselines, including PAIR, TAP, Base64, GPTFuzzer, DeepInception, DRA, ArtPrompt, PromptAttack, SelfCipher, CodeChameleon, and ReNeLLM.

    The target systems are GPT-3.5 Turbo, GPT-4 Turbo, GPT-4, GPT-4o, GPT-4o mini, Claude 3.5 Sonnet, LLaMA 3.1 405B, and Mixtral 8x22B. The primary metric is GPT-based attack success rate (ASR-GPT), in which GPT-4 judges whether the response both addresses the request and violates safety requirements. The paper also reports dictionary-based ASR, which marks a response as unsuccessful when it contains a predefined refusal phrase. On a 300-pair evaluation set labeled by three human experts, GPT-4 evaluation agrees with the human majority vote in 90.30%90.30\% of cases, whereas dictionary-based evaluation agrees in only 56.00%56.00\%; the paper therefore treats ASR-GPT as its primary metric.

  6. Knowl 6 — FlipAttack outperforms the evaluated attack baselines

    data/table

    On the full AdvBench evaluation, FlipAttack obtains the highest mean ASR-GPT among the 16 evaluated methods. Its average ASR is 78.97%78.97\%, exceeding the runner-up ReNeLLM at 56.64%56.64\% by 22.3322.33 percentage points. FlipAttack reaches 94.04%94.04\% on GPT-4 Turbo, 86.73%86.73\% on GPT-4, and 88.08%88.08\% on Claude 3.5 Sonnet. The table also shows weak transfer of several white-box attacks to commercial models; for example, GCG averages only 7.40%7.40\%.

    Could not parse LaTeX table

    Values are ASR-GPT percentages for GPT-3.5 Turbo, GPT-4 Turbo, GPT-4, GPT-4o, GPT-4o mini, Claude 3.5 Sonnet, LLaMA 3.1 405B, Mixtral 8x22B, and the cross-model average, respectively.

  7. Knowl 7 — Flipped prompts bypass tested guard models

    data/table

    FlipAttack is evaluated against one closed-source moderation endpoint and four open-source guard models. The bypass rate is the percentage of attack prompts not detected by a guard; therefore, a higher value indicates weaker guard performance. FlipAttack achieves a mean bypass rate of 98.08%98.08\%, with complete bypass on OpenAI's Moderation Endpoint and LLaMA Guard 2 8B.

    Could not parse LaTeX table
  8. Knowl 8 — Flipping increases guard-model perplexity and disrupts tokenization

    data/table

    The paper defines stealthiness operationally as high perplexity on guard models: a guard with high PPL is presumed to understand the concealed request less well. Across three LLMs and four guard LLMs, FlipAttack has a much higher mean PPL than the original harmful prompt and competing encodings. The authors also report that flipping increases token counts, especially when characters rather than words are reversed, suggesting that original words are split into less familiar token fragments.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The first table compares mean and standard-deviation PPL across the seven open-source evaluation models; the second reports average prompt-token counts under three tokenizers or APIs. FlipAttack has the highest reported PPL, while character-level modes use more than twice as many tokens as the original prompt.

  9. Knowl 9 — Victim models can usually recover the flipped request

    data/table

    To measure whether the guidance task itself is feasible, the paper applies flipping to 200 benign prompts from the Alpaca safe dataset and computes the exact match rate between each model's recovered sentence and the original sentence. Strong models such as GPT-4 Turbo, GPT-4o, and Claude 3.5 Sonnet exceed 95%95\% match rate without demonstrations. Task-oriented few-shot examples substantially help some weaker models, most notably increasing LLaMA 3.1 405B from 44.80%44.80\% to 90.46%90.46\%.

    Could not parse LaTeX table

    These results support the paper's design: flipping is difficult for some models to interpret directly, but recovering the original request is often easy enough for the victim model when appropriate guidance is supplied.

  10. Knowl 10 — Simple defenses provide limited protection

    empirical result

    The paper tests two defenses against FlipAttack. System Prompt Defense (SPD) adds a system-level instruction telling the victim to be safe and helpful. Perplexity-based Guardrail Filter (PGF) rejects prompts whose WildGuard 7B perplexity is at least 15001500; this threshold is selected to keep the benign rejection rate below 5%5\% in the authors' calibration set. SPD does not reliably reduce attack success and slightly increases the mean ASR in the reported experiment. PGF reduces mean ASR from 81.76%81.76\% to 74.40%74.40\%, a reduction the paper describes as about 7.167.16 percentage points, while rejecting 4%4\% of benign prompts.

    Could not parse LaTeX table

    The paper concludes that system instructions and simple perplexity thresholding are insufficient as standalone defenses against the flipping attack.

  11. Knowl 11 — Stated limitations of FlipAttack

    limitation

    The paper identifies three limitations. First, the iterative self-derived perturbation does not always produce the maximum possible perplexity; its effectiveness varies across prompts, so stronger perturbation construction remains open. Second, task-oriented few-shot guidance can reveal harmful words because demonstrations are derived from portions of the original request, causing some attacks to be detected; a more stealthy splitting and demonstration strategy is needed. Third, FlipAttack is reported to be less effective against models with strong reasoning capabilities, including OpenAI's o1, so bypassing reasoning-based safety systems remains unresolved.

Coverage note — Detailed prompt templates, individual case studies, secondary ASR-DICT/StrongREJECT comparisons, and full guard-category breakdowns were omitted because they instantiate or repeat the main method and headline evaluations rather than adding separate load-bearing contributions.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Alon, G. and Kamfonas, M. Detecting language model attacks with perplexity. arXiv preprint arXiv:2308.14132, 2023.
  3. 3.Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al. Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 1, 2023.
  4. 4.Bach, S. H., Sanh, V., Yong, Z.-X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Fevry, T., et al. Promptsource: An integrated development environment and repository for natural language prompts. arXiv preprint arXiv:2202.01279, 2022.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  6. 6.Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419, 2023.
  7. 7.Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramer, F., et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. arXiv preprint arXiv:2404.01318, 2024.
  8. 8.Chen, H., Zhang, Y., Dong, Y., Yang, X., Su, H., and Zhu, J. Rethinking model ensemble in transfer-based adversarial attacks. arXiv preprint arXiv:2303.09105, 2023.
  9. 9.Chen, Z., Zhao, Z., Qu, W., Wen, Z., Han, Z., Zhu, Z., Zhang, J., and Yao, H. Pandora: Detailed llm jailbreaking via collaborated phishing agents with decomposed reasoning. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models, 2024.
  10. 10.Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y. Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023.
  11. 11.Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., and Huang, S. A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily. arXiv preprint arXiv:2311.08268, 2023.
  12. 12.Ding, Y., Li, B., and Zhang, R. Eta: Evaluating then aligning safety of vision language models at inference time. arXiv preprint arXiv:2410.06625, 2024.
  13. 13.Ding, Y., Li, L., Cao, B., and Shao, J. Rethinking bottlenecks in safety fine-tuning of vision language models. arXiv preprint arXiv:2501.18533, 2025.
  14. 14.Duan, M., Li, Q., and He, B. ModelGo: A practical tool for machine learning license analysis. In Proceedings of the ACM Web Conference 2024, pp. 1158–1169, 2024. doi: 10.1145/3589334.3645520.
  15. 15.Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
  16. 16.Ethayarajh, K., Choi, Y., and Swayamdipta, S. Understanding dataset difficulty with mathcal v-usable information. In International Conference on Machine Learning. PMLR, 2022.
  17. 17.Fang, J., Jiang, H., Wang, K., Ma, Y., Jie, S., Wang, X., He, X., and Chua, T.-S. Alphaedit: Null-space constrained knowledge editing for language models. ICLR, 2025a.
  18. 18.Fang, J., Wang, Y., Wang, R., Yao, Z., Wang, K., Zhang, A., Wang, X., and Chua, T.-S. Safemlrm: Demystifying safety in multi-modal large reasoning models. arXiv preprint arXiv:2504.08813, 2025b.
  19. 19.Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858, 2022.
  20. 20.Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.-C., Wang, Q., Han, J., and Mao, Y. Mart: Improving llm safety with multi-round automatic red-teaming. arXiv preprint arXiv:2311.07689, 2023.
  21. 21.Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Del Giorno, A., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., et al. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023.
  22. 22.Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., and Dziri, N. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. arXiv preprint arXiv:2406.18495, 2024.
  23. 23.He, L., Xia, M., and Henderson, P. What’s in your ”safe” data?: Identifying benign data that breaks safety, 2024.
  24. 24.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  25. 25.Hu, L., Huang, T., Xie, H., Gong, X., Ren, C., Hu, Z., Yu, L., Ma, P., and Wang, D. Semi-supervised concept bottleneck models. arXiv preprint arXiv:2406.18992, 2024a.
  26. 26.Hu, L., Ren, C., Hu, Z., Lin, H., Wang, C.-L., Xiong, H., Zhang, J., and Wang, D. Editable concept bottleneck models. arXiv preprint arXiv:2405.15476, 2024b.
  27. 27.Hu, Z., Zhang, J., Wang, H., Liu, S., and Liang, S. Leveraging relational graph neural network for transductive model ensemble. In Proceedings of the 29th ACM SIGKDD Conference on knowledge discovery and data mining, pp. 775–787, 2023a.
  28. 28.Hu, Z., Zhang, J., Yu, Y., Zhuang, Y., and Xiong, H. How many validation labels do you need? exploring the design space of label-efficient model ranking. arXiv preprint arXiv:2312.01619, 2023b.
  29. 29.Hu, Z., Li, Y., Chen, Z., Wang, J., Liu, H., Lee, K., and Ding, K. Let’s ask gnn: Empowering large language model for graph in-context learning. arXiv preprint arXiv:2410.07074, 2024c.
  30. 30.Hu, Z., Song, L., Zhang, J., Xiao, Z., Chen, Z., and Xiong, H. Explaining length bias in llm-based preference evaluations. In ICLR 2025 Workshop on Navigating and Addressing Data Problems for Foundation Models, 2024d.
  31. 31.Hu, Z., Song, L., Zhang, J., Xiao, Z., Wang, J., Chen, Z., Zhao, J., and Xiong, H. Rethinking llm-based preference evaluation. arXiv e-prints, pp. arXiv–2407, 2024e.
  32. 32.Hu, Z., Zhang, J., Xiong, Z., Ratner, A., Xiong, H., and Krishna, R. Language model preference evaluation with multiple weak evaluators. arXiv preprint arXiv:2410.12869, 2024f.
  33. 33.Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D. Catastrophic jailbreak of open-source llms via exploiting generation. arXiv preprint arXiv:2310.06987, 2023.
  34. 34.Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Dang, K., et al. Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186, 2024.
  35. 35.Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023.
  36. 36.Ji, J., Hou, B., Robey, A., Pappas, G. J., Hassani, H., Zhang, Y., Wong, E., and Chang, S. Defending large language models against jailbreak attacks via semantic smoothing. arXiv preprint arXiv: 2402.16192, 2024.
  37. 37.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024a.
  38. 38.Jiang, F., Xu, Z., Niu, L., Xiang, Z., Ramasubramanian, B., Li, B., and Poovendran, R. Artprompt: Ascii art-based jailbreak attacks against aligned llms. arXiv preprint arXiv:2402.11753, 2024b.
  39. 39.Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T. Exploiting programmatic behavior of llms: Dual-use through standard security attacks. In 2024 IEEE Security and Privacy Workshops (SPW), pp. 132–143. IEEE, 2024.
  40. 40.Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E. Pretraining language models with human preferences. In International Conference on Machine Learning, pp. 17506–17533. PMLR, 2023.
  41. 41.Li, X., Zhou, Z., Zhu, J., Yao, J., Liu, T., and Han, B. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191, 2023.
  42. 42.Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H. Rain: Your language models can align themselves without fine-tuning. In International Conference on Learning Representations, 2024.
  43. 43.Li, Z.-Z., Zhang, D., Zhang, M.-L., Zhang, J., Liu, Z., Yao, Y., Xu, H., Zheng, J., Wang, P.-J., Chen, X., Zhang, Y., Yin, F., Dong, J., Li, Z., Bi, B.-L., Mei, L.-R., Fang, J., Guo, Z., Song, L., and Liu, C.-L. From system 1 to system 2: A survey of reasoning large language models, 2025. URL https://arxiv.org/abs/2502.17419.
  44. 44.Liu, B., Li, X., Zhang, J., Wang, J., He, T., Hong, S., Liu, H., Zhang, S., Song, K., Zhu, K., et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv preprint arXiv:2504.01990, 2025a.
  45. 45.Liu, T., Zhao, Z., Dong, Y., Meng, G., and Chen, K. Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. In 33rd USENIX Security Symposium (USENIX Security 24), pp. 4711–4728, 2024a.
  46. 46.Liu, X., Xu, N., Chen, M., and Xiao, C. Autodan: Generating stealthy jailbreak prompts on aligned large language models. In The Twelfth International Conference on Learning Representations, 2024b. URL https://openreview.net/forum?id=7Jwpw4qKkb.
  47. 47.Liu, Y., Gao, H., Zhai, S., Jun, X., Wu, T., Xue, Z., Chen, Y., Kawaguchi, K., Zhang, J., and Hooi, B. Guardreasoner: Towards reasoning-based llm safeguards. arXiv preprint arXiv:2501.18492, 2025b.
  48. 48.Liu, Y., Wu, J., He, Y., Gao, H., Chen, H., Bi, B., Zhang, J., Huang, Z., and Hooi, B. Efficient inference for large reasoning models: A survey. arXiv preprint arXiv:2503.23077, 2025c.
  49. 49.Liu, Y., Zhai, S., Du, M., Chen, Y., Cao, T., Gao, H., Wang, C., Li, X., Wang, K., Fang, J., Zhang, J., and Hooi, B. Guardreasoner-vl: Safeguarding vlms via reinforced reasoning. arXiv preprint arXiv:2505.11049, 2025d.
  50. 50.Lu, X., Huang, Z., Li, X., Xu, W., et al. Poex: Policy executable embodied ai jailbreak attacks. arXiv preprint arXiv:2412.16633, 2024.
  51. 51.Luo, H., Gu, J., Liu, F., and Torr, P. An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models. arXiv preprint arXiv:2403.09766, 2024.
  52. 52.Lv, H., Wang, X., Zhang, Y., Huang, C., Dou, S., Ye, J., Gui, T., Zhang, Q., and Huang, X. Codechameleon: Personalized encryption framework for jailbreaking large language models. arXiv preprint arXiv:2402.16717, 2024.
  53. 53.Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A. Tree of attacks: Jailbreaking black-box llms automatically. arXiv preprint arXiv:2312.02119, 2023.
  54. 54.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35, 2022.
  55. 55.Phute, M., Helbling, A., Hull, M., Peng, S., Szyller, S., Cornelius, C., and Chau, D. H. Llm self defense: By self examination, llms know they are being tricked. arXiv preprint arXiv:2308.07308, 2023.
  56. 56.Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
  57. 57.Qin, L., Welleck, S., Khashabi, D., and Choi, Y. Cold decoding: Energy-based constrained text generation with langevin dynamics. Advances in Neural Information Processing Systems, 35:9538–9551, 2022.
  58. 58.Ramesh, G., Dou, Y., and Xu, W. Gpt-4 jailbreaks itself with near-perfect success using self-explanation. arXiv preprint arXiv:2405.13077, 2024.
  59. 59.Rando, J. and Tramer, F. Universal jailbreak backdoors from poisoned human feedback. arXiv preprint arXiv:2311.14455, 2023.
  60. 60.Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  61. 61.Robey, A., Wong, E., Hassani, H., and Pappas, G. J. Smoothllm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684, 2023.
  62. 62.Shayegani, E., Dong, Y., and Abu-Ghazaleh, N. Jailbreak in pieces: Compositional adversarial attacks on multimodal language models. In The Twelfth International Conference on Learning Representations, 2023.
  63. 63.Solaiman, I. and Dennison, C. Process for adapting language models to society (palms) with values-targeted datasets. Advances in Neural Information Processing Systems, 34:5861–5873, 2021.
  64. 64.Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., et al. A strongreject for empty jailbreaks. arXiv preprint arXiv:2402.10260, 2024.
  65. 65.Team, A. The claude 3 model family: Opus, sonnet, haiku. https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model Card Claude 3.pdf, 2024.
  66. 66.Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W. Large language models in medicine. Nature medicine, 29(8):1930–1940, 2023.
  67. 67.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  68. 68.Wang, B., Ping, W., Xiao, C., Xu, P., Patwary, M., Shoeybi, M., Li, B., Anandkumar, A., and Catanzaro, B. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. Advances in Neural Information Processing Systems, 35:35811–35824, 2022a.
  69. 69.Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., et al. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. In NeurIPS, 2023.
  70. 70.Wang, C., Liu, Y., Li, B., Zhang, D., Li, Z., and Fang, J. Safety in large reasoning models: A survey. arXiv preprint arXiv:2504.17704, 2025a.
  71. 71.Wang, K., Zhang, G., Zhou, Z., Wu, J., Yu, M., Zhao, S., Yin, C., Fu, J., Yan, Y., Luo, H., et al. A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv:2504.15585, 2025b.
  72. 72.Wang, M., Liu, Y., Zhang, X., Li, S., Huang, Y., Zhang, C., Wang, D., Feng, S., and Li, J. Langgpt: Rethinking structured reusable prompt design framework for llms from the programming language. arXiv preprint arXiv:2402.16929, 2024a.
  73. 73.Wang, M., Zhang, N., Xu, Z., Xi, Z., Deng, S., Yao, Y., Zhang, Q., Yang, L., Wang, J., and Chen, H. Detoxifying large language models via knowledge editing. arXiv preprint arXiv:2403.14472, 2024b.
  74. 74.Wang, T., Zhan, Y., Lian, J., Hu, Z., Yuan, N. J., Zhang, Q., Xie, X., and Xiong, H. Llm-powered multi-agent framework for goal-oriented learning in intelligent tutoring system. arXiv preprint arXiv:2501.15749, 2025c.
  75. 75.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560, 2022b.
  76. 76.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. arXiv preprint arXiv:2204.07705, 2022c.
  77. 77.Wang, Y., Liu, X., Li, Y., Chen, M., and Xiao, C. Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield prompting. arXiv preprint arXiv:2403.09513, 2024c.
  78. 78.Wei, A., Haghtalab, N., and Steinhardt, J. Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.
  79. 79.Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445, 2021.
  80. 80.WitchBOT. You can use gpt-4 to create prompt injections against gpt-4. https://www.lesswrong.com/posts/bNCDexejSZpkuu3yz/you-can-use-gpt-4-to-create-prompt-injections-against-gpt-4, 2023.
  81. 81.Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P. Recursively summarizing books with human feedback. arXiv preprint arXiv:2109.10862, 2021.
  82. 82.Xie, Y., Yi, J., Shao, J., Curl, J., Lyu, L., Chen, Q., Xie, X., and Wu, F. Defending chatgpt against jailbreak attack via self-reminders. Nature Machine Intelligence, 5(12):1486–1496, 2023.
  83. 83.Xie, Y., Fang, M., Pi, R., and Gong, N. Gradsafe: Detecting unsafe prompts for llms via safety-critical gradient analysis. arXiv preprint arXiv:2402.13494, 2024.
  84. 84.Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079, 2020.
  85. 85.Xu, X., Kong, K., Liu, N., Cui, L., Wang, D., Zhang, J., and Kankanhalli, M. An llm can fool itself: A prompt-based adversarial attack. URL: http://arxiv.org/abs/2310.13345, 2023.
  86. 86.Xu, Z., Jiang, F., Niu, L., Jia, J., Lin, B. Y., and Poovendran, R. Safedecoding: Defending against jailbreak attacks via safety-aware decoding. arXiv preprint arXiv:2402.08983, 2024.
  87. 87.Yao, D., Zhang, J., Harris, I. G., and Carlsson, M. Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4485–4489. IEEE, 2024.
  88. 88.Yin, Z., Ye, M., Zhang, T., Du, T., Zhu, J., Liu, H., Chen, J., Wang, T., and Ma, F. Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models. Advances in Neural Information Processing Systems, 36, 2024.
  89. 89.Yu, J., Lin, X., and Xing, X. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253, 2023.
  90. 90.Yu, Z., Liu, X., Liang, S., Cameron, Z., Xiao, C., and Zhang, N. Don’t listen to me: Understanding and exploring jailbreak prompts of large language models. arXiv preprint arXiv:2403.17336, 2024.
  91. 91.Yuan, L., Li, X., Xu, C., Tao, G., Jia, X., Huang, Y., Dong, W., Liu, Y., Wang, X., and Li, B. Promptguard: Soft prompt-guided unsafe content moderation for text-to-image models. arXiv preprint arXiv:2501.03544, 2025.
  92. 92.Yuan, Y., Jiao, W., Wang, W., Huang, J.-t., He, P., Shi, S., and Tu, Z. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463, 2023.
  93. 93.Zhang, H., Guo, Z., Zhu, H., Cao, B., Lin, L., Jia, J., Chen, J., and Wu, D. Jailbreak open-sourced large language models via enforced decoding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5475–5493, 2024.
  94. 94.Zhang, J., Wang, B., Hu, Z., Koh, P. W. W., and Ratner, A. J. On the trade-off of intra-/inter-class diversity for supervised pre-training. Advances in Neural Information Processing Systems, 36:64193–64212, 2023a.
  95. 95.Zhang, Y. and Wei, Z. Boosting jailbreak attack with momentum. arXiv preprint arXiv:2405.01229, 2024.
  96. 96.Zhang, Z., Yang, J., Ke, P., and Huang, M. Defending large language models against jailbreaking attacks through goal prioritization. arXiv preprint arXiv:2311.09096, 2023b.
  97. 97.Zhao, H., Liu, Z., Wu, Z., Li, Y., Yang, T., Shu, P., Xu, S., Dai, H., Zhao, L., Mai, G., et al. Revolutionizing finance with llms: An overview of applications and insights. arXiv preprint arXiv:2401.11641, 2024.
  98. 98.Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N. On prompt-driven safeguarding for large language models. In Forty-first International Conference on Machine Learning, 2024a.
  99. 99.Zheng, X., Pang, T., Du, C., Liu, Q., Jiang, J., and Lin, M. Improved few-shot jailbreaking can circumvent aligned language models and their defenses. arXiv preprint arXiv:2406.01288, 2024b.
  100. 100.Zhou, Z., Yu, H., Zhang, X., Xu, R., Huang, F., Wang, K., Liu, Y., Fang, J., and Li, Y. On the role of attention heads in large language model safety. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=h0Ak8A5yqw.
  101. 101.Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T. Autodan: Interpretable gradient-based adversarial attacks on large language models. In First Conference on Language Modeling, 2023.
  102. 102.Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences. arXiv preprint arXiv:1909.08593, 2019.
  103. 103.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models, 2023.

Citation

MLA
Liu, Y., et al. “FlipAttack: Jailbreak LLMs via Flipping”. arXiv, 2024, http://arxiv.org/abs/2410.02832v2.
APA
Liu, Y., He, X., Xiong, M., Fu, J., Deng, S., Ma, Y., Zhang, J., & Hooi, B. (2024). FlipAttack: Jailbreak LLMs via Flipping. arXiv. http://arxiv.org/abs/2410.02832v2
Chicago
Liu, Y., X. He, M. Xiong, et al. 2024. “FlipAttack: Jailbreak LLMs via Flipping”. arXiv. http://arxiv.org/abs/2410.02832v2.
Harvard
Liu, Y. et al. (2024) “FlipAttack: Jailbreak LLMs via Flipping”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.02832v2.
Vancouver
1. Liu Y, He X, Xiong M, Fu J, Deng S, Ma Y, Zhang J, Hooi B (2024) FlipAttack: Jailbreak LLMs via Flipping. arXiv

BibTeX

@article{liu2024flipattack,
  title = {FlipAttack: Jailbreak LLMs via Flipping},
  author = {Liu, Yue and He, Xiaoxin and Xiong, Miao and Fu, Jinlan and Deng, Shumin and Ma, Yingwei and Zhang, Jiaheng and Hooi, Bryan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.02832v2},
  eprint = {2410.02832}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/