Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Bochuan CaoYuanpu CaoLu LinJinghui Chen
Proposes a plug-and-play defense mechanism that drops random tokens from input requests to invalidate jailbreak prompts, reducing attack success rates from nearly 100% to around 10% without retraining the target language model.
Large Language Models are increasingly integrated into critical commercial workflows, but they remain vulnerable to alignment-breaking attacks. In these exploits, adversaries bypass safety guardrails by appending crafted prompts or role-play scenarios to elicit harmful and toxic content. Existing defenses often depend on external discriminator models that are computationally expensive, prone to misclassifying benign user prompts, and narrow in scope.
The article demonstrates a defense mechanism termed Robustly Aligned Large Language Model (RA-LLM), designed to fortify existing safety safeguards without requiring model retraining or fine-tuning.
The approach operates on the insight that adversarial prompts are fragile to input disruptions, whereas underlying safety safeguards are robust. The system generates multiple sampled variations of an incoming prompt by randomly dropping a subset of its tokens and checking if the model’s internal alignment triggers a refusal. A prompt is accepted only if the majority of these sampled variations pass without activating safety refusals. The researchers validated the method using mathematical proofs alongside empirical evaluations on several open-source and commercial language models against state-of-the-art automated attacks and popular handcrafted exploits.
The findings show that the proposed defense reduces attack success rates from 80–99% down to 6–12% across varied model architectures. Benign request handling remains virtually unaffected, maintaining answering rates between 92% and 99.3%. In direct comparisons, alternative defenses like perplexity checks completely failed against handcrafted attacks, whereas RA-LLM defended against them reliably. In addition, implementation optimizations—such as limiting Monte Carlo generation lengths and implementing early exit rules—kept additional processing time below 20% relative to standard inference.
These results provide a low-overhead, plug-and-play defense strategy for organizations deploying language models. Organizations can significantly mitigate safety, legal, and compliance risks without undertaking expensive fine-tuning cycles or relying on fragile external filtering models.
For practical deployment, teams should implement this random-dropping verification layer with tunable thresholds calibrated to their specific risk tolerance, balancing security against user friction. Future work should focus on testing performance against extreme prompt lengths and refining alignment training to make safety refusals more distinct from benign clarification requests.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Introduces automated gradient-based adversarial suffixes on aligned LLMs, establishing the primary attack paradigm and benchmark that RA-LLM is engineered to defend against.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). Analyzes why standard safety alignment training fails against adversarial prompts, providing the foundational conceptual failure modes motivating post-hoc alignment defenses.
- Paper: Jailbreaking Black Box Large Language Models in Twenty Queries, Patrick Chao et al. (2023). Demonstrates efficient automated black-box jailbreaking techniques, contextualizing the broad spectrum of alignment-breaking attacks RA-LLM aims to mitigate.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). Presents input-output safety guardrail classification, serving as a core conceptual prerequisite for checking and filtering alignment violations during inference.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). Examines how downstream modifications easily compromise base model safety, motivating training-free defense mechanisms that safeguard aligned LLMs without fine-tuning.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). Details the standard reinforcement learning from human feedback pipeline for harmlessness, establishing the baseline aligned models evaluated in RA-LLM.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). Evaluates simple adaptive attacks against defended and safety-aligned LLMs, testing the boundaries of alignment checking mechanisms.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). Develops enhanced optimization-based jailbreaks that specifically counter mid-response refusals and safety-checking behaviors.
- Paper: COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability, Xingang Guo et al. (2024). Introduces controllable and fluent energy-based jailbreaks designed to evade prompt-level and perplexity-based alignment defenses.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). Explores genetic algorithm-based stealthy jailbreaks capable of bypassing automated input checks and content filters.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). Provides a comprehensive standardized benchmark (HarmBench) to evaluate and compare advanced red-teaming attacks and robust refusal defenses like RA-LLM.
- Paper: Aligning Large Language Models with Representation Editing: A Control Perspective, Lingkai Kong et al. (2024). Investigates representation editing as an alternative inference-time steering defense to maintain alignment without full retraining.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). Extends inference-stage alignment by enforcing multi-faceted safety constraints dynamically during the decoding process.
- Paper: FlipAttack: Jailbreak LLMs via Flipping, Yue Liu 0008 et al. (2025). Presents token-order transformation attacks that challenge safety checking functions by disguising harmful requests via left-to-right processing exploits.
