Fundamental Limitations of Alignment in Large Language Models
Yotam WolfNoam WiesOshri AvneryYoav LevineAmnon Shashua
Establishes a theoretical framework called Behavior Expectation Bounds to mathematically prove that standard alignment methods like RLHF cannot prevent adversarial jailbreaks as long as harmful behaviors retain a non-zero probability of being generated.
Large language models are increasingly deployed as public-facing assistants, yet they remain vulnerable to generating harmful, toxic, or biased content. While current safety alignment techniques—primarily reinforcement learning from human feedback and protective system prompts—aim to suppress these harmful behaviors, widespread real-world jailbreaks demonstrate that these defenses frequently fail. The article addresses this critical vulnerability by investigating whether current alignment practices can ever provide mathematically robust safety guarantees against adversarial prompting.
The main objective of the article is to establish a formal theoretical framework to evaluate the fundamental limitations of alignment in language models whose weights remain frozen at inference time. The authors demonstrate that if a model retains any non-zero probability of generating an undesired behavior, adversarial prompts can always be constructed to reliably trigger that behavior.
To conduct this evaluation, the authors develop a theoretical framework called Behavior Expectation Bounds. This framework models a language model's output distribution as a mixture of well-behaved and ill-behaved components, measuring the statistical distance and likelihood variance between them. The theoretical findings were supported empirically using models from the LLaMA 2 family across multiple evaluated behaviors, such as agreeableness and anti-immigration, utilizing specialized behavioral datasets to track distribution shifts and misaligning dynamics.
The analysis reveals several core findings. First, any alignment process that merely reduces the probability of a harmful behavior without eliminating it entirely is provably vulnerable to adversarial prompts, with the required prompt length scaling logarithmically with the rarity of the behavior. Second, while preset aligning system prompts provide some defense, an adversary can overcome them with an adversarial prompt whose length grows linearly with the preset prompt length. Third, in interactive multi-turn conversations, aligned models can initially resist misalignment through safe replies, but adversaries can still trigger harmful behavior by accumulating sufficient misaligning text over turns. Fourth, sampling multiple candidate outputs and selecting the most aligned response only delays misalignment logarithmically relative to the number of samples. Finally, empirical tests confirm that reinforcement learning from human feedback may unintentionally make undesired behaviors more statistically distinct, rendering them about five times more susceptible to targeted prompting than unaligned base models.
These findings imply that popular post-training alignment methods act only as temporary hurdles rather than absolute safety barriers. Relying exclusively on frozen-weight techniques like reinforcement learning from human feedback or system prompts leaves systems fundamentally vulnerable to adversarial manipulation, creating significant compliance, safety, and reputational risks for organizations deploying them in high-stakes environments.
Organizations developing and deploying language models should not rely solely on static alignment tuning or system prompts for robust security. Instead, teams should investigate runtime intervention mechanisms, such as representation engineering, which steer internal model activations during inference, and enforce strict constraints on input prompt lengths to bound adversarial attack capacity. Further research should focus on validating internal behavior superposition and establishing standardized, automated evaluations across broader behavioral categories.
The findings are bounded by the assumption that behavior can be scored reliably at the sentence level and that language model outputs can be approximated as mixtures of distinct behavioral components. While the bounds identify worst-case adversarial vulnerabilities rather than average user experiences, the theoretical derivations and accompanying empirical validations provide high confidence that static alignment cannot guarantee safety against determined adversarial prompting.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). This foundational work establishes instruction fine-tuning and reinforcement learning from human feedback as standard alignment techniques whose underlying safety limitations are theoretically analyzed in the source.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). It systematically investigates the conceptual failure modes of LLM safety training and adversarial prompt exploits, providing key empirical groundwork for the formal bounds developed in the source.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). It demonstrates automated, gradient-based adversarial suffix attacks on aligned LLMs, offering concrete empirical evidence for the vulnerability to adversarial prompting that the source models theoretically.
- Paper: Automatically Auditing Large Language Models via Discrete Optimization, Erik Jones et al. (2023). It introduces discrete optimization techniques to automatically audit and uncover undesired behaviors in language models, foundational to understanding how adversarial prompt triggers can be constructed.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). It highlights how fragile safety alignment is to minor perturbations and downstream modifications, motivating theoretical bounds on alignment retention.
- Paper: Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition, Sander Schulhoff et al. (2023). It provides large-scale empirical evidence of prompt hacking and systemic guardrail circumvention, contextualizing the fundamental limitations addressed in the source.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). It documents the baseline persistence of toxic degeneration in neural language models, which alignment procedures attempt to attenuate.
- Paper: What Makes and Breaks Safety Fine-tuning? A Mechanistic Study, Samyak Jain et al. (2024). It provides a complementary mechanistic investigation into how safety fine-tuning alters internal model representations and why those internal mechanisms fail under adversarial attacks.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). It introduces a standardized, comprehensive benchmarking framework to evaluate automated red-teaming and defense mechanisms against the alignment vulnerabilities formalized in the source.
- Paper: Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks, Maksym Andriushchenko et al. (2025). It extends the practical implications of alignment limitations by showing how simple adaptive attacks systematically jailbreak leading safety-aligned models.
- Paper: Improved Techniques for Optimization-Based Jailbreaking on Large Language Models, Xiaojun Jia et al. (2025). It builds upon optimization-based attack principles to deliver improved, highly reliable jailbreaks that exploit the theoretical weaknesses of alignment.
- Paper: Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM, Bochuan Cao et al. (2024). It develops a robust defense framework designed to prevent alignment-breaking prompt attacks by exploiting the fragility of adversarial inputs.
- Paper: COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability, Xingang Guo et al. (2024). It proposes controllable and fluent jailbreak generation methods to overcome standard perplexity defenses, demonstrating advanced exploits of alignment boundaries.
- Paper: Safe RLHF: Safe Reinforcement Learning from Human Feedback, Josef Dai et al. (2024). It explores safe reinforcement learning from human feedback by decoupling helpfulness and harmlessness objectives to mitigate alignment vulnerabilities.
- Paper: HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment, Shei Pern Chua et al. (2026). It applies internal representation analysis to couple harm recognition with refusal activation, offering a robust parameter-efficient alignment defense.
- Paper: Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing, Maryam Hashemzadeh et al. (2026). It proposes a dynamic routing architecture using expert adapters to rethink safety alignment beyond simple behavior erasure and refusal.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). It introduces decoding-time alignment mechanisms to dynamically enforce safety constraints, addressing the static limitations of fine-tuning identified in the source.
