keyword
behavior expectation bounds
Behavior expectation bounds is a theoretical framework in artificial intelligence safety used to mathematically analyze the characteristics and fundamental limits of aligning large language models with intended behavioral standards. The framework formalizes alignment by evaluating the expected score of a model's responses across different output distributions and prompts. Within this structure, it proves that if a model retains any non-zero probability of exhibiting an undesirable behavior, there exist adversarial prompts whose length enables the elicitation of that behavior with high probability. As a result, the framework demonstrates that alignment methods that merely suppress or lower the likelihood of unsafe outputs, rather than completely removing the underlying behavioral modes, remain inherently vulnerable to adversarial jailbreaks and targeted prompting attacks.
1 item

