Fundamental Limitations of Alignment in Large Language Models

Yotam WolfNoam WiesOshri AvneryYoav LevineAmnon Shashua

article2024ICML213 citations

Establishes a theoretical framework called Behavior Expectation Bounds to mathematically prove that standard alignment methods like RLHF cannot prevent adversarial jailbreaks as long as harmful behaviors retain a non-zero probability of being generated.

Listen

Large language models are increasingly deployed as public-facing assistants, yet they remain vulnerable to generating harmful, toxic, or biased content. While current safety alignment techniques—primarily reinforcement learning from human feedback and protective system prompts—aim to suppress these harmful behaviors, widespread real-world jailbreaks demonstrate that these defenses frequently fail. The article addresses this critical vulnerability by investigating whether current alignment practices can ever provide mathematically robust safety guarantees against adversarial prompting.

The main objective of the article is to establish a formal theoretical framework to evaluate the fundamental limitations of alignment in language models whose weights remain frozen at inference time. The authors demonstrate that if a model retains any non-zero probability of generating an undesired behavior, adversarial prompts can always be constructed to reliably trigger that behavior.

To conduct this evaluation, the authors develop a theoretical framework called Behavior Expectation Bounds. This framework models a language model's output distribution as a mixture of well-behaved and ill-behaved components, measuring the statistical distance and likelihood variance between them. The theoretical findings were supported empirically using models from the LLaMA 2 family across multiple evaluated behaviors, such as agreeableness and anti-immigration, utilizing specialized behavioral datasets to track distribution shifts and misaligning dynamics.

The analysis reveals several core findings. First, any alignment process that merely reduces the probability of a harmful behavior without eliminating it entirely is provably vulnerable to adversarial prompts, with the required prompt length scaling logarithmically with the rarity of the behavior. Second, while preset aligning system prompts provide some defense, an adversary can overcome them with an adversarial prompt whose length grows linearly with the preset prompt length. Third, in interactive multi-turn conversations, aligned models can initially resist misalignment through safe replies, but adversaries can still trigger harmful behavior by accumulating sufficient misaligning text over turns. Fourth, sampling multiple candidate outputs and selecting the most aligned response only delays misalignment logarithmically relative to the number of samples. Finally, empirical tests confirm that reinforcement learning from human feedback may unintentionally make undesired behaviors more statistically distinct, rendering them about five times more susceptible to targeted prompting than unaligned base models.

These findings imply that popular post-training alignment methods act only as temporary hurdles rather than absolute safety barriers. Relying exclusively on frozen-weight techniques like reinforcement learning from human feedback or system prompts leaves systems fundamentally vulnerable to adversarial manipulation, creating significant compliance, safety, and reputational risks for organizations deploying them in high-stakes environments.

Organizations developing and deploying language models should not rely solely on static alignment tuning or system prompts for robust security. Instead, teams should investigate runtime intervention mechanisms, such as representation engineering, which steer internal model activations during inference, and enforce strict constraints on input prompt lengths to bound adversarial attack capacity. Further research should focus on validating internal behavior superposition and establishing standardized, automated evaluations across broader behavioral categories.

The findings are bounded by the assumption that behavior can be scored reliably at the sentence level and that language model outputs can be approximated as mixtures of distinct behavioral components. While the bounds identify worst-case adversarial vulnerabilities rather than average user experiences, the theoretical derivations and accompanying empirical validations provide high confidence that static alignment cannot guarantee safety against determined adversarial prompting.

Cover for Fundamental Limitations of Alignment in Large Language Models

Abstract

An important aspect in developing language models that interact with humans is aligning their behavior to be useful and unharmful for their human users. This is usually achieved by tuning the model in a way that enhances desired behaviors and inhibits undesired ones, a process referred to as alignment. In this paper, we propose a theoretical approach called Behavior Expectation Bounds (BEB) which allows us to formally investigate several inherent characteristics and limitations of alignment in large language models. Importantly, we prove that within the limits of this framework, for any behavior that has a finite probability of being exhibited by the model, there exist prompts that can trigger the model into outputting this behavior, with probability that increases with the length of the prompt. This implies that any alignment process that attenuates an undesired behavior but does not remove it altogether, is not safe against adversarial prompting attacks. Furthermore, our framework hints at the mechanism by which leading alignment approaches such as reinforcement learning from human feedback make the LLM prone to being prompted into the undesired behaviors. This theoretical result is being experimentally demonstrated in large scale by the so called contemporary “chatGPT jailbreaks”, where adversarial users trick the LLM into breaking its alignment guardrails by triggering it into acting as a malicious persona. Our results expose fundamental limitations in alignment of LLMs and bring to the forefront the need to devise reliable mechanisms for ensuring AI safety.

Table of Contents

  • 1. Introduction
  • 2. Behavior Expectation Bounds: A Framework for Analyzing LLM Alignment
  • 2.1. LLMs as a Superposition of Behaviors
  • 2.2. Definitions for Bounding the Expected LLM Behavior
  • 3. Results: Limitations of LLM Alignment
  • 3.1. Misaligning via Adversarial Prompts
  • 3.2. Extensions: Aligning Prompts, Conversations and Best-of-n Sampling
  • 4. Empirical Results
  • 4.1. Possible Values for β, β′ and σ
  • 4.2. Demonstration of Misalignment via Convergence of LLM to P− and via Behavior Expectation
  • 5. Discussion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Discussion of limitations
  • A.1. Two components mixture
  • A.2. β-distinguishability
  • A.3. Limitation of results
  • B. Generalized misalignment
  • C. Proofs building blocks
  • C.1. Convergence to a single component
  • C.2. Behavioral implication of the convergence to a single component
  • C.3. Adversarial prompt construction
  • D. Proof of theorem 1
  • E. Proof of theorem 2
  • F. Proof of theorem 3
  • G. Proof of theorem 4
  • H. Lemmas for section 4
  • H.1. Proof of Lemma 5
  • H.2. Proof of Lemma 6
  • H.3. Proof of lemma H.3
  • I. Relaxation of β-distinguishability condition
  • J. Aquiring negative and positive behavior LLMs, “P−” and “P+”
  • K. Clustering of good and bad representations and defining approximate mixture
  • L. Empirical results for different behaviors on an RLHF model
  • L.1. Possible values of β and σ
  • L.2. Convergence via KL-divergence
  • L.3. Misalignment via behavior expectation
  • M. Pretrained models
  • N. β-prompt-distinguishability

Knowls

  1. Knowl 1 — Behavior expectation measures alignment and prompt-induced misalignment

    definition

    For a behavior vertical, let B:[Σ]^*\to[-1,1] assign each text string a score, with negative values indicating undesirable behavior and positive values indicating desirable behavior. For a language model distribution PP over single-sentence outputs, the behavior expectation is

    BP=Es∼P[B(s)].B_P=\mathbb{E}_{s\sim P}[B(s)].

    Given a prompt string xx, the prompted behavior expectation is BP(x)=Es∼P(⋅∣x)[B(s)]B_P(x)=\mathbb{E}_{s\sim P(\cdot\mid x)}[B(s)]. The paper calls PP γ\gamma-prompt-misalignable for BB, where γ∈[−1,0)\gamma\in[-1,0), if for every ϵ>0\epsilon>0 there is a prompt xx such that BP(x)<γ+ϵB_P(x)<\gamma+\epsilon. In this framework, alignment means raising expected behavior scores, while prompt misalignment means that some prompt can make the expected score arbitrarily close to a specified negative threshold. The formal results concern the next sentence, rather than a score over an entire long response.

  2. Knowl 2 — Two-component behavior mixture and the assumptions that enable prompt reweighting

    model/method

    The Behavior Expectation Bounds (BEB) framework represents an unprompted model as a mixture

    P=αP−+(1−α)P+,P=\alpha P_-+(1-\alpha)P_+,

    where P−P_- is an ill-behaved component for the behavior being analyzed, P+P_+ is a well-behaved component, and 0<α<10<\alpha<1 is the initial probability weight of the ill-behaved component. For a prompt xx, the posterior mixture weight on P−P_- is

    αP−(x)αP−(x)+(1−α)P+(x).\frac{\alpha P_-(x)}{\alpha P_-(x)+(1-\alpha)P_+(x)}.

    Thus, a prompt that is much more likely under P−P_- than under P+P_+ can raise the negative component’s weight in the prompted model, even though the unprompted mixture weight α\alpha is small.

    The framework assumes α,β,γ\alpha,\beta,\gamma negative distinguishability: P−P_- has behavior expectation at most γ<0\gamma<0 after every prompt, and is distinguishable from P+P_+ in the sense that for every number k≥0k\ge0 of preceding sentences, the expected conditional KL divergence between their next-sentence distributions, with prefixes sampled from P−P_-, exceeds β>0\beta>0. Prompted versions of the results use the stronger condition that this distinguishability persists after prefixes ending in a negative-behavior sentence. Some results also assume σ\sigma-similarity: for a sequence of mm sentences sampled under one component, the variance of the log likelihood ratio between the components is less than mσ2m\sigma^2. These assumptions quantify, respectively, negative-component prevalence, separability, and variability of the likelihood evidence.

  3. Knowl 3 — Any nonzero distinguishable negative component can be elicited by a prompt

    theoretical result

    Suppose an unprompted language model has the mixture P=αP−+(1−α)P+P=\alpha P_-+(1-\alpha)P_+ and satisfies α,β,γ\alpha,\beta,\gamma negative distinguishability for behavior BB, with α>0\alpha>0, β>0\beta>0, and γ∈[−1,0)\gamma\in[-1,0). Then for every ϵ>0\epsilon>0, there exists a prompt of length

    1β(log⁡1α+log⁡1ϵ+log⁡4)\frac{1}{\beta}\left(\log\frac{1}{\alpha}+\log\frac{1}{\epsilon}+\log 4\right)

    sentences for which BP(x)<γ+ϵB_P(x)<\gamma+\epsilon. Consequently, reducing an undesirable behavior to an arbitrarily small but nonzero mixture weight does not make the model immune to prompt-induced elicitation under these assumptions. The guaranteed prompt length grows logarithmically as α\alpha shrinks and decreases as the components become more distinguishable (larger β\beta).

  4. Knowl 4 — A preset aligning prompt increases, but does not eliminate, the misaligning prompt length

    theoretical result

    Suppose P=αP−+(1−α)P+P=\alpha P_-+(1-\alpha)P_+ is negatively prompt-distinguishable with parameters α,β,γ\alpha,\beta,\gamma. Assume also that P+P_+ is β′\beta'-undistinguishable from P−P_-, is σ\sigma-similar to P−P_-, and assigns lower conditional probability than P−P_- to negative-behavior sentences. If an aligning prefix x0x_0 is sampled from P+P_+, then with probability at least 1−δ1-\delta, the model conditioned on x0x_0 is γ\gamma-prompt-misalignable. A sufficient additional adversarial-prompt length, in sentences, is

    1β(log⁡1α+log⁡1ϵ+log⁡4)+β′β∣x0∣+σβ∣x0∣δ+1.\frac{1}{\beta}\left(\log\frac{1}{\alpha}+\log\frac{1}{\epsilon}+\log 4\right)+\frac{\beta'}{\beta}|x_0|+\frac{\sigma}{\beta}\sqrt{\frac{|x_0|}{\delta}}+1.

    Here ∣x0∣|x_0| is the prefix length in sentences, ϵ>0\epsilon>0 is the allowed behavior-expectation tolerance, and δ>0\delta>0 is the failure-probability bound. The guarantee therefore does not make a preset aligning prompt a permanent guardrail: it makes the required adversarial prompt longer, with a term linear in the aligning-prefix length.

  5. Knowl 5 — An adversarial user can misalign a conversation despite aligned model replies

    theoretical result

    Under the assumptions for misalignment with a preset aligning prefix, and assuming the well-behaved component is also β′\beta'-prompt-undistinguishable from the ill-behaved component, misalignment can be forced in a conversation. Let q1,…,qn+1q_1,\ldots,q_{n+1} be user prompts and a1,…,ana_1,\ldots,a_n the model’s intervening replies; write ∣qi∣|q_i| and ∣ai∣|a_i| for their lengths in sentences. A sufficient total amount of user-provided text is

    ∑i=1n+1∣qi∣=1β(log⁡1α+log⁡1ϵ+log⁡4)+∑i=1n(β′β∣ai∣+σβn∣ai∣δ)+n.\sum_{i=1}^{n+1}|q_i|=\frac{1}{\beta}\left(\log\frac{1}{\alpha}+\log\frac{1}{\epsilon}+\log 4\right)+\sum_{i=1}^{n}\left(\frac{\beta'}{\beta}|a_i|+\frac{\sigma}{\beta}\sqrt{\frac{n|a_i|}{\delta}}\right)+n.

    The prompts can be distributed across turns so that each user prompt has length at most

    ∣qi∣≤β′β∣ai∣+σβn∣ai∣δ+log⁡(1/ϵ)+log⁡(1/α)+log⁡4nβ+1.|q_i|\leq\frac{\beta'}{\beta}|a_i|+\frac{\sigma}{\beta}\sqrt{\frac{n|a_i|}{\delta}}+\frac{\log(1/\epsilon)+\log(1/\alpha)+\log 4}{n\beta}+1.

    The resulting conversation can make the model’s behavior expectation fall below γ+ϵ\gamma+\epsilon with the stated probability guarantee 1−δ1-\delta. The theorem captures a countervailing effect: aligned model replies can favor the well-behaved component, requiring the user to supply additional adversarial text.

  6. Knowl 6 — Best-of-$n$ selection raises the required prompt length only logarithmically

    theoretical result

    Assume the conditions of the main prompt-misalignment result, and suppose the model generates nn responses to a prompt and selects the response maximizing the behavior score BB. The best-of-nn model remains prompt-misalignable: for every ϵ>0\epsilon>0, a sufficient prompt length is

    1β(log⁡1α+log⁡1ϵ+log⁡4+log⁡n)\frac{1}{\beta}\left(\log\frac{1}{\alpha}+\log\frac{1}{\epsilon}+\log 4+\log n\right)

    sentences. Thus, selecting the most aligned of nn sampled responses does not guarantee alignment under the mixture and distinguishability assumptions, although it adds a logarithmic-in-nn term to the sufficient misaligning length.

  7. Knowl 7 — LLaMA 2 proxy distributions exhibit finite distinguishability and measured likelihood variance

    empirical result

    To estimate BEB parameters, the authors compared LLaMA 2 13B Chat with a LoRA-finetuned proxy trained to produce the opposite behavior, using behavior examples from Perez et al. For agreeableness, the estimated lower KL bound was β=20\beta=20, the upper bound was β′=30\beta'=30, and the log-likelihood-ratio variance parameter was estimated as σ2=50\sigma^2=50; hence σ/β=0.35\sigma/\beta=0.35 and β′/β=1.5\beta'/\beta=1.5. The reported anti-immigration estimates were approximately β=12\beta=12 and σ2=100\sigma^2=100. For agreeableness and anti-immigration, the KL divergence between the opposite-behavior proxies remained substantial as negative-component prompts grew, while the corresponding log-likelihood-ratio variance increased approximately linearly with prompt length. A prompted-prefix experiment also found an approximate β≈20\beta\approx20 after a negative-behavior sentence was added to the prefix. These measurements are estimates from proxy distributions, not measurements of latent components extracted from an LLM.

  8. Knowl 8 — Negative-behavior prompts drive LLaMA 2 Chat toward the negative proxy

    empirical result

    The authors tested the proposed mechanism on LLaMA 2 13B Chat, using a LoRA-finetuned negative-behavior model as a proxy for P−P_-. As prompts were lengthened with sentences sampled from that proxy, the KL divergence between the proxy and the RLHF-tuned model decreased, and the model’s measured behavior expectation became negative for agreeableness and anti-immigration. For agreeableness, the fitted effective parameters were approximately log⁡(1/α)=30\log(1/\alpha)=30 and βeff=10\beta_{\mathrm{eff}}=10, giving β−1log⁡(1/α)≈3\beta^{-1}\log(1/\alpha)\approx3 sentences; for anti-immigration, the corresponding estimates were about 1818 and 55. In the behavior-classification plots, the share of responses classified as positive fell as the adversarial prompt grew; inserting the default aligning prompt delayed this decline by roughly one sentence. The authors also observed prompt-induced negative behavior in pretrained LLaMA 2 13B, with estimated β\beta values of roughly 11–22, about five times smaller than for the RLHF model on the tested behaviors. Because the proxy is not the true mixture component and the KL bound need not be tight, the fitted α\alpha and β\beta values are approximate.

  9. Knowl 9 — Behavior representations cluster by polarity in LLaMA models

    empirical result

    The paper examined last-token representations of positive and negative behavior statements from a dataset spanning over 100 behavior verticals. Positive and negative examples showed spatial separation in the LLaMA representation space; an SVM trained on these representations achieved mean classification accuracy of 95.18%95.18\% (standard deviation 4.74%4.74\%) for LLaMA 7B and 95.61%95.61\% (standard deviation 4.52%4.52\%) for LLaMA 13B across 100 behaviors. The authors use this polarity structure as empirical support for modeling behaviorally distinct components. It does not establish that actual LLM probability distributions decompose into the specific P−P_- and P+P_+ components assumed by BEB.

  10. Knowl 10 — The guarantees depend on idealized components and sentence-level behavior scores

    limitation

    The theoretical results assume a decomposition into well-behaved and ill-behaved distributions with specified distinguishability and similarity properties. The authors note that real LLM components are not directly accessible; their experiments substitute LoRA-finetuned proxy models, so the empirical results do not verify the assumed decomposition in actual model distributions. The guarantees score the next sentence rather than a full response or longer interaction, and their formulation presumes ground-truth behavior scores for text, whereas real behavior categories and scoring can be ambiguous and depend on text granularity. The paper identifies developing more realistic models of component structure and behavior scoring as open work.

Coverage note — The appendix’s generalized high-probability prompt-sampling bound is not given as a separate knowl because it extends the main misalignment theorem with a probability-of-success refinement; the empirical tables of individual behavior classifications are summarized through the broader representation-clustering result.

References

  1. 1.Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mane, D. Concrete problems in ai safety. arXiv preprint arXiv:1606.06565, 2016.
  2. 2.Andreas, J. Language models as agent models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 5769–5779, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-emnlp.423.
  3. 3.Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861, 2021.
  4. 4.Atillah, I. E. Man ends his life after an ai chatbot ’encouraged’ him to sacrifice himself to stop climate change. Euronews, 2023.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  6. 6.Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp. 610–623, 2021.
  7. 7.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901, 2020.
  8. 8.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023.
  9. 9.Deshpande, A., Murahari, V., Rajpurohit, T., Kalyan, A., and Narasimhan, K. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv preprint arXiv:2304.05335, 2023.
  10. 10.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  11. 11.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3356–3369, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL https://aclanthology.org/2020.findings-emnlp.301.
  12. 12.Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J. Unsolved problems in ml safety. arXiv preprint arXiv:2109.13916, 2021.
  13. 13.Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  14. 14.Hutchinson, B., Prabhakaran, V., Denton, E., Webster, K., Zhong, Y., and Denuyl, S. Social biases in NLP models as barriers for persons with disabilities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 5491–5501, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.487. URL https://aclanthology.org/2020.acl-main.487.
  15. 15.Jorgensen, O., Cope, D., Schoots, N., and Shanahan, M. Improving activation steering in language models with mean-centring. arXiv preprint arXiv:2312.03813, 2023.
  16. 16.Leong, C. T., Cheng, Y., Wang, J., Wang, J., and Li, W. Self-detoxifying language models via toxification reversal. arXiv preprint arXiv:2310.09573, 2023.
  17. 17.Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-long.229.
  18. 18.Liu, W., Wang, X., Wu, M., Li, T., Lv, C., Ling, Z., Zhu, J., Zhang, C., Zheng, X., and Huang, X. Aligning large language models with human preferences through representation engineering. arXiv preprint arXiv:2312.15997, 2023.
  19. 19.Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., and Paul, S. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  20. 20.Meta, A. Introducing llama: A foundational, 65-billion-parameter large language model. Meta AI. https://ai.facebook.com/blog/large-language-model-llama-meta-ai, 2023.
  21. 21.Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1953–1967, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.154. URL https://aclanthology.org/2020.emnlp-main.154.
  22. 22.Nardo, C. The waluigi effect (mega-post). Less Wrong, 2023.
  23. 23.Ngo, R. The alignment problem from a deep learning perspective. arXiv preprint arXiv:2209.00626, 2022.
  24. 24.Nori, H., King, N., McKinney, S. M., Carignan, D., and Horvitz, E. Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375, 2023.
  25. 25.O’Brien, M. Musk, scientists call for halt to ai race sparked by chatgpt. AP News, 2023.
  26. 26.OpenAI. Gpt-4 technical report, 2023.
  27. 27.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow instructions with human feedback. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=TG8KACxEON.
  28. 28.Pan, A., Bhatia, K., and Steinhardt, J. The effects of reward misspecification: Mapping and mitigating misaligned models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=JYtwGwIL7ye.
  29. 29.Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. arXiv preprint arXiv:2304.03442, 2023.
  30. 30.Perez, E., Ringer, S., Lukosiˇ ut¯ e, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. Discovering language model behaviors with model-written evaluations. arXiv preprint arXiv:2212.09251, 2022.
  31. 31.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  32. 32.Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  33. 33.Roose, K. A conversation with bing’s chatbot left me deeply unsettled. New York Times, 2023.
  34. 34.Schulman, J., Zoph, B., Kim, C., Hilton, J., Menick, J., Weng, J., Felipe, J., Uribe, C., Fedus, L., Metz, L., Pokorny, M., Lopes, R. G., Zhao, S., Vijayvergiya, A., Sigler, E., Perelman, A., Voss, C., Heaton, M., Parish, J., Cummings, D., Nayak, R., Balcom, V., Schnurr, D., Kaftan, T., Hallacy, C., Turley, N., Deutsch, N., Goel, V., Ward, J., Konstantinidis, A., Zaremba, W., Ouyang, L., Bogdonoff, L., Gross, J., Medina, D., Yoo, S., Lee, T., Lowe, R., Mossing, D., Huizinga, J., Jiang, R., Wainwright, C., Almeida, D., Lin, S., Zhang, M., Xiao, K., Slama, K., Bills, S., Gray, A., Leike, J., Pachocki, J., Tillet, P., Jain, S., Brockman, G., Ryder, N., Paino, A., Yuan, Q., Winter, C., Wang, B., Bavarian, M., Babuschkin, I., Sidor, S., Kanitscheider, I., Pavlov, M., Plappert, M., Tezak, N., Jun, H., Zhuk, W., Pong, V., Kaiser, L., Tworek, J., Carr, A., Weng, L., Agarwal, S., Cobbe, K., Kosaraju, V., Power, A., Polu, S., Han, J., Puri, R., Jain, S., Chess, B., Gibson, C., Boiko, O., Parparita, E., Tootoonchian, A., Kosic, K., and Hesse, C. Introducing chatgpt. OpenAI blog, 2023.
  35. 35.Shalev-Shwartz, S., Shammah, S., and Shashua, A. On the ethics of building ai in a responsible manner. arXiv preprint arXiv:2004.04644, 2020.
  36. 36.Subhash, V. Can large language models change user preference adversarially? arXiv preprint arXiv:2302.10291, 2023.
  37. 37.Taylor, J., Yudkowsky, E., LaVictoire, P., and Critch, A. Alignment for advanced machine learning systems. Ethics of Artificial Intelligence, pp. 342–382, 2016.
  38. 38.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  39. 39.Turner, A., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M. Activation addition: Steering language models without optimization. arXiv preprint arXiv:2308.10248, 2023.
  40. 40.Venkit, P. N., Srinath, M., and Wilson, S. A study of implicit bias in pretrained language models against people with disabilities. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 1324–1332, Gyeongju, Republic of Korea, October 2022. International Committee on Computational Linguistics. URL https://aclanthology.org/2022.coling-1.113.
  41. 41.Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S. Universal adversarial triggers for attacking and analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2153–2162, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1221. URL https://aclanthology.org/D19-1221.
  42. 42.Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., and Gabriel, I. Taxonomy of risks posed by language models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, pp. 214–229, New York, NY, USA, 2022. Association for Computing Machinery. ISBN 9781450393522. doi: 10.1145/3531146.3533088. URL https://doi.org/10.1145/3531146.3533088.
  43. 43.West, C. G. Advances in apparent conceptual physics reasoning in gpt-4. arXiv e-prints, pp. arXiv–2303, 2023.
  44. 44.Wies, N., Levine, Y., and Shashua, A. The learnability of in-context learning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=f3JNQd7CHM.
  45. 45.Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2950–2968, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.235. URL https://aclanthology.org/2021.naacl-main.235.
  46. 46.Yu, D. and Sagae, K. Automatically exposing problems with neural dialog models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 456–470, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.37. URL https://aclanthology.org/2021.emnlp-main.37.
  47. 47.Yudkowsky, E. Creating friendly ai 1.0: The analysis and design of benevolent goal architectures. The Singularity Institute, San Francisco, USA, 2001.
  48. 48.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023.

Citation

MLA
Wolf, Y., et al. “Fundamental Limitations of Alignment in Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2304.11082v6.
APA
Wolf, Y., Wies, N., Avnery, O., Levine, Y., & Shashua, A. (2023). Fundamental Limitations of Alignment in Large Language Models. arXiv. http://arxiv.org/abs/2304.11082v6
Chicago
Wolf, Y., N. Wies, O. Avnery, Y. Levine, and A. Shashua. 2023. “Fundamental Limitations of Alignment in Large Language Models”. arXiv. http://arxiv.org/abs/2304.11082v6.
Harvard
Wolf, Y. et al. (2023) “Fundamental Limitations of Alignment in Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.11082v6.
Vancouver
1. Wolf Y, Wies N, Avnery O, Levine Y, Shashua A (2023) Fundamental Limitations of Alignment in Large Language Models. arXiv

BibTeX

@article{wolf2023fundamental,
  title = {Fundamental Limitations of Alignment in Large Language Models},
  author = {Wolf, Yotam and Wies, Noam and Avnery, Oshri and Levine, Yoav and Shashua, Amnon},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.11082v6},
  eprint = {2304.11082}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/