A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Andrew LeeXiaoyan BaiItamar PresMartin WattenbergJonathan K. KummerfeldRada Mihalcea

article2024ICML212 citations

Reveals that Direct Preference Optimization merely bypasses rather than eliminates toxic capabilities in language models, providing a mechanistic explanation for safety jailbreaks and enabling a simple method to reverse alignment.

Listen

Large language models often absorb toxic and biased behavior during pre-training on massive web datasets. While alignment techniques—specifically Direct Preference Optimization (DPO)—are widely deployed to steer models toward safe and desirable behavior, the internal mechanisms governing how these models suppress undesirable outputs remain poorly understood. This lack of transparency poses major safety concerns, especially given that aligned models can often be easily compromised or jailbroken.

The main objective of the article is to mechanistically evaluate how pre-trained models represent toxicity and how DPO alters internal model representations to avert toxic generations. Specifically, the article demonstrates whether alignment removes undesirable capabilities entirely or merely suppresses their expression.

To conduct this evaluation, the researchers analyzed two pre-trained models, GPT2-medium and Llama2-7b. They first trained a linear probe on the Jigsaw toxicity dataset (over 560,000 comments) to identify specific toxic vectors within the models' multi-layer feed-forward networks. Next, they aligned both models using DPO on a paired dataset of 24,576 toxic and non-toxic continuations generated from Wikipedia prompts. Finally, they evaluated changes in internal weights, activation paths, toxicity scores, and standard performance metrics using challenging toxicity evaluation prompts.

The investigation produced four central findings. First, alignment does not delete toxic capabilities; every model parameter exhibited a cosine similarity exceeding 0.99 before and after alignment, confirming that toxic vectors remain physically present within the model weights. Second, DPO reduces toxicity by steering internal processing away from these toxic vectors rather than modifying them. In GPT2, the model distributes subtle offsets across earlier layers to bypass toxic activation spaces, whereas in Llama2, internal gating mechanisms simply scale down and turn off toxic components. Third, DPO successfully reduced baseline toxicity from 0.453 to 0.208 in GPT2 and from 0.359 to 0.138 in Llama2, while preserving core language modeling quality and coherence. Fourth, because the underlying toxic circuits remain intact, alignment is exceptionally fragile: scaling up as few as 7 key vectors in GPT2 or toggling 8 gating units in Llama2 immediately reactivated toxic outputs, restoring toxicity to pre-alignment levels (0.458 and 0.244, respectively).

These findings provide direct mechanistic insight into why safety alignments are vulnerable to jailbreaks and adversarial fine-tuning. Alignment acts merely as a soft bypass rather than a permanent removal of harmful capabilities. For decision-makers and risk leaders, this demonstrates that relying solely on preference-based optimization leaves substantial latent security, safety, and compliance risks intact across deployed models.

Based on these results, the article suggests exploring more robust alignment architectures. Promising future directions include directly removing causal pathways responsible for unsafe behavior, selectively updating only problematic weights during alignment, or incorporating dedicated late-layer suppression heads. Before deploying models in high-risk environments, teams should test model vulnerability by auditing internal activation regions rather than relying strictly on standard output evaluations.

The study's primary limitation is its focus on toxicity within two specific architectures (GPT2 and Llama2) trained via DPO. Consequently, confidence is high regarding preference-based optimization mechanisms in similar transformer architectures, but caution is warranted when generalizing these conclusions to other alignment algorithms, fine-tuning setups, or broader safety domains without further empirical validation.

arXiv: 2401.01967
Cover for A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Abstract

While alignment algorithms are commonly used to tune pre-trained language models towards user preferences, we lack explanations for the underlying mechanisms in which models become “aligned”, thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in pre-trained language models (GPT2-medium, Llama2-7b). We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting models avert toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the models, reverting them back to their toxic behavior.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries
  • 3. Toxicity in Pre-trained Language Models
  • 3.1. Extracting Toxic Vectors
  • 3.2. Toxic Vectors in Vocabulary space
  • 3.3. Interventions Using Toxic Vectors
  • 4. Toxicity Alignment Using DPO
  • 4.1. Background: DPO
  • 4.2. Constructing Pairwise Toxic Data
  • 5. Toxicity After DPO
  • 5.1. Toxic Vectors Remain After DPO
  • 5.2. DPO Avoids MLP-k Toxic Regions
  • 6. Un-aligning DPO
  • 7. Discussion
  • 7.1. On Designing Robust Alignment Algorithms
  • 7.2. On the Role of KL-Divergence Regularization
  • 8. Related Work
  • 8.1. Alignment Algorithms
  • 8.2. Mechanistic Interpretability
  • 8.3. Jailbreaking Aligned Models
  • 9. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • A. Projecting Value Vectors onto Vocabulary Space
  • B. Additional Llama2 Results
  • C. Shift in Residual Streams
  • D. Shifts in Residual Streams vs. Shifts in MLP Value Vectors
  • E. Hyperparameters

Knowls

  1. Knowl 1 — GPT-2 DPO learns a distributed offset that bypasses toxic activation regions

    empirical result

    In GPT-2, DPO reduced toxic generations without deleting the feed-forward value vectors that promote toxic tokens. For a feed-forward key vector kiℓ∈Rdk_i^\ell\in\mathbb{R}^d at layer ℓ\ell, the paper defines its activation region as γ(kiℓ)={g∈Rd:σ(kiℓ⋅g)>0}\gamma(k_i^\ell)=\{g\in\mathbb{R}^d:\sigma(k_i^\ell\cdot g)>0\}, where gg is a residual-stream vector and σ\sigma is the model’s activation function. On prompts that elicited toxic outputs before alignment, the post-DPO residual stream shifted away from regions associated with toxic key vectors, lowering the activations of their corresponding value vectors.

    The authors attribute this shift to small changes spread across many earlier feed-forward value vectors. For a toxic vector at layer ℓ\ell, the shifts in preceding layers’ value vectors were often oriented opposite to the residual-stream shift. They report that these vectors usually had negative activations on the evaluated prompts; multiplying their changes by these negative activations makes their net contributions point toward the residual-stream shift. In GPT-2, this distributed offset lets the model avoid toxic-triggering regions while leaving the original toxic-generating capability in its weights.

  2. Knowl 2 — Llama 2 DPO suppresses toxic vectors through GLU components

    empirical result

    In Llama 2, DPO did not show the same distributed residual-stream bypass mechanism reported for GPT-2. Instead, the authors found that toxic feed-forward value vectors in Llama 2’s gated linear units (GLUs) were less active after DPO. A GLU scales its value-vector contributions using both a nonlinear gate, σ(W1x)\sigma(W_1x), and a linear projection, W2xW_2x, where xx is the residual-stream input and W1,W2W_1,W_2 are learned transformations. For examined toxic vectors, the mean activations of both components decreased after DPO. The authors therefore interpret Llama 2’s mechanism as turning toxic vectors off through its GLU components rather than removing the vectors.

  3. Knowl 3 — Toxic behavior can be reactivated in the DPO-aligned models

    empirical result

    The paper reactivated toxic behavior in both aligned models by changing the components that had made toxic vectors less likely to contribute. For GPT-2, the authors selected seven toxic key vectors with the highest cosine similarity to the toxicity probe and scaled each by 10×10\times. This increased their activation regions and raised the toxicity score from 0.208 for GPT-2-DPO to 0.458, with perplexity changing from 23.34 to 23.30 and F1 from 0.195 to 0.195. The unaligned GPT-2 baseline scored 0.453 toxicity, 21.7 perplexity, and 0.193 F1.

    For Llama 2, setting eight toxic-vector gate values to 1 raised toxicity from 0.138 to 0.217; perplexity changed from 6.587 to 6.596 and F1 from 0.194 to 0.195. Alternatively, scaling the relevant W2xW_2x component by 3×3\times raised toxicity to 0.244, with perplexity 6.648 and F1 0.194. The unaligned Llama 2 baseline scored 0.359 toxicity, 6.095 perplexity, and 0.227 F1. These interventions demonstrate that the aligned models’ toxic behavior could be recovered by restoring activation of retained toxic components.

  4. Knowl 4 — DPO leaves toxic-generating parameters nearly unchanged

    empirical result

    For GPT-2 and Llama 2, corresponding parameters before and after DPO had cosine similarity greater than 0.99, and the average parameter norm difference was less than 10−510^{-5}. This result also applied to the identified toxic key and value vectors: the toxic vectors themselves were not substantially changed by alignment. The GPT-2 unembedding layer was the stated exception to the norm-difference result, with a difference below 10−310^{-3}. Together with the reactivation experiments, these measurements support the paper’s account that DPO changes how toxic-generating components are used rather than erasing those components.

  5. Knowl 5 — DPO reduces toxicity scores in GPT-2 and Llama 2

    empirical result

    On the paper’s toxicity evaluation, DPO lowered the toxicity score for both model families. GPT-2 changed from 0.453 toxicity, 21.7 perplexity, and 0.193 F1 before DPO to 0.208 toxicity, 23.34 perplexity, and 0.195 F1 after DPO. Llama 2 changed from 0.359 toxicity, 6.095 perplexity, and 0.227 F1 to 0.138 toxicity, 6.587 perplexity, and 0.194 F1. Toxicity scores were assigned by Perspective API on prompts selected to elicit toxic outputs; perplexity was measured on WikiText-2; and F1 measured token overlap with Wikipedia continuations. The observed reduction in toxicity came with a modest perplexity increase in each model and a lower F1 for Llama 2.

  6. Knowl 6 — Toxic feed-forward vectors are identified using a linear toxicity probe

    model/method

    The authors trained a linear binary toxicity classifier on residual-stream representations from the final layer of GPT-2 and Llama 2. Representations were averaged across input timesteps, and the probe was trained on the Jigsaw toxic comment classification dataset, containing 561,808 labeled comments, with a 90:10 training-validation split. The probe achieved 94% validation accuracy. The authors treated its learned toxicity direction as an aggregate of signals used to classify toxic text, then searched feed-forward value vectors for high cosine similarity to that direction. They identified a set of 128 such vectors per model and applied singular value decomposition (SVD) to the stacked vectors to obtain directions they interpreted as a basis for dimensions of the models’ toxicity representations.

  7. Knowl 7 — Subtracting toxic directions suppresses toxic generations

    empirical result

    During generation, the authors subtracted a scaled toxicity-related direction from the final-layer residual stream and chose the scale so perplexity was comparable to that of the corresponding DPO-aligned model. For GPT-2, the unmodified model scored 0.453 toxicity, 21.7 perplexity, and 0.193 F1. Subtracting the toxicity probe direction yielded 0.245, 23.56, and 0.193; subtracting the value vector MLP.v at layer 19, index 770, yielded 0.305, 23.30, and 0.192; subtracting the first SVD direction yielded 0.268, 23.48, and 0.193.

    For Llama 2, the unmodified model scored 0.359 toxicity, 6.095 perplexity, and 0.227 F1. Subtracting its toxicity probe direction yielded 0.256, 6.523, and 0.225; subtracting GLU value vector 19:5447 yielded 0.171, 6.518, and 0.225; subtracting the first SVD direction yielded 0.246, 6.504, and 0.225. The results support a functional role for the identified directions: subtracting them reduced toxicity while keeping F1 close to the unmodified-model values.

  8. Knowl 8 — Pairwise toxicity data and DPO training procedure

    experimental setup

    The authors constructed preference pairs from WikiText-2 sentences used as prompts. GPT-2 generated preferred, nontoxic continuations using greedy sampling; PPLM generated nonpreferred, toxic continuations guided by the toxicity probe. The resulting dataset contained 24,576 pairs. DPO training used a beta of 0.1, learning rate 10−610^{-6}, batch size 4, RMSProp, one gradient-accumulation step, and maximum gradient norm 10. Validation used loss with patience 10; training reportedly converged after approximately 6,700 sample pairs.

    PPLM generation used step size 0.4, temperature 1, top-kk 10, 50 iterations, window length 0, horizon length 1, decay disabled, gamma 1, GM scale 0.95, and KL scale 0.1. The paper applied DPO to study toxicity alignment in GPT-2 and Llama 2.

  9. Knowl 9 — Toxicity directions capture varied lexical dimensions

    empirical result

    Projecting the identified feed-forward value vectors into vocabulary space showed that different vectors promoted different kinds of toxic-associated tokens. The GPT-2 examples included profanity, insults and disparagement, shame-related terms, and sexual vocabulary; the Llama 2 examples included profanity and insults as well as sexual terms. The SVD directions also had distinct lexical rankings: the paper reports that GPT-2’s second SVD direction was particularly gendered, which the authors attribute to the dataset and model used. These projections suggest that the extracted vectors represent multiple contexts or dimensions associated with toxicity rather than a single undifferentiated feature.

  10. Knowl 10 — Evaluation protocol for toxicity alignment and generation quality

    experimental setup

    Toxicity was evaluated by prompting models with the 1,199-prompt challenge subset of RealToxicityPrompts, chosen because it elicits extremely toxic outputs, and scoring generations with Perspective API. Perplexity was measured on WikiText-2. Generation quality was also evaluated using 2,000 Wikipedia sentences as prompts: precision was the fraction of generated tokens present in the reference Wikipedia continuation, recall was the fraction of reference continuation tokens present in the generation, and F1 was the harmonic mean of precision and recall. The same metrics were used to compare unmodified, DPO-aligned, and intervention-modified models.

Coverage note — The appendix’s additional layer-by-layer residual-shift plots and exhaustive token rankings are omitted because they reinforce the reported mechanisms without adding a distinct main result.

References

  1. 1.Adolphs, L., Gao, T., Xu, J., Shuster, K., Sukhbaatar, S., and Weston, J. The CRINGE loss: Learning what language not to model. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8854–8874, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.493. URL https://aclanthology.org/2023.acl-long.493.
  2. 2.Balestriero, R., Cosentino, R., and Shekkizhar, S. Characterizing large language model geometry solves toxicity detection and generation. arXiv preprint arXiv:2312.01648, 2023.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
  4. 4.Burns, C., Ye, H., Klein, D., and Steinhardt, J. Discovering latent knowledge in language models without supervision. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ETKGuby0hcs.
  5. 5.Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W., Ippolito, D., Tramer, F., and Schmidt, L. Are aligned neural networks adversarially aligned? In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=OQQoD8Vc3B.
  6. 6.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1–113, 2023.
  7. 7.cjadams, Sorensen, J., Elliott, J., Dixon, L., McDonald, M., nithum, and , Cukierski, W. Toxic comment classification challenge, 2017. URL https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge.
  8. 8.Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Gurevych, I. and Miyao, Y. (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2126–2136, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1198. URL https://aclanthology.org/P18-1198.
  9. 9.Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations, 2019.
  10. 10.Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D. Language modeling with gated convolutional networks. In International conference on machine learning, pp. 933–941. PMLR, 2017.
  11. 11.Dinan, E., Logacheva, V., Malykh, V., Miller, A., Shuster, K., Urbanek, J., Kiela, D., Szlam, A., Serban, I., Lowe, R., et al. The second conversational intelligence challenge (convai2). In The NeurIPS’18 Competition: From Machine Learning to Intelligent Conversations, pp. 187–208. Springer, 2020.
  12. 12.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. https://transformer-circuits.pub/2021/framework/index.html.
  13. 13.Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Cohn, T., He, Y., and Liu, Y. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 3356–3369, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.301. URL https://aclanthology.org/2020.findings-emnlp.301.
  14. 14.Geva, M., Schuster, R., Berant, J., and Levy, O. Transformer feed-forward layers are key-value memories. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 5484–5495, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.446. URL https://aclanthology.org/2021.emnlp-main.446.
  15. 15.Geva, M., Caciularu, A., Wang, K., and Goldberg, Y. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Goldberg, Y., Kozareva, Z., and Zhang, Y. (eds.), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 30–45, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.emnlp-main.3. URL https://aclanthology.org/2022.emnlp-main.3.
  16. 16.Geva, M., Bastings, J., Filippova, K., and Globerson, A. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767, 2023.
  17. 17.Hendel, R., Geva, M., and Globerson, A. In-context learning creates task vectors. In Bouamor, H., Pino, J., and Bali, K. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 9318–9333, Singapore, December 2023. Association for Computational Linguistics. URL https://aclanthology.org/2023.findings-emnlp.624.
  18. 18.Hendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  19. 19.Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D. Linearity of relation decoding in transformer language models. arXiv preprint arXiv:2308.09124, 2023.
  20. 20.Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations, 2022.
  21. 21.Jain, S., Kirk, R., Lubana, E. S., Dick, R. P., Tanaka, H., Grefenstette, E., Rocktaschel, T., and Krueger, D. S. Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks. arXiv preprint arXiv:2311.12786, 2023.
  22. 22.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  23. 23.Li, K., Hopkins, A. K., Bau, D., Viegas, F., Pfister, H., and Wattenberg, M. Emergent world representations: Exploring a sequence model trained on a synthetic task. In The Eleventh International Conference on Learning Representations, 2023a. URL https://openreview.net/forum?id=DeG07_TcZvT.
  24. 24.Li, K., Patel, O., Viegas, F., Pfister, H., and Wattenberg, M. Inference-time intervention: Eliciting truthful answers from a language model. arXiv preprint arXiv:2306.03341, 2023b.
  25. 25.Li, M., Davies, X., and Nadeau, M. Circuit breaking: Removing model behaviors with targeted ablation. arXiv preprint arXiv:2309.05973, 2023c.
  26. 26.Li, Z., You, C., Bhojanapalli, S., Li, D., Rawat, A. S., Reddi, S. J., Ye, K., Chern, F., Yu, F., Guo, R., and Kumar, S. The lazy neuron phenomenon: On emergence of activation sparsity in transformers. In The Eleventh International Conference on Learning Representations, 2023d. URL https://openreview.net/forum?id=TJ2nxciYCk-.
  27. 27.Meng, K., Bau, D., Andonian, A. J., and Belinkov, Y. Locating and editing factual associations in GPT. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=-h6WAS6eE4.
  28. 28.Merity, S., Xiong, C., Bradbury, J., and Socher, R. Pointer sentinel mixture models. In International Conference on Learning Representations, 2016.
  29. 29.Nanda, N., Lee, A., and Wattenberg, M. Emergent linear representations in world models of self-supervised sequence models. In Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pp. 16–30, 2023.
  30. 30.Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramer, F., and Lee, K. Scalable extraction of training data from (production) language models. arXiv preprint arXiv:2311.17035, 2023.
  31. 31.Nostalgebraist. Interpreting gpt: The logit lens, 2020. URL https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens.
  32. 32.Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. Zoom in: An introduction to circuits. Distill, 2020. doi: 10.23915/distill.00024.001. https://distill.pub/2020/circuits/zoom-in.
  33. 33.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., et al. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, 2022.
  34. 34.Park, K., Choe, Y. J., and Veitch, V. The linear representation hypothesis and the geometry of large language models. In Causal Representation Learning Workshop at NeurIPS 2023, 2023.
  35. 35.Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
  36. 36.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model, 2023.
  37. 37.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  38. 38.Shazeer, N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  39. 39.Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N. The woman worked as a babysitter: On biases in language generation. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3407–3412, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1339. URL https://aclanthology.org/D19-1339.
  40. 40.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021, 2020.
  41. 41.Tenney, I., Das, D., and Pavlick, E. BERT rediscovers the classical NLP pipeline. In Korhonen, A., Traum, D., and Márquez, L. (eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4593–4601, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1452. URL https://aclanthology.org/P19-1452.
  42. 42.Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. Function vectors in large language models. arXiv preprint arXiv:2310.15213, 2023.
  43. 43.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  44. 44.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I. Attention is all you need. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf.
  45. 45.Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S. Universal adversarial triggers for attacking and analyzing NLP. In Inui, K., Jiang, J., Ng, V., and Wan, X. (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2153–2162, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1221. URL https://aclanthology.org/D19-1221.
  46. 46.Wei, A., Haghtalab, N., and Steinhardt, J. Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=jA235JGM09.
  47. 47.Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J. Neural text generation with unlikelihood training. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SJeYe0NtvH.
  48. 48.Xu, J., Lee, A., Sukhbaatar, S., and Weston, J. Some things are more cringe than others: Preference optimization with the pairwise cringe loss. arXiv preprint arXiv:2312.16682, 2023.
  49. 49.Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D. Shadow alignment: The ease of subverting safely-aligned language models. arXiv preprint arXiv:2310.02949, 2023.
  50. 50.Zhang, Z., Lin, Y., Liu, Z., Li, P., Sun, M., and Zhou, J. MoEfication: Transformer feed-forward layers are mixtures of experts. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Findings of the Association for Computational Linguistics: ACL 2022, pp. 877–890, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.71. URL https://aclanthology.org/2022.findings-acl.71.
  51. 51.Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023a.
  52. 52.Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023b.

Citation

MLA
Lee, A., et al. “A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity”. arXiv, 2024, http://arxiv.org/abs/2401.01967v1.
APA
Lee, A., Bai, X., Pres, I., Wattenberg, M., Kummerfeld, J. K., & Mihalcea, R. (2024). A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv. http://arxiv.org/abs/2401.01967v1
Chicago
Lee, A., X. Bai, I. Pres, M. Wattenberg, J. K. Kummerfeld, and R. Mihalcea. 2024. “A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity”. arXiv. http://arxiv.org/abs/2401.01967v1.
Harvard
Lee, A. et al. (2024) “A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.01967v1.
Vancouver
1. Lee A, Bai X, Pres I, Wattenberg M, Kummerfeld JK, Mihalcea R (2024) A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv

BibTeX

@article{lee2024mechanistic,
  title = {A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity},
  author = {Lee, Andrew and Bai, Xiaoyan and Pres, Itamar and Wattenberg, Martin and Kummerfeld, Jonathan K. and Mihalcea, Rada},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.01967v1},
  eprint = {2401.01967}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/