Word Embeddings Are Steers for Language Models

Chi HanJialiang XuManling LiYi FungChenkai SunNan JiangTarek F. AbdelzaherHeng Ji

article2024ACL93 citations

Proposes LM-Steer, a lightweight approach that applies linear transformations to output word embeddings to achieve parameter-efficient, transferable, and continuous style control across language models without retraining internal weights.

Listen

Large language models often exhibit undesirable generation behaviors, including producing toxic language, social biases, or off-target styles inherited from their pre-training data. Controlling these behaviors typically requires computationally expensive model retraining, large-scale fine-tuning, or slow, complex external classifier guidance at runtime. The article investigates the theoretical and practical role of output word embeddings in model generation and introduces LM-Steer, an extremely lightweight method that steers generation style and safety by applying a simple linear transformation to output word embeddings.

The authors evaluated the approach across multiple open-source language model families ranging from 14 million to 7 billion parameters, including GPT-2, Pythia, GPT-J, and Llama-2. Training was conducted using standard benchmark datasets for detoxification and sentiment control. The primary assessments measured reduction in toxic output, sentiment polarity adherence, generation fluency (measured by text perplexity), generation diversity, and decoding speed relative to leading controlled-generation and fine-tuning baselines.

The findings establish that LM-Steer consistently achieves superior or competitive control compared to existing baselines while preserving text quality. On the language detoxification benchmark, LM-Steer reduced average maximum toxicity by more than 6 absolute percentage points relative to strong baselines while maintaining high fluency and diversity. The method proved exceptionally resource efficient: it requires training only about 0.2% of the original model parameters (less than one-tenth the parameters of low-rank adaptation, or LoRA) and achieves strong detoxification with as few as 30 training examples. Furthermore, LM-Steer allows continuous intensity adjustment, compositional multi-attribute control (such as simultaneously adjusting sentiment and toxicity), and direct transfer to other language models via an explicit mathematical transformation without retraining.

These results provide a low-cost, low-latency mechanism to enhance AI safety and content moderation in production environments. Because LM-Steer operates as an efficient output layer transformation, it incurs minimal computational overhead during decoding, making it practical for real-time deployment without altering core model parameters. Additionally, decomposing the learned steer matrix offers interpretability by highlighting specific vocabulary dimensions and text spans driving stylistic or toxic attributes.

Decision-makers should consider piloting LM-Steer for rapid, cost-effective safety alignment and style customization on self-hosted or open-source models, especially where training budgets or inference latencies are constrained. However, stakeholders should note key limitations: LM-Steer primarily operates at the lexical and wording level and is not designed for complex multi-step reasoning or deep syntactic structural changes. In addition, deployment is restricted to environments with direct access to model embedding weights, precluding its direct use with closed, black-box third-party model APIs.

Cover for Word Embeddings Are Steers for Language Models

Abstract

Recent studies have shown that large language models (LLMs), when augmented with additional weights through adapters—for instance, LoRA—can learn specialized tasks without updating model parameters—a paradigm known as either Instruction Tuning (IT) for individual instructions or Task Arithmetic (TA) for task-specific adaptation. In such scenarios, different instruction/task datasets induce distinct steering vectors within LLMs' embedding space, shaping behavior according to downstream objectives. Inspired by recent advances in model merging or parameter composition, we investigate whether multiple steering directions (each optimized toward generating responses on one dataset while discouraging outputs outside its distribution via unlearning-like objectives) can be combined into a single vector enabling zero-shot generalization across diverse tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 LM-Steer: Revealing Hidden Steers in Word Embeddings
  • 4 Steering Language Model Generation
  • 4.1 Language Detoxification
  • 4.2 Sentiment Control
  • 4.3 Continuous and Compositional Control
  • 4.4 Efficiency
  • 5 LM-Steers Connect Word Embeddings with the Text Distribution
  • 5.1 Interpreting Word Embeddings
  • 5.2 Highlighting Keywords in Styled Texts
  • 5.3 Transferring LM-Steer Between Models
  • 6 Conclusions
  • Limitations
  • Acknowledgement
  • References
  • A. Broader Impacts
  • B. Formal Statement of Theorem 1
  • C. Formal Statement and Proof of Theorem 2
  • D. Implementation Details
  • E. Hyperparameter Selection
  • F. Details of Transferring LM-Steer to Other Language Models
  • G. Details of Investigating Interpretability
  • H. Validity of Assumptions
  • I. Comparison with A (Soft) Word Blacklist
  • I.1 Formal statement of the universality of LM-Steer
  • I.2 Construction of a counterexample for SWB
  • J. Results of LM-Steer on Pythia Family
  • K. LoRA Configuration
  • L. Incorporating in LoRA
  • M. Results on GPT-J-6B
  • N. Effect of LM-Steer on Instruction Following
  • O. Embedding Tuning

Knowls

  1. Knowl 1 — Style changes can correspond to linear transformations of output embeddings

    theoretical result

    In the paper’s idealized correspondence between a language model and a hidden Markov model (HMM), changing the initial distribution over style-related conditions can be represented by a linear transformation of the language model’s output word embeddings. The result assumes that HMM states have representations divided into semantic and condition components; the two initial distributions have identical semantic components but different condition components, with every condition coordinate of the original initial representation nonzero. It also assumes the paper’s representation constraints: each state-representation dimension has the same total squared magnitude across states, distinct dimensions are pairwise orthogonal across states, and the specified mixed third-order sums involving condition dimensions vanish. Under these conditions, some embedding transformation changes the generated sequence distribution from the one associated with the first initial distribution to the one associated with the second. This is a theoretical existence result under the paper’s HMM and linear-probability formulation, not a guarantee that an arbitrary transformation will produce a desired style in an unrestricted pretrained language model.

  2. Knowl 2 — LM-Steer controls generation by transforming output word embeddings

    model/method

    LM-Steer keeps a language model’s original parameters fixed and learns a d×dd\times d matrix WW that modifies every output-token embedding. For a token vv with original embedding ev∈Rde_v\in\mathbb{R}^d, the modified embedding is

    ev′=(I+ϵW)ev,e'_v=(I+\epsilon W)e_v,

    where II is the d×dd\times d identity matrix and the scalar ϵ\epsilon sets steering direction and strength. Given context vector cc, the modified logit is c⊤ev′c^\top e'_v. The paper uses ϵ0=10−3\epsilon_0=10^{-3} as a default scale; positive and negative steering values can be trained on positively and negatively labeled text, respectively, by maximizing their likelihood under the correspondingly steered model. In the implementation, the context can instead be transformed to c′=c+ϵWcc'=c+\epsilon Wc before computing logits. Training uses Adam at learning rate 10−210^{-2} for 1,000 steps, with WW initialized from a zero-mean Gaussian of variance 10−310^{-3}. To accommodate the distribution gap between pretraining and task data, the implementation also uses a dummy steer: the positive model uses Pϵ0(W+Wdummy)P_{\epsilon_0(W+W_{\rm dummy})} and the negative model uses Pϵ0(−W+Wdummy)P_{\epsilon_0(-W+W_{\rm dummy})}. The base language model remains frozen.

  3. Knowl 3 — LM-Steer is more expressive than a context-independent word blacklist

    theoretical result

    A soft word blacklist (SWB) adds the same learned logit offset for a token at every generation position, so its preference for a token does not depend on the current context. In contrast, LM-Steer adds a context-dependent term, ϵc⊤Wev\epsilon c^\top W e_v, and can therefore change token preferences differently across contexts. The paper proves an existential universality result: for any two finite-length sequence distributions over a finite vocabulary, there exist a context-vector function, embeddings, and a steer matrix such that the negative and positive steers represent the two distributions, respectively. This result concerns the existence of a suitable constructed model, not universal control by a fixed pretrained model. The paper also gives an SWB counterexample: a model assigning all probability to the sequence “AB” cannot be converted by a context-independent blacklist into the distribution uniform over “AB” and “BA”.

  4. Knowl 4 — LM-Steer reduces toxicity across GPT-2 sizes while retaining generation quality

    empirical result

    For detoxification, the paper trains on the Jigsaw Unintended Bias in Toxicity Classification data and evaluates on 10,000 nontoxic RealToxicityPrompts prompts. It generates 25 continuations of up to 20 tokens per prompt with nucleus sampling at p=0.9p=0.9 and uses a steering value of 5ϵ05\epsilon_0. Toxicity is measured with Perspective API as average maximum toxicity per prompt and the probability of generating text with toxicity above 0.5; fluency is GPT-2-large perplexity, and diversity is the fraction of distinct nn-grams. On GPT-2 base, medium, and large, LM-Steer obtains average maximum toxicity of 0.296±0.0180.296\pm0.018, 0.215±0.0150.215\pm0.015, and 0.249±0.0070.249\pm0.007, and toxicity probabilities of 0.129±0.0120.129\pm0.012, 0.059±0.0290.059\pm0.029, and 0.089±0.0090.089\pm0.009, respectively. Their perplexities are 36.87, 43.56, and 28.26; Dist-1/2/3 are 0.54/0.86/0.86, 0.56/0.83/0.84, and 0.55/0.84/0.84. For comparison, GPT-2’s original model scores 0.527 and 0.520 on the two toxicity metrics, with perplexity 25.45; DExperts-base scores 0.302 and 0.118, with perplexity 38.20; MuCoLa scores 0.308 and 0.088, with perplexity 29.92; and the soft blacklist scores 0.270 and 0.154, with perplexity 18.28. Thus, LM-Steer medium gives the lowest average maximum toxicity and toxicity probability among these reported GPT-2 systems, while its perplexity and diversity show a trade-off rather than no quality cost. The paper also reports lower toxicity on GPT-J-6B after steering: average maximum toxicity changes from 0.364 to 0.265 and toxicity probability from 0.229 to 0.124, while perplexity changes from 18.70 to 18.26 and Dist-1/2/3 from 0.55/0.84/0.85 to 0.54/0.84/0.85. Human pairwise judgments found LM-Steer generations less toxic and more topical than the compared baselines, and more fluent than LoRA.

  5. Knowl 5 — LM-Steer controls sentiment with competitive results on positive and negative targets

    empirical result

    For sentiment control, the paper trains on SST-5, treating labels 1–2 as negative and 4–5 as positive, and evaluates generations with a Hugging Face sentiment classifier. Prompts are OpenWebText examples filtered by sentiment; each prompt receives 25 continuations of up to 20 tokens. Positive and negative control use 5ϵ05\epsilon_0 and −5ϵ0-5\epsilon_0, respectively. The reported positivity percentages are separated by whether the prompt is positive, neutral, or negative. Under positive steering, LM-Steer-medium scores 95.36%, 56.98%, and 67.68% on those three prompt groups; LM-Steer-base scores 90.46%, 57.26%, and 54.38%; and LM-Steer-large scores 90.70%, 41.23%, and 41.20%. DExperts-small scores 94.57%, 31.64%, and 42.08%. Under negative steering, LM-Steer-medium scores 52.32%, 7.10%, and 71.48%, while DExperts-large scores 35.99%, 3.77%, and 45.91%. The paper characterizes LM-Steer as first on the positive-control side and second or third on the negative-control side, with reasonable fluency and diversity. These results show that steering changes classifier-measured sentiment, though the magnitude varies by prompt group and target polarity.

  6. Knowl 6 — Steering strength enables continuous and compositional control

    model/method

    A trained LM-Steer can be strengthened, weakened, interpolated, or extrapolated by changing the scalar ϵ\epsilon without retraining the steer. The paper demonstrates sentiment control over approximately −5ϵ0-5\epsilon_0 to 5ϵ05\epsilon_0 and reports a corresponding shift in the generated sentiment distribution. Multiple controls can also be combined: if W1W_1 and W2W_2 are steer matrices for different tasks and ϵ1,ϵ2\epsilon_1,\epsilon_2 are their strengths, the combined model uses the summed transformation ϵ1W1+ϵ2W2\epsilon_1W_1+\epsilon_2W_2. The paper demonstrates joint sentiment and toxicity control, while noting that the factors can influence one another—for example, negative sentiment may also increase toxicity. In a toxicity example, increasing the steer from negative to positive values reduced both the number and intensity of toxic words in the generated continuations.

  7. Knowl 7 — A learned steer can expose style-related embedding dimensions and indicative spans

    model/method

    LM-Steer is also used as an interpretability tool. For detoxification, the authors apply singular value decomposition to the learned matrix WW and use its ranked, influential rows to identify word embeddings most associated with the learned transformation. Among the English-token groups they report are “stupid,” “idiot,” “jerk,” and “pathetic” for one dimension, and “bullshit,” “shit,” “lies,” and “manipulation” for another; three dimensions were discarded because they mainly matched non-English tokens. The paper also scores each token viv_i in a text by the change in its conditional log-likelihood, log⁡PϵW(vi∣v<i)−log⁡P0(vi∣v<i)\log P_{\epsilon W}(v_i\mid v_{<i})-\log P_0(v_i\mid v_{<i}), where PϵWP_{\epsilon W} is the steered model and P0P_0 is the original model. It then selects, by dynamic programming, the contiguous span of at most five tokens with the greatest cumulative score. On toxic prompts, the highlighted spans included insulting, profane, controversial, and sexually explicit language. The two analyses provide complementary views: dimensions associated with the learned style transformation and specific spans that distinguish a given text under that transformation.

  8. Knowl 8 — LM-Steer transfers between models through an embedding-space map

    model/method

    To transfer a steer from a source model to a target model, the paper first learns a linear map HH from the target model’s embedding space into the source model’s space. If eve_v and cc are source-model token and context vectors, and ev′e'_v and c′c' are their target-model counterparts, the transfer assumes ev=Hev′e_v=He'_v and c=Hc′c=Hc'. The source steer term then becomes c′⊤(H⊤WH)ev′c'^\top(H^\top W H)e'_v, so the target model can use the transformed steer matrix H⊤WHH^\top W H. The mapping is fitted using the top 4,000 words shared by the vocabularies, with Adam at learning rate 0.01 for 5,000 steps and a zero-centered Gaussian initialization of variance 10−310^{-3}. Because the mapping is approximate, the authors reduce the generation steer to 0.5ϵ00.5\epsilon_0 for transfer. A steer trained on GPT-2-large partially retained detoxification ability on other model sizes; the paper reports average maximum toxicity scores of 0.307 and 0.308 for transferred GPT-2 and GPT-2-medium, respectively, describing them as similar to the best baseline.

  9. Knowl 9 — LM-Steer uses few additional parameters and can learn from small datasets

    empirical result

    On GPT-2-large, the paper reports 1.6 million LM-Steer parameters, about 0.2% of the 762-million-parameter backbone and about 9% of the 18-million-parameter LoRA configuration. Its decoding-time ratio relative to the base model is 1.24. For comparison, the reported parameter counts and decoding-time ratios are: DAPT, 355M and 1.00; GeDi, 355M and 2.94; CTRL, 355M and 3.79; PPLM, 124M and 270.11; DExperts, 355M and 1.98; MuCoLa, 898M and 24.03; and LoRA, 18M and 1.00. In a detoxification data-size experiment, the paper varies training data from 30 to 10,000 examples and reports a toxicity score of 0.322 with only 30 examples. It reports a better balance between detoxification and generation quality once the training set exceeds 3,000 examples. The results indicate that the small steer can be trained with limited data and adds comparatively little decoding overhead.

  10. Knowl 10 — LM-Steer is limited to wording-related control and accessible output embeddings

    limitation

    LM-Steer acts on output word embeddings and therefore primarily controls styles expressed through word choice. The paper states that this restricts its ability to handle more complex targets, such as syntactic structures or persuasive techniques requiring logical reasoning. The method also requires direct access to a language model’s output embeddings, so it cannot be applied to language-model APIs that do not expose those embeddings.

Coverage note — Supplementary Pythia-family detoxification results, the prompted-generation and LoRA-combination studies, and detailed example generations were omitted because the main knowls capture the core mechanism, principal task results, flexibility, transfer, efficiency, interpretation, and stated scope limitation.

References

  1. 1.Carl Allen and Timothy Hospedales. 2019. Analogies explained: Towards understanding word embeddings. In International Conference on Machine Learning, pages 223–231. PMLR.
  2. 2.Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavaš. 2021. RedditBias: A real-world resource for bias evaluation and debiasing of conversational language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1941–1955, Online. Association for Computational Linguistics.
  3. 3.Leonard E Baum, Ted Petrie, George Soules, and Norman Weiss. 1970. A maximization technique occurring in the statistical analysis of probabilistic functions of markov chains. The annals of mathematical statistics, 41(1):164–171.
  4. 4.Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. 2023. Pythia: A suite for analyzing large language models across training and scaling.
  5. 5.Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  7. 7.Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-Graber, and David Blei. 2009. Reading tea leaves: How humans interpret topic models. Advances in neural information processing systems, 22.
  8. 8.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  9. 9.Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2020. Queens are powerful too: Mitigating gender bias in dialogue generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8173–8188, Online. Association for Computational Linguistics.
  10. 10.Kawin Ethayarajh. 2019. Rotate king to get queen: Word relationships as orthogonal transformations in embedding space. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3503–3508.
  11. 11.Manaal Faruqui, Jesse Dodge, Sujay K Jauhar, Chris Dyer, Eduard Hovy, and Noah A Smith. 2014. Retrofitting word vectors to semantic lexicons. arXiv preprint arXiv:1411.4166.
  12. 12.Yi Fung, Tuhin Chakrabarty, Hao Guo, Owen Rambow, Smaranda Muresan, and Heng Ji. 2023. NORMSAGE: Multi-lingual multi-cultural norm discovery from conversations on-the-fly. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15217–15230, Singapore. Association for Computational Linguistics.
  13. 13.Yi Fung, Ruining Zhao, Jae Doo, Chenkai Sun, and Heng Ji. 2024. Massively multi-cultural knowledge acquisition & lm benchmarking.
  14. 14.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369.
  15. 15.Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. 2019. Openwebtext corpus (2019). URL http://Skylion007. github. io/OpenWebTextCorpus.
  16. 16.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
  17. 17.John Hewitt, John Thickstun, Christopher Manning, and Percy Liang. 2023. Backpack language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9103–9125, Toronto, Canada. Association for Computational Linguistics.
  18. 18.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. In International Conference on Learning Representations.
  19. 19.Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2021a. Lora: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  20. 20.Xiaodan Hu, Pengfei Yu, Kevin Knight, Heng Ji, Bo Li, and Honghui Shi. 2021b. Muse: Textual attributes guided portrait painting generation. In Prof. The 3rd IEEE Workshop on Artificial Intelligence for Art Creation.
  21. 21.Ali Jahanian, Lucy Chai, and Phillip Isola. On the "steerability" of generative adversarial networks. In International Conference on Learning Representations.
  22. 22.Kyoung-Rok Jang and Sung-Hyon Myaeng. 2017. Elucidating conceptual properties from word embeddings. In Proceedings of the 1st Workshop on Sense, Concept and Entity Representations and their Applications, pages 91–95.
  23. 23.Masahiro Kaneko and Danushka Bollegala. 2021. Debiasing pre-trained contextualised embeddings. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1256–1266, Online. Association for Computational Linguistics.
  24. 24.Masahiro Kaneko, Danushka Bollegala, and Naoaki Okazaki. 2022. Debiasing isn’t enough! – on the effectiveness of debiasing MLMs and their social biases in downstream tasks. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1299–1310, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  25. 25.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  26. 26.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems.
  27. 27.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021. Gedi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4929–4952.
  28. 28.Sachin Kumar, Eric Malmi, Aliaksei Severyn, and Yulia Tsvetkov. 2021. Controlled text generation as continuous optimization with multiple constraints. Advances in Neural Information Processing Systems, 34:14542–14554.
  29. 29.Sachin Kumar, Biswajit Paria, and Yulia Tsvetkov. 2022. Gradient-based constrained sampling from language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2251–2277.
  30. 30.Ariel N Lee, Cole J Hunter, and Nataniel Ruiz. 2023. Platypus: Quick, cheap, and powerful refinement of llms. arXiv preprint arXiv:2308.07317.
  31. 31.Juncen Li, Robin Jia, He He, and Percy Liang. 2018. Delete, retrieve, generate: a simple approach to sentiment and style transfer. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1865–1874.
  32. 32.Sha Li, Chi Han, Pengfei Yu, Carl Edwards, Manling Li, Xingyao Wang, Yi Fung, Charles Yu, Joel Tetreault, Eduard Hovy, and Heng Ji. 2023a. Defining a new NLP playground. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 11932–11951, Singapore. Association for Computational Linguistics.
  33. 33.Sha Li, Ruining Zhao, Manling Li, Heng Ji, Chris Callison-Burch, and Jiawei Han. 2023b. Open-domain hierarchical event schema induction by incremental prompting and verification. In Proc. The 61st Annual Meeting of the Association for Computational Linguistics (ACL2023).
  34. 34.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597.
  35. 35.Yuling Li, Kui Yu, and Yuhong Zhang. 2021. Learning cross-lingual mappings in imperfectly isomorphic embedding spaces. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:2630–2642.
  36. 36.Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2020. Towards debiasing sentence representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5502–5515, Online. Association for Computational Linguistics.
  37. 37.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021. Dexperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706.
  38. 38.Kevin Lund and Curt Burgess. 1996. Producing high-dimensional semantic spaces from lexical co-occurrence. Behavior research methods, instruments, & computers, 28(2):203–208.
  39. 39.Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2022. An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language Models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1878–1898, Dublin, Ireland. Association for Computational Linguistics.
  40. 40.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
  41. 41.Brian Murphy, Partha Talukdar, and Tom Mitchell. 2012. Learning effective and interpretable semantic models using non-negative sparse embedding. In Proceedings of COLING 2012, pages 1933–1950.
  42. 42.Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5356–5371, Online. Association for Computational Linguistics.
  43. 43.Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1953–1967, Online. Association for Computational Linguistics.
  44. 44.Ali Omrani, Alireza Salkhordeh Ziabari, Charles Yu, Preni Golazizian, Brendan Kennedy, Mohammad Atari, Heng Ji, and Morteza Dehghani. 2023. Social-group-agnostic bias mitigation via the stereotype content model. In Proc. The 61st Annual Meeting of the Association for Computational Linguistics (ACL2023).
  45. 45.OpenAI. 2023. Gpt-4 technical report.
  46. 46.Abhishek Panigrahi, Harsha Vardhan Simhadri, and Chiranjib Bhattacharyya. 2019. Word2Sense: Sparse interpretable word embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5692–5705, Florence, Italy. Association for Computational Linguistics.
  47. 47.Sungjoon Park, JinYeong Bak, and Alice Oh. 2017. Rotated word vector representations and their interpretability. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 401–411, Copenhagen, Denmark. Association for Computational Linguistics.
  48. 48.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training.
  49. 49.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  50. 50.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  51. 51.Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020. Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7237–7256, Online. Association for Computational Linguistics.
  52. 52.Sascha Rothe and Hinrich Schütze. 2016. Word embedding calculus in meaningful ultradense subspaces. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 512–517.
  53. 53.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP. Transactions of the Association for Computational Linguistics, 9:1408–1424.
  54. 54.Lütfi Kerem Şenel, Furkan Şahinuç, Veysel Yücesoy, Hinrich Schütze, Tolga Çukur, and Aykut Koç. 2022. Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts. Information Processing & Management, 59(3):102925.
  55. 55.Lütfi Kerem Şenel, Ihsan Utlu, Veysel Yücesoy, Aykut Koc, and Tolga Cukur. 2018. Semantic structure and interpretability of word embeddings. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 26(10):1769–1779.
  56. 56.Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2019. The woman worked as a babysitter: On biases in language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3407–3412.
  57. 57.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.
  58. 58.Nishant Subramani, Nivedita Suresh, and Matthew E Peters. 2022. Extracting latent steering vectors from pretrained language models. In Findings of the Association for Computational Linguistics: ACL 2022, pages 566–581.
  59. 59.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  60. 60.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  61. 61.Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2021. Measuring and reducing gendered correlations in pre-trained models.
  62. 62.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38–45.
  63. 63.Ke Yang, Charles Yu, Yi Fung, Manling Li, and Heng Ji. 2023. Adept: A debiasing prompt framework. In Proc. Thirty-Seventh AAAI Conference on Artificial Intelligence (AAAI2023).
  64. 64.Kevin Yang and Dan Klein. 2021. Fudge: Controlled text generation with future discriminators. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3511–3535.
  65. 65.Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023. Unlearning bias in language models by partitioning gradients. In Proc. The 61st Annual Meeting of the Association for Computational Linguistics (ACL2023) Findings.
  66. 66.Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, and Mrinmaya Sachan. 2023. Controlled text generation with natural language instructions. arXiv preprint arXiv:2304.14293.

Citation

MLA
Han, C., et al. “Word Embeddings Are Steers for Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 16410–30, https://doi.org/10.18653/v1/2024.acl-long.864.
APA
Han, C., Xu, J., Li, M., Fung, Y., Sun, C., Jiang, N., Abdelzaher, T., & Ji, H. (2024). Word Embeddings Are Steers for Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16410–16430. https://doi.org/10.18653/v1/2024.acl-long.864
Chicago
Han, C., J. Xu, M. Li, et al. 2024. “Word Embeddings Are Steers for Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16410–30. https://doi.org/10.18653/v1/2024.acl-long.864.
Harvard
Han, C. et al. (2024) “Word Embeddings Are Steers for Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 16410–16430. Available at: https://doi.org/10.18653/v1/2024.acl-long.864.
Vancouver
1. Han C, Xu J, Li M, Fung Y, Sun C, Jiang N, Abdelzaher T, Ji H (2024) Word Embeddings Are Steers for Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 16410–16430

BibTeX

@inproceedings{han-etal-2024-word,
    title = "Word Embeddings Are Steers for Language Models",
    author = "Han, Chi  and
      Xu, Jialiang  and
      Li, Manling  and
      Fung, Yi  and
      Sun, Chenkai  and
      Jiang, Nan  and
      Abdelzaher, Tarek  and
      Ji, Heng",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.864/",
    doi = "10.18653/v1/2024.acl-long.864",
    pages = "16410--16430"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/