InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Pengyu WangDong ZhangLinyang LiChenkun TanXinghao WangMozhi ZhangKe RenBotian JiangXipeng Qiu

article2024EMNLP106 citations

Proposes InferAligner, an inference-time alignment method that transfers safety steering vectors from aligned models to target models during decoding, effectively blocking harmful and jailbreak prompts across domain-specific and multimodal architectures without retraining or sacrificing downstream task utility.

Listen

As organizations increasingly customize large language models for specialized domains such as finance, medicine, and mathematics, ensuring these systems remain safe and harmless is a critical priority. Traditional alignment techniques embed safety rules during the training phase using resource-intensive optimization methods. However, these conventional training-time approaches require massive amounts of curated data, demand heavy computational resources, and frequently trigger an "alignment tax" that degrades the model's specialized performance on downstream business tasks.

To address this challenge, the article introduces and evaluates InferAligner, an inference-time alignment method that decouples downstream capability training from safety enforcement. The system uses a "guidance gate" during deployment to evaluate the intent of incoming queries. When an input is benign, the model functions normally without intervention. When a harmful prompt or adversarial "jailbreak" attack is detected, the framework steers the target model's internal activations using safety steering vectors extracted from a safety-aligned reference model, effectively compelling the system to refuse harmful generation.

The authors conducted comprehensive evaluations across multiple language model families (including Llama 2, Llama 3, Qwen, and InternLM) across finance, medical, and mathematical tasks, as well as multimodal vision-language models such as LLaVA. The results show that InferAligner reduces the attack success rate of harmful instructions to 0% and jailbreak attack success rates from over 40% down to 0%–0.2%. Crucially, downstream task accuracy remained entirely preserved (e.g., maintaining 92.9% accuracy in finance and 42.7% in medicine), whereas training-time methods reduced downstream accuracy by several percentage points. Furthermore, the approach adds almost no latency overhead during inference.

These findings demonstrate that organizations can train specialized domain models purely for performance and apply cross-model safety guardrails during runtime without retraining or sacrificing accuracy. Decision-makers should consider piloting inference-time activation steering for domain-specific deployments to reduce safety risks and compute costs. Future work should focus on expanding this approach beyond harmlessness to other alignment objectives, such as honesty and helpfulness, while testing its scalability across broader operational environments.

Cover for InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Abstract

As large language models (LLMs) rapidly evolve, they are increasingly being customized through fine-tuning to suit the specific needs of various applications. A critical aspect of this advancement is the alignment process, which ensures that these models perform tasks in ways that align with human values and expectations. Current alignment methods, such as direct preference optimization (DPO) and reinforcement learning from human feedback (RLHF), focus primarily on alignment during training phase. However, these methods often involve complex and resource-intensive training processes, posing significant challenge for their implementation. Therefore, we propose InferAligner, a simple yet effective method for harmlessness alignment during inference phase. InferAligner decouples harmlessness from helpfulness. During the training phase, it focuses solely on enhancing the target model’s capabilities on downstream tasks. In the inference phase, it utilizes safety steering vectors extracted from the aligned model to guide the target model towards harmlessness alignment. Experimental results show that our method can be very effectively applied to domain-specific models in finance, medicine, and mathematics, as well as to multimodal large language models (MLLMs) such as LLaVA. It significantly diminishes the attack success rate (ASR) of both harmful instructions and jailbreak instructions, while maintaining almost unchanged performance in downstream tasks.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 LLMAlignment
  • 2.2 Safety Concerns of LLMs
  • 2.3 Activation Engineering
  • 3 Methodology
  • 3.1 Safety Related Vector
  • 3.2 Workflow of InferAligner
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Evaluation Metrics
  • 4.3 Implementation Details
  • 5 Experiments
  • 5.1 Baselines
  • 5.2 Main Results
  • 6 Analysis
  • 6.1 Results on Various LLMs
  • 6.2 Results on MLLMs
  • 6.3 Ablation Study
  • 6.4 The Effect of Intervention Strength
  • 7 Conclusion
  • Limitations
  • Ethical Concerns
  • Acknowledgements
  • References
  • A Glossary
  • B Experiments on Safety-Related Vectors
  • B.1 Intent Discernment
  • B.2 Extraction Methods
  • C Hyperparameters of InferAligner
  • D Links to Models
  • E Judgement Model for Harmfulness Evaluation
  • F Safety Score
  • G Case Study

Knowls

  1. Knowl 1 — Gated cross-model activation steering at inference time

    model/method

    InferAligner separates task capability from harmlessness alignment: a domain-specific target model is trained for its downstream task, while safety guidance comes from a separately aligned model. For an input prompt PP and transformer layer ll, a scalar gate uses the target model’s safety-related direction sls_l to decide whether the prompt is harmful. If it is harmful, the target model’s activations are shifted using the aligned model’s safety steering vector s^l\hat{s}_l; if it is harmless, the gate is zero and the target activations are unchanged.

    gl={1if al(P)⊤sl+bl>0,0otherwise,g_l = \begin{cases}1 & \text{if } a_l(P)^\top s_l + b_l > 0,\\0 & \text{otherwise,}\end{cases} xl=xl′+αgls^l,l∈LG.x_l = x'_l + \alpha g_l \hat{s}_l, \qquad l\in L_G.

    Here, al(P)a_l(P) is the target model’s last-token activation for prompt PP at layer ll; blb_l is the mean projection of the safety-vector extraction samples classified as negative onto sls_l; xl′x'_l and xlx_l are respectively the original and modified target-model activations at layer ll across all token positions; α\alpha is the intervention strength; and LGL_G is the set of layers to intervene on. The target model’s safety-related vector is used for intent detection, whereas the aligned model’s safety steering vector supplies the behavioral direction. The activation shift is broadcast across token positions.

  2. Knowl 2 — Mean-difference extraction of safety directions

    equation

    For a language model, InferAligner constructs a layer-specific safety-related vector (SRV) from matched sets of harmful and harmless prompts. Let D−D^- and D+D^+ each contain NN prompts, respectively formed from harmful and harmless instructions using the model’s conversation template. Let al(P)a_l(P) denote the last-token activation at transformer layer ll for prompt PP. The unnormalized vector is the difference between the harmful-prompt and harmless-prompt mean activations; the SRV is its unit-normalized version.

    vl′=1N∑i=1Nal(Pi−)−1N∑j=1Nal(Pj+),sl=vl′∥vl′∥.v'_l = \frac{1}{N}\sum_{i=1}^{N} a_l(P_i^-) - \frac{1}{N}\sum_{j=1}^{N} a_l(P_j^+), \qquad s_l = \frac{v'_l}{\lVert v'_l\rVert}.

    Here, Pi−∈D−P_i^-\in D^- and Pj+∈D+P_j^+\in D^+ are prompts, and ll indexes a transformer layer. Applying the same extraction to an aligned model yields the safety steering vector (SSV) used to steer the target model. The paper uses the target model’s SRV for intent detection and the aligned model’s SSV for intervention.

  3. Knowl 3 — Domain-specific evaluation and implementation protocol

    experimental setup

    The experiments fine-tuned Llama2-7B on finance, medicine, and mathematics data to create domain-specific target models; Llama2-7B-chat supplied the aligned-model safety directions in the main comparison. Finance training used financial instruction-tuning data plus 10,000 UltraChat conversations; medicine training used MEDQA data plus an equivalent amount of UltraChat conversations; mathematics training used GSM8K training examples with chain-of-thought answers plus an equivalent amount of UltraChat conversations. For safety-vector extraction, the authors sampled 64 harmful instructions from AdvBench’s 520 harmful instructions and 64 harmless instructions from a randomly selected set of 520 TruthfulQA questions. The remaining instructions formed the harmfulness test pool. The jailbreak test set combined 50 harmful instructions with 10 jailbreak prompt types, yielding 500 examples.

    For all evaluated models, intervention layers were chosen to discriminate intent in both the target and aligned models: LG=[12,24)L_G=[12,24) for 7B models and LG=[16,32)L_G=[16,32) for 13B models. Candidate intervention strengths were {1.0,2.0,3.0,4.0,6.0,8.0}\{1.0,2.0,3.0,4.0,6.0,8.0\}; the selected value was 4.04.0 except for InternLM, which used 8.08.0. Domain-specific fine-tuning used 8 NVIDIA A100 80G GPUs, batch size 128, maximum sequence length 2,048, AdamW, 10% warm-up, cosine learning-rate decay, and two training epochs. The maximum learning rates were 2×10−52\times10^{-5} for supervised fine-tuning and 5×10−65\times10^{-6} for DPO; evaluated generations used greedy decoding.

    Safety was measured by attack success rate (ASR), the proportion of harmful responses, with GPT-3.5 turbo judging LLM outputs and GPT-4V judging multimodal outputs. Utility was measured by accuracy on the downstream tasks. On 120 manually labeled LLM instruction-response pairs, GPT-3.5 turbo achieved 98.2% judgment accuracy, compared with 97.5% for GPT-4, 78.3% for a RoBERTa classifier, 57.5% for a BERT classifier, and 60.8% for rule matching. GPT-4V’s judgments on 40 manually labeled multimodal pairs matched human labels exactly.

  4. Knowl 4 — Main safety and utility results on three domains

    data/table

    The table reports harmfulness and downstream-task utility for Llama2-7B domain-specific models, comparing training-time and inference-time alignment baselines with InferAligner. Each domain reports harmfulness ASR, jailbreak ASR, and task accuracy (Acc.); lower ASR and higher accuracy are preferable. InferAligner reduced both ASRs to zero in medicine and mathematics and to 0.0 and 0.2 in finance, while retaining the unaligned domain-specific model’s accuracy values (92.9, 42.7, and 39.0, respectively). The DPO and safe-data baselines also reduced direct-harm ASR, but retained higher jailbreak ASRs in medicine and mathematics and showed lower utility in mathematics.

    Finance Medicine Mathematics
    Model ASR Jailbreak ASR Acc. ASR Jailbreak ASR Acc. ASR Jailbreak ASR Acc.
    DS-Safe-Llama2 0.7 13.4 92.9 0.0 0.6 40.1 0.2 14.0 36.7
    DS-Llama2-chat 0.7 1.0 93.7 0.2 1.4 40.6 0.7 2.6 36.8
    DS-Llama2 38.4 48.2 92.9 31.6 21.4 42.7 36.8 42.2 39.0
    DS-Llama2 + DPO 0.0 1.0 93.0 4.6 20.4 41.6 3.7 11.6 26.8
    DS-Llama2 + Self-Reminder 25.0 34.8 92.8 29.2 25.8 43.4 14.9 37.2 38.0
    DS-Llama2 + Goal Priority 21.3 25.8 92.4 11.0 13.6 43.8 7.5 4.2 39.3
    DS-Llama2 + InferAligner 0.0 0.2 92.9 0.0 0.0 42.7 0.0 0.0 39.0
  5. Knowl 5 — Transfer across LLM model families

    empirical result

    Beyond the Llama2 experiments, the authors applied InferAligner to domain-specific models based on Llama3, Qwen, and InternLM. Their reported comparisons show substantial safety improvement while downstream-task performance remains consistent with the corresponding models without InferAligner. The paper presents these outcomes graphically rather than reporting numerical values in the text, so no exact ASR or accuracy changes are specified here.

  6. Knowl 6 — Multimodal steering and MM-Harmful Bench

    empirical result

    InferAligner was applied to LLaVA-7B and LLaVA-13B using safety steering vectors extracted from Llama2-7B-chat, even though LLaVA incorporates visual information during training. The authors report that LLaVA refused all harmful multimodal instructions in their evaluation and gave coherent refusals that identified harmful aspects of the requests. The intervention did not increase context length and was reported to have almost no effect on inference time, unlike the longer-prompt Goal Priority baseline, which slowed inference.

    For this evaluation, the authors constructed MM-Harmful Bench, a set of 100 harmful instructions requiring both an image and text to respond. It covers ten malicious-intent categories: discrimination, sabotage, theft, defamation, illegal weapons, fraud, self-harm, psychological manipulation, misinformation, and cybercrime.

  7. Knowl 7 — The source of steering vectors determines whether guidance works

    empirical result

    In an ablation, safety-related vectors extracted from a domain-specific target model did not improve its safety and appeared to worsen it. The authors interpret this as evidence that detecting harmful intent in a target model does not necessarily give that model the ability to refuse harmful requests. By contrast, vectors extracted from SafeLIMA successfully guided target models toward harmlessness: SafeLIMA was trained with 1,000 LIMA instruction-tuning examples and 100 safe examples. The authors also report that increasing intervention strength could reinforce the control exerted by SafeLIMA’s vectors.

  8. Knowl 8 — Intervention strength controls the observed safety response

    empirical result

    The authors varied the intervention strength α\alpha and scored responses to harmful instructions on a five-point safety scale, where 5 means completely safe and 1 means highly unsafe. Across finance, medicine, and mathematics models, increasing α\alpha increased the reported safety score; at α=4.0\alpha=4.0, the score approached 5. Subtracting the safety steering vectors instead of adding them increased response harmfulness. The paper reports these trends graphically and does not provide the individual plotted scores as text.

  9. Knowl 9 — Alternative SRV extraction methods give the same intent accuracy

    empirical result

    The authors compared their mean-difference SRV extraction with PCA, associated with RepE, and MD, associated with contrastive activation addition. Using the 12th-layer guidance gate of DS-Llama2-7B, all three methods achieved 100% accuracy in determining intent on the evaluated test problems. The vectors extracted by the three methods had pairwise similarities exceeding 99.8%. Because mean difference was computationally simplest among the compared approaches, the authors selected it for InferAligner.

  10. Knowl 10 — Scope limitation to harmlessness alignment

    limitation

    The work evaluates InferAligner for harmlessness alignment. It does not establish that the method works for other preference-alignment goals; the authors identify extending it to more diverse preference alignments as future work.

Coverage note — The appendix’s detailed evaluator prompt templates and qualitative case-study transcripts are omitted because they illustrate the evaluation and example behaviors rather than add distinct generalizable methods or findings.

References

  1. 1.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.
  2. 2.Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609.
  3. 3.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  4. 4.Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2023. Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions. arXiv preprint arXiv:2309.07875.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  6. 6.Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  9. 9.Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233.
  10. 10.Evan Hernandez, Belinda Z Li, and Jacob Andreas. 2023. Inspecting and editing knowledge representations in language models. arXiv preprint arXiv:2304.00740.
  11. 11.Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, and Hongsheng Li. 2023a. Instruct2act: Mapping multi-modality instructions to robotic actions with large language model. arXiv preprint arXiv:2305.11176.
  12. 12.Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2023b. Catastrophic jailbreak of open-source llms via exploiting generation. arXiv preprint arXiv:2310.06987.
  13. 13.Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421.
  14. 14.Maxim Khanov, Jirayu Burapacheep, and Yixuan Li. 2024. Args: Alignment as reward-guided search. arXiv preprint arXiv:2402.01694.
  15. 15.Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song. 2023a. Multi-step jailbreaking privacy attacks on chatgpt. arXiv preprint arXiv:2304.05197.
  16. 16.Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023b. Inference-time intervention: Eliciting truthful answers from a language model, july 2023. URL http://arxiv.org/abs/2306.03341.
  17. 17.Linyang Li, Pengyu Wang, Ke Ren, Tianxiang Sun, and Xipeng Qiu. 2023c. Origin tracing and detecting of llms. arXiv preprint arXiv:2304.14072.
  18. 18.Tianlong Li, Xiaoqing Zheng, and Xuanjing Huang. 2024. Open the pandora’s box of llms: Jailbreaking llms through representation engineering. arXiv preprint arXiv:2401.06824.
  19. 19.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023d. Alpacaeval: An automatic evaluator of instruction-following models.
  20. 20.Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. 2023e. Rain: Your language models can align themselves without finetuning. arXiv preprint arXiv:2309.07124.
  21. 21.Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2023. The unlocking spell on base llms: Rethinking alignment via in-context learning. arXiv preprint arXiv:2312.01552.
  22. 22.Stephanie Lin, Jacob Hilton, and Owain Evans. 2021. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958.
  23. 23.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a. Visual instruction tuning.
  24. 24.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. Advances in neural information processing systems, 36.
  25. 25.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  26. 26.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023c. Jailbreaking chatgpt via prompt engineering: An empirical study. arXiv preprint arXiv:2305.13860.
  27. 27.Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. Www’18 open challenge: financial opinion mining and question answering. In Companion proceedings of the the web conference 2018, pages 1941–1942.
  28. 28.Pekka Malo, Ankur Sinha, Pekka Korhonen, Jyrki Wallenius, and Pyry Takala. 2014. Good debt or bad debt: Detecting semantic orientations in economic texts. Journal of the Association for Information Science and Technology, 65(4):782–796.
  29. 29.OpenAI. 2023. Gpt-4 technical report.
  30. 30.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  31. 31.Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021. Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193.
  32. 32.Andrew Peng, Michael Wu, John Allard, Logan Kilpatrick, and Steven Heidel. 2023. Gpt-3.5 turbo fine-tuning and api updates.
  33. 33.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
  34. 34.Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2023. Steering llama 2 via contrastive activation addition. arXiv preprint arXiv:2312.06681.
  35. 35.Nishant Subramani, Nivedita Suresh, and Matthew E Peters. 2022. Extracting latent steering vectors from pretrained language models. arXiv preprint arXiv:2205.05124.
  36. 36.Tianxiang Sun, Xiaotian Zhang, Zhengfu He, Peng Li, Qinyuan Cheng, Xiangyang Liu, Hang Yan, Yunfan Shao, Qiong Tang, Shiduo Zhang, Xingjian Zhao, Ke Chen, Yining Zheng, Zhejian Zhou, Ruixiao Li, Jun Zhan, Yunhua Zhou, Linyang Li, Xiaogui Yang, Lingling Wu, Zhangyue Yin, Xuanjing Huang, Yu-Gang Jiang, and Xipeng Qiu. 2024. Moss: An open conversational large language model. Machine Intelligence Research.
  37. 37.InternLM Team. 2023. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM.
  38. 38.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models.
  39. 39.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  40. 40.Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023. Activation addition: Steering language models without optimization. arXiv preprint arXiv:2308.10248.
  41. 41.Pengyu Wang, Linyang Li, Ke Ren, Botian Jiang, Dong Zhang, and Xipeng Qiu. 2023. Seqxgpt: Sentence-level ai-generated text detection. arXiv preprint arXiv:2310.08903.
  42. 42.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560.
  43. 43.Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023. Defending chatgpt against jailbreak attack via self-reminders. Nature Machine Intelligence, 5(12):1486–1496.
  44. 44.Hongyang Yang, Xiao-Yang Liu, and Christina Dan Wang. 2023. Fingpt: Open-source financial large language models. arXiv preprint arXiv:2306.06031.
  45. 45.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems, 36.
  46. 46.Jiahao Yu, Xingwei Lin, and Xinyu Xing. 2023. Gpt-fuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253.
  47. 47.Dong Zhang, Zhaowei Li, Pengyu Wang, Xin Zhang, Yaqian Zhou, and Xipeng Qiu. 2024. Speechagents: Human-communication simulation with multi-modal multi-agent systems. arXiv preprint arXiv:2401.03945.
  48. 48.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023a. Safety-bench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045.
  49. 49.Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023b. Defending large language models against jailbreaking attacks through goal prioritization. arXiv preprint arXiv:2311.09096.
  50. 50.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. 2023. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206.
  51. 51.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023a. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405.
  52. 52.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023b. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Citation

MLA
Wang, P., et al. “InferAligner: Inference-Time Alignment for Harmlessness Through Cross-Model Guidance”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 10460–79, https://doi.org/10.18653/V1/2024.EMNLP-MAIN.585.
APA
Wang, P., Zhang, D., Li, L., Tan, C., Wang, X., Zhang, M., Ren, K., Jiang, B., & Qiu, X. (2024). InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10460–10479. https://doi.org/10.18653/V1/2024.EMNLP-MAIN.585
Chicago
Wang, P., D. Zhang, L. Li, et al. 2024. “InferAligner: Inference-Time Alignment for Harmlessness Through Cross-Model Guidance”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 10460–79. https://doi.org/10.18653/V1/2024.EMNLP-MAIN.585.
Harvard
Wang, P. et al. (2024) “InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10460–10479. Available at: https://doi.org/10.18653/V1/2024.EMNLP-MAIN.585.
Vancouver
1. Wang P, Zhang D, Li L, Tan C, Wang X, Zhang M, Ren K, Jiang B, Qiu X (2024) InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10460–10479

BibTeX

@inproceedings{Wang_2024, title={InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance}, url={http://dx.doi.org/10.18653/V1/2024.EMNLP-MAIN.585}, DOI={10.18653/v1/2024.emnlp-main.585}, booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing}, publisher={Association for Computational Linguistics}, author={Wang, Pengyu and Zhang, Dong and Li, Linyang and Tan, Chenkun and Wang, Xinghao and Zhang, Mozhi and Ren, Ke and Jiang, Botian and Qiu, Xipeng}, year={2024}, pages={10460–10479} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/