InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
Pengyu WangDong ZhangLinyang LiChenkun TanXinghao WangMozhi ZhangKe RenBotian JiangXipeng Qiu
Proposes InferAligner, an inference-time alignment method that transfers safety steering vectors from aligned models to target models during decoding, effectively blocking harmful and jailbreak prompts across domain-specific and multimodal architectures without retraining or sacrificing downstream task utility.
As organizations increasingly customize large language models for specialized domains such as finance, medicine, and mathematics, ensuring these systems remain safe and harmless is a critical priority. Traditional alignment techniques embed safety rules during the training phase using resource-intensive optimization methods. However, these conventional training-time approaches require massive amounts of curated data, demand heavy computational resources, and frequently trigger an "alignment tax" that degrades the model's specialized performance on downstream business tasks.
To address this challenge, the article introduces and evaluates InferAligner, an inference-time alignment method that decouples downstream capability training from safety enforcement. The system uses a "guidance gate" during deployment to evaluate the intent of incoming queries. When an input is benign, the model functions normally without intervention. When a harmful prompt or adversarial "jailbreak" attack is detected, the framework steers the target model's internal activations using safety steering vectors extracted from a safety-aligned reference model, effectively compelling the system to refuse harmful generation.
The authors conducted comprehensive evaluations across multiple language model families (including Llama 2, Llama 3, Qwen, and InternLM) across finance, medical, and mathematical tasks, as well as multimodal vision-language models such as LLaVA. The results show that InferAligner reduces the attack success rate of harmful instructions to 0% and jailbreak attack success rates from over 40% down to 0%–0.2%. Crucially, downstream task accuracy remained entirely preserved (e.g., maintaining 92.9% accuracy in finance and 42.7% in medicine), whereas training-time methods reduced downstream accuracy by several percentage points. Furthermore, the approach adds almost no latency overhead during inference.
These findings demonstrate that organizations can train specialized domain models purely for performance and apply cross-model safety guardrails during runtime without retraining or sacrificing accuracy. Decision-makers should consider piloting inference-time activation steering for domain-specific deployments to reduce safety risks and compute costs. Future work should focus on expanding this approach beyond harmlessness to other alignment objectives, such as honesty and helpfulness, while testing its scalability across broader operational environments.
- Paper: Representation Engineering: A Top-Down Approach to AI Transparency, Andy Zou et al. (2023). It introduces representation engineering and activation steering techniques upon which InferAligner's cross-model steering approach is built.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). It details how downstream domain fine-tuning compromises safety alignment, motivating the need for InferAligner's decoupled inference-time guardrails.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). It establishes input-output moderation baselines that InferAligner builds upon and compares against for prompt gating and refusal enforcement.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). It formalizes the trade-offs of helpfulness versus harmlessness and the alignment tax problem that inference-time alignment seeks to eliminate.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). It analyzes the core failure modes and jailbreak vulnerabilities of standard safety training that runtime steering mechanisms are designed to mitigate.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). It introduces the adversarial jailbreak benchmarks and threat models used to evaluate InferAligner's robustness against transfer attacks.
- Paper: Aligning Large Language Models with Representation Editing: A Control Perspective, Lingkai Kong et al. (2024). It expands on inference-time activation steering by framing representation editing through dynamic mathematical control systems.
- Paper: HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment, Shei Pern Chua et al. (2026). It extends internal representation safety analysis by exploring directional coupling between harm recognition and refusal commitment across network layers.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). It generalizes inference-time alignment beyond safety steering vectors to multi-attribute decoding-time guided search.
- Paper: Weak-to-Strong Jailbreaking on Large Language Models, Xuandong Zhao et al. (2025). It demonstrates how inference-time cross-model logit guidance can conversely be exploited as an adversarial attack mechanism against aligned models.
- Paper: Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing, Maryam Hashemzadeh et al. (2026). It develops a dynamic routing architecture that balances safe responses and domain-specific knowledge during generation without total refusal.
- Paper: FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts, Yichen Gong et al. (2025). It tests multimodal safety vulnerabilities that challenge and extend cross-model alignment techniques in vision-language models.
