DeAL: Decoding-time Alignment for Large Language Models
James Y. HuangSailik SenguptaDaniele BonadimanYi-An LaiArshit GuptaNikolaos PappasSaab MansourKatrin KirchhoffDan Roth
Proposes DeAL, a framework that aligns large language models at decoding time by framing generation as heuristic search, enabling fine-grained control over customizable and modular reward functions without requiring model fine-tuning.
Modern large language models are expected to adhere to human preferences such as safety, factual accuracy, and task-specific constraints. Standard alignment methods primarily rely on fine-tuning models during training using human preference data, or providing alignment guidelines through system prompts. However, training-time alignment binds models to static, non-universal principles that require costly retraining to update, while prompting methods remain fragile and vulnerable to jailbreaks. The article introduces and evaluates DeAL, a framework that enforces custom, multi-faceted alignment objectives dynamically during decoding time without retraining the base model.
DeAL formulates text generation as a heuristic-guided search problem. During token generation, the framework evaluates candidate paths using lookahead mechanisms paired with scoring functions. These scoring functions can encompass both programmatic rules, such as keyword inclusion and word limits, and abstract preference models, such as helpfulness and harmlessness reward estimators. The authors evaluated the approach across multiple open-source language models on keyword generation, length-constrained summarization, and open-ended harmful prompt benchmarks.
Key findings show that DeAL substantially improves adherence to alignment constraints. On keyword generation, it increased full constraint satisfaction by an average of 17 percentage points over prompting baselines. In summarization tasks under strict length constraints, pairing prompt instructions with DeAL increased length compliance to between 53% and 73%, compared to 3% to 16% for prompting alone, while preserving summarization quality. For abstract safety objectives, DeAL achieved 100% harmlessness on malicious prompt benchmarks where standard safety prompting failed 37% of the time. When tested against adversarial continuation attacks, DeAL maintained a 73% harmlessness rate, whereas system prompt safeguards collapsed to 20% harmlessness. Furthermore, combining DeAL with reinforcement learning from human feedback yielded the highest overall safety and helpfulness scores.
These findings indicate that decoding-time alignment provides an effective, modular guardrail for real-time risk mitigation and compliance. It allows organizations to adjust trade-offs between competing goals dynamically without incurring the significant expense and time of retraining models. The results also show that fine-tuning alone can provide a false sense of security, as decoding-time guidance can easily override training-time behavior.
Organizations deploying language models in high-risk environments should consider implementing decoding-time verification alongside existing training safeguards rather than relying solely on prompting. When doing so, teams should evaluate latency trade-offs, as DeAL introduces a two- to five-fold slowdown during generation. Future work is needed to optimize efficiency through speculative decoding or compiled grammars and to test performance across proprietary platforms that currently restrict access to token probability outputs.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). This paper establishes the foundation for grammar-constrained decoding to force strict constraint satisfaction without fine-tuning, directly informing DeAL's decoding-time lookahead mechanics.
- Paper: Stay on Topic with Classifier-Free Guidance, Guillaume Sanchez et al. (2024). This work demonstrates how inference-time next-token distribution modifications steer model outputs without retraining, providing key precedent for decoding-level steering.
- Book: Safety Alignment Should Be Made More Than Just a Few Tokens Deep, Xiangyu Qi et al. (2025). This book reveals the fundamental vulnerability of shallow training-time safety alignment to continuation and prefilling exploits, highlighting the need for dynamic decoding-time interventions like DeAL.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). This study introduces automated adversarial attacks on aligned models, establishing the standard jailbreak threat model that DeAL is evaluated against.
- Paper: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Yuntao Bai et al. (2022). This foundational paper defines the helpful and harmless preference modeling framework that DeAL integrates as modular scoring functions during decoding.
- Paper: Safe RLHF: Safe Reinforcement Learning from Human Feedback, Josef Dai et al. (2024). This work formalizes the decoupling of helpfulness and harmlessness optimization objectives, motivating DeAL's multi-objective lookahead alignment design.
- Paper: Controlled Text Generation with Natural Language Instructions, Wangchunshu Zhou et al. (2023). This paper analyzes the trade-offs between prompting, decoding-time search, and fine-tuning for satisfying lexical and structural constraints.
- Paper: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!, Xiangyu Qi et al. (2023). This study highlights how training-time safety alignment degrades under slight modifications, motivating the requirement for runtime guardrails during generation.
- Paper: Control Illusion: The Failure of Instruction Hierarchies in Large Language Models, Yilin Geng et al. (2026). This paper examines how language models fail to respect system prompt instruction hierarchies when conflicts arise, directly extending the justification for decoding-time enforcement mechanisms like DeAL.
- Paper: GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization, Shih-Yang Liu et al. (2026). This work develops advanced group policy optimization for conflicting multi-reward alignment, offering a complementary multi-objective perspective to DeAL's inference-time multi-facet scoring.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). This work provides a theoretical unification of direct preference alignment methods, contextualizing the fundamental trade-offs between offline fine-tuning and online, runtime guidance.
