Controlled Decoding from Language Models
Sidharth MudgalJong LeeHarish GanapathyYaGuang LiTao WangYanping HuangZhifeng ChenHeng-Tze ChengMichael CollinsTrevor Strohman
Proposes a modular controlled decoding framework that aligns frozen language models at inference time using trained prefix scorers, enabling multi-objective control and zero-shot transfer across base models without retraining.
Aligning large language models with human preferences typically requires updating base model weights through reinforcement learning or preference optimization. While effective, these training-time interventions are computationally expensive, inflexible when reward criteria shift, and prone to policy degradation. Standard inference-time alternatives, such as sampling many full candidate outputs and picking the best one, improve safety and quality but introduce high computational latency and cost that make them impractical for real-time or streaming deployments.
The article introduces Controlled Decoding (CD), a modular framework that aligns language models during text generation while leaving the underlying base model entirely frozen. The objective is to formulate a step-by-step reinforcement learning objective and solve it at inference time using an auxiliary prefix scorer module that estimates expected future rewards for partially generated text. The researchers evaluate two training methods for the scorer—CD-FUDGE and an off-policy value-learning method called CD-Q—alongside two generation strategies: individual word-level scoring and a blockwise approach that samples and evaluates small chunks of text.
The authors evaluate the framework across dialogue length control, helpfulness and harmlessness benchmarks, and text summarization using PaLM 2 model variants. The key findings demonstrate strong performance and operational flexibility. First, blockwise CD-Q matches the output quality of selecting the best full response while requiring significantly fewer generation candidates—achieving equivalent reward and model stability at up to ten times smaller sample sizes (for example, evaluating 6 candidates blockwise versus 50 candidates at the full sequence level). Second, blockwise CD-Q significantly outperforms established model-tuning techniques like Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO) on quality-versus-drift tradeoffs. Third, the trained prefix scorer generalizes successfully to unseen base models without requiring retraining or fine-tuning. Fourth, modular prefix scorers can be blended dynamically at runtime to handle multi-objective goals, such as increasing helpfulness while simultaneously constraining output length.
These results demonstrate that organizations can achieve robust model alignment without repeatedly fine-tuning large base models for every new safety standard or user preference. Keeping the base model frozen reduces compute costs, shortens deployment timelines, and allows live policy customization across multiple business objectives. Furthermore, blockwise decoding lowers first-token latency relative to full-sequence ranking, making structured output control feasible for interactive streaming products.
Technical leaders should consider adopting blockwise controlled decoding as a flexible alignment layer, particularly when multiple reward objectives are involved or when base models are updated frequently. For best operational performance, teams can combine lightweight base-model tuning with blockwise decoding at inference time to minimize candidate generation overhead. Before production deployment in high-stakes safety domains, organizations should conduct additional empirical validation. The authors note that the prefix scorer achieved lower classification accuracy on complex safety tasks compared to standard reward models, indicating that value estimation in noisy reward environments requires further refinement.
- Paper: Fine-Tuning Language Models from Human Preferences, Daniel M. Ziegler et al. (2019). This foundational KL-regularized RLHF work establishes the human-preference reward and reference-policy constraint that Controlled Decoding reformulates at the token level.
- Paper: Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint, Wei Xiong et al. (2024). Its treatment of KL-constrained preference learning supplies the RL objective and divergence-constraint context behind Controlled Decoding’s tokenwise formulation.
- Paper: Contrastive Decoding: Open-ended Text Generation as Optimization, Xiang Lisa Li et al. (2023). Its optimization-based decoding framework provides useful groundwork for understanding how Controlled Decoding steers generation at inference time.
- Paper: CTRL: A Conditional Transformer Language Model for Controllable Generation, Nitish Shirish Keskar et al. (2019). CTRL’s explicit control of language-model generation offers an earlier foundation for the problem of steering outputs while preserving a pretrained model.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). DeAL continues inference-time alignment by using lookahead search and multiple scoring objectives rather than Controlled Decoding’s learned prefix value scorer.
- Paper: Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback, Yafu Li et al. (2025). Test-Time Preference Optimization extends inference-time alignment into iterative critique and revision driven by reward-model feedback.
- Paper: Theoretical guarantees on the best-of-n alignment policy, Ahmad Beirami et al. (2025). This work develops theoretical guarantees for best-of-n alignment, providing a formal continuation of Controlled Decoding’s bridge between sampling-based selection and tokenwise control.
