DeAL: Decoding-time Alignment for Large Language Models

James Y. HuangSailik SenguptaDaniele BonadimanYi-An LaiArshit GuptaNikolaos PappasSaab MansourKatrin KirchhoffDan Roth

article2025ACL62 citations

Proposes DeAL, a framework that aligns large language models at decoding time by framing generation as heuristic search, enabling fine-grained control over customizable and modular reward functions without requiring model fine-tuning.

Listen

Modern large language models are expected to adhere to human preferences such as safety, factual accuracy, and task-specific constraints. Standard alignment methods primarily rely on fine-tuning models during training using human preference data, or providing alignment guidelines through system prompts. However, training-time alignment binds models to static, non-universal principles that require costly retraining to update, while prompting methods remain fragile and vulnerable to jailbreaks. The article introduces and evaluates DeAL, a framework that enforces custom, multi-faceted alignment objectives dynamically during decoding time without retraining the base model.

DeAL formulates text generation as a heuristic-guided search problem. During token generation, the framework evaluates candidate paths using lookahead mechanisms paired with scoring functions. These scoring functions can encompass both programmatic rules, such as keyword inclusion and word limits, and abstract preference models, such as helpfulness and harmlessness reward estimators. The authors evaluated the approach across multiple open-source language models on keyword generation, length-constrained summarization, and open-ended harmful prompt benchmarks.

Key findings show that DeAL substantially improves adherence to alignment constraints. On keyword generation, it increased full constraint satisfaction by an average of 17 percentage points over prompting baselines. In summarization tasks under strict length constraints, pairing prompt instructions with DeAL increased length compliance to between 53% and 73%, compared to 3% to 16% for prompting alone, while preserving summarization quality. For abstract safety objectives, DeAL achieved 100% harmlessness on malicious prompt benchmarks where standard safety prompting failed 37% of the time. When tested against adversarial continuation attacks, DeAL maintained a 73% harmlessness rate, whereas system prompt safeguards collapsed to 20% harmlessness. Furthermore, combining DeAL with reinforcement learning from human feedback yielded the highest overall safety and helpfulness scores.

These findings indicate that decoding-time alignment provides an effective, modular guardrail for real-time risk mitigation and compliance. It allows organizations to adjust trade-offs between competing goals dynamically without incurring the significant expense and time of retraining models. The results also show that fine-tuning alone can provide a false sense of security, as decoding-time guidance can easily override training-time behavior.

Organizations deploying language models in high-risk environments should consider implementing decoding-time verification alongside existing training safeguards rather than relying solely on prompting. When doing so, teams should evaluate latency trade-offs, as DeAL introduces a two- to five-fold slowdown during generation. Future work is needed to optimize efficiency through speculative decoding or compiled grammars and to test performance across proprietary platforms that currently restrict access to token probability outputs.

arXiv: 2402.06147
Cover for DeAL: Decoding-time Alignment for Large Language Models

Abstract

Large Language Models (LLMs) are nowadays expected to generate content aligned with human preferences. Current work focuses on alignment at model training time, through techniques such as Reinforcement Learning with Human Feedback (RLHF). However, it is unclear if such methods are an effective choice to teach alignment objectives to the model. First, the inability to incorporate multiple, custom rewards and reliance on a model developer’s view of universal and static principles are key limitations. Second, the reliability of such approaches is also questionable (e.g. susceptibility to jailbreaking even after safety training). To address these issues, we propose DeAL, a framework that allows the user to customize reward functions and enables Decoding-time ALignment of LLMs. At its core, we view decoding as a heuristic-guided search process and facilitate the use of a wide variety of alignment objectives. Our experiments with programmatic constraints such as keyword and length constraints, and abstract alignment objectives such as harmlessness and helpfulness, show that we can DeAL with fine-grained trade-offs and improve adherence to alignment objectives. Lastly, we demonstrate that DeAL is largely complementary to existing alignment strategies, and can be effectively paired with RLHF and prompting techniques to achieve better alignment.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 The Search Problem
  • 3.2 The Search Agent
  • 3.2.1 Start-state Adaptation
  • 3.2.2 Action Selection
  • 4 Experiments
  • 4.1 Programmatically Verifiable Objectives
  • 4.1.1 Keyword Constraints
  • 4.1.2 Length-constrained Summarization
  • 4.2 Abstract Alignment Objectives
  • 4.2.1 Validating Alignment Adherence
  • 4.2.2 Calibrating Multiple Objectives
  • 4.2.3 Combination with RLHF
  • 4.3 Defense against Jailbreaking
  • 5 Conclusions
  • Limitations
  • Impact Statement
  • References
  • A Task Details and Examples
  • A.1 Keyword Constraints
  • A.1.1 Falcon-7B-Instruct
  • A.1.2 MPT-7B-Instruct
  • A.1.3 Dolly-v2-3B
  • A.2 Length-Constrained Summarization
  • A.2.1 Falcon-7B-Instruct
  • A.2.2 MPT-7B-Instruct
  • B Decoding-time Approaches for enabling a Helpful and Harmless Assistant
  • B.1 Comparison with Decoding-time Baselines
  • B.2 Combining Multiple Reward Functions
  • B.3 Working with fine-tuning approaches
  • C Continuation Attack Examples
  • D Breaking Fine-tuning time Alignment with DeAL

Knowls

  1. Knowl 1 — Search-Based Alignment Formulation in DeAL

    model/method

    Decoding-time Alignment (DeAL) formalizes aligning autoregressive Large Language Models (LLMs) to human preferences as a search problem defined by the tuple:

    ⟨S,V,T,Ra⟩\langle S, V, T, R_a \rangle

    where:

    • SS is the state space of token sequences ⟨v1,v2,… ⟩\langle v^1, v^2, \dots \rangle.
    • VV is the action space defined by the model vocabulary (typically ∣V∣≈30,000|V| \approx 30{,}000).
    • T:S×V→ST : S \times V \to S is a deterministic transition function appending an action token v′∈Vv' \in V to the current sequence state ⟨v1,…,vn⟩\langle v^1, \dots, v^n \rangle to yield ⟨v1,…,vn,v′⟩\langle v^1, \dots, v^n, v' \rangle.
    • RaR_a is an alignment reward function reflecting the target alignment objective.

    The initial start state or prompt p∈Sp \in S is partitioned into three distinct components:

    p=(pt,pa,pi)p = (p_t, p_a, p_i)

    where ptp_t denotes the primary task instruction (optionally containing in-context demonstrations), pap_a represents alignment/system instructions specifying constraints or behavioral guidelines (which can be empty ϕ\phi if private or unspecifiable in natural language), and pip_i is the task input text. The goal state corresponds to any valid sequence ending with the end-of-sentence token ⟨v1,…,∣eos∣⟩\langle v^1, \dots, |\text{eos}| \rangle.

  2. Knowl 2 — Lookahead-Guided Action Selection Criterion in DeAL

    equation

    To select the next token action during decoding without exhaustively traversing the full vocabulary space VV, DeAL restricts candidate generation at step tt to the top-kk tokens V′⊂VV' \subset V according to the autoregressive model's predicted next-token distribution. For each candidate token v′∈V′v' \in V', a greedy lookahead sequence of length ll is generated to provide a sufficiently complete context y1:t+ly_{1:t+l} for the alignment heuristic scoring function h(⋅)h(\cdot).

    The candidate action is selected to maximize the scoring criterion:

    c(yt)=log⁡P(y1:t∣p)+λh(y1:t+l,p)c(y_t) = \log P(y_{1:t} \mid p) + \lambda h(y_{1:t+l}, p)

    where:

    • y1:ty_{1:t} is the sequence generated up to step tt.
    • pp is the start-state prompt (pt,pa,pi)(p_t, p_a, p_i).
    • ll is the predetermined lookahead length (in tokens).
    • h(y1:t+l,p)h(y_{1:t+l}, p) is an alignment heuristic function (higher values indicate closer adherence to alignment objectives), which can be programmatic (e.g., keyword presence, length count) or parametric (e.g., reward model logits).
    • λ≥0\lambda \ge 0 is a scalar hyperparameter weighting the influence of the alignment heuristic relative to the language model probability.
  3. Knowl 3 — Modular Multi-Objective Reward Calibration at Decoding Time

    model/method

    DeAL enables modular and flexible multi-objective alignment during inference by ensembling distinct reward estimators without requiring model retraining or multi-task fine-tuning. For competing or complementary alignment goals (such as helpfulness and harmlessness), the heuristic function h(⋅)h(\cdot) can be composed via an affine combination of individual reward models:

    Ra=∑m=1MwmRmR_a = \sum_{m=1}^M w_m R_m

    where {Rm}m=1M\{R_m\}_{m=1}^M are individual reward models (e.g., RharmlessR_{\text{harmless}} and RhelpfulR_{\text{helpful}} trained on respective preference subsets), and wmw_m are user-specified non-negative scalar weights controlling the relative priority of each objective. Adjusting the weights {wm}\{w_m\} allows dynamic, post-hoc steerability between competing objectives (e.g., safety rejection versus detailed helpfulness) at generation time.

  4. Knowl 4 — Comparative Evaluation on Harmlessness and Helpfulness Across Decoding Strategies

    data/table

    The effectiveness of DeAL using parametric reward models (fine-tuned OPT-125M models: RharmlessR_{\text{harmless}}, RhelpfulR_{\text{helpful}}, and joint RhhR_{hh}) was evaluated against baseline prompting and reranking approaches using MPT-7B-Instruct on the out-of-domain malicious prompt benchmark HarmfulQ and the in-domain HH-RLHF test splits.

    Method HarmfulQ Harmless HH-RLHF Harmless HH-RLHF Helpful
    Base 0.43 0.40 0.33
    pap_a (for safety) 0.63 0.43 0.60
    Harmless rerank 0.40 0.47 0.53
    Helpful rerank 0.37 0.40 0.57
    DeAL w/ RharmlessR_{\text{harmless}} 1.00 0.57 0.23
    DeAL w/ RhelpfulR_{\text{helpful}} 0.20 0.37 0.77
    DeAL w/ RhhR_{hh} 1.00 0.67 0.67

    Safety prompting (pap_a) improves harmlessness and helpfulness over the unprompted base model but still produces harmful content 37% of the time on HarmfulQ and 57% of the time on the HH-RLHF harmless set. Standard candidate reranking at the end of generation performs poorly because misaligned generation paths are difficult to recover from post-hoc. DeAL w/ RhhR_{hh} achieves the best overall performance, outperforming safety prompting by 37%, 24%, and 7% on HarmfulQ Harmless, HH-RLHF Harmless, and HH-RLHF Helpful, respectively.

  5. Knowl 5 — Complementary Combination of Training-Time RLHF and Decoding-Time DeAL

    data/table

    To evaluate the synergy between training-time alignment (RLHF) and decoding-time alignment (DeAL), Dolly-v2-3B was evaluated on HarmfulQ and HH-RLHF using the joint reward model RhhR_{hh}.

    Method HarmfulQ Harmless HH-RLHF Helpful
    No RLHF, No DeAL 0.33 0.43
    DeAL w/ RharmlessR_{\text{harmless}} 0.83 0.33
    DeAL w/ RhelpfulR_{\text{helpful}} 0.10 0.70
    RLHF w/ RhhR_{hh} 0.80 0.70
    DeAL w/ RhhR_{hh} 0.83 0.53
    RLHF + DeAL 0.93 0.70

    Because RLHF operates during model training while DeAL operates during token generation, the two methods are orthogonal and complementary. Combining RLHF and DeAL achieves the highest performance on both harmlessness (0.93) and helpfulness (0.70), exceeding either RLHF alone (0.80 harmless) or DeAL alone (0.83 harmless).

  6. Knowl 6 — Robustness of Decoding-Time Alignment Against Continuation Jailbreak Attacks

    data/table

    Language models can be jailbroken by appending an affirmative continuation prefix such as "Assistant: To do that," at the end of harmful input prompts from HarmfulQ, exploiting the model's next-token autoregressive completion objective.

    Method HarmfulQ Harmless
    None 0.17
    pap_a (for safety) 0.20
    DeAL w/ RharmlessR_{\text{harmless}} 0.73

    Under this continuation attack, safety prompts (pap_a) are largely circumvented, reducing harmlessness from 0.63 down to 0.20 (comparable to the unprompted setting at 0.17). In contrast, DeAL guided by RharmlessR_{\text{harmless}} enforces constraints at the token selection level, preventing harmful continuations 73% of the time despite the forced affirmative prefix.

  7. Knowl 7 — Steerability of Harmlessness vs Helpfulness via Reward Weight Variation

    data/table

    By varying the relative heuristic weights (wharmless,whelpful)(w_{\text{harmless}}, w_{\text{helpful}}) of two distinct reward models (RharmlessR_{\text{harmless}} and RhelpfulR_{\text{helpful}}) in DeAL with MPT-7B-Instruct, the model's outputs can be calibrated smoothly across safety and utility objectives.

    Method (wharmless,whelpful)(w_{\text{harmless}}, w_{\text{helpful}}) HarmfulQ Harmless HH-RLHF Harmless HH-RLHF Helpful
    DeAL w/ RhhR_{hh} 1.00 0.67 0.67
    DeAL (1.00,0.00)(1.00, 0.00) 1.00 0.57 0.23
    DeAL (0.75,0.25)(0.75, 0.25) 1.00 0.57 0.34
    DeAL (0.50,0.50)(0.50, 0.50) 0.77 0.57 0.48
    DeAL (0.25,0.50)(0.25, 0.50) 0.43 0.40 0.67
    DeAL (0.00,1.00)(0.00, 1.00) 0.20 0.37 0.77

    As wharmlessw_{\text{harmless}} decreases and whelpfulw_{\text{helpful}} increases, the generation behavior systematically transitions from highly conservative safety compliance (1.00 harmless on HarmfulQ, 0.23 helpful on HH-RLHF) to maximum instruction fulfillment (0.20 harmless on HarmfulQ, 0.77 helpful on HH-RLHF).

  8. Knowl 8 — Keyword-Constrained Sentence Generation on CommonGen

    data/table

    DeAL was evaluated on CommonGen using three instruction-tuned LLMs with candidate size k=10k = 10, lookahead length l=32l = 32 tokens, and hard keyword coverage as the heuristic h(⋅)h(\cdot). Soft coverage measures the average fraction of keywords included per instance; hard coverage measures the fraction of instances containing all specified keywords.

    Model Method Coverage (soft) Coverage (hard)
    Falcon-7B-Instruct pap_a 0.88 0.62
    Falcon-7B-Instruct pa+DeALp_a + \text{DeAL} 0.94 0.80
    MPT-7B-Instruct pap_a 0.91 0.71
    MPT-7B-Instruct pa+DeALp_a + \text{DeAL} 0.96 0.85
    Dolly-v2-3B pap_a 0.65 0.30
    Dolly-v2-3B pa+DeALp_a + \text{DeAL} 0.79 0.51

    Applying DeAL consistently increases keyword coverage across all models, yielding an average improvement of 8% in soft coverage and 17% in hard coverage over prompt-only baselines (pap_a). Absolute gains in hard coverage are largest for models with weaker baseline instruction following (+21% for Dolly-v2-3B, +18% for Falcon-7B-Instruct, +14% for MPT-7B-Instruct).

  9. Knowl 9 — Length-Constrained Summarization on XSUM

    data/table

    DeAL was tested on length-constrained summarization ("≤10 words""\le 10\text{ words}") using 176 XSUM instances with human reference summaries ≤10\le 10 words. Search parameters were candidate size k=5k = 5, lookahead length l=32l = 32 tokens, and a binary length constraint satisfaction heuristic. Quality metrics evaluated by human annotators were Length Satisfaction (LS), Faithfulness (F, binary 0/1), Relevance (R, 1-5 Likert), and Coherence (C, 1-5 Likert).

    Model Method LS F R C
    Falcon-7B-Instruct pap_a 0.16 0.79 4.21 4.72
    Falcon-7B-Instruct DeAL 0.44 0.48 4.15 4.45
    Falcon-7B-Instruct pa+DeALp_a + \text{DeAL} 0.73 0.72 4.04 4.66
    MPT-7B-Instruct pap_a 0.03 0.86 4.66 4.93
    MPT-7B-Instruct DeAL 0.53 0.79 4.34 4.83
    MPT-7B-Instruct pa+DeALp_a + \text{DeAL} 0.53 0.86 4.31 4.97

    Prompting alone (pap_a) satisfies the length constraint in only 16% (Falcon) and 3% (MPT) of summaries. Combining pap_a with DeAL achieves the highest length satisfaction (0.73 on Falcon, 0.53 on MPT) with no statistically significant drop in Faithfulness, Relevance, or Coherence (p≫0.05p \gg 0.05, Wilcoxon-Mann-Whitney test). Omitting pap_a when running DeAL (setting pa=ϕp_a = \phi) causes quality degradation because instruction-tuned LLMs default to generating long summaries, meaning the unsteered top-k=5k = 5 candidate pool lacks high-quality short summary branches.

  10. Knowl 10 — Latency Overhead and API Access Limitations of DeAL

    limitation

    DeAL has two primary operational limitations:

    1. Computational Latency: Using lookahead generation and evaluating parametric reward models at each decoding step increases inference latency by 2×2\times to 5×5\times compared to standard greedy decoding without batch-inference optimizations.
    2. Logit Access Dependency: The search framework requires direct access to next-token logits and probability distributions over the vocabulary, preventing its application to closed proprietary LLM APIs that only return final text completions or provide limited biasing capabilities (e.g., logit_bias).

Coverage note — None was omitted; all key algorithmic formulations, mathematical scoring rules, and empirical results (keyword constraints, length constraints, multi-objective calibration, RLHF synergy, jailbreak defense, and limitations) were extracted.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022a. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
  2. 2.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022b. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
  3. 3.Amanda Bertsch, Alex Xie, Graham Neubig, and Matthew R Gormley. 2023. It’s mbr all the way down: Modern generation techniques through the lens of minimum bayes risk. arXiv preprint arXiv:2310.01387.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  6. 6.Haikang Deng and Colin Raffel. 2023. Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11781–11791, Singapore. Association for Computational Linguistics.
  7. 7.Daniel Deutsch, Shyam Upadhyay, and Dan Roth. 2019. A general-purpose algorithm for constrained sequential inference. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 482–492, Hong Kong, China. Association for Computational Linguistics.
  8. 8.Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023. Raft: Reward ranked finetuning for generative foundation model alignment. arXiv preprint arXiv:2304.06767.
  9. 9.Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388.
  10. 10.Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
  11. 11.Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics.
  12. 12.Saibo Geng, Martin Josifosky, Maxime Peyrard, and Robert West. 2023. Flexible grammar-based constrained decoding for language models. arXiv preprint arXiv:2305.13971.
  13. 13.Aria Haghighi, John DeNero, and Dan Klein. 2007. Approximate factoring for a* search. In Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference, pages 412–419.
  14. 14.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In International Conference on Learning Representations.
  15. 15.Mark Hopkins and Greg Langmead. 2009. Cube pruning as heuristic search. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing, pages 62–71.
  16. 16.Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825.
  17. 17.Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088.
  18. 18.J Joshua. 2023. Chatgpt api transition guide.
  19. 19.Muhammad Khalifa, Yogarshi Vyas, Shuai Wang, Graham Horwood, Sunil Mallya, and Miguel Ballesteros. 2023. Contrastive training improves zero-shot classification of semi-structured documents. In Findings of the Association for Computational Linguistics: ACL 2023, pages 7499–7508, Toronto, Canada. Association for Computational Linguistics.
  20. 20.Maxim Khanov, Jirayu Burapacheep, and Yixuan Li. 2024. ARGS: Alignment as reward-guided search. In The Twelfth International Conference on Learning Representations.
  21. 21.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021. GeDi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4929–4952, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  22. 22.Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019. Importance of search and evaluation strategies in neural dialogue modeling. In Proceedings of the 12th International Conference on Natural Language Generation, pages 76–87, Tokyo, Japan. Association for Computational Linguistics.
  23. 23.Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274–19286. PMLR.
  24. 24.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016a. A diversity-promoting objective function for neural conversation models. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119, San Diego, California. Association for Computational Linguistics.
  25. 25.Jiwei Li, Will Monroe, and Dan Jurafsky. 2016b. A simple, fast diverse decoding algorithm for neural generation. arXiv preprint arXiv:1611.08562.
  26. 26.Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. 2024. RAIN: Your language models can align themselves without finetuning. In The Twelfth International Conference on Learning Representations.
  27. 27.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text generation challenge for generative commonsense reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  28. 28.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment, may 2023. arXiv preprint arXiv:2303.16634.
  29. 29.Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, and Yejin Choi. 2022. NeuroLogic a*esque decoding: Constrained text generation with lookahead heuristics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 780–799, Seattle, United States. Association for Computational Linguistics.
  30. 30.Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. NeuroLogic decoding: (un)supervised neural text generation with predicate logic constraints. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4299, Online. Association for Computational Linguistics.
  31. 31.Clara Meister, Tim Vieira, and Ryan Cotterell. 2020. Best-first beam search. Transactions of the Association for Computational Linguistics, 8:795–809.
  32. 32.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  33. 33.Franz Josef Och, Nicola Ueffing, and Hermann Ney. 2001. An efficient a* search algorithm for statistical machine translation. In Proceedings of the ACL 2001 Workshop on Data-Driven Methods in Machine Translation.
  34. 34.OpenAI. 2023a. Using logit bias to define token probability | OpenAI Help Center — help.openai.com.
  35. 35.R OpenAI. 2023b. Gpt-4 technical report. arXiv, pages 2303–08774.
  36. 36.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  37. 37.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116.
  38. 38.Lianhui Qin, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022. Cold decoding: Energy-based constrained text generation with langevin dynamics. Advances in Neural Information Processing Systems, 35:9538–9551.
  39. 39.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  40. 40.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. arXiv preprint arXiv:2305.18290.
  41. 41.Shamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Mansour, and Arshit Gupta. 2024. Flap: Flow adhering planning with constrained decoding in llms. arXiv preprint arXiv:2403.05766.
  42. 42.Sailik Sengupta, He He, Batool Haider, Spandana Gella, and Mona Diab. 2019. Natural language generation with keyword constraints– a hybrid approach using supervised and reinforcement learning. West Coast NLP (WeCNLP).
  43. 43.Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4454–4470, Toronto, Canada. Association for Computational Linguistics.
  44. 44.Raphael Shu and Hideki Nakayama. 2018. Improving beam search by removing monotonic constraint for neural machine translation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 339–344, Melbourne, Australia. Association for Computational Linguistics.
  45. 45.Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023. Preference ranking optimization for human alignment. arXiv preprint arXiv:2306.17492.
  46. 46.MosaicML NLP Team. 2023. Introducing mpt-7b: A new standard for open-source, commercially usable llms.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  48. 48.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  49. 49.David Wan, Mengwen Liu, Kathleen McKeown, Markus Dreyer, and Mohit Bansal. 2023a. Faithfulness-aware decoding strategies for abstractive summarization. arXiv preprint arXiv:2303.03278.
  50. 50.David Wan, Mengwen Liu, Kathleen McKeown, Markus Dreyer, and Mohit Bansal. 2023b. Faithfulness-aware decoding strategies for abstractive summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2864–2880, Dubrovnik, Croatia. Association for Computational Linguistics.
  51. 51.David Wan, Mengwen Liu, Kathleen McKeown, Dreyer Markus, and Mohit Bansal. 2023c. Faithfulness-aware decoding strategies for abstractive summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics.
  52. 52.Bailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao, Rif A Saurous, and Yoon Kim. 2023a. Grammar prompting for domain-specific language generation with large language models. arXiv preprint arXiv:2305.19234.
  53. 53.Shufan Wang, Sebastien Jean, Sailik Sengupta, James Gung, Nikolaos Pappas, and Yi Zhang. 2023b. Measuring and mitigating constraint violations of in-context learning for utterance-to-api semantic parsing. arXiv preprint arXiv:2305.15338.
  54. 54.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483.
  55. 55.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837.
  56. 56.Sean Welleck, Jiacheng Liu, Jesse Michael Han, and Yejin Choi. 2021. Towards grounded natural language proof generation. In MathAI4Ed Workshop at NeurIPS.
  57. 57.Brandon T Willard and Rémi Louf. 2023. Efficient guided generation for large language models. arXiv e-prints, pages arXiv–2307.
  58. 58.Seungpil Won, Heeyoung Kwak, Joongbo Shin, Janghoon Han, and Kyomin Jung. 2023. BREAK: Breaking the dialogue state tracking barrier with beam search and re-ranking. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2832–2846, Toronto, Canada. Association for Computational Linguistics.
  59. 59.Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao, Min-Yen Kan, Junxian He, and Michael Xie. 2024. Self-evaluation guided beam search for reasoning. Advances in Neural Information Processing Systems, 36.
  60. 60.Kevin Yang and Dan Klein. 2021. FUDGE: Controlled text generation with future discriminators. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3511–3535, Online. Association for Computational Linguistics.
  61. 61.Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023. Rrhf: Rank responses to align language models with human feedback without tears. arXiv preprint arXiv:2304.05302.
  62. 62.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  63. 63.Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto. 2023. Benchmarking large language models for news summarization. arXiv preprint arXiv:2301.13848.
  64. 64.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

Citation

MLA
Huang, J. Y., et al. “DeAL: Decoding-time Alignment for Large Language Models”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 26280–300, https://doi.org/10.18653/v1/2025.acl-long.1274.
APA
Huang, J. Y., Sengupta, S., Bonadiman, D., Lai, Y.-A., Gupta, A., Pappas, N., Mansour, S., Kirchhoff, K., & Roth, D. (2025). DeAL: Decoding-time Alignment for Large Language Models. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 26280–26300. https://doi.org/10.18653/v1/2025.acl-long.1274
Chicago
Huang, J. Y., S. Sengupta, D. Bonadiman, et al. 2025. “DeAL: Decoding-time Alignment for Large Language Models”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 26280–300. https://doi.org/10.18653/v1/2025.acl-long.1274.
Harvard
Huang, J.Y. et al. (2025) “DeAL: Decoding-time Alignment for Large Language Models”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 26280–26300. Available at: https://doi.org/10.18653/v1/2025.acl-long.1274.
Vancouver
1. Huang JY, Sengupta S, Bonadiman D, Lai Y-A, Gupta A, Pappas N, Mansour S, Kirchhoff K, Roth D (2025) DeAL: Decoding-time Alignment for Large Language Models. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 26280–26300

BibTeX

@inproceedings{huang-etal-2025-deal,
    title = "{D}e{AL}: Decoding-time Alignment for Large Language Models",
    author = "Huang, James Y.  and
      Sengupta, Sailik  and
      Bonadiman, Daniele  and
      Lai, Yi-An  and
      Gupta, Arshit  and
      Pappas, Nikolaos  and
      Mansour, Saab  and
      Kirchhoff, Katrin  and
      Roth, Dan",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.1274/",
    doi = "10.18653/v1/2025.acl-long.1274",
    pages = "26280--26300",
    ISBN = "979-8-89176-251-0"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/