Controlled Decoding from Language Models

Sidharth MudgalJong LeeHarish GanapathyYaGuang LiTao WangYanping HuangZhifeng ChenHeng-Tze ChengMichael CollinsTrevor Strohman

article2024ICML157 citations

Proposes a modular controlled decoding framework that aligns frozen language models at inference time using trained prefix scorers, enabling multi-objective control and zero-shot transfer across base models without retraining.

Listen

Aligning large language models with human preferences typically requires updating base model weights through reinforcement learning or preference optimization. While effective, these training-time interventions are computationally expensive, inflexible when reward criteria shift, and prone to policy degradation. Standard inference-time alternatives, such as sampling many full candidate outputs and picking the best one, improve safety and quality but introduce high computational latency and cost that make them impractical for real-time or streaming deployments.

The article introduces Controlled Decoding (CD), a modular framework that aligns language models during text generation while leaving the underlying base model entirely frozen. The objective is to formulate a step-by-step reinforcement learning objective and solve it at inference time using an auxiliary prefix scorer module that estimates expected future rewards for partially generated text. The researchers evaluate two training methods for the scorer—CD-FUDGE and an off-policy value-learning method called CD-Q—alongside two generation strategies: individual word-level scoring and a blockwise approach that samples and evaluates small chunks of text.

The authors evaluate the framework across dialogue length control, helpfulness and harmlessness benchmarks, and text summarization using PaLM 2 model variants. The key findings demonstrate strong performance and operational flexibility. First, blockwise CD-Q matches the output quality of selecting the best full response while requiring significantly fewer generation candidates—achieving equivalent reward and model stability at up to ten times smaller sample sizes (for example, evaluating 6 candidates blockwise versus 50 candidates at the full sequence level). Second, blockwise CD-Q significantly outperforms established model-tuning techniques like Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO) on quality-versus-drift tradeoffs. Third, the trained prefix scorer generalizes successfully to unseen base models without requiring retraining or fine-tuning. Fourth, modular prefix scorers can be blended dynamically at runtime to handle multi-objective goals, such as increasing helpfulness while simultaneously constraining output length.

These results demonstrate that organizations can achieve robust model alignment without repeatedly fine-tuning large base models for every new safety standard or user preference. Keeping the base model frozen reduces compute costs, shortens deployment timelines, and allows live policy customization across multiple business objectives. Furthermore, blockwise decoding lowers first-token latency relative to full-sequence ranking, making structured output control feasible for interactive streaming products.

Technical leaders should consider adopting blockwise controlled decoding as a flexible alignment layer, particularly when multiple reward objectives are involved or when base models are updated frequently. For best operational performance, teams can combine lightweight base-model tuning with blockwise decoding at inference time to minimize candidate generation overhead. Before production deployment in high-stakes safety domains, organizations should conduct additional empirical validation. The authors note that the prefix scorer achieved lower classification accuracy on complex safety tasks compared to standard reward models, indicating that value estimation in noisy reward environments requires further refinement.

arXiv: 2310.17022
Cover for Controlled Decoding from Language Models

Abstract

KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called controlled decoding (CD). CD exerts control through a separate prefix scorer module, which is trained to learn a value function for the reward. The prefix scorer is used at inference time to control the generation from a frozen base model, provably sampling from a solution to the RL objective. We empirically demonstrate that CD is effective as a control mechanism on popular benchmarks. We also show that prefix scorers for multiple rewards may be combined at inference time, effectively solving a multi-objective RL problem with no additional training. We show that the benefits of applying CD transfer to an unseen base model with no further tuning as well. Finally, we show that CD can be applied in a blockwise decoding fashion at inference-time, essentially bridging the gap between the popular best-of-K strategy and tokenwise control through reinforcement learning. This makes CD a promising approach for alignment of language models.

Table of Contents

  • 1. Introduction
  • 2. KL-Regularized Reinforcement Learning
  • 3. Controlled Decoding
  • 3.1. Training the prefix scorer
  • 3.2. Inference-time sampling strategies
  • 4. Experimental Setup
  • 4.1. Datasets
  • 4.2. Reward Models
  • 4.3. Baselines
  • 4.4. Evaluation Metrics
  • 4.5. Training Details
  • 5. Experimental Results
  • 6. Related Work
  • 7. Concluding Remarks
  • Impact Statement
  • Acknowledgements
  • References
  • A. Additional details on experimental setup
  • Zero-shot prompts.
  • B. Additional experimental results
  • C. Proofs

Knowls

  1. Knowl 1 — Tokenwise KL-regularized alignment objective

    equation

    Let xx be a prompt, yty^t a partially decoded response of tt tokens, and Y\mathcal{Y} the token vocabulary. The frozen reference language model is πref\pi_{\mathrm{ref}}, and π\pi is a candidate distribution for the next token at a nonterminal prefix. The response receives a bounded-above scalar reward r([x,y])r([x,y]) only when generation ends: the tokenwise reward is r([x,yt])r([x,y^t]) if the last token is EOS and 00 otherwise. Define V∗([x,yt])V^*([x,y^t]) as the expected final reward when generation continues from that prefix under πref\pi_{\mathrm{ref}}. The advantage and next-token KL divergence are

    A([x,yt];π)=∑z∈Yπ(z∣[x,yt])V∗([x,yt,z])−V∗([x,yt]),D([x,yt];π)=KL⁡ ⁣(π(⋅∣[x,yt]) ∥ πref(⋅∣[x,yt])).A([x,y^t];\pi)=\sum_{z\in\mathcal{Y}}\pi(z\mid[x,y^t])V^*([x,y^t,z])-V^*([x,y^t]),\qquad D([x,y^t];\pi)=\operatorname{KL}\!\left(\pi(\cdot\mid[x,y^t])\,\middle\|\,\pi_{\mathrm{ref}}(\cdot\mid[x,y^t])\right).

    The proposed objective is Jλ([x,yt];π)=λA([x,yt];π)−D([x,yt];π)J_\lambda([x,y^t];\pi)=\lambda A([x,y^t];\pi)-D([x,y^t];\pi), where λ≥0\lambda\geq 0 sets the reward-versus-deviation tradeoff. At λ=0\lambda=0, the reference policy is optimal and has zero advantage and zero KL divergence.

  2. Knowl 2 — Closed-form optimal policy for the tokenwise objective

    theoretical result

    For a prompt xx and a nonterminal prefix yty^t, the unique next-token policy maximizing Jλ([x,yt];π)J_\lambda([x,y^t];\pi) is an exponential tilt of the frozen reference language model:

    πλ∗(z∣[x,yt])=πref(z∣[x,yt])exp⁡ ⁣(λV∗([x,yt,z]))∑z′∈Yπref(z′∣[x,yt])exp⁡ ⁣(λV∗([x,yt,z′])).\pi^*_\lambda(z\mid[x,y^t])= \frac{\pi_{\mathrm{ref}}(z\mid[x,y^t])\exp\!\left(\lambda V^*([x,y^t,z])\right)} {\sum_{z'\in\mathcal{Y}}\pi_{\mathrm{ref}}(z'\mid[x,y^t])\exp\!\left(\lambda V^*([x,y^t,z'])\right)}.

    Here z,z′z,z' range over the token vocabulary Y\mathcal{Y}, V∗V^* is the expected terminal reward under continuations from the extended prefix, and λ≥0\lambda\geq0 is the KL-regularization tradeoff. Thus a value scorer can control generation without changing the reference model’s parameters; the exact optimality claim uses the true value function.

  3. Knowl 3 — Controlled decoding as value-function learning with a frozen generator

    model/method

    Controlled decoding (CD) keeps the base language model πref\pi_{\mathrm{ref}} frozen and trains a separate prefix scorer Vθ([x,yt])V_\theta([x,y^t]) to approximate V∗([x,yt])V^*([x,y^t]), the expected terminal reward from prompt xx and partial response yty^t. The paper’s ideal fitting objective is

    L∗(θ)=Ex∼μ, y∼πref(⋅∣x)[12∑t∈[∣y∣](Vθ([x,yt])−V∗([x,yt]))2],L^*(\theta)=\mathbb{E}_{x\sim\mu,\,y\sim\pi_{\mathrm{ref}}(\cdot\mid x)}\left[\frac{1}{2}\sum_{t\in[|y|]}\left(V_\theta([x,y^t])-V^*([x,y^t])\right)^2\right],

    where μ\mu is the training-prompt distribution, yy is a response sampled from the reference model, and [∣y∣][|y|] indexes its decoded prefixes. The learned scorer is a modular inference-time controller; its quality depends on how well its training procedure estimates the value of partial responses.

  4. Knowl 4 — CD-FUDGE learns values from on-policy terminal rewards

    theoretical result

    CD-FUDGE trains a prefix scorer using complete responses sampled by rolling out the frozen reference model. For prompt xx and sampled response yy, its loss is

    LF(θ)=Ex∼μ, y∼πref(⋅∣x)[12∑t∈[∣y∣](Vθ([x,yt])−r([x,y]))2],L_F(\theta)=\mathbb{E}_{x\sim\mu,\,y\sim\pi_{\mathrm{ref}}(\cdot\mid x)}\left[\frac{1}{2}\sum_{t\in[|y|]}\left(V_\theta([x,y^t])-r([x,y])\right)^2\right],

    where r([x,y])r([x,y]) is the final response reward and every prefix on that rollout is trained toward that same reward. The paper shows that the expected gradient of this sampled loss equals the gradient of the ideal value-fitting objective. Under the stated regularity conditions—including a Lipschitz loss, a nonempty solution set, the Polyak–Łojasiewicz condition, and bounded expected squared gradients—stochastic gradient descent converges to a solution of the scorer objective. This gives a theoretical route from reference-model rollouts and terminal rewards to a value scorer usable for controlled decoding.

  5. Knowl 5 — CD-Q trains the prefix scorer with a Bellman target

    model/method

    CD-Q uses a bootstrapped value target rather than assigning the final reward to every prefix. For a nonterminal prefix [x,yt][x,y^t], its target is

    vt=∑z∈Yπref(z∣[x,yt])Vθ([x,yt,z]);v_t=\sum_{z\in\mathcal{Y}}\pi_{\mathrm{ref}}(z\mid[x,y^t])V_\theta([x,y^t,z]);

    at an EOS prefix, the target is the observed terminal reward r([x,yt])r([x,y^t]). The scorer minimizes the squared error between Vθ([x,yt])V_\theta([x,y^t]) and a stop-gradient version of this target, summed over prefixes in training responses. Responses may come from an off-policy dataset scored with the reward function, but computing each nonterminal target requires on-demand access to the reference model’s next-token distribution. CD-Q therefore learns values through policy evaluation while leaving the generator unchanged.

  6. Knowl 6 — Tokenwise controlled decoding uses value-weighted next-token sampling

    model/method

    At each nonterminal prefix [x,yt][x,y^t], tokenwise CD scores every possible next token with the learned prefix scorer and samples from

    πθ(z∣[x,yt])∝πref(z∣[x,yt])exp⁡ ⁣(λVθ([x,yt,z])),\pi_\theta(z\mid[x,y^t])\propto\pi_{\mathrm{ref}}(z\mid[x,y^t])\exp\!\left(\lambda V_\theta([x,y^t,z])\right),

    where zz is a vocabulary token and λ≥0\lambda\geq0 controls the strength of reward shaping. Equivalently, the reference-model logits are adjusted by the value scores before sampling. High-value continuations gain probability and low-value continuations lose probability. With an exact value scorer this is the policy that solves the paper’s tokenwise KL-regularized objective; in practice the scorer is learned and is therefore an approximation. The method requires scorer evaluation at each decoding step.

  7. Knowl 7 — Blockwise CD reranks sampled continuation blocks

    model/method

    Blockwise CD combines prefix-value control with best-of-KK selection. Given a current prompt and decoded prefix [x,yt][x,y^t], it samples KK independent continuation blocks zM(1),…,zM(K)z_M^{(1)},\ldots,z_M^{(K)}, each of length MM, from the frozen reference model. It appends the block with the largest scorer value, Vθ([x,yt,zM(k)])V_\theta([x,y^t,z_M^{(k)}]), then repeats from the extended prefix until an accepted block contains EOS. Thus the scorer ranks partial continuations rather than complete responses, and the procedure can return a response after decoding only a block instead of waiting for all KK full candidate responses. Its inference throughput is comparable to best-of-KK at the same KK, while its selection latency is a block of MM tokens.

  8. Knowl 8 — Controlled decoding improves reward tradeoffs across three evaluated tasks

    empirical result

    The experiments used PaLM 2-XXS (Gecko) as both the default frozen generator and the starting model for prefix scorers. They evaluated response length on Reddit conversations, helpfulness and harmlessness (HH) on dialogue prompts, and summarization quality on TL;DR posts. Length reward was log⁡(T/1024)\log(T/1024) for response length TT; HH and summarization scorers were trained from pairwise preferences, with larger PaLM 2-L (Unicorn) used for the reported preference evaluations.

    For response length, tokenwise CD-Q and CD-FUDGE had more favorable reward–KL tradeoffs than the training-time baselines PPO, DPO, and IPO. Blockwise CD-Q was on par with best-of-KK while using substantially fewer candidates: the paper reports comparable length and KL tradeoffs with K=6K=6 for blockwise CD-Q versus K=50K=50 for best-of-KK. Blockwise CD-Q also outperformed the training-time baselines in this task. For HH, tokenwise controllers yielded little improvement, while blockwise CD-Q and CD-FUDGE improved win rate substantially but did not match best-of-KK. For summarization, blockwise CD-Q outperformed IPO but did not match best-of-KK.

    The learned prefix scorers were also evaluated as HH preference classifiers on 1,500 ground-truth side-by-side examples. The reported training and test accuracies were:

    MethodTraining accuracyTest accuracy
    Reward-XXS0.8040.709
    CD-FUDGE0.6320.629
    CD-Q0.6240.631

    The scorer accuracies were weaker than those of Reward-XXS, indicating that value prediction quality—particularly with noisy preference rewards—remains a practical issue even when a scorer is useful for decoding.

  9. Knowl 9 — Reward combinations and model changes can be handled without retraining CD scorers

    empirical result

    The paper demonstrates two forms of inference-time modularity. First, separately trained prefix scorers can be combined by linearly combining their reward scores at decoding time. In blockwise CD-FUDGE, adding a negative length score to HH-based reranking counteracted the response-length increase caused by HH-only control; this preserved length at the cost of lower HH win rate. This changes the reward tradeoff without retraining the controller for each combination.

    Second, a CD-Q scorer trained with PaLM 2-XXS (Gecko) was applied without retraining to PaLM 2-S (Bison) for length control and PaLM 2-XS (Otter) for HH control. The paper reports performance on par with the strongest baseline, best-of-KK, in both transfers. It also applied the same CD-Q scorer to a DPO-finetuned base model it had not seen during scorer training: combining DPO with blockwise CD-Q gave similar or better HH tradeoffs with fewer candidates, including a reported comparison of K=8K=8 versus K=32K=32 for a similar tradeoff.

  10. Knowl 10 — Blockwise control has empirical block-size and objective limitations

    limitation

    The paper does not establish that blockwise CD optimizes the sequence-level KL-regularized objective used by methods such as PPO and DPO; the exact optimal-policy result applies to tokenwise control. In HH experiments, larger block sizes generally improved win-rate-versus-KL tradeoffs, but block sizes above M=32M=32 were not tested because the authors expected the efficiency advantage over best-of-KK to disappear. Fixed-throughput experiments found blockwise CD-Q superior to DPO plus best-of-KK at the tested budgets K∈{4,8,16}K\in\{4,8,16\}, while DPO plus blockwise CD-Q and DPO plus best-of-KK were comparable at K=4K=4; these findings do not establish a general throughput or latency advantage across settings. More broadly, the authors caution that alignment methods can behave inconsistently on safety and other socially consequential objectives, and describe their throughput, latency, and performance study as narrow rather than comprehensive.

Coverage note — No substantial contributed method or result was deliberately omitted; proof-only derivations, related work, and implementation details of evaluation prompts were excluded.

References

  1. 1.Adolphs, L., Gao, T., Xu, J., Shuster, K., Sukhbaatar, S., and Weston, J. The cringe loss: Learning what language not to model. arXiv preprint arXiv:2211.05826, 2022.
  2. 2.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. PaLM 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  3. 3.Arora, K., Shuster, K., Sukhbaatar, S., and Weston, J. Director: Generator-classifiers for supervised language modeling. arXiv preprint arXiv:2206.07694, 2022.
  4. 4.Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R. A general theoretical paradigm to understand learning from human preferences. arXiv preprint arXiv:2310.12036, 2023.
  5. 5.Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das-Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022.
  6. 6.Beirami, A., Agarwal, A., Berant, J., D’Amour, A., Eisenstein, J., Nagpal, C., and Suresh, A. T. Theoretical guarantees on the best-of-n alignment policy. arXiv preprint arXiv:2401.01879, 2024.
  7. 7.Chaffin, A., Claveau, V., and Kijak, E. PPL-MCTS: Constrained textual generation through discriminator-guided MCTS decoding. In Carpuat, M., de Marneffe, M.-C., and Meza Ruiz, I. V. (eds.), Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2953–2967, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.215. URL https://aclanthology.org/2022.naacl-main.215.
  8. 8.Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023.
  9. 9.Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017.
  10. 10.Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D., Fisch, A., Heller, K., Pfohl, S., Ramachandran, D., et al. Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking. arXiv preprint arXiv:2312.09244, 2023.
  11. 11.Gao, L., Schulman, J., and Hilton, J. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, pp. 10835–10866. PMLR, 2023.
  12. 12.Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375, 2022.
  13. 13.Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. Rainbow: Combining improvements in deep reinforcement learning. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018.
  14. 14.Karimi, H., Nutini, J., and Schmidt, M. Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2016, Riva del Garda, Italy, September 19-23, 2016, Proceedings, Part I 16, pp. 795–811. Springer, 2016.
  15. 15.Kim, M., Lee, H., Yoo, K. M., Park, J., Lee, H., and Jung, K. Critic-guided decoding for controlled text generation. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Findings of the Association for Computational Linguistics: ACL 2023, pp. 4598–4612, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.281. URL https://aclanthology.org/2023.findings-acl.281.
  16. 16.Korbak, T., Perez, E., and Buckley, C. RL with KL penalties is better viewed as Bayesian inference. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 1083–1091, 2022.
  17. 17.Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. GeDi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 4929–4952, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.424. URL https://aclanthology.org/2021.findings-emnlp.424.
  18. 18.Leviathan, Y., Kalman, M., and Matias, Y. Fast inference from transformers via speculative decoding. International Conference on Machine Learning, 2023.
  19. 19.Li, J., Monroe, W., and Jurafsky, D. Learning to decode for future success. arXiv preprint arXiv:1701.06549, 2017.
  20. 20.Lu, X., Welleck, S., West, P., Jiang, L., Kasai, J., Khashabi, D., Le Bras, R., Qin, L., Yu, Y., Zellers, R., Smith, N. A., and Choi, Y. NeuroLogic a*esque decoding: Constrained text generation with lookahead heuristics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 780–799, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.57. URL https://aclanthology.org/2022.naacl-main.57.
  21. 21.Meng, T., Lu, S., Peng, N., and Chang, K.-W. Controllable text generation with neurally-decomposed oracle. Advances in Neural Information Processing Systems, 35: 28125–28139, 2022.
  22. 22.Microsoft. DSTC8 Reddit Corpus. https://github.com/microsoft/dstc8-reddit-corpus/, 2019. Accessed: 2023-09-30.
  23. 23.Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  24. 24.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  25. 25.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  26. 26.Peng, X. B., Kumar, A., Zhang, G., and Levine, S. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019.
  27. 27.Qin, L., Welleck, S., Khashabi, D., and Choi, Y. COLD decoding: Energy-based constrained text generation with langevin dynamics. Neural Information Processing Systems (NeurIPS), 2022. URL https://openreview.net/forum?id=TiZYrQ-mPup.
  28. 28.Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2023.
  29. 29.Rame, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J. WARM: On the benefits of weight averaged reward models. arXiv preprint arXiv:2401.12187, 2024.
  30. 30.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
  31. 31.Scialom, T., Dray, P.-A., Staiano, J., Lamprier, S., and Piwowarski, B. To beam or not to beam: That is a question of cooperation for language gans. Advances in neural information processing systems, 34:26585–26597, 2021.
  32. 32.Snell, C. V., Kostrikov, I., Su, Y., Yang, S., and Levine, S. Offline rl for natural language generation with implicit language q learning. In The Eleventh International Conference on Learning Representations, 2023.
  33. 33.Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33: 3008–3021, 2020.
  34. 34.Sun, Z., Suresh, A. T., Ro, J. H., Beirami, A., Jain, H., and Yu, F. SpecTr: Fast speculative decoding via optimal transport. In Neural Information Processing Systems, 2023.
  35. 35.Sutton, R. S. and Barto, A. G. Reinforcement learning: An introduction. MIT press, 2018.
  36. 36.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  37. 37.Wang, Z. T. and Ueda, M. Convergent and efficient deep Q network algorithm. 2022.
  38. 38.Welleck, S., Kulikov, I., Roller, S., Dinan, E., Cho, K., and Weston, J. Neural text generation with unlikelihood training. International Conference on Learning Representations, 2020.
  39. 39.Yang, J. Q., Salamatian, S., Sun, Z., Suresh, A. T., and Beirami, A. Asymptotics of language model alignment. In IEEE International Symposium on Information Theory (ISIT), 2024.
  40. 40.Yang, K. and Klein, D. FUDGE: Controlled text generation with future discriminators. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 3511–3535, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.276. URL https://aclanthology.org/2021.naacl-main.276.
  41. 41.Zhang, H. and Song, D. Discup: Discriminator cooperative unlikelihood prompt-tuning for controllable text generation. EMNLP, 2022.
  42. 42.Zhao, Y., Khalman, M., Joshi, R., Narayan, S., Saleh, M., and Liu, P. J. Calibrating sequence likelihood improves conditional language generation. In The Eleventh International Conference on Learning Representations, 2022.

Citation

MLA
Mudgal, S., et al. “Controlled Decoding from Language Models”. arXiv, 2023, http://arxiv.org/abs/2310.17022v3.
APA
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., Chen, J., Beutel, A., & Beirami, A. (2023). Controlled Decoding from Language Models. arXiv. http://arxiv.org/abs/2310.17022v3
Chicago
Mudgal, S., J. Lee, H. Ganapathy, et al. 2023. “Controlled Decoding from Language Models”. arXiv. http://arxiv.org/abs/2310.17022v3.
Harvard
Mudgal, S. et al. (2023) “Controlled Decoding from Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.17022v3.
Vancouver
1. Mudgal S, Lee J, Ganapathy H, et al (2023) Controlled Decoding from Language Models. arXiv

BibTeX

@article{mudgal2023controlled,
  title = {Controlled Decoding from Language Models},
  author = {Mudgal, Sidharth and Lee, Jong and Ganapathy, Harish and Li, YaGuang and Wang, Tao and Huang, Yanping and Chen, Zhifeng and Cheng, Heng-Tze and Collins, Michael and Strohman, Trevor and Chen, Jilin and Beutel, Alex and Beirami, Ahmad},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.17022v3},
  eprint = {2310.17022}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/