Theoretical guarantees on the best-of-n alignment policy
Ahmad BeiramiAlekh AgarwalJonathan BerantAlexander Nicholas D'AmourJacob EisensteinChirag NagpalAnanda Theertha Suresh
Establishes rigorous theoretical guarantees for best-of-n alignment sampling by disproving a widely used closed-form KL divergence formula, bounding policy drift and win rates, and introducing a tighter KL estimator to guide test-time compute scaling.
Deploying generative artificial intelligence models safely requires alignment techniques that improve output quality and adhere to safety rules without degrading baseline capabilities. A widely adopted method is best-of-n sampling, where a system generates multiple candidate responses, scores them using a reward function, and selects the highest-scoring candidate. To monitor that the model does not drift too far from its original behavior, practitioners track statistical distance using Kullback-Leibler (KL) divergence, which measures distribution drift. Historically, the field relied on an analytical formula to estimate this drift as sample size grows. The article evaluates the theoretical foundations of best-of-n sampling to determine the mathematical accuracy of this formula and establish rigorous guarantees on distribution drift and win rates against the base model.
To conduct this evaluation, the article develops a closed-form probability mass function for the best-of-n policy under standard assumptions of finite outcomes and unique rewards. The authors derive theoretical upper and lower bounds for distribution drift and win rates, introducing a practical estimator for drift. They also compare best-of-n sampling against alternative rejection sampling mechanisms, such as rewind-and-repeat thresholding, and evaluate their findings through numerical simulations and empirical tests on language tasks using an instruction-tuned model.
The article demonstrates that the standard analytical formula used across the literature is not an exact equality but an upper bound on true distribution drift. When candidate outputs have low probabilities and sample sizes are small, the formula is reasonably close to reality; however, when the sample size is large or specific high-reward outputs have substantial probability mass, the formula drastically overestimates drift by an unbounded margin. Additionally, the article proves that the win rate of best-of-n sampling over the base model is strictly upper-bounded by the ratio of the sample size to the sample size plus one. The proposed drift estimator closely tracks actual behavior across tested regimes, and the analysis confirms that best-of-n sampling achieves near-optimal win rate versus drift tradeoffs at practical sample sizes below 1,000.
These findings have direct operational and governance implications for deploying generative models. Because existing literature relied on an overly conservative upper bound, best-of-n sampling actually preserves base model capabilities and safety guardrails significantly better than previously reported. Decision-makers can achieve competitive performance without the costly retraining required by complex reinforcement learning pipelines. However, this high alignment efficiency also implies that malicious actors can effectively repurpose best-of-n sampling to bypass safety guardrails if reward signals are unconstrained.
Organizations should adopt the article's proposed estimator to track distribution drift more accurately in deployment pipelines and leverage blockwise or standard best-of-n sampling within sample sizes below 1,000 for efficient inference-time alignment. For future research and development, teams should design hybrid approaches that balance compute cost, target reward, and capability preservation, while simultaneously engineering safeguards against adversarial test-time jailbreaking.
The findings carry high confidence due to formal mathematical proofs validated by extensive empirical simulations. Users should note the minor limitation that exact drift estimates can exhibit sample variance across individual draws, which requires averaging across small batches of prompts for reliable point estimates in production monitoring.
- Paper: Scaling Laws for Reward Model Overoptimization, Leo Gao et al. (2023). This foundational study derives empirical scaling laws for KL divergence drift and overoptimization in best-of-n sampling, directly providing the analytical baseline formulas and empirical behaviors that the source paper mathematically bounds and refines.
- Paper: Direct Preference Optimization: Your Language Model is Secretly a Reward Model, Rafael Rafailov et al. (2023). It introduces standard preference optimization frameworks and analyzes the fundamental trade-offs along reward-KL divergence frontiers that motivate test-time rejection sampling methods.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). It establishes the standard language model alignment paradigm using reward models and KL divergence tracking to constrain drift from the reference policy.
- Paper: Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs, Pranjal Aggarwal et al. (2023). It explores inference-time multi-candidate sampling and stopping criteria, offering practical context for evaluating compute-efficiency trade-offs in response selection.
- Paper: DeAL: Decoding-time Alignment for Large Language Models, James Y. Huang et al. (2025). It extends inference-time alignment beyond full-sequence best-of-n candidate selection by developing heuristic-guided search and lookahead alignment dynamically during decoding.
- Paper: A General Framework for Inference-time Scaling and Steering of Diffusion Models, Raghav Singhal et al. (2025). It generalizes inference-time reward steering beyond standard sequence-level best-of-n sampling by iteratively filtering and resampling intermediate generation trajectories.
- Paper: Reasoning with Sampling: Your Base Model is Smarter Than You Think, Aayush Karan et al. (2026). It builds on test-time sampling dynamics by demonstrating how sampling from base model probability distributions can unlock strong reasoning performance without full reinforcement learning pipelines.
- Paper: From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models, Tarun Raheja et al. (2026). It synthesizes preference learning techniques under a unified theoretical framework, examining policy drift regularization and distribution coverage issues analyzed in the source.
