Data-Efficient Policy Evaluation Through Behavior Policy Search
Josiah P. HannaYash ChandakPhilip S. ThomasMartha WhitePeter StoneScott Niekum
Develops algorithms that search for an optimal data-collection behavior policy to evaluate reinforcement learning policies with lower variance and mean squared error than standard on-policy rollouts.
Accurately evaluating decision-making policies before live deployment is a vital challenge across high-stakes domains such as healthcare, automated marketing, and robotics. The standard industry practice, known as on-policy evaluation, directly executes the target policy in the environment to observe outcomes. However, this approach is often data-inefficient and suffers from high variance—meaning it requires large, costly sample sizes to produce stable performance estimates. While using an alternate data-collection strategy (off-policy evaluation) has historically been viewed as higher variance, the article demonstrates that intentionally choosing a different data-collection policy can significantly reduce evaluation error.
The article's main objective is to formalize the behavior policy search problem and demonstrate that actively searching for and deploying an optimized data-collection policy yields lower mean squared error than standard on-policy evaluation, without sacrificing statistical unbiasedness.
The authors evaluated their approach through formal mathematical proofs alongside empirical simulations spanning discrete environments and complex continuous control benchmarks, including Cart Pole, Acrobot, and Hopper. The analysis derived the theoretical form of a minimal-variance data-collection policy and introduced two practical optimization algorithms: Behavior Policy Gradient on the Variance (which directly minimizes estimator variance via gradient descent) and Behavior Policy Gradient on the Kullback-Leibler Divergence (which minimizes the divergence between the current data-collection policy and the theoretical optimum). The framework was evaluated across tabular and continuous domains, parameter sensitivity sweeps, and combinations with predictive model control variates.
The investigation produced four central findings. First, despite data points being collected sequentially across changing data-collection policies, the resulting estimates remain strictly unbiased, statistically consistent, and uncorrelated. Second, both proposed search algorithms consistently reduced estimation error compared to standard on-policy Monte Carlo roll-outs, lowering mean squared error by up to an order of magnitude across both discrete and continuous domains. Third, the potential for error reduction is highest when the target policy has high variance or when critical high-reward outcomes are rare under normal execution; actively sampling those rare, high-impact paths dramatically reduces total estimation variance. Finally, combining behavior policy search with predictive environment models (doubly robust estimation) achieved an additional 56.9% error reduction over standard model-assisted baselines.
These results demonstrate that organizations can achieve substantially more reliable policy evaluations using fewer physical or simulated trials, directly mitigating the costs and risks of deploying poorly evaluated strategies. This finding challenges the conventional belief that on-policy data collection is always superior when feasible. Instead, deliberately steering exploration toward rare, high-magnitude scenarios yields faster convergence and greater precision.
Decision-makers and engineering teams should adopt behavior policy search in environments where on-policy outcomes exhibit high variance or rare high-impact events, provided the operational platform permits minor computational overhead during data collection. When applying these methods, teams should utilize gradient baselines to stabilize learning rates and consider incorporating model-based control variates when environment models are available. Further engineering exploration should examine extensions to multi-policy evaluation and safe exploration constraints.
Confidence in these findings is supported by theoretical convergence proofs and consistent performance across diverse benchmark tasks. However, practitioners should note key limitations: the methods require careful step-size tuning to avoid unstable optimization steps, and behavior policy search yields minimal benefit if policy outcomes are near-deterministic or if variance originates entirely from uncontrollable environmental dynamics rather than the policy's actions.
- Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). The policy-gradient theorem provides the foundation for understanding how the source optimizes a behavior policy by following gradients of its evaluation variance.
- Paper: Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation, Olivier Jeunen et al. (2026). It continues the source’s effort to reduce off-policy evaluation error by deriving additive control-variate estimators that can improve on standard variance-reduction approaches.
