PILCO: A Model-Based and Data-Efficient Approach to Policy Search

Marc Peter DeisenrothCarl Edward Rasmussen

article2011ICML1,827 citations

Proposes a model-based reinforcement learning framework that leverages Gaussian process dynamics and analytic policy gradients to dramatically reduce model bias and achieve unprecedented data efficiency on continuous control tasks.

Listen

Autonomous control systems and robotics frequently suffer from severe data inefficiency when learning control policies through trial and error. Conventional reinforcement learning often requires hundreds or thousands of physical attempts to master basic tasks, making deployment on real hardware impractical due to mechanical wear and tear. While model-based methods attempt to solve this by learning environmental dynamics, they typically suffer from model bias—overconfidently making arbitrary predictions when extrapolating outside limited training data.

The article demonstrates and evaluates PILCO (Probabilistic Inference for Learning Control), a model-based framework designed to dramatically improve data efficiency when learning continuous control tasks from scratch without prior expert knowledge.

The authors implemented a probabilistic dynamics model using Gaussian processes to quantify uncertainty across state transitions. Rather than relying on computationally heavy sampling or discrete value functions, the method propagates state distributions across time using deterministic approximate inference (moment matching) and calculates policy gradients analytically to directly optimize controller parameters. The approach was tested on physical hardware and simulated dynamic benchmarks, specifically a physical cart-pole swing-up, a cart-double-pendulum swing-up, and a 12-dimensional robotic unicycle.

The evaluations yielded several notable findings. First, PILCO achieved unprecedented data efficiency, solving the real-world cart-pole swing-up and balance task from scratch in under 10 trials with only 17.5 seconds of total physical interaction. Second, this represents an improvement in learning speed of at least one full order of magnitude (10x or greater) compared to previous methods in the literature. Third, the framework successfully scaled to complex, multi-variable control domains, autonomously solving the cart-double-pendulum task in 60–90 seconds of experience and stabilizing the 12-dimensional unicycle within roughly 20–30 seconds of experience, achieving a 93% balancing success rate across 1,000 test simulations. Finally, comparative tests confirmed that accounting for model uncertainty is strictly necessary; replacing the probabilistic dynamics with a deterministic model led to total failure in learning from scratch due to overconfident extrapolation.

These findings indicate that probabilistic modeling fundamentally overcomes the core barrier of model bias in robotic reinforcement learning. By reducing physical trial times from hours or days to mere seconds or minutes, the methodology drastically mitigates physical wear, hardware damage risk, and operational costs. It enables autonomous systems to safely learn continuous manipulation and balance tasks from scratch without requiring expensive human demonstrations or manually engineered physics equations.

Engineering teams and decision-makers working on autonomous systems should adopt probabilistic dynamics models rather than deterministic approximations when planning under scarce data conditions. PILCO's analytic gradient approach is recommended over sampling-based policy updates to keep computational optimization efficient. However, stakeholders should recognize key operational limitations: the method is an indirect policy search that finds locally optimal rather than globally optimal controllers, and optimization may stall if target reward functions are too narrowly specified or if initial models lack exploration. Additionally, the approach treats model uncertainty as temporally uncorrelated noise, which could theoretically underestimate model uncertainty in certain systems. Nevertheless, because moment-matching approximations provide conservative predictions, the framework demonstrates high empirical reliability and robust real-world applicability.

Cover for PILCO: A Model-Based and Data-Efficient Approach to Policy Search

Abstract

In this paper, we introduce PILCO, a practical, data-efficient model-based policy search method. PILCO reduces model bias, one of the key problems of model-based reinforcement learning, in a principled way. By learning a probabilistic dynamics model and explicitly incorporating model uncertainty into long-term planning, PILCO can cope with very little data and facilitates learning from scratch in only a few trials. Policy evaluation is performed in closed form using state-of-the-art approximate inference. Furthermore, policy gradients are computed analytically for policy improvement. We report unprecedented learning efficiency on challenging and high-dimensional control tasks.

Table of Contents

  • 1. Introduction and Related Work
  • 2. Model-based Indirect Policy Search
  • 2.1. Dynamics Model Learning
  • 2.2. Policy Evaluation
  • 2.2.1. Mean Prediction
  • 2.2.2. Covariance Matrix of the Prediction
  • 2.3. Analytic Gradients for Policy Improvement
  • 3. Experimental Results
  • 3.1. Cart-Pole Swing-up
  • 3.2. Cart-Double-Pendulum Swing-up
  • 3.3. Unicycle Riding
  • 3.4. Data Efficiency
  • 4. Discussion and Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — PILCO Model-Based Policy Search Algorithm

    algorithm

    PILCO (Probabilistic Inference for Learning Control) is a model-based policy search framework designed for data-efficient reinforcement learning in continuous state and action spaces without prior expert knowledge. It alternates between learning a non-parametric probabilistic dynamics model, computing analytic policy gradients by propagating state distributions through the dynamics model using deterministic approximate inference (moment matching), optimizing policy parameters, and applying the updated policy to collect new system interaction data.

    Input: Initial state distribution N(μ0,Σ0)\mathcal{N}(\boldsymbol{\mu}_0, \boldsymbol{\Sigma}_0), horizon length TT, cost function c(x)c(\mathbf{x}), convergence threshold ϵ\epsilon
    Output: Optimized policy parameters θ∗\boldsymbol{\theta}^*, learned dynamics model
    1: Initialize policy parameters θ∼N(0,I)\boldsymbol{\theta} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})
    2: Apply random control signals to the physical system and record observed transition data
    3: repeat
    4: Fit or update probabilistic Gaussian Process (GP) dynamics model p(xt∣xt−1,ut−1)p(\mathbf{x}_t | \mathbf{x}_{t-1}, \mathbf{u}_{t-1}) using all recorded data
    5: repeat
    6: Evaluate expected return Jπ(θ)=∑t=0TExt[c(xt)]J^\pi(\boldsymbol{\theta}) = \sum_{t=0}^T \mathbb{E}_{\mathbf{x}_t}[c(\mathbf{x}_t)] using analytic moment matching over cascaded predictions p(x1),…,p(xT)p(\mathbf{x}_1), \dots, p(\mathbf{x}_T)
    7: Compute analytic gradient dJπ(θ)dθ\frac{\mathrm{d} J^\pi(\boldsymbol{\theta})}{\mathrm{d} \boldsymbol{\theta}} via the chain rule through predictive distributions
    8: Update policy parameters θ\boldsymbol{\theta} using a gradient-based non-convex optimizer (e.g., Conjugate Gradients or L-BFGS)
    9: until optimization converges; return θ∗\boldsymbol{\theta}^*
    10: Set policy π∗←π(θ∗)\pi^* \leftarrow \pi(\boldsymbol{\theta}^*)
    11: Apply π∗\pi^* to the system for a single episode of length TT and append the new trajectory data to the dataset
    12: until the task is learned

    By leveraging analytical approximate inference and analytical policy gradients, PILCO eliminates the need for Monte Carlo trajectory sampling during planning and achieves high data efficiency.

  2. Knowl 2 — Analytic Moment Matching for Gaussian State Propagation Through GP Dynamics

    model/method

    In PILCO, evaluating the trajectory distribution under Gaussian process (GP) transition dynamics requires mapping an uncertain state-action input x~t−1=[xt−1⊤,ut−1⊤]⊤∼N(μ~t−1,Σ~t−1)\tilde{\mathbf{x}}_{t-1} = [\mathbf{x}_{t-1}^\top, \mathbf{u}_{t-1}^\top]^\top \sim \mathcal{N}(\tilde{\boldsymbol{\mu}}_{t-1}, \tilde{\boldsymbol{\Sigma}}_{t-1}) through the non-linear GP model to predict the distribution of state differences Δt=xt−xt−1∈RD\mathbf{\Delta}_t = \mathbf{x}_t - \mathbf{x}_{t-1} \in \mathbb{R}^D. Since the exact predictive distribution p(Δt)=∫p(f(x~t−1)∣x~t−1)p(x~t−1)dx~t−1p(\mathbf{\Delta}_t) = \int p(f(\tilde{\mathbf{x}}_{t-1}) | \tilde{\mathbf{x}}_{t-1}) p(\tilde{\mathbf{x}}_{t-1}) \mathrm{d}\tilde{\mathbf{x}}_{t-1} is analytically intractable, PILCO approximates p(Δt)p(\mathbf{\Delta}_t) as a Gaussian N(μΔ,ΣΔ)\mathcal{N}(\boldsymbol{\mu}_\Delta, \boldsymbol{\Sigma}_\Delta) using exact moment matching.

    For target dimension a∈{1,…,D}a \in \{1, \dots, D\}, the mean prediction μΔa\mu_\Delta^a is:

    μΔa=βa⊤qa\mu_\Delta^a = \boldsymbol{\beta}_a^\top \mathbf{q}_a

    where βa=(Ka+σεa2I)−1ya∈Rn\boldsymbol{\beta}_a = (\mathbf{K}_a + \sigma_{\varepsilon_a}^2 \mathbf{I})^{-1} \mathbf{y}_a \in \mathbb{R}^n, with ya\mathbf{y}_a being the training targets for dimension aa, Ka\mathbf{K}_a being the n×nn \times n kernel Gram matrix, and the entries of qa∈Rn\mathbf{q}_a \in \mathbb{R}^n given by:

    qai=αa2∣Σ~t−1Λa−1+I∣exp⁡(−12νi⊤(Σ~t−1+Λa)−1νi)q_{ai} = \frac{\alpha_a^2}{\sqrt{|\tilde{\boldsymbol{\Sigma}}_{t-1} \mathbf{\Lambda}_a^{-1} + \mathbf{I}|}} \exp\left(-\frac{1}{2} \boldsymbol{\nu}_i^\top (\tilde{\boldsymbol{\Sigma}}_{t-1} + \mathbf{\Lambda}_a)^{-1} \boldsymbol{\nu}_i\right)

    where νi=x~i−μ~t−1\boldsymbol{\nu}_i = \tilde{\mathbf{x}}_i - \tilde{\boldsymbol{\mu}}_{t-1}, x~i\tilde{\mathbf{x}}_i is the ii-th training input, αa2\alpha_a^2 is the GP signal variance, and Λa=diag⁡([ℓa,12,…,ℓa,D+F2])\mathbf{\Lambda}_a = \operatorname{diag}([\ell_{a,1}^2, \dots, \ell_{a,D+F}^2]) contains the squared length-scales.

    The covariance matrix ΣΔ∈RD×D\boldsymbol{\Sigma}_\Delta \in \mathbb{R}^{D \times D} has off-diagonal entries (a≠ba \neq b):

    σab2=βa⊤Qβb−μΔaμΔb\sigma_{ab}^2 = \boldsymbol{\beta}_a^\top \mathbf{Q} \boldsymbol{\beta}_b - \mu_\Delta^a \mu_\Delta^b

    where Q∈Rn×n\mathbf{Q} \in \mathbb{R}^{n \times n} has elements:

    Qij=ka(x~i,μ~t−1)kb(x~j,μ~t−1)∣R∣exp⁡(12zij⊤R−1Σ~t−1zij)Q_{ij} = \frac{k_a(\tilde{\mathbf{x}}_i, \tilde{\boldsymbol{\mu}}_{t-1}) k_b(\tilde{\mathbf{x}}_j, \tilde{\boldsymbol{\mu}}_{t-1})}{\sqrt{|\mathbf{R}|}} \exp\left(\frac{1}{2} \mathbf{z}_{ij}^\top \mathbf{R}^{-1} \tilde{\boldsymbol{\Sigma}}_{t-1} \mathbf{z}_{ij}\right)

    with R=Σ~t−1(Λa−1+Λb−1)+I\mathbf{R} = \tilde{\boldsymbol{\Sigma}}_{t-1}(\mathbf{\Lambda}_a^{-1} + \mathbf{\Lambda}_b^{-1}) + \mathbf{I} and zij=Λa−1νi+Λb−1νj\mathbf{z}_{ij} = \mathbf{\Lambda}_a^{-1} \boldsymbol{\nu}_i + \mathbf{\Lambda}_b^{-1} \boldsymbol{\nu}_j. For diagonal entries (a=ba = b):

    σaa2=αa2−tr⁡((Ka+σεa2I)−1Q)+βa⊤Qβa−(μΔa)2\sigma_{aa}^2 = \alpha_a^2 - \operatorname{tr}\left((\mathbf{K}_a + \sigma_{\varepsilon_a}^2 \mathbf{I})^{-1} \mathbf{Q}\right) + \boldsymbol{\beta}_a^\top \mathbf{Q} \boldsymbol{\beta}_a - (\mu_\Delta^a)^2

    The next state distribution p(xt)≈N(μt,Σt)p(\mathbf{x}_t) \approx \mathcal{N}(\boldsymbol{\mu}_t, \boldsymbol{\Sigma}_t) is obtained by:

    μt=μt−1+μΔ\boldsymbol{\mu}_t = \boldsymbol{\mu}_{t-1} + \boldsymbol{\mu}_\Delta

    Σt=Σt−1+ΣΔ+cov⁡[xt−1,Δt]+cov⁡[Δt,xt−1]\boldsymbol{\Sigma}_t = \boldsymbol{\Sigma}_{t-1} + \boldsymbol{\Sigma}_\Delta + \operatorname{cov}[\mathbf{x}_{t-1}, \mathbf{\Delta}_t] + \operatorname{cov}[\mathbf{\Delta}_t, \mathbf{x}_{t-1}]

  3. Knowl 3 — Probabilistic Dynamics Modeling with Gaussian Processes

    model/method

    PILCO models the unknown transition dynamics of a continuous system xt=f(xt−1,ut−1)\mathbf{x}_t = f(\mathbf{x}_{t-1}, \mathbf{u}_{t-1}) using non-parametric Gaussian Process (GP) regression. Instead of predicting the absolute state directly, the GP model predicts the state change:

    Δt=xt−xt−1+ε\mathbf{\Delta}_t = \mathbf{x}_t - \mathbf{x}_{t-1} + \boldsymbol{\varepsilon}

    where ε∼N(0,Σε)\boldsymbol{\varepsilon} \sim \mathcal{N}(\mathbf{0}, \boldsymbol{\Sigma}_\varepsilon) represents i.i.d. Gaussian measurement noise with diagonal covariance Σε=diag⁡([σε,12,…,σε,D2])\boldsymbol{\Sigma}_\varepsilon = \operatorname{diag}([\sigma_{\varepsilon,1}^2, \dots, \sigma_{\varepsilon,D}^2]).

    The training inputs are state-action pairs x~=[x⊤,u⊤]⊤∈RD+F\tilde{\mathbf{x}} = [\mathbf{x}^\top, \mathbf{u}^\top]^\top \in \mathbb{R}^{D+F}, where DD is the state dimensionality and FF is the control dimensionality. For each target dimension a∈{1,…,D}a \in \{1, \dots, D\}, a separate GP is trained assuming a prior mean function m≡0m \equiv 0 and a Squared Exponential (SE) covariance function with Automatic Relevance Determination (ARD):

    ka(x~,x~′)=αa2exp⁡(−12(x~−x~′)⊤Λa−1(x~−x~′))k_a(\tilde{\mathbf{x}}, \tilde{\mathbf{x}}') = \alpha_a^2 \exp\left(-\frac{1}{2} (\tilde{\mathbf{x}} - \tilde{\mathbf{x}}')^\top \mathbf{\Lambda}_a^{-1} (\tilde{\mathbf{x}} - \tilde{\mathbf{x}}')\right)

    where αa2\alpha_a^2 is the latent signal variance and Λa=diag⁡([ℓa,12,…,ℓa,D+F2])\mathbf{\Lambda}_a = \operatorname{diag}([\ell_{a,1}^2, \dots, \ell_{a,D+F}^2]) is the matrix of characteristic length-scales. All GP hyperparameters θGP={αa,Λa,σεa}a=1D\boldsymbol{\theta}_{GP} = \{\alpha_a, \mathbf{\Lambda}_a, \sigma_{\varepsilon_a}\}_{a=1}^D are learned from training data by maximizing the marginal likelihood (evidence).

  4. Knowl 4 — Analytic Policy Gradient Computation via Exact Derivative Propagation

    model/method

    PILCO computes the exact gradient of the expected return Jπ(θ)=∑t=0TExt[c(xt)]J^\pi(\boldsymbol{\theta}) = \sum_{t=0}^T \mathbb{E}_{\mathbf{x}_t}[c(\mathbf{x}_t)] with respect to policy parameters θ\boldsymbol{\theta} analytically by repeated application of the chain rule through the predictive state distributions p(xt)=N(μt,Σt)p(\mathbf{x}_t) = \mathcal{N}(\boldsymbol{\mu}_t, \boldsymbol{\Sigma}_t).

    Letting Et:=Ext[c(xt)]\mathcal{E}_t := \mathbb{E}_{\mathbf{x}_t}[c(\mathbf{x}_t)], the gradient contribution at time step tt is decomposed as:

    dEtdθ=∂Et∂μtdμtdθ+∂Et∂ΣtdΣtdθ\frac{\mathrm{d}\mathcal{E}_t}{\mathrm{d}\boldsymbol{\theta}} = \frac{\partial \mathcal{E}_t}{\partial \boldsymbol{\mu}_t} \frac{\mathrm{d}\boldsymbol{\mu}_t}{\mathrm{d}\boldsymbol{\theta}} + \frac{\partial \mathcal{E}_t}{\partial \boldsymbol{\Sigma}_t} \frac{\mathrm{d}\boldsymbol{\Sigma}_t}{\mathrm{d}\boldsymbol{\theta}}

    The total derivatives of the predictive moments with respect to θ\boldsymbol{\theta} obey temporal recurrence equations:

    dp(xt)dθ=∂p(xt)∂p(xt−1)dp(xt−1)dθ+∂p(xt)∂θ\frac{\mathrm{d} p(\mathbf{x}_t)}{\mathrm{d}\boldsymbol{\theta}} = \frac{\partial p(\mathbf{x}_t)}{\partial p(\mathbf{x}_{t-1})} \frac{\mathrm{d} p(\mathbf{x}_{t-1})}{\mathrm{d}\boldsymbol{\theta}} + \frac{\partial p(\mathbf{x}_t)}{\partial \boldsymbol{\theta}}

    For the predictive mean vector μt\boldsymbol{\mu}_t, this recurrence evaluates to:

    dμtdθ=∂μt∂μt−1dμt−1dθ+∂μt∂Σt−1dΣt−1dθ+∂μt∂θ\frac{\mathrm{d}\boldsymbol{\mu}_t}{\mathrm{d}\boldsymbol{\theta}} = \frac{\partial \boldsymbol{\mu}_t}{\partial \boldsymbol{\mu}_{t-1}} \frac{\mathrm{d}\boldsymbol{\mu}_{t-1}}{\mathrm{d}\boldsymbol{\theta}} + \frac{\partial \boldsymbol{\mu}_t}{\partial \boldsymbol{\Sigma}_{t-1}} \frac{\mathrm{d}\boldsymbol{\Sigma}_{t-1}}{\mathrm{d}\boldsymbol{\theta}} + \frac{\partial \boldsymbol{\mu}_t}{\partial \boldsymbol{\theta}}

    where the partial derivative with respect to θ\boldsymbol{\theta} is mapped through the control distribution moments μu\boldsymbol{\mu}_u and Σu\boldsymbol{\Sigma}_u:

    ∂μt∂θ=∂μΔ∂μu∂μu∂θ+∂μΔ∂Σu∂Σu∂θ\frac{\partial \boldsymbol{\mu}_t}{\partial \boldsymbol{\theta}} = \frac{\partial \boldsymbol{\mu}_\Delta}{\partial \boldsymbol{\mu}_u} \frac{\partial \boldsymbol{\mu}_u}{\partial \boldsymbol{\theta}} + \frac{\partial \boldsymbol{\mu}_\Delta}{\partial \boldsymbol{\Sigma}_u} \frac{\partial \boldsymbol{\Sigma}_u}{\partial \boldsymbol{\theta}}

    Because every term in these derivatives is available in closed form via the moment-matching expressions, policy gradients are computed analytically without sample-based gradient estimates, drastically reducing gradient variance.

  5. Knowl 5 — Saturating Exponential Cost Function for Target Stabilization

    equation

    PILCO formulates the immediate cost function c(x)∈[0,1]c(\mathbf{x}) \in [0, 1] as a continuous inverted Gaussian saturating penalty centered on a desired target state xtarget\mathbf{x}_{\text{target}}:

    c(x)=1−exp⁡(−∥x−xtarget∥2σc2)c(\mathbf{x}) = 1 - \exp\left(-\frac{\|\mathbf{x} - \mathbf{x}_{\text{target}}\|^2}{\sigma_c^2}\right)

    where x∈RD\mathbf{x} \in \mathbb{R}^D is the state vector, xtarget∈RD\mathbf{x}_{\text{target}} \in \mathbb{R}^D is the target state, and σc>0\sigma_c > 0 is a user-defined hyperparameter controlling the width of the target region.

    This cost function acts as a smooth, continuous approximation of a binary 0-1 cost (where cost is 0 within the target area and 1 elsewhere). Because c(x)c(\mathbf{x}) is expressed as a squared exponential, the expected immediate cost Ext[c(xt)]\mathbb{E}_{\mathbf{x}_t}[c(\mathbf{x}_t)] under a Gaussian state distribution p(xt)=N(μt,Σt)p(\mathbf{x}_t) = \mathcal{N}(\boldsymbol{\mu}_t, \boldsymbol{\Sigma}_t) can be evaluated analytically in closed form.

  6. Knowl 6 — Data Efficiency of PILCO Across Continuous Control Benchmarks

    data/table

    PILCO achieves learning from scratch on continuous state-action physical and simulated control tasks using minimal interaction time and very few trials, scaling across different state dimensionalities and parameter spaces.

    Task Cart-Pole Cart-Double-Pole Unicycle
    State space R4\mathbb{R}^4 R6\mathbb{R}^6 R12\mathbb{R}^{12}
    Number of trials ≤10\le 10 20–3020\text{--}30 ≈20\approx 20
    Total experience ≈20 s\approx 20\text{ s} ≈60 s–90 s\approx 60\text{ s}\text{--}90\text{ s} ≈20 s–30 s\approx 20\text{ s}\text{--}30\text{ s}
    Parameter space R305\mathbb{R}^{305} R1816\mathbb{R}^{1816} R28\mathbb{R}^{28}

    On the cart-pole task, PILCO learned a successful swing-up and balancing policy on a physical robot within 17.5 seconds of total system interaction (under 10 trials). When compared to other reinforcement learning algorithms from the literature learning cart-pole without prior expert knowledge—which require between 10210^2 and 10510^5 seconds of interaction time—PILCO improves data efficiency by at least one order of magnitude.

  7. Knowl 7 — Autonomous Balancing of a 5-DoF Robotic Unicycle

    empirical result

    PILCO was evaluated in simulation on a 5-degree-of-freedom robotic unicycle model governed by 12 coupled first-order ordinary differential equations (D=12D = 12) with a 2-dimensional control action space (F=2F = 2). The control inputs consisted of wheel torque ∣uw∣≤10 Nm|u_w| \le 10\text{ Nm} (longitudinal acceleration) and turntable flywheel torque ∣ut∣≤50 Nm|u_t| \le 50\text{ Nm} (lateral stability).

    A linear state-feedback controller π(x)=Ax+b\boldsymbol{\pi}(\mathbf{x}) = \mathbf{A}\mathbf{x} + \mathbf{b} with parameters θ={A,b}∈R28\boldsymbol{\theta} = \{\mathbf{A}, \mathbf{b}\} \in \mathbb{R}^{28} was trained to stabilize the unicycle upright from an initial state covariance Σ0=0.252I\boldsymbol{\Sigma}_0 = 0.25^2 \mathbf{I} (allowing angular offsets of roughly 30∘30^\circ). By jointly modeling the cross-correlations between all state and action dimensions, PILCO learned a balancing policy within approximately 20 trials (corresponding to 20–30 s20\text{--}30\text{ s} of total interaction time), achieving an empirical success rate of approximately 93% over 1,000 test trials.

  8. Knowl 8 — Nonlinear RBF Policy for Physical Cart-Pole Swing-Up and Balancing

    empirical result

    PILCO was applied to a real physical cart-pole setup consisting of a 0.7 kg0.7\text{ kg} cart on a track with an attached freely swinging 0.325 kg0.325\text{ kg} pendulum, with a continuous state space of dimension 4 (cart position, cart velocity, pendulum angle, and angular velocity) and horizontal force control inputs u∈[−10,10] Nu \in [-10, 10]\text{ N}.

    The policy was parameterized as a non-linear Radial Basis Function (RBF) network:

    π(x,θ)=∑i=1nwiexp⁡(−12(x−μi)⊤Λ−1(x−μi))\pi(\mathbf{x}, \boldsymbol{\theta}) = \sum_{i=1}^n w_i \exp\left(-\frac{1}{2}(\mathbf{x} - \boldsymbol{\mu}_i)^\top \mathbf{\Lambda}^{-1}(\mathbf{x} - \boldsymbol{\mu}_i)\right)

    with n=50n = 50 squared exponential basis functions centered at μi\boldsymbol{\mu}_i, yielding a parameter space θ={wi,Λ,μi}∈R305\boldsymbol{\theta} = \{w_i, \mathbf{\Lambda}, \boldsymbol{\mu}_i\} \in \mathbb{R}^{305}. PILCO learned both the dynamics model and the control policy completely from scratch to perform the swing-up and balance the inverted pendulum in the center of the track within 17.5 seconds of interaction time across under 10 trials.

  9. Knowl 9 — Single Nonlinear Controller for Cart-Double-Pendulum Swing-Up and Balancing

    empirical result

    PILCO was applied to a cart-double-pendulum system in simulation consisting of a cart (0.5 kg0.5\text{ kg}) and a two-link pendulum (each link 0.5 kg0.5\text{ kg}), characterized by a 6-dimensional continuous state space (x1,x˙1,θ2,θ˙2,θ3,θ˙3x_1, \dot{x}_1, \theta_2, \dot{\theta}_2, \theta_3, \dot{\theta}_3) and a 1-dimensional force action ∣u∣≤20 N|u| \le 20\text{ N}.

    Rather than using two separate hand-engineered controllers for swing-up and balance, PILCO optimized a single unified nonlinear RBF controller (n=200n = 200 basis functions, θ∈R1816\boldsymbol{\theta} \in \mathbb{R}^{1816}) that simultaneously solved the swing-up from hanging down and the inverted balancing at the target start position. The system learned the complete task fully autonomously from scratch within 20 to 30 trials, corresponding to 60–90 s60\text{--}90\text{ s} of total interaction time.

  10. Knowl 10 — Failure of Deterministic Dynamics Models Under Small-Sample Regimes

    empirical result

    To evaluate the necessity of modeling dynamics uncertainty, the PILCO policy search framework was tested on the simulated cart-pole swing-up task using a deterministic dynamics model consisting solely of the GP posterior mean function (discarding the GP predictive variance var⁡f[Δ]\operatorname{var}_f[\mathbf{\Delta}]).

    Policy search with this deterministic dynamics model failed to learn the task from scratch. Because initial training sets did not contain states near the inverted target region, deterministic long-term trajectory rollouts extrapolated into unexplored state space with zero variance and complete confidence. The predictive model eventually fell back to the uninformative prior mean function (m≡0m \equiv 0), generating useless predictions and misleading policy gradients. In contrast, the full probabilistic model correctly reflected high uncertainty in unvisited regions, penalizing unsupported state trajectories during policy optimization.

  11. Knowl 11 — Local Optimality, Flat Gradient Regions, and Temporal Independence Assumptions in PILCO

    limitation

    PILCO possesses three primary theoretical and practical limitations:

    1. Local Optimality: Because the expected return objective Jπ(θ)J^\pi(\boldsymbol{\theta}) is non-convex with respect to policy parameters θ\boldsymbol{\theta}, gradient-based optimization guarantees only locally optimal solutions, and the learned policy remains conditioned on the trajectories observed during exploration.
    2. Flat Gradient Regions with Peaked Costs: If the immediate cost function width σc\sigma_c is extremely small (approaching a sharp indicator function) and the initial random dynamics model is inaccurate, predicted state distributions p(x1),…,p(xT)p(\mathbf{x}_1), \dots, p(\mathbf{x}_T) may have zero overlap with the target region, causing analytic policy gradients to vanish (∇θJπ≈0\nabla_\theta J^\pi \approx \mathbf{0}) and trapping optimization in local minima.
    3. Uncorrelated Model Uncertainty: PILCO assumes model uncertainty is temporally uncorrelated across cascading time steps, treating epistemic uncertainty similarly to independent noise at each step. While moment matching tends to be conservative and partially offsets this issue, treating model errors as temporally independent can theoretically lead to an underestimation of model uncertainty over long planning horizons.

Coverage note — None was omitted; all key theoretical formulations, algorithmic steps, benchmark setups, comparative data, and limitations from the paper are fully covered.

References

  1. 1.Abbeel, P., Quigley, M., and Ng, A. Y. Using Inaccurate Models in Reinforcement Learning. In Proceedings of the ICML, pp. 1–8, 2006.
  2. 2.Atkeson, C. G. and Santamaría, J. C. A Comparison of Direct and Model-Based Reinforcement Learning. In Proceedings of the ICRA, 1997.
  3. 3.Bagnell, J. A. and Schneider, J. G. Autonomous Helicopter Control using Reinforcement Learning Policy Search Methods. In Proceedings of the ICRA, pp. 1615–1620, 2001.
  4. 4.Coulom, R. Reinforcement Learning Using Neural Networks, with Applications to Motor Control. PhD thesis, Institut National Polytechnique de Grenoble, 2002.
  5. 5.Deisenroth, M. P. Efficient Reinforcement Learning using Gaussian Processes. KIT Scientific Publishing, 2010. ISBN 978-3-86644-569-7.
  6. 6.Deisenroth, M. P., Rasmussen, C. E., and Peters, J. Gaussian Process Dynamic Programming. Neurocomputing, 72(7–9):1508–1524, 2009.
  7. 7.Deisenroth, M. P., Rasmussen, C. E., and Fox, D. Learning to Control a Low-Cost Manipulator using Data-Efficient Reinforcement Learning. In Proceedings of R:SS, 2011.
  8. 8.Doya, K. Reinforcement Learning in Continuous Time and Space. Neural Computation, 12(1):219–245, 2000. ISSN 0899-7667.
  9. 9.Engel, Y., Mannor, S., and Meir, R. Bayes Meets Bellman: The Gaussian Process Approach to Temporal Difference Learning. In Proceedings of the ICML, pp. 154–161, 2003.
  10. 10.Fabri, S. and Kadirkamanathan, V. Dual Adaptive Control of Nonlinear Stochastic Systems using Neural Networks. Automatica, 34(2):245–253, 1998.
  11. 11.Forster, D. Robotic Unicycle. Report, Department of Engineering, University of Cambridge, UK, 2009.
  12. 12.Kimura, H. and Kobayashi, S. Efficient Non-Linear Control by Combining Q-learning with Local Linear Controllers. In Proceedings of the ICML, pp. 210–219, 1999.
  13. 13.Ko, J., Klein, D. J., Fox, D., and Haehnel, D. Gaussian Processes and Reinforcement Learning for Identification and Control of an Autonomous Blimp. In ICRA, pp. 742–747, 2007.
  14. 14.Naveh, Y., Bar-Yoseph, P. Z., and Halevi, Y. Nonlinear Modeling and Control of a Unicycle. Journal of Dynamics and Control, 9(4):279–296, 1999.
  15. 15.Peters, J. and Schaal, S. Policy Gradient Methods for Robotics. In Proceedings of the IROS, pp. 2219–2225, 2006.
  16. 16.Quiñonero-Candela, J., Girard, A., Larsen, J., and Rasmussen, C. E. Propagation of Uncertainty in Bayesian Kernel Models—Application to Multiple-Step Ahead Forecasting. In Proceedings of the ICASSP, pp. 701–704, 2003.
  17. 17.Raiko, T. and Tornio, M. Variational Bayesian Learning of Nonlinear Hidden State-Space Models for Model Predictive Control. Neurocomputing, 72(16–18):3702–3712, 2009.
  18. 18.Rasmussen, C. E. and Kuss, M. Gaussian Processes in Reinforcement Learning. In NIPS, pp. 751–759. 2004.
  19. 19.Rasmussen, C. E. and Williams, C. K. I. Gaussian Processes for Machine Learning. The MIT Press, 2006.
  20. 20.Riedmiller, M. Neural Fitted Q Iteration—First Experiences with a Data Efficient Neural Reinforcement Learning Method. In Proceedings of the ECML, 2005.
  21. 21.Schaal, S. Learning From Demonstration. In NIPS, pp. 1040–1046. 1997.
  22. 22.Schneider, J. G. Exploiting Model Uncertainty Estimates for Safe Dynamic Control Learning. In NIPS. 1997.
  23. 23.van Hasselt, H. Insights in Reinforcement Learning. Wöhrmann Print Service, 2010. ISBN 978-90-39354964.
  24. 24.Wawrzynski, P. and Pacut, A. Model-free off-policy Reinforcement Learning in Continuous Environment. In Proceedings of the IJCNN, pp. 1091–1096, 2004.
  25. 25.Wilson, A., Fern, A., and Tadepalli, P. Incorporating Domain Models into Bayesian Optimization for RL. In ECML-PKDD, pp. 467–482, 2010.
  26. 26.Zhong, W. and Röck, H. Energy and Passivity Based Control of the Double Inverted Pendulum on a Cart. In Proceedings of the CCA, pp. 896–901, 2001.

Citation

MLA
Deisenroth, M. P., and C. E. Rasmussen. “PILCO: A Model-Based and Data-Efficient Approach to Policy Search”. Cambridge University Engineering Department Publications Database, 2011, pp. 465–72, http://publications.eng.cam.ac.uk/323736/.
APA
Deisenroth, M. P., & Rasmussen, C. E. (2011). PILCO: A Model-Based and Data-Efficient Approach to Policy Search. Cambridge University Engineering Department Publications Database, 465–472. http://publications.eng.cam.ac.uk/323736/
Chicago
Deisenroth, M. P., and C. E. Rasmussen. 2011. “PILCO: A Model-Based and Data-Efficient Approach to Policy Search”. Cambridge University Engineering Department Publications Database, 465–72. http://publications.eng.cam.ac.uk/323736/.
Harvard
Deisenroth, M.P. and Rasmussen, C.E. (2011) “PILCO: A Model-Based and Data-Efficient Approach to Policy Search”, Cambridge University Engineering Department Publications Database, pp. 465–472. Available at: http://publications.eng.cam.ac.uk/323736/.
Vancouver
1. Deisenroth MP, Rasmussen CE (2011) PILCO: A Model-Based and Data-Efficient Approach to Policy Search. Cambridge University Engineering Department Publications Database 465–472

BibTeX

@article{deisenroth2011pilco,
  title = {PILCO: A Model-Based and Data-Efficient Approach to Policy Search},
  author = {Deisenroth, Marc Peter and Rasmussen, Carl Edward},
  year = {2011},
  journal = {Cambridge University Engineering Department Publications Database},
  pages = {465-472},
  url = {http://publications.eng.cam.ac.uk/323736/}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors