Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms

Yanwei JiaXun Yu Zhou

article2022JMLR180 citations

Establishes a theoretical foundation for continuous-time and continuous-space policy gradient methods by reformulating the gradient estimation as a policy evaluation problem and deriving both offline and online actor-critic algorithms via martingale conditions.

Listen

Real-world decision-making systems—such as high-frequency trading platforms, autonomous vehicles, and industrial robotics—operate continuously in time and space rather than in artificial discrete intervals. Standard reinforcement learning techniques typically approximate reality by discretizing time upfront, which often produces algorithms that are unstable and highly sensitive to the chosen time step. The article establishes a rigorous mathematical foundation for continuous-time and continuous-space reinforcement learning, demonstrating model-free algorithms that learn optimal policies directly from observed sample data without relying on prior knowledge of environmental dynamics.

The research adopts an exploratory stochastic control framework that uses entropy regularization to balance exploration with exploitation. By developing the theoretical policy gradient in continuous time, the authors show that finding the policy gradient mathematically reduces to an auxiliary policy evaluation problem. Using stochastic calculus and martingale techniques—mathematical methods that characterize fair-game drift conditions—the article removes the need to know the underlying equations of the environment. The resulting framework enables two complementary actor-critic learning approaches: an offline method that updates decision rules across complete simulation episodes, and an online method that updates parameters in real time using only past observations.

The findings demonstrate the effectiveness and flexibility of these algorithms across both episodic and long-term average tasks. First, the derived continuous-time policy gradient naturally incorporates advantage-style baseline updates, improving learning stability without requiring manual baseline tuning. Second, in simulated financial portfolio selection over a 20-year horizon, the proposed offline actor-critic algorithm achieved significantly higher risk-adjusted returns (Sharpe ratios) than existing benchmark methods across various market conditions, while reliably hitting targeted return levels. Third, in continuous linear-quadratic control simulations, online real-time learning converged directly toward theoretical performance limits, with the remaining gap strictly accounting for the cost of active exploration.

These results provide a validated pathway for deploying model-free reinforcement learning in mission-critical, high-frequency environments where upfront time discretization is impractical or risky. Organizations facing stationary environments with historical trajectory data can use offline actor-critic algorithms to maximize sample efficiency and return stability. Conversely, for large-scale or non-stationary operations, online incremental updating reduces computational storage and continuously adapts to incoming data streams.

While the theoretical convergence and simulation evidence are strong, practitioners should exercise caution regarding specific operational constraints. Offline algorithms outperform online counterparts in finite datasets because they reuse sampled paths, whereas online methods require longer learning horizons to overcome early sub-optimal trials. Performance remains sensitive to the choice of exploration temperature and policy parameterizations. Future operational implementations should evaluate offline versus online trade-offs via domain-specific pilot testing and explore replay techniques to enhance online sample efficiency.

Jia et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms

Abstract

We study policy gradient (PG) for reinforcement learning in continuous time and space under the regularized exploratory formulation developed by Wang et al. (2020). We represent the gradient of the value function with respect to a given parameterized stochastic policy as the expected integration of an auxiliary running reward function that can be evaluated using samples and the current value function. This representation effectively turns PG into a policy evaluation (PE) problem, enabling us to apply the martingale approach recently developed by Jia and Zhou (2022a) for PE to solve our PG problem. Based on this analysis, we propose two types of actor-critic algorithms for RL, where we learn and update value functions and policies simultaneously and alternatingly. The first type is based directly on the aforementioned representation, which involves future trajectories and is offline. The second type, designed for online learning, employs the first-order condition of the policy gradient and turns it into martingale orthogonality conditions. These conditions are then incorporated using stochastic approximation when updating policies. Finally, we demonstrate the algorithms by simulations in two concrete examples.

Table of Contents

  • 1. Introduction
  • 2. Problem Formulation and Preliminaries
  • 3. Theoretical Foundation of Actor-Critic Algorithms
  • 3.1 Policy Evaluation
  • 3.2 Policy Gradient
  • 3.3 Actor-Critic Algorithms
  • 4. Extension to Ergodic Tasks
  • 5. Applications
  • 5.1 Mean-Variance Portfolio Selection
  • 5.2 Ergodic Linear-Quadratic Control
  • 6. Conclusion
  • Acknowledgments
  • Appendix A. Connections with Policy Gradient in Discrete Time
  • Appendix B. Theoretical Results Employed in Simulation Experiments
  • Appendix B1. Mean-Variance Portfolio Selection
  • Appendix B2. Ergodic Linear-Quadratic Control
  • Appendix C. Proofs of Statements
  • Proof of Lemma 2
  • Proof of Lemma 3
  • Proof of Theorem 4
  • Proof of Theorem 5
  • Proof of Theorem 6
  • Proof of Lemma 7
  • References

Knowls

  1. Knowl 1 — Sample-based representation of the continuous-time policy gradient

    theoretical result

    Let XπϕX^{\pi^\phi} be a state trajectory generated by a diffusion control system when actions asπϕa_s^{\pi^\phi} are sampled from a differentiable stochastic policy πϕ(⋅∣s,Xsπϕ)\pi^\phi(\cdot\mid s,X_s^{\pi^\phi}). Let J(t,x;πϕ)J(t,x;\pi^\phi) be the policy’s discounted value from state xx at time tt, with discount rate β≥0\beta\ge 0, running reward rr, regularizer pp, and regularization weight γ≥0\gamma\ge 0. Define q(t,x,a,ϕ)=∂ϕp(t,x,a,πϕ(⋅∣t,x))q(t,x,a,\phi)=\partial_\phi p(t,x,a,\pi^\phi(\cdot\mid t,x)). Under the paper’s regularity assumptions, the gradient g(t,x;ϕ)=∂ϕJ(t,x;πϕ)g(t,x;\phi)=\partial_\phi J(t,x;\pi^\phi) is

    g(t,x;ϕ)=E[∫tTe−β(s−t){∂ϕlog⁡πϕ(asπϕ∣s,Xsπϕ)[dJ(s,Xsπϕ;πϕ)+(r(s,Xsπϕ,asπϕ)+γp(s,Xsπϕ,asπϕ,πϕ)−βJ(s,Xsπϕ;πϕ))ds]+γq(s,Xsπϕ,asπϕ,ϕ)ds}∣Xtπϕ=x].g(t,x;\phi)=\mathbb E\left[\left.\int_t^T e^{-\beta(s-t)}\left\{\partial_\phi\log\pi^\phi(a_s^{\pi^\phi}\mid s,X_s^{\pi^\phi})\left[dJ(s,X_s^{\pi^\phi};\pi^\phi)+\big(r(s,X_s^{\pi^\phi},a_s^{\pi^\phi})+\gamma p(s,X_s^{\pi^\phi},a_s^{\pi^\phi},\pi^\phi)-\beta J(s,X_s^{\pi^\phi};\pi^\phi)\big)ds\right]+\gamma q(s,X_s^{\pi^\phi},a_s^{\pi^\phi},\phi)ds\right\}\right|X_t^{\pi^\phi}=x\right].

    Here dJ(s,Xsπϕ;πϕ)dJ(s,X_s^{\pi^\phi};\pi^\phi) is the Itô differential of the value function along the sampled trajectory, and πϕ\pi^\phi inside pp denotes its state- and time-conditioned action distribution. The expression uses observed state/action/reward samples and the current value function; it avoids requiring the system’s drift and diffusion coefficients. Thus policy-gradient estimation can be treated as evaluating an auxiliary expected reward. Since the expression depends on future observations after tt, it directly supports offline rather than purely online updates.

  2. Knowl 2 — Online policy optimality as martingale orthogonality

    theoretical result

    Suppose an interior parameter ϕ∗\phi^* maximizes J(0,x;πϕ)J(0,x;\pi^\phi) for every initial state xx, so the first-order condition holds. For a trajectory generated under πϕ∗\pi^{\phi^*}, let q=∂ϕp(t,x,a,πϕ(⋅∣t,x))q=\partial_\phi p(t,x,a,\pi^\phi(\cdot\mid t,x)) and let ηs\eta_s be any square-integrable scalar process adapted to the observed state history, while ζs\zeta_s is any square-integrable, policy-parameter-dimensional adapted process. Then the optimality condition implies

    E[∫0Tηs{[∂ϕlog⁡πϕ∗(as∣s,Xsπϕ∗)+ζs][dJ(s,Xsπϕ∗;πϕ∗)+(r(s,Xsπϕ∗,as)+γp(s,Xsπϕ∗,as,πϕ∗)−βJ(s,Xsπϕ∗;πϕ∗))ds]+γq(s,Xsπϕ∗,as,ϕ∗)ds}∣X0πϕ∗=x]=0.\mathbb E\left[\left.\int_0^T\eta_s\left\{\left[\partial_\phi\log\pi^{\phi^*}(a_s\mid s,X_s^{\pi^{\phi^*}})+\zeta_s\right]\left[dJ(s,X_s^{\pi^{\phi^*}};\pi^{\phi^*})+\big(r(s,X_s^{\pi^{\phi^*}},a_s)+\gamma p(s,X_s^{\pi^{\phi^*}},a_s,\pi^{\phi^*})-\beta J(s,X_s^{\pi^{\phi^*}};\pi^{\phi^*})\big)ds\right]+\gamma q(s,X_s^{\pi^{\phi^*}},a_s,\phi^*)ds\right\}\right|X_0^{\pi^{\phi^*}}=x\right]=0.

    The equality is a necessary condition, not a sufficiency guarantee. Choosing test processes that use only data already observed converts it into equations that can be solved incrementally by stochastic approximation; in particular, setting the multiplier to zero after the current time makes the update depend only on the past trajectory. This is the paper’s route to online policy improvement without estimating a future-trajectory gradient.

  3. Knowl 3 — Martingale characterization of policy evaluation under stochastic policies

    theoretical result

    For a stochastic feedback policy π\pi, write XπX^\pi for the observable state trajectory and atπ∼π(⋅∣t,Xtπ)a_t^\pi\sim\pi(\cdot\mid t,X_t^\pi) for its sampled action. Let J(t,x;π)J(t,x;\pi) be the policy value, hh the terminal reward, and β\beta the discount rate. Under the stated regularity conditions, policy evaluation is characterized by the terminal condition J(T,x;π)=h(x)J(T,x;\pi)=h(x) together with the following orthogonality condition for every square-integrable test process ξt\xi_t adapted to the observed state history:

    E[∫0Tξt{dJ(t,Xtπ;π)+[r(t,Xtπ,atπ)+γp(t,Xtπ,atπ,π(⋅∣t,Xtπ))−βJ(t,Xtπ;π)]dt}]=0.\mathbb E\left[\int_0^T\xi_t\left\{dJ(t,X_t^\pi;\pi)+\big[r(t,X_t^\pi,a_t^\pi)+\gamma p(t,X_t^\pi,a_t^\pi,\pi(\cdot\mid t,X_t^\pi))-\beta J(t,X_t^\pi;\pi)\big]dt\right\}\right]=0.

    Equivalently, along the averaged state process induced by the stochastic policy, the discounted value plus accumulated discounted regularized reward is a martingale. This extends the martingale-based policy-evaluation characterization to randomized policies and supplies the critic condition used by the actor–critic methods.

  4. Knowl 4 — Relaxed diffusion model for randomized continuous-time control

    model/method

    The analysis represents a stochastic policy by its action density π(a∣t,x)\pi(a\mid t,x) and averages the controlled dynamics over that distribution. If the sampled-action system has drift b(t,x,a)b(t,x,a) and diffusion matrix σ(t,x,a)\sigma(t,x,a), the relaxed state X~π\widetilde X^\pi satisfies

    dX~sπ=b~(s,X~sπ,π)ds+σ~(s,X~sπ,π)dWs,b~(t,x,π)=∫Ab(t,x,a)π(a∣t,x)da,σ~(t,x,π)=(∫Aσ2(t,x,a)π(a∣t,x)da)1/2.d\widetilde X_s^\pi=\widetilde b(s,\widetilde X_s^\pi,\pi)ds+\widetilde\sigma(s,\widetilde X_s^\pi,\pi)dW_s, \quad \widetilde b(t,x,\pi)=\int_A b(t,x,a)\pi(a\mid t,x)da, \quad \widetilde\sigma(t,x,\pi)=\left(\int_A\sigma^2(t,x,a)\pi(a\mid t,x)da\right)^{1/2}.

    Here WW is the environmental Brownian motion, and σ2=σσ⊤\sigma^2=\sigma\sigma^\top. The relaxed trajectory averages over policy randomization; it is not itself an observed sample path. Nevertheless, its expected objective equals the sampled-path objective

    J(t,x;π)=E[∫tTe−β(s−t)∫A(r(s,X~sπ,a)+γp(s,X~sπ,a,π(⋅∣s,X~sπ)))π(a∣s,X~sπ)da ds+e−β(T−t)h(X~Tπ)∣X~tπ=x].J(t,x;\pi)=\mathbb E\left[\left.\int_t^T e^{-\beta(s-t)}\int_A\big(r(s,\widetilde X_s^\pi,a)+\gamma p(s,\widetilde X_s^\pi,a,\pi(\cdot\mid s,\widetilde X_s^\pi))\big)\pi(a\mid s,\widetilde X_s^\pi)da\,ds+e^{-\beta(T-t)}h(\widetilde X_T^\pi)\right|\widetilde X_t^\pi=x\right].

    The regularizer pp scores an action under the current policy and γ≥0\gamma\ge0 weights it; differential entropy is one example. The relaxed formulation makes the value function amenable to PDE and martingale analysis, while algorithms use observable sampled trajectories.

  5. Knowl 5 — Offline episodic actor–critic updates

    algorithm

    The offline method repeatedly samples complete episodes, then updates the critic and actor from the stored trajectories. Use a horizon TT, step size Δt\Delta t, grid ti=iΔtt_i=i\Delta t for i=0,…,Ki=0,\ldots,K with K=⌊T/Δt⌋K=\lfloor T/\Delta t\rfloor, episode count NN, value approximator JθJ^\theta, policy πϕ\pi^\phi, regularizer pp, temperature γ\gamma, discount rate β\beta, learning rates αθ,αϕ\alpha_\theta,\alpha_\phi, and schedule l(j)l(j) at episode jj. The critic test function ξi\xi_i and actor test function ζi\zeta_i may depend on the trajectory history through time tit_i; ξi\xi_i has the dimension of θ\theta and ζi\zeta_i the dimension of ϕ\phi.

    Input: Initial state x0, T, Δt, N, Jθ, πφ, p, γ, β, αθ, αφ, schedule l, test functions ξ and ζ, and an environment simulator.
    Initialize θ and φ.
    For episode j = 1,...,N:
        Set x0 to the observed initial state.
        For i = 0,...,K-1:
            Evaluate ξi and ζi from the trajectory history through ti.
            Sample action ai from πφ(·|ti, xi).
            Apply ai; observe next state xi+1 and instantaneous reward ri.
            Compute δi = Jθ(ti+1, xi+1) - Jθ(ti, xi)
                       + ri Δt + γ p(ti, xi, ai, πφ(·|ti, xi)) Δt
                       - β Jθ(ti, xi) Δt.
        Compute Δθ = sum over i of ξi δi.
        Compute Δφ = sum over i of exp(-β ti) times
                     [(∂φ log πφ(ai|ti, xi) + ζi) δi
                      + γ ∂φ p(ti, xi, ai, πφ(·|ti, xi)) Δt].
        Update θ ← θ + l(j) αθ Δθ.
        Update φ ← φ + l(j) αφ Δφ.
    Output: Updated θ and φ.

    The critic increment is a discretization of the martingale orthogonality residual. The actor increment discretizes the sample-based policy-gradient representation; the whole future trajectory is available before the update, which is why this version is offline.

  6. Knowl 6 — Online incremental actor–critic updates

    algorithm

    The online algorithm updates after each observed transition and needs no future part of an episode. Let tk=kΔtt_k=k\Delta t, with K=⌊T/Δt⌋K=\lfloor T/\Delta t\rfloor per episode. Inputs are a value approximator JθJ^\theta, stochastic policy πϕ\pi^\phi, regularizer pp, discount rate β\beta, temperature γ\gamma, learning rates αθ,αϕ\alpha_\theta,\alpha_\phi, episode schedule l(j)l(j), and adapted test functions ξk\xi_k (critic dimension), ηk\eta_k (scalar actor multiplier), and ζk\zeta_k (policy-parameter dimension). Initialize θ,ϕ\theta,\phi and observe the initial state. At each step, sample an action, observe the next state and reward, and compute

    δk=Jθ(tk+1,xk+1)−Jθ(tk,xk)+rkΔt+γp(tk,xk,ak,πϕ)Δt−βJθ(tk,xk)Δt.\delta_k=J^\theta(t_{k+1},x_{k+1})-J^\theta(t_k,x_k)+r_k\Delta t+\gamma p(t_k,x_k,a_k,\pi^\phi)\Delta t-\beta J^\theta(t_k,x_k)\Delta t.
    For each episode j = 1,2,...:
        Observe the episode's initial state x0.
        For k = 0,...,K-1:
            Evaluate ξk, ηk, and ζk using the observed history through tk.
            Sample ak from πφ(·|tk, xk).
            Apply ak; observe xk+1 and instantaneous reward rk.
            Compute δk as defined in the accompanying equation.
            Compute Δθ = ξk δk.
            Compute Δφ = ηk [(∂φ log πφ(ak|tk, xk) + ζk) δk
                             + γ ∂φ p(tk, xk, ak, πφ(·|tk, xk)) Δt].
            Update θ ← θ + l(j) αθ Δθ.
            Update φ ← φ + l(j) αφ Δφ.
    Output: Continuously updated θ and φ.

    The policy score and regularizer derivative are evaluated at the current state, action, and policy. Compared with the offline method, this update uses only the current transition and past history; its test functions determine the particular temporal-difference variant.

  7. Knowl 7 — Ergodic policy evaluation and policy-gradient representation

    theoretical result

    For a time-homogeneous stochastic policy π\pi in an ergodic diffusion, the average regularized reward is a scalar V(π)V(\pi), independent of the initial state. A state function J(x;π)J(x;\pi) and this average reward satisfy the Poisson equation

    ∫A[LaJ(x;π)+r(x,a)+γp(x,a,π(⋅∣x))]π(a∣x)da−V(π)=0,\int_A\left[\mathcal L^a J(x;\pi)+r(x,a)+\gamma p(x,a,\pi(\cdot\mid x))\right]\pi(a\mid x)da-V(\pi)=0,

    where Laf=b(x,a)⋅∇f+12σ2(x,a):∇2f\mathcal L^a f=b(x,a)\cdot\nabla f+\tfrac12\sigma^2(x,a):\nabla^2 f is the diffusion generator. The state function is defined only up to an additive constant. Along the sampled state trajectory, J(Xtπ;π)+∫0t[r(Xsπ,asπ)+γp(Xsπ,asπ,π(⋅∣Xsπ))−V(π)]dsJ(X_t^\pi;\pi)+\int_0^t[r(X_s^\pi,a_s^\pi)+\gamma p(X_s^\pi,a_s^\pi,\pi(\cdot\mid X_s^\pi))-V(\pi)]ds is a martingale.

    For a parameterized policy πϕ\pi^\phi, let q(x,a,ϕ)=∂ϕp(x,a,πϕ(⋅∣x))q(x,a,\phi)=\partial_\phi p(x,a,\pi^\phi(\cdot\mid x)). Under the paper’s ergodicity and regularity conditions, its average-reward gradient can be represented as

    ∂ϕV(πϕ)=lim inf⁡T→∞1TE[∫0T{∂ϕlog⁡πϕ(at∣Xtπ)[dJ(Xtπ;πϕ)+(r(Xtπ,at)+γp(Xtπ,at,πϕ)−V(πϕ))dt]+γq(Xtπ,at,ϕ)dt}∣X0π=x].\partial_\phi V(\pi^\phi)=\liminf_{T\to\infty}\frac1T\mathbb E\left[\left.\int_0^T\left\{\partial_\phi\log\pi^\phi(a_t\mid X_t^\pi)\left[dJ(X_t^\pi;\pi^\phi)+\big(r(X_t^\pi,a_t)+\gamma p(X_t^\pi,a_t,\pi^\phi)-V(\pi^\phi)\big)dt\right]+\gamma q(X_t^\pi,a_t,\phi)dt\right\}\right|X_0^\pi=x\right].

    The long-run average permits incremental gradient estimation from a single trajectory after the state process approaches its stationary regime; an optimality-based orthogonality form also permits stochastic-approximation policy updates.

  8. Knowl 8 — Incremental actor–critic algorithm for ergodic tasks

    algorithm

    For continuing ergodic learning, the paper updates a critic Jθ(x)J^\theta(x), an average-reward estimate VV, and policy parameters ϕ\phi from one trajectory. Choose step size Δt\Delta t, learning rates αθ,αϕ,αV\alpha_\theta,\alpha_\phi,\alpha_V, time-based schedule l(t)l(t), policy πϕ\pi^\phi, regularizer pp, temperature γ\gamma, critic test function ξk\xi_k, actor multiplier ηk\eta_k, and actor test function ζk\zeta_k, all evaluated from history through step kk. Initialize x0,θ,ϕ,Vx_0,\theta,\phi,V.

    Repeat indefinitely:
        Evaluate ξk, ηk, and ζk from the observed state history.
        Sample ak from πφ(·|xk).
        Apply ak; observe next state xk+1 and instantaneous reward rk.
        Compute δk = Jθ(xk+1) - Jθ(xk) + rk Δt
                    + γ p(xk, ak, πφ(·|xk)) Δt - V Δt.
        Compute Δθ = ξk δk.
        Compute ΔV = δk.
        Compute Δφ = ηk [(∂φ log πφ(ak|xk) + ζk) δk
                         + γ ∂φ p(xk, ak, πφ(·|xk)) Δt].
        Update θ ← θ + l(k Δt) αθ Δθ.
        Update V ← V + l(k Δt) αV ΔV.
        Update φ ← φ + l(k Δt) αφ Δφ.
        Set xk ← xk+1 and advance k.

    The temporal-difference residual subtracts the current average-reward estimate, so the critic estimates the relative-value function while VV tracks long-run reward. The actor update uses the ergodic policy-gradient orthogonality condition.

  9. Knowl 9 — Mean–variance portfolio experiments: offline and online performance

    empirical result

    The portfolio experiments used a pre-committed agent targeting expected terminal wealth z=1.4z=1.4, initial wealth x0=1x_0=1, horizon T=1T=1 year, time step 1/2521/252, and entropy weight γ=0.1\gamma=0.1. Training observations were generated from geometric Brownian stock prices for market drift μ∈{−0.5,−0.3,−0.1,0,0.1,0.3,0.5}\mu\in\{-0.5,-0.3,-0.1,0,0.1,0.3,0.5\} and volatility σ∈{0.1,0.2,0.3,0.4}\sigma\in\{0.1,0.2,0.3,0.4\}. The offline training used 20 years of simulated data, 128 sampled one-year trajectories per iteration, and 2×1042\times10^4 iterations; the online method updated per step over the sequential 20-year sample. Both methods were repeated 100 times, and the reported out-of-sample measures were terminal-wealth mean, variance, and Sharpe ratio.

    For μ=0.5,σ=0.1\mu=0.5,\sigma=0.1, the prior method of Wang and Zhou reported Sharpe ratio 66 (standard deviation 0.0870.087), while the paper’s offline and online methods reported 7.357.35 (0.0550.055) and 10.8210.82 (0.150.15), respectively. For μ=−0.5,σ=0.1\mu=-0.5,\sigma=0.1, the corresponding Sharpe ratios were 6.696.69 (0.0960.096), 8.158.15 (0.060.06), and 7.437.43 (0.040.04). These examples illustrate the paper’s finding that its offline method outperformed the earlier method on Sharpe ratio in most tested market settings, while online learning sometimes achieved higher Sharpe ratios than offline learning.

    Performance was less favorable and more variable in several low-drift cases (μ=0\mu=0 or ±0.1\pm0.1), especially at higher volatility; for example, at μ=0,σ=0.4\mu=0,\sigma=0.4, offline learning’s reported Sharpe ratio was 0.010.01 with standard deviation 0.0490.049, and online learning’s was 0.010.01 with standard deviation 0.0480.048. The authors report offline learning as more reliable for reaching the target return, while noting that large variability in some cases made method-to-method differences statistically inconclusive.

  10. Knowl 10 — Ergodic LQ simulation demonstrates policy and reward convergence

    empirical result

    The single-trajectory demonstration used the ergodic linear–quadratic system dXt=(AXt+Bat)dt+(CXt+Dat)dWtdX_t=(AX_t+Ba_t)dt+(CX_t+Da_t)dW_t with A=−1A=-1, B=C=0B=C=0, D=1D=1, initial state x0=0x_0=0, reward r(x,a)=−(Mx2/2+Rxa+Na2/2+Px+Qa)r(x,a)=-(Mx^2/2+Rxa+Na^2/2+Px+Qa), and M=N=Q=2M=N=Q=2, R=P=1R=P=1. The entropy-regularization weight was γ=0.1\gamma=0.1 and the time step was Δt=0.01\Delta t=0.01. The policy was parameterized as a Gaussian with mean ϕ1x+ϕ2\phi_1x+\phi_2 and variance eϕ3e^{\phi_3}; the critic was quadratic in state and included a learned average reward. The actor–critic used TD(0), with critic test function ξt=∂θJθ(Xt)\xi_t=\partial_\theta J^\theta(X_t), actor multiplier ηt=1\eta_t=1, and ζt=0\zeta_t=0. Parameters started at zero, the initial learning rate was 0.0010.001, and the schedule was l(t)=1/max⁡{1,log⁡t}l(t)=1/\max\{1,\log t\}. The experiment was repeated 100 times.

    The plotted single-trajectory policy parameters approach the theoretically optimal parameters for the entropy-regularized problem. The running average reward initially falls, which the authors attribute as a possibility to the state process not yet being stationary and early policies being far from optimal; it then rises toward the optimal average reward under the exploration constraint. It approaches more slowly than the policy parameters because the early poor rewards continue to affect the cumulative average. The learned reward is expected to approach the omniscient optimum less the exploration cost, not the unregularized optimum, because the learning policy remains stochastic.

Coverage note — Proof details and appendix derivations are omitted because they support the stated characterizations rather than adding separate results; background and discrete-time literature comparisons are also excluded.

References

  1. 1.V. Aleksandrov, V. Sysoev, and V. Shemeneva. Stochastic optimization. Engineering Cybernetics, 5:11–16, 1968.
  2. 2.R. Aragon-Gómez and J. B. Clempner. Traffic-signal control reinforcement learning approach for continuous-time Markov games. Engineering Applications of Artificial Intelligence, 89:103415, 2020.
  3. 3.A. Arapostathis, V. S. Borkar, and M. K. Ghosh. Ergodic control of diffusion processes, volume 143. Cambridge University Press, 2012.
  4. 4.L. C. Baird. Advantage updating. Technical report, Write Lab Wright-Patterson Air Force Base, OH 45433-7301, USA, 1993.
  5. 5.A. G. Barto, R. S. Sutton, and C. W. Anderson. Neuronlike adaptive elements that can solve difficult learning control problems. IEEE Transactions on Systems, Man, and Cybernetics, (5):834–846, 1983.
  6. 6.S. Basak and G. Chabakauri. Dynamic mean-variance asset allocation. The Review of Financial Studies, 23(8):2970–3016, 2010.
  7. 7.M. Basei, X. Guo, A. Hu, and Y. Zhang. Logarithmic regret for episodic continuous-time linear-quadratic reinforcement learning over a finite-time horizon. arXiv preprint arXiv:2006.15316, 2020.
  8. 8.C. Beck, M. Hutzenthaler, and A. Jentzen. On nonlinear Feynman–Kac formulas for viscosity solutions of semilinear parabolic partial differential equations. Stochastics and Dynamics, page 2150048, 2021.
  9. 9.A. Bensoussan and J. Frehse. On Bellman equations of ergodic control in Rn. In Applied Stochastic Analysis, pages 21–29. Springer, 1992.
  10. 10.S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee. Natural actor–critic algorithms. Automatica, 45(11):2471–2482, 2009.
  11. 11.T. Björk, A. Murgoci, and X. Y. Zhou. Mean–variance portfolio optimization with state-dependent risk aversion. Mathematical Finance, 24(1):1–24, 2014.
  12. 12.V. S. Borkar and M. K. Ghosh. Ergodic control of multidimensional diffusions I: The existence results. SIAM Journal on Control and Optimization, 26(1):112–126, 1988.
  13. 13.V. S. Borkar and M. K. Ghosh. Ergodic control of multidimensional diffusions II: Adaptive control. Applied Mathematics and Optimization, 21(1):191–220, 1990.
  14. 14.S. J. Bradtke and A. G. Barto. Linear least-squares algorithms for temporal difference learning. Machine Learning, 22(1):33–57, 1996.
  15. 15.M. Dai, Y. Dong, and Y. Jia. Learning equilibrium mean-variance strategy. SSRN preprint SSRN:3770818, 2020.
  16. 16.M. Dai, H. Jin, S. Kou, and Y. Xu. A dynamic mean-variance analysis for log returns. Management Science, 67(2):1093–1108, 2021.
  17. 17.T. Degris, P. M. Pilarski, and R. S. Sutton. Model-free reinforcement learning with continuous action in practice. In 2012 American Control Conference (ACC), pages 2177–2182. IEEE, 2012.
  18. 18.K. Doya. Reinforcement learning in continuous time and space. Neural Computation, 12(1):219–245, 2000.
  19. 19.Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel. Benchmarking deep reinforcement learning for continuous control. In International Conference on Machine Learning, pages 1329–1338. PMLR, 2016.
  20. 20.I. Ekeland and A. Lazrak. Being serious about non-commitment: subgame perfect equilibrium in continuous time. arXiv preprint math/0604264, 2006.
  21. 21.W. Fedus, P. Ramachandran, R. Agarwal, Y. Bengio, H. Larochelle, M. Rowland, and W. Dabney. Revisiting fundamentals of experience replay. In International Conference on Machine Learning, pages 3061–3071. PMLR, 2020.
  22. 22.W. H. Fleming and H. M. Soner. Controlled Markov Processes and Viscosity Solutions, volume 25. Springer Science & Business Media, 2006.
  23. 23.N. Frémaux, H. Sprekeler, and W. Gerstner. Reinforcement learning using a continuous time actor-critic framework with spiking neurons. PLoS Computational Biology, 9(4):e1003024, 2013.
  24. 24.X. Gao, Z. Q. Xu, and X. Y. Zhou. State-dependent temperature control for Langevin diffusions. SIAM Journal on Control and Optimization, 60(3):1250–1268, 2022.
  25. 25.P. W. Glynn. Likelihood ratio gradient estimation for stochastic systems. Communications of the ACM, 33(10):75–84, 1990.
  26. 26.X. Guo, R. Xu, and T. Zariphopoulou. Entropy regularization for mean field games with learning. Mathematics of Operations Research, 2022.
  27. 27.T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al. Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905, 2018.
  28. 28.Y. Jia and X. Y. Zhou. Policy evaluation and temporal-difference learning in continuous time and space: A martingale approach. Journal of Machine Learning Research, 23(154):1–55, 2022a.
  29. 29.Y. Jia and X. Y. Zhou. q-Learning in continuous time. arXiv preprint; http://arxiv.org/abs/2207.00713, 2022b.
  30. 30.I. Karatzas and S. Shreve. Brownian motion and stochastic calculus, volume 113. Springer, 2014.
  31. 31.J. Kim, J. Shin, and I. Yang. Hamilton-Jacobi deep Q-Learning for deterministic continuous-time systems with Lipschitz continuous controls. Journal of Machaine Learning Research, 22:206–1, 2021.
  32. 32.B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. Al Sallab, S. Yogamani, and P. Pérez. Deep reinforcement learning for autonomous driving: A survey. IEEE Transactions on Intelligent Transportation Systems, 2021.
  33. 33.J. Lee and R. S. Sutton. Policy iterations for reinforcement learning problems in continuous time and space—Fundamental theory and methods. Automatica, 126:109421, 2021.
  34. 34.D. Li and W.-L. Ng. Optimal dynamic portfolio selection: Multiperiod mean-variance formulation. Mathematical Finance, 10(3):387–406, 2000.
  35. 35.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  36. 36.H. R. Maei, C. Szepesvari, S. Bhatnagar, D. Precup, D. Silver, and R. S. Sutton. Convergent temporal-difference learning with arbitrary smooth function approximation. In NIPS, pages 1204–1212, 2009.
  37. 37.P. Marbach and J. N. Tsitsiklis. Simulation-based optimization of Markov reward processes. IEEE Transactions on Automatic Control, 46(2):191–209, 2001.
  38. 38.S. P. Meyn and R. L. Tweedie. Markov Chains and Stochastic Stability. Springer Science & Business Media, 2012.
  39. 39.V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning, pages 1928–1937. PMLR, 2016.
  40. 40.R. Munos. Policy gradient in continuous time. Journal of Machine Learning Research, 7:771–791, 2006.
  41. 41.R. Munos and P. Bourgine. Reinforcement learning for continuous stochastic control problems. Advances in Neural Information Processing Systems, 10, 1997.
  42. 42.N. Sandrić. A note on the Birkhoff ergodic theorem. Results in Mathematics, 72(1):715–730, 2017.
  43. 43.J. Schulman, X. Chen, and P. Abbeel. Equivalence between policy gradients and soft Q-learning. arXiv preprint arXiv:1704.06440, 2017a.
  44. 44.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017b.
  45. 45.D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller. Deterministic policy gradient algorithms. In International Conference on Machine Learning, pages 387–395. PMLR, 2014.
  46. 46.D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al. Mastering the game of go without human knowledge. Nature, 550(7676):354–359, 2017.
  47. 47.R. H. Strotz. Myopia and inconsistency in dynamic utility maximization. The Review of Economic Studies, 23(3):165–180, 1955.
  48. 48.R. S. Sutton. Learning to predict by the methods of temporal differences. Machine Learning, 3(1):9–44, 1988.
  49. 49.R. S. Sutton and A. G. Barto. Reinforcement Learning: An Introduction. Cambridge, MA: MIT Press, 2018.
  50. 50.R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour. Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems, 12, 1999.
  51. 51.R. S. Sutton, S. Singh, and D. McAllester. Comparing policy-gradient algorithms. IEEE Transactions on Systems, Man, and Cybernetics, 2000.
  52. 52.R. S. Sutton, C. Szepesvári, and H. R. Maei. A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation. In NIPS, 2008.
  53. 53.R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora. Fast gradient-descent methods for temporal-difference learning with linear function approximation. In Proceedings of the 26th Annual International Conference on Machine Learning, pages 993–1000, 2009.
  54. 54.L. Szpruch, T. Treetanthiploet, and Y. Zhang. Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models. arXiv preprint arXiv:2112.10264, 2021.
  55. 55.C. Tallec, L. Blier, and Y. Ollivier. Making deep Q-learning methods robust to time discretization. In International Conference on Machine Learning, pages 6096–6104. PMLR, 2019.
  56. 56.W. Tang, P. Y. Zhang, and X. Y. Zhou. Exploratory HJB equations and their convergence. arXiv preprint arXiv:2109.10269, 2021.
  57. 57.K. G. Vamvoudakis and F. L. Lewis. Online actor–critic algorithm to solve the continuous-time infinite horizon optimal control problem. Automatica, 46(5):878–888, 2010.
  58. 58.H. Wang and X. Y. Zhou. Continuous-time mean–variance portfolio selection: A reinforcement learning framework. Mathematical Finance, 30(4):1273–1308, 2020.
  59. 59.H. Wang, T. Zariphopoulou, and X. Y. Zhou. Reinforcement learning in continuous time and space: A stochastic control approach. Journal of Machine Learning Research, 21(198):1–34, 2020.
  60. 60.W. Wang, J. Han, Z. Yang, and Z. Wang. Global convergence of policy gradient for linear-quadratic mean-field control/game in continuous time. In International Conference on Machine Learning, pages 10772–10782. PMLR, 2021.
  61. 61.P. Wawrzynski. Learning to control a 6-degree-of-freedom walking robot. In EUROCON 2007-The International Conference on” Computer as a Tool”, pages 698–705. IEEE, 2007.
  62. 62.R. J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3):229–256, 1992.
  63. 63.C. Yildiz, M. Heinonen, and H. Lähdesmäki. Continuous-time model-based reinforcement learning. In International Conference on Machine Learning, pages 12009–12018. PMLR, 2021.
  64. 64.J. Yong and X. Y. Zhou. Stochastic Controls: Hamiltonian Systems and HJB Equations. New York, NY: Spinger, 1999.
  65. 65.D. Zambrano, P. R. Roelfsema, and S. M. Bohte. Continuous-time on-policy neural reinforcement learning of working memory tasks. In 2015 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2015.
  66. 66.S. Zhang and R. Sutton. A deeper look at experience replay. arXiv preprint arXiv:1712.01275, 2017.
  67. 67.T. Zhao, H. Hachiya, G. Niu, and M. Sugiyama. Analysis and improvement of policy gradient estimation. In NIPS, pages 262–270. Citeseer, 2011.
  68. 68.X. Y. Zhou and D. Li. Continuous-time mean-variance portfolio selection: A stochastic LQ framework. Applied Mathematics and Optimization, 42(1):19–33, 2000.

Citation

MLA
Jia, Y., and X. Y. Zhou. “Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms”. Journal of Machine Learning Research, vol. 23, no. 275, 2022, pp. 1–0, https://www.jmlr.org/papers/v23/21-1387.html.
APA
Jia, Y., & Zhou, X. Y. (2022). Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms. Journal of Machine Learning Research, 23(275), 1–50. https://www.jmlr.org/papers/v23/21-1387.html
Chicago
Jia, Y., and X. Y. Zhou. 2022. “Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms”. Journal of Machine Learning Research 23 (275): 1–50. https://www.jmlr.org/papers/v23/21-1387.html.
Harvard
Jia, Y. and Zhou, X.Y. (2022) “Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms”, Journal of Machine Learning Research, 23(275), pp. 1–50. Available at: https://www.jmlr.org/papers/v23/21-1387.html.
Vancouver
1. Jia Y, Zhou XY (2022) Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms. Journal of Machine Learning Research 23:1–50

BibTeX

@article{JMLR:v23:21-1387,
  author  = {Yanwei Jia and Xun Yu Zhou},
  title   = {Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms},
  journal = {Journal of Machine Learning Research},
  year    = {2022},
  volume  = {23},
  number  = {275},
  pages   = {1--50},
  url     = {http://jmlr.org/papers/v23/21-1387.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/