Policy Gradient and Actor-Critic Learning in Continuous Time and Space: Theory and Algorithms
Yanwei JiaXun Yu Zhou
Establishes a theoretical foundation for continuous-time and continuous-space policy gradient methods by reformulating the gradient estimation as a policy evaluation problem and deriving both offline and online actor-critic algorithms via martingale conditions.
Real-world decision-making systems—such as high-frequency trading platforms, autonomous vehicles, and industrial robotics—operate continuously in time and space rather than in artificial discrete intervals. Standard reinforcement learning techniques typically approximate reality by discretizing time upfront, which often produces algorithms that are unstable and highly sensitive to the chosen time step. The article establishes a rigorous mathematical foundation for continuous-time and continuous-space reinforcement learning, demonstrating model-free algorithms that learn optimal policies directly from observed sample data without relying on prior knowledge of environmental dynamics.
The research adopts an exploratory stochastic control framework that uses entropy regularization to balance exploration with exploitation. By developing the theoretical policy gradient in continuous time, the authors show that finding the policy gradient mathematically reduces to an auxiliary policy evaluation problem. Using stochastic calculus and martingale techniques—mathematical methods that characterize fair-game drift conditions—the article removes the need to know the underlying equations of the environment. The resulting framework enables two complementary actor-critic learning approaches: an offline method that updates decision rules across complete simulation episodes, and an online method that updates parameters in real time using only past observations.
The findings demonstrate the effectiveness and flexibility of these algorithms across both episodic and long-term average tasks. First, the derived continuous-time policy gradient naturally incorporates advantage-style baseline updates, improving learning stability without requiring manual baseline tuning. Second, in simulated financial portfolio selection over a 20-year horizon, the proposed offline actor-critic algorithm achieved significantly higher risk-adjusted returns (Sharpe ratios) than existing benchmark methods across various market conditions, while reliably hitting targeted return levels. Third, in continuous linear-quadratic control simulations, online real-time learning converged directly toward theoretical performance limits, with the remaining gap strictly accounting for the cost of active exploration.
These results provide a validated pathway for deploying model-free reinforcement learning in mission-critical, high-frequency environments where upfront time discretization is impractical or risky. Organizations facing stationary environments with historical trajectory data can use offline actor-critic algorithms to maximize sample efficiency and return stability. Conversely, for large-scale or non-stationary operations, online incremental updating reduces computational storage and continuously adapts to incoming data streams.
While the theoretical convergence and simulation evidence are strong, practitioners should exercise caution regarding specific operational constraints. Offline algorithms outperform online counterparts in finite datasets because they reuse sampled paths, whereas online methods require longer learning horizons to overcome early sub-optimal trials. Performance remains sensitive to the choice of exploration temperature and policy parameterizations. Future operational implementations should evaluate offline versus online trade-offs via domain-specific pilot testing and explore replay techniques to enhance online sample efficiency.
- Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). Its policy-gradient theorem supplies the foundational gradient formulation that this paper recasts as policy evaluation in continuous time and space.
- Paper: Actor-Critic Algorithms, Vijay R. Konda et al. (1999). Its two-time-scale actor-critic framework provides the discrete-time actor/critic foundations for understanding the paper’s simultaneous policy and value updates.
No sufficiently relevant recommendations were found.
