The Option-Critic Architecture
Pierre-Luc BaconJean HarbDoina Precup
Develops the option-critic architecture by extending policy gradient theorems to temporal abstractions, enabling reinforcement learning agents to learn hierarchical options and their termination conditions end-to-end without engineered subgoals.
Temporal abstraction—the ability of an automated agent to plan and act over extended time horizons rather than step-by-step—is essential for scaling artificial intelligence to complex, high-dimensional tasks. Historically, discovering these multi-step subroutines, known as options, required manual engineering, pre-specified subgoals, extra reward signals, or expert demonstrations. These traditional methods are computationally intensive, often as costly as solving the target problem itself, and scale poorly to continuous or large-scale environments.
The article develops and demonstrates an end-to-end framework, called the option-critic architecture, that autonomously learns both internal option policies and their termination conditions directly from environmental rewards. The objective is to enable simultaneous, gradient-based learning of sub-behaviors and high-level strategy without requiring human guidance or auxiliary rewards.
To accomplish this, the authors derive theoretical policy gradient theorems tailored to intra-option policies and termination functions within a two-timescale actor-critic architecture. The system updates high-level option values at a fast rate while refining lower-level option behaviors and stopping rules at a slower rate. The framework was evaluated across discrete navigation tasks (the Four-Rooms domain), continuous control simulations (the Pinball domain), and high-dimensional visual environments via deep neural networks across four Atari 2600 games in the Arcade Learning Environment.
The evaluations yielded several key findings. First, the option-critic framework learns effective sub-behaviors entirely from scratch without suffering the learning slowdowns typical of previous option discovery methods. Second, in dynamic environments where objectives change abruptly, the learned options enabled agents to adapt and recover significantly faster than standard algorithms relying solely on primitive actions. Third, in the continuous Pinball domain, the system learned near-optimal behaviors within 40 episodes, avoiding the mandatory warm-up delays required by earlier baselines. Finally, when integrated with deep neural networks on Atari benchmarks, the architecture successfully learned eight specialized options end-to-end, outperforming baseline Deep Q-Networks on three out of four tested games (Asterix, Seaquest, and Zaxxon) within 200 training episodes.
These findings show that autonomous temporal abstraction can be achieved efficiently using standard policy gradient optimization, eliminating the need for expensive combinatorial subgoal searches. For applied decision-making and deployment, this reduces the engineering overhead and domain expertise needed to build hierarchical decision-making systems. It also improves flexibility in non-stationary environments where goals shift over time.
Organizations implementing this approach should apply regularization techniques identified in the article, such as entropy regularization to avoid deterministic policy collapse and margin-based advantage adjustments to prevent learned options from prematurely shrinking into single primitive steps. Future work and pilot implementations should investigate methods to autonomously restrict initiation sets—specifying where particular options can start—to ensure computational efficiency as state spaces expand.
A primary limitation of this framework is the operational assumption that every option is available in every state, which differs from biological and classical hierarchical models where specific skills are restricted to relevant contexts. Additionally, gradient estimators in discounted reinforcement learning introduce mild theoretical bias, though empirical results demonstrate robust convergence and performance.
- Paper: Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning, Richard S. Sutton et al. (1999). It introduces the fundamental options framework for temporal abstraction and semi-Markov decision processes upon which the Option-Critic architecture directly builds.
- Paper: Actor-Critic Algorithms, Vijay Konda et al. (1999). It establishes the foundational two-time-scale actor-critic framework and gradient convergence results adapted by Option-Critic to optimize hierarchical option policies.
- Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). It derives the standard Policy Gradient Theorem with function approximation that the authors generalize to intra-option policies and option termination functions.
- Paper: Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition, Thomas G. Dietterich (1999). It provides foundational background on hierarchical reinforcement learning and value function decomposition across temporal abstractions.
- Paper: A Natural Policy Gradient, Sham M. Kakade (2001). It introduces natural policy gradient updates that motivate stable policy optimization in parameterized actor-critic architectures.
- Paper: Learning to Predict by the Methods of Temporal Differences, Richard S. Sutton (1988). It provides the foundational temporal-difference learning methods used to evaluate value functions and train critics in reinforcement learning.
- Paper: Deep Reinforcement Learning: An Overview, Yuxi Li (2017). It surveys the broader landscape of deep reinforcement learning algorithms, framing how hierarchical actor-critic methods fit into end-to-end representation learning.
- Paper: Addressing Function Approximation Error in Actor-Critic Methods, Scott Fujimoto et al. (2018). It introduces techniques like clipped double Q-learning and delayed updates to address value overestimation issues common in continuous actor-critic models like Option-Critic.
- Paper: Soft Actor-Critic Algorithms and Applications, Tuomas Haarnoja et al. (2018). It develops maximum-entropy actor-critic algorithms that can be incorporated into hierarchical frameworks to improve exploration and stability in continuous control.
- Paper: Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor, Tuomas Haarnoja et al. (2018). It formulates off-policy soft actor-critic optimization with stochastic actors, offering an alternative paradigm for robust policy gradient learning.
- Paper: Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments, Ryan Lowe et al. (2017). It extends centralized-critic policy gradient architectures to multi-agent environments with mixed cooperative and competitive dynamics.
