The Option-Critic Architecture

Pierre-Luc BaconJean HarbDoina Precup

article2016AAAI1,352 citations

Develops the option-critic architecture by extending policy gradient theorems to temporal abstractions, enabling reinforcement learning agents to learn hierarchical options and their termination conditions end-to-end without engineered subgoals.

Listen

Temporal abstraction—the ability of an automated agent to plan and act over extended time horizons rather than step-by-step—is essential for scaling artificial intelligence to complex, high-dimensional tasks. Historically, discovering these multi-step subroutines, known as options, required manual engineering, pre-specified subgoals, extra reward signals, or expert demonstrations. These traditional methods are computationally intensive, often as costly as solving the target problem itself, and scale poorly to continuous or large-scale environments.

The article develops and demonstrates an end-to-end framework, called the option-critic architecture, that autonomously learns both internal option policies and their termination conditions directly from environmental rewards. The objective is to enable simultaneous, gradient-based learning of sub-behaviors and high-level strategy without requiring human guidance or auxiliary rewards.

To accomplish this, the authors derive theoretical policy gradient theorems tailored to intra-option policies and termination functions within a two-timescale actor-critic architecture. The system updates high-level option values at a fast rate while refining lower-level option behaviors and stopping rules at a slower rate. The framework was evaluated across discrete navigation tasks (the Four-Rooms domain), continuous control simulations (the Pinball domain), and high-dimensional visual environments via deep neural networks across four Atari 2600 games in the Arcade Learning Environment.

The evaluations yielded several key findings. First, the option-critic framework learns effective sub-behaviors entirely from scratch without suffering the learning slowdowns typical of previous option discovery methods. Second, in dynamic environments where objectives change abruptly, the learned options enabled agents to adapt and recover significantly faster than standard algorithms relying solely on primitive actions. Third, in the continuous Pinball domain, the system learned near-optimal behaviors within 40 episodes, avoiding the mandatory warm-up delays required by earlier baselines. Finally, when integrated with deep neural networks on Atari benchmarks, the architecture successfully learned eight specialized options end-to-end, outperforming baseline Deep Q-Networks on three out of four tested games (Asterix, Seaquest, and Zaxxon) within 200 training episodes.

These findings show that autonomous temporal abstraction can be achieved efficiently using standard policy gradient optimization, eliminating the need for expensive combinatorial subgoal searches. For applied decision-making and deployment, this reduces the engineering overhead and domain expertise needed to build hierarchical decision-making systems. It also improves flexibility in non-stationary environments where goals shift over time.

Organizations implementing this approach should apply regularization techniques identified in the article, such as entropy regularization to avoid deterministic policy collapse and margin-based advantage adjustments to prevent learned options from prematurely shrinking into single primitive steps. Future work and pilot implementations should investigate methods to autonomously restrict initiation sets—specifying where particular options can start—to ensure computational efficiency as state spaces expand.

A primary limitation of this framework is the operational assumption that every option is available in every state, which differs from biological and classical hierarchical models where specific skills are restricted to relevant contexts. Additionally, gradient estimators in discounted reinforcement learning introduce mild theoretical bias, though empirical results demonstrate robust convergence and performance.

arXiv: 1609.05140
Cover for The Option-Critic Architecture

Abstract

Temporal abstraction is key to scaling up learning and planning in reinforcement learning. While planning with temporally extended actions is well understood, creating such abstractions autonomously from data has remained challenging. We tackle this problem in the framework of options [Sutton, Precup & Singh, 1999; Precup, 2000]. We derive policy gradient theorems for options and propose a new option-critic architecture capable of learning both the internal policies and the termination conditions of options, in tandem with the policy over options, and without the need to provide any additional rewards or subgoals. Experimental results in both discrete and continuous environments showcase the flexibility and efficiency of the framework.

Table of Contents

  • Introduction
  • Preliminaries and Notation
  • Learning Options
  • Algorithms and Architecture
  • Experiments
  • Pinball Domain
  • Arcade Learning Environment
  • Related Work
  • Discussion
  • Acknowledgements
  • Appendix
  • References

Knowls

  1. Knowl 1 — Intra-Option Policy Gradient Theorem

    theoretical result

    Let a Markov Decision Process be defined by state space S\mathcal{S}, action space A\mathcal{A}, transition model P(s′∣s,a)P(s' \mid s, a), reward function r(s,a)r(s, a), and discount factor γ∈[0,1)\gamma \in [0, 1). Let Ω\Omega be a set of Markov options available everywhere, where each option ω∈Ω\omega \in \Omega has a stochastic intra-option policy πω,θ(a∣s)\pi_{\omega,\theta}(a \mid s) parameterized by differentiable parameter vector θ\theta, a termination function βω,ϑ(s)∈[0,1]\beta_{\omega,\vartheta}(s) \in [0, 1], and a policy over options πΩ(ω∣s)\pi_\Omega(\omega \mid s).

    The gradient of the expected discounted return from designated starting state s0s_0 and starting option ω0\omega_0, ρ(Ω,θ,ϑ,s0,ω0)=EΩ,θ,ϑ[∑t=0∞γtrt+1∣s0,ω0]\rho(\Omega, \theta, \vartheta, s_0, \omega_0) = \mathbb{E}_{\Omega,\theta,\vartheta}\left[\sum_{t=0}^\infty \gamma^t r_{t+1} \mid s_0, \omega_0\right], with respect to the intra-option policy parameters θ\theta is:

    ∂ρ(Ω,θ,ϑ,s0,ω0)∂θ=∑s∈S,ω∈ΩμΩ(s,ω∣s0,ω0)∑a∈A∂πω,θ(a∣s)∂θQU(s,ω,a)\frac{\partial \rho(\Omega, \theta, \vartheta, s_0, \omega_0)}{\partial \theta} = \sum_{s \in \mathcal{S}, \omega \in \Omega} \mu_\Omega(s, \omega \mid s_0, \omega_0) \sum_{a \in \mathcal{A}} \frac{\partial \pi_{\omega,\theta}(a \mid s)}{\partial \theta} Q_U(s, \omega, a)

    where μΩ(s,ω∣s0,ω0)=∑t=0∞γtP(st=s,ωt=ω∣s0,ω0)\mu_\Omega(s, \omega \mid s_0, \omega_0) = \sum_{t=0}^\infty \gamma^t P(s_t = s, \omega_t = \omega \mid s_0, \omega_0) is the discounted weighting of state-option pairs along trajectories generated under call-and-return option execution, and QU(s,ω,a)=r(s,a)+γ∑s′∈SP(s′∣s,a)U(ω,s′)Q_U(s, \omega, a) = r(s, a) + \gamma \sum_{s' \in \mathcal{S}} P(s' \mid s, a) U(\omega, s') is the value of executing primitive action aa in state ss under option ω\omega, with U(ω,s′)U(\omega, s') denoting the option-value function upon arrival in state s′s'.

  2. Knowl 2 — Termination Gradient Theorem

    theoretical result

    Let a Markov Decision Process have state space S\mathcal{S}, action space A\mathcal{A}, and discount factor γ∈[0,1)\gamma \in [0, 1). Let Ω\Omega be a set of Markov options, where each option ω∈Ω\omega \in \Omega has a stochastic termination function βω,ϑ(s)∈[0,1]\beta_{\omega,\vartheta}(s) \in [0, 1] parameterized by differentiable parameter vector ϑ\vartheta, an intra-option policy πω,θ(a∣s)\pi_{\omega,\theta}(a \mid s), and a policy over options πΩ(ω∣s)\pi_\Omega(\omega \mid s).

    The gradient of the expected discounted return with respect to ϑ\vartheta, conditioned on arriving at next state s1s_1 with currently executing option ω0\omega_0, is:

    ∂U(ω0,s1)∂ϑ=−∑s′∈S,ω∈ΩμΩ(s′,ω∣s1,ω0)∂βω,ϑ(s′)∂ϑAΩ(s′,ω)\frac{\partial U(\omega_0, s_1)}{\partial \vartheta} = - \sum_{s' \in \mathcal{S}, \omega \in \Omega} \mu_\Omega(s', \omega \mid s_1, \omega_0) \frac{\partial \beta_{\omega,\vartheta}(s')}{\partial \vartheta} A_\Omega(s', \omega)

    where AΩ(s′,ω)=QΩ(s′,ω)−VΩ(s′)A_\Omega(s', \omega) = Q_\Omega(s', \omega) - V_\Omega(s') is the option advantage function measuring the value of continuing option ω\omega in state s′s' relative to the state value VΩ(s′)=∑ωˉ∈ΩπΩ(ωˉ∣s′)QΩ(s′,ωˉ)V_\Omega(s') = \sum_{\bar{\omega} \in \Omega} \pi_\Omega(\bar{\omega} \mid s') Q_\Omega(s', \bar{\omega}), and μΩ(s′,ω∣s1,ω0)=∑t=0∞γtP(st+1=s′,ωt=ω∣s1,ω0)\mu_\Omega(s', \omega \mid s_1, \omega_0) = \sum_{t=0}^\infty \gamma^t P(s_{t+1} = s', \omega_t = \omega \mid s_1, \omega_0) is the discounted weighting of state-option pairs shifted by one time step.

    Because the advantage AΩ(s′,ω)A_\Omega(s', \omega) is negative when option ω\omega performs worse than average option selection at s′s', the negative sign in the gradient drives an increase in termination probability βω,ϑ(s′)\beta_{\omega,\vartheta}(s') whenever an option becomes suboptimal.

  3. Knowl 3 — Option-Value Function Upon Arrival and State-Option Value Functions

    definition

    In the call-and-return option execution model, an agent selects option ω∈Ω\omega \in \Omega from policy over options πΩ\pi_\Omega, executes intra-option policy πω,θ\pi_{\omega,\theta} until option termination according to βω,ϑ\beta_{\omega,\vartheta}, and re-evaluates option choice. The value functions are defined over the augmented state-option space as follows:

    1. The option-value function upon arrival U:Ω×S→RU: \Omega \times \mathcal{S} \to \mathbb{R} evaluates entering state s′s' while executing option ω\omega: U(ω,s′)=(1−βω,ϑ(s′))QΩ(s′,ω)+βω,ϑ(s′)VΩ(s′)U(\omega, s') = (1 - \beta_{\omega,\vartheta}(s')) Q_\Omega(s', \omega) + \beta_{\omega,\vartheta}(s') V_\Omega(s') where VΩ(s′)=∑ωˉ∈ΩπΩ(ωˉ∣s′)QΩ(s′,ωˉ)V_\Omega(s') = \sum_{\bar{\omega} \in \Omega} \pi_\Omega(\bar{\omega} \mid s') Q_\Omega(s', \bar{\omega}).

    2. The state-option-action value function QU:S×Ω×A→RQ_U: \mathcal{S} \times \Omega \times \mathcal{A} \to \mathbb{R} evaluates executing primitive action aa in state ss under active option ω\omega: QU(s,ω,a)=r(s,a)+γ∑s′∈SP(s′∣s,a)U(ω,s′)Q_U(s, \omega, a) = r(s, a) + \gamma \sum_{s' \in \mathcal{S}} P(s' \mid s, a) U(\omega, s')

    3. The option value function QΩ:S×Ω→RQ_\Omega: \mathcal{S} \times \Omega \to \mathbb{R} is the expectation over primitive actions chosen by the intra-option policy: QΩ(s,ω)=∑a∈Aπω,θ(a∣s)QU(s,ω,a)Q_\Omega(s, \omega) = \sum_{a \in \mathcal{A}} \pi_{\omega,\theta}(a \mid s) Q_U(s, \omega, a)

  4. Knowl 4 — The Option-Critic Architecture

    model/method

    The Option-Critic architecture is a two-timescale framework for learning hierarchical temporal abstractions end-to-end without subgoals, intrinsic rewards, or demonstrations. It splits learning into an actor component and a critic component:

    1. Actor Component: Consists of parameterized intra-option policies πω,θ(a∣s)\pi_{\omega,\theta}(a \mid s), parameterized termination functions βω,ϑ(s)\beta_{\omega,\vartheta}(s), and a policy over options πΩ(ω∣s)\pi_\Omega(\omega \mid s). Intra-option policies and termination functions are updated via stochastic gradient descent at a slow timescale using the intra-option policy gradient theorem and termination gradient theorem.

    2. Critic Component: Estimates the value functions QU(s,ω,a)Q_U(s, \omega, a) and AΩ(s,ω)=QΩ(s,ω)−VΩ(s)A_\Omega(s, \omega) = Q_\Omega(s, \omega) - V_\Omega(s) at a faster timescale using temporal difference methods (such as intra-option Q-learning).

    The execution operates via a call-and-return switch: an option ωt\omega_t persists across steps until a termination event drawn from βωt,ϑ(st+1)\beta_{\omega_t,\vartheta}(s_{t+1}) triggers the selection of a new option from πΩ(⋅∣st+1)\pi_\Omega(\cdot \mid s_{t+1}).

  5. Knowl 5 — Tabular Option-Critic with Intra-Option Q-Learning

    algorithm

    Tabular Option-Critic learns intra-option policy parameters θ\theta, termination parameters ϑ\vartheta, and critic values QU(s,ω,a)Q_U(s, \omega, a) online from single-step transition samples (s,a,r,s′)(s, a, r, s').

    Input: Discount factor γ∈[0,1)\gamma \in [0, 1), learning rates α\alpha (critic), αθ\alpha_\theta (policies), αϑ\alpha_\vartheta (terminations), exploration rate ϵ\epsilon
    Initialize critic QU(s,ω,a)Q_U(s, \omega, a), policy parameters θ\theta, and termination parameters ϑ\vartheta
    Initialize state s←s0s \leftarrow s_0
    Select initial option ω\omega according to ϵ\epsilon-soft policy πΩ(s)\pi_\Omega(s)
    repeat
        Select primitive action a∼πω,θ(⋅∣s)a \sim \pi_{\omega,\theta}(\cdot \mid s)
        Execute action aa, observe transition to next state s′s' and reward rr
        δ←r−QU(s,ω,a)\delta \leftarrow r - Q_U(s, \omega, a)
        if s′s' is non-terminal then
            δ←δ+γ(1−βω,ϑ(s′))QΩ(s′,ω)+γβω,ϑ(s′)max⁡ωˉQΩ(s′,ωˉ)\delta \leftarrow \delta + \gamma (1 - \beta_{\omega,\vartheta}(s')) Q_\Omega(s', \omega) + \gamma \beta_{\omega,\vartheta}(s') \max_{\bar{\omega}} Q_\Omega(s', \bar{\omega})
        end if
        QU(s,ω,a)←QU(s,ω,a)+αδQ_U(s, \omega, a) \leftarrow Q_U(s, \omega, a) + \alpha \delta
        θ←θ+αθ∂log⁡πω,θ(a∣s)∂θQU(s,ω,a)\theta \leftarrow \theta + \alpha_\theta \frac{\partial \log \pi_{\omega,\theta}(a \mid s)}{\partial \theta} Q_U(s, \omega, a)
        ϑ←ϑ−αϑ∂βω,ϑ(s′)∂ϑ(QΩ(s′,ω)−VΩ(s′))\vartheta \leftarrow \vartheta - \alpha_\vartheta \frac{\partial \beta_{\omega,\vartheta}(s')}{\partial \vartheta} (Q_\Omega(s', \omega) - V_\Omega(s'))
        if βω,ϑ(s′)\beta_{\omega,\vartheta}(s') triggers termination then
            Select new option ω\omega according to ϵ\epsilon-soft policy πΩ(s′)\pi_\Omega(s')
        end if
        s←s′s \leftarrow s'
    until s′s' is terminal
  6. Knowl 6 — Termination Advantage Regularization via Duration Margin

    model/method

    When optimizing options strictly for return, the termination gradient tends to shorten options toward single-step primitive actions because primitive actions are mathematically sufficient for optimal control in an MDP. To prevent option collapse and incentivize temporal duration, a constant margin ξ>0\xi > 0 is added to the advantage function within the termination update:

    AΩ(s,ω)+ξ=QΩ(s,ω)−VΩ(s)+ξA_\Omega(s, \omega) + \xi = Q_\Omega(s, \omega) - V_\Omega(s) + \xi

    This modification ensures that as long as the value of the currently executing option is within ξ\xi of the optimal option value VΩ(s)V_\Omega(s), the effective advantage remains strictly positive, driving the termination gradient to decrease βω,ϑ(s)\beta_{\omega,\vartheta}(s) and thereby extending the temporal duration of the option.

  7. Knowl 7 — Deep Option-Critic Architecture for High-Dimensional State Spaces

    model/method

    For high-dimensional observation spaces such as raw video frames, the option-critic architecture shares convolutional representation layers across all components:

    1. Network Architecture: A history of 4 stacked image frames passes through three convolutional layers (layer 1: 32 filters 8×88\times 8 with stride 4; layer 2: 64 filters 4×44\times 4 with stride 2; layer 3: 64 filters 3×33\times 3 with stride 1) followed by a shared dense layer of 512 units. The shared representation feeds four output heads: the policy over options πΩ(⋅∣s)\pi_\Omega(\cdot \mid s), termination probabilities {βω(s)}\{\beta_\omega(s)\} with sigmoid activations, intra-option action distributions {πω(⋅∣s)}\{\pi_\omega(\cdot \mid s)\} with linear-softmax activations, and option-value functions QΩ(s,⋅)Q_\Omega(s, \cdot).

    2. Value Estimation: Rather than maintaining a separate parameterization for QU(s,ω,a)Q_U(s, \omega, a), QUQ_U is estimated on-the-fly via the one-step sample target: gt(1)=rt+1+γ[(1−βωt,ϑ(st+1))QΩ(st+1,ωt)+βωt,ϑ(st+1)max⁡ωQΩ(st+1,ω)]g_t^{(1)} = r_{t+1} + \gamma \left[ (1 - \beta_{\omega_t,\vartheta}(s_{t+1})) Q_\Omega(s_{t+1}, \omega_t) + \beta_{\omega_t,\vartheta}(s_{t+1}) \max_\omega Q_\Omega(s_{t+1}, \omega) \right]

    3. Training & Stability: The critic network is trained off-policy using intra-option Q-learning with experience replay and RMSProp. Intra-option policy parameters and terminations are updated online. Intra-option policy variance and premature determinism are countered by subtracting the baseline QΩ(s,ω)Q_\Omega(s, \omega) and adding an entropy regularization penalty.

  8. Knowl 8 — Autonomous Skill Acquisition in Continuous Pinball Navigation

    empirical result

    In the continuous-state Pinball maze domain (continuous position and velocity in R4\mathbb{R}^4, 5 primitive actions, elastic polygon collisions, drag coefficient 0.995, step penalties −5-5 for thrust and −1-1 for null action, +10000+10000 goal reward), option-critic was evaluated with 2, 3, and 4 options using order-3 Fourier basis linear function approximation for the critic and linear-sigmoid/Boltzmann policies for the actor.

    Option-critic learned near-optimal navigation policies across all option counts within 40 episodes. Unlike prior skill discovery approaches that require explicit gestation periods (such as 10-episode delays before using discovered skills) or predefined subgoals, option-critic learned temporally extended options continuously from episode zero, exhibiting distinct trajectory specialization where specific learned options consistently handled navigation near the goal.

  9. Knowl 9 — Doorway Bottleneck Specialization and Rapid Transfer in Four-Rooms Domain

    empirical result

    In the discrete Four-Rooms gridworld environment (stochastic grid transitions failing with probability 1/31/3, goal reward +1+1, γ=0.99\gamma = 0.99), option-critic agents with 4 and 8 options (initialized to zero weights) were trained for 1000 episodes with an east doorway goal, after which the goal was relocated to a random location in the lower-right room.

    Across 350 independent runs, option-critic learned initial solutions from scratch at rates matching primitive actor-critic and SARSA(0), but adapted significantly faster upon sudden goal relocation. Without any heuristic subgoal discovery mechanisms or pre-specified subgoals, the learned option termination probabilities autonomously concentrated around doorway bottlenecks connecting the rooms.

  10. Knowl 10 — End-to-End Deep Option Learning on Atari 2600 Benchmarks

    empirical result

    In the Arcade Learning Environment (ALE), deep option-critic was trained end-to-end with 8 options on Asterix, Ms. Pacman, Seaquest, and Zaxxon using identical hyperparameters (learning rate 0.00025, ξ=0.01\xi = 0.01 termination margin, 0.010.01 entropy penalty, test ϵ=0.05\epsilon = 0.05).

    Deep option-critic learned successful behaviors from scratch in all four games within 200 epochs, surpassing the score of the baseline Deep Q-Network (DQN) trained on primitive actions in Asterix, Seaquest, and Zaxxon. In Seaquest experiments with 2 options, options automatically specialized into directional primitives (one option controlling upward movement toward the surface and the other controlling downward submarine navigation).

  11. Knowl 11 — Universal Option Availability Assumption and Ergodicity Constraints in Option Discovery

    limitation

    The option-critic theoretical framework assumes universal availability of options across the state space: ∀s∈S,∀ω∈Ω:s∈Iω\forall s \in \mathcal{S}, \forall \omega \in \Omega : s \in \mathcal{I}_\omega, where Iω\mathcal{I}_\omega is the initiation set of option ω\omega.

    Extending option-critic to learn restricted initiation sets Iω\mathcal{I}_\omega under function approximation requires maintaining ergodicity in the Markov chain over the augmented state-option space S×Ω\mathcal{S} \times \Omega. This requires enforcing an explicit flow balance constraint linking initiation sets and termination functions. Furthermore, under function approximation, evaluating whether an option is initiable via a feature classifier incurs a computational cost comparable to evaluating a policy over options, eliminating the computational savings that sparse initiation sets provide in tabular settings.

Coverage note — None omitted; all primary theoretical derivations (Theorems 1 and 2), architecture specifications, algorithmic procedures, regularization methods, empirical benchmark results, and stated theoretical limitations are fully covered.

References

  1. 1.Baird, L. C. 1993. Advantage updating. Technical Report WL–TR-93-1146, Wright Laboratory.
  2. 2.Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research 47:253–279.
  3. 3.Comanici, G., and Precup, D. 2010. Optimal policy switching algorithms for reinforcement learning. In AAMAS, 709–714.
  4. 4.S¸ims¸ek, O., and Barto, A. G. 2009. Skill characterization based on betweenness. In NIPS 21, 1497–1504.
  5. 5.Daniel, C.; van Hoof, H.; Peters, J.; and Neumann, G. 2016. Probabilistic inference for determining options in reinforcement learning. Machine Learning, Special Issue 104(2):337–357.
  6. 6.Harb, J. 2016. Learning options in deep reinforcement learning. Master’s thesis, McGill University.
  7. 7.Konda, V. R., and Tsitsiklis, J. N. 2000. Actor-critic algorithms. In NIPS 12, 1008–1014.
  8. 8.Konidaris, G., and Barto, A. 2009. Skill discovery in continuous reinforcement learning domains using skill chaining. In NIPS 22, 1015–1023.
  9. 9.Konidaris, G.; Kuindersma, S.; Grupen, R. A.; and Barto, A. G. 2011. Autonomous skill acquisition on a mobile manipulator. In AAAI.
  10. 10.Krishnamurthy, R.; Lakshminarayanan, A. S.; Kumar, P.; and Ravindran, B. 2016. Hierarchical reinforcement learning using spatio-temporal abstractions and deep neural networks. CoRR abs/1605.05359.
  11. 11.Kulkarni, T.; Narasimhan, K.; Saeedi, A.; and Tenenbaum, J. 2016. Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation. In NIPS 29.
  12. 12.Levy, K. Y., and Shimkin, N. 2011. Unified inter and intra options learning using policy gradient methods. In EWRL, 153–164.
  13. 13.Mankowitz, D. J.; Mann, T. A.; and Mannor, S. 2016. Adaptive skills, adaptive partitions (ASAP). In NIPS 29.
  14. 14.Mann, T. A.; Mankowitz, D. J.; and Mannor, S. 2014. Time-regularized interrupting options (TRIO). In ICML, 1350–1358.
  15. 15.Mann, T. A.; Mannor, S.; and Precup, D. 2015. Approximate value iteration with temporally extended actions. Journal of Artificial Intelligence Research 53:375–438.
  16. 16.McGovern, A., and Barto, A. G. 2001. Automatic discovery of subgoals in reinforcement learning using diverse density. In ICML, 361–368.
  17. 17.Menache, I.; Mannor, S.; and Shimkin, N. 2002. Q-cut - dynamic discovery of sub-goals in reinforcement learning. In ECML, 295–306.
  18. 18.Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. A. 2013. Playing atari with deep reinforcement learning. CoRR abs/1312.5602.
  19. 19.Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T. P.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016. Asynchronous methods for deep reinforcement learning. In ICML.
  20. 20.Niekum, S. 2013. Semantically Grounded Learning from Unstructured Demonstrations. Ph.D. Dissertation, University of Massachusetts, Amherst.
  21. 21.Precup, D. 2000. Temporal abstraction in reinforcement learning. Ph.D. Dissertation, University of Massachusetts, Amherst.
  22. 22.Puterman, M. L. 1994. Markov Decision Processes: Discrete Stochastic Dynamic Programming. John Wiley & Sons, Inc.
  23. 23.Silver, D., and Ciosek, K. 2012. Compositional planning using optimal option models. In ICML.
  24. 24.Sorg, J., and Singh, S. P. 2010. Linear options. In AAMAS, 31–38.
  25. 25.Stolle, M., and Precup, D. 2002. Learning options in reinforcement learning. In Abstraction, Reformulation and Approximation, 5th International Symposium, SARA Proceedings, 212–223.
  26. 26.Sutton, R. S.; McAllester, D. A.; Singh, S. P.; and Mansour, Y. 2000. Policy gradient methods for reinforcement learning with function approximation. In NIPS 12. 1057–1063.
  27. 27.Sutton, R. S.; Precup, D.; and Singh, S. P. 1999. Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence 112(1-2):181–211.
  28. 28.Sutton, R. S. 1984. Temporal Credit Assignment in Reinforcement Learning. Ph.D. Dissertation.
  29. 29.Thomas, P. 2014. Bias in natural actor-critic algorithms. In ICML, 441–448.
  30. 30.Vezhnevets, A. S.; Mnih, V.; Agapiou, J.; Osindero, S.; Graves, A.; Vinyals, O.; and Kavukcuoglu, K. 2016. Strategic attentive writer for learning macro-actions. In NIPS 29.

Citation

MLA
Bacon, P.-L., et al. “The Option-Critic Architecture”. arXiv, 2016, http://arxiv.org/abs/1609.05140v2.
APA
Bacon, P.-L., Harb, J., & Precup, D. (2016). The Option-Critic Architecture. arXiv. http://arxiv.org/abs/1609.05140v2
Chicago
Bacon, P.-L., J. Harb, and D. Precup. 2016. “The Option-Critic Architecture”. arXiv. http://arxiv.org/abs/1609.05140v2.
Harvard
Bacon, P.-L., Harb, J. and Precup, D. (2016) “The Option-Critic Architecture”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1609.05140v2.
Vancouver
1. Bacon P-L, Harb J, Precup D (2016) The Option-Critic Architecture. arXiv

BibTeX

@article{bacon2016the,
  title = {The Option-Critic Architecture},
  author = {Bacon, Pierre-Luc and Harb, Jean and Precup, Doina},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1609.05140v2},
  eprint = {1609.05140}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF