Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation

Tejas D. KulkarniKarthik R. NarasimhanArdavan SaeediJoshua B. Tenenbaum

article2016NeurIPS1,319 citations

Proposes hierarchical-DQN, a framework that integrates multi-timescale value functions with intrinsic motivation to solve sparse-reward reinforcement learning problems like Montezuma's Revenge by decoupling goal selection from action execution.

Listen

Reinforcement learning systems often struggle in environments where meaningful feedback is delayed or sparse. Standard exploration techniques rely on random, low-level trial and error, which makes it nearly impossible for an automated agent to discover complex sequences of actions needed to achieve long-term objectives. The article introduces hierarchical Deep Q-Networks (h-DQN) to evaluate whether integrating hierarchical value functions operating over different time scales with intrinsic motivation can enable efficient exploration and long-range planning.

The framework divides decision-making into a two-level hierarchy. A top-level module (the meta-controller) selects intrinsic goals to maximize long-term external rewards from the environment, while a lower-level module (the controller) generates primitive actions to accomplish the chosen goal and receives intrinsic rewards from an internal critic. The authors tested this system on two challenging domains with delayed rewards: a six-state stochastic decision process and the complex video game 'Montezuma's Revenge', where goals were parameterized using objects and relational entities detected in visual scenes.

The empirical findings demonstrate significant performance gains over traditional reinforcement learning approaches. In the stochastic environment, the hierarchical method learned the optimal exploratory strategy to achieve an average reward of about 0.13, whereas standard baseline methods converged to a suboptimal reward of 0.01. In 'Montezuma's Revenge', where standard deep reinforcement learning models score 0 and state-of-the-art parallel approaches score only around 4.16, the proposed framework consistently achieved scores of approximately 400 per episode. The system successfully discovered critical intermediate sub-goals, such as retrieving keys and navigating ladders, progressing systematically from simpler milestones to more complex objectives.

These results demonstrate that temporal abstraction and goal-driven intrinsic exploration can dramatically reduce the sample complexity and exploration bottlenecks of artificial intelligence systems in sparse environments. By organizing exploration around entities and intermediate objectives rather than raw actions, artificial intelligence can navigate complex, multi-stage workflows that were previously intractable. Organizations seeking to deploy autonomous agents in sparse-feedback environments should explore hierarchical goal-setting architectures to improve sample efficiency and task success.

To scale these capabilities further, future development should focus on integrating automatic, unsupervised object discovery directly from raw sensory inputs and incorporating flexible short-term memory to handle long-range dependencies and non-Markovian settings. Decision-makers should note that current implementations rely on custom, pre-defined object detectors to specify the candidate goal space, meaning broad generalization to entirely unstructured environments remains an ongoing area of research.

  • Paper: Exploration by Random Network Distillation, Yuri Burda et al. (2019). Building on the challenge of sparse-feedback benchmarks like Montezuma's Revenge explored by h-DQN, this work introduces Random Network Distillation to scale intrinsic novelty rewards in complex deep RL environments.
  • Paper: Curiosity-Driven Exploration by Self-Supervised Prediction, Deepak Pathak et al. (2017). This paper advances intrinsic motivation by using self-supervised prediction error as curiosity rewards, offering an unsupervised alternative to explicit multi-level goal specifications.
  • Paper: The Option-Critic Architecture, Pierre-Luc Bacon et al. (2016). This work extends temporal abstraction by learning options and their termination conditions end-to-end via policy gradients, removing the need to manually define intrinsic goal spaces.
  • Paper: Hindsight Experience Replay, Marcin Andrychowicz et al. (2017). This work addresses goal-conditioned reinforcement learning in sparse-reward settings by replaying trajectories with hindsight goals, providing a complementary approach to hierarchical exploration.
  • Paper: Diversity is All You Need: Learning Skills without a Reward Function, Benjamin Eysenbach et al. (2018). This paper generalizes intrinsic motivation by learning task-agnostic, diverse repertoires of skills entirely without external reward functions for downstream hierarchical transfer.
Cover for Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation

Abstract

Learning goal-directed behavior in environments with sparse feedback is a major challenge for reinforcement learning algorithms. The primary difficulty arises due to insufficient exploration, resulting in an agent being unable to learn robust value functions. Intrinsically motivated agents can explore new behavior for its own sake rather than to directly solve problems. Such intrinsic behaviors could eventually help the agent solve tasks posed by the environment. We present hierarchical-DQN (h-DQN), a framework to integrate hierarchical value functions, operating at different temporal scales, with intrinsically motivated deep reinforcement learning. A top-level value function learns a policy over intrinsic goals, and a lower-level function learns a policy over atomic actions to satisfy the given goals. h-DQN allows for flexible goal specifications, such as functions over entities and relations. This provides an efficient space for exploration in complicated environments. We demonstrate the strength of our approach on two problems with very sparse, delayed feedback: (1) a complex discrete stochastic decision process, and (2) the classic ATARI game `Montezuma's Revenge'.

Table of Contents

  • 1 Introduction
  • 2 Literature Review
  • 2.1 Reinforcement Learning with Temporal Abstractions
  • 2.2 Intrinsically motivated RL
  • 2.3 Object-based RL
  • 2.4 Deep Reinforcement Learning
  • 2.5 Cognitive Science and Neuroscience
  • 3 Model
  • 4 Experiments
  • 4.1 Discrete stochastic decision process
  • 4.2 ATARI game with delayed rewards
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Hierarchical Deep Q-Network (h-DQN) Framework

    model/method

    The hierarchical Deep Q-Network (h-DQN) framework integrates temporal abstraction with intrinsic motivation to solve reinforcement learning tasks characterized by sparse, delayed feedback. The architecture operates over two levels of hierarchy functioning at different temporal scales:

    1. Meta-Controller: Operates at a coarse temporal scale. Given the environment state st∈Ss_t \in \mathcal{S}, it selects a discrete intrinsic goal gt∈Gg_t \in \mathcal{G} using an action-value function Q2(st,gt;θ2)Q_2(s_t, g_t; \theta_2). The meta-controller's objective is to maximize the cumulative discounted extrinsic reward Ft=∑t′=t∞γt′−tft′F_t = \sum_{t'=t}^{\infty} \gamma^{t'-t} f_{t'} received from the external environment.

    2. Controller: Operates at a single-step temporal scale. Conditioned on the state sts_t and the selected goal gtg_t, it selects a primitive action at∈Aa_t \in \mathcal{A} using an action-value function Q1(st,at;θ1,gt)Q_1(s_t, a_t; \theta_1, g_t). The controller's objective is to maximize cumulative discounted intrinsic rewards Rt(g)=∑t′=t∞γt′−trt′(g)R_t(g) = \sum_{t'=t}^{\infty} \gamma^{t'-t} r_{t'}(g).

    3. Internal Critic: A module that evaluates the agent's progress towards goal gg and provides the intrinsic reward signal rt(g)r_t(g) to the controller. The controller persists in pursuing gtg_t for NN steps until gtg_t is achieved or a terminal state is reached, at which point the meta-controller observes next state st+Ns_{t+N} and selects a new goal.

  2. Knowl 2 — Bellman Equations and Optimization Objectives in h-DQN

    equation

    In the h-DQN framework, the controller and meta-controller optimize distinct Bellman equations over different temporal intervals and reward structures.

    For the controller, the optimal action-value function Q1∗(s,a;g)Q_1^*(s, a; g) under a policy πag=P(a∣s,g)\pi_{ag} = P(a \mid s, g) satisfies: Q1∗(s,a;g)=max⁡πagE[rt+γmax⁡at+1Q1∗(st+1,at+1;g) ∣ st=s,at=a,gt=g,πag]Q_1^*(s, a; g) = \max_{\pi_{ag}} \mathbb{E}\left[r_t + \gamma \max_{a_{t+1}} Q_1^*(s_{t+1}, a_{t+1}; g) \,\Big|\, s_t = s, a_t = a, g_t = g, \pi_{ag}\right] where rtr_t is the intrinsic reward provided by the internal critic for transitioning toward goal gg, and γ∈[0,1)\gamma \in [0, 1) is the discount factor.

    For the meta-controller, which selects goals for durations of NN steps under goal policy πg=P(g∣s)\pi_g = P(g \mid s), the semi-Markov action-value function Q2∗(s,g)Q_2^*(s, g) satisfies: Q2∗(s,g)=max⁡πgE[∑t′=tt+Nft′+γmax⁡g′Q2∗(st+N,g′) ∣ st=s,gt=g,πg]Q_2^*(s, g) = \max_{\pi_g} \mathbb{E}\left[\sum_{t'=t}^{t+N} f_{t'} + \gamma \max_{g'} Q_2^*(s_{t+N}, g') \,\Big|\, s_t = s, g_t = g, \pi_g\right] where ft′f_{t'} represents the extrinsic environment reward at step t′t', and NN is the number of steps until the controller terminates the sub-policy for goal gg.

    The parameterized networks θ1\theta_1 and θ2\theta_2 are trained by gradient descent on the mean squared temporal-difference errors with target network parameters θ1,i−1\theta_{1,i-1} and θ2,i−1\theta_{2,i-1}: L1(θ1,i)=E(s,a,g,r,s′)∼D1[(r+γmax⁡a′Q1(s′,a′;θ1,i−1,g)−Q1(s,a;θ1,i,g))2]L_1(\theta_{1,i}) = \mathbb{E}_{(s, a, g, r, s') \sim \mathcal{D}_1}\left[\left(r + \gamma \max_{a'} Q_1(s', a'; \theta_{1,i-1}, g) - Q_1(s, a; \theta_{1,i}, g)\right)^2\right] L2(θ2,i)=E(s,g,F,s′)∼D2[(F+γmax⁡g′Q2(s′,g′;θ2,i−1)−Q2(s,g;θ2,i))2]L_2(\theta_{2,i}) = \mathbb{E}_{(s, g, F, s') \sim \mathcal{D}_2}\left[\left(F + \gamma \max_{g'} Q_2(s', g'; \theta_{2,i-1}) - Q_2(s, g; \theta_{2,i})\right)^2\right] where D1\mathcal{D}_1 and D2\mathcal{D}_2 are disjoint experience replay memories storing atomic transitions ({s,g},a,r,{s′,g})(\{s, g\}, a, r, \{s', g\}) and multi-step meta-transitions (st,gt,Ft,st+N)(s_t, g_t, F_t, s_{t+N}), respectively.

  3. Knowl 3 — h-DQN Training with Dual Replay and Adaptive Exploration

    algorithm

    The h-DQN algorithm simultaneously trains the controller over primitive actions and the meta-controller over intrinsic goals using two distinct replay buffers and decoupled time scales.

    Input: Set of primitive actions A\mathcal{A}, set of goals G\mathcal{G}, discount factor γ\gamma, learning rates, replay buffer capacities
    Output: Controller parameters θ1\theta_1 and meta-controller parameters θ2\theta_2
    Initialize experience replay memories D1\mathcal{D}_1 and D2\mathcal{D}_2
    Initialize network parameters θ1\theta_1 and θ2\theta_2
    Initialize meta-controller exploration probability ϵ2←1.0\epsilon_2 \leftarrow 1.0
    Initialize controller exploration probabilities ϵ1,g←1.0\epsilon_{1,g} \leftarrow 1.0 for all g∈Gg \in \mathcal{G}
    for episode = 1 to num_episodes do
        Initialize environment and obtain initial state ss
        Sample goal g∼epsGreedy(s,G,ϵ2,Q2)g \sim \text{epsGreedy}(s, \mathcal{G}, \epsilon_2, Q_2)
        while ss is not terminal do
            F←0F \leftarrow 0
            s0←ss_0 \leftarrow s
            while not (ss is terminal or goal gg is reached) do
                Sample action a∼epsGreedy({s,g},A,ϵ1,g,Q1)a \sim \text{epsGreedy}(\{s, g\}, \mathcal{A}, \epsilon_{1,g}, Q_1)
                Execute aa, observe next state s′s' and extrinsic reward ff
                Obtain intrinsic reward r←critic(s,a,s′,g)r \leftarrow \text{critic}(s, a, s', g)
                Store transition ({s,g},a,r,{s′,g})(\{s, g\}, a, r, \{s', g\}) in D1\mathcal{D}_1
                Sample mini-batch from D1\mathcal{D}_1 and perform gradient descent on L1(θ1)L_1(\theta_1)
                Sample mini-batch from D2\mathcal{D}_2 and perform gradient descent on L2(θ2)L_2(\theta_2)
                F←F+fF \leftarrow F + f
                s←s′s \leftarrow s'
            Store multi-step transition (s0,g,F,s′)(s_0, g, F, s') in D2\mathcal{D}_2
            if ss is not terminal then
                Sample goal g∼epsGreedy(s,G,ϵ2,Q2)g \sim \text{epsGreedy}(s, \mathcal{G}, \epsilon_2, Q_2)
        Anneal ϵ2\epsilon_2 and adaptively anneal each ϵ1,g\epsilon_{1,g} based on empirical success rate of reaching gg

    The exploration probability ϵ2\epsilon_2 of the meta-controller is annealed gradually from 1.01.0, while ϵ1,g\epsilon_{1,g} for each goal gg is dynamically annealed based on the empirical success rate of the controller in reaching gg.

  4. Knowl 4 — Entity-Relation Parameterization for Intrinsic Goal Spaces

    model/method

    In complex visual reinforcement learning domains, structuring the intrinsic goal space around detected objects and relational interactions constrains exploration to semantically relevant state spaces.

    1. Object Detection: A vision front-end provides bounding boxes or candidate spatial locations of entities present in the scene (such as key, doors, ladders, or enemies).
    2. Relational Internal Critic: The internal critic evaluates transitions over structured relational tuples of the form ⟨entity1,relation,entity2⟩\langle \text{entity}_1, \text{relation}, \text{entity}_2 \rangle, where entity1\text{entity}_1 is the controllable agent, entity2\text{entity}_2 is a candidate object chosen by the meta-controller, and relation\text{relation} is a spatial condition (such as reaching or touching entity2\text{entity}_2).
    3. Goal Specification: When the meta-controller chooses a target entity g=entity2g = \text{entity}_2, the internal critic yields an intrinsic reward r(s,a,s′)=1r(s, a, s') = 1 when the agent satisfies the relational predicate with the target entity, and 00 otherwise.
  5. Knowl 5 — Hierarchical Exploration on Stochastic Chain MDP

    empirical result

    On a 6-state discrete stochastic chain Markov Decision Process with states {s1,s2,s3,s4,s5,s6}\{s_1, s_2, s_3, s_4, s_5, s_6\}, the agent starts at s2s_2 with s1s_1 being the terminal state. Choosing the action 'left' moves left deterministically, whereas choosing 'right' succeeds with probability 0.50.5 (and moves left with probability 0.50.5). The extrinsic reward at s1s_1 is 1.01.0 if state s6s_6 was visited earlier in the trajectory, and 0.010.01 if s1s_1 is reached without having visited s6s_6.

    • Standard Q-learning Baseline: With ϵ\epsilon-greedy exploration annealed from 11 to 0.10.1 over 50,000 steps, standard flat Q-learning converges to the suboptimal policy of navigating immediately from s2s_2 to s1s_1, yielding an average extrinsic reward of 0.010.01.
    • Hierarchical Q-learning (h-DQN framework): Defining each state {s1,…,s6}\{s_1, \dots, s_6\} as a possible goal and providing positive intrinsic rewards when reaching the selected goal state enables the meta-controller to learn to select intermediate goals s4,s5,s_4, s_5, and s6s_6. This leads the agent to regularly visit s6s_6 before terminating at s1s_1, achieving an average extrinsic reward of approximately 0.130.13 over 10 random runs.
  6. Knowl 6 — Neural Network Architecture and Two-Phase Training on Montezuma's Revenge

    experimental setup

    For the visual domain of Montezuma's Revenge on the Arcade Learning Environment:

    1. Input Representation:

      • Meta-Controller: Four consecutive grayscale frames resized to 84×8484 \times 84 pixels.
      • Controller: The four consecutive 84×8484 \times 84 frames concatenated with a binary spatial mask channel corresponding to the chosen goal object's location in image coordinates (5 channels total).
    2. Network Architecture:

      • Three convolutional layers: Conv1 (32 filters of size 8×88 \times 8, stride 4, ReLU), Conv2 (64 filters of size 4×44 \times 4, stride 2, ReLU), Conv3 (64 filters of size 3×33 \times 3, stride 1, ReLU).
      • Fully connected layer of 512 ReLU units followed by a linear output layer producing Q-values (∣A∣|\mathcal{A}| outputs for the controller Q1Q_1, and ∣G∣|\mathcal{G}| outputs for the meta-controller Q2Q_2).
    3. Hyperparameters and Two-Phase Protocol:

      • Optimizer: RMSprop with learning rate 2.5×10−42.5 \times 10^{-4} and discount factor γ=0.99\gamma = 0.99.
      • Replay buffer capacities: ∣D1∣=106|\mathcal{D}_1| = 10^6 transitions for the controller; ∣D2∣=5×104|\mathcal{D}_2| = 5 \times 10^4 transitions for the meta-controller.
      • Phase 1 (Pre-training): ϵ2\epsilon_2 is fixed to 1.01.0 for approximately 2.3×1062.3 \times 10^6 steps so the meta-controller samples goals uniformly at random while the controller learns low-level sub-policies to reach accessible entities.
      • Phase 2 (Joint Training): Controller and meta-controller are jointly trained for an additional 2.0×1062.0 \times 10^6 steps with ϵ2\epsilon_2 annealed.
  7. Knowl 7 — Performance on Montezuma's Revenge with Sparse Delayed Rewards

    empirical result

    In the ATARI game Montezuma's Revenge, which requires navigating ladders, avoiding obstacles (a skull), acquiring a key (+100+100 reward), and unlocking a door (+300+300 reward):

    • Baseline DQN: Achieves an average score of 00 because random primitive exploration cannot discover the multi-step action sequence needed to collect the key.
    • Gorila DQN (Massively Parallel DQN): Achieves an average score of only 4.164.16.
    • h-DQN: Following the two-phase training protocol (2.3×1062.3 \times 10^6 pre-training steps followed by 2.0×1062.0 \times 10^6 joint training steps), h-DQN consistently reaches a score of approximately +400+400 per episode by learning to navigate to the key and subsequently opening the door.

    During training progression, the meta-controller first masters choosing easier proximal goals (such as the middle ladder and right door) before increasing its selection and success rate for harder distant goals (such as the key and bottom ladders).

  8. Knowl 8 — Limitations of the h-DQN Architecture

    limitation

    The h-DQN framework has three primary stated limitations:

    1. Manual / Custom Object Proposal: Entity and goal proposals rely on a customized object detector rather than end-to-end unsupervised visual representation learning or automated object discovery from raw video frames.
    2. Fixed Option Termination: The low-level controller executes until the target goal is reached or the episode terminates, lacking dynamic, intermittent termination mechanisms to abort suboptimal sub-policies mid-execution.
    3. Lack of Episodic Memory: Because the policy relies on feed-forward convolutions over 4 stacked frames, it lacks an explicit short-term or episodic memory module to track history across long sequences of sub-goals in non-Markovian environments.

Coverage note — None was omitted; all core contributions, theoretical formulations, algorithms, experimental setups, empirical results, and stated limitations from the paper are represented.

References

  1. 1.A. G. Barto and S. Mahadevan. Recent advances in hierarchical reinforcement learning. Discrete Event Dynamic Systems, 13(4):341–379, 2003.
  2. 2.M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 2012.
  3. 3.M. M. Botvinick, Y. Niv, and A. C. Barto. Hierarchically organized behavior and its neural foundations: A reinforcement learning perspective. Cognition, 113(3):262–280, 2009.
  4. 4.L. C. Cobo, C. L. Isbell, and A. L. Thomaz. Object focused q-learning for autonomous agents. In Proceedings of the 2013 international conference on Autonomous agents and multi-agent systems, pages 1061–1068. International Foundation for Autonomous Agents and Multiagent Systems, 2013.
  5. 5.P. Dayan. Improving generalization for temporal difference learning: The successor representation. Neural Computation, 5(4):613–624, 1993.
  6. 6.P. Dayan and G. E. Hinton. Feudal reinforcement learning. In Advances in neural information processing systems, pages 271–271. Morgan Kaufmann Publishers, 1993.
  7. 7.T. G. Dietterich. Hierarchical reinforcement learning with the maxq value function decomposition. J. Artif. Intell. Res.(JAIR), 13:227–303, 2000.
  8. 8.C. Diuk, A. Cohen, and M. L. Littman. An object-oriented representation for efficient reinforcement learning. In Proceedings of the 25th international conference on Machine learning, pages 240–247. ACM, 2008.
  9. 9.S. Eslami, N. Heess, T. Weber, Y. Tassa, K. Kavukcuoglu, and G. E. Hinton. Attend, infer, repeat: Fast scene understanding with generative models. arXiv preprint arXiv:1603.08575, 2016.
  10. 10.K. Fragkiadaki, P. Arbelaez, P. Felsen, and J. Malik. Learning to segment moving objects in videos. In Computer Vision and Pattern Recognition (CVPR), 2015 IEEE Conference on, pages 4083–4090. IEEE, 2015.
  11. 11.M. Frank, J. Leitner, M. Stollenga, A. Förster, and J. Schmidhuber. Curiosity driven reinforcement learning for motion planning on humanoids. Intrinsic motivations and open-ended development in animals, humans, and robots, page 245, 2015.
  12. 12.S. J. Gershman, C. D. Moore, M. T. Todd, K. A. Norman, and P. B. Sederberg. The successor representation and temporal context. Neural Computation, 24(6):1553–1568, 2012.
  13. 13.S. Goel and M. Huber. Subgoal discovery for hierarchical reinforcement learning using learned policies. In FLAIRS conference, pages 346–350, 2003.
  14. 14.K. Greff, R. K. Srivastava, and J. Schmidhuber. Binding via reconstruction clustering. arXiv preprint arXiv:1511.06418, 2015.
  15. 15.K. Gregor, I. Danihelka, A. Graves, and D. Wierstra. Draw: A recurrent neural network for image generation. arXiv preprint arXiv:1502.04623, 2015.
  16. 16.C. Guestrin, D. Koller, C. Gearhart, and N. Kanodia. Generalizing plans to new environments in relational mdps. In Proceedings of the 18th international joint conference on Artificial intelligence, pages 1003–1010. Morgan Kaufmann Publishers Inc., 2003.
  17. 17.C. Guestrin, D. Koller, R. Parr, and S. Venkataraman. Efficient solution algorithms for factored mdps. Journal of Artificial Intelligence Research, pages 399–468, 2003.
  18. 18.M. Hausknecht and P. Stone. Deep recurrent q-learning for partially observable mdps. arXiv preprint arXiv:1507.06527, 2015.
  19. 19.N. Hernandez-Gardiol and S. Mahadevan. Hierarchical memory-based reinforcement learning. Advances in Neural Information Processing Systems, pages 1047–1053, 2001.
  20. 20.J. Huang and K. Murphy. Efficient inference in occlusion-aware generative models of images. arXiv preprint arXiv:1511.06362, 2015.
  21. 21.J. Koutník, J. Schmidhuber, and F. Gomez. Evolving deep unsupervised convolutional networks for vision-based reinforcement learning. In Proceedings of the 2014 conference on Genetic and evolutionary computation, pages 541–548. ACM, 2014.
  22. 22.T. D. Kulkarni, W. F. Whitney, P. Kohli, and J. Tenenbaum. Deep convolutional inverse graphics network. In Advances in Neural Information Processing Systems, pages 2530–2538, 2015.
  23. 23.B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. Building machines that learn and think like people. arXiv preprint arXiv:1604.00289, 2016.
  24. 24.M. C. Machado and M. Bowling. Learning purposeful behaviour in the absence of rewards. arXiv preprint arXiv:1605.07700, 2016.
  25. 25.S. Mannor, I. Menache, A. Hoze, and U. Klein. Dynamic abstraction in reinforcement learning via clustering. In Proceedings of the twenty-first international conference on Machine learning, page 71. ACM, 2004.
  26. 26.A. McGovern and A. G. Barto. Automatic discovery of subgoals in reinforcement learning using diverse density. Computer Science Department Faculty Publication Series, page 8, 2001.
  27. 27.I. Menache, S. Mannor, and N. Shimkin. Q-cutdynamic discovery of sub-goals in reinforcement learning. In Machine Learning: ECML 2002, pages 295–306. Springer, 2002.
  28. 28.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015.
  29. 29.S. Mohamed and D. J. Rezende. Variational information maximisation for intrinsically motivated reinforcement learning. In Advances in Neural Information Processing Systems, pages 2116–2124, 2015.
  30. 30.A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, et al. Massively parallel methods for deep reinforcement learning. arXiv preprint arXiv:1507.04296, 2015.
  31. 31.K. Narasimhan, T. Kulkarni, and R. Barzilay. Language understanding for text-based games using deep reinforcement learning. arXiv preprint arXiv:1506.08941, 2015.
  32. 32.I. Osband, C. Blundell, A. Pritzel, and B. Van Roy. Deep exploration via bootstrapped dqn. arXiv preprint arXiv:1602.04621, 2016.
  33. 33.D. J. Rezende, S. Mohamed, I. Danihelka, K. Gregor, and D. Wierstra. One-shot generalization in deep generative models. arXiv preprint arXiv:1603.05106, 2016.
  34. 34.T. Schaul, D. Horgan, K. Gregor, and D. Silver. Universal value function approximators. In Proceedings of the 32nd International Conference on Machine Learning (ICML-15), pages 1312–1320, 2015.
  35. 35.T. Schaul, J. Quan, I. Antonoglou, and D. Silver. Prioritized experience replay. arXiv preprint arXiv:1511.05952, 2015.
  36. 36.J. Schmidhuber. Formal theory of creativity, fun, and intrinsic motivation (1990–2010). Autonomous Mental Development, IEEE Transactions on, 2(3):230–247, 2010.
  37. 37.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. Mastering the game of go with deep neural networks and tree search. Nature, 529(7587):484–489, 2016.
  38. 38.Ö. Şimşek, A. Wolfe, and A. Barto. Identifying useful subgoals in reinforcement learning by local graph partitioning. In Proceedings of the International conference on Machine learning, pages 816–823, 2005.
  39. 39.S. Singh, R. L. Lewis, and A. G. Barto. Where do rewards come from. In Proceedings of the annual conference of the cognitive science society, pages 2601–2606, 2009.
  40. 40.S. Singh, R. L. Lewis, A. G. Barto, and J. Sorg. Intrinsically motivated reinforcement learning: An evolutionary perspective. Autonomous Mental Development, IEEE Transactions on, 2(2):70–82, 2010.
  41. 41.S. P. Singh, A. G. Barto, and N. Chentanez. Intrinsically motivated reinforcement learning. In Advances in neural information processing systems, pages 1281–1288, 2004.
  42. 42.J. Sorg and S. Singh. Linear options. In Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems: Volume 1 - Volume 1, AAMAS ’10, pages 31–38, Richland, SC, 2010. International Foundation for Autonomous Agents and Multiagent Systems.
  43. 43.E. S. Spelke and K. D. Kinzler. Core knowledge. Developmental science, 10(1):89–96, 2007.
  44. 44.K. L. Stachenfeld, M. Botvinick, and S. J. Gershman. Design principles of the hippocampal cognitive map. In Advances in neural information processing systems, pages 2528–2536, 2014.
  45. 45.B. C. Stadie, S. Levine, and P. Abbeel. Incentivizing exploration in reinforcement learning with deep predictive models. arXiv preprint arXiv:1507.00814, 2015.
  46. 46.R. S. Sutton and A. G. Barto. Introduction to reinforcement learning, volume 135. MIT Press Cambridge, 1998.
  47. 47.R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup. Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2, pages 761–768. International Foundation for Autonomous Agents and Multiagent Systems, 2011.
  48. 48.R. S. Sutton, D. Precup, and S. Singh. Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning. Artificial intelligence, 112(1):181–211, 1999.
  49. 49.C. Szepesvari, R. S. Sutton, J. Modayil, S. Bhatnagar, et al. Universal option models. In Advances in Neural Information Processing Systems, pages 990–998, 2014.
  50. 50.W. F. Whitney, M. Chang, T. Kulkarni, and J. B. Tenenbaum. Understanding visual concepts with continuation learning. arXiv preprint arXiv:1602.06822, 2016.

Citation

MLA
Kulkarni, T. D., et al. “Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation”. arXiv, 2016, http://arxiv.org/abs/1604.06057v2.
APA
Kulkarni, T. D., Narasimhan, K. R., Saeedi, A., & Tenenbaum, J. B. (2016). Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation. arXiv. http://arxiv.org/abs/1604.06057v2
Chicago
Kulkarni, T. D., K. R. Narasimhan, A. Saeedi, and J. B. Tenenbaum. 2016. “Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation”. arXiv. http://arxiv.org/abs/1604.06057v2.
Harvard
Kulkarni, T.D. et al. (2016) “Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1604.06057v2.
Vancouver
1. Kulkarni TD, Narasimhan KR, Saeedi A, Tenenbaum JB (2016) Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation. arXiv

BibTeX

@article{kulkarni2016hierarchical,
  title = {Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation},
  author = {Kulkarni, Tejas D. and Narasimhan, Karthik R. and Saeedi, Ardavan and Tenenbaum, Joshua B.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1604.06057v2},
  eprint = {1604.06057}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission