Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
Tejas D. KulkarniKarthik R. NarasimhanArdavan SaeediJoshua B. Tenenbaum
Proposes hierarchical-DQN, a framework that integrates multi-timescale value functions with intrinsic motivation to solve sparse-reward reinforcement learning problems like Montezuma's Revenge by decoupling goal selection from action execution.
Reinforcement learning systems often struggle in environments where meaningful feedback is delayed or sparse. Standard exploration techniques rely on random, low-level trial and error, which makes it nearly impossible for an automated agent to discover complex sequences of actions needed to achieve long-term objectives. The article introduces hierarchical Deep Q-Networks (h-DQN) to evaluate whether integrating hierarchical value functions operating over different time scales with intrinsic motivation can enable efficient exploration and long-range planning.
The framework divides decision-making into a two-level hierarchy. A top-level module (the meta-controller) selects intrinsic goals to maximize long-term external rewards from the environment, while a lower-level module (the controller) generates primitive actions to accomplish the chosen goal and receives intrinsic rewards from an internal critic. The authors tested this system on two challenging domains with delayed rewards: a six-state stochastic decision process and the complex video game 'Montezuma's Revenge', where goals were parameterized using objects and relational entities detected in visual scenes.
The empirical findings demonstrate significant performance gains over traditional reinforcement learning approaches. In the stochastic environment, the hierarchical method learned the optimal exploratory strategy to achieve an average reward of about 0.13, whereas standard baseline methods converged to a suboptimal reward of 0.01. In 'Montezuma's Revenge', where standard deep reinforcement learning models score 0 and state-of-the-art parallel approaches score only around 4.16, the proposed framework consistently achieved scores of approximately 400 per episode. The system successfully discovered critical intermediate sub-goals, such as retrieving keys and navigating ladders, progressing systematically from simpler milestones to more complex objectives.
These results demonstrate that temporal abstraction and goal-driven intrinsic exploration can dramatically reduce the sample complexity and exploration bottlenecks of artificial intelligence systems in sparse environments. By organizing exploration around entities and intermediate objectives rather than raw actions, artificial intelligence can navigate complex, multi-stage workflows that were previously intractable. Organizations seeking to deploy autonomous agents in sparse-feedback environments should explore hierarchical goal-setting architectures to improve sample efficiency and task success.
To scale these capabilities further, future development should focus on integrating automatic, unsupervised object discovery directly from raw sensory inputs and incorporating flexible short-term memory to handle long-range dependencies and non-Markovian settings. Decision-makers should note that current implementations rely on custom, pre-defined object detectors to specify the candidate goal space, meaning broad generalization to entirely unstructured environments remains an ongoing area of research.
- Paper: Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning, Richard S. Sutton et al. (1999). This seminal paper introduces the options framework and semi-Markov decision processes, establishing the theoretical foundations of temporal abstraction that h-DQN integrates into deep architectures.
- Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). This foundational work demonstrates human-level control from raw pixels using Deep Q-Networks, providing the core value-function approximation algorithm extended by h-DQN.
- Paper: Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition, Thomas G. Dietterich (1999). This work establishes the MAXQ value function decomposition for hierarchical reinforcement learning, providing key concepts for separating top-level goal selection from subtask execution.
- Paper: Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping, Andrew Y. Ng et al. (1999). This paper formalizes potential-based reward shaping to guide exploration in sparse-reward tasks, laying theoretical groundwork for intrinsic reward and goal motivation mechanisms.
- Paper: Exploration by Random Network Distillation, Yuri Burda et al. (2019). Building on the challenge of sparse-feedback benchmarks like Montezuma's Revenge explored by h-DQN, this work introduces Random Network Distillation to scale intrinsic novelty rewards in complex deep RL environments.
- Paper: Curiosity-Driven Exploration by Self-Supervised Prediction, Deepak Pathak et al. (2017). This paper advances intrinsic motivation by using self-supervised prediction error as curiosity rewards, offering an unsupervised alternative to explicit multi-level goal specifications.
- Paper: The Option-Critic Architecture, Pierre-Luc Bacon et al. (2016). This work extends temporal abstraction by learning options and their termination conditions end-to-end via policy gradients, removing the need to manually define intrinsic goal spaces.
- Paper: Hindsight Experience Replay, Marcin Andrychowicz et al. (2017). This work addresses goal-conditioned reinforcement learning in sparse-reward settings by replaying trajectories with hindsight goals, providing a complementary approach to hierarchical exploration.
- Paper: Diversity is All You Need: Learning Skills without a Reward Function, Benjamin Eysenbach et al. (2018). This paper generalizes intrinsic motivation by learning task-agnostic, diverse repertoires of skills entirely without external reward functions for downstream hierarchical transfer.
