Diversity is All You Need: Learning Skills without a Reward Function

Benjamin EysenbachAbhishek GuptaJulian IbarzSergey Levine

article2018ICLR1,332 citations

Introduces DIAYN, an unsupervised reinforcement learning method that discovers diverse, functional motor skills by maximizing mutual information without task rewards, enabling effective exploration and pretraining for sparse-reward environments.

Listen

Modern artificial intelligence often struggles to learn complex physical behaviors in environments where clear feedback or performance scores are unavailable or expensive to collect. Designing custom reward signals for every unique task requires extensive human effort and can lead to unintended behaviors. The article introduces a method called Diversity is All You Need (DIAYN), which enables automated systems to independently explore simulated environments and acquire a wide library of useful motor skills without requiring any reward signals, external supervision, or human intervention.

The research evaluated DIAYN across several benchmark continuous-control environments, ranging from simple navigation systems to complex robotic bodies such as simulated walkers and an 111-dimensional quadruped. The core objective was to demonstrate that an information-theoretic training mechanism—pairing a skill-prediction classifier with an exploration-driven reinforcement learning algorithm—can autonomously generate diverse, stable movement behaviors that directly accelerate solving downstream tasks.

The findings show that this unsupervised objective produces rich behavioral repertoires, including walking, jumping, balancing, and forward and backward running, with several discovered skills solving standard benchmarks entirely by chance without ever observing a task reward. When these pretrained skills were used to initialize neural network policies for specific downstream goals, training proceeded substantially faster than starting from scratch. Furthermore, organizing these learned skills under a higher-level coordinator allowed robots to complete complex, sparse-reward objectives—such as navigating sequentially through multiple waypoints and clearing hurdles—outperforming established reinforcement learning and intrinsic exploration baselines. In imitation benchmarks across 600 trials, the method reliably matched synthetic expert demonstrations more closely than alternative baselines, while its internal scoring served as an effective confidence metric for tracking accuracy.

These results demonstrate that unsupervised skill discovery can significantly reduce the data collection, engineering time, and human supervision typically needed to train robotic agents. Organizations can amortize the initial cost of exploratory simulation across multiple downstream tasks, effectively turning complex low-level control problems into manageable high-level skill selection. For practical deployment, practitioners should adopt DIAYN as a modular pretraining step for sparse-reward control problems, optionally biasing skill discovery by focusing the classifier on specific state variables when prior domain knowledge is available. However, because the framework relies on finite skill libraries and simulated physics, teams should validate coverage on target physical hardware and conduct pilot evaluations before deploying autonomous controllers in safety-critical settings.

arXiv: 1802.06070
  • Paper: Reinforcement Learning with Deep Energy-Based Policies, Tuomas Haarnoja et al. (2017). Establishes the maximum entropy reinforcement learning framework that DIAYN directly relies on to maximize policy entropy and encourage diverse skill discovery.
  • Paper: The Option-Critic Architecture, Pierre-Luc Bacon et al. (2016). Introduces the theoretical foundations of learning temporal abstractions and option policies in reinforcement learning, which motivates DIAYN's focus on unsupervised skill acquisition.
  • Paper: Curiosity-Driven Exploration by Self-Supervised Prediction, Deepak Pathak et al. (2017). Presents intrinsic curiosity and self-supervised prediction errors as reward-free exploration mechanisms, offering vital context for DIAYN's information-theoretic alternative to unsupervised skill learning.
  • Paper: Unifying Count-Based Exploration and Intrinsic Motivation, Marc G. Bellemare et al. (2016). Provides fundamental background on intrinsic motivation and exploration bonuses in high-dimensional state spaces without external rewards.
  • Paper: Maximum Entropy Inverse Reinforcement Learning, Brian D. Ziebart et al. (2008). Introduces the foundational maximum entropy formulation for decision-making and trajectory distributions underlying modern entropy-regularized RL methods.
Cover for Diversity is All You Need: Learning Skills without a Reward Function

Abstract

Intelligent creatures can explore their environments and learn useful skills without supervision. In this paper, we propose DIAYN ('Diversity is All You Need'), a method for learning useful skills without a reward function. Our proposed method learns skills by maximizing an information theoretic objective using a maximum entropy policy. On a variety of simulated robotic tasks, we show that this simple objective results in the unsupervised emergence of diverse skills, such as walking and jumping. In a number of reinforcement learning benchmark environments, our method is able to learn a skill that solves the benchmark task despite never receiving the true task reward. We show how pretrained skills can provide a good parameter initialization for downstream tasks, and can be composed hierarchically to solve complex, sparse reward tasks. Our results suggest that unsupervised discovery of skills can serve as an effective pretraining mechanism for overcoming challenges of exploration and data efficiency in reinforcement learning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Diversity is All You Need
  • 3.1 How it Works
  • 3.2 Implementation
  • 3.3 Stability
  • 4 Experiments
  • 4.1 Analysis of Learned Skills
  • 4.2 Harnessing Learned Skills
  • 4.2.1 Accelerating Learning with Policy Initialization
  • 4.2.2 Using Skills for Hierarchical RL
  • 4.2.3 Imitating an Expert
  • 5 Conclusion
  • References
  • A Pseudo-Reward
  • B Optimum for Gridworlds
  • C Experimental Details
  • C.1 Environments
  • C.2 Hierarchical RL Experiment
  • D More Analysis of DIAYN Skills
  • D.1 Training Objectives
  • D.2 Effect of Entropy Regularization
  • D.3 Distribution over Task Reward
  • D.4 Exploration
  • E Learning p⁡(z)p(z)
  • E.1 How to Learn p⁡(z)p(z)
  • E.2 Effect of Learning p⁡(z)p(z)
  • F Visualizing Learned Skills
  • F.1 Classic Control Tasks
  • F.2 Simulated Robot Tasks
  • G Imitation Learning
  • G.1 Imitation Learning Experiments

Knowls

  1. Knowl 1 — Information-Theoretic Skill Discovery Objective

    model/method

    Diversity is All You Need (DIAYN) formulates unsupervised skill learning as an information-theoretic optimization problem. Let SS and AA denote random variables for states and actions, and let Z∼p(z)Z \sim p(z) be a latent skill variable conditioned upon by policy πθ(a∣s,z)\pi_\theta(a \mid s, z). The objective F(θ)\mathcal{F}(\theta) maximizes the mutual information between states and skills I(S;Z)I(S; Z) so that skills dictate visited states, maximizes the action entropy given states H[A∣S]H[A \mid S] to encourage exploration, and minimizes the conditional mutual information I(A;Z∣S)I(A; Z \mid S) to ensure skills are distinguished by states rather than redundant actions:

    F(θ)≜I(S;Z)+H[A∣S]−I(A;Z∣S)=H[Z]−H[Z∣S]+H[A∣S,Z]\mathcal{F}(\theta) \triangleq I(S; Z) + H[A \mid S] - I(A; Z \mid S) = H[Z] - H[Z \mid S] + H[A \mid S, Z]

    Here, H[⋅]H[\cdot] and I(⋅;⋅)I(\cdot; \cdot) denote Shannon entropy and mutual information computed with base ee. Because the true posterior distribution p(z∣s)p(z \mid s) is intractable over continuous state spaces, it is approximated with a learned discriminator qϕ(z∣s)q_\phi(z \mid s). By Jensen's inequality, replacing p(z∣s)p(z \mid s) with qϕ(z∣s)q_\phi(z \mid s) yields a variational lower bound G(θ,ϕ)≤F(θ)\mathcal{G}(\theta, \phi) \le \mathcal{F}(\theta):

    G(θ,ϕ)≜H[A∣S,Z]+Ez∼p(z),s∼πθ(z)[log⁡qϕ(z∣s)−log⁡p(z)]\mathcal{G}(\theta, \phi) \triangleq H[A \mid S, Z] + \mathbb{E}_{z \sim p(z), s \sim \pi_\theta(z)} [\log q_\phi(z \mid s) - \log p(z)]

    Fixing p(z)p(z) to be a uniform distribution over a discrete set of skills guarantees maximum prior entropy H[Z]H[Z].

  2. Knowl 2 — DIAYN Skill Discovery Algorithm

    algorithm

    DIAYN trains skill-conditioned policies πθ(a∣s,z)\pi_\theta(a \mid s, z) concurrently with a skill discriminator qϕ(z∣s)q_\phi(z \mid s) using Soft Actor-Critic (SAC). At the start of each episode, a skill index zz is sampled from a fixed uniform prior p(z)p(z) and held constant throughout the episode. At each transition (st,at,st+1)(s_t, a_t, s_{t+1}), the policy receives an intrinsic pseudo-reward:

    rz(st,at)=log⁡qϕ(z∣st+1)−log⁡p(z)r_z(s_t, a_t) = \log q_\phi(z \mid s_{t+1}) - \log p(z)

    Subtracting log⁡p(z)\log p(z) acts as a baseline ensuring non-negative rewards under a discriminator that performs better than chance (qϕ(z∣s)≥p(z)q_\phi(z \mid s) \ge p(z)). Policy parameters θ\theta are updated via SAC with an action-entropy regularization weight α=0.1\alpha = 0.1, while the discriminator parameters ϕ\phi are updated using stochastic gradient descent on the cross-entropy loss E[−log⁡qϕ(z∣st+1)]\mathbb{E}[-\log q_\phi(z \mid s_{t+1})].

    Input: Fixed prior distribution p(z)p(z), initial policy parameters θ\theta, discriminator parameters ϕ\phi, entropy scale α\alpha
    while not converged do
        Sample skill z∼p(z)z \sim p(z) and initial state s0∼p0(s)s_0 \sim p_0(s)
        for t←0t \leftarrow 0 to steps_per_episode - 1 do
            Sample action at∼πθ(at∣st,z)a_t \sim \pi_\theta(a_t \mid s_t, z)
            Step environment: st+1∼p(st+1∣st,at)s_{t+1} \sim p(s_{t+1} \mid s_t, a_t)
            Compute discriminator probability qϕ(z∣st+1)q_\phi(z \mid s_{t+1})
            Set intrinsic reward rt←log⁡qϕ(z∣st+1)−log⁡p(z)r_t \leftarrow \log q_\phi(z \mid s_{t+1}) - \log p(z)
            Update policy parameters θ\theta to maximize rt+αH[A∣st,z]r_t + \alpha H[A \mid s_t, z] via SAC
            Update discriminator parameters ϕ\phi via SGD on log⁡qϕ(z∣st+1)\log q_\phi(z \mid s_{t+1})
        end for
    end while
  3. Knowl 3 — State Space Partitioning and Bottleneck Preference in Gridworlds

    theoretical result

    In an N×NN \times N gridworld with state coordinates (x,y)∈{1,…,N}2(x, y) \in \{1, \dots, N\}^2 and deterministic cardinal transitions, the global optimum of the unregularized DIAYN objective (entropy scale α=0\alpha = 0) for K=2K=2 skills corresponds to an equal partition of the state space. Specifically, skills π1\pi_1 and π2\pi_2 with stationary distributions

    ρπ1(x,y)=2N2I(y≤N/2)andρπ2(x,y)=2N2I(y>N/2)\rho_{\pi_1}(x, y) = \frac{2}{N^2} \mathbb{I}(y \le N/2) \quad \text{and} \quad \rho_{\pi_2}(x, y) = \frac{2}{N^2} \mathbb{I}(y > N/2)

    achieve exact discriminability (H[Z∣S]=0H[Z \mid S] = 0) and maximal prior entropy (H[Z]=log⁡2H[Z] = \log 2).

    When action entropy regularization is present (α>0\alpha > 0), conflicts between action entropy H[A∣S,Z]H[A \mid S, Z] and discriminability −H[Z∣S]-H[Z \mid S] occur exclusively along partition boundaries. The stationary distribution over the equal split achieves an objective within log⁡42N=O(1/N)\frac{\log 4}{2N} = \mathcal{O}(1/N) of the global optimum. Because the boundary conflict incurs an entropy penalty proportional to the number of border states, the DIAYN objective strictly prefers state partitions with minimal boundary length, naturally placing skill boundaries across bottleneck states in environments with non-uniform geometry.

  4. Knowl 4 — Downstream Task Adaptation via Policy and Critic Initialization

    model/method

    Pretrained DIAYN skills provide parameter initializations for accelerated learning on downstream tasks with extrinsic reward functions rtask(s,a)r_{\text{task}}(s, a). Given KK unsupervised skills, the skill policy πθ(a∣s,z)\pi_\theta(a \mid s, z) with the highest zero-shot cumulative task return is selected. Both the actor network πθ\pi_\theta and the critic network QQ learned during pretraining are used to initialize downstream RL training.

    Although the pretrained critic estimates the pseudo-reward rz(s,a)=log⁡qϕ(z∣s)−log⁡p(z)r_z(s, a) = \log q_\phi(z \mid s) - \log p(z) rather than rtaskr_{\text{task}}, initializing both the critic and the actor improves sample efficiency on continuous control benchmarks (HalfCheetah, Hopper, Ant) relative to initializing from scratch.

  5. Knowl 5 — Hierarchical Reinforcement Learning with Discovered Skills

    model/method

    DIAYN skills can serve as motion primitives for hierarchical reinforcement learning (HRL). A high-level meta-controller policy is trained to emit a discrete action corresponding to a skill choice z∈{1,…,K}z \in \{1, \dots, K\} every kk environment steps (e.g., k=10k=10 for cheetah hurdle tasks, k=100k=100 for ant navigation tasks). During the kk intervening steps, the low-level policy πθ(a∣s,z)\pi_\theta(a \mid s, z) executes continuous control conditioned on the selected zz.

    The meta-controller observes the environment state ss and is trained directly with task rewards, replacing high-frequency continuous action spaces with temporally extended, diverse behavioral options.

  6. Knowl 6 — Expert Imitation via Discriminator M-Projection

    model/method

    Given an unlabelled, state-only expert trajectory τ∗=(s1,s2,…,sT)\tau^* = (s_1, s_2, \dots, s_T) without action observations, an imitation skill z^\hat{z} is retrieved from a discrete set of pretrained DIAYN skills by evaluating the learned discriminator qϕ(z∣s)q_\phi(z \mid s):

    z^=arg⁡max⁡z∏t=1Tqϕ(z∣st)\hat{z} = \arg\max_z \prod_{t=1}^T q_\phi(z \mid s_t)

    Under a uniform prior p(z)p(z) and an optimal discriminator qϕ(z∣s)=p(z∣s)q_\phi(z \mid s) = p(z \mid s), applying Bayes' rule gives p(s∣z)∝qϕ(z∣s)p(s \mid z) \propto q_\phi(z \mid s). Maximizing the product of discriminator probabilities corresponds to finding the M-projection of the empirical expert state distribution p∗p^* onto the family of skill state distributions P={pz}z=1K\mathcal{P} = \{p^z\}_{z=1}^K:

    z^=arg⁡min⁡pz∈PDKL(p∗∥pz)\hat{z} = \arg\min_{p^z \in \mathcal{P}} D_{\text{KL}}(p^* \parallel p^z)

    The M-projection (zero-avoiding KL divergence) penalizes the policy if it fails to cover states visited by the expert.

  7. Knowl 7 — Biasing Skill Discovery with Feature-Conditioned Discriminators

    model/method

    When state observation spaces are high-dimensional, the skill discriminator can be conditioned on a chosen feature function f(s)f(s) rather than the full state ss. The discriminator then optimizes the log-likelihood:

    Ez∼p(z),s∼πθ(z)[log⁡qϕ(z∣f(s))]\mathbb{E}_{z \sim p(z), s \sim \pi_\theta(z)} [\log q_\phi(z \mid f(s))]

    This modification biases skill discovery toward specific degrees of freedom. For instance, setting f(s)f(s) to the (x,y)(x, y) center-of-mass coordinates in locomotion environments forces skills to diversify their center-of-mass trajectories across the 2D plane while ignoring limb configurations.

  8. Knowl 8 — Fixed vs. Learned Prior Distribution and Skill Collapse

    empirical result

    Learning the skill prior distribution p(z)p(z) dynamically during unsupervised skill optimization causes severe skill collapse due to the Matthew effect: skills that become discriminable early receive more probability mass under p(z)∝exp⁡(Es∼pz(s)[log⁡q(z∣s)])p(z) \propto \exp(\mathbb{E}_{s \sim p_z(s)}[\log q(z \mid s)]), starving less-developed skills of gradient updates.

    Measuring the effective number of active skills as exp⁡(H[Z])\exp(H[Z]), models with a learned prior (such as Variational Intrinsic Control) suffer a tenfold drop in effective skill count across HalfCheetah, MountainCar, and InvertedPendulum. Fixing p(z)p(z) to a uniform categorical distribution maintains maximum skill entropy exp⁡(H[Z])=K\exp(H[Z]) = K throughout training and ensures all KK skills continue to receive training signals.

  9. Knowl 9 — Hierarchical Exploration and Task Performance on Sparse Benchmarks

    empirical result

    In long-horizon, sparse-reward continuous control tasks, composing DIAYN skills with a high-level meta-controller outperforms standard model-free and intrinsic exploration baselines:

    1. Cheetah Hurdle: In HalfCheetah modified with vertical hurdles spaced every 3 meters, a meta-controller switching DIAYN skills every 10 steps learns to bound over obstacles, reaching rewards above 3 within 15 hours of training, while non-hierarchical algorithms (TRPO, SAC, and VIME) fail to progress.
    2. Ant Navigation: In an Ant continuous-control maze requiring sequential visits to 5 distant waypoints for sparse +1+1 rewards, a meta-controller selecting DIAYN skills every 100 steps achieves total rewards near the maximum score of +5+5, whereas flat RL baselines and single-policy curiosity exploration (VIME) obtain scores near 0.
  10. Knowl 10 — State-Space Coverage Comparison Against Single-Policy Exploration

    empirical result

    Evaluating unsupervised DIAYN skills against single-policy exploration methods (VIME) across standard continuous control benchmarks (HalfCheetah, Hopper, Ant) on three downstream metrics—running (maximizing X coordinate), jumping (maximizing Z coordinate), and displacement (maximizing L2 distance from origin)—demonstrates that DIAYN discovers policies with specialized performance on each metric without seeing their rewards during pretraining. In contrast, curiosity-driven exploration bonuses that train a single policy fail to maintain high-performing coverage across distinct, specialized behaviors.

Coverage note — None was omitted; all primary contributions—including the mutual information objective, variational lower bound, SAC algorithm implementation, gridworld theoretical proofs, prior fixing analysis, fine-tuning, hierarchical control, feature biasing, and imitation via M-projection—are captured in self-contained knowls.

References

  1. 1.Joshua Achiam, Harrison Edwards, Dario Amodei, and Pieter Abbeel. Variational autoencoding learning of options by reinforcement. NIPS Deep Reinforcement Learning Symposium, 2017.
  2. 2.David Barber Felix Agakov. The im algorithm: a variational approach to information maximization. Advances in Neural Information Processing Systems, 16:201, 2004.
  3. 3.Pierre-Luc Bacon, Jean Harb, and Doina Precup. The option-critic architecture. In AAAI, pp. 1726–1734, 2017.
  4. 4.Adrien Baranes and Pierre-Yves Oudeyer. Active learning of inverse models with intrinsically motivated goal exploration in robots. Robotics and Autonomous Systems, 61(1):49–73, 2013.
  5. 5.Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos. Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems, pp. 1471–1479, 2016.
  6. 6.Christopher M Bishop. Pattern Recognition and Machine Learning. Springer-Verlag New York, 2016.
  7. 7.Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. arXiv preprint arXiv:1606.01540, 2016.
  8. 8.Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems, pp. 4302–4310, 2017.
  9. 9.Peter Dayan and Geoffrey E Hinton. Feudal reinforcement learning. In Advances in neural information processing systems, pp. 271–278, 1993.
  10. 10.Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. Benchmarking deep reinforcement learning for continuous control. In International Conference on Machine Learning, pp. 1329–1338, 2016.
  11. 11.Carlos Florensa, Yan Duan, and Pieter Abbeel. Stochastic neural networks for hierarchical reinforcement learning. arXiv preprint arXiv:1704.03012, 2017.
  12. 12.Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman. Meta learning shared hierarchies. arXiv preprint arXiv:1710.09767, 2017.
  13. 13.Justin Fu, John Co-Reyes, and Sergey Levine. Ex2: Exploration with exemplar models for deep reinforcement learning. In Advances in Neural Information Processing Systems, pp. 2574–2584, 2017.
  14. 14.Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra. Variational intrinsic control. arXiv preprint arXiv:1611.07507, 2016.
  15. 15.Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine. Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In Robotics and Automation (ICRA), 2017 IEEE International Conference on, pp. 3389–3396. IEEE, 2017.
  16. 16.Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. Reinforcement learning with deep energy-based policies. arXiv preprint arXiv:1702.08165, 2017.
  17. 17.Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. arXiv preprint arXiv:1801.01290, 2018.
  18. 18.Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan. Inverse reward design. In Advances in Neural Information Processing Systems, pp. 6768–6777, 2017.
  19. 19.Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller. Learning an embedding space for transferable robot skills. International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rk07ZXZRb.
  20. 20.Nicolas Heess, Greg Wayne, Yuval Tassa, Timothy Lillicrap, Martin Riedmiller, and David Silver. Learning and transfer of modulated locomotor controllers. arXiv preprint arXiv:1610.05182, 2016.
  21. 21.Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. arXiv preprint arXiv:1709.06560, 2017.
  22. 22.Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. Vime: Variational information maximizing exploration. In Advances in Neural Information Processing Systems, pp. 1109–1117, 2016.
  23. 23.Tobias Jung, Daniel Polani, and Peter Stone. Empowerment for continuous agent—environment systems. Adaptive Behavior, 19(1):16–39, 2011.
  24. 24.Sanjay Krishnan, Roy Fox, Ion Stoica, and Ken Goldberg. Ddco: Discovery of deep continuous options for robot learning from demonstrations. In Conference on Robot Learning, pp. 418–437, 2017.
  25. 25.Joel Lehman and Kenneth O Stanley. Abandoning objectives: Evolution through the search for novelty alone. Evolutionary computation, 19(2):189–223, 2011a.
  26. 26.Joel Lehman and Kenneth O Stanley. Evolving a diversity of virtual creatures through novelty search and local competition. In Proceedings of the 13th annual conference on Genetic and evolutionary computation, pp. 211–218. ACM, 2011b.
  27. 27.Robert K Merton. The matthew effect in science: The reward and communication systems of science are considered. Science, 159(3810):56–63, 1968.
  28. 28.Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al. Learning to navigate in complex environments. arXiv preprint arXiv:1611.03673, 2016.
  29. 29.Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602, 2013.
  30. 30.Shakir Mohamed and Danilo Jimenez Rezende. Variational information maximisation for intrinsically motivated reinforcement learning. In Advances in neural information processing systems, pp. 2125–2133, 2015.
  31. 31.Jean-Baptiste Mouret and Stéphane Doncieux. Overcoming the bootstrap problem in evolutionary robotics using behavioral diversity. In Evolutionary Computation, 2009. CEC’09. IEEE Congress on, pp. 1161–1168. IEEE, 2009.
  32. 32.Kevin P Murphy. Machine Learning: A Probabilistic Perspective. MIT Press, 2012.
  33. 33.Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans. Bridging the gap between value and policy based reinforcement learning. In Advances in Neural Information Processing Systems, pp. 2772–2782, 2017.
  34. 34.Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner. Intrinsic motivation systems for autonomous mental development. IEEE transactions on evolutionary computation, 11(2):265–286, 2007.
  35. 35.Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. arXiv preprint arXiv:1705.05363, 2017.
  36. 36.Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine. Temporal difference models: Model-free deep rl for model-based control. arXiv preprint arXiv:1802.09081, 2018.
  37. 37.Justin K Pugh, Lisa B Soros, and Kenneth O Stanley. Quality diversity: A new frontier for evolutionary computation. Frontiers in Robotics and AI, 3:40, 2016.
  38. 38.Richard M Ryan and Edward L Deci. Intrinsic and extrinsic motivations: Classic definitions and new directions. Contemporary educational psychology, 25(1):54–67, 2000.
  39. 39.Jürgen Schmidhuber. Formal theory of creativity, fun, and intrinsic motivation. IEEE Transactions on Autonomous Mental Development, 2(3):230–247, 2010.
  40. 40.John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. Trust region policy optimization. In International Conference on Machine Learning, pp. 1889–1897, 2015a.
  41. 41.John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438, 2015b.
  42. 42.John Schulman, Pieter Abbeel, and Xi Chen. Equivalence between policy gradients and soft q-learning. arXiv preprint arXiv:1704.06440, 2017.
  43. 43.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  44. 44.David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484–489, 2016.
  45. 45.Kenneth O Stanley and Risto Miikkulainen. Evolving neural networks through augmenting topologies. Evolutionary computation, 10(2):99–127, 2002.
  46. 46.Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Kenneth O Stanley, and Jeff Clune. Deep neuroevolution: Genetic algorithms are a competitive alternative for training deep neural networks for reinforcement learning. arXiv preprint arXiv:1712.06567, 2017.
  47. 47.Sainbayar Sukhbaatar, Zeming Lin, Ilya Kostrikov, Gabriel Synnaeve, Arthur Szlam, and Rob Fergus. Intrinsic motivation and automatic curricula via asymmetric self-play. arXiv preprint arXiv:1703.05407, 2017.
  48. 48.Brian G Woolley and Kenneth O Stanley. On the deleterious effects of a priori objectives on evolution and representation. In Proceedings of the 13th annual conference on Genetic and evolutionary computation, pp. 957–964. ACM, 2011.
  49. 49.Yuke Zhu, Roozbeh Mottaghi, Eric Kolve, Joseph J Lim, Abhinav Gupta, Li Fei-Fei, and Ali Farhadi. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In Robotics and Automation (ICRA), 2017 IEEE International Conference on, pp. 3357–3364. IEEE, 2017.
  50. 50.Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, and Anind K Dey. Maximum entropy inverse reinforcement learning. In AAAI, volume 8, pp. 1433–1438. Chicago, IL, USA, 2008.

Citation

MLA
Eysenbach, B., et al. “Diversity Is All You Need: Learning Skills Without a Reward Function”. arXiv, 2018, http://arxiv.org/abs/1802.06070v6.
APA
Eysenbach, B., Gupta, A., Ibarz, J., & Levine, S. (2018). Diversity is All You Need: Learning Skills without a Reward Function. arXiv. http://arxiv.org/abs/1802.06070v6
Chicago
Eysenbach, B., A. Gupta, J. Ibarz, and S. Levine. 2018. “Diversity Is All You Need: Learning Skills Without a Reward Function”. arXiv. http://arxiv.org/abs/1802.06070v6.
Harvard
Eysenbach, B. et al. (2018) “Diversity is All You Need: Learning Skills without a Reward Function”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1802.06070v6.
Vancouver
1. Eysenbach B, Gupta A, Ibarz J, Levine S (2018) Diversity is All You Need: Learning Skills without a Reward Function. arXiv

BibTeX

@article{eysenbach2018diversity,
  title = {Diversity is All You Need: Learning Skills without a Reward Function},
  author = {Eysenbach, Benjamin and Gupta, Abhishek and Ibarz, Julian and Levine, Sergey},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1802.06070v6},
  eprint = {1802.06070}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission