A comprehensive survey on safe reinforcement learning

Javier GarcíaFernando Fernández

article2015JMLR2,141 citations
Listen

As autonomous systems and physical robots are deployed in increasingly complex, real-world environments, traditional machine learning methods present serious operational hazards. Standard reinforcement learning trains an autonomous agent by rewarding success and penalizing failure through trial and error. However, this process traditionally relies on random exploration, which can lead to catastrophic damage or severe financial losses when applied to physical systems like robotic aircraft or industrial machinery. Because training solely in digital simulators often fails to reflect real-world physics, algorithms must be able to learn safely during live deployment without causing catastrophic failures.

The article establishes a structured taxonomy of safe reinforcement learning methods and evaluates their capacity to prevent dangerous outcomes while maintaining high system performance. To accomplish this, the authors conduct a comprehensive survey analyzing dozens of academic approaches published across control engineering, robotics, and artificial intelligence.

The findings show that safe learning methods fall into two core categories: modifying the mathematical objective or transforming the exploration process. Under the first category, systems modify their target criteria by planning for worst-case outcomes, penalizing return variance, or setting hard constraints. While these methods produce cautious policies once fully trained, they often lead to extreme pessimism and still require systems to experience hazardous states repeatedly before recognizing their danger. Under the second category, systems alter exploration by incorporating external knowledge or applying risk metrics. Utilizing expert demonstrations to initialize learning or applying teacher advice allows autonomous agents to safely navigate complex environments from the very start of training.

These insights demonstrate that mathematical safety objectives alone are insufficient for physical deployments. Modifying optimization criteria ensures long-term risk avoidance but leaves systems unprotected during early learning phases. In contrast, incorporating external guidance successfully mitigates immediate risk and reduces costly trial-and-error sampling, directly lowering operational risk, hardware damage, and equipment downtime.

Organizations deploying autonomous learning systems should adopt hybrid architectures rather than relying on a single safety technique. Specifically, teams should implement automated teacher-advice mechanisms to safeguard early exploratory actions while simultaneously applying risk-sensitive or constrained criteria to ensure long-term stability. Development roadmaps should prioritize automated risk detection based on state novelty rather than relying on continuous, subjective human monitoring.

Decision-makers must note that most surveyed techniques have only been validated in discrete, small-scale simulations or controlled robotic benchmarks. Real-world environments with continuous, high-dimensional action spaces remain a significant technical challenge. While there is high confidence in the foundational taxonomy and the identified failure modes of pure trial-and-error learning, deploying these methods in safety-critical production settings still requires rigorous piloting and domain-specific safety boundaries.

García et al (2015).pdf
  • Paper: Reinforcement Learning: A Survey, Leslie Pack Kaelbling et al. (1996). This foundational survey establishes the core Markov decision process formulation, exploration-exploitation trade-offs, and basic reinforcement learning algorithms that the safe RL literature directly builds upon.
  • Paper: Technical Note: Q-Learning, CHRISTOPHER J.C.H. WATKINS et al. (2004). This seminal note establishes the mathematical formulation and convergence proof of Q-learning, providing the baseline unconstrained algorithmic foundation that safe reinforcement learning aims to constrain and safeguard.
  • Paper: Policy Gradient Methods for Reinforcement Learning with Function Approximation, Richard S. Sutton et al. (1999). This classic work derives the fundamental policy gradient theorem and optimization principles that underpin policy search methods and modern constrained policy optimization in safe RL.
  • Paper: A robust layered control system for a mobile robot, Rodney A. Brooks (1986). This landmark paper introduces layered behavioral architectures for robotic control, establishing the foundational paradigm of external safety supervisors and layered behavioral competence in physical agents.
  • Paper: Using Confidence Bounds for Exploitation-Exploration Trade-offs, P. Auer (2003). This paper establishes formal confidence bounds and risk-aware exploration frameworks that directly motivate risk-sensitive objectives and optimism-under-uncertainty exploration analyzed in the survey.
  • Paper: Deterministic Policy Gradient Algorithms, David Silver et al. (2014). This foundational paper develops deterministic policy gradients for continuous control, formulating the baseline continuous actor-critic framework that safe RL adapts for physical systems.

Table of Contents

  • 1. Introduction
  • 2. Overview of Safe Reinforcement Learning
  • 3. Modifying the Optimization Criterion
  • 3.1 Worst-Case Criterion
  • 3.1.1 Worst-Case Criterion under Inherent Uncertainty
  • 3.1.2 Worst-Case Criterion under Parameter Uncertainty
  • 3.2 Risk-Sensitive Criterion
  • 3.2.1 Risk-Sensitive Based on Exponential Functions
  • 3.2.2 Risk-Sensitive RL Based on the Weighted Sum of Return and Risk
  • 3.3 Constrained Criterion
  • 3.4 Other Optimization Criteria
  • 4. Modifying the Exploration Process
  • 4.1 Incorporating External Knowledge
  • 4.1.1 Providing Initial Knowledge
  • 4.1.2 Deriving a Policy from a Finite Set of Demonstrations
  • 4.1.3 Using Teacher Advice
  • 4.1.3.1 The Learner Agent Asks for Advice
  • 4.1.3.2 The Teacher Provides Advice
  • 4.1.3.3 Other Approaches
  • 4.2 Risk-directed Exploration
  • 5. Discussion and Open Issues
  • 5.1 Characterization of Safe RL Algorithms
  • 5.1.1 Allowed Learner
  • 5.1.2 Space Complexity
  • 5.1.3 Risk
  • 5.1.4 Exploration
  • 5.2 Discussion
  • 5.2.1 Selection of the Risk Metric
  • 5.2.2 Selection of the Optimization Criterion
  • 5.2.3 Selection of the Mechanism for Risk Detection
  • 5.2.4 Selection of the Learning Schema
  • 5.2.5 Selection of the Exploration Strategy
  • 6. Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Definition and Two-Branch Taxonomy of Safe Reinforcement Learning

    definition

    Safe Reinforcement Learning (Safe RL) is defined as the process of learning control policies that maximize the expectation of the return in Markov Decision Processes (MDPs) while ensuring acceptable system performance and/or satisfying safety constraints during both the learning and deployment phases.

    Safe RL methodologies are organized into two primary branches:

    1. Modifying the Optimization Criterion: The standard risk-neutral objective (maximizing expected discounted future return) is replaced or augmented to incorporate safety considerations. This branch comprises:

      • Worst-Case Criteria: Optimizing performance for the worst-case trajectory under inherent stochastic uncertainty, or for the worst-case transition dynamics model under parameter uncertainty.
      • Risk-Sensitive Criteria: Controlling sensitivity to risk through exponential utility functions or linear combinations balancing return and risk metrics (such as return variance, temporal-difference error volatility, or the probability of entering error states).
      • Constrained Criteria: Maximizing return subject to bounds on performance thresholds, variance, or ergodicity preservation.
      • Other Criteria: Applying risk formulations derived from financial engineering, such as Value-at-Risk (VaR), Conditional Value-at-Risk (CVaR), and Sharpe ratio.
    2. Modifying the Exploration Process: The standard optimization criterion is retained, but the action-selection mechanism during learning is altered to prevent catastrophic or unrecoverable states. This branch comprises:

      • Incorporating External Knowledge: Using prior information, demonstration data sets, or interactive teacher advising (agent-initiated requests for help, teacher-initiated interventions, or blended/supervised action selection).
      • Risk-Directed Exploration: Employing an explicit risk measure (such as temporal-difference controllability or entropy-based uncertainty) to guide exploration toward safe state-space regions.
  2. Knowl 2 — Worst-Case Minimax Criterion under Inherent Uncertainty

    model/method

    Under inherent stochastic uncertainty in a Markov Decision Process (MDP) defined by state space SS, action space AA, transition probability function TT, and reward function RR, the worst-case or minimax criterion identifies a policy πΠ\pi \in \Pi that maximizes the expected discounted return along the worst possible trajectory realization. Let Ωπ\Omega^\pi denote the set of possible trajectories w=(s0,a0,s1,a1,)w = (s_0, a_0, s_1, a_1, \dots) generated under policy π\pi, and let γ[0,1]\gamma \in [0, 1] denote the discount factor. The minimax objective is expressed as:

    maxπΠminwΩπEπ,w(t=0γtrt)\max_{\pi \in \Pi} \min_{w \in \Omega^\pi} E_{\pi,w}\left(\sum_{t=0}^\infty \gamma^t r_t\right)

    where Eπ,w()E_{\pi,w}(\cdot) denotes the expectation with respect to policy π\pi along trajectory ww, and rtr_t is the reward at time step tt.

    In model-free learning, this objective is approximated via Q^\hat{Q}-learning, which maintains a lower bound on state-action value estimates through the update rule:

    Q^(st,at)min(Q^(st,at),rt+1+γmaxat+1AQ^(st+1,at+1))\hat{Q}(s_t, a_t) \leftarrow \min\left(\hat{Q}(s_t, a_t), \, r_{t+1} + \gamma \max_{a_{t+1} \in A} \hat{Q}(s_{t+1}, a_{t+1})\right)

    To mitigate the extreme pessimism of Q^\hat{Q}-learning, β\beta-pessimistic Q-learning interpolates between standard Q-learning and minimax evaluation using a risk parameter β[0,1]\beta \in [0, 1] and learning rate α(0,1]\alpha \in (0, 1]:

    Qβ(st,at)Qβ(st,at)+α[rt+1+γ((1β)maxat+1AQβ(st+1,at+1)+βminat+1AQβ(st+1,at+1))Qβ(st,at)]Q_\beta(s_t, a_t) \leftarrow Q_\beta(s_t, a_t) + \alpha \left[ r_{t+1} + \gamma \left((1 - \beta) \max_{a_{t+1} \in A} Q_\beta(s_{t+1}, a_{t+1}) + \beta \min_{a_{t+1} \in A} Q_\beta(s_{t+1}, a_{t+1})\right) - Q_\beta(s_t, a_t) \right]

  3. Knowl 3 — Worst-Case Criterion under Parameter Uncertainty in Robust MDPs

    model/method

    When dynamic parameters of a Markov Decision Process (MDP) are estimated from noisy or limited empirical data, the true transition probability matrix is assumed to lie within an uncertainty set P\mathcal{P}. In Robust MDPs, the objective is to determine a policy πΠ\pi \in \Pi that maximizes the expected discounted return under the worst possible transition model pPp \in \mathcal{P}:

    maxπΠminpPEπ,p(t=0γtrt)\max_{\pi \in \Pi} \min_{p \in \mathcal{P}} E_{\pi,p}\left(\sum_{t=0}^\infty \gamma^t r_t\right)

    where Eπ,p()E_{\pi,p}(\cdot) denotes the expectation with respect to policy π\pi under transition distribution pp, rtr_t is the reward received at time step tt, and γ[0,1]\gamma \in [0, 1] is the discount factor.

    This optimization guarantees performance bounds against model inaccuracies, preventing deployment failures caused by discrepancies between estimated training dynamics and real-world execution.

  4. Knowl 4 — Risk-Sensitive Reinforcement Learning via Exponential Utility Functions

    model/method

    Risk-sensitive control based on exponential utility functions transforms the cumulative discounted return R=t=0γtrtR = \sum_{t=0}^\infty \gamma^t r_t into an exponential objective controlled by a scalar risk-sensitivity parameter β\beta:

    maxπΠ1βlogEπ(exp(βR))=maxπΠ1βlogEπ(exp(βt=0γtrt))\max_{\pi \in \Pi} \frac{1}{\beta} \log E_\pi\left(\exp(\beta R)\right) = \max_{\pi \in \Pi} \frac{1}{\beta} \log E_\pi\left(\exp\left(\beta \sum_{t=0}^\infty \gamma^t r_t\right)\right)

    where γ[0,1]\gamma \in [0, 1] is the discount factor and rtr_t is the reward at step tt.

    A Taylor series expansion demonstrates the direct dependence on the return's mean and variance:

    1βlogEπ(exp(βR))=Eπ(R)+β2Var(R)+O(β2)\frac{1}{\beta} \log E_\pi\left(\exp(\beta R)\right) = E_\pi(R) + \frac{\beta}{2} \operatorname{Var}(R) + O(\beta^2)

    where Var(R)\operatorname{Var}(R) is the variance of the return under policy π\pi:

    • Setting β<0\beta < 0 penalizes return variance, producing risk-averse behavior.
    • Setting β>0\beta > 0 rewards return variance, producing risk-seeking behavior.
    • In the limit β0\beta \to 0, the objective converges to the risk-neutral expected return maxπΠEπ(R)\max_{\pi \in \Pi} E_\pi(R).

    While maximizing expected exponential utility is mathematically equivalent to solving a worst-case robust MDP, optimal policies under this criterion are generally time-dependent, posing challenges for formulating standard model-free temporal-difference or Q-learning algorithms.

  5. Knowl 5 — Risk-Sensitive Reinforcement Learning via Weighted Sum of Return and Risk Metrics

    model/method

    A standard approach to risk-sensitive reinforcement learning expresses the objective function as a weighted linear combination of expected return and an explicit risk metric ω\omega:

    maxπΠ(Eπ(R)βω)\max_{\pi \in \Pi} \left( E_\pi(R) - \beta \omega \right)

    where Eπ(R)=Eπ(t=0γtrt)E_\pi(R) = E_\pi\left(\sum_{t=0}^\infty \gamma^t r_t\right) is the expected discounted return, β\beta is a risk-sensitivity parameter, and ω\omega quantifies risk. Prominent formulations of ω\omega include:

    1. Variance of the Return (Markowitz Mean-Variance): ω=Var(R)\omega = \operatorname{Var}(R). This penalizes return variability. However, variance penalizes positive and negative deviations equally, fails to account for heavy-tailed distributions, and results in an NP-hard optimization problem in general MDPs.
    2. Temporal Difference Errors: ω\omega is derived from the volatility of temporal-difference error signals δt\delta_t, over-weighting transitions to states yielding worse-than-average returns and under-weighting transitions yielding higher-than-average returns.
    3. Probability of Reaching Error States: ω=ρπ(s)\omega = \rho^\pi(s), where ρπ(s)\rho^\pi(s) represents the probability that a state trajectory starting at ss under policy π\pi terminates in an absorbing failure/error state:

    ρπ(s)=E(i=0γirˉi)\rho^\pi(s) = E\left( \sum_{i=0}^\infty \gamma^i \bar{r}_i \right)

    where rˉ=1\bar{r} = 1 if an error state is visited and rˉ=0\bar{r} = 0 otherwise. A drawback of estimating ρπ(s)\rho^\pi(s) via temporal-difference learning is that the agent must repeatedly experience catastrophic states during training to learn the risk function.

  6. Knowl 6 — Constrained MDP Criterion for Safe Reinforcement Learning

    model/method

    In the Constrained Markov Decision Process (CMDP) framework, the agent optimizes expected discounted return subject to functional constraints ciCc_i \in C:

    maxπΠEπ(R)subject to ciC,ci={hiαi}\max_{\pi \in \Pi} E_\pi(R) \quad \text{subject to } c_i \in C, \quad c_i = \{ h_i \le \alpha_i \}

    where hih_i is a function related to policy performance or safety metrics, αi\alpha_i is a defined threshold bound, and the inequality may be \le or \ge. The feasible policy set is Γ={πΠci is satisfied ciC}\Gamma = \{ \pi \in \Pi \mid c_i \text{ is satisfied } \forall c_i \in C \}, reducing the optimization problem to:

    maxπΓEπ(R)\max_{\pi \in \Gamma} E_\pi(R)

    Core constraint formulations in Safe RL include:

    • Return Threshold Constraint: Requiring E(R)αE(R) \ge \alpha (or chance constraints P(E(R)α)1ϵP(E(R) \ge \alpha) \ge 1 - \epsilon) so that policy exploration does not degrade performance below a safety margin.
    • Variance Constraint: Requiring Var(R)α\operatorname{Var}(R) \le \alpha to limit return volatility, commonly converted to an unconstrained problem via Lagrangian multipliers or penalty functions.
    • Ergodicity Preservation Constraint: Restricting exploration to policies that preserve state ergodicity with user-defined probability α\alpha, ensuring that the agent only visits states from which recovery to the initial state is possible.
  7. Knowl 7 — Modifying Safe Exploration via External Knowledge and Teacher Advising

    model/method

    Modifying the exploration process using external knowledge prevents catastrophic exploratory actions during learning while preserving the standard risk-neutral optimization criterion. External knowledge is integrated via three primary mechanisms:

    1. Providing Initial Knowledge (Bootstrapping): Teacher demonstrations or baseline controllers initialize the value function, policy representation, or evolutionary population. This biases exploration toward viable state-space regions from the beginning of learning, bypassing extensive random exploration.
    2. Deriving Policies from Demonstrations (Apprenticeship Learning): Demonstration trajectories recorded from a teacher are used to learn a dynamics model offline, from which an optimal policy is computed and deployed without exploratory risk on the physical system.
    3. Interactive Teacher Advising: An external teacher (human or controller) guides action selection during online learning through three interaction models:
      • Agent-Initiated ("Ask for Help"): The agent tracks its confidence (e.g., via state novelty, distance to known safe states in a case base, or similarity among action QQ-values) and queries teacher advice whenever confidence falls below a threshold.
      • Teacher-Initiated: The teacher monitors the agent and supplies actions, feedback rewards, or rule constraints whenever dangerous states are approached.
      • Shared / Blended Control: The executed action aa is a convex combination of teacher advice aTa_T and the agent's exploratory action aE=aA+N(0,σ)a_E = a_A + \mathcal{N}(0, \sigma):

    a=kaE+(1k)aTa = k a_E + (1 - k) a_T

    where k[0,1]k \in [0, 1] represents the agent's autonomy level based on relative confidence.

  8. Knowl 8 — Risk-Directed Exploration via Controllability and Entropy Metrics

    model/method

    Risk-directed exploration alters action selection probabilities using a risk metric while keeping the underlying optimization criterion unchanged:

    1. Temporal-Difference Controllability Metric: Controllability C(s,a)C(s, a) measures the volatility of the temporal-difference error δt=rt+1+γmaxaQ(st+1,a)Q(st,at)\delta_t = r_{t+1} + \gamma \max_{a'} Q(s_{t+1}, a') - Q(s_t, a_t). High absolute error indicates low controllability. Controllability is updated online as:

    C(st,at)C(st,at)α(δt+C(st,at))C(s_t, a_t) \leftarrow C(s_t, a_t) - \alpha' \left( |\delta_t| + C(s_t, a_t) \right)

    where α\alpha' is a learning rate. Actions are selected greedily with a controllability bonus: argmaxa[Q(st,a)+wC(st,a)]\arg\max_a \left[ Q(s_t, a) + w C(s_t, a) \right], directing the agent toward predictable and controllable state regions.

    1. Entropy and Expected Return Combination: The risk metric U(s,a)U(s, a) computes the weighted sum of transition entropy (stochastic outcome spread) and normalized expected return. A risk-adjusted utility combines safety and expected payoff:

    Uadj(st,at)=p(1U(st,at))+(1p)Q(st,at)U_{\text{adj}}(s_t, a_t) = p \left(1 - U(s_t, a_t)\right) + (1 - p) Q(s_t, a_t)

    where p[0,1]p \in [0, 1]. Action selection is executed by substituting UadjU_{\text{adj}} into a Boltzmann distribution.

    A fundamental limitation is that risk avoidance depends on correctly approximating C(s,a)C(s, a) or U(s,a)U(s, a) through experience; until these functions converge, early-stage exploratory actions remain susceptible to catastrophe.

  9. Knowl 9 — Failure Modes of Variance and Worst-Case Minimax Criteria in Risky Environments

    theoretical result

    Evaluating risk strictly through return variance Var(R)\operatorname{Var}(R) or through the worst-case minimax return minwΩπEπ,w(R)\min_{w \in \Omega^\pi} E_{\pi,w}(R) fails to provide safety guarantees in environments where catastrophic policies have low variance or where all policies share identical worst outcomes.

    Consider a stochastic grid-world featuring normal states, absorbing error states (terminating with reward 0), and absorbing goal states (terminating with reward 1), with intermediate step rewards of 0 and stochastic transition slippage (e.g., 21% probability of moving perpendicular to the chosen action):

    • Failure of Variance Metric: A policy that drives the agent directly and rapidly into an error state produces a sequence of zero rewards with low or zero variance, which is equal to or lower than the variance of a policy navigating safely to the goal. Consequently, minimizing return variance does not penalize policies leading directly to catastrophe.
    • Failure of Minimax Criterion: Because any policy has a non-zero probability of slipping into an error state under stochastic transitions, the worst-case outcome for all policies is identical (0 return). The minimax criterion therefore cannot distinguish safe policies from hazardous ones.

    This counterexample demonstrates that variance-based and minimax criteria are not universally applicable risk metrics across arbitrary MDP domains.

  10. Knowl 10 — Trade-Off Analysis of Safe Reinforcement Learning Paradigms

    data/table

    A comparative synthesis of Safe Reinforcement Learning paradigms reveals fundamental operational trade-offs between long-term risk optimization, online exploration safety, and computational tractability:

    Approach Category Main Advantages Main Drawbacks
    Worst-Case Criterion Effective when avoiding rare occurrences of large negative returns is critical. Overly pessimistic; minimax and variance metrics do not generalize to arbitrary domains; distorts true action utilities; fails to detect risk during early exploration steps.
    Risk-Sensitive Criterion Allows continuous tuning between risk-averse and risk-seeking behavior; addresses long-term risk. Conservative parameters produce overly pessimistic policies; true action utilities are lost; approximating risk functions requires repeatedly experiencing hazardous states.
    Constrained Criterion Natural formulation for safe exploration by restricting search strictly to the allowable policy space Γ\Gamma. Often computationally intractable for large state spaces; sensitive to threshold selection; return or variance bounds do not prevent short-term catastrophic failures.
    Providing Initial Knowledge Bootstraps value functions and guides exploration to relevant state regions from step one. Biased initialization can converge to suboptimal policies; unvisited states encountered after initialization remain dangerous; difficult to implement for complex representations.
    Deriving Policy from Demonstrations Learns model and policy offline, avoiding dangerous online trial-and-error exploration. Performance is strictly bounded by teacher demonstration quality and coverage; action selection in unrepresented states is undefined.
    Using Teacher Advice Prevents catastrophic actions from early learning steps; supports automatic risk detection in ask-for-help schemes. Ask-for-help detects immediate risk but may lack long-term foresight; teacher-initiated advice relies on subjective human judgment and requires continuous, costly monitoring.
    Risk-Directed Exploration Preserves the true value function while using explicit risk metrics to guide exploration. Requires the risk metric function to be accurately estimated before safety is achieved, leaving initial exploration vulnerable.

    Approaches modifying the optimization criterion focus on optimizing long-term risk in the final converged policy but do not guarantee safety during early exploration. In contrast, approaches modifying exploration via external knowledge or teacher advice protect the agent against immediate short-term risk from early learning steps.

Coverage note — Individual application-specific implementations and algorithms cited from the literature (such as d-SARSA with CVaR, HEDGER, RATLE, TAMER, and PI-SRL) were synthesized into the overarching methodological knowls rather than extracted as separate standalone knowls.

References

  1. 1.Pieter Abbeel. Apprenticeship Learning and Reinforcement Learning with Application to Robotic Control. PhD thesis, Stanford, CA, USA, 2008. AAI3332983.
  2. 2.Pieter Abbeel and Andrew Y. Ng. Exploration and apprenticeship learning in reinforcement learning. In Proceedings of the 22nd International Conference on Machine Learning, 2005.
  3. 3.Pieter Abbeel, Adam Coates, Timothy Hunter, and Andrew Y. Ng. Autonomous autorotation of an rc helicopter. In Experimental Robotics, volume 54 of Springer Tracts in Advanced Robotics, pages 385–394. Springer Berlin Heidelberg, 2009.
  4. 4.Pieter Abbeel, Adam Coates, and Andrew Y. Ng. Autonomous helicopter aerobatics through apprenticeship learning. International Journal of Robotic Research, 29(13):1608–1639, 2010.
  5. 5.Naoki Abe, Prem Melville, Cezar Pendus, Chandan K. Reddy, David L. Jensen, Vince P. Thomas, James J. Bennett, Gary F. Anderson, Brent R. Cooley, Melissa Kowalczyk, Mark Domick, and Timothy Gardinier. Optimizing debt collections using constrained reinforcement learning. In Proceedings of the 16th international conference on Knowledge discovery and data mining, pages 75–84, New York, NY, USA, 2010. ACM. ISBN 978-1-4503-0055-1.
  6. 6.Eitan Altman. Asymptotic properties of constrained markov decision processes. Rapport de recherche RR-1598, INRIA, 1992.
  7. 7.Brenna Argall, Sonia Chernova, Manuela Veloso, and Brett Browning. A survey of robot learning from demonstration. Robotics and Autonomous Systems, 57(5):469–483, May 2009. ISSN 09218890.
  8. 8.Drew Bagnell. Learning Decisions: Robustness, Uncertainty, and Approximation. PhD thesis, Robotics Institute, Carnegie Mellon University, Pittsburgh, PA, August 2004.
  9. 9.Drew Bagnell and Jeff Schneider. Robustness and exploration in policy-search based reinforcement learning. In Proceedings of the 25th International Conference on Machine Learning, pages 544–551, New York, NY, USA, 2008. ACM. ISBN 978-1-60558-205-4.
  10. 10.Drew Bagnell, Andrew Ng, and Jeff Schneider. Solving uncertain markov decision problems. Technical report, Robotics Institute Carnegie Mellon, 2001.
  11. 11.A. Baranes and P. Y. Oudeyer. R-IAC: Robust intrinsically motivated exploration and active learning. Autonomous Mental Development, IEEE Transactions on, 1(3):155–169, October 2009. ISSN 1943-0604.
  12. 12.Arnab Basu, Tirthankar Bhattacharyya, and Vivek S. Borkar. A learning algorithm for risk-sensitive cost. Mathematics of Operational Research, 33(4):880–898, 2008.
  13. 13.Vivek S. Borkar. A sensitivity formula for risk-sensitive cost and the actor-critic algorithm. Systems & Control Letters, 44:339–346, 2001.
  14. 14.Vivek S. Borkar. Q-learning for risk-sensitive control. Mathematics of Operations Research, 27(2):294–311, May 2002. ISSN 0364-765X.
  15. 15.Ronen I. Brafman and Moshe Tennenholtz. R-max - a general polynomial time algorithm for near-optimal reinforcement learning. Journal of Machine Learning Research, 3:213–231, March 2003. ISSN 1532-4435.
  16. 16.Andriy Burkov and Brahim Chaib-draa. Reducing the complexity of multiagent reinforcement learning. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems, page 44, 2007.
  17. 17.Pedro Campos and Thibault Langlois. Abalearn: Efficient self-play learning of the game abalone. In INESC-ID, Neural Networks and Signal Processing Group, 2003.
  18. 18.Dotan Di Castro, Aviv Tamar, and Shie Mannor. Policy gradients with variance related risk criteria. In Proceedings of the 29th International Conference on Machine Learning, Edinburgh, Scotland, UK, 2012.
  19. 19.Victor Uc Cetina. Autonomous agent learning using an actor-critic algorithm and behavior models. In Proceedings of the 7th International Conference on Autonomous Agents and Multi-Agent Systems, Estoril, Portugal, pages 1353–1356, 2008.
  20. 20.Suman Chakravorty and David C. Hyland. Minimax reinforcement learning. In Proceedings of the AIAA Guidance, Navigation, and Control Conference and Exhibit, Austin, Texas, USA, 2003.
  21. 21.Yin Chang-Ming, Han-Xing Wang, and Fei Zhao. Risk-sensitive reinforcement learning algorithms with generalized average criterion. Applied Mathematics and Mechanics, 28 (3):405–416, March 2007. ISSN 0253-4827.
  22. 22.Sonia Chernova and Manuela M. Veloso. Interactive policy learning through confidence-based autonomy. Journal of Artificial Intelligence Research, 34:1–25, 2009.
  23. 23.Kun-Jen Chung and Matthew J. Sobel. Discounted mdps: distribution functions and exponential utility maximization. SIAM Journal on Control Optimization, 25(1):49–62, January 1987. ISSN 0363-0129.
  24. 24.Jeffery A. Clouse. On integrating apprentice learning and reinforcement learning. Technical report, Amherst, MA, USA, 1997.
  25. 25.Jeffery A. Clouse and Paul E. Utgoff. A teaching method for reinforcement learning. In ML, pages 92–110. Morgan Kaufmann, 1992. ISBN 1-55860-247-X.
  26. 26.Stefano P. Coraluppi. Optimal control of markov decision processes for performance and robustness. University of Maryland, College Park, Md., 1997.
  27. 27.Stefano P. Coraluppi and Steven I. Marcus. Risk-sensitive and minimax control of discrete-time, finite-state markov decision processes. Automatica, 35:301–309, 1999.
  28. 28.Stefano P. Coraluppi and Steven I. Marcus. Mixed risk-neutral/minimax control of markov decision processes. IEEE Transactions on Automatic Control, 45(3):528–532, 2000.
  29. 29.Erick Delage and Shie Mannor. Percentile optimization for markov decision processes with parameter uncertainty. Operations Research, 58(1):203–213, January 2010.
  30. 30.Kurt Driessens and Saˇso Dˇzeroski. Integrating guidance into relational reinforcement learning. Machine Learning, 57(3):271–304, December 2004. ISSN 0885-6125.
  31. 31.Fernando Fern´andez and Manuela Veloso. Probabilistic policy reuse in a reinforcement learning agent. In Proceedings of the 5th International Joint Conference on Autonomous Agents and Multi-Agent Systems, Hakodate, Japan, May 2006.
  32. 32.Fernando Fern´andez, Javier Garc´ıa, and Manuela M. Veloso. Probabilistic policy reuse for inter-task transfer learning. Robotics and Autonomous Systems, 58(7):866–871, 2010.
  33. 33.Javier Garc´ıa and Fernando Fern´andez. Safe reinforcement learning in high-risk tasks through policy improvement. In Proceedings of the IEEE Symposium on Adaptive Dynamic Programming And Reinforcement Learning, pages 76–83. IEEE, 2011.
  34. 34.Javier Garc´ıa and Fernando Fern´andez. Safe exploration of state and action spaces in reinforcement learning. Journal of Artificial Intelligence Research, 45:515–564, December 2012.
  35. 35.Javier Garc´ıa, Daniel Acera, and Fernando Fern´andez. Safe reinforcement learning through probabilistic policy reuse. In Proceedings of the 1st Multidisciplinary Conference on Reinforcement Learning and Decision Making, October 2013.
  36. 36.Chris Gaskett. Reinforcement learning under circumstances beyond its control. In Proceedings of the International Conference on Computational Intelligence for Modelling Control and Automation, 2003.
  37. 37.Clement Gehring and Doina Precup. Smart exploration in reinforcement learning using absolute temporal difference errors. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems, Saint Paul, MN, USA, pages 1037–1044, 2013.
  38. 38.Peter Geibel. Reinforcement learning for mdps with constraints. In Proceedings of the 17th European Conference on Machine Learning, Berlin, Germany,, volume 4212 of Lecture Notes in Computer Science, pages 646–653. Springer, 2006. ISBN 3-540-45375-X.
  39. 39.Peter Geibel and Fritz Wysotzki. Risk-sensitive reinforcement learning applied to control under constraints. Journal of Artificial Intelligence Research, 24:81–108, 2005.
  40. 40.Alborz Geramifard. Practical Reinforcement Learning using Representation Learning and Safe Exploration for Large Scale Markov Decision Processes. PhD thesis, Massachusetts Institute of Technology, Department of Aeronautics and Astronautics, February 2012.
  41. 41.Alborz Geramifard, Joshua Redding, Nicholas Roy, and Jonathan P. How. UAV cooperative control with stochastic risk models. In Proceedings of the American Control Conference, pages 3393 – 3398, June 2011.
  42. 42.Alborz Geramifard, Joshua Redding, and JonathanP. How. Intelligent cooperative control architecture: A framework for performance improvement using safe learning. Journal of Intelligent & Robotic Systems, 72(1):83–103, 2013. ISSN 0921-0296.
  43. 43.Abhijit Gosavi. Reinforcement learning for model building and variance-penalized control. In Proceedings of the Winter Simulation Conference, pages 373–379. WSC, 2009.
  44. 44.Getachew Hailu and Gerald Sommer. Learning by biasing. In Proceedings of the International Conference on Robotics and Automation, pages 2168–2173. IEEE Computer Society, 1998. ISBN 0-7803-4301-8.
  45. 45.Alexander Hans, Daniel Schneegass, Anton M. Sch¨afer, and Steffen Udluft. Safe Exploration for Reinforcement Learning. In Proceedings of the European Symposium on Artificial Neural Network, pages 143–148, 2008.
  46. 46.Matthias Heger. Risk and reinforcement learning: concepts and dynamic programming. ZKW-Bericht. ZKW, 1994a.
  47. 47.Matthias Heger. Consideration of risk in reinforcement learning. In Proceedings of the 11th International Conference on Machine Learning, pages 105–111, 1994b.
  48. 48.Alfredo Garc´ıa Hern´andez-D´ıaz, Carlos A. Coello Coello, Fatima Perez, Rafael Caballero, Juli´an Molina Luque, and Luis V. Santana-Quintero. Seeding the initial population of a multi-objective evolutionary algorithm using gradient-based information. In Proceedings of the IEEE Congress on Evolutionary Computation, Hong Kong, China, pages 1617–1624, 2008.
  49. 49.Todd Hester and Peter Stone. TEXPLORE: Real-time sample-efficient reinforcement learning for robots. Machine Learning, 90(3), 2013.
  50. 50.Ronald A. Howard and James E. Matheson. Risk-sensitive markov decision processes. Management Science, 18(7):356–369, 1972.
  51. 51.Marcus Hutter. Self-optimizing and pareto-optimal policies in general environments based on bayes-mixtures. In Proceedings of the 15th Annual Conference on Computational Learning Theory, Sydney, Australia, 2002.
  52. 52.Roberto Iglesias, Carlos V. Regueiro, J. Correa, E. Sanchez, and Senen Barro. Improving wall following behaviour in a mobile robot using reinforcement learning. In Proceedings of the International symposium on engineering of intelligent systems, Tenerife (Espa˜na), February 1998a. ISBN 3-906454-12-6.
  53. 53.Roberto Iglesias, Carlos V. Regueiro, J.Correa, and Senen Barro. Supervised reinforcement learning: Application to a wall following behaviour in a mobile robot. In Methodology and tools in knowledge-based systems, pages 300–309, Castellon (Espa˜na), June 1998b. Lecture notes in artificial intelligence 1415. ISBN 3-540-64574-8.
  54. 54.Garud N. Iyengar. Robust dynamic programming. Mathematics of Operations Research, 30:257–280, 2004.
  55. 55.Koen Hermans Jessica Vleugel, Michelle Hoogwout and Imre Gelens. Reinforcement learning with avoidance of unsafe regions. BSc Project, 2011.
  56. 56.Guofei Jiang, Cang-Pu Wu, and George Cybenko. Minimax-based reinforcement learning with state aggregation. In Proceedings of the 37th IEEE Conference on Decision & Control, Tampa, Florida, USA, 1998.
  57. 57.Kshitij Judah, Saikat Roy, Alan Fern, and Thomas G. Dietterich. Reinforcement learning via practice and critique advice. In Proceedings of the 24th AAAI Conference on Artificial Intelligence, Atlanta, Georgia, USA, 2010.
  58. 58.Yoshinobu Kadota, Masami Kurano, and Masami Yasuda. Discounted markov decision processes with utility constraints. Computers & Mathematics with Applications, 51(2): 279–284, 2006.
  59. 59.Hisashi Kashima. Risk-sensitive learning via minimization of empirical conditional value-at-risk. IEICE Transactions, 90-D(12):2043–2052, 2007.
  60. 60.Michael Kearns and Satinder Singh. Near-optimal reinforcement learning in polynomial time. Machine Learning, 49(2-3):209–232, 2002. ISSN 0885-6125.
  61. 61.W. Bradley Knox and Peter Stone. Interactively shaping agents via human reinforcement: the tamer framework. In Proceedings of the 5th International Conference on Knowledge Capture, September 2009.
  62. 62.W. Bradley Knox and Peter Stone. Combining manual feedback with subsequent mdp reward signals for reinforcement learning. In Proceedings of 9th International Conference on Autonomous Agents and Multiagent Systems, May 2010.
  63. 63.W. Bradley Knox, Matthew E. Taylor, and Peter Stone. Understanding human teaching modalities in reinforcement learning environments: A preliminary report. In Proceedings of the Agents Learning Interactively from Human Teachers Workshop, July 2011.
  64. 64.Rogier Koppejan and Shimon Whiteson. Neuroevolutionary reinforcement learning for generalized helicopter control. In Proceedings of the Genetic and Evolutionary Computation Conference, pages 145–152, July 2009.
  65. 65.Rogier Koppejan and Shimon Whiteson. Neuroevolutionary reinforcement learning for generalized control of simulated helicopters. Evolutionary Intelligence, 4:219–241, 2011.
  66. 66.Gregory Kuhlmann, Peter Stone, Raymond J. Mooney, and Jude W. Shavlik. Guiding a reinforcement learner with natural language advice: Initial results in robocup soccer. In Proceedings of the AAAI-2004 Workshop on Supervisory Control of Learning and Adaptive Systems, July 2004.
  67. 67.Edith L.M. Law. Risk-directed exploration in reinforcement learning. McGill University, 2005.
  68. 68.Long Ji Lin. Programming robots using reinforcement learning and teaching. In Proceedings of the 9th National Conference on Artificial Intelligence, Anaheim, CA, USA, July 14-19, 1991, Volume 2, pages 781–786, 1991.
  69. 69.Long-Ji Lin. Self-improving reactive agents based on reinforcement learning, planning and teaching. Machine Learning, 8(3–4):293–321, 1992.
  70. 70.Yaxin Liu, Richard Goodwin, and Sven Koenig. Risk-averse auction agents. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems, pages 353–360. ACM, 2003. ISBN 1-58113-683-8.
  71. 71.David G. Luenberger. Investment science. Oxford University Press, Incorporated, 2013.
  72. 72.Richard Maclin and Jude W. Shavlik. Creating advice-taking reinforcement learners. Machine Learning, 22(1-3):251–281, 1996. doi: 10.1023/A:1018020625251.
  73. 73.Richard Maclin, Jude Shavlik, Lisa Torrey, Trevor Walker, and Edward Wild. Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression. In Proceedings of the 20th National Conference on Artificial Intelligence, 2005a.
  74. 74.Richard Maclin, Jude Shavlik, Trevor Walker, and Lisa Torrey. Knowledge-based support-vector regression for reinforcement learning. In Proceedings of the IJCAI’05 Workshop on Reasoning, Representation, and Learning in Computer Games, 2005b.
  75. 75.Frederic Maire. Apprenticeship learning for initial value functions in reinforcement learning. In Proceedings of the IJCAI’05 Workshop on Planning and Learning in A Priori Unknown or Dynamic Domains, pages 23–28, 2005.
  76. 76.Shie Mannor and John N. Tsitsiklis. Mean-variance optimization in markov decision processes. In Proceedings of the 28th International Conference on Machine Learning, Bellevue, Washington, USA, pages 177–184, 2011.
  77. 77.Harry Markowitz. Portfolio selection. In Journal of Finance, volume 7, pages 77–91, 1952.
  78. 78.Jos´e Antonio Mart´ın H. and Javier Lope. Learning autonomous helicopter flight with evolutionary reinforcement learning. In Proceedings of the 12th International Conference on Computer Aided Systems Theory, pages 75–82, 2009. ISBN 978-3-642-04771-8.
  79. 79.Helmut Mausser and Dan Rosen. Beyond var: From measuring risk to managing risk. ALGO Research Quarterly, 1(2):5–20, 1998.
  80. 80.John Mccarthy. Programs with common sense. In Semantic Information Processing, pages 403–418. MIT Press, 1959.
  81. 81.Oliver Mihatsch and Ralph Neuneier. Risk-sensitive reinforcement learning. Machine Learning, 49(2-3):267–290, 2002. ISSN 0885-6125.
  82. 82.Teodor Mihai Moldovan and Pieter Abbeel. Safe exploration in markov decision processes. In Proceedings of NIPS Workshop on Bayesian Optimization, Experimental Design and Bandits, 2011.
  83. 83.Teodor Mihai Moldovan and Pieter Abbeel. Safe exploration in markov decision processes. In Proceedings of the 29th International Conference on Machine Learning, Edinburgh, Scotland, UK, 2012a.
  84. 84.Teodor Mihai Moldovan and Pieter Abbeel. Risk aversion in markov decision processes via near optimal chernoff bounds. In Advances in Neural Information Processing Systems 25, Lake Tahoe, Nevada, United States., pages 3140–3148, 2012b.
  85. 85.David L. Moreno, Carlos V. Regueiro, Roberto Iglesias, and Senen Barro. Using prior knowledge to improve reinforcement learning in mobile robotics. In Proceedings fo the Conference Towards Autonomous Robotics Systems, Bath (Reino Unido), September 2004.
  86. 86.Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka. Nonparametric return distribution approximation for reinforcement learning. In Proceedings of the 27th International Conference on Machine Learning, pages 799–806, 2010a.
  87. 87.Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka. Parametric return density estimation for reinforcement learning. In Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, pages 368–375, Catalina Island, California, USA, Jul. 8–11 2010b.
  88. 88.Arnab Nilim and Laurent El Ghaoui. Robust control of markov decision processes with uncertain transition matrices. Operational Research, 53(5):780–798, September 2005. ISSN 0030-364X.
  89. 89.Ali Nouri. Efficient Model-Based Exploration in Continuous State-Space Environments. PhD thesis, New Brunswick, NJ, USA, 2011. AAI3444957.
  90. 90.Ali Nouri and Michael L. Littman. Multi-resolution exploration in continuous spaces. In Advances in Neural Information Processing Systems 21, pages 1209–1216, 2008.
  91. 91.Takayuki Osogami. Robustness and risk-sensitivity in markov decision processes. In Advances in Neural Information Processing Systems 25, Lake Tahoe, Nevada, United States, pages 233–241, 2012.
  92. 92.Stephen D. Patek. On terminating markov decision processes with a risk-averse objective function. Automatica, 37(9):1379–1386, 2001.
  93. 93.Frederick Philip Klahr Hayes-Roth and David J. Mostow. Advice-taking and knowledge refinement: An iterative view of skill acquisition. Cognitive Skills and Their Acquisition, 1981.
  94. 94.Sameera S. Ponda, Luke B. Johnson, and Jonathan P. How. Risk allocation strategies for distributed chance-constrained task allocation. In American Control Conference, June 2013.
  95. 95.Martin L. Putterman. Markov decision processes: Discrete stochastic dynamic programming. Jhon Wiley & Sons, Inc, 1994.
  96. 96.Michael T. Rosenstein and Andrew G. Barto. Supervised learning combined with an actor-critic architecture. Technical report, Amherst, MA, USA, 2002.
  97. 97.Michael T. Rosenstein and Andrew G. Barto. Supervised actor-critic reinforcement learning. Wiley-IEEE Press, 2004.
  98. 98.Daniil Ryabko and Marcus Hutter. Theorical Computer Science, (3):274–284.
  99. 99.Makoto Sato, Hajime Kimura, and Shigenobu Kobayashi. TD algorithm for the variance of return and mean-variance reinforcement learning. Transactions of the Japanese Society for Artificial Intelligence, 16:353–362, 2002.
  100. 100.Nils T. Siebel and Gerald Sommer. Evolutionary reinforcement learning of artificial neural networks. International Journal of Hybrid Intelligent Systems, 4:171–183, August 2007. ISSN 1448-5869.
  101. 101.William D. Smart and Leslie Pack Kaelbling. Practical reinforcement learning in continuous spaces. In Artificial Intelligence, pages 903–910. Morgan Kaufmann, 2000.
  102. 102.Alice Smith, Alice E. Smith, David W. Coit, Thomas Baeck, David Fogel, and Zbigniew Michalewicz. Penalty functions. Oxford University Press and Institute of Physics Publishing, 1997.
  103. 103.Yong Song, Yi bin Li, Cai hong Li, and Gui fang Zhang. An efficient initialization approach of q-learning for mobile robots. International Journal of Control, Automation and Systems, 10(1):166–172, 2012.
  104. 104.Alexander L. Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L. Littman. PAC model-free reinforcement learning. In Proceedings of the 23rd International Conference on Machine Learning, ICML ’06, pages 881–888, New York, NY, USA, 2006. ACM. ISBN 1-59593-383-2.
  105. 105.Halit B. Suay and Sonia Chernova. Effect of human guidance and state space size on interactive reinforcement learning. In Proceedings of the IEEE International Symposium on Robot and Human Interactive Communication, pages 1–6. IEEE, July 2011. ISBN 978-1-4577-1571-6.
  106. 106.Richard S. Sutton and Andrew G. Barto. Reinforcement learning: an introduction. The MIT Press, March 1998. ISBN 0262193981.
  107. 107.Giorgio Szeg¨o. Measures of risk. European Journal of Operational Research, 163(1):5–19, 2005.
  108. 108.Hamdy A. Taha. Operations research: an introduction. Number 1. Macmillan Publishing Company, 1992. ISBN 9780024189752.
  109. 109.Aviv Tamar, Huan Xu, and Shie Mannor. Scaling Up Robust MDPs by Reinforcement Learning. Computing Research Repository, abs/1306.6189, 2013.
  110. 110.Jie Tang, Arjun Singh, Nimbus Goehausen, and Pieter Abbeel. Parameterized maneuver learning for autonomous helicopter flight. In International Conference on Robotics and Automation, 2010.
  111. 111.Matthew E. Taylor and Peter Stone. Representation transfer for reinforcement learning. In Fall Symposium on Computational Approaches to Representation Change during Learning and Development, November 2007.
  112. 112.Matthew E. Taylor and Peter Stone. Transfer learning for reinforcement learning domains: A survey. Journal of Machine Learning Research, 10(1):1633–1685, 2009.
  113. 113.Matthew E. Taylor, Peter Stone, and Yaxin Liu. Transfer learning via inter-task mappings for temporal difference learning. Journal of Machine Learning Research, 8(1):2125–2167, 2007.
  114. 114.Andrea L. Thomaz and Cynthia Breazeal. Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance. In Proceedings of the 21st National Conference on Artificial Intelligence, AAAI’06, pages 1000–1005. AAAI Press, 2006. ISBN 978-1-57735-281-5.
  115. 115.Andrea Lockerd Thomaz and Cynthia Breazeal. Teachable robots: Understanding human teaching behavior to build more effective robot learners. Artificial Intelligence, 172(6-7): 716–737, 2008.
  116. 116.Lisa Torrey and Matthew E. Taylor. Help an agent out: Student/teacher learning in sequential decision tasks. In Proceedings of the AAMAS Workshop Adaptive and Learning Agents, June 2012.
  117. 117.Lisa Torrey, Trevor Walker, Jude Shavlik, and Richard Maclin. Using advice to transfer knowledge acquired in one reinforcement learning task to another. Machine Learning: ECML 2005, pages 412–424, 2005.
  118. 118.Paul E. Utgoff and Jeffrey A. Clouse. Two kinds of training information for evaluation function learning. In Proceedings of the 9th National Conference on Artificial Intelligence, Anaheim, CA, USA, July 14-19, 1991, Volume 2, pages 596–600, 1991.
  119. 119.Pablo Quint´ıa Vidal, Roberto Iglesias Rodr´ıguez, Miguel Rodr´ıguez Gonz´alez, and Carlos V´azquez Regueiro. Learning on real robots from experience and simple user feedback. Journal of Physical Agents, 7(1), 2013. ISSN 1888-0258.
  120. 120.Pradyot Korupolu VN and Balaraman Ravindran. Beyond rewards: Learning from richer supervision. In Proceedings of the 9th European Workshop on Reinforcement Learning, Athens Greece, September 2011.
  121. 121.Thomas J. Walsh, Daniel Hewlett, and Clayton T. Morrison. Blending autonomous exploration and apprenticeship learning. In Proceedings of the Conference Advances in Neural Information Processing Systems 24, Granada, Spain, pages 2258–2266, 2011.
  122. 122.Christopher Watkins. Learning from Delayed Rewards. PhD thesis, King’s College, Cambridge, UK, May 1989.
  123. 123.Kemin Zhou, John C. Doyle, and Keith Glover. Robust and Optimal Control. Prentice-Hall, Inc., Upper Saddle River, NJ, USA, 1996. ISBN 0-13-456567-3.

Citation

MLA
García, J., and F. Fernández. “A Comprehensive Survey on Safe Reinforcement Learning”. Journal of Machine Learning Research, vol. 16, no. 42, 2015, pp. 1437–80, https://www.jmlr.org/papers/v16/garcia15a.html.
APA
García, J., & Fernández, F. (2015). A Comprehensive Survey on Safe Reinforcement Learning. Journal of Machine Learning Research, 16(42), 1437–1480. https://www.jmlr.org/papers/v16/garcia15a.html
Chicago
García, J., and F. Fernández. 2015. “A Comprehensive Survey on Safe Reinforcement Learning”. Journal of Machine Learning Research 16 (42): 1437–80. https://www.jmlr.org/papers/v16/garcia15a.html.
Harvard
García, J. and Fernández, F. (2015) “A Comprehensive Survey on Safe Reinforcement Learning”, Journal of Machine Learning Research, 16(42), pp. 1437–1480. Available at: https://www.jmlr.org/papers/v16/garcia15a.html.
Vancouver
1. García J, Fernández F (2015) A Comprehensive Survey on Safe Reinforcement Learning. Journal of Machine Learning Research 16:1437–1480

BibTeX

@article{JMLR:v16:garcia15a,
  author  = {Javier García and Fernando Fernández},
  title   = {A Comprehensive Survey on Safe Reinforcement Learning},
  journal = {Journal of Machine Learning Research},
  year    = {2015},
  volume  = {16},
  number  = {42},
  pages   = {1437--1480},
  url     = {http://jmlr.org/papers/v16/garcia15a.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF