Deep Reinforcement Learning for Autonomous Driving: A Survey

Bangalore Ravi KiranIbrahim SobhVictor TalpaertPatrick MannionAhmad A. Al SallabSenthil YogamaniPatrick Pérez

article2020IEEE transactions on intelligent transportation systems (Print)2,483 citations

Systematizes deep reinforcement learning techniques across core autonomous driving tasks, addressing key implementation hurdles from simulation and imitation learning to real-world validation.

Listen

Autonomous driving requires vehicles to handle sequential decision-making in complex, unpredictable, and highly dynamic environments. While traditional supervised machine learning has achieved high accuracy in perception tasks like object detection, it struggles with planning, real-time decision-making, and low-level vehicle control where an agent's actions directly alter future sensory inputs and environment states. As a result, deep reinforcement learningan approach where an autonomous agent learns optimal actions through trial-and-error interactions with its environmenthas emerged as a vital technology for self-driving systems.

The article systematically evaluates the core theoretical foundations, practical applications, and operational challenges of applying deep reinforcement learning to autonomous driving. By reviewing key algorithmic paradigms, simulation platforms, adjacent techniques like imitation learning, and real-world trials, the article establishes a comprehensive taxonomy of automated driving tasks and identifies critical barriers to production deployment.

The analysis reveals several key findings across algorithmic performance, system design, and deployment feasibility. First, continuous-action algorithms consistently produce smoother, more stable vehicle trajectories than discrete-action methods, though they require longer training times and stricter operational boundaries. Second, pure reinforcement learning suffers from severe sample inefficiency, requiring millions of interactions that make training directly on real-world vehicles hazardous and economically impractical. Third, bridging the simulation-to-reality gap using domain adaptation and randomized simulator dynamics enables models trained in synthetic environments to transfer effectively to physical vehicles, cutting real-world sample requirements by up to 50 times. Fourth, hybrid learning strategiesspecifically bootstrapping reinforcement learning with expert human demonstrationsresolve critical exploration bottlenecks that neither pure imitation learning nor pure reinforcement learning can solve alone.

These findings have direct strategic implications for automotive safety, system performance, development budgets, and testing timelines. Pure imitation learning remains brittle when encountering rare or unseen edge cases, creating substantial safety risks. Conversely, unconstrained reinforcement learning exploration in real environments is unacceptable for road safety. Integrating safety-based supervisory control layers, multi-fidelity simulation pipelines, and multi-agent coordination frameworks allows engineering teams to train and validate complex policies safely while significantly reducing physical testing costs.

Organizations developing autonomous driving systems should adopt hybrid development architectures. Technical leaders should initialize decision policies using human demonstrations, refine them inside high-fidelity multi-agent simulators using principled reward shaping, and transfer them to hardware via robust domain adaptation. Crucially, autonomous systems must pair learned policies with hard safety constraints and negative-avoidance functions to prevent hazardous maneuvers. Future work must prioritize standardizing validation benchmarks, improving multi-agent interaction modeling, and advancing sample-efficient meta-learning before full real-world deployment can occur.

While confidence in deep reinforcement learning for simulated and structured driving tasks is high, significant limitations remain regarding reproducibility, sensitivity to hyper-parameters, and behavioral verification in unstructured, real-world edge cases. Decision-makers should maintain cautious oversight, validating all reinforcement learning policies against rigorous simulated adversarial testbeds and standardized baseline frameworks prior to public road trials.

arXiv: 2002.00444
  • Paper: Playing Atari with Deep Reinforcement Learning, Volodymyr Mnih et al. (2013). This foundational paper introduces the Deep Q-Network algorithm, providing the core value-based methodology essential for understanding deep reinforcement learning in complex environments.
  • Paper: Human-level control through deep reinforcement learning, Volodymyr Mnih et al. (2015). This seminal study establishes human-level control from raw pixels, serving as the primary benchmark and algorithmic baseline for modern deep reinforcement learning surveys.
  • Paper: Trust Region Policy Optimization, John Schulman et al. (2015). This paper presents Trust Region Policy Optimization, providing the theoretical guarantees and policy gradient foundations needed to understand stable continuous control.
  • Paper: Proximal Policy Optimization Algorithms, John Schulman et al. (2017). This work introduces Proximal Policy Optimization, an algorithm widely adopted for stabilizing policy updates and foundational for modern reinforcement learning workflows.
  • Paper: Continuous control with deep reinforcement learning, T. Lillicrap et al. (2015). This paper establishes Deep Deterministic Policy Gradient algorithms, directly enabling continuous control in robotics and driving simulations discussed in the survey.
Cover for Deep Reinforcement Learning for Autonomous Driving: A Survey

Abstract

With the development of deep representation learning, the domain of reinforcement learning (RL) has become a powerful learning framework now capable of learning complex policies in high dimensional environments. This review summarises deep reinforcement learning (DRL) algorithms and provides a taxonomy of automated driving tasks where (D)RL methods have been employed, while addressing key computational challenges in real world deployment of autonomous driving agents. It also delineates adjacent domains such as behavior cloning, imitation learning, inverse reinforcement learning that are related but are not classical RL algorithms. The role of simulators in training agents, methods to validate, test and robustify existing solutions in RL are discussed.

Table of Contents

  • I Introduction
  • II Components of AD System
  • II-A Scene Understanding
  • II-B Localization and Mapping
  • II-C Planning and Driving policy
  • II-D Control
  • III Reinforcement learning
  • III-A Value-based methods
  • III-B Policy-based methods
  • III-C Actor-critic methods
  • III-D Model-based (vs. Model-free) & On/Off Policy methods
  • III-E Deep reinforcement learning (DRL)
  • IV Extensions to reinforcement learning
  • IV-A Reward shaping
  • IV-B Multi-agent reinforcement learning (MARL)
  • IV-C Multi-objective reinforcement learning
  • IV-D State Representation Learning (SRL)
  • IV-E Learning from Demonstrations
  • V Reinforcement learning for Autonomous driving tasks
  • V-A State Spaces, Action Spaces and Rewards
  • V-B Motion Planning & Trajectory optimization
  • V-C Simulator & Scenario generation tools
  • V-D LfD and IRL for AD applications
  • VI Real world challenges and future perspectives
  • VI-A Validating RL systems
  • VI-B Bridging the simulation-reality gap
  • VI-C Sample efficiency
  • VI-D Exploration issues with Imitation
  • VI-E Intrinsic Reward functions
  • VI-F Incorporating safety in DRL
  • VI-G Multi-agent reinforcement learning
  • VII Conclusion
  • References

Knowls

  1. Knowl 1 — Taxonomy of Autonomous Driving Tasks Formulated with Deep Reinforcement Learning

    data/table

    Autonomous driving tasks can be categorized by the specific driving behavior to be learned, the corresponding reinforcement learning (RL) or deep reinforcement learning (DRL) algorithm applied, and the associated algorithmic tradeoffs:

    AD Task DRL Method Description Improvements Tradeoffs
    Lane Keeping Discrete actions via Deep Q-Networks (DQN) or continuous actions via Deep Deterministic Actor-Critic (DDAC) in driving simulators (e.g., TORCS) to follow lanes and maximize average velocity. Continuous action spaces yield smoother vehicle trajectories but can induce slower convergence and more restrictive termination conditions. Removing experience replay in DQNs can accelerate convergence, whereas discrete one-hot action spaces result in abrupt steering.
    Lane Changing Q-learning to select discrete tactical maneuvers: maintain lane (no-op), change left/right, accelerate, or decelerate. Increases robustness over classical rule-based trajectories with pre-computed fixed waypoints, velocity profiles, and fixed curvature.
    Ramp Merging Recurrent neural architectures (e.g., LSTMs) combined with DRL to model long-term temporal dependencies when merging into highway traffic. Retaining past history of surrounding traffic states enables safer and more robust merging decisions under dense traffic.
    Overtaking Multi-goal RL via Q-learning or Double Action Q-Learning (DAQL) determining maneuvers based on dynamic interaction with adjacent vehicles. Improves overtaking efficiency and lane-keeping speed while avoiding collisions.
    Intersections DQN evaluating state-action values under occluded or unsignalized intersection scenarios with defined "Creep-Go" actions. "Creep-Go" action abstractions enable vehicles to inch forward into restricted visibility zones and negotiate turns safely.
    Motion Planning Deep neural networks learning heuristic guidance functions integrated into the AA^* search algorithm over obstacle map representations. Produces smoother control behaviors and superior path planning performance compared to standard multi-step DQN formulations.
  2. Knowl 2 — State, Action, and Reward Design Spaces for Reinforcement Learning in Autonomous Driving

    model/method

    Formulating autonomous driving within a Markov Decision Process (MDP) or Partially Observable MDP (POMDP) framework requires defining domain-specific state spaces, action spaces, and reward functions:

    1. State Space Representations:

      • Ego & Object Dynamics: Continuous kinematic state vectors comprising ego-vehicle pose (x,y,θ)(x, y, \theta), velocity, longitudinal metrics such as Time-to-Collision (TTC), and bounding boxes or trajectories of surrounding obstacles.
      • Occupancy Grids & Bird's-Eye View (BEV): 2D Cartesian or polar top-down occupancy grids encoding drivable areas, lane boundaries, past/future vehicle trajectory rollouts, and semantic overlays (e.g., traffic light states). BEV representations provide sensor-agnostic mid-level abstractions that preserve spatial metric relationships without the dimensional instability of raw perception inputs.
      • Raw Sensor Feeds: Multi-camera images, LiDAR point clouds, or radar returns processed via convolutional or recurrent networks.
    2. Action Space Representations:

      • Continuous Actuation: Low-level direct control vectors [asteer,athrottle,abrake]TR3[a_{\text{steer}}, a_{\text{throttle}}, a_{\text{brake}}]^T \in \mathbb{R}^3 executed via policy-gradient or actor-critic architectures (such as DDPG or SAC).
      • Discretized Actions: Uniform binning or log-scale discretization (concentrating resolution around small steering angles near zero) to accommodate value-based methods (such as DQN). Large discretization bins cause jerky actuation, whereas fine discretizations inflate action-space dimensionality.
      • Temporal Abstractions / Options: High-level discrete meta-actions (e.g., "follow lane", "yield", "overtake", "creep") sustained across multiple time steps before selecting a subsequent macro-action.
    3. Reward Function Formulations: Composite reward signals rtr_t balance multiple competing objectives: rt=wprogrprogress+wvelrvelocitywcolrcollisionwinfrinfractionwjerkrcomfortr_t = w_{\text{prog}} r_{\text{progress}} + w_{\text{vel}} r_{\text{velocity}} - w_{\text{col}} r_{\text{collision}} - w_{\text{inf}} r_{\text{infraction}} - w_{\text{jerk}} r_{\text{comfort}} where components evaluate forward distance progressed along the global route (rprogressr_{\text{progress}}), adherence to the target speed limit (rvelocityr_{\text{velocity}}), severe penalties for colliding with obstacles or pedestrians (rcollisionr_{\text{collision}}), penalties for sidewalk incursions or red light violations (rinfractionr_{\text{infraction}}), and regularization penalizing abrupt acceleration, steering jerk, or vehicle instability (rcomfortr_{\text{comfort}}).

  3. Knowl 3 — Simulation Environments for Autonomous Driving Reinforcement Learning

    data/table

    Simulation platforms provide the primary testbeds for training and evaluating reinforcement learning policies in autonomous driving prior to physical deployment:

    Simulator Description and Capabilities
    CARLA Open-source urban driving simulator supporting flexible sensor suites (RGB cameras, LiDAR point clouds, depth maps, semantic segmentation) and environmental controls for weather, lighting, dynamic actors, and traffic rule benchmark evaluation.
    TORCS Open racing car simulator with 3D physics and visual streams, widely used for benchmarking continuous and discrete vehicle control and trajectory tracking algorithms.
    AirSim High-fidelity simulation platform built on Unreal Engine providing physically and visually realistic environments with camera, depth, and semantic streams for autonomous ground vehicles and drones.
    Gazebo (ROS) Multi-robot rigid-body physics simulator integrated with the Robot Operating System (ROS), employed for kinematics, 2D/3D map navigation, and vehicle control validation.
    SUMO Microscopic and macroscopic traffic simulator modeling city-scale road networks, multi-lane dynamics, and multi-vehicle traffic flow optimization.
    DeepDrive Unreal Engine-based driving simulator providing multi-camera (up to eight views) sensor streaming with synchronized depth buffers for end-to-end RL control.
    NVIDIA DRIVE Constellation Proprietary hardware-in-the-loop simulation platform delivering photorealistic sensor simulation (camera, LiDAR, radar) for automotive validation.
    MADRaS Multi-agent autonomous driving simulation platform built upon TORCS designed for evaluating multi-vehicle competitive and cooperative interactions.
    Flow Computational framework integrating SUMO and Aimsun microscopic traffic simulators with deep RL libraries for macro-scale autonomous traffic control.
    Highway-env OpenAI Gym-compatible environment providing minimal, fast 2D kinematic simulation for highway tactical decision-making, lane changes, and ramp merges.
    Carcraft Proprietary multi-agent simulation framework utilized by Waymo for large-scale scenario testing and virtual policy rollouts.
  4. Knowl 4 — Bridging the Reality Gap in Autonomous Driving Reinforcement Learning

    model/method

    Direct deployment of driving policies trained in simulation to physical vehicles suffers from performance degradation due to discrepancies in visual appearance, sensor noise distributions, and vehicle dynamics (the sim-to-real gap). Three main paradigms address this transfer problem:

    1. Dynamics Randomization: Perturbing physical parameters during simulation training (such as surface friction, vehicle mass, suspension stiffness, actuator latency, and steering calibration). This forces the RL policy to be robust across a distribution of transition dynamics, enabling transfer to real vehicles without requiring retraining on physical hardware.

    2. Pixel-Level and Feature-Level Domain Adaptation: Generative adversarial domain translation frameworks map visual modalities between domains:

      • Sim-to-Real Translation: Synthetic simulator renderings are mapped to photorealistic styles matching real-world automotive camera distributions before feeding the visual encoder of the driving policy.
      • Real-to-Sim Adaptation ("Virtual Goggles"): Real camera streams captured during physical vehicle operation are mapped into the synthetic domain of the simulator, projecting unfamiliar physical sensor inputs back into the representation space where the driving policy was trained.
    3. Multi-Fidelity Reinforcement Learning (MFRL): A cascade of simulators with progressively increasing physical and sensory fidelity (and corresponding computational expense) is used to guide exploration. Low-fidelity models establish coarse baseline policies and heuristics, while high-fidelity simulation and sparse real-world rollouts fine-tune optimal driving policies with minimal costly real-world vehicle interactions.

  5. Knowl 5 — Methods for Sample Efficiency Enhancement in Driving RL

    model/method

    Training deep reinforcement learning agents directly in driving environments requires an impractical number of state transitions due to delayed returns and the sparse nature of critical safety events. Sample efficiency is enhanced via the following strategies:

    1. Imitation Learning Bootstrapping: An initial driving policy is trained offline via supervised imitation learning (or behavior cloning) using human driving demonstration datasets. The pre-trained network weights serve as the policy initialization for subsequent online DRL fine-tuning through active environment interaction.

    2. Specialized Experience Replay:

      • Prioritized Experience Replay (PER): Transitions with higher temporal difference (TD) error δt\delta_t are sampled with higher probability from the replay memory, prioritizing unexpected or high-information events.
      • Dual-Bucket / Trauma Memory: Segregating experience memory into separate storage pools for nominal driving vs. rare safety-critical events (e.g., collisions, near-misses, or lane departures). Sampling fixed proportions from both buckets ensures regular gradient updates on dangerous edge cases.
    3. Latent State Representation Learning (SRL) and World Models: High-dimensional visual observations oto_t are compressed into low-dimensional latent state vectors sts_t via unsupervised models (such as Variational Autoencoders or predictive recurrent transition models). Driving policies π(atst)\pi(a_t | s_t) are learned over the compact latent space, reducing sample complexity and isolating policy training from raw visual noise.

    4. Meta-Learning and Policy Composition: Model-Agnostic Meta-Learning (MAML) and recurrent meta-RL architectures optimize network parameter initializations that adapt rapidly to new driving conditions (e.g., changing weather, road geometry, or regional traffic customs) within few fine-tuning iterations.

  6. Knowl 6 — Safety Mechanisms and Constrained Decision-Making in Autonomous Driving DRL

    model/method

    Deploying unconstrained deep reinforcement learning models for physical driving risks catastrophic failures during exploration and policy updates. Safety-integrated RL architectures combine constrained optimization with classical safety controllers:

    1. Constrained MDPs and Survival-Oriented RL (SORL): Driving is formulated as a Constrained Markov Decision Process where survival and constraint satisfaction override unconstrained reward maximization: maxπEπ[t=0Hγtr(st,at)]subject toEπ[t=0Hck(st,at)]dk,k\max_\pi \mathbb{E}_\pi \left[ \sum_{t=0}^H \gamma^t r(s_t, a_t) \right] \quad \text{subject to} \quad \mathbb{E}_\pi \left[ \sum_{t=0}^H c_k(s_t, a_t) \right] \le d_k, \quad \forall k where ck(st,at)c_k(s_t, a_t) denotes cost functions for critical safety infractions (such as collisions or boundary violations) and dkd_k is the tolerated violation budget. Negative-avoidance functions update policies directly from failure boundaries to prevent recurrence of hazardous states.

    2. Hierarchical Control and Hybrid Safety Overrides: Decoupling the driving pipeline into a high-level DRL decision policy (optimizing passenger comfort, route efficiency, and tactical maneuvers) and a low-level deterministic safety verification layer. If the RL policy selects an action leading to imminent hazard, an analytical override—such as Artificial Potential Fields (APF) or Model Predictive Control (MPC) barrier certificates—takes precedence to ensure collision avoidance.

    3. Safe DAgger: An auxiliary safety network predicts whether the primary learned policy will deviate dangerously from an expert reference trajectory based on partial state observations. The safety network acts as an automatic query trigger, transferring vehicle actuation back to a reference supervisor only when deviation probability exceeds a defined threshold.

  7. Knowl 7 — Distribution Shift and Compounding Errors in Imitation Learning for Autonomous Driving

    limitation

    Behavior Cloning (BC) and standard Learning from Demonstrations (LfD) train driving policies by minimizing supervised loss on static expert datasets: LBC(θ)=E(s,a)Dexpert[(πθ(s),a)]\mathcal{L}_{\text{BC}}(\theta) = \mathbb{E}_{(s, a) \sim \mathcal{D}_{\text{expert}}} \left[ \ell(\pi_\theta(s), a) \right] This paradigm suffers from fundamental operational failures when deployed in closed-loop autonomous driving systems:

    1. Violation of the i.i.d. Assumption and Error Cascades: Standard supervised learning assumes independent and identically distributed training samples. In autonomous driving, the agent's action ata_t determines the subsequent state st+1s_{t+1}. Minor steering prediction errors drive the ego-vehicle into off-trajectory states not represented in the expert dataset Dexpert\mathcal{D}_{\text{expert}}. Once out of distribution, the policy makes larger errors, leading to compounding drift and inevitable lane departure or collision.

    2. Data Coverage Bottleneck: Static expert demonstration datasets lack negative recovery examples (how to steer back into lane from near-curb angles or avoid collisions after an evasion maneuver), because human experts rarely traverse failure trajectories. Scaling dataset size alone is insufficient: even tens of millions of real-world state-action samples fail to cover the tail distribution of rare perturbations.

    3. Interactive Data Aggregation Solutions: Mitigating this shift requires online interactive training algorithms (such as DAgger or SMILE), where the learned policy executes rollouts in simulation, encounters perturbation states, queries an expert for corrective actions at those novel states, and aggregates them into the training set iteratively. Alternatively, synthetic perturbation injection (synthesizing off-center camera views with corrective steering labels) must be artificially added to the demonstration distribution.

  8. Knowl 8 — Multi-Agent Reinforcement Learning Formulation for Autonomous Driving

    model/method

    Autonomous driving involves concurrent interactions among multiple non-stationary entities (ego-vehicle, surrounding vehicles, pedestrians, cyclists), which invalidates single-agent Markov decision process assumptions. Multi-Agent Reinforcement Learning (MARL) formalizes this environment as a Stochastic Game (Markov Game):

    N,S,{Ai}i=1N,T,{Ri}i=1N\langle N, \mathcal{S}, \{\mathcal{A}_i\}_{i=1}^N, \mathcal{T}, \{\mathcal{R}_i\}_{i=1}^N \rangle where:

    • NN is the number of active traffic participants.
    • S\mathcal{S} is the global environment state space.
    • Ai\mathcal{A}_i is the discrete or continuous action set available to agent ii, with the joint action defined as a=(a1,,aN)A=i=1NAi\mathbf{a} = (a_1, \dots, a_N) \in \mathcal{A} = \prod_{i=1}^N \mathcal{A}_i.
    • T:S×A×S[0,1]\mathcal{T}: \mathcal{S} \times \mathcal{A} \times \mathcal{S} \to [0, 1] is the state transition probability function conditioned on the joint action a\mathbf{a}.
    • Ri:S×AR\mathcal{R}_i: \mathcal{S} \times \mathcal{A} \to \mathbb{R} is the reward function specific to agent ii.
    • siSis_i \in \mathcal{S}_i denotes the local observation of agent ii, reflecting partial observability in real traffic.

    MARL architectures in autonomous driving address two distinct functional roles:

    1. Cooperative Tactical Coordination: Vehicles in dense traffic environments (e.g., unsignalized intersection crossings, cooperative highway lane merging, or platoon management) optimize shared objectives to maximize collective traffic throughput and avoid gridlocks without centralized traffic lights.
    2. Adversarial Policy Stress-Testing: Surrounding vehicles are trained as adversarial RL agents whose explicit objective is to discover rare edge-case driving maneuvers, unexpected pedestrian incursions, or non-compliant road behaviors that cause the ego-vehicle's driving policy to fail.
  9. Knowl 9 — Multi-Objective Reinforcement Learning Framework for Autonomous Vehicle Decision-Making

    model/method

    Autonomous vehicle decision-making requires navigating trade-offs among mutually conflicting criteria—such as trip efficiency, rider comfort, collision avoidance, and energy consumption. Multi-Objective Reinforcement Learning (MORL) formalizes this through Multi-Objective Markov Decision Processes (MOMDPs).

    In a MOMDP, the scalar reward function is replaced by a vector-valued reward function: R(s,a)=[R1(s,a),R2(s,a),,Rm(s,a)]TRm\mathbf{R}(s, a) = [R_1(s, a), R_2(s, a), \dots, R_m(s, a)]^T \in \mathbb{R}^m where each component Rc(s,a)R_c(s, a) corresponds to an individual objective c{1,,m}c \in \{1, \dots, m\} (e.g., RprogressR_{\text{progress}}, RsafetyR_{\text{safety}}, RcomfortR_{\text{comfort}}, RefficiencyR_{\text{efficiency}}).

    The expected vector return for a policy π\pi is defined as: Vπ(s)=Eπ[k=0H1γkrk+1|s0=s]\mathbf{V}^\pi(s) = \mathbb{E}_\pi \left[ \sum_{k=0}^{H-1} \gamma^k \mathbf{r}_{k+1} \,\middle|\, s_0 = s \right] where γ[0,1]\gamma \in [0, 1] is the discount factor and HH is the horizon length.

    Because no single policy typically maximizes all mm objectives simultaneously, solutions are characterized by Pareto dominance:

    • A policy π1\pi_1 dominates π2\pi_2 (denoted Vπ1Vπ2\mathbf{V}^{\pi_1} \succ \mathbf{V}^{\pi_2}) if c{1,,m},Vcπ1(s)Vcπ2(s)\forall c \in \{1, \dots, m\}, V_c^{\pi_1}(s) \ge V_c^{\pi_2}(s) and c\exists c' such that Vcπ1(s)>Vcπ2(s)V_{c'}^{\pi_1}(s) > V_{c'}^{\pi_2}(s).
    • MORL algorithms learn or approximate the Pareto front (the set of non-dominated policies), allowing runtime driving style adaptation (e.g., selecting conservative vs. agile driving modes) by varying preference weights wRm\mathbf{w} \in \mathbb{R}^m applied to the objective vector.
  10. Knowl 10 — Validation Challenges and Hyperparameter Sensitivity in Continuous Control DRL

    limitation

    Empirical validation and reproducibility of deep reinforcement learning policies for continuous vehicle control (e.g., steering and acceleration via PPO, TRPO, DDPG, or SAC) face systematic evaluation pitfalls:

    1. Sensitivity to Implementation and Hyperparameters: Nominally identical continuous policy gradient algorithms yield widely varying driving performance across differing random seeds, network architectures, activation functions, learning rate schedules, and replay buffer sampling parameters. Slight differences in low-level codebases alter policy convergence and stability.

    2. Rollout Selection Bias and Reporting Inconsistency: Selecting the top-kk evaluation rollouts or reporting only maximum episode returns without comprehensive confidence bounds produces unrepresentative estimates of safety and generalization. In autonomous driving, average return is insufficient because safety depends strictly on the worst-case failure probability across diverse test scenarios.

    3. Benchmarking Deficits: Evaluating driving agents solely on fixed nominal circuits fails to test robustness against rare environmental conditions, occlusions, and adversarial agent behaviors. Robust validation requires standardized evaluation protocols, such as parameterizing surrounding vehicle and pedestrian policies in high-fidelity simulators to actively discover edge-case pre-crash topologies.

Coverage note — Standard textbook RL formulas (e.g., standard scalar Bellman equations, vanilla DQN architecture layers, basic REINFORCE derivations, and author biographies) were deliberately omitted to focus exclusively on the paper's substantive contributions to autonomous driving domain formulations, taxonomies, simulation platforms, and deployment challenges.

References

  1. 1.R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction (Second Edition). MIT Press, 2018.
  2. 2.V. Talpaert., I. Sobh., B. R. Kiran., P. Mannion., S. Yogamani., A. El-Sallab., and P. Perez., “Exploring applications of deep reinforcement learning for real-world autonomous driving systems,” in Proceedings of the 14th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 5 VISAPP: VISAPP,, INSTICC. SciTePress, 2019, pp. 564–572.
  3. 3.M. Siam, S. Elkerdawy, M. Jagersand, and S. Yogamani, “Deep semantic segmentation for automated driving: Taxonomy, roadmap and challenges,” in 2017 IEEE 20th international conference on intelligent transportation systems (ITSC). IEEE, 2017, pp. 1–8.
  4. 4.K. El Madawi, H. Rashed, A. El Sallab, O. Nasr, H. Kamel, and S. Yogamani, “Rgb and lidar fusion based 3d semantic segmentation for autonomous driving,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE, 2019, pp. 7–12.
  5. 5.M. Siam, H. Mahgoub, M. Zahran, S. Yogamani, M. Jagersand, and A. El-Sallab, “Modnet: Motion and appearance based moving object detection network for autonomous driving,” in 2018 21st International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2018, pp. 2859–2864.
  6. 6.V. R. Kumar, S. Milz, C. Witt, M. Simon, K. Amende, J. Petzold, S. Yogamani, and T. Pech, “Monocular fisheye camera depth estimation using sparse lidar supervision,” in 2018 21st International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2018, pp. 2853–2858.
  7. 7.M. Uˇricᡠr, P. Kˇrížek, G. Sistu, and S. Yogamani, “Soilingnet: Soiling detection on automotive surround-view cameras,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE, 2019, pp. 67–72.
  8. 8.G. Sistu, I. Leang, S. Chennupati, S. Yogamani, C. Hughes, S. Milz, and S. Rawashdeh, “Neurall: Towards a unified visual perception model for automated driving,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE, 2019, pp. 796–803.
  9. 9.S. Yogamani, C. Hughes, J. Horgan, G. Sistu, P. Varley, D. O’Dea, M. Uricár, S. Milz, M. Simon, K. Amende et al., “Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving,” in Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 9308–9318.
  10. 10.S. Milz, G. Arbeiter, C. Witt, B. Abdallah, and S. Yogamani, “Visual slam for automated driving: Exploring the applications of deep learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 247–257.
  11. 11.S. M. LaValle, Planning Algorithms. New York, NY, USA: Cambridge University Press, 2006.
  12. 12.S. M. LaValle and J. James J. Kuffner, “Randomized kinodynamic planning,” The International Journal of Robotics Research, vol. 20, no. 5, pp. 378–400, 2001.
  13. 13.T. D. Team. Dimensions publication trends. [Online]. Available: https://app.dimensions.ai/discover/publication
  14. 14.Y. Kuwata, J. Teo, G. Fiore, S. Karaman, E. Frazzoli, and J. P. How, “Real-time motion planning with applications to autonomous urban driving,” IEEE Transactions on Control Systems Technology, vol. 17, no. 5, pp. 1105–1118, 2009.
  15. 15.B. Paden, M. Cáp, S. Z. Yong, D. Yershov, and E. Frazzoli, “A survey of ˇ motion planning and control techniques for self-driving urban vehicles,” IEEE Transactions on intelligent vehicles, vol. 1, no. 1, pp. 33–55, 2016.
  16. 16.W. Schwarting, J. Alonso-Mora, and D. Rus, “Planning and decision-making for autonomous vehicles,” Annual Review of Control, Robotics, and Autonomous Systems, no. 0, 2018.
  17. 17.S. Kuutti, R. Bowden, Y. Jin, P. Barber, and S. Fallah, “A survey of deep learning applications to autonomous vehicle control,” IEEE Transactions on Intelligent Transportation Systems, 2020.
  18. 18.T. M. Mitchell, Machine learning, ser. McGraw-Hill series in computer science. Boston (Mass.), Burr Ridge (Ill.), Dubuque (Iowa): McGraw-Hill, 1997.
  19. 19.S. J. Russell and P. Norvig, Artificial intelligence: a modern approach (3rd edition). Prentice Hall, 2009.
  20. 20.Z.-W. Hong, T.-Y. Shann, S.-Y. Su, Y.-H. Chang, T.-J. Fu, and C.-Y. Lee, “Diversity-driven exploration strategy for deep reinforcement learning,” in Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., 2018, pp. 10 489–10 500.
  21. 21.M. Wiering and M. van Otterlo, Eds., Reinforcement Learning: State-of-the-Art. Springer, 2012.
  22. 22.M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming, 1st ed. New York, NY, USA: John Wiley & Sons, Inc., 1994.
  23. 23.C. J. Watkins and P. Dayan, “Technical note: Q-learning,” Machine Learning, vol. 8, no. 3-4, 1992.
  24. 24.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, 2015.
  25. 25.C. J. C. H. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, King’s College, Cambridge, 1989.
  26. 26.D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in ICML, 2014.
  27. 27.R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, pp. 229–256, 1992.
  28. 28.J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning, 2015, pp. 1889–1897.
  29. 29.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
  30. 30.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning.” in 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, Y. Bengio and Y. LeCun, Eds., 2016.
  31. 31.V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning, 2016.
  32. 32.T. Haarnoja, H. Tang, P. Abbeel, and S. Levine, “Reinforcement learning with deep energy-based policies,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70, ser. ICML’17. JMLR.org, 2017, pp. 1352–1361.
  33. 33.T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel et al., “Soft actor-critic algorithms & applications,” arXiv:1812.05905, 2018.
  34. 34.R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine Learning Proceedings 1990. Elsevier, 1990.
  35. 35.R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research, vol. 3, no. Oct, 2002.
  36. 36.D. Silver, R. S. Sutton, and M. Müller, “Sample-based learning and search with permanent and transient memories,” in Proceedings of the 25th international conference on Machine learning. ACM, 2008, pp. 968–975.
  37. 37.G. A. Rummery and M. Niranjan, “On-line Q-learning using connectionist systems,” Cambridge University Engineering Department, Cambridge, England, Tech. Rep. TR 166, 1994.
  38. 38.R. S. Sutton and A. G. Barto, “Reinforcement learning an introduction–second edition, in progress (draft),” 2015.
  39. 39.R. Bellman, Dynamic Programming. Princeton, NJ, USA: Princeton University Press, 1957.
  40. 40.G. Tesauro, “Td-gammon, a self-teaching backgammon program, achieves master-level play,” Neural Computing, vol. 6, no. 2, Mar. 1994.
  41. 41.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al., “Mastering the game of go with deep neural networks and tree search,” nature, vol. 529, no. 7587, pp. 484–489, 2016.
  42. 42.K. Narasimhan, T. Kulkarni, and R. Barzilay, “Language understanding for text-based games using deep reinforcement learning,” in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2015, pp. 1–11.
  43. 43.T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” arXiv preprint arXiv:1511.05952, 2015.
  44. 44.H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning.” in AAAI, vol. 16, 2016, pp. 2094–2100.
  45. 45.Z. Wang, T. Schaul, M. Hessel, H. Van Hasselt, M. Lanctot, and N. De Freitas, “Dueling network architectures for deep reinforcement learning,” arXiv preprint arXiv:1511.06581, 2015.
  46. 46.M. Hausknecht and P. Stone, “Deep recurrent q-learning for partially observable mdps,” CoRR, abs/1507.06527, 2015.
  47. 47.B. F. Skinner, The behavior of organisms: An experimental analysis. Appleton-Century, 1938.
  48. 48.E. Wiewiora, “Reward shaping,” in Encyclopedia of Machine Learning and Data Mining, C. Sammut and G. I. Webb, Eds. Boston, MA: Springer US, 2017, pp. 1104–1106.
  49. 49.J. Randløv and P. Alstrøm, “Learning to drive a bicycle using reinforcement learning and shaping,” in Proceedings of the Fifteenth International Conference on Machine Learning, ser. ICML ’98. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1998, pp. 463–471.
  50. 50.D. H. Wolpert, K. R. Wheeler, and K. Tumer, “Collective intelligence for control of distributed dynamical systems,” EPL (Europhysics Letters), vol. 49, no. 6, p. 708, 2000.
  51. 51.A. Y. Ng, D. Harada, and S. J. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in Proceedings of the Sixteenth International Conference on Machine Learning, ser. ICML ’99, 1999, pp. 278–287.
  52. 52.S. Devlin and D. Kudenko, “Theoretical considerations of potential-based reward shaping for multi-agent systems,” in Proceedings of the 10th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2011.
  53. 53.P. Mannion, S. Devlin, K. Mason, J. Duggan, and E. Howley, “Policy invariance under reward transformations for multi-objective reinforcement learning,” Neurocomputing, vol. 263, 2017.
  54. 54.M. Colby and K. Tumer, “An evolutionary game theoretic analysis of difference evaluation functions,” in Proceedings of the 2015 Annual Conference on Genetic and Evolutionary Computation. ACM, 2015, pp. 1391–1398.
  55. 55.P. Mannion, J. Duggan, and E. Howley, “A theoretical and empirical analysis of reward transformations in multi-objective stochastic games,” in Proceedings of the 16th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2017.
  56. 56.L. Bu¸soniu, R. Babuška, and B. Schutter, “Multi-agent reinforcement learning: An overview,” in Innovations in Multi-Agent Systems and Applications - 1, ser. Studies in Computational Intelligence, D. Srinivasan and L. Jain, Eds. Springer Berlin Heidelberg, 2010, vol. 310.
  57. 57.P. Mannion, K. Mason, S. Devlin, J. Duggan, and E. Howley, “Multi-objective dynamic dispatch optimisation using multi-agent reinforcement learning,” in Proceedings of the 15th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), 2016.
  58. 58.K. Mason, P. Mannion, J. Duggan, and E. Howley, “Applying multi-agent reinforcement learning to watershed management,” in Proceedings of the Adaptive and Learning Agents workshop (at AAMAS 2016), 2016.
  59. 59.V. Pareto, Manual of political economy. OUP Oxford, 1906.
  60. 60.D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley, “A survey of multi-objective sequential decision-making,” Journal of Artificial Intelligence Research, vol. 48, pp. 67–113, 2013.
  61. 61.R. Radulescu, P. Mannion, D. M. Roijers, and A. Nowé, “Multi-objective multi-agent decision making: a utility-based analysis and survey,” Autonomous Agents and Multi-Agent Systems, vol. 34, no. 1, p. 10, 2020.
  62. 62.T. Lesort, N. Diaz-Rodriguez, J.-F. Goudou, and D. Filliat, “State representation learning for control: An overview,” Neural Networks, vol. 108, pp. 379 – 392, 2018.
  63. 63.A. Raffin, A. Hill, K. R. Traoré, T. Lesort, N. D. Rodríguez, and D. Filliat, “Decoupling feature extraction from policy learning: assessing benefits of state representation learning in goal based robotics,” CoRR, vol. abs/1901.08651, 2019.
  64. 64.W. Böhmer, J. T. Springenberg, J. Boedecker, M. Riedmiller, and K. Obermayer, “Autonomous learning of state representations for control: An emerging field aims to autonomously learn state representations for reinforcement learning agents from their real-world sensor observations,” KI-Künstliche Intelligenz, vol. 29, no. 4, pp. 353–362, 2015.
  65. 65.D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton et al., “Mastering the game of go without human knowledge,” Nature, vol. 550, no. 7676, p. 354, 2017.
  66. 66.P. Abbeel and A. Y. Ng, “Exploration and apprenticeship learning in reinforcement learning,” in Proceedings of the 22nd international conference on Machine learning. ACM, 2005, pp. 1–8.
  67. 67.B. Kang, Z. Jie, and J. Feng, “Policy optimization with demonstrations,” in International Conference on Machine Learning, 2018, pp. 2474–2483.
  68. 68.T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband et al., “Deep q-learning from demonstrations,” in Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  69. 69.S. Ibrahim and D. Nevin, “End-to-end framework for fast learning asynchronous agents,” in the 32nd Conference on Neural Information Processing Systems, Imitation Learning and its Challenges in Robotics workshop, 2018.
  70. 70.P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the twenty-first international conference on Machine learning. ACM, 2004, p. 1.
  71. 71.A. Y. Ng, S. J. Russell et al., “Algorithms for inverse reinforcement learning.” in ICML, 2000.
  72. 72.J. Ho and S. Ermon, “Generative adversarial imitation learning,” in Advances in Neural Information Processing Systems, 2016, pp. 4565–4573.
  73. 73.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27, 2014, pp. 2672–2680.
  74. 74.M. Uˇricᡠr, P. Kˇrížek, D. Hurych, I. Sobh, S. Yogamani, and P. Denny, “Yes, we gan: Applying adversarial techniques for autonomous driving,” Electronic Imaging, vol. 2019, no. 15, pp. 48–1, 2019.
  75. 75.E. Leurent, Y. Blanco, D. Efimov, and O.-A. Maillard, “A survey of state-action representations for autonomous driving,” HAL archives, 2018.
  76. 76.H. Xu, Y. Gao, F. Yu, and T. Darrell, “End-to-end learning of driving models from large-scale video datasets,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 2174–2182.
  77. 77.R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence, vol. 112, no. 1-2, pp. 181–211, 1999.
  78. 78.A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in Proceedings of the 1st Annual Conference on Robot Learning, 2017, pp. 1–16.
  79. 79.C. Li and K. Czarnecki, “Urban driving with multi-objective deep reinforcement learning,” in Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 2019, pp. 359–367.
  80. 80.S. Kardell and M. Kuosku, “Autonomous vehicle control via deep reinforcement learning,” Master’s thesis, Chalmers University of Technology, 2017.
  81. 81.J. Chen, B. Yuan, and M. Tomizuka, “Model-free deep reinforcement learning for urban autonomous driving,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE, 2019, pp. 2765–2771.
  82. 82.A. E. Sallab, M. Abdou, E. Perot, and S. Yogamani, “End-to-end deep reinforcement learning for lane keeping assist,” in MLITS, NIPS Workshop, vol. 2, 2016.
  83. 83.A.-E. Sallab, M. Abdou, E. Perot, and S. Yogamani, “Deep reinforcement learning framework for autonomous driving,” Electronic Imaging, vol. 2017, no. 19, pp. 70–76, 2017.
  84. 84.P. Wang, C.-Y. Chan, and A. de La Fortelle, “A reinforcement learning based approach for automated lane change maneuvers,” in 2018 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2018, pp. 1379–1384.
  85. 85.P. Wang and C.-Y. Chan, “Formulation of deep reinforcement learning architecture toward autonomous driving for on-ramp merge,” in Intelligent Transportation Systems (ITSC), 2017 IEEE 20th International Conference on. IEEE, 2017, pp. 1–6.
  86. 86.D. C. K. Ngai and N. H. C. Yung, “A multiple-goal reinforcement learning method for complex vehicle overtaking maneuvers,” IEEE Transactions on Intelligent Transportation Systems, vol. 12, no. 2, pp. 509–522, 2011.
  87. 87.D. Isele, R. Rahimi, A. Cosgun, K. Subramanian, and K. Fujimura, “Navigating occluded intersections with autonomous vehicles using deep reinforcement learning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018, pp. 2034–2039.
  88. 88.A. Keselman, S. Ten, A. Ghazali, and M. Jubeh, “Reinforcement learning with a* and a deep heuristic,” arXiv preprint arXiv:1811.07745, 2018.
  89. 89.W. Zhan, L. Sun, D. Wang, H. Shi, A. Clausse, M. Naumann, J. Kümmerle, H. Königshof, C. Stiller, A. de La Fortelle, and M. Tomizuka, “INTERACTION Dataset: An INTERnational, Adversarial and Cooperative moTION Dataset in Interactive Driving Scenarios with Semantic Maps,” arXiv:1910.03088 [cs, eess], 2019.
  90. 90.A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J.-M. Allen, V.-D. Lam, A. Bewley, and A. Shah, “Learning to drive in a day,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 8248–8254.
  91. 91.M. Watter, J. Springenberg, J. Boedecker, and M. Riedmiller, “Embed to control: A locally linear latent dynamics model for control from raw images,” in Advances in neural information processing systems, 2015.
  92. 92.N. Wahlström, T. B. Schön, and M. P. Deisenroth, “Learning deep dynamical models from image pixels,” IFAC-PapersOnLine, vol. 48, no. 28, pp. 1059–1064, 2015.
  93. 93.S. Chiappa, S. Racanière, D. Wierstra, and S. Mohamed, “Recurrent environment simulators,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017.
  94. 94.B. Recht, “A tour of reinforcement learning: The view from continuous control,” Annual Review of Control, Robotics, and Autonomous Systems, 2008.
  95. 95.H. Mania, A. Guy, and B. Recht, “Simple random search of static linear policies is competitive for reinforcement learning,” in Advances in Neural Information Processing Systems 31, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds., 2018, pp. 1800–1809.
  96. 96.B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner, “Torcs, the open racing car simulator,” Software available at http://torcs. sourceforge. net, vol. 4, 2000.
  97. 97.S. Shah, D. Dey, C. Lovett, and A. Kapoor, “Airsim: High-fidelity visual and physical simulation for autonomous vehicles,” in Field and Service Robotics. Springer, 2018, pp. 621–635.
  98. 98.N. Koenig and A. Howard, “Design and use paradigms for gazebo, an open-source multi-robot simulator,” in 2004 International Conference on Intelligent Robots and Systems (IROS), vol. 3. IEEE, 2004, pp. 2149–2154.
  99. 99.P. A. Lopez, M. Behrisch, L. Bieker-Walz, J. Erdmann, Y.-P. Flötteröd, R. Hilbrich, L. Lücken, J. Rummel, P. Wagner, and E. Wießner, “Microscopic traffic simulation using sumo,” in The 21st IEEE International Conference on Intelligent Transportation Systems. IEEE, 2018.
  100. 100.C. Quiter and M. Ernst, “deepdrive/deepdrive: 2.0,” Mar. 2018. [Online]. Available: https://doi.org/10.5281/zenodo.1248998
  101. 101.Nvidia, “Drive Constellation now available,” https://blogs.nvidia.com/blog/2019/03/18/drive-constellation-now-available/, 2019, [accessed 14-April-2019].
  102. 102.A. S. et al., “Multi-Agent Autonomous Driving Simulator built on top of TORCS,” https://github.com/madras-simulator/MADRaS, 2019, [Online; accessed 14-April-2019].
  103. 103.C. Wu, A. Kreidieh, K. Parvate, E. Vinitsky, and A. M. Bayen, “Flow: Architecture and benchmarking for reinforcement learning in traffic control,” CoRR, vol. abs/1710.05465, 2017.
  104. 104.E. Leurent, “A collection of environments for autonomous driving and tactical decision-making tasks,” https://github.com/eleurent/highway-env, 2019, [Online; accessed 14-April-2019].
  105. 105.F. Rosique, P. J. Navarro, C. Fernández, and A. Padilla, “A systematic review of perception system and simulators for autonomous vehicles research,” Sensors, vol. 19, no. 3, p. 648, 2019.
  106. 106.M. Cutler, T. J. Walsh, and J. P. How, “Reinforcement learning with multi-fidelity simulators,” in 2014 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2014, pp. 3888–3895.
  107. 107.F. C. German Ros, Vladlen Koltun and A. M. Lopez, “Carla autonomous driving challenge,” https://carlachallenge.org/, 2019, [Online; accessed 14-April-2019].
  108. 108.W. G. Najm, J. D. Smith, M. Yanagisawa et al., “Pre-crash scenario typology for crash avoidance research,” United States. National Highway Traffic Safety Administration, Tech. Rep., 2007.
  109. 109.D. A. Pomerleau, “Alvinn: An autonomous land vehicle in a neural network,” in Advances in neural information processing systems, 1989.
  110. 110.D. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,” Neural Computation, vol. 3, no. 1, 1991.
  111. 111.M. Bojarski, D. Del Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang et al., “End to end learning for self-driving cars,” in NIPS 2016 Deep Learning Symposium, 2016.
  112. 112.M. Bojarski, P. Yeres, A. Choromanska, K. Choromanski, B. Firner, L. Jackel, and U. Muller, “Explaining how a deep neural network trained with end-to-end learning steers a car,” arXiv preprint arXiv:1704.07911, 2017.
  113. 113.M. Kuderer, S. Gulati, and W. Burgard, “Learning driving styles for autonomous vehicles from demonstration,” in Robotics and Automation (ICRA), 2015 IEEE International Conference on. IEEE, 2015, pp. 2641–2646.
  114. 114.S. Sharifzadeh, I. Chiotellis, R. Triebel, and D. Cremers, “Learning to drive using inverse reinforcement learning and deep q-networks,” in NIPS Workshops, December 2016.
  115. 115.P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Thirty-Second AAAI Conference on Artificial Intelligence, 2018.
  116. 116.Y. Abeysirigoonawardena, F. Shkurti, and G. Dudek, “Generating adversarial driving scenarios in high-fidelity simulators,” in 2019 IEEE International Conference on Robotics and Automation (ICRA). ICRA, 2019.
  117. 117.K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige et al., “Using simulation and domain adaptation to improve efficiency of deep robotic grasping,” in 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018, pp. 4243–4250.
  118. 118.X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” in 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2018, pp. 1–8.
  119. 119.Z. W. Xinlei Pan, Yurong You and C. Lu, “Virtual to real reinforcement learning for autonomous driving,” in Proceedings of the British Machine Vision Conference (BMVC), G. B. Tae-Kyun Kim, Stefanos Zafeiriou and K. Mikolajczyk, Eds. BMVA Press, September 2017.
  120. 120.A. Bewley, J. Rigley, Y. Liu, J. Hawke, R. Shen, V.-D. Lam, and A. Kendall, “Learning to drive from simulation without real world labels,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 4818–4824.
  121. 121.J. Zhang, L. Tai, P. Yun, Y. Xiong, M. Liu, J. Boedecker, and W. Burgard, “Vr-goggles for robots: Real-to-sim domain adaptation for visual control,” IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 1148–1155, 2019.
  122. 122.H. Chae, C. M. Kang, B. Kim, J. Kim, C. C. Chung, and J. W. Choi, “Autonomous braking system via deep reinforcement learning,” 2017 IEEE 20th International Conference on Intelligent Transportation Systems (ITSC), pp. 1–6, 2017.
  123. 123.Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas, “Sample efficient actor-critic with experience replay,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017.
  124. 124.R. Liaw, S. Krishnan, A. Garg, D. Crankshaw, J. E. Gonzalez, and K. Goldberg, “Composing meta-policies for autonomous driving using hierarchical deep reinforcement learning,” arXiv preprint arXiv:1711.01503, 2017.
  125. 125.M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research, vol. 10, no. Jul, pp. 1633–1685, 2009.
  126. 126.D. Isele and A. Cosgun, “Transferring autonomous driving knowledge on simulated and real intersections,” in Lifelong Learning: A Reinforcement Learning Approach,ICML WORKSHOP 2017, 2017.
  127. 127.J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick, “Learning to reinforcement learn,” Complete CogSci 2017 Proceedings, 2016.
  128. 128.Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel, “Fast reinforcement learning via slow reinforcement learning,” arXiv preprint arXiv:1611.02779, 2016.
  129. 129.C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70, ser. ICML’17. JMLR.org, 2017, p. 1126–1135.
  130. 130.A. Nichol, J. Achiam, and J. Schulman, “On first-order meta-learning algorithms,” CoRR, abs/1803.02999, 2018.
  131. 131.M. Al-Shedivat, T. Bansal, Y. Burda, I. Sutskever, I. Mordatch, and P. Abbeel, “Continuous adaptation via meta-learning in nonstationary and competitive environments,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net, 2018.
  132. 132.D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,” in Advances in Neural Information Processing Systems, 2018.
  133. 133.S. Ross and D. Bagnell, “Efficient reductions for imitation learning,” in Proceedings of the thirteenth international conference on artificial intelligence and statistics, 2010, pp. 661–668.
  134. 134.M. Bansal, A. Krizhevsky, and A. Ogale, “Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst,” in Robotics: Science and Systems XV, 2018.
  135. 135.T. Buhet, E. Wirbel, and X. Perrotton, “Conditional vehicle trajectories prediction in carla urban environment,” in Proceedings of the IEEE International Conference on Computer Vision Workshops, 2019.
  136. 136.A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML, vol. 99, 1999, pp. 278–287.
  137. 137.P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in Proceedings of the Twenty-first International Conference on Machine Learning, ser. ICML ’04. ACM, 2004.
  138. 138.N. Chentanez, A. G. Barto, and S. P. Singh, “Intrinsically motivated reinforcement learning,” in Advances in neural information processing systems, 2005, pp. 1281–1288.
  139. 139.D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell, “Curiosity-driven exploration by self-supervised prediction,” in International Conference on Machine Learning (ICML), vol. 2017, 2017.
  140. 140.Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, and A. A. Efros, “Large-scale study of curiosity-driven learning,” arXiv preprint arXiv:1808.04355, 2018.
  141. 141.J. Zhang and K. Cho, “Query-efficient imitation learning for end-to-end simulated driving,” in Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, California, USA., 2017, pp. 2891–2897.
  142. 142.S. Shalev-Shwartz, S. Shammah, and A. Shashua, “Safe, multi-agent, reinforcement learning for autonomous driving,” arXiv preprint arXiv:1610.03295, 2016.
  143. 143.X. Xiong, J. Wang, F. Zhang, and K. Li, “Combining deep reinforcement learning and safety based control for autonomous driving,” arXiv preprint arXiv:1612.00147, 2016.
  144. 144.C. Ye, H. Ma, X. Zhang, K. Zhang, and S. You, “Survival-oriented reinforcement learning model: An effcient and robust deep reinforcement learning algorithm for autonomous driving problem,” in International Conference on Image and Graphics. Springer, 2017, pp. 417–429.
  145. 145.J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research, vol. 16, no. 1, pp. 1437–1480, 2015.
  146. 146.P. Palanisamy, “Multi-agent connected autonomous driving using deep reinforcement learning,” in 2020 International Joint Conference on Neural Networks (IJCNN). IEEE, 2020, pp. 1–7.
  147. 147.S. Bhalla, S. Ganapathi Subramanian, and M. Crowley, “Deep multi agent reinforcement learning for autonomous driving,” in Advances in Artificial Intelligence, C. Goutte and X. Zhu, Eds. Cham: Springer International Publishing, 2020, pp. 67–78.
  148. 148.A. Wachi, “Failure-scenario maker for rule-based agent using multi-agent adversarial reinforcement learning and its application to autonomous driving,” arXiv preprint arXiv:1903.10654, 2019.
  149. 149.C. Yu, X. Wang, X. Xu, M. Zhang, H. Ge, J. Ren, L. Sun, B. Chen, and G. Tan, “Distributed multiagent coordinated learning for autonomous driving in highways based on dynamic coordination graphs,” IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 2, pp. 735–748, 2020.
  150. 150.P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, Y. Wu, and P. Zhokhov, “Openai baselines,” https://github.com/openai/baselines, 2017.
  151. 151.A. Juliani, V.-P. Berges, E. Vckay, Y. Gao, H. Henry, M. Mattar, and D. Lange, “Unity: A general platform for intelligent agents,” arXiv preprint arXiv:1809.02627, 2018.
  152. 152.I. Caspi, G. Leibovich, G. Novik, and S. Endrawis, “Reinforcement learning coach,” Dec. 2017. [Online]. Available: https://doi.org/10.5281/zenodo.1134899
  153. 153.Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Neal Wu, Chris Harris, Vincent Vanhoucke, Eugene Brevdo, “TF-Agents: A library for reinforcement learning in tensorflow,” https://github.com/tensorflow/agents, 2018, [Online; accessed 25-June-2019]. [Online]. Available: https://github.com/tensorflow/agents
  154. 154.A. Stooke and P. Abbeel, “rlpyt: A research code base for deep reinforcement learning in pytorch,” arXiv preprint arXiv:1909.01500, 2019.
  155. 155.I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepezvari, S. Singh et al., “Behaviour suite for reinforcement learning,” arXiv preprint arXiv:1908.03568, 2019.

Citation

MLA
Kiran, B. R., et al. “Deep Reinforcement Learning for Autonomous Driving: A Survey”. arXiv, 2020, http://arxiv.org/abs/2002.00444v2.
APA
Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Sallab, A. A. A., Yogamani, S., & Pérez, P. (2020). Deep Reinforcement Learning for Autonomous Driving: A Survey. arXiv. http://arxiv.org/abs/2002.00444v2
Chicago
Kiran, B. R., I. Sobh, V. Talpaert, et al. 2020. “Deep Reinforcement Learning for Autonomous Driving: A Survey”. arXiv. http://arxiv.org/abs/2002.00444v2.
Harvard
Kiran, B.R. et al. (2020) “Deep Reinforcement Learning for Autonomous Driving: A Survey”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2002.00444v2.
Vancouver
1. Kiran BR, Sobh I, Talpaert V, Mannion P, Sallab AAA, Yogamani S, Pérez P (2020) Deep Reinforcement Learning for Autonomous Driving: A Survey. arXiv

BibTeX

@article{kiran2020deep,
  title = {Deep Reinforcement Learning for Autonomous Driving: A Survey},
  author = {Kiran, B Ravi and Sobh, Ibrahim and Talpaert, Victor and Mannion, Patrick and Sallab, Ahmad A. Al and Yogamani, Senthil and Pérez, Patrick},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2002.00444v2},
  eprint = {2002.00444}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF