A survey of deep learning techniques for autonomous driving

Sorin GrigorescuBogdan TrasneaTiberiu CociasGigel Macesanu

article2019J. Field Robotics1,784 citations

Presents a structured comparative analysis of modular and end-to-end deep learning methods across perception, path planning, and control to guide system design and address key safety and hardware constraints in autonomous vehicles.

Listen

Autonomous vehicles have shifted from controlled research environments to real-world deployment on public roads, driven by the need to decrease traffic accidents, relieve congestion, and enhance urban mobility. While traditional perception and control methods effectively handle standard driving situations, they fail in complex corner cases that require human-like reasoning. This article evaluates the current state of deep learning technologies across autonomous driving architectures, examining perception, path planning, behavior arbitration, and motion control, while addressing systemic challenges in safety compliance, data collection, and embedded computing hardware.

The authors conducted a comprehensive survey of modern artificial intelligence methodologies applied to self-driving systems, comparing modular pipelines—where discrete tasks are executed by specialized deep learning or classical algorithms—against unified end-to-end systems that directly translate sensor streams into control signals. The survey examines core algorithmic paradigms including Convolutional Neural Networks for spatial understanding, Recurrent Neural Networks for temporal sequence tracking, and Deep Reinforcement Learning for decision-making policies. In addition, the article benchmarks object detection, segmentation, and localization models, while analyzing public datasets and target compute platforms like Graphics Processing Units and Field-Programmable Gate Arrays.

The findings establish that modular perception-planning-action architectures remain the dominant and most viable deployment model because they allow developers to incorporate established safety constraints and verify individual subsystems, unlike fully end-to-end models whose decision processes lack interpretability. In perception, single-stage object detectors deliver real-time operational speeds but sacrifice accuracy compared to double-stage detectors, while emerging pseudo-LiDAR techniques demonstrate the potential to derive three-dimensional spatial data from lower-cost camera sensors. In vehicle control, integrating machine learning with traditional methods—such as Model Predictive Control—yields superior stability and safety compared to pure reinforcement learning, which often learns biased behaviors when transferred from simulations to the real world. Furthermore, while Graphics Processing Units lead raw artificial intelligence processing, Field-Programmable Gate Arrays consume up to ten times less power and offer higher architectural flexibility and functional safety compliance for embedded in-vehicle integration.

These insights indicate that current industrial safety frameworks, such as the standard ISO 26262 for automotive functional safety, are structurally unequipped to certify machine learning systems that learn from data rather than explicit software specifications. This mismatch heightens operational risk, as real-world edge cases (often termed "Black Swans") absent from training sets can lead to catastrophic failures and severe liability. Hardware and sensor selection also directly impact production cost and scalability; while high-resolution LiDAR ensures robust three-dimensional perception, its high price tag currently hinders widespread commercial adoption compared to camera-first approaches.

To advance the safe commercialization of autonomous vehicles, industry stakeholders and policymakers should prioritize developing updated functional safety standards tailored to deep learning systems. Engineering teams should deploy hybrid architectures that wrap learning-based models with non-learnable deterministic safety monitors and explore low-power embedded hardware like Field-Programmable Gate Arrays for production deployment. Research efforts must focus on improving data collection for rare corner cases, advancing few-shot learning techniques, and refining camera-based three-dimensional depth perception to bridge simulation-to-reality performance gaps. Confidence in these findings is high regarding algorithmic trade-offs and compute characteristics, though caution is warranted regarding long-term reliability due to the fragmented nature of available training datasets and the persistent lack of explainability in deep neural networks.

arXiv: 1910.07738
Cover for A survey of deep learning techniques for autonomous driving

Abstract

The last decade witnessed increasingly rapid progress in self-driving vehicle technology, mainly backed up by advances in the area of deep learning and artificial intelligence. The objective of this paper is to survey the current state-of-the-art on deep learning technologies used in autonomous driving. We start by presenting AI-based self-driving architectures, convolutional and recurrent neural networks, as well as the deep reinforcement learning paradigm. These methodologies form a base for the surveyed driving scene perception, path planning, behavior arbitration and motion control algorithms. We investigate both the modular perception-planning-action pipeline, where each module is built using deep learning methods, as well as End2End systems, which directly map sensory information to steering commands. Additionally, we tackle current challenges encountered in designing AI architectures for autonomous driving, such as their safety, training data sources and computational hardware. The comparison presented in this survey helps to gain insight into the strengths and limitations of deep learning and AI approaches for autonomous driving and assist with design choices

Table of Contents

  • 1 Introduction
  • 2 Deep Learning based Decision-Making Architectures for Self-Driving Cars
  • 3 Overview of Deep Learning Technologies
  • 3.1 Deep Convolutional Neural Networks
  • 3.2 Recurrent Neural Networks
  • 3.3 Deep Reinforcement Learning
  • 4 Deep Learning for Driving Scene Perception and Localization
  • 4.1 Sensing Hardware: Camera vs. LiDAR Debate
  • 4.2 Driving Scene Understanding
  • 4.2.1 Bounding-Box-Like Object Detectors
  • 4.2.2 Semantic and Instance Segmentation
  • 4.2.3 Localization
  • 4.3 Perception using Occupancy Maps
  • 5 Deep Learning for Path Planning and Behavior Arbitration
  • 6 Motion Controllers for AI-based Self-Driving Cars
  • 6.1 Learning Controllers
  • 6.2 End2End Learning Control
  • 7 Safety of Deep Learning in Autonomous Driving
  • 8 Data Sources for Training Autonomous Driving Systems
  • 9 Computational Hardware and Deployment
  • 10 Discussion and Conclusions
  • 10.1 Final Notes
  • References

Knowls

  1. Knowl 1 — Modular Perception-Planning-Action Pipeline vs. End-to-End Decision Architectures

    model/method

    Autonomous vehicle decision-making systems process multi-sensor data streams (cameras, LiDARs, radars, ultrasonic sensors, GPS, IMUs) using two principal architectural paradigms:

    1. Modular Perception-Planning-Action Pipeline: The driving task is hierarchically partitioned into four sequential stages:

      • Perception and Localization: Generates an environmental representation (e.g., bounding boxes, semantic segmentations, occupancy grids) and calculates ego-vehicle pose.
      • High-Level Path Planning: Computes a global route through the road network between origin and destination coordinates.
      • Behavior Arbitration (Low-Level Path Planning): Resolves multi-agent driving interactions (e.g., lane changing, overtaking, giving way) and calculates local collision-free trajectories over a finite prediction horizon.
      • Motion Controllers: Computes longitudinal and lateral actuator commands (steering, throttle, brake) to track the planned trajectory and compensate for execution errors. Each module can be independently designed using deep learning methods or classical algorithms, supervised by dedicated safety monitors.
    2. End-to-End (End2End) Learning Systems: Sensory observations (e.g., camera images, LiDAR point clouds) are mapped directly to vehicle control commands (e.g., steering angle, velocity) via a unified deep neural network or deep reinforcement learning policy, bypassing intermediate symbolic representations.

  2. Knowl 2 — Partially Observable Markov Decision Process Formulation for Deep Reinforcement Learning in Autonomous Driving

    model/method

    The autonomous driving task in Deep Reinforcement Learning (DRL) is formalized as a Partially Observable Markov Decision Process (POMDP) defined by the 6-tuple M = (I, S, A, T, R, oldsymbol{\gamma}):

    • II is the observation space, where I⟨t⟩∈II^{\langle t \rangle} \in I denotes the sensory observation of the environment at discrete time step tt.
    • SS is the finite state space, where s⟨t⟩∈Ss^{\langle t \rangle} \in S represents the internal vehicle state (position, heading, and velocity).
    • AA is the action space containing commands a⟨t⟩∈Aa^{\langle t \rangle} \in A available to the vehicle.
    • T:S×A×S→[0,1]T: S \times A \times S \to [0, 1] is the stochastic transition probability function, where T(s⟨t+1⟩∣s⟨t⟩,a⟨t⟩)T(s^{\langle t+1 \rangle} \mid s^{\langle t \rangle}, a^{\langle t \rangle}) gives the probability of reaching state s⟨t+1⟩s^{\langle t+1 \rangle} after executing action a⟨t⟩a^{\langle t \rangle} in state s⟨t⟩s^{\langle t \rangle}.
    • R:S×A×S→RR: S \times A \times S \to \mathbb{R} is the scalar reward function, where R(s⟨t⟩,a⟨t⟩,s⟨t+1⟩)∈RR(s^{\langle t \rangle}, a^{\langle t \rangle}, s^{\langle t+1 \rangle}) \in \mathbb{R} evaluates the transition quality toward a destination state sdest⟨t+k⟩s_{\text{dest}}^{\langle t+k \rangle}.
    • γ∈[0,1]\gamma \in [0, 1] is the temporal discount factor.

    For a trajectory of length kk, the cumulative future discounted return from time index t^\hat{t} is:

    R⟨t^⟩=∑t=t^kγt−t^r⟨t⟩R^{\langle \hat{t} \rangle} = \sum_{t=\hat{t}}^{k} \gamma^{t - \hat{t}} r^{\langle t \rangle}

    where r⟨t⟩r^{\langle t \rangle} is the immediate reward. The optimal action-value function Q∗(s,a)=max⁡πE[R⟨t^⟩∣s⟨t^⟩=s,a⟨t^⟩=a,π]Q^*(s, a) = \max_{\pi} \mathbb{E}[R^{\langle \hat{t} \rangle} \mid s^{\langle \hat{t} \rangle} = s, a^{\langle \hat{t} \rangle} = a, \pi] satisfies the Bellman optimality equation:

    Q∗(s,a)=∑s′T(s′∣s,a)[R(s,a,s′)+γmax⁡a′Q∗(s′,a′)]Q^*(s, a) = \sum_{s'} T(s' \mid s, a) \left[ R(s, a, s') + \gamma \max_{a'} Q^*(s', a') \right]

    In Deep Q-Networks (DQN), Q∗(s,a)Q^*(s, a) is approximated with neural network parameters Θ\Theta. The network is trained by minimizing the mean squared Bellman error:

    min⁡ΘiEs,a,r,s′[(y−Q(s,a;Θi))2]\min_{\Theta_i} \mathbb{E}_{s, a, r, s'} \left[ \left( y - Q(s, a; \Theta_i) \right)^2 \right]

    with target y=R(s,a,s′)+γmax⁡a′Q(s′,a′;Θˉi)y = R(s, a, s') + \gamma \max_{a'} Q(s', a'; \bar{\Theta}_i), where Θˉi\bar{\Theta}_i represents target network parameters from a prior iteration.

  3. Knowl 3 — Learning-Augmented State-Space Dynamics for Autonomous Vehicle Motion Control

    equation

    To account for unmodeled physical disturbances, tire-road friction shifts, and dynamic platform changes during autonomous vehicle control, the true discrete-time nonlinear state transition function z⟨t+1⟩=ftrue(z⟨t⟩,u⟨t⟩)z^{\langle t+1 \rangle} = f_{\text{true}}(z^{\langle t \rangle}, u^{\langle t \rangle}) with state vector z⟨t⟩∈Rnz^{\langle t \rangle} \in \mathbb{R}^n and control vector u⟨t⟩∈Rmu^{\langle t \rangle} \in \mathbb{R}^m is modeled as the combination of an analytical baseline and a learned data-driven model:

    z⟨t+1⟩=f(z⟨t⟩,u⟨t⟩)⏟a-priori model+h(z⟨t⟩)⏟learned modelz^{\langle t+1 \rangle} = \underbrace{f(z^{\langle t \rangle}, u^{\langle t \rangle})}_{\text{a-priori model}} + \underbrace{h(z^{\langle t \rangle})}_{\text{learned model}}

    where:

    • z⟨t⟩∈Rnz^{\langle t \rangle} \in \mathbb{R}^n is the vehicle state at time tt (e.g., pose, linear and angular velocities).
    • u⟨t⟩∈Rmu^{\langle t \rangle} \in \mathbb{R}^m is the control input at time tt (e.g., steering wheel angle, throttle, brake pressure).
    • f(z⟨t⟩,u⟨t⟩):Rn×Rm→Rnf(z^{\langle t \rangle}, u^{\langle t \rangle}): \mathbb{R}^n \times \mathbb{R}^m \to \mathbb{R}^n is a deterministic, physics-based a-priori model governing nominal vehicle kinematics.
    • h(z⟨t⟩):Rn→Rnh(z^{\langle t \rangle}): \mathbb{R}^n \to \mathbb{R}^n is a non-linear learned function approximator (e.g., Gaussian Process regression, Support Vector Regression, or deep neural networks) trained on past state-observation sequences to predict residual nonlinear dynamic disturbances.

    This hybrid formulation preserves established control-theoretic stability analysis methods within Model Predictive Control (MPC) and Iterative Learning Control (ILC) while improving path-tracking fidelity under handling limits.

  4. Knowl 4 — Autonomy Performance Metric for Evaluating End-to-End Driving Systems

    equation

    The autonomy metric measures the proportion of total operating time during which an autonomous end-to-end driving network maintains vehicle control without requiring human safety takeover:

    autonomy=(1−(no. of interventions)×6 selapsed time [s])×100\text{autonomy} = \left( 1 - \frac{(\text{no. of interventions}) \times 6\text{ s}}{\text{elapsed time [s]}} \right) \times 100

    where:

    • no. of interventions\text{no. of interventions} is the total count of instances where the human safety driver takes manual control of the vehicle.
    • An intervention condition is triggered when the autonomous vehicle deviates by more than 1 m1\text{ m} from the road centerline.
    • 6 s6\text{ s} is the standardized penalty duration assigned for a human driver to retake control, stabilize the vehicle trajectory, and return the car to the target operating state.
    • elapsed time [s]\text{elapsed time [s]} is the total duration of the autonomous test drive in seconds.
  5. Knowl 5 — Taxonomy of End-to-End Deep Learning Control Methods for Autonomous Vehicles

    data/table

    End-to-End driving methods bypass modular pipelines to predict control commands directly from sensory streams. The approaches span supervised learning on human driving data and deep reinforcement learning in simulation.

    System Problem Space Architecture Sensory Input Key Description
    ALVINN Road following 3-layer backprop Camera, laser range finder Pioneer road follower trained on simulated road images
    DAVE DARPA challenge 6-layer CNN Stereo camera images 50cm off-road mobile robot learning obstacle avoidance
    NVIDIA PilotNet Real traffic driving CNN Front camera images Direct mapping from raw pixels to human steering angles
    Novel FCN-LSTM Ego-motion prediction FCN-LSTM Large-scale video data Learns generic vehicle motion model distribution
    Novel C-LSTM Steering angle control C-LSTM Camera frames, steering angle Captures visual and dynamic temporal dependencies
    Drive360 Steering and velocity control CNN + FC + LSTM Surround cameras, CAN bus Integrates 360∘360^\circ surround visual feed with route planner
    DNN policy Steering angle control CNN + FC Front camera images Compares direct camera-to-steering net with human driver
    DeepPicar Steering angle control CNN Web camera images Real-time embedded implementation of DAVE-2 on Raspberry Pi 3
    TORCS DRL Lane keeping, obstacle avoidance DQN + RNN + CNN TORCS simulator images Handles partial observability in racing simulator
    TORCS E2E Simulated steering control CNN TORCS simulator images Evaluates sky, roadside, and road-related visual features
    Agile Auto. Driving Aggressive driving control CNN Raw camera images Learns aggressive control trajectories demonstrated by MPC
    WRC6 AD Racing game driving CNN + LSTM Encoder WRC6 racing game Asynchronous Advantage Actor-Critic (A3C) rally driving

    The progression demonstrates a transition from early shallow networks (ALVINN) to deep CNN feature extractors (PilotNet), recurrent spatial-temporal models (C-LSTM, FCN-LSTM), surround-view integration (Drive360), and policy learning guided by Model Predictive Control (MPC) trajectory optimizers.

  6. Knowl 6 — Statistical Risk Minimization and the Epistemic Safety Gap in Deep Learning

    theoretical result

    In statistical machine learning, an algorithm's hypothesis hh minimizes the empirical risk Remp(h)R_{\text{emp}}(h) over a training dataset of mm samples drawn from space X×YX \times Y:

    Remp(h)=1m∑i=1mL(x(i),h(x(i)),y(i))R_{\text{emp}}(h) = \frac{1}{m} \sum_{i=1}^m L(x^{(i)}, h(x^{(i)}), y^{(i)})

    as an estimator of the true statistical risk under probability distribution P(X,Y)P(X, Y):

    R(f)=∫X×YL(x,f(x),y) dP(x,y)R(f) = \int_{X \times Y} L(x, f(x), y) \, dP(x, y)

    where L:X×Y×Y→RL: X \times Y \times Y \to \mathbb{R} measures prediction loss. In automotive safety standards (such as ISO 26262), component verification requires demonstrating conformity against complete behavioral specifications.

    Deep learning violates this paradigm because:

    1. The system behavior is learned inductively from training data rather than deduced from explicit functional specifications.
    2. The true environmental probability distribution P(X,Y)P(X, Y) is unknown and non-stationary (subject to distributional shift and domain transfer errors).
    3. Rare, catastrophic real-world edge cases ("Black Swan" events) are under-represented or absent in finite training and test datasets, meaning empirical risk minimization cannot guarantee safe operational bounds in safety-critical situations.
  7. Knowl 7 — Public Benchmark Datasets for Autonomous Driving System Development

    data/table

    Public datasets provide essential training and evaluation benchmarks for autonomous driving sub-tasks, spanning perception, SLAM, and behavioral prediction.

    Dataset Problem Space Sensor Setup Size Location Traffic License
    NuScenes 3D tracking, 3D detection Radar, LiDAR, EgoData, GPS, IMU, Cam 345 GB Boston, Singapore Urban CC BY-NC-SA 3.0
    AMUSE SLAM Omnidirectional cam, IMU, EgoData, GPS 1 TB Los Angeles Urban CC BY-NC-ND 3.0
    Ford 3D tracking, 3D detection Omnidirectional cam, IMU, LiDAR, GPS 100 GB Michigan Urban Not specified
    KITTI 3D tracking, 3D detection, SLAM Stereo/Mono cams, IMU, LiDAR, GPS 180 GB Karlsruhe Urban, Rural CC BY-NC-SA 3.0
    Udacity 3D tracking, 3D detection Mono cams, IMU, LiDAR, GPS, EgoData 220 GB Mountain View Rural MIT
    Cityscapes Semantic understanding Color stereo cameras 63 GB 50 Cities (EU) Urban CC BY-NC-SA 3.0
    Oxford 3D tracking, 3D detection, SLAM Stereo/Mono cams, GPS, LiDAR, IMU 23 TB Oxford Urban, Highway CC BY-NC-SA 3.0
    CamVid Object detection, segmentation Monocular color camera 8 GB Cambridge Urban N/A
    Daimler Ped. Pedestrian detection, segmentation Stereo and monocular cameras 91 GB Amsterdam, Beijing Urban N/A
    Caltech Tracking, segmentation, detection Monocular camera 11 GB Los Angeles Urban N/A

    Sensor configurations vary widely across datasets: LiDAR and high-precision GPS/IMUs dominate 3D spatial tasks (KITTI, NuScenes, Oxford), multi-sensor radar feeds are rare (present in NuScenes), and stereo video dominates pixel-level segmentation benchmarks (Cityscapes).

  8. Knowl 8 — Hardware Acceleration Technologies for Embedded On-Vehicle Deep Learning

    model/method

    Deploying deep learning in autonomous driving requires balancing raw compute throughput, latency, power consumption, and automotive functional safety (ISO 26262):

    • Automotive GPU Systems-on-Chip (SoCs): NVIDIA DRIVE PX utilizes dual Tegra X2 SoCs (each combining two Denver cores, four ARM Cortex-A57 cores, and a Pascal GPU). The subsequent DRIVE AGX Xavier delivers 30 trillion operations per second (TOPS) under ISO 26262 compliance, while DRIVE AGX Pegasus integrates dual Xavier SoCs with two TensorCore GPUs for Level 5 compute demands.
    • Dedicated Automotive Vision SoCs: Renesas R-Car V3H utilizes four 1.0 GHz Arm Cortex-A53 cores with hardware accelerators optimized specifically for convolutional neural networks, stereo vision, dense optical flow, and object classification. The R-Car H3 incorporates Arm Cortex-A57 cores to support broader deep learning topologies under functional safety constraints.
    • Field-Programmable Gate Arrays (FPGAs): FPGAs offer high on-chip Block RAM (BRAM), reducing external memory bandwidth bottlenecks. Studies indicate FPGAs can achieve comparable deep learning inference throughput while consuming up to 10×10\times less power than GPUs, natively supporting sparse networks, custom numeric precisions, and automotive functional safety standards.
  9. Knowl 9 — Camera vs. LiDAR Trade-Offs and Pseudo-LiDAR 3D Perception

    model/method

    Environmental perception relies on complementary sensor modalities with distinct physical trade-offs:

    • LiDAR: Emits pulsed laser beams to directly measure 3D spatial point clouds with up to 3 cm3\text{ cm} accuracy, 360∘360^\circ horizontal field-of-view, and ranges exceeding 200 m200\text{ m}. It operates effectively in complete darkness but suffers significant degradation in adverse weather (heavy rain, fog) and incurs high financial costs.
    • Camera: Provides rich semantic, color, and texture information at low hardware cost, but loses explicit depth during perspective projection onto 2D image sensors and is vulnerable to poor lighting and obstructed lenses.
    • Pseudo-LiDAR: Bridges the cost and modality gap by estimating dense depth maps from stereo or monocular visual images using deep networks, and reprojecting the predicted depth into 3D metric coordinates relative to the camera sensor. The resulting synthetic point cloud is processed directly using 3D point cloud architectures (e.g., PointNet, AVOD), approaching LiDAR-level 3D bounding box detection accuracy without requiring physical LiDAR sensors.
  10. Knowl 10 — Seven Open Challenges for Deep Learning in Autonomous Driving

    limitation

    The integration of deep learning across autonomous vehicle systems faces seven fundamental challenges:

    1. Perception Granularity: Scaling detection beyond simple 2D/3D bounding boxes to real-time, highly granular semantic tracking of complex road scenes, while closing the performance gap between camera and LiDAR 3D estimation.
    2. Short- to Middle-Term Temporal Reasoning: Transitioning from static scene perception to multi-agent negotiation and trajectory planning over millisecond-to-minute horizons.
    3. Training Data Scalability: Managing the collection, storage, and training parallelization of petabyte-scale on-vehicle sensor data, alongside mitigating sim-to-real domain gaps.
    4. Handling Driving Corner Cases: Enabling robust out-of-distribution generalization and developing few-shot or one-shot learning algorithms for rare, hazardous driving scenarios ("Black Swans").
    5. Learning-Based Control Integration: Combining data-driven system identification with established control theory (e.g., MPC) to ensure dynamic model adaptivity while maintaining hard stability and safety bounds.
    6. Functional Safety and Verification: Formulating explainability, verification, and fault-tolerance frameworks that reconcile stochastic, empirical neural networks with automotive functional safety standards (ISO 26262 / ASIL).
    7. Real-Time Embedded Compute and 5G Communication: Fulfilling ultra-low-latency processing constraints on resource-constrained embedded automotive hardware while enabling over-the-air parameter updates.

Coverage note — None was omitted; the knowls fully capture the paper's core survey taxonomy, mathematical models, controller formulations, datasets, hardware comparisons, safety frameworks, and open research challenges.

References

  1. 1.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Advances in Neural Information Processing Systems 25, F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2012, pp. 1097–1105.
  2. 2.M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba, “Learning Dexterous In-Hand Manipulation,” CoRR, vol. abs/1808.00177, August 2018. [Online]. Available: https://arxiv.org/abs/1808.00177
  3. 3.Y. Goldberg, Neural Network Methods for Natural Language Processing, ser. Synthesis Lectures on Human Language Technologies. Morgan & Claypool, 2017, vol. 37.
  4. 4.SAE Committee, “Taxonomy and Definitions for Terms Related to On-road Motor Vehicle Automated Driving Systems,” 2014.
  5. 5.E. Dickmanns and V. Graefe, “Dynamic Monocular Machine Vision,” Machine vision and applications, vol. 1, pp. 223–240, 1988.
  6. 6.B. Paden, M. Cáp, S. Z. Yong, D. S. Yershov, and E. Frazzoli, “A Survey of Motion Planning and Control Techniques for Self-Driving Urban Vehicles,” IEEE Trans. Intelligent Vehicles, vol. 1, no. 1, pp. 33–55, 2016.
  7. 7.Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based Learning Applied to Document Recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, Nov 1998.
  8. 8.Y. Bengio, A. Courville, and P. Vincent, “Representation Learning: A Review and New Perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 8, pp. 1798–1828, Aug 2013.
  9. 9.P. A. Viola and M. J. Jones, “Rapid Object Detection using a Boosted Cascade of Simple Features,” in 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2001), with CD-ROM, 8-14 December 2001, Kauai, HI, USA, 2001, pp. 511–518.
  10. 10.T. Ojala, M. Pietikäinen, and D. Harwood, “A Comparative Study of Texture Measures with Classification Based on Featured Distributions,” Pattern Recognition, vol. 29, no. 1, pp. 51–59, Jan. 1996.
  11. 11.N. Dalal and B. Triggs, “Histograms of Oriented Gradients for Human Detection,” in In CVPR, 2005, pp. 886–893.
  12. 12.D. H. Hubel and T. N.Wiesel, “Shape and Arrangement of Columns in Cats Striate Cortex,” The Journal of Physiology, vol. 165, no. 3, p. 559568, 1963.
  13. 13.M. A. Goodale and A. Milner, “Separate Visual Pathways for Perception and Action,” Trends in Neurosciences, vol. 15, no. 1, pp. 20 – 25, 1992.
  14. 14.D. E. Rumelhart, J. L. McClelland, and C. PDP Research Group, Eds., Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1: Foundations. Cambridge, MA, USA: MIT Press, 1986.
  15. 15.D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in 3rd Int. Conf. on Learning Representations, ICLR 2015, San Diego, CA, USA, May 2015.
  16. 16.E. H. J. Duchi and Y. Singer, “Adaptive Subgradient Methods for Online Learning and Stochastic Optimization,” Journal of Machine Learning Research, vol. 12, pp. 2121–2159, 2011.
  17. 17.S. Hochreiter and J. Schmidhuber, “Long Short-term Memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997.
  18. 18.R. Sutton and A. Barto, Introduction to Reinforcement Learning. MIT Press, 1998.
  19. 19.R. Bellman, Dynamic Programming. Princeton University Press, 1957.
  20. 20.C. Watkins and P. Dayan, “Q-Learning,” Machine Learning, vol. 8, no. 3, p. 279292, 1992.
  21. 21.V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level Control Through Deep Reinforcement Learning,” Nature, vol. 518, no. 7540, pp. 529–533, Feb. 2015.
  22. 22.M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver, “Rainbow: Combining Improvements in Deep Reinforcement Learning,” 2017.
  23. 23.A. E. Sallab, M. Abdou, E. Perot, and S. Yogamani, “Deep Reinforcement Learning framework for Autonomous Driving,” CoRR, vol. abs/1704.02532, 2017.
  24. 24.T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous Control with Deep Reinforcement Learning,” 2-4 May 2016.
  25. 25.S. Gu, T. Lillicrap, I. Sutskever, and S. Levine, “Continuous Deep Q-Learning with Model-based Acceleration,” in Int. Conf. on Machine Learning ICML 2016, vol. 48, Jun. 2016, pp. 2829–2838.
  26. 26.M. Jaritz, R. de Charette, M. Toromanoff, E. Perot, and F. Nashashibi, “End-to-End Race Driving with Deep Reinforcement Learning,” 2018 IEEE Int. Conf. on Robotics and Automation (ICRA), pp. 2070–2075, 2018.
  27. 27.M. Wulfmeier, D. Z. Wang, and I. Posner, “Watch This: Scalable Cost-Function Learning for Path Planning in Urban Environments,” 2016 IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS), vol. abs/1607.02329, 2016. [Online]. Available: http://arxiv.org/abs/1607.02329
  28. 28.H. Zhu, K.-V. Yuen, L. S. Mihaylova, and H. Leung, “Overview of Environment Perception for Intelligent Vehicles,” IEEE Transactions on Intelligent Transportation Systems, vol. 18, pp. 2584–2601, 2017.
  29. 29.J. Janai, F. Guney, A. Behl, and A. Geiger, “Computer Vision for Autonomous Vehicles: Problems, Datasets and State-of-the-Art,” 04 2017.
  30. 30.S. O’Kane, “How Tesla and Waymo are Tackling a Major Problem for Self-Driving Cars: Data,” Transportation, 2018.
  31. 31.S. Hasirlioglu, A. Kamann, I. Doric, and T. Brandmeier, “Test Methodology for Rain Influence on Automotive Surround Sensors,” in 2016 IEEE 19th Int. Conf. on Intelligent Transportation Systems (ITSC), Nov 2016, pp. 2242–2247.
  32. 32.Y. Wang, W.-L. Chao, D. Garg, B. Hariharan, M. Campbell, and K. Weinberger, “Pseudo-LiDAR from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving,” in IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) 2019, June 2019.
  33. 33.C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation,” in IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) 2017, July 2017.
  34. 34.J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. L. Waslander, “Joint 3D Proposal Generation and Object Detection from View Aggregation,” in IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS) 2018. IEEE, 2018.
  35. 35.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei, “ImageNet Large Scale Visual Recognition Challenge,” Int. Journal of Computer Vision (IJCV), vol. 115, no. 3, pp. 211–252, 2015.
  36. 36.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You Only Look Once: Unified, Real-time Object Detection,” in Proceedings of the IEEE Conf. on computer vision and pattern recognition, 2016, pp. 779–788.
  37. 37.H. Law and J. Deng, “Cornernet: Detecting Objects as Paired Keypoints,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 734–750.
  38. 38.S. Zhang, L. Wen, X. Bian, Z. Lei, and S. Z. Li, “Single-shot Refinement Neural Network for Object Detection,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  39. 39.R. Girshick, “Fast R-CNN,” in Proceedings of the IEEE Int. Conf. on computer vision, 2015, pp. 1440–1448.
  40. 40.F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “Squeezenet: Alexnet-level Accuracy with 50x Fewer Parameters and¡ 0.5 Mb Model Size,” arXiv preprint arXiv:1602.07360, 2016.
  41. 41.J. Dai, Y. Li, K. He, and J. Sun, “R-fcn: Object Detection via Region-based Fully Convolutional Networks,” in Advances in neural information processing systems, 2016, pp. 379–387.
  42. 42.V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, 2017.
  43. 43.H. Zhao, X. Qi, X. Shen, J. Shi, and J. Jia, “Icnet for Real-time Semantic Segmentation on High-resolution Images,” European Conference on Computer Vision, pp. 418–434, 2018.
  44. 44.M. Treml, J. A. Arjona-Medina, T. Unterthiner, R. Durgesh, F. Friedmann, P. Schuberth, A. Mayr, M. Heusel, M. Hofmarcher, M. Widrich, B. Nessler, and S. Hochreiter, “Speeding up Semantic Segmentation for Autonomous Driving,” 2016.
  45. 45.K. He, G. Gkioxari, P. Dollar, and R. B. Girshick, “Mask R-CNN,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 2980–2988, 2017.
  46. 46.Y. Zhou and O. Tuzel, “VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection,” IEEE Conf. on Computer Vision and Pattern Recognition 2018, pp. 4490–4499, 2018.
  47. 47.W. Luo, B. Yang, and R. Urtasun, “Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting With a Single Convolutional Net,” in IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) 2018, June 2018.
  48. 48.C. R. Qi, W. Liu, C. Wu, H. Su, and L. J. Guibas, “Frustum PointNets for 3D Object Detection from RGB-D Data,” in IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) 2018, June 2018.
  49. 49.X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-View 3D Object Detection Network for Autonomous Driving,” in IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) 2017, July 2017.
  50. 50.J. Redmon and A. Farhadi, “YOLO9000: Better, Faster, Stronger,” IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017.
  51. 51.——, “Yolov3: An Incremental Improvement,” arXiv preprint arXiv:1804.02767, 2018.
  52. 52.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single Shot Multibox Detector,” in European conference on computer vision. Springer, 2016, pp. 21–37.
  53. 53.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation,” in Proceedings of the 2014 IEEE Conf. on Computer Vision and Pattern Recognition, ser. CVPR ’14. Washington, DC, USA: IEEE Computer Society, 2014, pp. 580–587.
  54. 54.S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-time Object Detection with Region Proposal Networks,” IEEE Transactions on Pattern Analysis & Machine Intelligence, no. 6, pp. 1137–1149, 2017.
  55. 55.J. Li, K. Peng, and C.-C. Chang, “An Efficient Object Detection Algorithm Based on Compressed Networks,” Symmetry, vol. 10, no. 7, p. 235, 2018.
  56. 56.K. Shin, Y. P. Kwon, and M. Tomizuka, “RoarNet: A Robust 3D Object Detection based on RegiOn Approximation Refinement,” CoRR, vol. abs/1811.03818, 2018.
  57. 57.A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “Enet: A Deep Neural Network Architecture for Real-time Semantic Segmentation,” arXiv preprint arXiv:1606.02147, 2016.
  58. 58.A. Valada, J. Vertens, A. Dhall, and W. Burgard, “AdapNet: Adaptive Semantic Segmentation in Adverse Environmental Conditions,” 2017 IEEE Int. Conf. on Robotics and Automation (ICRA), pp. 4644–4651, 2017.
  59. 59.K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-scale Image Recognition,” arXiv preprint arXiv:1409.1556, 2014.
  60. 60.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper with Convolutions,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015.
  61. 61.K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE Conf. on computer vision and pattern recognition, 2016, pp. 770–778.
  62. 62.D. Barnes, W. Maddern, G. Pascoe, and I. Posner, “Driven to Distraction: Self-Supervised Distractor Learning for Robust Monocular Visual Odometry in Urban Environments,” in 2018 IEEE Int. Conf. on Robotics and Automation (ICRA). IEEE, 2018.
  63. 63.G. Bresson, Z. Alsayed, L. Yu, and S. Glaser, “Simultaneous Localization and Mapping: A Survey of Current Trends in Autonomous Driving,” IEEE Transactions on Intelligent Vehicles, vol. 2, no. 3, pp. 194–220, Sep 2017.
  64. 64.A. Kendall, M. Grimes, and R. Cipolla, “PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization,” in Proceedings of the 2015 IEEE Int. Conf. on Computer Vision (ICCV). Washington, DC, USA: IEEE Computer Society, 2015, pp. 2938–2946.
  65. 65.N. Radwan, A. Valada, and W. Burgard, “VLocNet++: Deep Multitask Learning for Semantic Visual Localization and Odometry,” IEEE Robotics and Automation Letters, Sep 2018.
  66. 66.F. Walch, C. Hazirbas, L. Leal-Taixé, T. Sattler, S. Hilsenbeck, and D. Cremers, “Image-Based Localization Using LSTMs for Structured Feature Correlation,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 627–637, 2017.
  67. 67.I. Melekhov, J. Ylioinas, J. Kannala, and E. Rahtu, “Image-Based Localization Using Hourglass Networks,” 2017 IEEE Int. Conf. on Computer Vision Workshops (ICCVW), pp. 870–877, 2017.
  68. 68.Z. Laskar, I. Melekhov, S. Kalia, and J. Kannala, “Camera Relocalization by Computing Pairwise Relative Poses Using Convolutional Neural Network,” in The IEEE Int. Conf. on Computer Vision (ICCV), Oct 2017.
  69. 69.E. Brachmann and C. Rother, “Learning Less is More 6D Camera Localization via 3D Surface Regression,” in IEEE Conf. on Computer Vision and Pattern Recognition (CVPR) 2018, June 2018.
  70. 70.P. Sarlin, F. Debraine, M. Dymczyk, R. Siegwart, and C. Cadena, “Leveraging Deep Visual Descriptors for Hierarchical Efficient Localization,” in Proc. of the 2nd Conf. on Robot Learning (CoRL), Oct 2018.
  71. 71.I. A. Barsan, S. Wang, A. Pokrovsky, and R. Urtasun, “Learning to Localize Using a LiDAR Intensity Map,” in Proc. of the 2nd Conf. on Robot Learning (CoRL), Oct 2018.
  72. 72.O. Garcia-Favrot and M. Parent, “Laser Scanner Based SLAM in Real Road and Traffic Environment,” in IEEE Int. Conf. Robotics and Automation (ICRA09). Workshop on Safe navigation in open and dynamic environments Application to autonomous vehicles, 2009.
  73. 73.A. K. Ushani and R. M. Eustice, “Feature Learning for Scene Flow Estimation from LIDAR,” in Proc. of the 2nd Conf. on Robot Learning (CoRL), vol. 87, Oct 2018, pp. 283–292.
  74. 74.Cityscapes, “Cityscapes Data Collection,” https://www.cityscapes-dataset.com/, 2018.
  75. 75.S. Thrun, W. Burgard, and D. Fox, “Probabilistic Robotics (Intelligent Robotics and Autonomous Agents),” in Cambridge: The MIT Press, 2005.
  76. 76.P. Ondruska, J. Dequaire, D. Z. Wang, and I. Posner, “End-to-End Tracking and Semantic Segmentation Using Recurrent Neural Networks,” CoRR, vol. abs/1604.05091, 2016.
  77. 77.S. Hoermann, M. Bach, and K. Dietmayer, “Dynamic Occupancy Grid Prediction for Urban Autonomous Driving: Deep Learning Approach with Fully Automatic Labeling,” IEEE Int. Conf. on Robotics and Automation (ICRA), 2017.
  78. 78.S. Ramos, S. K. Gehrig, P. Pinggera, U. Franke, and C. Rother, “Detecting Unexpected Obstacles for Self-Driving Cars: Fusing Deep Learning and Geometric Modeling,” IEEE Intelligent Vehicles Symposium, vol. 4, 2016.
  79. 79.C. Seeger, A. Mller, and L. Schwarz, “Towards Road Type Classification with Occupancy Grids,” in Intelligent Vehicles Symposium - Workshop: DeepDriving - Learning Representations for Intelligent Vehicles, IEEE, Gothenburg, Sweden, July 2016.
  80. 80.L. Marina, B. Trasnea, T. Cocias, A. Vasilcoi, F. Moldoveanu, and S. Grigorescu, “Deep Grid Net (DGN): A Deep Learning System for Real-Time Driving Context Understanding,” in Int. Conf. on Robotic Computing IRC 2019, Naples, Italy, 25-27 February 2019.
  81. 81.S. Shalev-Shwartz, S. Shammah, and A. Shashua, “Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving,” 2016.
  82. 82.S. D. Pendleton, H. Andersen, X. Du, X. Shen, M. Meghjani, Y. H. Eng, D. Rus, and M. H. Ang, “Perception, Planning, Control, and Coordination for Autonomous Vehicles,” Machines, vol. 5, no. 1, p. 6, 2017.
  83. 83.E. Rehder, J. Quehl, and C. Stiller, “Driving Like a Human: Imitation Learning for Path Planning using Convolutional Neural Networks,” in Int. Conf. on Robotics and Automation Workshops, 2017.
  84. 84.L. Sun, C. Peng, W. Zhan, and M. Tomizuka, “A Fast Integrated Planning and Control Framework for Autonomous Driving via Imitation Learning,” ASME 2018 Dynamic Systems and Control Conference, vol. 3, 2018. [Online]. Available: https://arxiv.org/pdf/1707.02515.pdf
  85. 85.S. Grigorescu, B. Trasnea, L. Marina, A. Vasilcoi, and T. Cocias, “NeuroTrajectory: A Neuroevolutionary Approach to Local State Trajectory Learning for Autonomous Vehicles,” IEEE Robotics and Automation Letters, vol. 4, no. 4, pp. 3441–3448, October 2019.
  86. 86.L. Yu, X. Shao, Y. Wei, and K. Zhou, “Intelligent Land-Vehicle Model Transfer Trajectory Planning Method Based on Deep Reinforcement Learning,” Sensors (Basel, Switzerland), vol. 18, 09 2018.
  87. 87.C. Paxton, V. Raman, G. D. Hager, and M. Kobilarov, “Combining Neural Networks and Tree Search for Task and Motion Planning in Challenging Environments,” 2017 IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS), vol. abs/1703.07887, 2017. [Online]. Available: http://arxiv.org/abs/1703.07887
  88. 88.W. Schwarting, J. Alonso-Mora, and D. Rus, “Planning and Decision-Making for Autonomous Vehicles,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 1, 05 2018.
  89. 89.T. Gu, J. M. Dolan, and J. Lee, “Human-like Planning of Swerve Maneuvers for Autonomous Vehicles,” in 2016 IEEE Intelligent Vehicles Symposium (IV), June 2016, pp. 716–721.
  90. 90.A. I. Panov, K. S. Yakovlev, and R. Suvorov, “Grid Path Planning with Deep Reinforcement Learning: Preliminary Results,” Procedia Computer Science, vol. 123, pp. 347 – 353, 2018, 8th Annual Int. Conf. on Biologically Inspired Cognitive Architectures, BICA 2017. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S1877050918300553
  91. 91.C. J. Ostafew, J. Collier, A. P. Schoellig, and T. D. Barfoot, “Learning-based Nonlinear Model Predictive Control to Improve Vision-based Mobile Robot Path Tracking,” Journal of Field Robotics, vol. 33, no. 1, pp. 133–152, 2015.
  92. 92.P. J. Nguyen-Tuong D and S. M, “Local Gaussian Process Regression for Real Time Online Model Learning,” in Proceedings of the neural information processing systems Conference, 2008, pp. 1193–1200.
  93. 93.H. P. Meier F and S. S, “Efficient Bayesian Local Model Learning for Control,” in IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS) 2016. IEEE, 2014, pp. 2244–2249.
  94. 94.C. J. Ostafew, A. P. Schoellig, and T. D. Barfoot, “Robust Constrained Learning-Based NMPC Enabling Reliable Mobile Robot Path Tracking,” Int. Journal of Robotics Research, vol. 35, no. 13, pp. 1547–1563, 2016.
  95. 95.O. Sigaud, C. Salaun, and V. Padois, “On-line Regression Algorithms for Learning Mechanical Models of Robots: A Survey,” Robotics and Autonomous Systems, vol. 59, no. 12, pp. 1115–1129, Dec. 2011.
  96. 96.C. Ostafew, A. Schoellig, and T. D. Barfoot, “Visual Teach and Repeat, Repeat, Repeat: Iterative Learning Control to Improve Mobile Robot Path Tracking in Challenging Outdoor Environments,” 11 2013, pp. 176–181.
  97. 97.B. Panomruttanarug, “Application of Iterative Learning Control in Tracking a Dubin’s Path in Parallel Parking,” Int. Journal of Automotive Technology, vol. 18, no. 6, pp. 1099–1107, Dec 2017.
  98. 98.N. R. Kapania and J. C. Gerdes, “Path Tracking of Highly Dynamic Autonomous Vehicle Trajectories via Iterative Learning Control,” in 2015 American Control Conference (ACC), July 2015, pp. 2753–2758.
  99. 99.Z. Yang, F. Zhou, Y. Li, and Y. Wang, “A Novel Iterative Learning Path-tracking Control for Nonholonomic Mobile Robots Against Initial Shifts,” Int. Journal of Advanced Robotic Systems, vol. 14, p. 172988141771063, 05 2017.
  100. 100.S. Lefvre, A. Carvalho, and F. Borrelli, “A Learning-Based Framework for Velocity Control in Autonomous Driving,” IEEE Transactions on Automation Science and Engineering, vol. 13, no. 1, pp. 32–42, Jan 2016.
  101. 101.S. Lefevre, A. Carvalho, and F. Borrelli, “Autonomous Car Following: A Learning-based Approach,” in 2015 IEEE Intelligent Vehicles Symposium (IV), June 2015, pp. 920–926.
  102. 102.P. Drews, G. Williams, B. Goldfain, E. A Theodorou, and J. M Rehg, “Aggressive Deep Driving: Combining Convolutional Neural Networks and Model Predictive Control,” 01 2017, pp. 133–142.
  103. 103.P. Drews, G. Williams, B. Goldfain, E. A. Theodorou, and J. M. Rehg, “Aggressive Deep Driving: Model Predictive Control with a CNN Cost Model,” CoRR, vol. abs/1707.05303, 2017.
  104. 104.U. Rosolia, A. Carvalho, and F. Borrelli, “Autonomous Racing using Learning Model Predictive Control,” in 2017 American Control Conference (ACC), May 2017, pp. 5115–5120.
  105. 105.Y. Pan, C.-A. Cheng, K. Saigol, K. Lee, X. Yan, E. A. Theodorou, and B. Boots, “Learning Deep Neural Network Control Policies for Agile Off-Road Autonomous Driving,” 2017.
  106. 106.Y. Pan, C. Cheng, K. Saigol, K. Lee, X. Yan, E. Theodorou, and B. Boots, “Agile Off-Road Autonomous Driving Using End-to-End Deep Imitation Learning,” Robotics: Science and Systems 2018, 2018.
  107. 107.J. Rawlings and D. Mayne, Model Predictive Control: Theory and Design. Nob Hill Pub., 2009.
  108. 108.M. Kamel, A. Hafez, and X. Yu, “A Review on Motion Control of Unmanned Ground and Aerial Vehicles Based on Model Predictive Control Techniques,” Engineering Science and Military Technologies, vol. 2, pp. 10–23, 03 2018.
  109. 109.M. Brunner, U. Rosolia, J. Gonzales, and F. Borrelli, “Repetitive Learning Model Predictive Control: An Autonomous Racing Example,” in 2017 IEEE 56th Annual Conference on Decision and Control (CDC), Dec 2017, pp. 2545–2550.
  110. 110.D. A. Pomerleau, “Alvinn: An autonomous Land Vehicle in a Neural Network,” in Advances in neural information processing systems, 1989, pp. 305–313.
  111. 111.U. Muller, J. Ben, E. Cosatto, B. Flepp, and Y. L. Cun, “Off-road Obstacle Avoidance through End-to-End Learning,” in Advances in neural information processing systems, 2006, pp. 739–746.
  112. 112.M. Bojarski, P. Yeres, A. Choromanska, K. Choromanski, B. Firner, L. Jackel, and U. Muller, “Explaining How a Deep Neural Network Trained with End-to-End Learning Steers a Car,” arXiv preprint arXiv:1704.07911, 2017.
  113. 113.H. Xu, Y. Gao, F. Yu, and T. Darrell, “End-to-End Learning of Driving Models from Large-scale Video Datasets,” IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017.
  114. 114.H. M. Eraqi, M. N. Moustafa, and J. Honer, “End-to-end Deep Learning for Steering Autonomous Vehicles Considering Temporal Dependencies,” arXiv preprint arXiv:1710.03804, 2017.
  115. 115.S. Hecker, D. Dai, and L. Van Gool, “End-to-End Learning of Driving Models with Surround-view Cameras and Route Planners,” in European Conference on Computer Vision (ECCV), 2018.
  116. 116.V. Rausch, A. Hansen, E. Solowjow, C. Liu, E. Kreuzer, and J. K. Hedrick, “Learning a Deep Neural Net Policy for End-to-End Control of Autonomous Vehicles,” in 2017 American Control Conference (ACC), May 2017, pp. 4914–4919.
  117. 117.M. G. Bechtel, E. McEllhiney, and H. Yun, “DeepPicar: A Low-cost Deep Neural Network-based Autonomous Car,” in The 24th IEEE Inter. Conf. on Embedded and Real-Time Computing Systems and Applications (RTCSA), August 2018, pp. 1–12.
  118. 118.S. Yang, W. Wang, C. Liu, K. Deng, and J. K. Hedrick, “Feature Analysis and Selection for Training an End-to-End Autonomous Vehicle Controller Using the Deep Learning Approach,” 2017 IEEE Intelligent Vehicles Symposium, vol. 1, 2017. [Online]. Available: http://arxiv.org/abs/1703.09744
  119. 119.M. Bojarski, D. D. Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Zieba, “End to End Learning for Self-Driving Cars,” CoRR, vol. abs/1604.07316, 2016. [Online]. Available: http://arxiv.org/abs/1604.07316
  120. 120.L. Fridman, D. E. Brown, M. Glazer, W. Angell, S. Dodd, B. Jenik, J. Terwilliger, J. Kindelsberger, L. Ding, S. Seaman, H. Abraham, A. Mehler, A. Sipperley, A. Pettinato, L. Angell, B. Mehler, and B. Reimer, “MIT Autonomous Vehicle Technology Study: Large-Scale Deep Learning Based Analysis of Driver Behavior and Interaction with Automation,” IEEE Access 2017, 2017. [Online]. Available: https://arxiv.org/abs/1711.06976
  121. 121.C. Chen, A. Seff, A. L. Kornhauser, and J. Xiao, “DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving,” 2015 IEEE Int. Conf. on Computer Vision (ICCV), pp. 2722–2730, 2015.
  122. 122.E. Perot, M. Jaritz, M. Toromanoff, and R. D. Charette, “End-to-End Driving in a Realistic Racing Game with Deep Reinforcement Learning,” in 2017 IEEE Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW), July 2017, pp. 474–475.
  123. 123.Wayve. (2018) Learning to Drive in a Day. [Online]. Available: https://wayve.ai/blog/learning-to-drive-in-a-day-with-reinforcement-learning
  124. 124.T. Zhang, G. Kahn, S. Levine, and P. Abbeel, “Learning Deep Control Policies for Autonomous Aerial Vehicles with MPC-guided Policy Search,” 2016 IEEE Int. Conf. on Robotics and Automation (ICRA), May 2016. [Online]. Available: http://dx.doi.org/10.1109/ICRA.2016.7487175
  125. 125.T. Ferrel, “Engineering Safety-critical Systems in the 21st Century,” 2010.
  126. 126.H. C. Burton S., Gauerhof L., “Making the Case for Safety of Machine Learning in Highly Automated Driving,” Lecture Notes in Computer Science, vol. 10489, 2017.
  127. 127.K. R. Varshney, “Engineering Safety in Machine Learning,” in 2016 Information Theory and Applications Workshop (ITA), Jan 2016, pp. 1–5.
  128. 128.D. Amodei, C. Olah, J. Steinhardt, P. F. Christiano, J. Schulman, and D. Mané, “Concrete Problems in AI Safety,” CoRR, vol. abs/1606.06565, 2016.
  129. 129.N. Möller, The Concepts of Risk and Safety. Springer Netherlands, 2012.
  130. 130.R. Salay, R. Queiroz, and K. Czarnecki, “An Analysis of ISO 26262: Using Machine Learning Safely in Automotive Software,” CoRR, vol. abs/1709.02435, 2017. [Online]. Available: http://arxiv.org/abs/1709.02435
  131. 131.S. Bernd, R. Detlev, E. Susanne, W. Ulf, B. Wolfgang, Patz, and Carsten, “Challenges in Applying the ISO 26262 for Driver Assistance Systems,” in Schwerpunkt Vernetzung, 5. Tagung Fahrerassistenz, 2012.
  132. 132.R. Parasuraman and V. Riley, “Humans and Automation: Use, Misuse, Disuse, Abuse,” Human Factors, vol. 39, no. 2, pp. 230–253, 1997.
  133. 133.F. Jose, Safety-Critical Systems, 2018.
  134. 134.H. Daumé, III and D. Marcu, “Domain Adaptation for Statistical Classifiers,” J. Artif. Int. Res., vol. 26, no. 1, pp. 101–126, May 2006.
  135. 135.R. Caruana, Y. Lou, J. Gehrke, P. Koch, M. Sturm, and N. Elhadad, “Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission,” in Proceedings of the 21th ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining, 2015, pp. 1721–1730.
  136. 136.K. R. Varshney and H. Alemzadeh, “On the Safety of Machine Learning: Cyber-Physical Systems, Decision Sciences, and Data Products,” Big data, vol. 5, 10 2016.
  137. 137.S. Levin, “Tesla Fatal Crash: ’Autopilot’ Mode Sped up Car Before Driver Killed, Report Finds,” The Guardian, 2018.
  138. 138.P. Koopman, “Challenges in Autonomous Vehicle Validation: Keynote Presentation Abstract,” in Proceedings of the 1st Int. Workshop on Safe Control of Connected and Autonomous Vehicles, 2017.
  139. 139.Z. Kurd, T. Kelly, and J. Austin, “Developing Artificial Neural Networks for Safety Critical Systems,” Neural Computing and Applications, vol. 16, no. 1, pp. 11–19, Jan 2007.
  140. 140.M. Harris, “Google Reports Self-driving Car Mistakes: 272 Failures and 13 Near Misses,” The Guardian, 2016.
  141. 141.J. McPherson, “How Uber’s Self-Driving Technology Could Have Failed In The Fatal Tempe Crash,” Forbes, 2018.
  142. 142.A. Chakarov, A. Nori, S. Rajamani, S. Sen, and D. Vijaykeerthy, “Debugging Machine Learning Tasks,” arXiv preprint arXiv:1603.07292, 2018.
  143. 143.B. Nushi, E. Kamar, E. Horvitz, and D. Kossmann, “On Human Intellect and Machine Failures: Troubleshooting Integrative Machine Learning Systems,” in AAAI, 2017.
  144. 144.I. Takanami, M. Sato, and Y. P. Yang, “A Fault-value Injection Approach for Multiple-weight-fault Tolerance of MNNs,” in Proceedings of the IEEE-INNS-ENNS, 2000, pp. 515–520 vol.3.
  145. 145.G. Katz, C. W. Barrett, D. L. Dill, K. Julian, and M. J. Kochenderfer, “Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks,” in CAV, 2017.
  146. 146.H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuScenes: A multimodal Dataset for Autonomous Driving,” arXiv preprint arXiv:1903.11027, 2019.
  147. 147.H. Yin and C. Berger, “When to Use what Data Set for Your Self-driving Car Algorithm: An Overview of Publicly Available Driving Datasets,” in 2017 IEEE 20th Int. Conf. on Intelligent Transportation Systems (ITSC), Oct 2017, pp. 1–8.
  148. 148.F. Yu, W. Xian, Y. Chen, F. Liu, M. Liao, V. Madhavan, and T. Darrell, “BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling,” CoRR, vol. abs/1805.04687, 2018.
  149. 149.P. Koschorrek, T. Piccini, P. berg, M. Felsberg, L. Nielsen, and R. Mester, “A Multi-sensor Traffic Scene Dataset with Omnidirectional Video,” in Ground Truth - What is a good dataset? CVPR Workshop 2013, 2013.
  150. 150.G. Pandey, J. R. McBride, and R. M. Eustice, “Ford Campus Vision and Lidar Data Set ,” Int. Journal of Robotics Research, vol. 30, no. 13, pp. 1543–1552, 2011.
  151. 151.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision Meets Robotics: The KITTI Dataset,” The Int. Journal of Robotics Research, vol. 32, no. 11, pp. 1231–1237, 2013.
  152. 152.Udacity, “Udacity Data Collection,” http://academictorrents.com/collection/self-driving-cars, 2018.
  153. 153.W. Maddern, G. Pascoe, C. Linegar, and P. Newman, “1 Year, 1000km: The Oxford RobotCar Dataset,” The Int. Journal of Robotics Research (IJRR), vol. 36, no. 1, pp. 3–15, 2017.
  154. 154.G. J. Brostow, J. Fauqueur, and R. Cipolla, “Semantic Object Classes in Video: A High-definition Ground Truth Database,” Pattern Recognition Letters, vol. 30, pp. 88–97, 2009.
  155. 155.F. Flohr and D. M. Gavrila, “Daimler Pedestrian Segmentation Benchmark Dataset,” in Proc. of the British Machine Vision Conference, 2013.
  156. 156.P. Dollar, C. Wojek, B. Schiele, and P. Perona, “Pedestrian Detection: A Benchmark,” in 2009 IEEE Conf. on Computer Vision and Pattern Recognition, 2009, pp. 304–311.
  157. 157.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid Scene Parsing Network,” in 2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 6230–6239.
  158. 158.S. Liu, J. Jia, S. Fidler, and R. Urtasun, “SGN: Sequential Grouping Networks for Instance Segmentation,” pp. 3516–3524, 10 2017.
  159. 159.X. Li, F. Flohr, Y. Yang, H. Xiong, M. Braun, S. Pan, K. Li, and D. M. Gavrila, “A New Benchmark for Vision-based Cyclist Detection,” in 2016 IEEE Intelligent Vehicles Symposium (IV), 2016, pp. 1028–1033.
  160. 160.Velodyne, “Velodyne LiDAR for Data Collection,” https://velodynelidar.com/, 2018.
  161. 161.Sick, “Sick LiDAR for Data Collection,” https://www.sick.com/, 2018.
  162. 162.NVIDIA, “NVIDIA AI Car Computer Drive PX,” https://www.nvidia.com/en-au/self-driving-cars/drive-px/.
  163. 163.——, “Tegra X2,” https://devblogs.nvidia.com/jetson-tx2-delivers-twice-intelligence-edge/.
  164. 164.——, “Denver Core,” https://en.wikichip.org/wiki/nvidia/microarchitectures/denver.
  165. 165.——, “Pascal Microarchitecture,” https://www.nvidia.com/en-us/data-center/pascal-gpu-architecture/.
  166. 166.——, “NVIDIA Drive AGX,” https://www.nvidia.com/en-us/self-driving-cars/drive-platform/hardware/.
  167. 167.——, “NVIDIA Volta,” https://www.nvidia.com/en-us/data-center/volta-gpu-architecture/.
  168. 168.Renesas, “R-Car V3H,” https://www.renesas.com/eu/en/solutions/automotive/soc/r-car-v3h.html/.
  169. 169.——, “R-Car H3,” https://www.renesas.com/sg/en/solutions/automotive/soc/r-car-h3.html/.
  170. 170.E. Nurvitadhi, G. Venkatesh, J. Sim, D. Marr, R. Huang, J. Ong Gee Hock, Y. T. Liew, K. Srivatsan, D. Moss, S. Subhaschandra, and G. Boudoukh, “Can FPGAs Beat GPUs in Accelerating Next-Generation Deep Neural Networks?” in Proceedings of the 2017 ACM/SIGDA Int. Symposium on Field-Programmable Gate Arrays, ser. FPGA ’17. New York, NY, USA: ACM, 2017, pp. 5–14. [Online]. Available: http://doi.acm.org/10.1145/3020078.3021740
  171. 171.K. Ovtcharov, O. Ruwase, J.-Y. Kim, J. Fowers, K. Strauss, and E. Chung, “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware,” February 2015.
  172. 172.J. Cong, Z. Fang, M. Lo, H. Wang, J. Xu, and S. Zhang, “Understanding Performance Differences of FPGAs and GPUs: (Abtract Only),” in Proceedings of the 2018 ACM/SIGDA Int. Symposium on Field-Programmable Gate Arrays, ser. FPGA ’18. New York, NY, USA: ACM, 2018, pp. 288–288. [Online]. Available: http://doi.acm.org/10.1145/3174243.3174970
  173. 173.Z.-Q. Zhao, P. Zheng, S.-t. Xu, and X. Wu, “Object Detection with Deep Learning: A Review,” IEEE transactions on neural networks and learning systems, 2018.
  174. 174.C. J. Ostafew, “Learning-based Control for Autonomous Mobile Robots,” Ph.D. dissertation, University of Toronto, 2016.

Citation

MLA
Grigorescu, S., et al. “A Survey of Deep Learning Techniques for Autonomous Driving”. Journal of Field Robotics, vol. 37, no. 3, 2019, pp. 362–86, https://doi.org/10.1002/rob.21918.
APA
Grigorescu, S., Trasnea, B., Cocias, T., & Macesanu, G. (2019). A survey of deep learning techniques for autonomous driving. Journal of Field Robotics, 37(3), 362–386. https://doi.org/10.1002/rob.21918
Chicago
Grigorescu, S., B. Trasnea, T. Cocias, and G. Macesanu. 2019. “A Survey of Deep Learning Techniques for Autonomous Driving”. Journal of Field Robotics 37 (3): 362–86. https://doi.org/10.1002/rob.21918.
Harvard
Grigorescu, S. et al. (2019) “A survey of deep learning techniques for autonomous driving”, Journal of Field Robotics, 37(3), pp. 362–386. Available at: https://doi.org/10.1002/rob.21918.
Vancouver
1. Grigorescu S, Trasnea B, Cocias T, Macesanu G (2019) A survey of deep learning techniques for autonomous driving. Journal of Field Robotics 37:362–386

BibTeX

@article{Grigorescu_2019, title={A survey of deep learning techniques for autonomous driving}, volume={37}, ISSN={1556-4967}, url={http://dx.doi.org/10.1002/rob.21918}, DOI={10.1002/rob.21918}, number={3}, journal={Journal of Field Robotics}, publisher={Wiley}, author={Grigorescu, Sorin and Trasnea, Bogdan and Cocias, Tiberiu and Macesanu, Gigel}, year={2019}, month=Nov, pages={362–386} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF