A survey of deep learning techniques for autonomous driving
Sorin GrigorescuBogdan TrasneaTiberiu CociasGigel Macesanu
Presents a structured comparative analysis of modular and end-to-end deep learning methods across perception, path planning, and control to guide system design and address key safety and hardware constraints in autonomous vehicles.
Autonomous vehicles have shifted from controlled research environments to real-world deployment on public roads, driven by the need to decrease traffic accidents, relieve congestion, and enhance urban mobility. While traditional perception and control methods effectively handle standard driving situations, they fail in complex corner cases that require human-like reasoning. This article evaluates the current state of deep learning technologies across autonomous driving architectures, examining perception, path planning, behavior arbitration, and motion control, while addressing systemic challenges in safety compliance, data collection, and embedded computing hardware.
The authors conducted a comprehensive survey of modern artificial intelligence methodologies applied to self-driving systems, comparing modular pipelines—where discrete tasks are executed by specialized deep learning or classical algorithms—against unified end-to-end systems that directly translate sensor streams into control signals. The survey examines core algorithmic paradigms including Convolutional Neural Networks for spatial understanding, Recurrent Neural Networks for temporal sequence tracking, and Deep Reinforcement Learning for decision-making policies. In addition, the article benchmarks object detection, segmentation, and localization models, while analyzing public datasets and target compute platforms like Graphics Processing Units and Field-Programmable Gate Arrays.
The findings establish that modular perception-planning-action architectures remain the dominant and most viable deployment model because they allow developers to incorporate established safety constraints and verify individual subsystems, unlike fully end-to-end models whose decision processes lack interpretability. In perception, single-stage object detectors deliver real-time operational speeds but sacrifice accuracy compared to double-stage detectors, while emerging pseudo-LiDAR techniques demonstrate the potential to derive three-dimensional spatial data from lower-cost camera sensors. In vehicle control, integrating machine learning with traditional methods—such as Model Predictive Control—yields superior stability and safety compared to pure reinforcement learning, which often learns biased behaviors when transferred from simulations to the real world. Furthermore, while Graphics Processing Units lead raw artificial intelligence processing, Field-Programmable Gate Arrays consume up to ten times less power and offer higher architectural flexibility and functional safety compliance for embedded in-vehicle integration.
These insights indicate that current industrial safety frameworks, such as the standard ISO 26262 for automotive functional safety, are structurally unequipped to certify machine learning systems that learn from data rather than explicit software specifications. This mismatch heightens operational risk, as real-world edge cases (often termed "Black Swans") absent from training sets can lead to catastrophic failures and severe liability. Hardware and sensor selection also directly impact production cost and scalability; while high-resolution LiDAR ensures robust three-dimensional perception, its high price tag currently hinders widespread commercial adoption compared to camera-first approaches.
To advance the safe commercialization of autonomous vehicles, industry stakeholders and policymakers should prioritize developing updated functional safety standards tailored to deep learning systems. Engineering teams should deploy hybrid architectures that wrap learning-based models with non-learnable deterministic safety monitors and explore low-power embedded hardware like Field-Programmable Gate Arrays for production deployment. Research efforts must focus on improving data collection for rare corner cases, advancing few-shot learning techniques, and refining camera-based three-dimensional depth perception to bridge simulation-to-reality performance gaps. Confidence in these findings is high regarding algorithmic trade-offs and compute characteristics, though caution is warranted regarding long-term reliability due to the fragmented nature of available training datasets and the persistent lack of explainability in deep neural networks.
- Paper: End to End Learning for Self-Driving Cars, Mariusz Bojarski et al. (2016). This seminal work demonstrates the practical realization of end-to-end convolutional neural networks mapping raw camera pixels to steering control, establishing a foundational paradigm surveyed in the source.
- Paper: CARLA: An Open Urban Driving Simulator, Alexey Dosovitskiy et al. (2017). This paper introduces the primary open-source simulation benchmark used across autonomous driving research to evaluate both modular pipelines and end-to-end learning agents.
- Paper: DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving, Chenyi Chen et al. (2015). This work introduces the direct perception paradigm via affordance learning, bridging modular scene understanding and end-to-end reflex driving surveyed in the source.
- Paper: A Survey of Motion Planning and Control Techniques for Self-Driving Urban Vehicles, Brian Paden et al. (2016). This survey provides essential background on classical motion planning, vehicle dynamics, and behavioral control hierarchies that form the baseline for deep learning-based planning systems.
- Paper: Multi-view 3D Object Detection Network for Autonomous Driving, Xiaozhi Chen et al. (2017). This foundational paper presents the multi-view deep sensor fusion framework for 3D bounding box detection, serving as a core perception technique surveyed in the source.
- Paper: Deep Reinforcement Learning: An Overview, Yuxi Li (2017). This overview introduces key deep reinforcement learning algorithms and stabilization strategies that underpin autonomous vehicle decision-making and control algorithms.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). This landmark paper establishes fully convolutional networks for dense pixel prediction, which forms the basis for driving scene semantic segmentation modules.
- Paper: Object Detection With Deep Learning: A Review, Zhong-Qiu Zhao et al. (2018). This review surveys modern convolutional object detection architectures, providing crucial prerequisite context on the vision backbones used in autonomous perception.
- Paper: The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes, Germán Ros et al. (2016). This work details the SYNTHIA synthetic dataset and establishes the methodology of using simulated urban imagery to train deep perception networks.
- Paper: Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey, Naveed Akhtar et al. (2018). This survey details adversarial vulnerabilities and safety risks in computer vision models, motivating the safety and robustness challenges analyzed in the source.
- Paper: Deep Reinforcement Learning for Autonomous Driving: A Survey, B Ravi Kiran et al. (2020). This survey specifically expands the deep reinforcement learning aspects of autonomous driving by evaluating sequential decision-making taxonomies, sim-to-real transfer, and hybrid policy deployments.
- Paper: Deep Learning for 3D Point Clouds: A Survey, Yulan Guo et al. (2019). This survey provides an advanced, dedicated examination of deep learning architectures tailored for raw and irregular 3D point cloud processing in robotic perception.
- Paper: A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges, M. Abdar et al. (2020). This review investigates uncertainty quantification methodologies in deep models, directly addressing the safety-critical challenges and failure modes highlighted in autonomous driving.
- Paper: Image Segmentation Using Deep Learning: A Survey, Shervin Minaee et al. (2020). This survey delivers an exhaustive follow-up on recent architectural advances in deep image and instance segmentation for autonomous scene parsing.
- Paper: A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects, Zewen Li et al. (2020). This paper expands on newer convolutional neural network designs, loss functions, and optimization advancements developed after the initial wave of driving perception models.
- Paper: A survey on large language model based autonomous agents, Lei Wang et al. (2023). This survey extends autonomous decision-making from specialized deep networks to generalized, large language model-driven planning and reasoning agents.
- Paper: Vision-Language-Action Models: Concepts, Progress, Applications and Challenges, Ranjan Sapkota et al. (2025). This work explores the frontier of vision-language-action foundation models, generalizing end-to-end robotic and vehicular control paradigms.
