DeepTest: Automated Testing of Deep-Neural-Network-Driven Autonomous Cars
Yuchi TianKexin PeiSuman JanaBaishakhi Ray
Proposes an automated testing tool that generates realistic driving conditions like rain and fog to maximize neuron coverage, successfully exposing thousands of dangerous corner-case failures in deep-learning-based autonomous vehicle models.
Autonomous vehicles rely heavily on deep neural networks to process sensor data and make real-time driving decisions. While these systems perform well under typical conditions, they frequently suffer from unexpected corner-case failures that can result in fatal collisions. Traditional testing methods depend on manually collecting driving data or unguided simulations, which are cost-prohibitive, miss rare environmental conditions, and fail to systematically evaluate the complex internal logic of deep learning software.
The article evaluates DeepTest, an automated testing tool designed to systematically detect erroneous behaviors in neural-network-driven autonomous vehicles. The tool assesses how realistic environmental changes trigger safety-critical errors and demonstrates how to generate high-coverage test cases without manual intervention.
The approach uses neuron coverage—the proportion of network units activated during execution—as a systematic guide to explore the decision-making logic of driving models. DeepTest generates realistic synthetic driving scenes by applying nine physical transformations to baseline camera images, including shifts in brightness, contrast, camera perspective, blurring, fog, and rain. The tool employs a coverage-guided search to stack these transformations and systematically activate untested parts of the network. Because human specifications are impractical to hand-code for every scenario, DeepTest identifies defects using domain-specific metamorphic relations, flagging instances where environmental changes cause the vehicle's steering angle to deviate significantly from the expected trajectory.
The evaluation across three top-performing models from the Udacity self-driving car competition yielded several key findings. First, neuron coverage strongly correlates with vehicle maneuvers such as steering angle and direction, proving it is an effective metric for guiding test generation. Second, combining image transformations via guided search improved neuron coverage by approximately 100% on average compared to original test images, increasing coverage by up to 22% over unguided transformations. Third, DeepTest successfully identified thousands of erroneous behaviors—including 6,339 violations under a balanced test setting—that could cause vehicles to drift off the road or crash. Finally, retraining one of the vulnerable models using the synthesized error cases reduced its prediction error by up to 46% under challenging weather conditions without sacrificing baseline performance.
These findings indicate that current autonomous vehicle models have severe, unaddressed vulnerabilities to routine environmental changes such as fog and rain. Uncovering these silent failures before public deployment is vital for mitigating safety risks, preventing costly real-world collisions, and ensuring compliance with emerging transportation safety standards. The results demonstrate that systematic, coverage-guided synthetic testing provides a scalable alternative to millions of manual physical test miles.
Organizations developing autonomous systems should integrate coverage-guided synthetic testing and metamorphic relation checks into their continuous validation pipelines. In addition, development teams should incorporate detected failure cases directly back into training datasets to harden model robustness against edge cases before on-road testing.
Confidence in these findings is strong regarding camera-based steering logic, supported by statistical validation and manual review that confirmed low false-positive rates. However, the study evaluated model responses to synthetic visual effects rather than live physical weather, and the current scope focuses solely on camera inputs and steering angles rather than braking, acceleration, or multi-sensor systems like LiDAR.
- Paper: DeepXplore: Automated Whitebox Testing of Deep Learning Systems, Kexin Pei et al. (2017). DeepXplore introduced the concept of neuron coverage and automated whitebox testing for deep learning systems, which DeepTest directly adapts and specializes for autonomous driving scenarios.
- Paper: End to End Learning for Self-Driving Cars, Mariusz Bojarski et al. (2016). This seminal paper demonstrated end-to-end deep learning for self-driving cars mapping camera images to steering angles, providing the foundational DNN driving architecture tested by DeepTest.
- Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). Szegedy et al. established the foundational vulnerability of deep neural networks to unexpected input perturbations and corner cases, motivating systematic testing frameworks like DeepTest.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This work explains how small, systematic modifications to inputs can cause dramatic failures in neural networks, underlying DeepTest's approach to synthesizing realistic corner-case conditions.
- Paper: DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving, Chenyi Chen et al. (2015). DeepDriving established direct perception and affordance learning paradigms for vision-based autonomous driving models evaluated in automated testing research.
- Paper: Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, Dan Hendrycks et al. (2019). This work systematically formalizes and benchmarks neural network robustness against common real-world corruptions and weather perturbations, generalizing the synthetic transformations pioneered by DeepTest.
- Paper: Model Assertions for Monitoring and Improving ML Models, Daniel Kang et al. (2020). Model assertions build on automated testing principles to establish continuous runtime monitoring and oracle checks for detecting behavioral bugs in vision pipelines.
- Paper: Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift, Stephan Rabanser et al. (2019). This study analyzes methods for detecting dataset shifts and silent failures in deployed models, addressing the real-world distributional changes exposed by DeepTest.
- Paper: BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning, Fisher Yu et al. (2020). BDD100K provides a large-scale real-world dataset covering diverse weather and lighting conditions to evaluate driving model performance under the exact scenarios DeepTest synthesizes.
- Paper: A survey of deep learning techniques for autonomous driving, Sorin Grigorescu et al. (2019). This comprehensive survey provides an overview of deep learning techniques and safety challenges across autonomous driving architectures following early automated testing advances.
- Paper: A Survey of Autonomous Driving: Common Practices and Emerging Technologies, Ekim Yurtsever et al. (2019). This survey evaluates broad autonomous driving practices, system-level safety issues, and testing bottlenecks in modern intelligent vehicles.
