DeepTest: Automated Testing of Deep-Neural-Network-Driven Autonomous Cars

Yuchi TianKexin PeiSuman JanaBaishakhi Ray

article2017International Conference on Software Engineering1,533 citations

Proposes an automated testing tool that generates realistic driving conditions like rain and fog to maximize neuron coverage, successfully exposing thousands of dangerous corner-case failures in deep-learning-based autonomous vehicle models.

Listen

Autonomous vehicles rely heavily on deep neural networks to process sensor data and make real-time driving decisions. While these systems perform well under typical conditions, they frequently suffer from unexpected corner-case failures that can result in fatal collisions. Traditional testing methods depend on manually collecting driving data or unguided simulations, which are cost-prohibitive, miss rare environmental conditions, and fail to systematically evaluate the complex internal logic of deep learning software.

The article evaluates DeepTest, an automated testing tool designed to systematically detect erroneous behaviors in neural-network-driven autonomous vehicles. The tool assesses how realistic environmental changes trigger safety-critical errors and demonstrates how to generate high-coverage test cases without manual intervention.

The approach uses neuron coverage—the proportion of network units activated during execution—as a systematic guide to explore the decision-making logic of driving models. DeepTest generates realistic synthetic driving scenes by applying nine physical transformations to baseline camera images, including shifts in brightness, contrast, camera perspective, blurring, fog, and rain. The tool employs a coverage-guided search to stack these transformations and systematically activate untested parts of the network. Because human specifications are impractical to hand-code for every scenario, DeepTest identifies defects using domain-specific metamorphic relations, flagging instances where environmental changes cause the vehicle's steering angle to deviate significantly from the expected trajectory.

The evaluation across three top-performing models from the Udacity self-driving car competition yielded several key findings. First, neuron coverage strongly correlates with vehicle maneuvers such as steering angle and direction, proving it is an effective metric for guiding test generation. Second, combining image transformations via guided search improved neuron coverage by approximately 100% on average compared to original test images, increasing coverage by up to 22% over unguided transformations. Third, DeepTest successfully identified thousands of erroneous behaviors—including 6,339 violations under a balanced test setting—that could cause vehicles to drift off the road or crash. Finally, retraining one of the vulnerable models using the synthesized error cases reduced its prediction error by up to 46% under challenging weather conditions without sacrificing baseline performance.

These findings indicate that current autonomous vehicle models have severe, unaddressed vulnerabilities to routine environmental changes such as fog and rain. Uncovering these silent failures before public deployment is vital for mitigating safety risks, preventing costly real-world collisions, and ensuring compliance with emerging transportation safety standards. The results demonstrate that systematic, coverage-guided synthetic testing provides a scalable alternative to millions of manual physical test miles.

Organizations developing autonomous systems should integrate coverage-guided synthetic testing and metamorphic relation checks into their continuous validation pipelines. In addition, development teams should incorporate detected failure cases directly back into training datasets to harden model robustness against edge cases before on-road testing.

Confidence in these findings is strong regarding camera-based steering logic, supported by statistical validation and manual review that confirmed low false-positive rates. However, the study evaluated model responses to synthetic visual effects rather than live physical weather, and the current scope focuses solely on camera inputs and steering angles rather than braking, acceleration, or multi-sensor systems like LiDAR.

  • Paper: DeepXplore: Automated Whitebox Testing of Deep Learning Systems, Kexin Pei et al. (2017). DeepXplore introduced the concept of neuron coverage and automated whitebox testing for deep learning systems, which DeepTest directly adapts and specializes for autonomous driving scenarios.
  • Paper: End to End Learning for Self-Driving Cars, Mariusz Bojarski et al. (2016). This seminal paper demonstrated end-to-end deep learning for self-driving cars mapping camera images to steering angles, providing the foundational DNN driving architecture tested by DeepTest.
  • Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). Szegedy et al. established the foundational vulnerability of deep neural networks to unexpected input perturbations and corner cases, motivating systematic testing frameworks like DeepTest.
  • Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This work explains how small, systematic modifications to inputs can cause dramatic failures in neural networks, underlying DeepTest's approach to synthesizing realistic corner-case conditions.
  • Paper: DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving, Chenyi Chen et al. (2015). DeepDriving established direct perception and affordance learning paradigms for vision-based autonomous driving models evaluated in automated testing research.
Cover for DeepTest: Automated Testing of Deep-Neural-Network-Driven Autonomous Cars

Abstract

Recent advances in Deep Neural Networks (DNNs) have led to the development of DNN-driven autonomous cars that, using sensors like camera, LiDAR, etc., can drive without any human intervention. Most major manufacturers including Tesla, GM, Ford, BMW, and Waymo/Google are working on building and testing different types of autonomous vehicles. The lawmakers of several US states including California, Texas, and New York have passed new legislation to fast-track the process of testing and deployment of autonomous vehicles on their roads.

However, despite their spectacular progress, DNNs, just like traditional software, often demonstrate incorrect or unexpected corner case behaviors that can lead to potentially fatal collisions. Several such real-world accidents involving autonomous cars have already happened including one which resulted in a fatality. Most existing testing techniques for DNN-driven vehicles are heavily dependent on the manual collection of test data under different driving conditions which become prohibitively expensive as the number of test conditions increases.

In this paper, we design, implement and evaluate DeepTest, a systematic testing tool for automatically detecting erroneous behaviors of DNN-driven vehicles that can potentially lead to fatal crashes. First, our tool is designed to automatically generated test cases leveraging real-world changes in driving conditions like rain, fog, lighting conditions, etc. DeepTest systematically explores different parts of the DNN logic by generating test inputs that maximize the numbers of activated neurons. DeepTest found thousands of erroneous behaviors under different realistic driving conditions (e.g., blurring, rain, fog, etc.) many of which lead to potentially fatal crashes in three top performing DNNs in the Udacity self-driving car challenge.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Deep Learning for Autonomous Driving
  • 2.2 Different DNN Architectures
  • 3 Methodology
  • 3.1 Systematic Testing with Neuron Coverage
  • 3.2 Increasing Coverage with Synthetic Images
  • 3.3 Combining Transformations to Increase Coverage
  • 3.4 Creating a Test Oracle with Metamorphic Relations
  • 4 Implementation
  • 5 Results
  • 6 Threats to Validity
  • 7 Related Work
  • 8 Conclusion
  • 9 Acknowledgements
  • References

Knowls

  1. Knowl 1 — Neuron Coverage Formulation for Feed-Forward, Convolutional, and Recurrent Neural Networks

    definition

    Neuron coverage measures the proportion of internal computing units (neurons) in a deep neural network (DNN) activated by a given test suite:

    Neuron Coverage=∣Activated Neurons∣∣Total Neurons∣\text{Neuron Coverage} = \frac{|\text{Activated Neurons}|}{|\text{Total Neurons}|}

    Activation is evaluated based on the layer and architecture type:

    • Fully Connected Layers: A neuron produces a scalar activation output yy. The neuron is classified as activated if yy, scaled relative to the layer's output distribution, exceeds a predefined DNN activation threshold τ=0.2\tau = 0.2.
    • Convolutional Layers: A convolutional neuron produces a multidimensional feature map of dimensions H×WH \times W by convolving its kernel across the input space. To evaluate activation, the feature map is reduced to a single scalar by taking the arithmetic mean of all H×WH \times W elements, which is then compared against the activation threshold τ=0.2\tau = 0.2.
    • Recurrent Layers (RNN/LSTM): Recurrent architectures with internal loops across sequence steps are unrolled across the sequence of time steps. Every intermediate neuron at each unrolled time step is evaluated as an independent individual neuron for the calculation of total neuron coverage.
  2. Knowl 2 — Metamorphic Testing Oracle and Transformation Filter for Autonomous Steering

    equation

    In end-to-end autonomous driving systems lacking exact ground-truth outputs for synthesized driving environments, metamorphic testing relations relaxed by baseline model error serve as automated test oracles.

    Let IoI_o denote a seed image with human ground-truth steering angle θ^\hat{\theta}, and let θo\theta_o denote the model's predicted steering angle for IoI_o. For a set of nn original seed images, the baseline Mean Squared Error (MSEorig\text{MSE}_{orig}) is defined as:

    MSEorig=1n∑i=1n(θ^i−θoi)2\text{MSE}_{orig} = \frac{1}{n} \sum_{i=1}^n (\hat{\theta}_i - \theta_{oi})^2

    When a realistic transformation tt (such as lighting adjustment, weather synthesis, or mild affine transformation) is applied to IoI_o to produce synthetic image ItI_t, the car's predicted steering angle θt\theta_t should remain consistent. An erroneous behavior (metamorphic violation) on transformed input ItiI_{ti} is reported if:

    (θ^i−θti)2>λ MSEorig(\hat{\theta}_i - \theta_{ti})^2 > \lambda \, \text{MSE}_{orig}

    where λ≥1\lambda \ge 1 is a configurable multiplier balancing false positives and false negatives (e.g., λ=5\lambda = 5).

    To prevent false positives from transformations that naturally require altered driving trajectories (such as large image rotations), a transformation filter criterion is applied:

    ∣MSE(trans,param)−MSEorig∣≤ϵ|\text{MSE}_{(trans, param)} - \text{MSE}_{orig}| \le \epsilon

    where MSE(trans,param)\text{MSE}_{(trans, param)} is the mean squared error over the test set under transformation transtrans with parameter paramparam, and ϵ\epsilon is an error tolerance threshold (e.g., ϵ=0.03\epsilon = 0.03). Only transformations satisfying this constraint are flagged as bugs when violating the metamorphic relation.

  3. Knowl 3 — Neuron-Coverage-Guided Greedy Search for Combining Image Transformations

    algorithm

    DeepTest combines multiple parameterized image transformations to maximize deep neural network neuron coverage using a depth-first greedy search guided by coverage increments.

    Input: Transformations TT, Seed images II, Maximum failed attempts maxFailedTriesmaxFailedTries
    Output: Synthetically generated test images genTestsgenTests
    Stack S←∅S \leftarrow \emptyset
    for each image img∈Iimg \in I do
        S.push(img)S.push(img)
    end for
    genTests←∅genTests \leftarrow \emptyset
    while SS is not empty do
        img←S.pop()img \leftarrow S.pop()
        TransformationQueue Tqueue←∅Tqueue \leftarrow \emptyset
        numFailedTries←0numFailedTries \leftarrow 0
        while numFailedTries≤maxFailedTriesnumFailedTries \le maxFailedTries do
            if TqueueTqueue is not empty then
                T1←Tqueue.dequeue()T_1 \leftarrow Tqueue.dequeue()
            else
                T1←Randomly select from TT_1 \leftarrow \text{Randomly select from } T
            end if
            P1←Randomly select valid parameter for T1P_1 \leftarrow \text{Randomly select valid parameter for } T_1
            T2←Randomly select from TT_2 \leftarrow \text{Randomly select from } T
            P2←Randomly select valid parameter for T2P_2 \leftarrow \text{Randomly select valid parameter for } T_2
            newImage←ApplyTransforms(img,T1,P1,T2,P2)newImage \leftarrow \text{ApplyTransforms}(img, T_1, P_1, T_2, P_2)
            if CoverageIncreases(newImage)\text{CoverageIncreases}(newImage) then
                Tqueue.enqueue(T1)Tqueue.enqueue(T_1)
                Tqueue.enqueue(T2)Tqueue.enqueue(T_2)
                UpdateCoverage()\text{UpdateCoverage}()
                genTests←genTests∪{newImage}genTests \leftarrow genTests \cup \{newImage\}
                S.push(newImage)S.push(newImage)
            else
                numFailedTries←numFailedTries+1numFailedTries \leftarrow numFailedTries + 1
            end if
        end while
    end while
    return genTestsgenTests

    The procedure maintains a queue TqueueTqueue to prioritize transformation types that previously succeeded in increasing neuron coverage on the current image path before exploring randomized transformations.

  4. Knowl 4 — Realistic Image Transformation Parameterizations for Autonomous Driving Testing

    model/method

    DeepTest parameterizes nine realistic image transformations across three categories to systematically synthesize driving variations:

    1. Linear Transformations:

      • Brightness Adjustment: Pixel intensity modified by bias parameter β∈[10,100]\beta \in [10, 100] in steps of 10 (Inew=I+βI_{new} = I + \beta).
      • Contrast Adjustment: Pixel intensity scaled by gain factor α∈[1.2,3.0]\alpha \in [1.2, 3.0] in steps of 0.2 (Inew=α⋅II_{new} = \alpha \cdot I).
    2. Affine Transformations: Represented by a 2×32 \times 3 transformation matrix MM applied to image coordinates:

      • Translation: Horizontal and vertical shifts (tx,ty)∈[(10,10),(100,100)](t_x, t_y) \in [(10, 10), (100, 100)] with step (10,10)(10, 10), using matrix M=[10tx01ty]M = \begin{bmatrix} 1 & 0 & t_x \\ 0 & 1 & t_y \end{bmatrix}.
      • Scaling: Zoom factors (sx,sy)∈[(1.5,1.5),(6.0,6.0)](s_x, s_y) \in [(1.5, 1.5), (6.0, 6.0)] with step (0.5,0.5)(0.5, 0.5), using matrix M=[sx000sy0]M = \begin{bmatrix} s_x & 0 & 0 \\ 0 & s_y & 0 \end{bmatrix}.
      • Horizontal Shearing: Shear factors (sx,sy)∈[(−1.0,0),(−0.1,0)](s_x, s_y) \in [(-1.0, 0), (-0.1, 0)] with step (0.1,0)(0.1, 0), using matrix M=[1sx0sy10]M = \begin{bmatrix} 1 & s_x & 0 \\ s_y & 1 & 0 \end{bmatrix}.
      • Rotation: Rotation angle q∈[3∘,30∘]q \in [3^\circ, 30^\circ] with step 3∘3^\circ, using matrix M=[cos⁡q−sin⁡q0sin⁡qcos⁡q0]M = \begin{bmatrix} \cos q & -\sin q & 0 \\ \sin q & \cos q & 0 \end{bmatrix}.
    3. Convolutional Transformations:

      • Blur Filters: Averaging (kernel sizes 3×3,4×4,5×5,6×63\times3, 4\times4, 5\times5, 6\times6), Gaussian (kernel sizes 3×3,5×5,7×73\times3, 5\times5, 7\times7), Median (aperture linear sizes 3, 5), and Bilateral Filtering (diameter =9= 9, σColor=75\sigma_{\text{Color}} = 75, σSpace=75\sigma_{\text{Space}} = 75).
      • Weather Synthesis: Multi-filter compositions simulating realistic rain streaks and fog/mist attenuation.
  5. Knowl 5 — Statistical Correlation between Neuron Coverage and Steering Output Variations

    empirical result

    Neuron coverage exhibits statistically significant correlations with the behavioral outputs of autonomous vehicle DNNs, demonstrating its validity as a test guidance metric:

    • Steering Angle (Continuous Variable): Evaluating the Spearman rank correlation ρ\rho between neuron coverage and continuous steering angle yielded statistically significant associations (p<2.2×10−16p < 2.2 \times 10^{-16}) across models:

      • Chauffeur model (CNN + LSTM): overall ρ=−0.10\rho = -0.10 (CNN component: ρ=0.28\rho = 0.28; LSTM component: ρ=−0.10\rho = -0.10).
      • Rambo model (ensemble of 3 CNNs): overall ρ=−0.11\rho = -0.11 (Sub-model S1: ρ=−0.19\rho = -0.19; S2: ρ=0.10\rho = 0.10; S3: ρ=−0.11\rho = -0.11).
      • Epoch model (single CNN): ρ=0.78\rho = 0.78.
    • Steering Direction (Binary Variable): Using the Wilcoxon signed-rank test, neuron coverage distributions differed significantly between left turns (positive steering angles) and right turns (negative steering angles) with p<2.2×10−16p < 2.2 \times 10^{-16} for Chauffeur, Epoch, and Rambo overall. In Rambo, sub-model S1 exhibited a large effect size (Cohen's dd), while sub-models S2 and S3 showed negligible direction correlation, demonstrating that specific sub-networks dominate directional decisions.

  6. Knowl 6 — Neuron Activation Diversity across Distinct Image Transformations

    empirical result

    Different synthetic image transformations applied to the same seed image activate substantially distinct subsets of neurons within autonomous driving neural networks.

    Evaluating pairs of transformations (T1,T2T_1, T_2) across 1,000 seed images (70,000 total transformed images) using the Jaccard distance between activated neuron sets N1N_1 and N2N_2:

    1−∣N1∩N2∣∣N1∪N2∣1 - \frac{|N_1 \cap N_2|}{|N_1 \cup N_2|}

    yielded high median dissimilarity across CNN architectures:

    • Chauffeur-CNN: median Jaccard distance of 0.53
    • Epoch CNN: median Jaccard distance of 0.67
    • Rambo-S1: median Jaccard distance of 0.12
    • Rambo-S2: median Jaccard distance of 0.17
    • Rambo-S3: median Jaccard distance of 0.30
    • Chauffeur-LSTM: median Jaccard distance of 0.002 (low variation caused by sequence-level state retention across frames).

    Additionally, evaluating cumulative neuron activations per seed image as successive transformations are applied demonstrates monotonic increases in cumulative coverage across all tested architectures.

  7. Knowl 7 — Neuron Coverage Improvements from Cumulative and Coverage-Guided Transformations

    data/table

    Combining image transformations systematically improves cumulative neuron coverage over baseline seed images, with coverage-guided search achieving superior coverage using substantially fewer test images.

    Model Baseline Coverage Cumulative Trans. Guided Generation % Incr. vs Baseline % Incr. vs Cumulative
    Chauffeur-CNN 658 (46%) 1,065 (75%) 1,250 (88%) 90% 17%
    Epoch 621 (25%) 1,034 (41%) 1,266 (51%) 104% 22%
    Rambo-S1 710 (44%) 929 (57%) 1,043 (64%) 47% 12%
    Rambo-S2 1,146 (30%) 2,210 (58%) 2,676 (70%) 134% 21%
    Rambo-S3 13,008 (97%) 13,080 (97%) 13,150 (98%) 1.1% 0.5%

    Results were evaluated on 100 seed images. Cumulative Transformation applies 7 distinct transformations with 10 parameter values (7,000 synthetic images). Guided Generation uses coverage-directed search to generate 254 images for Chauffeur-CNN, 221 for Epoch, and 864 for Rambo, increasing neuron coverage by up to 104% over the seed baseline and 12% to 22% over unguided cumulative transformations.

  8. Knowl 8 — Erroneous Behavior Detection and False Positive Rates across Driving Models

    empirical result

    Using the metamorphic testing oracle with error multiplier λ=5\lambda = 5 and filter threshold ϵ=0.03\epsilon = 0.03, DeepTest detected 6,339 unique erroneous behaviors across three Udacity competition models:

    • Simple Transformations (330 errors total):

      • Blur: 3 in Chauffeur, 27 in Epoch, 11 in Rambo (41 total)
      • Brightness: 97 in Chauffeur, 32 in Epoch, 15 in Rambo (144 total)
      • Contrast: 31 in Chauffeur, 12 in Epoch, 0 in Rambo (43 total)
      • Rotation: 0 in Chauffeur, 13 in Epoch, 0 in Rambo (13 total)
      • Scale: 0 in Chauffeur, 10 in Epoch, 0 in Rambo (10 total)
      • Shear: 0 in Chauffeur, 0 in Epoch, 23 in Rambo (23 total)
      • Translation: 21 in Chauffeur, 35 in Epoch, 0 in Rambo (56 total)
    • Composite and Guided Transformations (6,009 errors total):

      • Rain: 650 in Chauffeur, 64 in Epoch, 27 in Rambo (741 total)
      • Fog: 201 in Chauffeur, 135 in Epoch, 4,112 in Rambo (4,448 total)
      • Guided Search: 89 in Chauffeur, 65 in Epoch, 666 in Rambo (820 total)

    Manual inspection of all reported errors under λ=5,ϵ=0.03\lambda = 5, \epsilon = 0.03 identified 130 false positives (14 in Epoch, 26 in Chauffeur, 90 in Rambo), confirming low false positive rates.

  9. Knowl 9 — Model Robustness and Accuracy Improvement via Synthetic Test Retraining

    data/table

    Retraining self-driving DNNs on synthetic driving condition test cases generated by DeepTest significantly improves model accuracy under degraded environmental conditions without hurting clean-data accuracy.

    Test Set Original MSE Retrained MSE
    Original images 0.10 0.09
    With fog 0.18 0.10
    With rain 0.13 0.07

    The Epoch driving model was retrained using its original training data augmented with 66% randomly sampled synthetic fog and rain images generated from the Udacity HMB_3.bag dataset. Evaluated on the held-out 34% test set, retraining reduced Mean Squared Error on foggy scenes from 0.18 to 0.10 (a 44.4% reduction) and on rainy scenes from 0.13 to 0.07 (a 46.2% reduction), while baseline error on clean images improved slightly from 0.10 to 0.09.

  10. Knowl 10 — Evaluated Autonomous Driving DNN Architectures and Dataset Configuration

    experimental setup

    DeepTest evaluated three top-ranking driving models from the Udacity Self-Driving Car Challenge 2:

    1. Chauffeur (3rd place): Comprises a CNN visual feature extractor (1,427 neurons) and an LSTM sequence network (513 neurons). The CNN extracts visual features per image, and 100 features extracted from 100 consecutive frames are concatenated as input to the LSTM to predict the steering angle. Implemented in Keras and TensorFlow.
    2. Rambo (2nd place): An ensemble of three CNN sub-models: S1 (1,625 neurons), S2 (3,801 neurons), and S3 (13,473 neurons), merged via a final output layer. Rambo takes pixel differences across three consecutive frames as input. Implemented in Keras and Theano.
    3. Epoch (6th place): A single CNN model (2,500 neurons) trained on the Udacity CH2_002 dataset using Keras and TensorFlow.

    Steering Prediction Setup: Models take camera images and output a steering angle normalized by 1/251/25 to the range [−1.0,1.0][-1.0, 1.0], corresponding to maximum hardware steering angles of [−25∘,+25∘][-25^\circ, +25^\circ] (negative values indicate turning left, positive values indicate turning right). Evaluation is performed on Udacity dataset HMB_3.bag.

  11. Knowl 11 — Testing and Realism Limitations in DeepTest

    limitation

    The testing methodology of DeepTest is subject to three primary validity constraints:

    1. Control Action Scope: DeepTest evaluates only steering angle prediction accuracy because the open-source Udacity benchmark models do not implement braking, acceleration, or throttle control.
    2. Synthetic Image Fidelity: Synthetic digital transformations (e.g., Photoshop-based fog and rain filters) approximate weather and lighting conditions but do not simulate complex physical effects such as ray scattering, lens droplets, puddle reflections, or changing sun angles.
    3. Transformation Coverage: The set of nine linear, affine, and convolutional transformations does not provide exhaustive coverage of all possible real-world driving environments and hardware sensor anomalies.

Coverage note — None was omitted; all key contributed definitions, algorithms, metamorphic relations, empirical findings, and experimental configurations are fully represented.

References

  1. 1.
    1. Add Dramatic Rain to a Photo in Photoshop. https://design.tutsplus.com/tutorials/add-dramatic-rain-to-a-photo-in-photoshop--psd-29536. (2013).
  2. 2.
    1. How to create mist: Photoshop effects for atmospheric landscapes. http://www.techradar.com/how-to/photography-video-capture/cameras/how-to-create-mist-photoshop-effects-for-atmospheric-landscapes-1320997. (2013).
  3. 3.
    1. The OpenCV Reference Manual (2.4.9.0 ed.).
  4. 4.
    1. This Is How Bad Self-Driving Cars Suck In The Rain. http://jalopnik.com/this-is-how-bad-self-driving-cars-suck-in-the-rain-1666268433. (2014).
  5. 5.
    1. Affine Transformation. https://www.mathworks.com/discovery/affine-transformation.html. (2015).
  6. 6.
    1. Affine Transformations. http://docs.opencv.org/3.1.0/d4/d61/tutorial_warp_affine.html. (2015).
  7. 7.
    1. Open Source Computer Vision Library. https://github.com/itseez/opencv. (2015).
  8. 8.
    1. Chauffeur model. https://github.com/udacity/self-driving-car/tree/master/steering-models/community-models/chauffeur. (2016).
  9. 9.
    1. comma.ai’s steering model. https://github.com/commaai/research/blob/master/train_steering_model.py. (2016).
  10. 10.
    1. Epoch model. https://github.com/udacity/self-driving-car/tree/master/steering-models/community-models/cg23. (2016).
  11. 11.
    1. Google Auto Waymo Disengagement Report for Autonomous Driving. https://www.dmv.ca.gov/portal/wcm/connect/946b3502-c959-4e3b-b119-91319c27788f/GoogleAutoWaymo_disengage_report_2016.pdf?MOD=AJPERES. (2016).
  12. 12.
    1. Google’s Self-Driving Car Caused Its First Crash. https://www.wired.com/2016/02/googles-self-driving-car-may-caused-first-crash/. (2016).
  13. 13.
    1. Rambo model. https://github.com/udacity/self-driving-car/tree/master/steering-models/community-models/rambo. (2016).
  14. 14.
    1. Tesla Autopilot. https://www.tesla.com/autopilot. (2016).
  15. 15.
    1. Udacity self driving car challenge 2. https://github.com/udacity/self-driving-car/tree/master/challenges/challenge-2. (2016).
  16. 16.
    1. Udacity self driving car challenge 2 dataset. https://github.com/udacity/self-driving-car/tree/master/datasets/CH2. (2016).
  17. 17.
    1. Who’s responsible when an autonomous car crashes? http://money.cnn.com/2016/07/07/technology/tesla-liability-risk/index.html. (2016).
  18. 18.
    1. Autonomous Vehicles Enacted Legislation. http://www.ncsl.org/research/transportation/autonomous-vehicles-self-driving-vehicles-enacted-legislation.aspx. (2017).
  19. 19.
    1. Baidu Apollo. https://github.com/ApolloAuto/apollo. (2017).
  20. 20.
    1. Inside Waymo’s Secret World for Training Self-Driving Cars. https://www.theatlantic.com/technology/archive/2017/08/inside-waymos-secret-testing-and-simulation-facilities/537648/. (2017).
  21. 21.
    1. The Numbers Don’t Lie: Self-Driving Cars Are Getting Good. https://www.wired.com/2017/02/california-dmv-autonomous-car-disengagement. (2017).
  22. 22.
    1. Software 2.0. https://medium.com/@karpathy/software-2-0-a64152b37c35. (2017).
  23. 23.
    1. Tesla’s Self-Driving System Cleared in Deadly Crash. https://www.nytimes.com/2017/01/19/business/tesla-model-s-autopilot-fatal-crash.html. (2017).
  24. 24.Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al. 2016. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467 (2016).
  25. 25.Raja Ben Abdessalem, Shiva Nejati, Lionel C Briand, and Thomas Stifter. 2016. Testing advanced driver assistance systems using multi-objective search and neural networks. In Automated Software Engineering (ASE), 2016 31st IEEE/ACM International Conference on. IEEE, 63–74.
  26. 26.an Goodfellow and Nicolas Papernot. 2017. The challenge of verification and testing of machine learning. http://www.cleverhans.io/security/privacy/ml/2017/06/14/verification.html. (2017).
  27. 27.Saswat Anand, Edmund K Burke, Tsong Yueh Chen, John Clark, Myra B Cohen, Wolfgang Grieskamp, Mark Harman, Mary Jean Harrold, Phil Mcminn, Antonia Bertolino, et al. 2013. An orchestrated survey of methodologies for automated software test case generation. Journal of Systems and Software 86, 8 (2013), 1978–2001.
  28. 28.Hyrum Anderson. 2017. Evading Next-Gen AV using A.I. https://www.defcon.org/html/defcon-25/dc-25-index.html. (2017).
  29. 29.Osbert Bastani, Yani Ioannou, Leonidas Lampropoulos, Dimitrios Vytiniotis, Aditya Nori, and Antonio Criminisi. 2016. Measuring neural net robustness with constraints. In Advances in Neural Information Processing Systems. 2613–2621.
  30. 30.Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle. 2007. Greedy layer-wise training of deep networks. In Advances in neural information processing systems. 153–160.
  31. 31.Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al. 2016. End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316 (2016).
  32. 32.Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In Security and Privacy (SP), 2017 IEEE Symposium on. IEEE, 39–57.
  33. 33.Tsong Y Chen, Shing C Cheung, and Shiu Ming Yiu. 1998. Metamorphic testing: a new approach for generating next test cases. Technical Report. Technical Report HKUST-CS98-01, Department of Computer Science, Hong Kong University of Science and Technology, Hong Kong.
  34. 34.François Chollet et al. 2015. Keras. https://github.com/fchollet/keras. (2015).
  35. 35.Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier. 2017. Parseval networks: Improving robustness to adversarial examples. In International Conference on Machine Learning. 854–863.
  36. 36.California DMV. 2016. Autonomous Vehicle Disengagement Reports. https://www.dmv.ca.gov/portal/dmv/detail/vr/autonomous/disengagement_report_2016. (2016).
  37. 37.Ivan Evtimov, Kevin Eykholt, Earlence Fernandes, Tadayoshi Kohno, Bo Li, Atul Prakash, Amir Rahmati, and Dawn Song. 2017. Robust Physical-World Attacks on Machine Learning Models. arXiv preprint arXiv:1707.08945 (2017).
  38. 38.Reuben Feinman, Ryan R Curtin, Saurabh Shintre, and Andrew B Gardner. 2017. Detecting Adversarial Samples from Artifacts. arXiv preprint arXiv:1703.00410 (2017).
  39. 39.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. http://www.deeplearningbook.org Book in preparation for MIT Press.
  40. 40.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR).
  41. 41.Kathrin Grosse, Praveen Manoharan, Nicolas Papernot, Michael Backes, and Patrick McDaniel. 2017. On the (statistical) detection of adversarial examples. arXiv preprint arXiv:1702.06280 (2017).
  42. 42.Kathrin Grosse, Nicolas Papernot, Praveen Manoharan, Michael Backes, and Patrick D. McDaniel. 2017. Adversarial Perturbations Against Deep Neural Networks for Malware Classification. In Proceedings of the 2017 European Symposium on Research in Computer Security.
  43. 43.Shixiang Gu and Luca Rigazio. 2015. Towards deep neural network architectures robust to adversarial examples. In International Conference on Learning Representations (ICLR).
  44. 44.Jan Hauke and Tomasz Kossowski. 2011. Comparison of values of Pearson’s and Spearman’s correlation coefficients on the same sets of data. Quaestiones geographicae 30, 2 (2011), 87.
  45. 45.Samer Hijazi, Rishi Kumar, and Chris Rowen. 2015. Using convolutional neural networks for image recognition. Technical Report. Tech. Rep., 2015.[Online]. Available: http://ip.cadence.com/uploads/901/cnn-wp-pdf.
  46. 46.Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al. 2001. Gradient flow in recurrent nets: the difficulty of learning long-term dependencies. (2001).
  47. 47.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735–1780.
  48. 48.Xiaowei Huang, Marta Kwiatkowska, Sen Wang, and Min Wu. 2017. Safety verification of deep neural networks. In International Conference on Computer Aided Verification. Springer, 3–29.
  49. 49.L. C. Jain and L. R. Medsker. 1999. Recurrent Neural Networks: Design and Applications (1st ed.). CRC Press, Inc., Boca Raton, FL, USA.
  50. 50.Andrej Karpathy. [n. d.]. Convolutional neural networks. http://cs231n.github.io/convolutional-networks/. ([n. d.]).
  51. 51.Guy Katz, Clark Barrett, David L. Dill, Kyle Julian, and Mykel J. Kochenderfer. 2017. Reluplex: An Efficient SMT Solver for Verifying Deep Neural Networks. Springer International Publishing, Cham, 97–117.
  52. 52.Jernej Kos, Ian Fischer, and Dawn Song. 2017. Adversarial examples for generative models. arXiv preprint arXiv:1702.06832 (2017).
  53. 53.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097–1105.
  54. 54.Pavel Laskov et al. 2014. Practical evasion of a learning-based classifier: A case study. In Security and Privacy (SP), 2014 IEEE Symposium on. IEEE, 197–211.
  55. 55.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Xiaodong Song. 2017. Delving into Transferable Adversarial Examples and Black-box Attacks. In International Conference on Learning Representations (ICLR).
  56. 56.Phil McMinn. 2004. Search-based software test data generation: a survey. Software testing, Verification and reliability 14, 2 (2004), 105–156.
  57. 57.Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff. 2017. On detecting adversarial perturbations. In International Conference on Learning Representations (ICLR).
  58. 58.Thomas M. Mitchell. 1997. Machine Learning (1 ed.). McGraw-Hill, Inc., New York, NY, USA.
  59. 59.Takeru Miyato, Andrew M Dai, and Ian Goodfellow. 2016. Adversarial Training Methods for Semi-Supervised Text Classification. In Proceedings of the International Conference on Learning Representations (ICLR).
  60. 60.Christian Murphy, Gail E Kaiser, Lifeng Hu, and Leon Wu. 2008. Properties of Machine Learning Applications for Use in Metamorphic Testing.. In SEKE, Vol. 8. 867–872.
  61. 61.Vinod Nair and Geoffrey E Hinton. 2010. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10). 807–814.
  62. 62.Nina Narodytska and Shiva Prasad Kasiviswanathan. 2016. Simple black-box adversarial perturbations for deep networks. In Workshop on Adversarial Training, NIPS 2016.
  63. 63.Anh Nguyen, Jason Yosinski, and Jeff Clune. 2015. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 427–436.
  64. 64.Nicolas Papernot and Patrick McDaniel. 2017. Extending Defensive Distillation. arXiv preprint arXiv:1705.05264 (2017).
  65. 65.Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami. 2017. Practical black-box attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security. ACM, 506–519.
  66. 66.Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. 2016. The limitations of deep learning in adversarial settings. In 2016 IEEE European Symposium on Security and Privacy (EuroS&P). IEEE, 372–387.
  67. 67.Nicolas Papernot, Patrick McDaniel, Ananthram Swami, and Richard Harang. 2016. Crafting adversarial input sequences for recurrent neural networks. In Military Communications Conference, MILCOM 2016-2016 IEEE. IEEE, 49–54.
  68. 68.Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. 2016. Distillation as a defense to adversarial perturbations against deep neural networks. In Security and Privacy (SP), 2016 IEEE Symposium on. IEEE, 582–597.
  69. 69.Corina S Păsăreanu and Willem Visser. 2009. A survey of new trends in symbolic execution for software testing and analysis. International Journal on Software Tools for Technology Transfer (STTT) 11, 4 (2009), 339–353.
  70. 70.Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017. DeepXplore: Automated Whitebox Testing of Deep Learning Systems. arXiv preprint arXiv:1705.06640 (2017).
  71. 71.Luca Pulina and Armando Tacchella. 2010. An abstraction-refinement approach to verification of artificial neural networks. In Computer Aided Verification. Springer, 243–257.
  72. 72.David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1988. Learning representations by back-propagating errors. Cognitive modeling 5, 3 (1988), 1.
  73. 73.D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young. 2014. Machine Learning: The High Interest Credit Card of Technical Debt.
  74. 74.Uri Shaham, Yutaro Yamada, and Sahand Negahban. 2015. Understanding adversarial training: Increasing local stability of neural nets through robust optimization. arXiv preprint arXiv:1511.05432 (2015).
  75. 75.Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K Reiter. 2016. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. ACM, 1528–1540.
  76. 76.Charles Spearman. 1904. The proof and measurement of association between two things. The American journal of psychology 15, 1 (1904), 72–101.
  77. 77.Jacob Steinhardt, Pang Wei Koh, and Percy Liang. 2017. Certified Defenses for Data Poisoning Attacks. arXiv preprint arXiv:1706.03691 (2017).
  78. 78.C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. 2014. Intriguing properties of neural networks. In International Conference on Learning Representations (ICLR).
  79. 79.Theano Development Team. 2016. Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs/1605.02688 (May 2016). http://arxiv.org/abs/1605.02688
  80. 80.Gang Wang, Tianyi Wang, Haitao Zheng, and Ben Y Zhao. 2014. Man vs. Machine: Practical Adversarial Detection of Malicious Crowdsourcing Workers.. In USENIX Security Symposium. 239–254.
  81. 81.Michael J Wilber, Vitaly Shmatikov, and Serge Belongie. 2016. Can we still avoid automatic face detection?. In Applications of Computer Vision (WACV), 2016 IEEE Winter Conference on. IEEE, 1–9.
  82. 82.Ian H Witten, Eibe Frank, Mark A Hall, and Christopher J Pal. 2016. Data Mining: Practical machine learning tools and techniques. Morgan Kaufmann.
  83. 83.Xiaoyuan Xie, Joshua Ho, Christian Murphy, Gail Kaiser, Baowen Xu, and Tsong Yueh Chen. 2009. Application of metamorphic testing to supervised classifiers. In Quality Software, 2009. QSIC’09. 9th International Conference on. IEEE, 135–144.
  84. 84.Weilin Xu, David Evans, and Yanjun Qi. 2017. Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks. arXiv preprint arXiv:1704.01155 (2017).
  85. 85.Weilin Xu, Yanjun Qi, and David Evans. 2016. Automatically evading classifiers. In Proceedings of the 2016 Network and Distributed Systems Symposium.
  86. 86.Stephan Zheng, Yang Song, Thomas Leung, and Ian Goodfellow. 2016. Improving the robustness of deep neural networks via stability training. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4480–4488.
  87. 87.Zhi Quan Zhou, DH Huang, TH Tse, Zongyuan Yang, Haitao Huang, and TY Chen. 2004. Metamorphic testing and its applications. In Proceedings of the 8th International Symposium on Future Software Technology (ISFST 2004). 346–351.

Citation

MLA
Tian, Y., et al. “DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars”. arXiv, 2017, http://arxiv.org/abs/1708.08559v2.
APA
Tian, Y., Pei, K., Jana, S., & Ray, B. (2017). DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars. arXiv. http://arxiv.org/abs/1708.08559v2
Chicago
Tian, Y., K. Pei, S. Jana, and B. Ray. 2017. “DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars”. arXiv. http://arxiv.org/abs/1708.08559v2.
Harvard
Tian, Y. et al. (2017) “DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1708.08559v2.
Vancouver
1. Tian Y, Pei K, Jana S, Ray B (2017) DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars. arXiv

BibTeX

@article{tian2017deeptest,
  title = {DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars},
  author = {Tian, Yuchi and Pei, Kexin and Jana, Suman and Ray, Baishakhi},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1708.08559v2},
  eprint = {1708.08559}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF