End-to-End Driving Via Conditional Imitation Learning

Felipe CodevillaMatthias MüllerAntonio LópezVladlen KoltunAlexey Dosovitskiy

article2017IEEE International Conference on Robotics and Automation1,353 citations

Proposes conditional imitation learning to make end-to-end autonomous driving controllable at test time, enabling vision-based policies to execute high-level directional commands across simulated urban environments and physical robotic platforms.

Listen

Autonomous driving systems trained through imitation learning learn to drive by mimicking human demonstrations. However, standard end-to-end models struggle in complex urban environments because camera images alone cannot tell a vehicle which way to turn at an intersection. Without knowing the driver's goal, these models either hesitate, make arbitrary choices, or cannot be directed toward a specific destination. The article introduces and evaluates conditional imitation learning, a method that incorporates high-level navigational commands—such as turning left, turning right, or continuing straight—into the training and operational phases of an autonomous driving network.

To evaluate this approach, the authors tested two deep neural network designs: one that treats commands as an auxiliary input to a single controller, and a branched architecture where high-level commands route visual data through specialized sub-networks. The systems were evaluated across 50 point-to-point routes in CARLA, a realistic urban driving simulator, and validated in the physical world on a one-fifth scale robotic truck navigating a residential neighborhood with 14 intersections. The training data comprised two hours of human driving augmented with visual perturbations and simulated drift to teach the models how to recover from disturbances.

The findings show that conditional imitation learning dramatically improves navigation compared to traditional methods. In the simulation, the branched network completed 88% of routes in the training town and 64% in a previously unseen town, whereas standard imitation learning succeeded only 20% to 26% of the time, and a goal-vector baseline succeeded only 24% to 30% of the time. The branched design consistently outperformed the single-input architecture, which achieved success rates of 78% and 52% in the two simulation environments. In physical track testing, the branched model completed all turns with only 0.67 manual interventions per run, whereas a model without disturbance-recovery training missed roughly 24% of turns and required 8.67 interventions.

These results demonstrate that separating high-level route planning from low-level vehicle control enables safer and more controllable end-to-end autonomous driving. Equipping the network with specialized command branches resolves navigational ambiguity and prevents catastrophic failures like off-road shortcutting. Furthermore, incorporating recovery data and image augmentation is vital for policy stability and real-world transfer, drastically reducing costly and dangerous driving interventions.

Organizations developing vision-based autonomous systems should adopt branched command-conditional architectures rather than simple end-to-end or raw goal-direction mappings. Training pipelines must include noise injection or recovery demonstrations alongside aggressive data augmentation to ensure stability against real-world lighting and trajectory drifts. Future engineering efforts should focus on expanding network capacity, collecting larger multi-condition datasets, and integrating unstructured natural language instructions to expand passenger communication beyond a discrete vocabulary.

While the approach demonstrates strong generalization to unseen simulated towns and new physical environments, performance remains limited by the small size of the training datasets (two hours of driving) and simple discrete command vocabularies. The physical experiments were also conducted on a sub-scale robotic platform under overcast conditions, meaning caution is warranted before directly deploying this specific implementation to full-scale vehicles without extensive real-world validation.

arXiv: 1710.02410
Cover for End-to-End Driving Via Conditional Imitation Learning

Abstract

Deep networks trained on demonstrations of human driving have learned to follow roads and avoid obstacles. However, driving policies trained via imitation learning cannot be controlled at test time. A vehicle trained end-to-end to imitate an expert cannot be guided to take a specific turn at an upcoming intersection. This limits the utility of such systems. We propose to condition imitation learning on high-level command input. At test time, the learned driving policy functions as a chauffeur that handles sensorimotor coordination but continues to respond to navigational commands. We evaluate different architectures for conditional imitation learning in vision-based driving. We conduct experiments in realistic three-dimensional simulations of urban driving and on a 1/5 scale robotic truck that is trained to drive in a residential area. Both systems drive based on visual input yet remain responsive to high-level navigational commands. The supplementary video can be viewed at this https URL

Table of Contents

  • I Introduction
  • II Related Work
  • III Conditional Imitation Learning
  • IV Methodology
  • IV-A Network Architecture
  • IV-B Network Details
  • IV-C Training Data Distribution
  • IV-D Data Augmentation
  • V System Setup
  • V-A Simulated Environment
  • V-B Physical System
  • V-B1 Data collection
  • V-B2 Model evaluation
  • VI Experiments
  • VI-A Simulated Environment
  • VI-A1 Experimental setup
  • VI-A2 Results
  • VI-B Physical System
  • VI-B1 Experimental setup
  • VI-B2 Main results
  • VI-B3 Generalization to new environments
  • VII Discussion
  • VIII Acknowledgements
  • References

Knowls

  1. Knowl 1 — Command-Conditional Imitation Learning Formulation

    model/method

    Standard imitation learning trains a parameterized policy F(o;θ)F(o; \theta) to mimic an expert E(o)E(o) mapping perceptual observations oo to control actions aa over a dataset of demonstrations D={(oi,ai)}i=1N\mathcal{D} = \{(o_i, a_i)\}_{i=1}^N via empirical risk minimization:

    min⁡θ∑i=1Nℓ(F(oi;θ),ai)\min_\theta \sum_{i=1}^N \ell(F(o_i; \theta), a_i)

    In complex tasks such as urban driving at intersections, the mapping from sensory observations oo alone to optimal control actions aa is ill-posed and non-injective, because the expert's action depends on an unobserved latent internal state hh (e.g., destination intent): a=E(o,h)a = E(o, h).

    Command-conditional imitation learning exposes this latent intent by conditioning the policy on a high-level navigational command c=c(h)c = c(h) provided during demonstration and inference (e.g., turn indicators or topological planner instructions). The training dataset comprises tuples D={(oi,ci,ai)}i=1N\mathcal{D} = \{(o_i, c_i, a_i)\}_{i=1}^N, and the optimization objective is defined as:

    min⁡θ∑i=1Nℓ(F(oi,ci;θ),ai)\min_\theta \sum_{i=1}^N \ell(F(o_i, c_i; \theta), a_i)

    where oio_i represents the sensory observation at step ii, cic_i is the high-level directional command, aia_i is the demonstrated low-level continuous control action, θ\theta denotes policy network parameters, and ℓ(⋅,⋅)\ell(\cdot, \cdot) is a per-sample loss function.

  2. Knowl 2 — Branched Policy Architecture for Command Conditioning

    model/method

    The branched architecture integrates discrete high-level commands c∈C={c0,c1,…,cK}c \in \mathcal{C} = \{c^0, c^1, \dots, c^K\} by acting as an architectural switch over specialized control sub-modules. The network splits into a shared perception backbone and specialized output branches.

    The perception backbone extracts a joint representation from sensory image ii and auxiliary vehicle measurements mm (such as vehicle speed):

    j=J(i,m)=⟨I(i),M(m)⟩j = J(i, m) = \langle I(i), M(m) \rangle

    where I(i)I(i) is a convolutional neural network image module and M(m)M(m) is a fully-connected measurement module.

    Rather than passing cc as a feature vector to a single controller, the policy instantiates K+1K+1 dedicated multilayer perceptron control branches {A0,A1,…,AK}\{A^0, A^1, \dots, A^K\}, each dedicated to a distinct command ckc^k (e.g., c0c^0 for default lane following/continue, c1c^1 for left turn, c2c^2 for straight, c3c^3 for right turn). The network output is computed by routing the shared perceptual embedding J(i,m)J(i, m) through the specific branch indexed by the active command ckc^k:

    F(i,m,ck)=Ak(J(i,m))F(i, m, c^k) = A^k(J(i, m))

    This structural separation forces individual branches to learn specialized sensorimotor behaviors without letting dominant driving behaviors suppress turn-conditional control signals.

  3. Knowl 3 — Perturbation Recovery via Temporally Correlated Steering Noise Injection

    model/method

    Policies trained strictly on clean expert demonstrations often suffer from compounding drift when encountering unfamiliar off-trajectory states at execution time. To supply demonstrations of recovery without human-in-the-loop active data relabeling schemes, synthetic disturbances are injected into the physical steering control loop during data collection.

    A temporally correlated triangular noise signal is added to the expert driver's steering input before sending the combined control signal to the vehicle actuators. This perturbation causes the vehicle to gradually drift away from the center of the lane. As drift occurs, the human driver naturally applies corrective counter-steering controls to bring the vehicle back onto the correct heading.

    The recorded training action is the human driver's corrective control signal, not the disturbed composite signal sent to the steering actuator. Consequently, the dataset contains rich demonstrations of recovery from off-center and drifted states without containing demonstrations of erratic driving.

  4. Knowl 4 — End-to-End Sensorimotor Driving Network Architecture and Objective

    model/method

    The vision-based conditional driving policy maps an RGB input image i∈R200×88×3i \in \mathbb{R}^{200 \times 88 \times 3} and a scalar speed measurement mm to a continuous two-dimensional control action vector a=⟨s,α⟩a = \langle s, \alpha \rangle, where s∈[−1,1]s \in [-1, 1] represents steering angle and α∈[−1,1]\alpha \in [-1, 1] represents acceleration/braking.

    The image module I(i)I(i) comprises 8 convolutional layers followed by 2 fully-connected layers of 512 units each. The first convolutional layer uses a 5×55 \times 5 kernel, and all subsequent convolutional layers use 3×33 \times 3 kernels. Strides are set to 2 in convolutional layers 1, 3, and 5, with channel depth increasing progressively from 32 in the first layer to 256 in the final convolutional layer. Batch normalization is applied after all convolutional layers. ReLU nonlinearities are used after all hidden layers. Dropout is applied at a rate of 20% after convolutional layers and 50% after fully-connected hidden layers.

    The model is trained using the Adam optimizer with an initial learning rate of 0.00020.0002 and minibatches of 120 samples. Minibatches are sampled with a uniform distribution across discrete command classes. The per-sample loss function between predicted action a=⟨s,α⟩a = \langle s, \alpha \rangle and ground truth action agt=⟨sgt,αgt⟩a_{\mathrm{gt}} = \langle s_{\mathrm{gt}}, \alpha_{\mathrm{gt}} \rangle is:

    ℓ(a,agt)=∥s−sgt∥2+λa∥α−αgt∥2\ell(a, a_{\mathrm{gt}}) = \|s - s_{\mathrm{gt}}\|^2 + \lambda_a \|\alpha - \alpha_{\mathrm{gt}}\|^2

    where λa\lambda_a balances the steering and acceleration loss components.

  5. Knowl 5 — Urban Driving Performance Across Architectures in CARLA Simulation

    data/table

    Quantitative performance of conditional and baseline imitation learning models was evaluated across 50 start-goal navigation episodes (routes of at least 1 km) in the CARLA simulator. Town 1 was used during training, and Town 2 was reserved strictly for evaluating generalization under novel visual layout and geometry.

    Success rate (%) Km per infraction
    Model Town 1 Town 2 Town 1 Town 2
    Non-conditional 20% 26% 5.76 0.89
    Goal-conditional 24% 30% 1.87 1.22
    Ours branched 88% 64% 2.34 1.18
    Ours cmd. input 78% 52% 3.97 1.30
    Ours no noise 56% 22% 1.31 0.54
    Ours no aug. 80% 0% 4.03 0.36
    Ours shallow net 46% 14% 0.96 0.42

    The branched architecture outperforms all baselines, achieving an 88% success rate in Town 1 and 64% in the unseen Town 2. Goal-vector conditioning fails to provide substantial benefit over non-conditional imitation learning (24%/30% vs 20%/26%). Ablation shows that removing perturbation noise drops Town 2 success from 64% to 22%, omitting visual data augmentation drops Town 2 success to 0%, and utilizing a shallower convolutional network drops Town 2 success to 14%.

  6. Knowl 6 — Real-World Robotic Vehicle Navigation Performance Across Architectures and Ablations

    data/table

    Physical deployment was evaluated on a 1/5-scale Traxxas Maxx truck equipped with an Nvidia TX2 embedded computer, Holybro Pixhawk controller, and three webcams. The vehicle navigated an outdoor residential route containing 14 intersections under overcast weather (trained on data collected primarily in sunny weather). Metrics report the percentage of missed intersection turns, average human interventions per run, and route completion time, averaged across 3 runs.

    Model Missed turns Interventions Time
    Ours branched 0% 0.67 2:19
    Ours cmd. input 11.1% 2.33 4:13
    Ours no noise 24.4% 8.67 4:39
    Ours no aug. 73% 39.00 10:41

    The branched conditional model navigated the physical course with 0% missed turns and an average of 0.67 interventions. The command input architecture missed 11.1% of turns and required 2.33 interventions. Omitting noise-injected drift recovery degraded policy stability, increasing interventions to 8.67 and missed turns to 24.4%. Omitting visual data augmentation caused severe policy failure due to domain shift between sunny training and overcast testing conditions, resulting in 39 interventions and 73% missed turns.

  7. Knowl 7 — Command-Input Architecture for Goal and Command Conditioning

    model/method

    In the command-input architecture, commands and perceptual features are concatenated into a single shared intermediate representation before being passed to a common control module. The architecture processes the input image ii via a convolutional module I(i)I(i), vehicle speed measurements mm via a fully-connected module M(m)M(m), and a categorical or continuous command cc via a fully-connected module C(c)C(c).

    The module outputs are concatenated into a joint latent vector:

    j=J(i,m,c)=⟨I(i),M(m),C(c)⟩j = J(i, m, c) = \langle I(i), M(m), C(c) \rangle

    A single fully-connected control module A(j)A(j) maps this joint embedding to predicted driving control actions a=A(j)a = A(j). While this architecture accommodates both discrete categorical vectors and continuous vectors (such as spatial goal directions), the shared controller is not structurally constrained to attend to the command channel C(c)C(c), leading to lower compliance with directional commands compared to branched architectures.

  8. Knowl 8 — Failure Mode of Feed-Forward Goal-Vector Conditioning

    empirical result

    Conditioning an end-to-end feed-forward reactive driving policy on a continuous vector pointing directly to a destination coordinate (in the vehicle's coordinate frame) does not lead to effective urban navigation. In multi-intersection simulation benchmarks in CARLA, goal-conditioned policies achieve only 24% completion in Town 1 and 30% in Town 2, performing comparably to a purely non-conditional baseline (20% and 26%).

    The failure arises because a feed-forward reactive policy cannot translate a straight-line vector pointing toward a distant goal into the sequence of discrete, road-compliant intersection turns required to reach it. Instead, the vehicle attempts to drive directly along the goal vector, veering off the drivable road surface, crossing sidewalks, and entering opposing traffic lanes to shortcut toward the coordinate.

  9. Knowl 9 — Online Photometric and Noise Data Augmentation Protocol

    experimental setup

    To enable zero-shot visual generalization across weather conditions, illumination changes, and unseen environments, input camera images are processed through an online data augmentation pipeline during training. Each presented frame is transformed by a randomly selected subset of operations with sampled magnitudes:

    1. Photometric adjustments: perturbations to image contrast, brightness, and color tone.
    2. Noise and blur corruption: addition of Gaussian noise, Gaussian blur, and salt-and-pepper noise.
    3. Region dropout: random spatial masking of multiple rectangular patches, where each masked rectangle covers approximately 1% of the total image area.

    Geometric transformations (such as horizontal image flipping, random rotations, or 2D translations) are strictly excluded from the augmentation pipeline because vehicular control targets (specifically steering angles) are non-invariant to spatial geometric transformations.

  10. Knowl 10 — Categorical Vocabulary and Policy Scalability Boundaries

    limitation

    The conditional imitation learning approach has several explicit design boundaries:

    1. Fixed command vocabulary: High-level guidance is restricted to a small predefined discrete set of directional maneuvers (follow road, turn left, turn right, go straight) rather than natural language instructions or dialogues.
    2. Model capacity and scalability: Scaling end-to-end conditional imitation learning to complex, large-scale autonomous urban driving requires substantially higher-capacity neural architectures and larger demonstration datasets than the 8-layer convolutional networks tested.
    3. Semantic traffic rules: The policy learns reactive sensorimotor control and obstacle avoidance from demonstrations but does not incorporate explicit modular reasoning for complex semantic driving rules such as traffic light states and stop sign priority.

Coverage note — None was omitted; all key contributions including the conditional imitation learning formulation, branched and command-input architectures, perturbation injection data collection method, network specifications, simulation and physical experimental results, and stated limitations are covered.

References

  1. 1.P. Abbeel, A. Coates, and A. Y. Ng. Autonomous helicopter aerobatics through apprenticeship learning. International Journal of Robotics Research, 29(13), 2010.
  2. 2.B. Argall, S. Chernova, M. M. Veloso, and B. Browning. A survey of robot learning from demonstration. Robotics and Autonomous Systems, 57(5), 2009.
  3. 3.A. G. Barto and S. Mahadevan. Recent advances in hierarchical reinforcement learning. Discrete Event Dynamic Systems, 13(1-2), 2003.
  4. 4.M. Bojarski, D. D. Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, and K. Zieba. End to end learning for self-driving cars. arXiv:1604.07316, 2016.
  5. 5.A. Broad, J. Arkin, N. Ratliff, T. Howard, and B. Argall. Real-time natural language corrections for assistive robotic manipulators. International Journal of Robotics Research, 2017.
  6. 6.C. Chen, A. Seff, A. L. Kornhauser, and J. Xiao. DeepDriving: Learning affordance for direct perception in autonomous driving. In ICCV, 2015.
  7. 7.B. C. da Silva, G. Konidaris, and A. G. Barto. Learning parameterized skills. In ICML, 2012.
  8. 8.M. P. Deisenroth, P. Englert, J. Peters, and D. Fox. Multi-task policy search for robotics. In ICRA, 2014.
  9. 9.A. Dosovitskiy and V. Koltun. Learning to act by predicting the future. In ICLR, 2017.
  10. 10.A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun. ´ CARLA: An open urban driving simulator. In Conference on Robot Learning (CoRL), 2017.
  11. 11.P. Englert, A. Paraschos, J. Peters, and M. P. Deisenroth. Model-based imitation learning by probabilistic trajectory matching. In ICRA, 2013.
  12. 12.U. Franke. Autonomous driving. In Computer Vision in Vehicle Technology. 2017.
  13. 13.A. Giusti, J. Guzzi, D. Ciresan, F.-L. He, J. P. Rodriguez, F. Fontana, M. Faessler, C. Forster, J. Schmidhuber, G. Di Caro, D. Scaramuzza, and L. Gambardella. A machine learning approach to visual perception of forest trails for mobile robots. IEEE Robotics and Automation Letters, 2016.
  14. 14.S. Hemachandra, F. Duvallet, T. M. Howard, N. Roy, A. Stentz, and M. R. Walter. Learning models for following natural language directions in unknown environments. In ICRA, 2015.
  15. 15.S. Javdani, S. S. Srinivasa, and J. A. Bagnell. Shared autonomy via hindsight optimization. In RSS, 2015.
  16. 16.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  17. 17.J. Kober, J. A. Bagnell, and J. Peters. Reinforcement learning in robotics: A survey. International Journal of Robotics Research, 32(11), 2013.
  18. 18.J. Kober, A. Wilhelm, E. Oztop, and J. Peters. Reinforcement learning to adjust parametrized motor primitives to new situations. Autonomous Robots, 33(4), 2012.
  19. 19.G. Konidaris, S. Kuindersma, R. A. Grupen, and A. G. Barto. Robot learning from demonstration by constructing skill trees. International Journal of Robotics Research, 31(3), 2012.
  20. 20.T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. B. Tenenbaum. Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation. In NIPS, 2016.
  21. 21.M. Laskey, A. Dragan, J. Lee, K. Goldberg, and R. Fox. Dart: Optimizing noise injection in imitation learning. In Conference on Robot Learning (CoRL), 2017.
  22. 22.Y. LeCun, U. Muller, J. Ben, E. Cosatto, and B. Flepp. Off-road obstacle avoidance through end-to-end learning. In NIPS, 2005.
  23. 23.S. Levine and V. Koltun. Guided policy search. In ICML, 2013.
  24. 24.C. Matuszek, L. Bo, L. Zettlemoyer, and D. Fox. Learning from unscripted deictic gesture and language for human-robot interactions. In AAAI, 2014.
  25. 25.B. Paden, M. Cap, S. Z. Yong, D. S. Yershov, and E. Frazzoli. A ´ survey of motion planning and control techniques for self-driving urban vehicles. IEEE Transactions on Intelligent Vehicles, 1(1), 2016.
  26. 26.P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal. Learning and generalization of motor skills by learning from demonstration. In ICRA, 2009.
  27. 27.D. Pomerleau. ALVINN: An autonomous land vehicle in a neural network. In NIPS, 1988.
  28. 28.N. D. Ratliff, J. A. Bagnell, and S. S. Srinivasa. Imitation learning for locomotion and manipulation. In International Conference on Humanoid Robots, 2007.
  29. 29.S. Ross, G. J. Gordon, and J. A. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS, 2011.
  30. 30.S. Ross, N. Melik-Barkhudarov, K. S. Shankar, A. Wendel, D. Dey, J. A. Bagnell, and M. Hebert. Learning monocular reactive UAV control in cluttered natural environments. In ICRA, 2013.
  31. 31.T. Schaul, D. Horgan, K. Gregor, and D. Silver. Universal value function approximators. In ICML, 2015.
  32. 32.D. Silver, J. A. Bagnell, and A. Stentz. Learning from demonstration for autonomous navigation in complex unstructured terrain. International Journal of Robotics Research, 29(12), 2010.
  33. 33.R. S. Sutton, D. Precup, and S. P. Singh. Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112(1-2), 1999.
  34. 34.S. Tellex, T. Kollar, S. Dickerson, M. R. Walter, A. G. Banerjee, S. J. Teller, and N. Roy. Understanding natural language commands for robotic navigation and mobile manipulation. In AAAI, 2011.
  35. 35.M. R. Walter, S. Hemachandra, B. Homberg, S. Tellex, and S. J. Teller. Learning semantic maps from natural language descriptions. In RSS, 2013.
  36. 36.J. Zhang and K. Cho. Query-efficient imitation learning for end-to-end simulated driving. In AAAI, 2017.
  37. 37.B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey. Maximum entropy inverse reinforcement learning. In AAAI, 2008.
  38. 38.B. D. Ziebart, A. L. Maas, A. K. Dey, and J. A. Bagnell. Navigate like a cabbie: Probabilistic reasoning from observed context-aware behavior. In UbiComp, 2008.

Citation

MLA
Codevilla, F., et al. “End-to-end Driving via Conditional Imitation Learning”. arXiv, 2017, http://arxiv.org/abs/1710.02410v2.
APA
Codevilla, F., Müller, M., López, A., Koltun, V., & Dosovitskiy, A. (2017). End-to-end Driving via Conditional Imitation Learning. arXiv. http://arxiv.org/abs/1710.02410v2
Chicago
Codevilla, F., M. Müller, A. López, V. Koltun, and A. Dosovitskiy. 2017. “End-to-end Driving via Conditional Imitation Learning”. arXiv. http://arxiv.org/abs/1710.02410v2.
Harvard
Codevilla, F. et al. (2017) “End-to-end Driving via Conditional Imitation Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1710.02410v2.
Vancouver
1. Codevilla F, Müller M, López A, Koltun V, Dosovitskiy A (2017) End-to-end Driving via Conditional Imitation Learning. arXiv

BibTeX

@article{codevilla2017end,
  title = {End-to-end Driving via Conditional Imitation Learning},
  author = {Codevilla, Felipe and Müller, Matthias and López, Antonio and Koltun, Vladlen and Dosovitskiy, Alexey},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1710.02410v2},
  eprint = {1710.02410}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF