End-to-End Driving Via Conditional Imitation Learning
Felipe CodevillaMatthias MüllerAntonio LópezVladlen KoltunAlexey Dosovitskiy
Proposes conditional imitation learning to make end-to-end autonomous driving controllable at test time, enabling vision-based policies to execute high-level directional commands across simulated urban environments and physical robotic platforms.
Autonomous driving systems trained through imitation learning learn to drive by mimicking human demonstrations. However, standard end-to-end models struggle in complex urban environments because camera images alone cannot tell a vehicle which way to turn at an intersection. Without knowing the driver's goal, these models either hesitate, make arbitrary choices, or cannot be directed toward a specific destination. The article introduces and evaluates conditional imitation learning, a method that incorporates high-level navigational commands—such as turning left, turning right, or continuing straight—into the training and operational phases of an autonomous driving network.
To evaluate this approach, the authors tested two deep neural network designs: one that treats commands as an auxiliary input to a single controller, and a branched architecture where high-level commands route visual data through specialized sub-networks. The systems were evaluated across 50 point-to-point routes in CARLA, a realistic urban driving simulator, and validated in the physical world on a one-fifth scale robotic truck navigating a residential neighborhood with 14 intersections. The training data comprised two hours of human driving augmented with visual perturbations and simulated drift to teach the models how to recover from disturbances.
The findings show that conditional imitation learning dramatically improves navigation compared to traditional methods. In the simulation, the branched network completed 88% of routes in the training town and 64% in a previously unseen town, whereas standard imitation learning succeeded only 20% to 26% of the time, and a goal-vector baseline succeeded only 24% to 30% of the time. The branched design consistently outperformed the single-input architecture, which achieved success rates of 78% and 52% in the two simulation environments. In physical track testing, the branched model completed all turns with only 0.67 manual interventions per run, whereas a model without disturbance-recovery training missed roughly 24% of turns and required 8.67 interventions.
These results demonstrate that separating high-level route planning from low-level vehicle control enables safer and more controllable end-to-end autonomous driving. Equipping the network with specialized command branches resolves navigational ambiguity and prevents catastrophic failures like off-road shortcutting. Furthermore, incorporating recovery data and image augmentation is vital for policy stability and real-world transfer, drastically reducing costly and dangerous driving interventions.
Organizations developing vision-based autonomous systems should adopt branched command-conditional architectures rather than simple end-to-end or raw goal-direction mappings. Training pipelines must include noise injection or recovery demonstrations alongside aggressive data augmentation to ensure stability against real-world lighting and trajectory drifts. Future engineering efforts should focus on expanding network capacity, collecting larger multi-condition datasets, and integrating unstructured natural language instructions to expand passenger communication beyond a discrete vocabulary.
While the approach demonstrates strong generalization to unseen simulated towns and new physical environments, performance remains limited by the small size of the training datasets (two hours of driving) and simple discrete command vocabularies. The physical experiments were also conducted on a sub-scale robotic platform under overcast conditions, meaning caution is warranted before directly deploying this specific implementation to full-scale vehicles without extensive real-world validation.
- Paper: End to End Learning for Self-Driving Cars, Mariusz Bojarski et al. (2016). This seminal work demonstrates end-to-end convolutional networks mapping raw front-camera video directly to vehicle steering commands, establishing the core imitation learning baseline that conditional imitation learning directly aims to make controllable.
- Paper: DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving, Chenyi Chen et al. (2015). This paper establishes direct perception and affordance learning for vision-based vehicle control, providing vital foundation for understanding how vision representations can guide autonomous driving decisions.
- Paper: A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning, Stephane Ross et al. (2010). This paper formalizes the theoretical foundations of imitation learning and covariate shift mitigation via the DAgger algorithm, which informs robust policy training from human demonstrations.
- Paper: Target-driven visual navigation in indoor scenes using deep reinforcement learning, Yuke Zhu et al. (2016). This work introduces target conditioning for end-to-end visual navigation policies, directly inspiring methods that condition continuous visuomotor control on navigational objectives.
- Paper: A Survey of Motion Planning and Control Techniques for Self-Driving Urban Vehicles, Brian Paden et al. (2016). This comprehensive survey provides the necessary background on classical motion planning, feedback control, and behavioral decision-making hierarchies that end-to-end learning paradigms seek to reformulate.
- Paper: A survey of deep learning techniques for autonomous driving, Sorin Grigorescu et al. (2019). This survey extensively contextualizes and compares end-to-end conditional imitation architectures against modern modular deep learning pipelines for autonomous vehicles.
- Paper: A Survey of Autonomous Driving: Common Practices and Emerging Technologies, Ekim Yurtsever et al. (2019). This review evaluates modern autonomous driving software architectures, tracking how end-to-end neural driving policies integrate into broader production vehicular stacks.
- Paper: Deep Reinforcement Learning for Autonomous Driving: A Survey, Bangalore Ravi Kiran et al. (2020). This survey explores advanced reinforcement and hybrid imitation learning paradigms in autonomous driving, extending beyond standard supervised behavioral cloning formulations.
- Paper: Planning-oriented Autonomous Driving, Yi Hu et al. (2022). This work advances end-to-end driving architectures by unifying vision perception, multi-agent forecasting, and conditional trajectory planning within an interpretable transformer framework.
- Paper: DeepTest: Automated Testing of Deep-Neural-Network-Driven Autonomous Cars, Yuchi Tian et al. (2017). This paper introduces automated metamorphic testing techniques specifically tailored to detect safety-critical failure cases in end-to-end vision-based driving networks.
- Paper: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion et al. (2020). This work extends camera-based end-to-end motion planning by learning explicit bird's-eye-view representations from multi-camera setups for trajectory scoring.
