DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving
Chenyi ChenAri SeffAlain KornhauserJianxiong Xiao
Proposes a direct perception framework for autonomous driving that bridges the gap between full scene parsing and end-to-end control by learning compact driving affordances that enable simple controllers to drive across diverse simulated and real-world environments.
Autonomous driving systems typically rely on one of two paradigms: mediated perception, which constructs a full, complex model of the entire visual scene, or behavior reflex, which directly maps raw images to steering and throttle commands. Mediated perception demands heavy computation and expensive sensor suites to solve complex visual recognition tasks, whereas behavior reflex approaches lack scene understanding, produce erratic behavior in multi-lane traffic, and struggle with conflicting human driving demonstrations. The article addresses this trade-off by proposing a third paradigm, termed direct perception, which balances task complexity and actionable scene understanding.
The objective of the article is to demonstrate that a deep convolutional neural network can directly estimate a compact set of task-critical driving affordance indicators from single forward-facing images, providing sufficient information for a simple controller to steer and regulate vehicle speed safely in highway scenarios.
To evaluate this framework, the authors trained a deep convolutional neural network using 484,815 image frames captured during 12 hours of human driving across customized tracks and diverse traffic configurations within the open-source racing simulator TORCS. They defined 13 direct affordance indicators—including vehicle heading angle relative to the road, lateral distances to lane markings, and longitudinal distances to preceding vehicles across three lanes. A simple rule-based controller translated these indicators into steering, lane-change, and speed commands. To assess real-world viability, the authors tested the trained network on smartphone road video footage and trained a dual-network configuration on 61,894 real-world images from the KITTI benchmark dataset to predict close- and far-range vehicle coordinates.
The key findings demonstrate clear advantages over traditional baselines. First, the direct perception network achieved smooth, collision-free autonomous driving across novel simulated tracks and traffic conditions, successfully following lanes and executing safe overtaking maneuvers. Second, the proposed method reduced lane marking distance estimation errors to approximately 0.16–0.32 meters, substantially outperforming both a standard mediated lane detector (errors of 0.90–1.67 meters) and traditional handcrafted GIST feature baselines (errors of 0.54–1.55 meters). Third, on car distance estimation, the deep direct model achieved errors around 4.7–10.7 meters, outperforming GIST baselines that exhibited errors between 12.7 and 31.4 meters. Fourth, on the real-world KITTI benchmark, the direct perception network matched the overall distance accuracy of a state-of-the-art deformable part model baseline (overall Euclidean error of roughly 6.3 meters for both) while outperforming it when excluding false positives (4.67 meters versus 5.33 meters) without relying on assumptions of flat ground geometry.
These results indicate that direct affordance estimation significantly simplifies system architecture by bypassing the overhead of complete 3D scene reconstruction while avoiding the instability of end-to-end reflex learning. This compact representation reduces computational demands, lowers hardware sensor costs, and improves explainability, as the internal neural activations directly reflect meaningful physical features like road edges and vehicle positions.
Organizations developing autonomous driving software should consider implementing direct affordance layers as an intermediate representation to enhance control interpretability and robustness. Before transitioning this technology to production or on-road deployment, further engineering is required to integrate backward-looking sensors, expand training across diverse real-world weather and lighting conditions, and validate closed-loop vehicle control on physical testbeds.
Current limitations include restricted sensing range, as vehicle distance estimation becomes noisy beyond 30 meters at low image resolutions, a lack of rear-view awareness requiring artificial time-delay assumptions during lane changes, and a higher false-positive rate in cluttered real-world environments due to limited real training samples. While confidence is high in the theoretical and experimental validity within multi-lane highway settings, cautious expansion is necessary before applying this architecture to complex urban intersections.
- Paper: Depth Map Prediction from a Single Image using a Multi-Scale Deep Network, David Eigen et al. (2014). Introduces foundational multi-scale deep architectures and scale-invariant loss formulations for monocular depth estimation on datasets like KITTI, which underpin direct perception representations.
- Paper: Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-scale Convolutional Architecture, David Eigen et al. (2014). Demonstrates predicting geometric and semantic scene affordances directly from monocular images using unified convolutional architectures, establishing the core technical precedent for direct perception.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). Establishes fully convolutional neural networks for pixel-level dense prediction, providing the architectural foundation for intermediate visual affordance estimation.
- Paper: Intelligence Without Reason, Rodney A. Brooks (1991). Provides the conceptual foundation for bypassing full-scene world modeling in favor of action-oriented perceptual representations.
- Paper: End to End Learning for Self-Driving Cars, Mariusz Bojarski et al. (2016). Provides the canonical pure end-to-end reflex baseline directly contrasting DeepDriving's affordance-based direct perception paradigm.
- Paper: CARLA: An Open Urban Driving Simulator, Alexey Dosovitskiy et al. (2017). Establishes an open-source urban simulation platform and benchmarks modular, direct perception, and end-to-end driving paradigms.
- Paper: Deep Reinforcement Learning for Autonomous Driving: A Survey, B Ravi Kiran et al. (2020). Surveys the evolution of visual decision-making paradigms in autonomous driving, placing affordance learning and direct perception within modern learning-based control frameworks.
- Paper: A Survey of Autonomous Driving: Common Practices and Emerging Technologies, Ekim Yurtsever et al. (2019). Synthesizes modular versus end-to-end architectural paradigms and real-world perception systems across autonomous driving technologies.
- Paper: Unsupervised Monocular Depth Estimation with Left-Right Consistency, Clément Godard et al. (2016). Advances monocular spatial perception by learning geometric depth indicators without expensive supervised depth labels.
- Paper: Multi-view 3D Object Detection Network for Autonomous Driving, Xiaozhi Chen et al. (2017). Extends intermediate driving perception by predicting accurate 3D oriented bounding boxes directly from sensory feeds.
