DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving

Chenyi ChenAri SeffAlain KornhauserJianxiong Xiao

article2015ICCV1,873 citations

Proposes a direct perception framework for autonomous driving that bridges the gap between full scene parsing and end-to-end control by learning compact driving affordances that enable simple controllers to drive across diverse simulated and real-world environments.

Listen

Autonomous driving systems typically rely on one of two paradigms: mediated perception, which constructs a full, complex model of the entire visual scene, or behavior reflex, which directly maps raw images to steering and throttle commands. Mediated perception demands heavy computation and expensive sensor suites to solve complex visual recognition tasks, whereas behavior reflex approaches lack scene understanding, produce erratic behavior in multi-lane traffic, and struggle with conflicting human driving demonstrations. The article addresses this trade-off by proposing a third paradigm, termed direct perception, which balances task complexity and actionable scene understanding.

The objective of the article is to demonstrate that a deep convolutional neural network can directly estimate a compact set of task-critical driving affordance indicators from single forward-facing images, providing sufficient information for a simple controller to steer and regulate vehicle speed safely in highway scenarios.

To evaluate this framework, the authors trained a deep convolutional neural network using 484,815 image frames captured during 12 hours of human driving across customized tracks and diverse traffic configurations within the open-source racing simulator TORCS. They defined 13 direct affordance indicators—including vehicle heading angle relative to the road, lateral distances to lane markings, and longitudinal distances to preceding vehicles across three lanes. A simple rule-based controller translated these indicators into steering, lane-change, and speed commands. To assess real-world viability, the authors tested the trained network on smartphone road video footage and trained a dual-network configuration on 61,894 real-world images from the KITTI benchmark dataset to predict close- and far-range vehicle coordinates.

The key findings demonstrate clear advantages over traditional baselines. First, the direct perception network achieved smooth, collision-free autonomous driving across novel simulated tracks and traffic conditions, successfully following lanes and executing safe overtaking maneuvers. Second, the proposed method reduced lane marking distance estimation errors to approximately 0.16–0.32 meters, substantially outperforming both a standard mediated lane detector (errors of 0.90–1.67 meters) and traditional handcrafted GIST feature baselines (errors of 0.54–1.55 meters). Third, on car distance estimation, the deep direct model achieved errors around 4.7–10.7 meters, outperforming GIST baselines that exhibited errors between 12.7 and 31.4 meters. Fourth, on the real-world KITTI benchmark, the direct perception network matched the overall distance accuracy of a state-of-the-art deformable part model baseline (overall Euclidean error of roughly 6.3 meters for both) while outperforming it when excluding false positives (4.67 meters versus 5.33 meters) without relying on assumptions of flat ground geometry.

These results indicate that direct affordance estimation significantly simplifies system architecture by bypassing the overhead of complete 3D scene reconstruction while avoiding the instability of end-to-end reflex learning. This compact representation reduces computational demands, lowers hardware sensor costs, and improves explainability, as the internal neural activations directly reflect meaningful physical features like road edges and vehicle positions.

Organizations developing autonomous driving software should consider implementing direct affordance layers as an intermediate representation to enhance control interpretability and robustness. Before transitioning this technology to production or on-road deployment, further engineering is required to integrate backward-looking sensors, expand training across diverse real-world weather and lighting conditions, and validate closed-loop vehicle control on physical testbeds.

Current limitations include restricted sensing range, as vehicle distance estimation becomes noisy beyond 30 meters at low image resolutions, a lack of rear-view awareness requiring artificial time-delay assumptions during lane changes, and a higher false-positive rate in cluttered real-world environments due to limited real training samples. While confidence is high in the theoretical and experimental validity within multi-lane highway settings, cautious expansion is necessary before applying this architecture to complex urban intersections.

Cover for DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving

Abstract

Today, there are two major paradigms for vision-based autonomous driving systems: mediated perception approaches that parse an entire scene to make a driving decision, and behavior reflex approaches that directly map an input image to a driving action by a regressor. In this paper, we propose a third paradigm: a direct perception approach to estimate the affordance for driving. We propose to map an input image to a small number of key perception indicators that directly relate to the affordance of a road/traffic state for driving. Our representation provides a set of compact yet complete descriptions of the scene to enable a simple controller to drive autonomously. Falling in between the two extremes of mediated perception and behavior reflex, we argue that our direct perception representation provides the right level of abstraction. To demonstrate this, we train a deep Convolutional Neural Network using recording from 12 hours of human driving in a video game and show that our model can work well to drive a car in a very diverse set of virtual environments. We also train a model for car distance estimation on the KITTI dataset. Results show that our direct perception approach can generalize well to real driving images. Source code and data are available on our project website.

Table of Contents

  • 1 Introduction
  • 1.1 Related work
  • 2 Learning affordance for driving perception
  • 2.1 Mapping from an image to affordance
  • 2.2 Mapping from affordance to action
  • 3 Implementation
  • 4 TORCS evaluation
  • 4.1 Qualitative assessment
  • 4.2 Comparison with baselines
  • 5 Testing on real-world data
  • 5.1 Smartphone video
  • 5.2 Car distance estimation on the KITTI dataset
  • 5.3 Comparison with DPM-based baseline
  • 6 Visualization
  • 7 Conclusions
  • References

Knowls

  1. Knowl 1 — Direct Perception Affordance Representation for Autonomous Driving

    definition

    Direct perception for autonomous driving formulates driving scene understanding as predicting a compact set of 13 task-specific affordance indicators directly from a single front-facing monocular image, rather than constructing a full 3D semantic geometric world model (mediated perception) or directly regressing steering wheel angles from pixels (behavior reflex). The state space models highway driving with up to three lanes (the host car's lane and two adjacent lanes) using two coordinated reference frames:

    1. Global Indicator (Always Active):

      • angle: Heading angle between the host car's longitudinal direction and the road centerline tangent, with angle∈[−π,π]\text{angle} \in [-\pi, \pi].
    2. In-Lane Coordinate System (Active when driving within a lane):

      • toMarking_LL: Lateral distance from the car to the left boundary of the left adjacent lane.
      • toMarking_ML: Lateral distance from the car to the left boundary of the current lane.
      • toMarking_MR: Lateral distance from the car to the right boundary of the current lane.
      • toMarking_RR: Lateral distance from the car to the right boundary of the right adjacent lane.
      • dist_LL: Longitudinal distance to the closest preceding vehicle in the left lane.
      • dist_MM: Longitudinal distance to the closest preceding vehicle in the host car's current lane.
      • dist_RR: Longitudinal distance to the closest preceding vehicle in the right lane.
    3. On-Marking Coordinate System (Active when the car is straddling a lane boundary during lane changes):

      • toMarking_L: Lateral distance to the left adjacent lane marking.
      • toMarking_M: Lateral distance to the central lane marking currently being crossed.
      • toMarking_R: Lateral distance to the right adjacent lane marking.
      • dist_L: Longitudinal distance to the preceding vehicle in the lane to the left of the crossed marking.
      • dist_R: Longitudinal distance to the preceding vehicle in the lane to the right of the crossed marking.

    To ensure smooth transitions during lane changes, an overlapping activation region exists between the in-lane and on-marking systems where both sets of indicators are simultaneously evaluated. If an adjacent lane does not exist (e.g., when traveling on an outermost boundary lane), the corresponding indicators output an inactive state.

  2. Knowl 2 — Autonomous Driving Controller for Direct Perception Affordances

    algorithm

    A rule-based driving controller maps estimated affordance indicators to steering, acceleration, and braking commands. Steering minimizes the displacement between the vehicle's position and the target lane centerline. Longitudinal velocity is governed by a baseline cruise speed, curvature-based speed reductions, and optimal velocity car-following constraints.

    Input: Affordance indicators from perception network, current host vehicle speed vcurrentv_{\text{current}}
    Output: Steering command steerCmd\text{steerCmd}, acceleration / brake commands
    while in autonomous driving mode do
        Read affordance indicators from perception network
        Check availability of left and right lanes from indicator active states
        
        // Discrete decision making
        if approaching preceding car in current lane then
            if left lane exists and left lane is clear and lane change is allowable then
                decision ←\leftarrow "left lane change"
            else if right lane exists and right lane is clear and lane change is allowable then
                decision ←\leftarrow "right lane change"
            else
                decision ←\leftarrow "slow down"
            end if
        else
            decision ←\leftarrow "normal driving"
        end if
        
        // Steering control
        if decision = "normal driving" then
            target_center_line←\text{target\_center\_line} \leftarrow centerline of current lane
        else if decision = "left lane change" or decision = "right lane change" then
            target_center_line←\text{target\_center\_line} \leftarrow centerline of objective lane
        end if
        
        dist_center←\text{dist\_center} \leftarrow distance from host car to target_center_line\text{target\_center\_line}
        steerCmd←C⋅(angle−dist_center/road_width)\text{steerCmd} \leftarrow C \cdot (\text{angle} - \text{dist\_center} / \text{road\_width})
        
        // Longitudinal speed control
        vdesired←72 km/hv_{\text{desired}} \leftarrow 72\text{ km/h}
        if host car is turning then
            vdesired←vdesired−speed_drop(past steering commands)v_{\text{desired}} \leftarrow v_{\text{desired}} - \text{speed\_drop}(\text{past steering commands})
        end if
        if decision = "slow down" then
            vfollow←vmax⁡⋅(1−exp⁡(−cvmax⁡⋅dist_MM−d))v_{\text{follow}} \leftarrow v_{\max} \cdot \left(1 - \exp\left(-\frac{c}{v_{\max}} \cdot \text{dist\_MM} - d\right)\right)
            vdesired←min⁡(vdesired,vfollow)v_{\text{desired}} \leftarrow \min(v_{\text{desired}}, v_{\text{follow}})
        end if
        
        Compute acceleration and brake commands to drive vcurrentv_{\text{current}} toward vdesiredv_{\text{desired}}
    end while

    In the steering control law, CC is a gain coefficient calibrated for vehicle dynamics, angle\text{angle} is the vehicle heading relative to the road tangent, and road_width\text{road\_width} is the lane width. In the car-following model, vmax⁡v_{\max} is the maximum permissible speed, dist_MM\text{dist\_MM} is the distance to the preceding vehicle in the host car's lane, and c,dc, d are calibration constants.

  3. Knowl 3 — Direct Affordance Estimation ConvNet Architecture and Training

    model/method

    The direct perception model uses a deep Convolutional Neural Network based on the AlexNet architecture trained to perform multi-task regression from monocular road images directly to driving affordance indicators.

    • Architecture: The network consists of 5 convolutional layers followed by 4 fully connected layers with output dimensions of 4096, 4096, 256, and 13, respectively. The final layer outputs the 13 affordance indicators.
    • Input Resolution: Non-square monocular images of size 280×210×3280 \times 210 \times 3 RGB without image cropping or horizontal flipping.
    • Loss Function: Euclidean loss (mean squared error) over all active output indicators: L=12N∑i=1N∑k∈Ai(y^i,k−yi,k)2L = \frac{1}{2N} \sum_{i=1}^N \sum_{k \in \mathcal{A}_i} (\hat{y}_{i,k} - y_{i,k})^2 where Ai\mathcal{A}_i denotes the set of active affordance indicator indices for training sample ii, y^i,k\hat{y}_{i,k} is the predicted indicator value, and yi,ky_{i,k} is the normalized ground truth.
    • Target Normalization: Ground-truth affordance values spanning different physical units (angles in radians, distances in meters) are linearly scaled into the normalized range [0.1,0.9][0.1, 0.9].
    • Training Configuration: Trained from scratch using mini-batch stochastic gradient descent with a batch size of 64 and an initial learning rate of 0.01 for 140,000 iterations on 484,815 synthetic frames collected from the TORCS simulator across 7 tracks and 22 vehicle models.
  4. Knowl 4 — Dual-ConvNet Structure for Direct 3D Car Coordinate Estimation on KITTI

    model/method

    To apply direct perception to real-world driving scenes lacking explicit lane markings, the 3D space ahead of the vehicle on the ground plane (z=0z = 0) is represented using a 2D Cartesian coordinate frame where the origin is the center of the host car, the positive yy-axis points along the host car's forward heading, and the positive xx-axis points to the right.

    The forward driving space up to 55 m55\text{ m} ahead is partitioned laterally into three distinct semantic regions:

    1. Left Area: x∈[−12,−1.6) mx \in [-12, -1.6)\text{ m}
    2. Central Area: x∈[−1.6,1.6] mx \in [-1.6, 1.6]\text{ m}
    3. Right Area: x∈(1.6,12] mx \in (1.6, 12]\text{ m}

    The network estimates the 2D coordinate pair (x,y)(x, y) of the single closest vehicle in each of the three lateral zones, resulting in 6 continuous regression outputs (xl,yl,xm,ym,xr,yr)(x_l, y_l, x_m, y_m, x_r, y_r).

    Because distant vehicles appear at very low pixel resolution, a dual-ConvNet architecture is used:

    • Close-Range ConvNet: Covers distances y∈[2,25] my \in [2, 25]\text{ m}. Its input is the full monocular frame resized to 497×150497 \times 150 pixels.
    • Far-Range ConvNet: Covers distances y∈[15,55] my \in [15, 55]\text{ m}. Its input is a central crop of the monocular frame covering the central 497×150497 \times 150 region.

    The outputs of the close-range and far-range networks are fused across their overlapping range (15–25 m15\text{--}25\text{ m}) to produce the final continuous vehicle position estimates.

  5. Knowl 5 — Direct Perception vs. GIST and Handcrafted Perception Baselines

    data/table

    The direct perception ConvNet was evaluated against alternative vision pipelines for predicting affordance indicators in TORCS: a Caltech lane detector baseline using Support Vector Regression/Classification (SVR/SVC) and hand-crafted GIST descriptors mapped via SVR/SVC. Performance is measured by Mean Absolute Error (MAE; heading angle in radians, all distances in meters).

    Comparison of lane-related affordance MAE against the Caltech lane detector baseline (evaluated on a zero-traffic test track at 640×480640 \times 480 resolution):

    Model angle to_LL to_ML to_MR to_RR to_L to_M to_R
    Caltech lane detector + SVR/SVC 0.048 1.673 1.179 1.084 1.220 1.113 1.060 0.895
    Direct Perception ConvNet (Full) 0.025 0.260 0.197 0.179 0.239 0.291 0.262 0.231

    Comparison of all 13 affordance indicators against GIST features on a two-lane testing set (8,639 samples with traffic; vehicle distances evaluated within [2,50] m[2, 50]\text{ m}):

    Model angle to_LL to_ML to_MR to_RR dist_LL dist_MM dist_RR to_L to_M to_R dist_L dist_R
    GIST whole 0.051 1.033 0.596 0.598 1.140 18.561 13.081 20.542 1.201 1.310 1.462 30.164 30.138
    GIST half 0.055 1.052 0.547 0.544 1.238 17.643 12.749 22.229 1.156 1.377 1.549 29.484 31.394
    ConvNet sub (86k train) 0.043 0.253 0.180 0.193 0.289 6.168 8.608 9.839 0.345 0.336 0.345 12.681 14.782
    ConvNet full (485k train) 0.033 0.188 0.155 0.159 0.183 5.085 4.738 7.983 0.316 0.308 0.294 8.784 10.740

    The learned ConvNet features substantially outperform hand-crafted GIST descriptors across both lane boundaries and car distance estimations, and training on the full 485k dataset yields significant error reductions over the 86k subset.

  6. Knowl 6 — Quantitative Comparison of Direct Distance Estimation with Mediated DPM on KITTI

    data/table

    The dual-ConvNet direct perception model for vehicle distance estimation was evaluated on 2,200 real-world testing images from the KITTI dataset and compared against a mediated perception baseline comprising a Deformable Part Model (DPM) car detector paired with calibrated ground plane projection (z=0z = 0). Errors are reported as Mean Absolute Error (MAE, in meters) for lateral coordinate xx, longitudinal distance yy, and Euclidean distance d=x2+y2d = \sqrt{x^2 + y^2} for traffic vehicles within 50 m50\text{ m}.

    Penalizing False Positives Without Penalizing False Positives
    Method yy (m) xx (m) dd (m) y\FPy\backslash\text{FP} (m) x\FPx\backslash\text{FP} (m) d\FPd\backslash\text{FP} (m)
    Direct Perception ConvNet 5.832 1.565 6.299 4.332 1.097 4.669
    DPM + Ground Plane Projection 5.824 1.502 6.271 5.000 1.214 5.331

    When false positives are penalized (by setting the ground truth to 50 m50\text{ m} when no car is present), the ConvNet performs comparably to DPM + Projection (6.299 m6.299\text{ m} vs 6.271 m6.271\text{ m} total distance error). When false positives are not penalized, the ConvNet achieves lower error (d=4.669 md = 4.669\text{ m} vs 5.331 m5.331\text{ m}), indicating that direct regression of distances for visible vehicles is more accurate than 2D bounding box detection followed by planar projection.

  7. Knowl 7 — Steering and Longitudinal Car-Following Control Laws

    equation

    The direct perception driving controller executes continuous lateral steering and longitudinal velocity control using closed-form kinematic models driven directly by affordance indicators.

    The lateral steering command steerCmd\text{steerCmd} is given by: steerCmd=C⋅(angle−dist_centerroad_width)\text{steerCmd} = C \cdot \left(\text{angle} - \frac{\text{dist\_center}}{\text{road\_width}}\right) where:

    • steerCmd∈[−1,1]\text{steerCmd} \in [-1, 1] is the normalized steering command.
    • CC is a gain coefficient calibrated for driving conditions.
    • angle∈[−π,π]\text{angle} \in [-\pi, \pi] is the vehicle heading angle relative to the road centerline tangent.
    • dist_center\text{dist\_center} is the lateral distance from the host car's current position to the center line of the target lane.
    • road_width\text{road\_width} is the lateral width of the lane in meters.

    The longitudinal target speed v(t)v(t) when following a leading vehicle is governed by Newell's optimal velocity model: v(t)=vmax⁡(1−exp⁡(−cvmax⁡(dist(t)−d)))v(t) = v_{\max} \left(1 - \exp\left(-\frac{c}{v_{\max}} (\text{dist}(t) - d)\right)\right) where:

    • v(t)v(t) is the desired vehicle speed at time tt in km/h\text{km/h} (or m/s\text{m/s}).
    • vmax⁡v_{\max} is the maximum allowable cruising speed (baseline 72 km/h72\text{ km/h}).
    • dist(t)\text{dist}(t) is the estimated distance to the leading vehicle in meters.
    • cc and dd are positive calibration constants determining deceleration slope and safety stopping distance.
  8. Knowl 8 — Inherent Failure Modes of Behavior Reflex Models in Multi-Vehicle Driving

    empirical result

    When evaluated in the TORCS simulation environment, behavior reflex ConvNets (which directly regress steering angles from raw camera images without an intermediate affordance representation) exhibit distinct failure modes when navigating multi-vehicle traffic compared to empty tracks:

    1. Ill-Posed Training from Multimodal Human Demonstrations: In identical traffic configurations (such as approaching a slower vehicle ahead), human drivers make diverse choices (following behind, overtaking left, or overtaking right). Training a direct regressor on such multimodal data causes target averaging, leading to erratic trajectories and frequent collisions.
    2. Inadequate Abstraction Level: High-level tactical maneuvers (such as overtaking and merging back) are decomposed into sequences of low-level steering commands over time, preventing the reflex model from capturing the spatial-temporal structure of the maneuver.
    3. Weak Supervisory Gradient: Scalar steering angles provide insufficient supervisory signal to force convolutional feature extractors to selectively attend to critical environmental affordances (such as lane markings and distant obstacles) rather than background image clutter.

    Empirically, while a reflex model trained on empty tracks (>60,000 images) successfully maintains lane tracking, the same architecture trained in traffic (>80,000 images) displays erratic steering and repeatedly collides with preceding cars on the training track.

  9. Knowl 9 — Interpretability of Intermediate ConvNet Features for Driving Affordances

    empirical result

    Analysis of the internal representations learned by the direct perception ConvNet reveals that task-specific visual structures emerge spontaneously across different layers without explicit bounding-box or segmentation annotations:

    1. First Fully Connected Layer (fc1\text{fc}_1) Activation Patterns: Averaging the top 100 images that maximally activate individual neurons in fc1\text{fc}_1 (out of 4,096 neurons evaluated over 21,100 samples) shows selective tuning to specific vehicle heading angles, lane boundary locations, and relative vehicle positions.
    2. Convolutional Feature Response Maps: Visualizing the maximum activation value across all feature channels at each spatial pixel location in the fourth convolutional layer (conv4\text{conv}_4) demonstrates strong spatial localization:
      • For networks trained on the KITTI dataset, peak filter responses concentrate tightly on the bounding contours of nearby and preceding vehicles.
      • For networks trained on TORCS, peak filter responses concentrate sharply on lane markings and road edges.

Coverage note — Qualitative video demonstration results on car-mounted smartphone footage and specific TORCS asset replacement counts were omitted in favor of self-contained quantitative and methodological knowls.

References

  1. 1.M. Aly. Real time detection of lane markers in urban streets. In Intelligent Vehicles Symposium, 2008 IEEE, pages 7–12. IEEE, 2008. 2, 5
  2. 2.C.-C. Chang and C.-J. Lin. Libsvm: a library for support vector machines. ACM Transactions on Intelligent Systems and Technology (TIST), 2(3):27, 2011. 5
  3. 3.D. Erhan, C. Szegedy, A. Toshev, and D. Anguelov. Scalable object detection using deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014. 3
  4. 4.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan. Object detection with discriminatively trained part-based models. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 32(9):1627–1645, 2010. 2
  5. 5.A. Geiger, M. Lauer, C. Wojek, C. Stiller, and R. Urtasun. 3d traffic scene understanding from movable platforms. Pattern Analysis and Machine Intelligence (PAMI), 2014. 2, 7
  6. 6.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 2013. 1, 2, 3, 6
  7. 7.J. J. Gibson. The ecological approach to visual perception. Psychology Press, 1979. 2
  8. 8.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014. 3
  9. 9.R. Hadsell, P. Sermanet, J. Ben, A. Erkan, M. Scoffier, K. Kavukcuoglu, U. Muller, and Y. LeCun. Learning long-range vision for autonomous off-road driving. Journal of Field Robotics, 26(2):120–144, 2009. 3
  10. 10.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. arXiv preprint arXiv:1408.5093, 2014. 4
  11. 11.J. Koutn4ık, G. Cuccu, J. Schmidhuber, and F. J. Gomez. Evolving large-scale neural networks for vision-based torcs. In FDG, pages 206–212, 2013. 2
  12. 12.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012. 3, 4
  13. 13.P. Lenz, J. Ziegler, A. Geiger, and M. Roser. Sparse scene flow segmentation for moving object detection in urban environments. In Intelligent Vehicles Symposium (IV), 2011 IEEE, pages 926–932. IEEE, 2011. 2
  14. 14.U. Muller, J. Ben, E. Cosatto, B. Flepp, and Y. L. Cun. Off-road obstacle avoidance through end-to-end learning. In Advances in neural information processing systems, pages 739–746, 2005. 3
  15. 15.G. F. Newell. Nonlinear effects in the dynamics of car following. Operations research, 9(2):209–229, 1961. 4
  16. 16.A. Oliva and A. Torralba. Modeling the shape of the scene: A holistic representation of the spatial envelope. International journal of computer vision, 42(3):145–175, 2001. 3
  17. 17.D. A. Pomerleau. Alvinn: An autonomous land vehicle in a neural network. Technical report, DTIC Document, 1989. 1, 2
  18. 18.D. A. Pomerleau. Neural network perception for mobile robot guidance. Technical report, DTIC Document, 1992. 1, 2
  19. 19.S. Ullman. Against direct perception. Behavioral and Brain Sciences, 3(03):373–381, 1980. 1
  20. 20.P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid. Deepflow: Large displacement optical flow with deep matching. In Computer Vision (ICCV), 2013 IEEE International Conference on, pages 1385–1392. IEEE, 2013. 3
  21. 21.B. Wymann, E. Espié, C. Guionneau, C. Dimitrakakis, R. Coulom, and A. Sumner. TORCS, The Open Racing Car Simulator. http://www.torcs.org, 2014. 3
  22. 22.H. Zhang, A. Geiger, and R. Urtasun. Understanding high-level semantics by modeling traffic patterns. In Computer Vision (ICCV), 2013 IEEE International Conference on, pages 3056–3063. IEEE, 2013. 2

Citation

MLA
Chen, C., et al. “DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving”. 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 2722–30, https://doi.org/10.1109/ICCV.2015.312.
APA
Chen, C., Seff, A., Kornhauser, A., & Xiao, J. (2015). DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving. 2015 IEEE International Conference on Computer Vision (ICCV), 2722–2730. https://doi.org/10.1109/ICCV.2015.312
Chicago
Chen, C., A. Seff, A. Kornhauser, and J. Xiao. 2015. “DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving”. 2015 IEEE International Conference on Computer Vision (ICCV), 2722–30. https://doi.org/10.1109/ICCV.2015.312.
Harvard
Chen, C. et al. (2015) “DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving”, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, pp. 2722–2730. Available at: https://doi.org/10.1109/ICCV.2015.312.
Vancouver
1. Chen C, Seff A, Kornhauser A, Xiao J (2015) DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving. In: 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, pp 2722–2730

BibTeX

@inproceedings{Chen_2015, title={DeepDriving: Learning Affordance for Direct Perception in Autonomous Driving}, url={http://dx.doi.org/10.1109/ICCV.2015.312}, DOI={10.1109/iccv.2015.312}, booktitle={2015 IEEE International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Chen, Chenyi and Seff, Ari and Kornhauser, Alain and Xiao, Jianxiong}, year={2015}, month=Dec, pages={2722–2730} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE