Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges

Di FengChristian Haase-SchützLars RosenbaumHeinz HertleinClaudius GlaeserFabian TimmWerner WiesbeckKlaus Dietmayer

article2019IEEE transactions on intelligent transportation systems (Print)1,381 citations

Systematizes multi-sensor fusion methodologies across camera, LiDAR, and radar modalities to provide clear architectural guidelines on what, when, and how to integrate data for autonomous object detection and semantic segmentation.

Listen

Autonomous vehicles must operate reliably in complex environments, where minor perception errors can result in catastrophic accidents. To achieve safe navigation, perception systems require precise environmental data, robust handling of adverse conditions, and real-time execution speeds. Relying on single-sensor solutions such as cameras or Light Detection and Ranging (LiDAR) introduces failure modes under varying weather, lighting, and sparse depth conditions. Fusing complementary modalities through deep learning has become the primary avenue to overcome these individual sensor limitations.

The article systematically reviews and evaluates deep multi-modal object detection and semantic segmentation methods for autonomous driving. It examines sensor setups, public datasets, fusion architectures, and performance metrics while identifying critical challenges in multi-modal system design.

To conduct this evaluation, the authors analyzed public autonomous driving datasets published between 2013 and 2019 and reviewed dozens of deep learning fusion architectures. The analysis categorized methods by input representations, architectural fusion stages (early, middle, or late), and mathematical fusion mechanisms (such as concatenation, addition, ensembling, and Mixture of Experts).

The article highlights five key findings. First, combining camera images and LiDAR point clouds consistently outperforms single-modality baselines in detection accuracy, while the top-performing models across major benchmarks rely universally on deep learning. Second, there is no universally optimal fusion architecture: performance varies significantly depending on network topologies, sensor representations, and application domains. Third, multi-sensor datasets expanded by two orders of magnitude between 2014 and 2019, yet they remain small compared to traditional computer vision datasets and suffer from heavy class imbalances and limited environmental diversity. Fourth, research on integrating Radar and Ultrasonic sensors with deep learning remains sparse despite their low cost and robustness in adverse conditions. Finally, most current fusion networks rely on rigid, empirical operations that fail to quantify sensor uncertainty or dynamically weight reliable modalities when a sensor degrades.

These findings indicate that relying solely on static, empirical fusion designs increases system-level safety risks, particularly under open-set and adverse conditions. Standard evaluation metrics focus almost entirely on predictive accuracy, masking vulnerabilities such as sensor misalignment and failure modes. Improving overall system safety requires explicitly modeling uncertainty and dynamically adapting sensor weights, which directly affects downstream planning, safety certification, and commercial deployment.

Organizations developing autonomous systems should prioritize incorporating probabilistic uncertainty estimation, such as Bayesian neural networks and Mixture of Experts gating mechanisms, into perception stacks. Stakeholders should also invest in automated sensor calibration tools and efficient data labeling frameworks, while expanding research into Radar fusion and simulation-based data augmentation. Because standardized hardware benchmarks and temporal-aware fusion frameworks remain underdeveloped, technical teams should evaluate multi-modal models under simulated sensor faults and realistic automotive processing constraints before committing to fixed hardware architectures.

Cover for Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges

Abstract

Recent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs, Radars), and multiple sensing modalities can be fused to exploit their complementary properties. In this context, many methods have been proposed for deep multi-modal perception problems. However, there is no general guideline for network architecture design, and questions of "what to fuse", "when to fuse", and "how to fuse" remain open. This review paper attempts to systematically summarize methodologies and discuss challenges for deep multi-modal object detection and semantic segmentation in autonomous driving. To this end, we first provide an overview of on-board sensors on test vehicles, open datasets, and background information for object detection and semantic segmentation in autonomous driving research. We then summarize the fusion methodologies and discuss challenges and open questions. In the appendix, we provide tables that summarize topics and methods. We also provide an interactive online platform to navigate each reference: this https URL.

Table of Contents

  • I Introduction
  • I-A Related Works
  • I-B Contributions
  • II Background
  • II-A Sensing Modalities for Autonomous Driving
  • II-A1 Visual and Thermal Cameras
  • II-A2 LiDARs
  • II-A3 Radars
  • II-A4 Ultrasonics
  • II-A5 GNSS and HD Maps
  • II-A6 IMU and Odometers
  • II-B Test Vehicle Setup
  • II-C Deep Object Detection
  • II-C1 Two-stage Object Detection
  • II-C2 One-stage Object Detection
  • II-D Deep Semantic Segmentation
  • III Multi-modal Datasets
  • III-A Sensing Modalities
  • III-B Recording Conditions
  • III-C Dataset Size
  • III-D Labels
  • IV Deep Multi-modal Perception Problems for Autonomous Driving
  • IV-A Deep Multi-modal Object Detection
  • IV-A1 Sensing Modalities
  • IV-A2 2D or 3D Detection
  • IV-A3 What to detect
  • IV-B Deep Multi-modal Semantic Segmentation
  • V Methodology
  • V-A What to Fuse
  • V-A1 LiDAR Point Clouds
  • V-A2 Camera Images
  • V-A3 Processing LiDAR Points and Camera Images in Deep Multi-modal Perception
  • V-A4 Radar Signals
  • V-B How to Fuse
  • V-B1 Addition or Average Mean
  • V-B2 Concatenation
  • V-B3 Ensemble
  • V-B4 Mixture of Experts
  • V-C When to Fuse
  • V-C1 Early Fusion
  • V-C2 Late Fusion
  • V-C3 Middle Fusion
  • V-C4 Fusion in Object Detection Networks
  • V-C5 Fusion Operation and Fusion Scheme
  • VI Challenges and Open Questions
  • VI-A Multi-modal Data Preparation
  • VI-A1 Data Diversity
  • VI-A2 Data Quality and Alignment
  • VI-B Fusion Methodology
  • VI-B1 What to Fuse
  • VI-B2 How to Fuse
  • VI-B3 When to Fuse
  • VI-B4 Real-time Consideration
  • VI-C Others
  • VI-C1 Evaluation Metrics
  • VI-C2 More Network Architectures
  • VII Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — Mathematical Formulations of Multi-Modal Fusion Stages

    equation

    In a deep neural network processing two distinct sensing modalities MiM_i and MjM_j, let flMif_l^{M_i} and flMjf_l^{M_j} denote the intermediate feature representations in layer l∈{0,1,…,L}l \in \{0, 1, \dots, L\}, where l=0l=0 corresponds to the input representation and LL denotes the final output layer. Let Gl(⋅)G_l(\cdot) be the layer-wise feature transformation at layer ll, and let ⊕\oplus denote a general fusion operation (such as concatenation, addition, or gating).

    1. Early Fusion: Sensor inputs or low-level feature representations are combined prior to subsequent transformations in an (L+1)(L+1)-layer network: fL=GL(GL−1(…G2(G1(f0Mi⊕f0Mj))… ))f_L = G_L\left(G_{L-1}\left(\dots G_2\left(G_1\left(f_0^{M_i} \oplus f_0^{M_j}\right)\right)\dots\right)\right)

    2. Late Fusion: Decision outputs from modality-specific networks are integrated at the final layer: fL=GLMi(GL−1Mi(…G1Mi(f0Mi)… ))⊕GLMj(GL−1Mj(…G1Mj(f0Mj)… ))f_L = G_L^{M_i}\left(G_{L-1}^{M_i}\left(\dots G_1^{M_i}\left(f_0^{M_i}\right)\dots\right)\right) \oplus G_L^{M_j}\left(G_{L-1}^{M_j}\left(\dots G_1^{M_j}\left(f_0^{M_j}\right)\dots\right)\right)

    3. Single-Layer Middle Fusion: Modality branches process inputs independently up to layer l∗l^*, where features are combined once before passing through downstream shared layers: fL=GL(…Gl∗+1(Gl∗Mi(…G1Mi(f0Mi)… )⊕Gl∗Mj(…G1Mj(f0Mj)… ))… )f_L = G_L\left(\dots G_{l^*+1}\left(G_{l^*}^{M_i}\left(\dots G_1^{M_i}\left(f_0^{M_i}\right)\dots\right) \oplus G_{l^*}^{M_j}\left(\dots G_1^{M_j}\left(f_0^{M_j}\right)\dots\right)\right)\dots\right)

    4. Deep Hierarchical Middle Fusion: Features are fused repeatedly across consecutive intermediate layers starting at layer l∗l^*: fl∗+1=fl∗Mi⊕fl∗Mjf_{l^*+1} = f_{l^*}^{M_i} \oplus f_{l^*}^{M_j} fk+1=GkMi(fk)⊕GkMj(fk),∀k∈{l∗+1,…,L−1}f_{k+1} = G_k^{M_i}(f_k) \oplus G_k^{M_j}(f_k), \quad \forall k \in \{l^*+1, \dots, L-1\}

    5. Short-Cut Middle Fusion: Features are combined with residual connections bridging earlier non-consecutive representations: fl+1=flMi⊕flMjf_{l+1} = f_l^{M_i} \oplus f_l^{M_j} fk+1=fk⊕fk∗Mi⊕fk∗Mj,∀k∈{l+1,…,L−1},∃k∗∈{1,…,l−1}f_{k+1} = f_k \oplus f_{k^*}^{M_i} \oplus f_{k^*}^{M_j}, \quad \forall k \in \{l+1, \dots, L-1\}, \quad \exists k^* \in \{1, \dots, l-1\}

  2. Knowl 2 — Mixture of Experts Gating Mechanism for Deep Multi-Modal Perception

    model/method

    The Mixture of Experts (MoE) fusion method explicitly weights the informativeness of individual sensing modalities rather than assuming equal contributions across changing environmental conditions (e.g., daylight versus night driving).

    Given two sensing modalities MiM_i and MjM_j with feature representations fl−1Mif_{l-1}^{M_i} and fl−1Mjf_{l-1}^{M_j} at layer l−1l-1, each modality representation is processed by a domain-specific expert subnetwork. A gating network takes the combined expert feature representations hh (typically obtained via feature concatenation) as input and outputs modality weights wMiw_{M_i} and wMjw_{M_j} constrained such that: wMi+wMj=1,wMi,wMj≥0w_{M_i} + w_{M_j} = 1, \quad w_{M_i}, w_{M_j} \ge 0

    The fused feature map flf_l at layer ll transformed by the layer operator Gl(⋅)G_l(\cdot) is computed as: fl=Gl(wMi⋅fl−1Mi+wMj⋅fl−1Mj)f_l = G_l\left(w_{M_i} \cdot f_{l-1}^{M_i} + w_{M_j} \cdot f_{l-1}^{M_j}\right)

    This formulation dynamically attenuates degraded sensor streams (e.g., camera channels in unlit scenarios or fog) while upweighting complementary, reliable sensor streams (e.g., LiDAR or thermal signals).

  3. Knowl 3 — LiDAR Point Cloud Representations for Deep Perception

    model/method

    In deep multi-modal perception systems, 3D LiDAR point clouds—which provide 3D spatial coordinates [x,y,z][x, y, z] and reflectance intensity—are processed using three main representation strategies:

    1. 3D Voxelization: The continuous 3D space is discretized into regular volumetric voxel grids. Voxels preserve fine 3D geometric structures but introduce extreme sparsity and high memory/computational costs when using dense 3D convolutions. Methods mitigate this using Voxel Feature Encoding (VFE) layers or sparse convolutional layers that only execute computations on non-empty voxels.

    2. Direct Vector-Space Learning (Continuous Space): PointNet-based architectures process raw, un-voxelized point sets [xi,yi,zi,ri][x_i, y_i, z_i, r_i] using permutation-invariant operations (point-wise multilayer perceptrons followed by symmetrical pooling operators), continuous convolutions, or χ\chi-transformations to aggregate local geometric neighborhoods directly in metric space.

    3. 2D Projected Grid Representations:

    • Bird's Eye View (BEV) Maps: Points are projected onto the ground plane (x,y)(x, y) and discretized into 2D grid cells encoding height slices, point density (count per column), and maximum intensity. BEV eliminates perspective occlusions and preserves true physical object dimensions and metric distances along the ground plane.
    • Camera-Plane Maps (CPM): Points are projected onto the front-facing camera imaging plane using sensor extrinsic and intrinsic calibration matrices. CPM shares the camera resolution and spatial coordinate system but produces sparse depth grids requiring upsampling via bilateral filtering, mean pooling, or nearest neighbors.
    • Spherical / Range Maps: Points are mapped to a 2D cylindrical or spherical grid parameterized by azimuth and zenith angles, yielding dense 2D representations suitable for range-based segmentation, though spatial dimensions diverge from front-view camera frames.
  4. Knowl 4 — Architectural Paradigms for Two-Stage Multi-Modal 3D Object Detection

    model/method

    Two-stage multi-modal 3D object detection frameworks combine visual images and LiDAR point clouds using distinct region proposal (RP) and feature aggregation strategies:

    1. BEV-Driven Multi-View Fusion (e.g., MV3D): Generates 3D candidate bounding box proposals exclusively from LiDAR Bird's Eye View (BEV) feature representations via a 3D Region Proposal Network (RPN). Each 3D proposal is then projected onto LiDAR BEV maps, LiDAR front-view cylindrical maps, and RGB camera images. Regional feature vectors extracted via RoI pooling across all views are combined using deep hierarchical fusion.

    2. Joint-Feature RPN Fusion (e.g., AVOD): Fuses high-resolution feature maps from RGB images and LiDAR BEV representations prior to proposal generation. A shared RPN operates on the combined feature space to output 3D oriented bounding box candidates, and a subsequent detection network refines proposals using RoI pooling on the multi-modal feature maps.

    3. 2D-Driven Frustum Proposal Networks (e.g., Frustum PointNet, RoarNet): Employs a pre-trained 2D CNN detector on high-resolution RGB camera images to produce 2D bounding boxes. Each 2D box is extruded into 3D space using camera projection geometry to define a 3D viewing frustum in the LiDAR point cloud. Points inside the frustum are isolated and passed directly to PointNet-based architectures for point-cloud segmentation and 3D bounding box regression.

    4. Independent Proposal Ensemble Networks: Generates candidate 2D/3D region proposals separately from camera images and LiDAR range/depth images using independent region-proposal pipelines (e.g., selective search, point-cloud Euclidean clustering, or modality-specific RPNs). The union of candidate regions is merged via non-maximum suppression or ensemble scoring prior to regional classification.

  5. Knowl 5 — Operational Mechanisms for Multi-Modal Feature Fusion

    definition

    In deep multi-modal architectures, combining feature maps flMif_l^{M_i} and flMjf_l^{M_j} from modalities MiM_i and MjM_j at network layer ll is achieved through four primary mathematical operations:

    1. Element-wise Addition and Average Mean: Computes the sum or mean across aligned spatial and channel dimensions: fl=Gl−1(fl−1Mi+fl−1Mj)orfl=Gl−1(12(fl−1Mi+fl−1Mj))f_l = G_{l-1}\left(f_{l-1}^{M_i} + f_{l-1}^{M_j}\right) \quad \text{or} \quad f_l = G_{l-1}\left(\frac{1}{2}\left(f_{l-1}^{M_i} + f_{l-1}^{M_j}\right)\right) This requires identical spatial dimensions and channel depths.

    2. Concatenation: Stacks feature maps along the channel (depth) axis before advancing to subsequent convolutional filters, or concatenates flattened feature vectors in fully-connected layers: fl=Gl−1([fl−1Mi  ∥  fl−1Mj])f_l = G_{l-1}\left([f_{l-1}^{M_i} \; \Vert \; f_{l-1}^{M_j}]\right) where ∥\Vert represents channel-wise concatenation.

    3. Ensemble Operation: Unifies predicted outputs or candidate region proposals generated independently by each modality branch: fl=Gl−1(fl−1Mi)∪Gl−1(fl−1Mj)f_l = G_{l-1}\left(f_{l-1}^{M_i}\right) \cup G_{l-1}\left(f_{l-1}^{M_j}\right)

    4. Mixture of Experts: Combines modality streams via dynamic convex combinations using weights generated by an auxiliary gating subnetwork conditioned on the joint input features.

  6. Knowl 6 — Trade-offs Among Early, Middle, and Late Fusion Architectures

    model/method

    The choice of fusion stage introduces structural trade-offs across efficiency, modularity, and sensitivity:

    • Early Fusion: Merges raw or preprocessed multi-modal inputs at the network entry. It has the lowest computational cost and memory footprint because a single unified backbone processes all signals. However, it exhibits strict architectural inflexibility (altering input channels requires full network retraining) and high sensitivity to spatial-temporal sensor calibration errors, differing sampling rates, or missing sensor channels.
    • Late Fusion: Combines decision-level outputs or class score distributions from independent single-modality models. It provides maximum modularity (modalities can be added, removed, or updated without retraining other branches) and high robustness to individual sensor dropouts, but it prevents the network from learning cross-modal feature correlations at intermediate abstraction levels and incurs high cumulative computational and memory overhead.
    • Middle Fusion: Combines intermediate layer representations across modality-specific branches. It allows the model to learn cross-modal representations at varying semantic depths and receptive fields. While it balances representational power and computational efficiency, identifying the optimal intermediate layers and fusion operations requires extensive architecture search.
  7. Knowl 7 — Benchmark Results for 3D Object Detection on the KITTI Test Benchmark

    data/table

    The table below compares 3D object detection average precision (AP in %) across Car, Pedestrian, and Cyclist classes under Moderate, Easy, and Hard difficulty levels, along with reported inference runtime and computing environments on the KITTI vision benchmark test dataset:

    Reference Car Pedestrian Cyclist Runtime Environment
    Mod. Easy Hard Mod. Easy Hard Mod. Easy Hard
    Liang et al., 2019 76.75% 86.81% 68.41% 45.61% 52.37% 41.49% 64.68% 79.58% 57.03% 0.08 s GPU @ 2.5 GHz (Python)
    Wang et al., 2019 76.51% 85.88% 68.08% - - - - - - 0.47 s GPU @ 2.5 GHz (Python + C/C++)
    Sindagi et al., 2019 72.70% 83.20% 65.12% - - - - - - - -
    Shin et al., 2018 73.04% 83.71% 59.16% - - - - - - - GPU Titan X (not Pascal)
    Du et al., 2018 73.80% 84.33% 64.83% - - - - - - 0.50 s GPU @ 2.5 GHz (Matlab + C/C++)
    Liang et al., 2018 66.22% 82.54% 64.04% - - - - - - 0.06 s GPU @ 2.5 GHz (Python)
    Ku et al., 2017 71.88% 81.94% 66.38% 42.81% 50.80% 40.88% 52.18% 64.00% 46.61% 0.10 s GPU Titan X (Pascal)
    Qi et al., 2017 70.39% 81.20% 62.19% 44.89% 51.21% 40.23% 56.77% 71.96% 50.39% 0.17 s GPU @ 3.0 GHz (Python)
    Chen et al., 2017 62.35% 71.09% 55.12% - - - - - - 0.36 s GPU @ 2.5 GHz (Python + C/C++)

    The benchmark results show that multi-task multi-sensor fusion (Liang et al., 2019) achieves the highest 3D detection AP across all classes (76.75% Moderate Car AP, 45.61% Pedestrian, 64.68% Cyclist) at 80 ms latency, while continuous fusion (Liang et al., 2018) demonstrates the fastest inference at 60 ms.

  8. Knowl 8 — Benchmark Results for Urban Road Semantic Segmentation on the KITTI Test Benchmark

    data/table

    The table below summarizes performance metrics—Maximum F1-measure (MaxF in %), Average Precision (AP in %), Precision (PRE in %), Recall (REC in %), False Positive Rate (FPR in %), False Negative Rate (FNR in %), reported runtime, and compute environments—for camera-LiDAR urban road semantic segmentation on the KITTI test set:

    Method MaxF AP PRE REC FPR FNR Runtime Environment
    Chen et al., 2019 97.03% 94.03% 97.19% 96.88% 1.54% 3.12% 0.16 s GPU
    Caltagirone et al., 2019 96.03% 93.93% 96.23% 95.83% 2.07% 4.17% 0.15 s GPU
    Gu et al., 2018 95.22% 89.31% 94.69% 95.76% 2.96% 4.24% 0.07 s CPU
    Lv et al., 2018 94.48% 93.65% 94.28% 94.69% 3.17% 5.31% - GPU Titan X
    Yang et al., 2018 91.40% 84.22% 89.09% 93.84% 6.33% 6.16% - GPU

    The results demonstrate that middle-level cross-fusion combining RGB imagery with LiDAR altitude difference maps (Chen et al., 2019) yields the leading MaxF of 97.03% and AP of 94.03% with a 1.54% false positive rate at 160 ms per frame.

  9. Knowl 9 — Radar Representations and Processing in Deep Perception Networks

    model/method

    Automotive Radar signals—which provide target range, azimuth angle, reflection amplitude, and radial velocity via the Doppler effect—are processed in deep perception models through four primary data structures:

    1. Radar Grid Maps: Accumulate radar reflection data across temporal windows into 2D spatial occupancy grid maps processed by 2D CNNs for static object classification and semantic freespace segmentation.
    2. Range-Velocity / Range-Doppler Spectra: Post-processed 2D range-Doppler or range-velocity spectrograms are fed into 2D CNNs, stacked autoencoders, or Convolutional Recurrent Neural Networks (CRNNs) to classify moving obstacles and human motion.
    3. Radar Point Sets: Radar target reflections are structured directly as sparse 3D point sets [x,y,z,vr][x, y, z, v_r] containing radial Doppler velocities vrv_r and processed using PointNet++ or cluster-based feature encoders for dynamic object segmentation.
    4. Camera-Plane Radar Projections: Radar points are projected onto the 2D front-view camera plane to construct dense range-velocity feature maps aligned with RGB pixel coordinates for distant vehicle detection.
  10. Knowl 10 — Uncertainty Quantification in Multi-Modal Deep Perception

    model/method

    To prevent catastrophic perception errors under sensor degradation, adverse weather, or out-of-distribution (open-set) conditions, multi-modal networks can model two complementary forms of predictive uncertainty:

    1. Epistemic Uncertainty (Model Uncertainty): Represents uncertainty in the network parameters due to limited training data. It is estimated using Bayesian Neural Networks (BNNs) through variational inference over weight posteriors, Monte Carlo dropout sampling at inference time, batch normalization uncertainty, or noise injection.
    2. Aleatoric Uncertainty (Observation/Data Noise): Captures irreducible sensor measurement noise (e.g., thermal sensor bloom, LiDAR backscatter in fog, camera sensor glare). It is modeled by training the network to directly output the parameters of a predictive probability distribution (e.g., predicting mean μ\mu and variance σ2\sigma^2 under a Gaussian or Laplacian likelihood) via heteroscedastic loss functions.

    Propagating these uncertainty quantities into downstream tracking, maneuver forecasting, and path-planning modules allows autonomous vehicles to adjust confidence thresholds and trigger safety fallbacks during sensor failures.

Coverage note — No substantial contributed material was omitted. The survey's foundational taxonomies (what/when/how to fuse), mathematical formulations of fusion schemes and gating networks, LiDAR and Radar representation strategies, 3D detection paradigms, benchmark tables on KITTI, and open challenges (uncertainty quantification, real-time trade-offs) are fully covered.

References

  1. 1.E. D. Dickmanns and B. D. Mysliwetz, “Recursive 3-d road and relative ego-state recognition,” IEEE Trans. Pattern Anal. Mach. Intell., no. 2, pp. 199–213, 1992.
  2. 2.C. Urmson et al., “Autonomous driving in urban environments: Boss and the urban challenge,” J. Field Robotics, vol. 25, no. 8, pp. 425–466, 2008.
  3. 3.R. Berger, “Autonomous driving,” Think Act, 2014. [Online]. Available: http://www.rolandberger.ch/media/pdf/Roland_Berger_TABAutonomousDrivingfinal20141211
  4. 4.G. Neuhold, T. Ollmann, S. R. Bulo, and P. Kontschieder, “The Mapillary Vistas dataset for semantic understanding of street scenes,” in Proc. IEEE Conf. Computer Vision, Oct. 2017, pp. 5000–5009.
  5. 5.Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, p. 436, 2015.
  6. 6.A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the KITTI vision benchmark suite,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2012.
  7. 7.H. Yin and C. Berger, “When to use what data set for your self-driving car algorithm: An overview of publicly available driving datasets,” in IEEE 20th Int. Conf. Intelligent Transportation Systems, 2017, pp. 1–8.
  8. 8.D. Ramachandram and G. W. Taylor, “Deep multimodal learning: A survey on recent advances and trends,” IEEE Signal Process. Mag., vol. 34, no. 6, pp. 96–108, 2017.
  9. 9.J. Janai, F. Güney, A. Behl, and A. Geiger, “Computer vision for autonomous vehicles: Problems, datasets and state-of-the-art,” arXiv:1704.05519 [cs.CV], 2017.
  10. 10.E. Arnold, O. Y. Al-Jarrah, M. Dianati, S. Fallah, D. Oxtoby, and A. Mouzakitis, “A survey on 3d object detection methods for autonomous driving applications,” IEEE Trans. Intell. Transp. Syst., pp. 1–14, 2019.
  11. 11.L. Liu et al., “Deep learning for generic object detection: A survey,” arXiv:1809.02165 [cs.CV], 2018.
  12. 12.A. Garcia-Garcia, S. Orts-Escolano, S. Oprea, V. Villena-Martinez, and J. Garcia-Rodriguez, “A review on deep learning techniques applied to semantic segmentation,” Applied Soft Computing, 2017.
  13. 13.K. Bengler, K. Dietmayer, B. Farber, M. Maurer, C. Stiller, and H. Winner, “Three decades of driver assistance systems: Review and future perspectives,” IEEE Intell. Transp. Syst. Mag., vol. 6, no. 4, pp. 6–22, 2014.
  14. 14.Waymo. (2017) Waymo safety report: On the road to fully self-driving. [Online]. Available: https://waymo.com/safety
  15. 15.M. Aeberhard et al., “Experience, results and lessons learned from automated driving on Germany’s highways,” IEEE Intell. Transp. Syst. Mag., vol. 7, no. 1, pp. 42–57, 2015.
  16. 16.J. Ziegler et al., “Making Bertha drive – an autonomous journey on a historic route,” IEEE Intell. Transp. Syst. Mag., vol. 6, no. 2, pp. 8–20, 2014.
  17. 17.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman, “The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results,” http://www.pascalnetwork.org/challenges/VOC/voc2007/workshop/index.html.
  18. 18.T.-Y. Lin et al., “Microsoft COCO: Common objects in context,” in Proc. Eur. Conf. Computer Vision. Springer, 2014, pp. 740–755.
  19. 19.M. Weber, P. Wolf, and J. M. Zöllner, “DeepTLR: A single deep convolutional network for detection and classification of traffic lights,” in IEEE Intelligent Vehicles Symp., 2016, pp. 342–348.
  20. 20.J. Müller and K. Dietmayer, “Detecting traffic lights by single shot detection,” in 21st Int. Conf. Intelligent Transportation Systems. IEEE, 2016, pp. 342–348.
  21. 21.M. Bach, S. Reuter, and K. Dietmayer, “Multi-camera traffic light recognition using a classifying labeled multi-bernoulli filter,” in IEEE Intelligent Vehicles Symp., 2017, pp. 1045–1051.
  22. 22.K. Behrendt, L. Novak, and R. Botros, “A deep learning approach to traffic lights: Detection, tracking, and classification,” in IEEE Int. Conf. Robotics and Automation, 2017, pp. 1370–1377.
  23. 23.Z. Zhu, D. Liang, S. Zhang, X. Huang, B. Li, and S. Hu, “Traffic-sign detection and classification in the wild,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 2110–2118.
  24. 24.H. S. Lee and K. Kim, “Simultaneous traffic sign detection and boundary estimation using convolutional neural network,” IEEE Trans. Intell. Transp. Syst., 2018.
  25. 25.H. Luo, Y. Yang, B. Tong, F. Wu, and B. Fan, “Traffic sign recognition using a multi-task convolutional neural network,” IEEE Trans. Intell. Transp. Syst., vol. 19, no. 4, pp. 1100–1111, 2018.
  26. 26.S. Zhang, R. Benenson, M. Omran, J. Hosang, and B. Schiele, “Towards reaching human performance in pedestrian detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 4, pp. 973–986, 2018.
  27. 27.L. Zhang, L. Lin, X. Liang, and K. He, “Is Faster R-CNN doing well for pedestrian detection?” in Proc. Eur. Conf. Computer Vision. Springer, 2016, pp. 443–457.
  28. 28.X. Chen, K. Kundu, Y. Zhu, H. Ma, S. Fidler, and R. Urtasun, “3d object proposals using stereo imagery for accurate object class detection,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 5, pp. 1259–1272, 2018.
  29. 29.B. Li, “3d fully convolutional network for vehicle detection in point cloud,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2017, pp. 1513–1518.
  30. 30.B. Li, T. Zhang, and T. Xia, “Vehicle detection from 3d lidar using fully convolutional network,” in Proc. Robotics: Science and Systems, Jun. 2016.
  31. 31.X. Chen, K. Kundu, Z. Zhang, H. Ma, S. Fidler, and R. Urtasun, “Monocular 3d object detection for autonomous driving,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 2147–2156.
  32. 32.J. Fang, Y. Zhou, Y. Yu, and S. Du, “Fine-grained vehicle model recognition using a coarse-to-fine convolutional neural network architecture,” IEEE Trans. Intell. Transp. Syst., vol. 18, no. 7, pp. 1782–1792, 2017.
  33. 33.A. Mousavian, D. Anguelov, J. Flynn, and J. Kosecká, “3d bounding box estimation using deep learning and geometry,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2017, pp. 5632–5640.
  34. 34.P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “OverFeat: Integrated recognition, localization and detection using convolutional networks,” in Int. Conf. Learning Representations, 2013.
  35. 35.R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2014, pp. 580–587.
  36. 36.K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 37, no. 9, pp. 1904–1916, 2015.
  37. 37.R. Girshick, “Fast R-CNN,” in Proc. IEEE Conf. Computer Vision, 2015, pp. 1440–1448.
  38. 38.K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv:1409.1556 [cs.CV], 2014.
  39. 39.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 770–778.
  40. 40.C. Szegedy et al., “Going deeper with convolutions,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2015, pp. 1–9.
  41. 41.S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in Advances in Neural Information Processing Systems, 2015, pp. 91–99.
  42. 42.J. Dai, Y. Li, K. He, and J. Sun, “R-FCN: Object detection via region-based fully convolutional networks,” in Advances in Neural Information Processing Systems, 2016, pp. 379–387.
  43. 43.C. Szegedy, A. Toshev, and D. Erhan, “Deep neural networks for object detection,” in Advances in Neural Information Processing Systems, 2013, pp. 2553–2561.
  44. 44.J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 779–788.
  45. 45.W. Liu et al., “SSD: Single shot multibox detector,” in Proc. Eur. Conf. Computer Vision. Springer, 2016, pp. 21–37.
  46. 46.J. Huang et al., “Speed/accuracy trade-offs for modern convolutional object detectors,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, vol. 4, 2017.
  47. 47.A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, “Panoptic segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018.
  48. 48.A. Kirillov, R. Girshick, K. He, and P. Dollár, “Panoptic feature pyramid networks,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 6399–6408.
  49. 49.Y. Xiong et al., “Upsnet: A unified panoptic segmentation network,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 8818–8826.
  50. 50.L. Porzi, S. R. Bulo, A. Colovic, and P. Kontschieder, “Seamless scene segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 8277–8286.
  51. 51.B. Wu, A. Wan, X. Yue, and K. Keutzer, “SqueezeSeg: Convolutional neural nets with recurrent CRF for real-time road-object segmentation from 3d lidar point cloud,” in IEEE Int. Conf. Robotics and Automation, May 2018, pp. 1887–1893.
  52. 52.L. Caltagirone, S. Scheidegger, L. Svensson, and M. Wahde, “Fast lidar-based road detection using fully convolutional neural networks,” in IEEE Intelligent Vehicles Symp., 2017, pp. 1019–1024.
  53. 53.Q. Huang, W. Wang, and U. Neumann, “Recurrent slice networks for 3d segmentation of point clouds,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 2626–2635.
  54. 54.A. Dewan, G. L. Oliveira, and W. Burgard, “Deep semantic classification for 3d lidar data,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2017, pp. 3544–3549.
  55. 55.A. Dewan and W. Burgard, “DeepTemporalSeg: Temporally consistent semantic segmentation of 3d lidar scans,” arXiv preprint arXiv:1906.06962, 2019.
  56. 56.A. Milioto, I. Vizzo, J. Behley, and C. Stachniss, “RangeNet++: Fast and Accurate LiDAR Semantic Segmentation,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2019.
  57. 57.M. Cordts et al., “The Cityscapes dataset for semantic urban scene understanding,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 3213–3223.
  58. 58.S. Wang et al., “TorontoCity: Seeing the world with a million eyes,” in Proc. IEEE Conf. Computer Vision, 2017, pp. 3028–3036.
  59. 59.G. Neuhold, T. Ollmann, S. Rota Bulo, and P. Kontschieder, “The mapillary vistas dataset for semantic understanding of street scenes,” in Proc. IEEE Conf. Computer Vision, 2017. [Online]. Available: https://www.mapillary.com/dataset/vistas
  60. 60.X. Huang et al., “The ApolloScape dataset for autonomous driving,” in Workshop Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 954–960.
  61. 61.L. Schneider et al., “Multimodal neural networks: RGB-D for semantic segmentation and object detection,” in Scandinavian Conf. Image Analysis. Springer, 2017, pp. 98–109.
  62. 62.V. Badrinarayanan, A. Kendall, and R. Cipolla, “SegNet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., no. 12, pp. 2481–2495, 2017.
  63. 63.M. Teichmann, M. Weber, M. Zoellner, R. Cipolla, and R. Urtasun, “MultiNet: Real-time joint semantic reasoning for autonomous driving,” in IEEE Intelligent Vehicles Symp., 2018.
  64. 64.K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in Proc. IEEE Conf. Computer Vision, 2017, pp. 2980–2988.
  65. 65.J. Uhrig, E. Rehder, B. Fröhlich, U. Franke, and T. Brox, “Box2Pix: Single-shot instance segmentation by assigning pixels to object boxes,” in IEEE Intelligent Vehicles Symp., 2018.
  66. 66.S. Gupta, R. Girshick, P. Arbelaez, and J. Malik, “Learning rich features from RGB-D images for object detection and segmentation,” in Proc. Eur. Conf. Computer Vision. Springer, 2014, pp. 345–360.
  67. 67.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik, “Simultaneous detection and segmentation,” in Proc. Eur. Conf. Computer Vision. Springer, 2014, pp. 297–312.
  68. 68.J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2015, pp. 3431–3440.
  69. 69.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 4, pp. 834–848, 2018.
  70. 70.A. Paszke, A. Chaurasia, S. Kim, and E. Culurciello, “ENet: A deep neural network architecture for real-time semantic segmentation,” arXiv:1606.02147 [cs.CV], 2016.
  71. 71.A. Roy and S. Todorovic, “A multi-scale CNN for affordance segmentation in RGB images,” in Proc. Eur. Conf. Computer Vision. Springer, 2016, pp. 186–201.
  72. 72.S. Zheng et al., “Conditional random fields as recurrent neural networks,” in Proc. IEEE Conf. Computer Vision, 2015, pp. 1529–1537.
  73. 73.M. Siam, M. Gamal, M. Abdel-Razek, S. Yogamani, M. Jagersand, and H. Zhang, “A comparative study of real-time semantic segmentation for autonomous driving,” in Workshop Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 587–597.
  74. 74.W. Maddern, G. Pascoe, C. Linegar, and P. Newman, “1 year, 1000 km: The Oxford RobotCar dataset,” Int. J. Robotics Research, vol. 36, no. 1, pp. 3–15, 2017.
  75. 75.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The KITTI dataset,” Int. J. Robotics Research, 2013.
  76. 76.J.-L. Blanco-Claraco, F.-A. Moreno-Dueñas, and J. González-Jiménez, “The Malaga urban dataset: High-rate stereo and lidar in a realistic urban scenario,” Int. J. Robotics Research, vol. 33, no. 2, pp. 207–214, 2014.
  77. 77.H. Jung, Y. Oto, O. M. Mozos, Y. Iwashita, and R. Kurazume, “Multimodal panoramic 3d outdoor datasets for place categorization,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2016, pp. 4545–4550.
  78. 78.Y. Chen et al., “Lidar-video driving dataset: Learning driving policies effectively,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 5870–5878.
  79. 79.A. Patil, S. Malla, H. Gang, and Y.-T. Chen, “The H3D dataset for full-surround 3D multi-object detection and tracking in crowded urban scenes,” in IEEE Int. Conf. Robotics and Automation, 2019.
  80. 80.X. Jianru et al., “BLVD: Building a large-scale 5D semantics benchmark for autonomous driving,” in IEEE Int. Conf. Robotics and Automation, 2019.
  81. 81.R. Kesten et al. (2019) Lyft level 5 AV dataset 2019. [Online]. Available: https://level5.lyft.com/dataset/
  82. 82.M.-F. Chang et al., “Argoverse: 3D tracking and forecasting with rich maps,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, June 2019.
  83. 83.(2019) PandaSet: Public large-scale dataset for autonomous driving. [Online]. Available: https://scale.com/open-datasets/pandaset
  84. 84.(2019) Waymo open dataset: An autonomous driving dataset. [Online]. Available: https://waymo.com/open
  85. 85.D. Barnes, M. Gadd, P. Murcutt, P. Newman, and I. Posner, “The Oxford radar RobotCar dataset: A radar extension to the Oxford RobotCar dataset,” arXiv preprint arXiv: 1909.01300, 2019. [Online]. Available: https://arxiv.org/pdf/1909.01300
  86. 86.Q.-H. Pham et al., “A*3D Dataset: Towards autonomous driving in challenging environments,” arXiv preprint arXiv: 1909.07541, 2019.
  87. 87.J. Geyer et al. (2019) A2D2: AEV autonomous driving dataset. [Online]. Available: https://www.audi-electronics-venture.de/aev/web/en/driving-dataset.html
  88. 88.M. Braun, S. Krebs, F. B. Flohr, and D. M. Gavrila, “EuroCity Persons: A novel benchmark for person detection in traffic scenes,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–1, 2019.
  89. 89.H. Caesar et al., “nuScenes: A multimodal dataset for autonomous driving,” arXiv preprint arXiv:1903.11027, 2019.
  90. 90.S. Hwang, J. Park, N. Kim, Y. Choi, and I. So Kweon, “Multispectral pedestrian detection: Benchmark dataset and baseline,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2015, pp. 1037–1045.
  91. 91.K. Takumi, K. Watanabe, Q. Ha, A. Tejero-De-Pablos, Y. Ushiku, and T. Harada, “Multispectral object detection for autonomous vehicles,” in Proc. Thematic Workshops of ACM Multimedia, 2017, pp. 35–43.
  92. 92.Q. Ha, K. Watanabe, T. Karasawa, Y. Ushiku, and T. Harada, “MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2017, pp. 5108–5115.
  93. 93.Y. Choi et al., “KAIST multi-spectral day/night data set for autonomous and assisted driving,” IEEE Trans. Intell. Transp. Syst., vol. 19, no. 3, pp. 934–948, 2018.
  94. 94.M. Meyer and G. Kuschk, “Automotive radar dataset for deep learning based 3d object detection,” in Proceedings of the 16th European Radar Conference, 2019.
  95. 95.D. Kondermann et al., “Stereo ground truth with error bars,” in 12th Asian Conf. on Computer Vision. Springer, 2014, pp. 595–610.
  96. 96.M. Larsson, E. Stenborg, L. Hammarstrand, T. Sattler, M. Pollefeys, and F. Kahl, “A cross-season correspondence dataset for robust semantic segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019.
  97. 97.T. Sattler et al., “Benchmarking 6DOF outdoor visual localization in changing conditions,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 8601–8610.
  98. 98.X. Chen, H. Ma, J. Wan, B. Li, and T. Xia, “Multi-view 3D object detection network for autonomous driving,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2017, pp. 6526–6534.
  99. 99.A. Asvadi, L. Garrote, C. Premebida, P. Peixoto, and U. J. Nunes, “Multimodal vehicle detection: fusing 3d-lidar and color camera data,” Pattern Recognition Lett., 2017.
  100. 100.S.-I. Oh and H.-B. Kang, “Object detection and classification by decision-level fusion for intelligent vehicle systems,” Sensors, vol. 17, no. 1, p. 207, 2017.
  101. 101.J. Schlosser, C. K. Chow, and Z. Kira, “Fusing lidar and images for pedestrian detection using convolutional neural networks,” in IEEE Int. Conf. Robotics and Automation, 2016, pp. 2198–2205.
  102. 102.Z. Wang, W. Zhan, and M. Tomizuka, “Fusing bird view lidar point cloud and front view camera image for deep object detection,” in IEEE Intelligent Vehicles Symp., 2018.
  103. 103.J. Ku, M. Mozifian, J. Lee, A. Harakeh, and S. Waslander, “Joint 3d proposal generation and object detection from view aggregation,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, Oct. 2018, pp. 1–8.
  104. 104.D. Xu, D. Anguelov, and A. Jain, “PointFusion: Deep sensor fusion for 3D bounding box estimation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018.
  105. 105.C. R. Qi, W. Liu, C. Wu, H. Su, and L. J. Guibas, “Frustum PointNets for 3d object detection from RGB-D data,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018.
  106. 106.X. Du, M. H. Ang, and D. Rus, “Car detection for autonomous vehicle: Lidar and vision fusion approach through deep learning framework,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2017, pp. 749–754.
  107. 107.X. Du, M. H. Ang Jr., S. Karaman, and D. Rus, “A general pipeline for 3d detection of vehicles,” in IEEE Int. Conf. Robotics and Automation, 2018.
  108. 108.D. Matti, H. K. Ekenel, and J.-P. Thiran, “Combining lidar space clustering and convolutional neural networks for pedestrian detection,” in 14th IEEE Int. Conf. Advanced Video and Signal Based Surveillance, 2017, pp. 1–6.
  109. 109.T. Kim and J. Ghosh, “Robust detection of non-motorized road users using deep learning on optical and lidar data,” in IEEE 19th Int. Conf. Intelligent Transportation Systems, 2016, pp. 271–276.
  110. 110.J. Kim, J. Koh, Y. Kim, J. Choi, Y. Hwang, and J. W. Choi, “Robust deep multi-modal learning based on gated information fusion network,” in Asian Conf. Computer Vision, 2018.
  111. 111.A. Pfeuffer and K. Dietmayer, “Optimal sensor data fusion architecture for object detection in adverse weather conditions,” in Proc. 21st Int. Conf. Information Fusion. IEEE, 2018, pp. 2592–2599.
  112. 112.M. Bijelic, F. Mannan, T. Gruber, W. Ritter, K. Dietmayer, and F. Heide, “Seeing through fog without seeing fog: Deep sensor fusion in the absence of labeled training data,” in Proc. IEEE Conf. Computer Vision, 2019.
  113. 113.V. A. Sindagi, Y. Zhou, and O. Tuzel, “MVX-Net: Multimodal voxelnet for 3D object detection,” in IEEE Int. Conf. Robotics and Automation, 2019.
  114. 114.J. Dou, J. Xue, and J. Fang, “SEG-VoxelNet for 3D vehicle detection from rgb and lidar data,” in IEEE Int. Conf. Robotics and Automation. IEEE, 2019, pp. 4362–4368.
  115. 115.Z. Wang and K. Jia, “Frustum convnet: Sliding frustums to aggregate local point-wise features for amodal 3D object detection,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems. IEEE, 2019.
  116. 116.M. Liang, B. Yang, Y. Chen, R. Hu, and R. Urtasun, “Multi-task multisensor fusion for 3D object detection,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 7345–7353.
  117. 117.J. Wagner, V. Fischer, M. Herman, and S. Behnke, “Multispectral pedestrian detection using deep fusion convolutional neural networks,” in 24th Eur. Symp. Artificial Neural Networks, Computational Intelligence and Machine Learning, 2016, pp. 509–514.
  118. 118.S. W. Jingjing Liu, Shaoting Zhang and D. Metaxas, “Multispectral deep neural networks for pedestrian detection,” in Proc. British Machine Vision Conf., Sep. 2016, pp. 73.1–73.13.
  119. 119.D. Guan, Y. Cao, J. Liang, Y. Cao, and M. Y. Yang, “Fusion of multispectral data through illumination-aware deep neural networks for pedestrian detection,” Information Fusion, vol. 50, pp. 148–157, 2019.
  120. 120.O. Mees, A. Eitel, and W. Burgard, “Choosing smartly: Adaptive multimodal fusion for object detection in changing environments,” in IEEE/RSJ Int. Conf. Intelligent Robots and Systems, 2016, pp. 151–156.
  121. 121.B. Yang, M. Liang, and R. Urtasun, “HDNET: Exploiting HD maps for 3D object detection,” in Proc. 2nd Annu. Conf. Robot Learning, 2018, pp. 146–155.
  122. 122.S. Casas, W. Luo, and R. Urtasun, “IntentNet: Learning to predict intention from raw sensor data,” in Proc. 2nd Annu. Conf. Robot Learning, 2018, pp. 947–956.
  123. 123.D.-K. Kim, D. Maturana, M. Uenoyama, and S. Scherer, “Season-invariant semantic segmentation with a deep multimodal network,” in Field and Service Robotics. Springer, 2018, pp. 255–270.
  124. 124.Y. Sun, W. Zuo, and M. Liu, “RTFNet: Rgb-thermal fusion network for semantic segmentation of urban scenes,” IEEE Robotics and Automation Letters, 2019.
  125. 125.A. Valada, G. L. Oliveira, T. Brox, and W. Burgard, “Deep multi-spectral semantic scene understanding of forested environments using multimodal fusion,” in Int. Symp. Experimental Robotics. Springer, 2016, pp. 465–477.
  126. 126.A. Valada, J. Vertens, A. Dhall, and W. Burgard, “AdapNet: Adaptive semantic segmentation in adverse environmental conditions,” in IEEE Int. Conf. Robotics and Automation, 2017, pp. 4644–4651.
  127. 127.A. Valada, R. Mohan, and W. Burgard, “Self-supervised model adaptation for multimodal semantic segmentation,” Int. J. Computer Vision, 2018.
  128. 128.F. Yang, J. Yang, Z. Jin, and H. Wang, “A fusion model for road detection based on deep learning and fully connected CRF,” in 13th Annu. Conf. System of Systems Engineering. IEEE, 2018, pp. 29–36.
  129. 129.L. Caltagirone, M. Bellone, L. Svensson, and M. Wahde, “Lidar-camera fusion for road detection using fully convolutional neural networks,” Robotics and Autonomous Systems, vol. 111, pp. 125–131, 2019.
  130. 130.X. Lv, Z. Liu, J. Xin, and N. Zheng, “A novel approach for detecting road based on two-stream fusion fully convolutional network,” in IEEE Intelligent Vehicles Symp., 2018, pp. 1464–1469.
  131. 131.F. Wulff, B. Schäufele, O. Sawade, D. Becker, B. Henke, and I. Radusch, “Early fusion of camera and lidar for robust road detection based on U-Net FCN,” in IEEE Intelligent Vehicles Symp., 2018, pp. 1426–1431.
  132. 132.Z. Chen, J. Zhang, and D. Tao, “Progressive lidar adaptation for road detection,” IEEE/CAA Journal of Automatica Sinica, vol. 6, no. 3, pp. 693–702, 2019.
  133. 133.F. Piewak et al., “Boosting lidar-based semantic labeling by cross-modal training data generation,” in Workshop Proc. Eur. Conf. Computer Vision, 2018.
  134. 134.S. Chadwick, W. Maddern, and P. Newman, “Distant vehicle detection using radar and vision,” in IEEE Int. Conf. Robotics and Automation, 2019.
  135. 135.Y. Zhou and O. Tuzel, “VoxelNet: End-to-end learning for point cloud based 3d object detection,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018.
  136. 136.M. Engelcke, D. Rao, D. Z. Wang, C. H. Tong, and I. Posner, “Vote3Deep: Fast object detection in 3d point clouds using efficient convolutional neural networks,” in IEEE Int. Conf. Robotics and Automation, 2017, pp. 1355–1361.
  137. 137.S. Shi, Z. Wang, X. Wang, and H. Li, “Part −A^2 Net: 3d part-aware and aggregation neural network for object detection from point cloud,” arXiv preprint arXiv:1907.03670, 2019.
  138. 138.Y. Yan, Y. Mao, and B. Li, “Second: Sparsely embedded convolutional detection,” Sensors, vol. 18, no. 10, p. 3337, 2018.
  139. 139.C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “PointNet: Deep learning on point sets for 3d classification and segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, Jul. 2017, pp. 77–85.
  140. 140.C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “PointNet++: Deep hierarchical feature learning on point sets in a metric space,” in Advances in Neural Information Processing Systems, 2017, pp. 5099–5108.
  141. 141.K. Shin, Y. P. Kwon, and M. Tomizuka, “RoarNet: A robust 3D object detection based on region approximation refinement,” in IEEE Intelligent Vehicles Symp., 2018.
  142. 142.S. Wang, S. Suo, M. Wei-Chiu, A. Pokrovsky, and R. Urtasun, “Deep parametric continuous convolutional neural networks,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 2589–2597.
  143. 143.Y. Li, R. Bu, M. Sun, W. Wu, X. Di, and B. Chen, “PointCNN: Convolution on χ-transformed points,” in Advances in Neural Information Processing Systems, 2018, pp. 826–836.
  144. 144.A. Asvadi, L. Garrote, C. Premebida, P. Peixoto, and U. J. Nunes, “DepthCN: Vehicle detection using 3d-lidar and ConvNet,” in IEEE 20th Int. Conf. Intelligent Transportation Systems, 2017.
  145. 145.C. Premebida, L. Garrote, A. Asvadi, A. P. Ribeiro, and U. Nunes, “High-resolution lidar-based depth mapping using bilateral filter,” in IEEE 19th Int. Conf. Intelligent Transportation Systems, Nov. 2016, pp. 2469–2474.
  146. 146.A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “PointPillars: Fast encoders for object detection from point clouds,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018.
  147. 147.T. Roddick, A. Kendall, and R. Cipolla, “Orthographic feature transform for monocular 3d object detection,” in Proc. British Machine Vision Conf., 2019.
  148. 148.Y. Wang, W.-L. Chao, D. Garg, B. Hariharan, M. Campbell, and K. Weinberger, “Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019.
  149. 149.Y. You et al., “Pseudo-lidar++: Accurate depth for 3d object detection in autonomous driving,” arXiv preprint arXiv:1906.06310, 2019.
  150. 150.M. Liang, B. Yang, S. Wang, and R. Urtasun, “Deep continuous fusion for multi-sensor 3d object detection,” in Proc. Eur. Conf. Computer Vision, 2018, pp. 641–656.
  151. 151.K. Werber et al., “Automotive radar gridmap representations,” in IEEE MTT-S Int. Conf. Microwaves for Intelligent Mobility, 2015, pp. 1–4.
  152. 152.J. Lombacher, M. Hahn, J. Dickmann, and C. Wöhler, “Potential of radar for static object classification using deep learning methods,” in IEEE MTT-S Int. Conf. Microwaves for Intelligent Mobility, 2016, pp. 1–4.
  153. 153.J. Lombacher, K. Laudt, M. Hahn, J. Dickmann, and C. Wöhler, “Semantic radar grids,” in IEEE Intelligent Vehicles Symp., 2017, pp. 1170–1175.
  154. 154.T. Visentin, A. Sagainov, J. Hasch, and T. Zwick, “Classification of objects in polarimetric radar images using cnns at 77 ghz,” in 2017 IEEE Asia Pacific Microwave Conference (APMC). IEEE, 2017, pp. 356–359.
  155. 155.S. Kim, S. Lee, S. Doo, and B. Shim, “Moving target classification in automotive radar systems using convolutional recurrent neural networks,” in 26th Eur. Signal Processing Conf. IEEE, 2018, pp. 1482–1486.
  156. 156.M. G. Amin and B. Erol, “Understanding deep neural networks performance for radar-based human motion recognition,” in IEEE Radar Conf., 2018, pp. 1461–1465.
  157. 157.O. Schumann, M. Hahn, J. Dickmann, and C. Wöhler, “Semantic segmentation on radar point clouds,” in Proc. 21st Int. Conf. Information Fusion. IEEE, 2018, pp. 2179–2186.
  158. 158.C. Wöhler, O. Schumann, M. Hahn, and J. Dickmann, “Comparison of random forest and long short-term memory network performances in classification tasks using radar,” in Sensor Data Fusion: Trends, Solutions, Applications. IEEE, 2017, pp. 1–6.
  159. 159.R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural Computation, vol. 3, pp. 79–87, 1991.
  160. 160.D. Eigen, M. Ranzato, and I. Sutskever, “Learning factored representations in a deep mixture of experts,” in Workshop Proc. Int. Conf. Learning Representations, 2014.
  161. 161.M. Bloesch, J. Czarnowski, R. Clark, S. Leutenegger, and A. J. Davison, “Codeslam: learning a compact, optimisable representation for dense visual slam,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 2560–2568.
  162. 162.J. Wang, Z. Wei, T. Zhang, and W. Zeng, “Deeply-fused nets,” arXiv:1605.07716 [cs.CV], 2016.
  163. 163.O. Russakovsky et al., “ImageNet large scale visual recognition challenge,” Int. J. Computer Vision, vol. 115, no. 3, pp. 211–252, 2015.
  164. 164.J. Ngiam et al., “Starnet: Targeted computation for object detection in point clouds,” arXiv preprint arXiv:1908.11069, 2019.
  165. 165.A. Gaidon, Q. Wang, Y. Cabon, and E. Vig, “Virtual worlds as proxy for multi-object tracking analysis,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016.
  166. 166.S. R. Richter, V. Vineet, S. Roth, and V. Koltun, “Playing for data: Ground truth from computer games,” in Proc. Eur. Conf. Computer Vision, 2016, pp. 102–118.
  167. 167.G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez, “The SYNTHIA dataset: A large collection of synthetic images for semantic segmentation of urban scenes,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 3234–3243.
  168. 168.S. R. Richter, Z. Hayder, and V. Koltun, “Playing for benchmarks,” in Proc. IEEE Conf. Computer Vision, Oct. 2017, pp. 2232–2241.
  169. 169.X. Yue, B. Wu, S. A. Seshia, K. Keutzer, and A. L. Sangiovanni-Vincentelli, “A lidar point cloud generator: from a virtual world to autonomous driving,” in Proc. ACM Int. Conf. Multimedia Retrieval. ACM, 2018, pp. 458–464.
  170. 170.M. Wrenninge and J. Unger, “Synscapes: A photorealistic synthetic dataset for street scene parsing,” arXiv preprint arXiv:1810.08705, 2018.
  171. 171.P.-H. Huang, K. Matzen, J. Kopf, N. Ahuja, and J.-B. Huang, “Deepmvs: Learning multi-view stereopsis,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018.
  172. 172.D. Griffiths and J. Boehm, “SynthCity: A large scale synthetic point cloud,” in ArXiv preprint, 2019.
  173. 173.A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in Proc. 1st Annu. Conf. Robot Learning, 2017, pp. 1–16.
  174. 174.B. Hurl, K. Czarnecki, and S. Waslander, “Precise synthetic image and lidar (presil) dataset for autonomous vehicle perception,” arXiv preprint arXiv:1905.00160, 2019.
  175. 175.J. Lee, S. Walsh, A. Harakeh, and S. L. Waslander, “Leveraging pretrained 3d object detection models for fast ground truth generation,” in 21st Int. Conf. Intelligent Transportation Systems. IEEE, Nov. 2018, pp. 2504–2510.
  176. 176.J. Mei, B. Gao, D. Xu, W. Yao, X. Zhao, and H. Zhao, “Semantic segmentation of 3d lidar data in dynamic scene using semi-supervised learning,” IEEE Trans. Intell. Transp. Syst., 2018.
  177. 177.R. Mackowiak, P. Lenz, O. Ghori, F. Diego, O. Lange, and C. Rother, “CEREALS – cost-effective region-based active learning for semantic segmentation,” in Proc. British Machine Vision Conf., 2018.
  178. 178.S. Roy, A. Unmesh, and V. P. Namboodiri, “Deep active learning for object detection,” in Proc. British Machine Vision Conf., 2018, p. 91.
  179. 179.D. Feng, X. Wei, L. Rosenbaum, A. Maki, and K. Dietmayer, “Deep active learning for efficient training of a lidar 3d object detector,” in IEEE Intelligent Vehicles Symp., 2019.
  180. 180.S. J. Pan, Q. Yang et al., “A survey on transfer learning,” IEEE Trans. Knowl. Data Eng., vol. 22, no. 10, pp. 1345–1359, 2010.
  181. 181.V. M. Patel, R. Gopalan, R. Li, and R. Chellappa, “Visual domain adaptation: A survey of recent advances,” IEEE Signal Process. Mag., vol. 32, no. 3, pp. 53–69, 2015.
  182. 182.Y. Chen, W. Li, X. Chen, and L. V. Gool, “Learning semantic segmentation from synthetic data: A geometrically guided input-output adaptation approach,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 1841–1850.
  183. 183.Y. Chen, W. Li, C. Sakaridis, D. Dai, and L. Van Gool, “Domain adaptive faster r-cnn for object detection in the wild,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 3339–3348.
  184. 184.K.-H. Lee, G. Ros, J. Li, and A. Gaidon, “Spigan: Privileged adversarial learning from simulation,” in Proc. Int. Conf. Learning Representations, 2019.
  185. 185.J. Tremblay et al., “Training deep networks with synthetic data: Bridging the reality gap by domain randomization,” in Workshop Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 969–977.
  186. 186.D. P. Kingma, S. Mohamed, D. J. Rezende, and M. Welling, “Semisupervised learning with deep generative models,” in Advances in Neural Information Processing Systems, 2014, pp. 3581–3589.
  187. 187.G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks, 2019.
  188. 188.Y. Wang et al., “Iterative learning with open-set noisy labels,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 8688–8696.
  189. 189.M. Ren, W. Zeng, B. Yang, and R. Urtasun, “Learning to reweight examples for robust deep learning,” in Int. Conf. Machine Learning, 2018.
  190. 190.L. Jiang, Z. Zhou, T. Leung, L.-J. Li, and L. Fei-Fei, “MentorNet: Learning data-driven curriculum for very deep neural networks on corrupted labels,” in Int. Conf. Machine Learning, 2018, pp. 2309–2318.
  191. 191.X. Ma et al., “Dimensionality-driven learning with noisy labels,” in Int. Conf. Machine Learning, 2018, pp. 3361–3370.
  192. 192.A. Zlateski, R. Jaroensri, P. Sharma, and F. Durand, “On the importance of label quality for semantic segmentation,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 1479–1487.
  193. 193.P. Meletis and G. Dubbelman, “On boosting semantic street scene segmentation with weak supervision,” in IEEE Intelligent Vehicles Symp., 2019.
  194. 194.C. Haase-Schütz, H. Hertlein, and W. Wiesbeck, “Estimating labeling quality with deep object detectors,” in IEEE Intelligent Vehicles Symp., June 2019, pp. 33–38.
  195. 195.S. Chadwick and P. Newman, “Training object detectors with noisy data,” in IEEE Intelligent Vehicles Symp., June 2019, pp. 1319–1325.
  196. 196.T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” IEEE Trans. Pattern Anal. Mach. Intell., 2018.
  197. 197.M. Giering, V. Venugopalan, and K. Reddy, “Multi-modal sensor registration for vehicle perception via deep neural networks,” in IEEE High Performance Extreme Computing Conf., 2015, pp. 1–6.
  198. 198.N. Schneider, F. Piewak, C. Stiller, and U. Franke, “RegNet: Multimodal sensor registration using deep neural networks,” in IEEE Intelligent Vehicles Symp., 2017, pp. 1803–1810.
  199. 199.S. Ramos, S. Gehrig, P. Pinggera, U. Franke, and C. Rother, “Detecting unexpected obstacles for self-driving cars: Fusing deep learning and geometric modeling,” in IEEE Intelligent Vehicles Symp., 2017, pp. 1025–1032.
  200. 200.H. Banzhaf, M. Dolgov, J. Stellet, and J. M. Zöllner, “From footprints to beliefprints: Motion planning under uncertainty for maneuvering automated vehicles in dense scenarios,” in 21st Int. Conf. Intelligent Transportation Systems. IEEE, 2018, pp. 1680–1687.
  201. 201.D. J. C. MacKay, “A practical Bayesian framework for backpropagation networks,” Neural Computation, vol. 4, no. 3, pp. 448–472, 1992.
  202. 202.G. E. Hinton and D. Van Camp, “Keeping the neural networks simple by minimizing the description length of the weights,” in Proc. 6th Annu. Conf. Computational Learning Theory. ACM, 1993, pp. 5–13.
  203. 203.Y. Gal, “Uncertainty in deep learning,” Ph.D. dissertation, University of Cambridge, 2016.
  204. 204.A. Graves, “Practical variational inference for neural networks,” in Advances in Neural Information Processing Systems, 2011, pp. 2348–2356.
  205. 205.S. Mandt, M. D. Hoffman, and D. M. Blei, “Stochastic gradient descent as approximate Bayesian inference,” J. Machine Learning Research, vol. 18, no. 1, pp. 4873–4907, 2017.
  206. 206.M. Teye, H. Azizpour, and K. Smith, “Bayesian uncertainty estimation for batch normalized deep networks,” in Int. Conf. Machine Learning, 2018.
  207. 207.J. Postels, F. Ferroni, H. Coskun, N. Navab, and F. Tombari, “Sampling-free epistemic uncertainty estimation using approximated variance propagation,” in Proc. IEEE Conf. Computer Vision, 2019.
  208. 208.A. Kendall, V. Badrinarayanan, and R. Cipolla, “Bayesian SegNet: Model uncertainty in deep convolutional encoder-decoder architectures for scene understanding,” in Proc. British Machine Vision Conf., 2017.
  209. 209.D. Miller, L. Nicholson, F. Dayoub, and N. Sünderhauf, “Dropout sampling for robust object detection in open-set conditions,” in IEEE Int. Conf. Robotics and Automation, 2018.
  210. 210.D. Miller, F. Dayoub, M. Milford, and N. Sünderhauf, “Evaluating merging strategies for sampling-based uncertainty techniques in object detection,” in IEEE Int. Conf. Robotics and Automation, 2018.
  211. 211.A. Kendall and Y. Gal, “What uncertainties do we need in Bayesian deep learning for computer vision?” in Advances in Neural Information Processing Systems, 2017, pp. 5574–5584.
  212. 212.E. Ilg et al., “Uncertainty estimates and multi-hypotheses networks for optical flow,” in Proc. Eur. Conf. Computer Vision, 2018.
  213. 213.D. Feng, L. Rosenbaum, and K. Dietmayer, “Towards safe autonomous driving: Capture uncertainty in the deep neural network for lidar 3d vehicle detection,” in 21st Int. Conf. Intelligent Transportation Systems, Nov. 2018, pp. 3266–3273.
  214. 214.D. Feng, L. Rosenbaum, F. Timm, and K. Dietmayer, “Leveraging heteroscedastic aleatoric uncertainties for robust real-time lidar 3d object detection,” in IEEE Intelligent Vehicles Symp., 2019.
  215. 215.G. P. Meyer, A. Laddha, E. Kee, C. Vallespi-Gonzalez, and C. K. Wellington, “Lasernet: An efficient probabilistic 3d object detector for autonomous driving,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 12 677–12 686.
  216. 216.S. Wirges, M. Reith-Braun, M. Lauer, and C. Stiller, “Capturing object detection uncertainty in multi-layer grid maps,” in IEEE Intelligent Vehicles Symp., 2019.
  217. 217.M. T. Le, F. Diehl, T. Brunner, and A. Knol, “Uncertainty estimation for deep neural object detectors in safety-critical applications,” in IEEE 21st Int. Conf. Intelligent Transportation Systems. IEEE, 2018, pp. 3873–3878.
  218. 218.D. Feng, L. Rosenbaum, C. Gläser, F. Timm, and K. Dietmayer, “Can we trust you? on calibration of a probabilistic object detector for autonomous driving,” arXiv:1909.12358 [cs.RO], 2019.
  219. 219.D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in Int. Conf. Learning Representations, 2014.
  220. 220.I. Goodfellow et al., “Generative adversarial nets,” in Advances in Neural Information Processing Systems, 2014, pp. 2672–2680.
  221. 221.A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, “Autoencoding beyond pixels using a learned similarity metric,” in Int. Conf. Machine Learning, 2016, pp. 1558–1566.
  222. 222.A. Deshpande, J. Lu, M.-C. Yeh, M. J. Chong, and D. A. Forsyth, “Learning diverse image colorization,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 2877–2885.
  223. 223.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2017, pp. 5967–5976.
  224. 224.T. A. Wheeler, M. Holder, H. Winner, and M. J. Kochenderfer, “Deep stochastic radar models,” in IEEE Intelligent Vehicles Symp., 2017, pp. 47–53.
  225. 225.X. Han, J. Lu, C. Zhao, S. You, and H. Li, “Semi-supervised and weakly-supervised road detection based on generative adversarial networks,” IEEE Signal Process. Lett., 2018.
  226. 226.J. L. Elman, “Learning and development in neural networks: The importance of starting small,” Cognition, vol. 48, no. 1, pp. 71–99, 1993.
  227. 227.J. Feng and T. Darrell, “Learning the structure of deep convolutional networks,” in Proc. IEEE Conf. Computer Vision, 2015, pp. 2749–2757.
  228. 228.D. Ramachandram, M. Lisicki, T. J. Shields, M. R. Amer, and G. W. Taylor, “Structure optimization for deep multimodal fusion networks using graph-induced kernels,” in 25th Eur. Symp. Artificial Neural Networks, Computational Intelligence and Machine Learning, 2017.
  229. 229.D. Whitley, T. Starkweather, and C. Bogart, “Genetic algorithms and neural networks: Optimizing connections and connectivity,” Parallel computing, vol. 14, no. 3, pp. 347–361, 1990.
  230. 230.B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” arXiv:1611.01578 [cs.LG], 2016.
  231. 231.P. Kulkarni, J. Zepeda, F. Jurie, P. Perez, and L. Chevallier, “Learning the structure of deep architectures using l1 regularization,” in Proc. British Machine Vision Conf., 2015.
  232. 232.C. Murdock, Z. Li, H. Zhou, and T. Duerig, “Blockout: Dynamic model selection for hierarchical deep networks,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, pp. 2583–2591.
  233. 233.F. Li, N. Neverova, C. Wolf, and G. Taylor, “Modout: Learning multimodal architectures by stochastic regularization,” in 12th IEEE Int. Conf. Automatic Face & Gesture Recognition, 2017, pp. 422–429.
  234. 234.A. Bilal, A. Jourabloo, M. Ye, X. Liu, and L. Ren, “Do convolutional neural networks learn class hierarchy?” IEEE Trans. Vis. Comput. Graphics, vol. 24, no. 1, pp. 152–162, 2018.
  235. 235.M. Liu, J. Shi, Z. Li, C. Li, J. Zhu, and S. Liu, “Towards better analysis of deep convolutional neural networks,” IEEE Trans. Vis. Comput. Graphics, vol. 23, no. 1, pp. 91–100, 2017.
  236. 236.S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” in Int. Conf. Learning Representations, 2015.
  237. 237.A. G. Howard et al., “MobileNets: Efficient convolutional neural networks for mobile vision applications,” arXiv:1704.04861 [cs.CV], 2017.
  238. 238.Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “A survey of model compression and acceleration for deep neural networks,” IEEE Signal Process. Mag., 2017.
  239. 239.L. Enderich, F. Timm, L. Rosenbaum, and W. Burgard, “Learning multimodal fixed-point weights using gradient descent,” in 27th Eur. Symp. Artificial Neural Networks, Computational Intelligence and Machine Learning, 2019.
  240. 240.M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The PASCAL visual object classes (VOC) challenge,” Int. J. Computer Vision, vol. 88, no. 2, pp. 303–338, 2010.
  241. 241.A. P. Dawid, “The well-calibrated Bayesian,” J. American Statistical Association, vol. 77, no. 379, pp. 605–610, 1982.
  242. 242.S. Liang, Y. Li, and R. Srikant, “Enhancing the reliability of out-of-distribution image detection in neural networks,” in Int. Conf. Learning Representations, 2017.
  243. 243.D. Hall, F. Dayoub, J. Skinner, P. Corke, G. Carneiro, and N. Sünderhauf, “Probability-based detection quality (PDQ): A probabilistic approach to detection evaluation,” arXiv:1811.10800 [cs.CV], 2018.
  244. 244.W. Luo, B. Yang, and R. Urtasun, “Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 3569–3577.
  245. 245.M. Bojarski et al., “End to end learning for self-driving cars,” arXiv:1604.07316 [cs.CV], 2016.
  246. 246.G.-H. Liu, A. Siravuru, S. Prabhakar, M. Veloso, and G. Kantor, “Learning end-to-end multimodal sensor policies for autonomous navigation,” in Proc. 1st Annu. Conf. Robot Learning, 2017, pp. 249–261.
  247. 247.M. Bansal, A. Krizhevsky, and A. Ogale, “Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst,” in Proc. Robotics: Science and Systems, 2019.
  248. 248.A. Sauer, N. Savinov, and A. Geiger, “Conditional affordance learning for driving in urban environments,” in Proc. 2st Annu. Conf. Robot Learning, 2018, pp. 237–252.
  249. 249.C. Chen, A. Seff, A. Kornhauser, and J. Xiao, “DeepDriving: Learning affordance for direct perception in autonomous driving,” in Proc. IEEE Conf. Computer Vision, 2015, pp. 2722–2730.
  250. 250.B. Yang, W. Luo, and R. Urtasun, “PIXOR: Real-time 3d object detection from point clouds,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, pp. 7652–7660.
  251. 251.O. Erkent, C. Wolf, C. Laugier, D. S. González, and V. R. Cano, “Semantic grid estimation with a hybrid bayesian and deep neural network approach,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2018, pp. 888–895.
  252. 252.S. Gu, T. Lu, Y. Zhang, J. Alvarez, J. Yang, and H. Kong, “3D lidar+ monocular camera: an inverse-depth induced fusion framework for urban road detection,” IEEE Transactions on Intelligent Vehicles, 2018.
  253. 253.Y. Cai, D. Li, X. Zhou, and X. Mou, “Robust drivable road region detection for fixed-route autonomous vehicles using map-fusion images,” Sensors, vol. 18, no. 12, p. 4158, 2018.

Citation

MLA
Feng, D., et al. “Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges”. IEEE Transactions on Intelligent Transportation Systems, vol. 22, no. 3, 2021, pp. 1341–60, https://doi.org/10.1109/TITS.2020.2972974.
APA
Feng, D., Haase-Schutz, C., Rosenbaum, L., Hertlein, H., Glaser, C., Timm, F., Wiesbeck, W., & Dietmayer, K. (2021). Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges. IEEE Transactions on Intelligent Transportation Systems, 22(3), 1341–1360. https://doi.org/10.1109/TITS.2020.2972974
Chicago
Feng, D., C. Haase-Schutz, L. Rosenbaum, et al. 2021. “Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges”. IEEE Transactions on Intelligent Transportation Systems 22 (3): 1341–60. https://doi.org/10.1109/TITS.2020.2972974.
Harvard
Feng, D. et al. (2021) “Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges”, IEEE Transactions on Intelligent Transportation Systems, 22(3), pp. 1341–1360. Available at: https://doi.org/10.1109/TITS.2020.2972974.
Vancouver
1. Feng D, Haase-Schutz C, Rosenbaum L, Hertlein H, Glaser C, Timm F, Wiesbeck W, Dietmayer K (2021) Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges. IEEE Transactions on Intelligent Transportation Systems 22:1341–1360

BibTeX

@article{Feng_2021, title={Deep Multi-Modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges}, volume={22}, ISSN={1558-0016}, url={http://dx.doi.org/10.1109/TITS.2020.2972974}, DOI={10.1109/tits.2020.2972974}, number={3}, journal={IEEE Transactions on Intelligent Transportation Systems}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Feng, Di and Haase-Schutz, Christian and Rosenbaum, Lars and Hertlein, Heinz and Glaser, Claudius and Timm, Fabian and Wiesbeck, Werner and Dietmayer, Klaus}, year={2021}, month=Mar, pages={1341–1360} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF