SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences

Jens BehleyMartin GarbadeAndres MiliotoJan QuenzelSven BehnkeCyrill StachnissJuergen Gall

article2019ICCV2,526 citations

Introduces SemanticKITTI, a large-scale dataset providing dense, point-wise 360-degree semantic annotations for LiDAR sequences from the KITTI benchmark alongside standard baseline tasks for single-scan segmentation, multi-scan segmentation, and semantic scene completion in autonomous driving.

Listen

Self-driving cars require detailed understanding of their surroundings to navigate safely, but existing datasets for semantic segmentation of LiDAR point clouds remain small and lack sequential information from automotive sensors. This gap hinders development of methods that can handle real-world driving scenarios, including changes in the environment and unmapped areas.

The article introduces SemanticKITTI, a large annotated dataset derived from the KITTI Odometry Benchmark, to support three tasks: semantic segmentation from single scans, segmentation from multiple past scans, and semantic scene completion that predicts future scene structure.

Researchers annotated over 43,000 scans across 22 sequences with 28 classes using a custom labeling tool that aggregates scans via SLAM for consistency. They evaluated multiple state-of-the-art point cloud segmentation methods as baselines on the training and test splits.

The best single-scan method achieved 49.9% mean intersection-over-union across 19 classes, with performance declining sharply at greater distances due to sparsity. Multi-scan approaches showed limited gains in distinguishing moving from static objects. Semantic scene completion reached 50.6% completion IoU and 17.7% semantic mIoU only when using a higher-resolution backbone.

These results indicate that current models lack sufficient capacity and mechanisms to exploit temporal data or handle sparse distant points effectively, which limits reliable perception for autonomous driving. The dataset enables reproducible progress and new directions such as semantic SLAM.

Future work should develop architectures that process sequential inputs more explicitly and produce higher-resolution outputs. Instance-level annotations over time are planned to support additional tasks.

The findings rest on a single sensor type and location, with class imbalance and limited test evaluations that may affect generalization; results should be validated on diverse environments before deployment decisions.

Cover for SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences

Abstract

Semantic scene understanding is important for various applications. In particular, self-driving cars need a fine-grained understanding of the surfaces and objects in their vicinity. Light detection and ranging (LiDAR) provides precise geometric information about the environment and is thus a part of the sensor suites of almost all self-driving cars. Despite the relevance of semantic scene understanding for this application, there is a lack of a large dataset for this task which is based on an automotive LiDAR.

In this paper, we introduce a large dataset to propel research on laser-based semantic segmentation. We annotated all sequences of the KITTI Vision Odometry Benchmark and provide dense point-wise annotations for the complete 360o360^{o} field-of-view of the employed automotive LiDAR. We propose three benchmark tasks based on this dataset: (i) semantic segmentation of point clouds using a single scan, (ii) semantic segmentation using multiple past scans, and (iii) semantic scene completion, which requires to anticipate the semantic scene in the future. We provide baseline experiments and show that there is a need for more sophisticated models to efficiently tackle these tasks. Our dataset opens the door for the development of more advanced methods, but also provides plentiful data to investigate new research directions.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The SemanticKITTI Dataset
  • 3.1 Labeling Process
  • 3.2 Dataset Statistics
  • 4 Evaluation of Semantic Segmentation
  • 4.1 Single Scan Experiments
  • 4.2 Multiple Scan Experiments
  • 5 Evaluation of Semantic Scene Completion
  • 6 Conclusion and Outlook
  • References
  • A Consistent Labels for LiDAR Sequences
  • B Basis of the Dataset
  • C Class Definition
  • D Baseline Setup
  • E Results using Multiple Scans
  • F Semantic Scene Completion
  • G Qualitative Results
  • H Dataset and Baseline Access API

Knowls

  1. Knowl 1 — SemanticKITTI Dataset Specification and Annotation Schema

    experimental setup

    SemanticKITTI is a large-scale dataset providing dense, point-wise semantic annotations for all 22 sequences of the KITTI Vision Odometry Benchmark. The data is acquired using a roof-mounted Velodyne HDL-64E rotating LiDAR providing a 360-degree horizontal field-of-view. The dataset contains 43,552 individual scans comprising over 4.549 billion points, split into 23,201 scans across sequences 00 to 10 for training and validation (with sequence 08 designated as the standard validation split) and 20,351 scans across sequences 11 to 21 for testing.

    The annotation taxonomy consists of 28 fine-grained classes:

    • Ground-related classes: road, sidewalk, parking, and other-ground.
    • Structure classes: building and other-structure.
    • Vehicle classes: car, truck, bicycle, motorcycle, and other-vehicle.
    • Nature classes: vegetation, trunk, and terrain.
    • Human classes: person, bicyclist, and motorcyclist.
    • Object classes: fence, pole, traffic sign, and other-object.
    • Noise class: outlier (spurious points from reflections or deskewing).

    For 6 movable categories (car, truck, other-vehicle, person, bicyclist, motorcyclist), moving and non-moving states are distinguished; an entity receives the moving tag if it changed position at any point during observation. Bicyclists and motorcyclists encompass both the human rider and the vehicle within arm's reach due to the inability of single LiDAR scans to separate the rider from the frame.

  2. Knowl 2 — Single-Scan LiDAR Semantic Segmentation Benchmark Performance

    data/table

    The single-scan semantic segmentation benchmark evaluates models on the task of predicting point-wise labels for individual 360360^\circ LiDAR sweeps across 19 classes (moving and non-moving variants merged; outlier, other-structure, and other-object excluded). Models are trained on sequences 00 to 10 (excluding validation sequence 08) and evaluated on test sequences 11 to 21.

    Approach mIoU road side. park. oth-g. bldg. car truck bicy. moto. oth-v. veg. trunk terr. pers. bic-r. mot-r. fenc. pole traf.
    PointNet 14.6 61.6 35.7 15.8 1.4 41.4 46.3 0.1 1.3 0.3 0.8 31.0 4.6 17.6 0.2 0.2 0.0 12.9 2.4 3.7
    SPGraph 17.4 45.0 28.5 0.6 0.6 64.3 49.3 0.1 0.2 0.2 0.8 48.9 27.2 24.6 0.3 2.7 0.1 20.8 15.9 0.8
    SPLATNet 18.4 64.6 39.1 0.4 0.0 58.3 58.2 0.0 0.0 0.0 0.0 71.1 9.9 19.3 0.0 0.0 0.0 23.1 5.6 0.0
    PointNet++ 20.1 72.0 41.8 18.7 5.6 62.3 53.7 0.9 1.9 0.2 0.2 46.5 13.8 30.0 0.9 1.0 0.0 16.9 6.0 8.9
    SqueezeSeg 29.5 85.4 54.3 26.9 4.5 57.4 68.8 3.3 16.0 4.1 3.6 60.0 24.3 53.7 12.9 13.1 0.9 29.0 17.5 24.5
    SqueezeSegV2 39.7 88.6 67.6 45.8 17.7 73.7 81.8 13.4 18.5 17.9 14.0 71.8 35.8 60.2 20.1 25.1 3.9 41.1 20.2 36.3
    TangentConv 40.9 83.9 63.9 33.4 15.4 83.4 90.8 15.2 2.7 16.5 12.1 79.5 49.3 58.1 23.0 28.4 8.1 49.0 35.8 28.5
    DarkNet21Seg 47.4 91.4 74.0 57.0 26.4 81.9 85.4 18.6 26.2 26.5 15.6 77.6 48.4 63.6 31.8 33.6 4.0 52.3 36.0 50.0
    DarkNet53Seg 49.9 91.8 74.6 64.8 27.9 84.1 86.4 25.5 24.5 32.7 22.6 78.3 50.1 64.0 36.2 33.6 4.7 55.0 38.9 52.2

    Spherical range projection approaches paired with high-capacity 2D backbones (DarkNet53Seg at 49.9% mIoU) substantially outperform direct point-based networks (PointNet at 14.6%, PointNet++ at 20.1%) and sparse lattice convolutions (SPLATNet at 18.4%), while Tangent Convolutions directly on surfaces reaches 40.9% mIoU.

  3. Knowl 3 — Mean Intersection-over-Union Metric for Point-Wise LiDAR Segmentation

    equation

    LiDAR point cloud semantic segmentation accuracy across multiple classes is quantified using the mean Jaccard Index, also known as the mean intersection-over-union (mIoU):

    mIoU=1Cc=1CTPcTPc+FPc+FNc\text{mIoU} = \frac{1}{C} \sum_{c=1}^{C} \frac{\text{TP}_c}{\text{TP}_c + \text{FP}_c + \text{FN}_c}

    where:

    • CNC \in \mathbb{N} denotes the total number of evaluated semantic categories (C=19C = 19 for single-scan benchmarks and semantic scene completion; C=25C = 25 for sequential multi-scan segmentation with moving state separation).
    • TPcN0\text{TP}_c \in \mathbb{N}_0 is the count of true positive point predictions for class cc.
    • FPcN0\text{FP}_c \in \mathbb{N}_0 is the count of false positive point predictions for class cc.
    • FNcN0\text{FN}_c \in \mathbb{N}_0 is the count of false negative point predictions for class cc.

    In standard benchmark evaluations, points labeled as outlier, other-structure, and other-object are ignored during both training and evaluation due to high intra-class variance or measurement noise.

  4. Knowl 4 — Real-World LiDAR Semantic Scene Completion Dataset Generation

    model/method

    Semantic scene completion ground truth for outdoor automotive LiDAR is constructed by superimposing future point cloud scans along the vehicle's SLAM trajectory. As the sensor vehicle moves through the environment, subsequent scans capture previously occluded regions and object backsides.

    The volume of interest is defined in the sensor frame as:

    • Length: 51.2 m51.2\text{ m} forward along the driving direction.
    • Width: 25.6 m25.6\text{ m} to each side (51.2 m51.2\text{ m} total lateral width).
    • Height: 6.4 m6.4\text{ m} vertical span.
    • Grid resolution: Uniform isotropic voxels of 0.2 m×0.2 m×0.2 m0.2\text{ m} \times 0.2\text{ m} \times 0.2\text{ m}, yielding a target grid of 256×256×32256 \times 256 \times 32 voxels.

    Voxel semantic labels are assigned via majority vote across all aggregated labeled 3D points within each voxel cell; voxels containing no points are designated as empty.

    To handle occluded space, ray-tracing is conducted from every sensor pose along the vehicle path. Voxels that are never intersected by any sensor ray across all poses (e.g., solid interior volume of buildings or underground regions) are marked as permanently unobserved and excluded from training loss and benchmark evaluation.

    The resulting dataset contains 19,130 training pairs, 815 validation pairs, and 3,992 testing pairs of raw single-scan voxel inputs and completed dense voxel targets.

  5. Knowl 5 — Semantic Scene Completion Benchmark Evaluation

    data/table

    The 3D semantic scene completion benchmark evaluates the capacity of neural networks to infer occupied geometry (binary completion) and predict semantic labels across 19 classes from an incomplete single-scan voxel input.

    Approach Precision (%) Recall (%) Completion IoU (%) Semantic Scene Completion mIoU (%)
    SSCNet 31.71 83.40 29.83 9.53
    TS3D 31.58 84.18 29.81 9.54
    TS3D + DarkNet53Seg 25.85 88.25 24.99 10.19
    TS3D + DarkNet53Seg + SATNet 80.52 57.65 50.60 17.70

    Approaches based on SSCNet perform a 4-fold spatial downsampling in the forward pass, limiting resolution on fine structures. Exchanging the backbone with SATNet preserves native voxel output resolution via subvolume division (90×138×3290 \times 138 \times 32 overlapping inference blocks fused together), yielding a +20.77% gain in completion IoU and +7.51% gain in semantic completion mIoU.

  6. Knowl 6 — Multi-Scan Sequential Semantic Segmentation and Motion Disambiguation

    data/table

    Exploiting temporal information by combining 5 consecutive raw LiDAR scans (current scan at timestamp tt and 4 preceding scans at t1,t2,t3,t4t-1, t-2, t-3, t-4) allows models to distinguish moving from non-moving objects across 25 total classes.

    Approach State car truck other-vehicle person bicyclist motorcyclist mIoU
    TangentConv Non-Moving 84.9 21.1 18.5 1.6 0.0 0.0 34.1
    Moving 40.3 42.2 30.1 6.4 1.1 1.9
    DarkNet53Seg Non-Moving 84.1 20.0 20.7 7.5 0.0 0.0 41.6
    Moving 61.5 37.8 28.9 15.2 14.1 0.2

    While DarkNet53Seg reaches 41.6% overall mIoU, simply concatenating past raw point clouds into a single aggregated input creates severe difficulty in correctly classifying sparse non-moving instances: non-moving bicyclist and motorcyclist achieve 0.0% IoU for both networks due to point sparsity and temporal blurring.

  7. Knowl 7 — Tile-Based Annotation Pipeline for Sequential LiDAR Scans

    model/method

    To achieve temporally consistent point-wise labels over driving sequences with loop closures, the dataset utilizes a tile-based spatial subdivision workflow:

    1. Sequence Registration: Scans are globally registered and loop-closed using surfel-based LiDAR SLAM, with manual loop closure constraints applied where inertial navigation drift causes height discrepancies.
    2. Spatial Tiling: Global point clouds are partitioned into spatial tiles of 100 m×100 m100\text{ m} \times 100\text{ m}. Each tile aggregates all points from any scan that intersects its spatial boundary, along with an overlapping boundary margin shared with adjacent tiles to enforce cross-tile label continuity.
    3. GPU-Accelerated Labeling: Points are projected to screen coordinates in an interactive OpenGL viewer supporting over 20 million points. Annotators apply polygon and brush selection tools to projectively assign labels, followed by label-hiding filters that isolate unlabeled geometry (e.g., road beneath vehicles) and prevent accidental overwriting.
    4. Hierarchical Pass: Moving objects are labeled and hidden scan-by-scan or via height filtering before static infrastructure and nature categories are annotated across the combined tile.

    The overall dataset required over 1,700 hours of annotation and verification across 518 spatial tiles.

  8. Knowl 8 — Spherical Range Image Projection for Rotating LiDAR Semantic Segmentation

    model/method

    To apply 2D fully convolutional network architectures to unstructured 3D LiDAR point clouds, single-turn scans from rotating sensors are mapped onto a spherical range image grid.

    For the 64-beam Velodyne HDL-64E:

    • Vertical resolution: 64 rows, matching the sensor's individual laser diodes.
    • Horizontal resolution: 2048 columns, covering the complete 360360^\circ azimuth.
    • Point mapping: Each 3D point p=(x,y,z)\mathbf{p} = (x, y, z) is mapped to image coordinates (u,v)(u, v) via spherical coordinates (azimuth and elevation angles).
    • Duplicate resolution: If multiple points map to the same (u,v)(u, v) grid cell, the point with the closest range value r=x2+y2+z2r = \sqrt{x^2 + y^2 + z^2} is retained in the image.
    • Feature channels: Each pixel encodes range rr, 3D coordinates (x,y,z)(x, y, z), and sensor remission intensity.
    • In DarkNet-based segmentation backbones (DarkNet21Seg and DarkNet53Seg), vertical downsampling operations are removed to preserve the 64-row resolution, and 2D pixel-level predictions are re-mapped back to original 3D point indices during inference.
  9. Knowl 9 — Computational and Memory Profile of LiDAR Segmentation Architectures

    data/table

    The computational complexity, model capacity, training duration, and per-scan inference latency for point cloud semantic segmentation baselines on SemanticKITTI are summarized below:

    Approach Model Parameters (×106\times 10^6) Training Time (GPU hours / epoch) Inference Time (seconds / point cloud)
    PointNet 3.0 4 0.5
    PointNet++ 6.0 16 5.9
    SPGraph 0.25 6 5.2
    TangentConv 0.4 6 3.0
    SPLATNet 0.8 8 1.0
    SqueezeSeg 1.0 0.5 0.015
    SqueezeSegV2 1.0 0.6 0.020
    DarkNet21Seg 25.0 2 0.055
    DarkNet53Seg 50.0 3 0.100

    Spherical projection networks (SqueezeSeg, SqueezeSegV2, DarkNet21Seg, DarkNet53Seg) operate orders of magnitude faster (15 ms to 100 ms per scan) than direct point-based (PointNet++ at 5.9 s) or graph-based (SPGraph at 5.2 s) approaches. Direct point networks require aggressive downsampling (to 50,000 points per scan) due to GPU memory limitations, whereas range projection networks natively process complete scans.

  10. Knowl 10 — Distance-Dependent Sparsity Effect on Point Cloud Segmentation

    empirical result

    Semantic segmentation accuracy across all evaluated baseline architectures exhibits a monotonic decline as the Euclidean distance from the LiDAR sensor increases from 10 m10\text{ m} to 50 m50\text{ m}.

    • Cause: The physical divergence of laser beams in rotating LiDAR causes point density to decrease quadratically with range, leaving distant objects with very few sample points.
    • Projection Networks: Range-projection architectures (DarkNet53Seg, DarkNet21Seg, SqueezeSegV2) achieve the highest absolute mIoU in the near field (10 m10\text{ m} to 25 m25\text{ m}) due to dense range image neighborhoods, but degrade steeply as range increases.
    • Graph Networks: Superpoint Graph (SPGraph) exhibits a substantially flatter degradation curve across distance compared to pure projection or point-based methods, indicating that geometric superpoint clustering offers greater robustness to distance-induced point sparsity.

Coverage note — None was omitted; all primary contributions, benchmark definitions, quantitative results across single-scan, multi-scan, and semantic scene completion tasks, and dataset generation pipelines are represented.

References

  1. 1.Anuraag Agrawal, Atsushi Nakazawa, and Haruo Takemura. MMM-classification of 3D Range Data. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2009.
  2. 2.Dragomir Anguelov, Ben Taskar, Vassil Chatalbashev, Daphne Koller, Dinkar Gupta, Geremy Heitz, and Andrew Ng. Discriminative Learning of Markov Random Fields for Segmentation of 3D Scan Data. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 169–176, 2005.
  3. 3.Iro Armeni, Alexander Sax, Amir R. Zamir, and Silvio Savarese. Joint 2D-3D-Semantic Data for Indoor Scene Understanding. arXiv preprint, 2017.
  4. 4.Jens Behley, Kristian Kersting, Dirk Schulz, Volker Steinhage, and Armin B. Cremers. Learning to Hash Logistic Regression for Fast 3D Scan Point Classification. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), pages 5960–5965, 2010.
  5. 5.Jens Behley and Cyrill Stachniss. Efficient Surfel-Based SLAM using 3D Laser Range Data in Urban Environments. In Proc. of Robotics: Science and Systems (RSS), 2018.
  6. 6.Jens Behley, Volker Steinhage, and Armin B. Cremers. Performance of Histogram Descriptors for the Classification of 3D Laser Range Data in Urban Environments. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2012.
  7. 7.Alexandre Boulch, Joris Guerry, Bertrand Le Saux, and Nicolas Audebert. SnapNet: 3D point cloud semantic labeling with 2D deep segmentation networks. Computers & Graphics, 2017.
  8. 8.Angel X. Chang, Thomas Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. ShapeNet: An Information-Rich 3D Model Repository. Technical Report arXiv:1512.03012 [cs.GR], Stanford University and Princeton University and Toyota Technological Institute at Chicago, 2015.
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. DeepLab: Semantic Image Segmentation withDeep Convolutional Nets, Atrous Convolution,and Fully Connected CRFs. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 40(4):834–848, 2018.
  10. 10.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016.
  11. 11.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2009.
  12. 12.Angela Dai, Daniel Ritchie, Martin Bokeloh, Scott Reed, Jurgen Sturm, and Matthias Nießner. ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  13. 13.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2009.
  14. 14.Francis Engelmann, Theodora Kontogianni, Jonas Schult, and Bastian Leibe. Know What Your Neighbors Do: 3D Semantic Segmentation of Point Clouds. arXiv preprint, 2018.
  15. 15.Mark Everingham, S.M. Ali Eslami, Luc van Gool, Christopher K.I. Williams, John Winn, and Andrew Zisserman. The Pascal Visual Object Classes Challenge a Retrospective. International Journal on Computer Vision (IJCV), 111(1):98–136, 2015.
  16. 16.Michael Firman, Oisin Mac Aodha, Simon Julier, and Gabriel J. Brostow. Structured Prediction of Unobserved Voxels From a Single Depth Image. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 5431–5440, 2016.
  17. 17.Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. Virtual Worlds as Proxy for Multi-Object Tracking Analysis. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016.
  18. 18.Martin Garbade, Yueh-Tung Chen, J. Sawatzky, and Juergen Gall. Two Stream 3D Semantic Scene Completion. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019.
  19. 19.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 3354–3361, 2012.
  20. 20.Andres Geiger and Chaohui Wang. Joint 3d Object and Layout Inference from a single RGB-D Image. In Proc. of the German Conf. on Pattern Recognition (GCPR), pages 183–195, 2015.
  21. 21.Benjamin Graham, Martin Engelcke, and Laurens van der Maaten. 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  22. 22.Fabian Groh, Patrick Wieschollek, and Hendrik Lensch. Flex-Convolution (Million-Scale Pointcloud Learning Beyond Grid-Worlds). In Proc. of the Asian Conf. on Computer Vision (ACCV), Dezember 2018.
  23. 23.Timo Hackel, Nikolay Savinov, Lubor Ladicky, Jan D. Wegner, Konrad Schindler, and Marc Pollefeys. SEMANTIC3D.NET: A new large-scale point cloud classification benchmark. In ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, volume IV-1-W1, pages 91–98, 2017.
  24. 24.Binh-Son Hua, Quang-Hieu Pham, Duc Thanh Nguyen, Minh-Khoi Tran, Lap-Fai Yu, and Sai-Kit Yeung. SceneNN: A Scene Meshes Dataset with aNNotations. In Proc. of the Intl. Conf. on 3D Vision (3DV), 2016.
  25. 25.Binh-Son Hua, Minh-Khoi Tran, and Sai-Kit Yeung. Point-wise Convolutional Neural Networks. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  26. 26.Jing Huang and Suya You. Point Cloud Labeling using 3D Convolutional Neural Network. In Proc. of the Intl. Conf. on Pattern Recognition (ICPR), 2016.
  27. 27.Varun Jampani, Martin Kiefel, and Peter V. Gehler. Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs and Bilateral Neural Networks. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016.
  28. 28.Mingyang Jiang, Yiran Wu, and Cewu Lu. PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation. arXiv preprint, 2018.
  29. 29.Andrew E. Johnson and Martial Hebert. Using spin images for effcient object recognition in cluttered 3D scenes. Trans. on Pattern Analysis and Machine Intelligence (TPAMI), 21(5):433–449, 1999.
  30. 30.Roman Klukov and Victor Lempitsky. Escape from Cells: Deep Kd-Networks for the Recognition of 3D Point Cloud Models. In Proc. of the IEEE Intl. Conf. on Computer Vision (ICCV), 2017.
  31. 31.Loic Landrieu and Martin Simonovsky. Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  32. 32.Wenbin Li, Sajad Saeedi, John McCormac, Ronald Clark, Dimos Tzoumanikas, Qing Ye, Yuzhong Huang, Rui Tang, and Stefan Leutenegger. InteriorNet: Mega-scale Multi-sensor Photo-realistic Indoor Scenes Dataset. In Proc. of the British Machine Vision Conference (BMVC), 2018.
  33. 33.Shice Liu, Yu Hu, Yiming Zeng, Qiankun Tang, Beibei Jin, Yainhe Han, and Xiaowei Li. See and Think: Disentangling Semantic Scene Completion. In Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), pages 261–272, 2018.
  34. 34.Daniel Maturana and Sebastian Scherer. VoxNet: A 3D Convolutional Neural Network for Real-Time Object Recognition. In Proc. of the IEEE/RSJ Intl. Conf. on Intelligent Robots and Systems (IROS), 2015.
  35. 35.John McCormac, Ankur Handa, Stefan Leutenegger, and Andrew J. Davison. SceneNet RGB-D: Can 5M Synthetic Images Beat Generic ImageNet Pre-training on Indoor Segmentation? In Proc. of the IEEE Intl. Conf. on Computer Vision (ICCV), 2017.
  36. 36.Daniel Munoz, J. Andrew Bagnell, Nicolas Vandapel, and Martial Hebert. Contextual Classification with Functional Max-Margin Markov Networks. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2009.
  37. 37.Daniel Munoz, Nicholas Vandapel, and Marial Hebert. Directional Associative Markov Network for 3-D Point Cloud Classification. In Proc. of the International Symposium on 3D Data Processing, Visualization and Transmission (3DPVT), pages 63–70, 2008.
  38. 38.Daniel Munoz, Nicholas Vandapel, and Martial Hebert. On-board Contextual Classification of 3-D Point Clouds with Learned High-order Markov Random Fields. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2009.
  39. 39.Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder. The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes. In Proc. of the IEEE Intl. Conf. on Computer Vision (ICCV), 2017.
  40. 40.Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017.
  41. 41.Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. In Proc. of the Conf. on Neural Information Processing Systems (NeurIPS), 2017.
  42. 42.Joseph Redmon and Ali Farhadi. YOLOv3: An Incremental Improvement. arXiv preprint, 2018.
  43. 43.Dario Rethage, Johanna Wald, Jurgen Sturm, Nassir Navab, and Frederico Tombari. Fully-Convolutional Point Networks for Large-Scale Point Clouds. Proc. of the European Conf. on Computer Vision (ECCV), 2018.
  44. 44.Gernot Riegler, Ali Osman Ulusoy, and Andreas Geiger. OctNet: Learning Deep 3D Representations at High Resolutions. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017.
  45. 45.Jason Rock, Tanmay Gupta, Justin Thorsen, JunYoung Gwak, Daeyun Shin, and Derek Hoiem. Completing 3D Object Shape from One Depth Image. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2015.
  46. 46.German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio Lopez. The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 2016.
  47. 47.Xavier Roynard, Jean-Emmanuel Deschaud, and Francois Goulette. Paris-Lille-3D: A large and high-quality ground-truth urban point cloud dataset for automatic segmentation and classification. Intl. Journal of Robotics Research (IJRR), 37(6):545–557, 2018.
  48. 48.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor Segmentation and Support Inference from RGBD Images. In Proc. of the European Conf. on Computer Vision (ECCV), 2012.
  49. 49.Shuran Song, Fisher Yu, Andy Zeng, Angel X. Chang, Manolis Savva, and Thomas Funkhouser. Semantic Scene Completion from a Single Depth Image. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017.
  50. 50.Bastian Steder, Giorgio Grisetti, and Wolfram Burgard. Robust Place Recognition for 3D Range Data based on Point Features. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2010.
  51. 51.Hang Su, Varun Jampani, Deqing Sun, Subhransu Maji, Evangelos Kalogerakis, Ming-Hsuan Yang, and Jan Kautz. SPLATNet: Sparse Lattice Networks for Point Cloud Processing. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  52. 52.Maxim Tatarchenko, Jaesik Park, Vladen Koltun, and Qian-Yi Zhou. Tangent Convolutions for Dense Prediction in 3D. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  53. 53.Lyne P. Tchapmi, Christopher B. Choy, Iro Armeni, Jun Young Gwak, and Silvio Savarese. SEGCloud: Semantic Segmentation of 3D Point Clouds. In Proc. of the Intl. Conf. on 3D Vision (3DV), 2017.
  54. 54.Gusi Te, Wei Hu, Zongming Guo, and Amin Zheng. RGCNN: Regularized Graph CNN for Point Cloud Segmentation. arXiv preprint, 2018.
  55. 55.Antonio Torralba and Alexei A. Efros. Unbiased Look at Dataset Bias. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2011.
  56. 56.Rudolph Triebel, Krisitian Kersting, and Wolfram Burgard. Robust 3D Scan Point Classification using Associative Markov Networks. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), pages 2603–2608, 2006.
  57. 57.Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, and Raquel Urtasun. Deep Parametric Continuous Convolutional Neural Networks. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  58. 58.Yida Wang, Davod Tan Joseph, Nassir Navab, and Frederico Tombari. Adversarial Semantic Scene Completion from a Single Depth Image. In Proc. of the Intl. Conf. on 3D Vision (3DV), pages 426–434, 2018.
  59. 59.Zongji Wang and Feng Lu. VoxSegNet: Volumetric CNNs for Semantic Part Segmentation of 3D Shapes. arXiv preprint, 2018.
  60. 60.Bichen Wu, Alvin Wan, Xiangyu Yue, and Kurt Keutzer. SqueezeSeg: Convolutional Neural Nets with Recurrent CRF for Real-Time Road-Object Segmentation from 3D LiDAR Point Cloud. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2018.
  61. 61.Bichen Wu, Xuanyu Zhou, Sicheng Zhao, Xiangyu Yue, and Kurt Keutzer. SqueezeSegV2: Improved Model Structure and Unsupervised Domain Adaptation for Road-Object Segmentation from a LiDAR Point Cloud. Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2019.
  62. 62.Jun Xie, Martin Kiefel, Ming-Ting Sun, and Andreas Geiger. Semantic Instance Annotation of Street Scenes by 3D to 2D Label Transfer. In Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016.
  63. 63.Xuehan Xiong, Daniel Munoz, J. Andrew Bagnell, and Martial Hebert. 3-D Scene Analysis via Sequenced Predictions over Points and Regions. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), pages 2609–2616, 2011.
  64. 64.Wei Zeng and Theo Gevers. 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues. arXiv preprint, 2017.
  65. 65.Jiahui Zhang, Hao Zhao, Anbang Yao, Yurong Chen, Li Zhang, and Hongen Liao. Efficient Semantic Scene Completion Network with Spatial Group Convolution. In Proc. of the European Conf. on Computer Vision (ECCV), pages 733–749, 2018.
  66. 66.Richard Zhang, Stefan A. Candra, Kai Vetter, and Avideh Zakhor. Sensor Fusion for Semantic Segmentation of Urban Scenes. In Proc. of the IEEE Intl. Conf. on Robotics & Automation (ICRA), 2015.

Citation

MLA
Behley, J., et al. “SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences”. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 9296–306, https://doi.org/10.1109/ICCV.2019.00939.
APA
Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., & Gall, J. (2019). SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 9296–9306. https://doi.org/10.1109/ICCV.2019.00939
Chicago
Behley, J., M. Garbade, A. Milioto, et al. 2019. “SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences”. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), 9296–9306. https://doi.org/10.1109/ICCV.2019.00939.
Harvard
Behley, J. et al. (2019) “SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp. 9296–9306. Available at: https://doi.org/10.1109/ICCV.2019.00939.
Vancouver
1. Behley J, Garbade M, Milioto A, Quenzel J, Behnke S, Stachniss C, Gall J (2019) SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In: 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, pp 9296–9306

BibTeX

@inproceedings{Behley_2019, title={SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences}, url={http://dx.doi.org/10.1109/ICCV.2019.00939}, DOI={10.1109/iccv.2019.00939}, booktitle={2019 IEEE/CVF International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Behley, Jens and Garbade, Martin and Milioto, Andres and Quenzel, Jan and Behnke, Sven and Stachniss, Cyrill and Gall, Jurgen}, year={2019}, month=Oct, pages={9296–9306} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE