keyword
KITTI dataset
The KITTI dataset is a widely used multimodal benchmark dataset in computer vision and autonomous driving research, created by the Karlsruhe Institute of Technology and the Toyota Technological Institute at Chicago. Captured from an instrumented vehicle driven through urban, rural, and highway environments, the dataset provides synchronized high-resolution stereo camera images, 3D LiDAR point clouds, and GPS and inertial measurement data paired with accurate ground-truth annotations. It serves as a foundational benchmark for developing and evaluating algorithms in perception and spatial understanding, including 2D and 3D object detection, depth prediction, visual odometry, optical flow, and multi-object tracking.
3 items

Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point Clouds
Chenhang He, Ruihuang Li, Shuai Li, Lei Zhang
Why you should read this
Proposes Voxel Set Transformer, a linear-complexity backbone that processes arbitrary-sized 3D point clusters in parallel through latent-code induced cross-attentions to achieve efficient and competitive 3D object detection on large-scale point clouds.
Transformer has demonstrated promising performance in many 2D vision tasks. However, it is cumbersome to compute the self-attention on large-scale point cloud data because point cloud is a long sequence and unevenly distributed in 3D space. To solve this issue, existing methods usually compute self-attention locally by grouping the points into clusters of the same size, or perform convolutional self-attention on a discretized representation. However, the former results in stochastic point dropout, while the latter typically has narrow attention fields. In this paper, we propose a novel voxel-based architecture, namely Voxel Set Transformer (VoxSeT), to detect 3D objects from point clouds by means of set-to-set translation. VoxSeT is built upon a voxel-based set attention (VSA) module, which reduces the self-attention in each voxel by two cross-attentions and models features in a hidden space induced by a group of latent codes. With the VSA module, VoxSeT can manage voxelized point clusters with arbitrary size in a wide range, and process them in parallel with linear complexity. The proposed VoxSeT integrates the high performance of transformer with the efficiency of voxel-based model, which can be used as a good alternative to the convolutional and point-based backbones. VoxSeT reports competitive results on the KITTI and Waymo detection benchmarks. The source codes can be found at https://github.com/skyhehe123/VoxSeT.
Added
2026-10-05

OcTr: Octree-Based Transformer for 3D Object Detection
Chao Zhou, Yanan Zhang, Jiaxin Chen, Di Huang
Why you should read this
Proposes an octree-based Transformer architecture that captures coarse-to-fine global context through hierarchical dynamic self-attention, reducing computational complexity to linear while achieving state-of-the-art 3D object detection on large-scale LiDAR point clouds.
A key challenge for LiDAR-based 3D object detection is to capture sufficient features from large scale 3D scenes especially for distant or/and occluded objects. Albeit recent efforts made by Transformers with the long sequence modeling capability, they fail to properly balance the accuracy and efficiency, suffering from inadequate receptive fields or coarse-grained holistic correlations. In this paper, we propose an Octree-based Transformer, named OcTr, to address this issue. It first constructs a dynamic octree on the hierarchical feature pyramid through conducting self-attention on the top level and then recursively propagates to the level below restricted by the octants, which captures rich global context in a coarse-to-fine manner while maintaining the computational complexity under control. Furthermore, for enhanced foreground perception, we propose a hybrid positional embedding, composed of the semantic-aware positional embedding and attention mask, to fully exploit semantic and geometry clues. Extensive experiments are conducted on the Waymo Open Dataset and KITTI Dataset, and OcTr reaches newly state-of-the-art results.
Added
2026-09-26

Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
David Eigen, Christian Puhrsch, Rob Fergus
Why you should read this
Proposes a multi-scale deep network and a scale-invariant training objective that estimates accurate dense depth from a single image by integrating coarse global context with fine local refinements.
Predicting depth is an essential component in understanding the 3D geometry of a scene. While for stereo images local correspondence suffices for estimation, finding depth relations from a single image is less straightforward, requiring integration of both global and local information from various cues. Moreover, the task is inherently ambiguous, with a large source of uncertainty coming from the overall scale. In this paper, we present a new method that addresses this task by employing two deep network stacks: one that makes a coarse global prediction based on the entire image, and another that refines this prediction locally. We also apply a scale-invariant error to help measure depth relations rather than scale. By leveraging the raw datasets as large sources of training data, our method achieves state-of-the-art results on both NYU Depth and KITTI, and matches detailed depth boundaries without the need for superpixelation.
Added
2026-09-10
