Tri-Perspective view Decomposition for Geometry-Aware Depth Completion
Zhiqiang YanYuankai LinKun WangYupeng ZhengYufei WangZhenyu ZhangJun LiJian Yang
Proposes a tri-perspective view decomposition framework that models structural 3D geometry across top, front, and side projections using spherical convolutions and geometric spatial propagation to achieve superior depth completion on standard benchmarks and real-world smartphone data.
Accurate 3D depth perception is critical for autonomous driving, robotics, and mobile augmented reality. However, real-world depth sensors such as LiDAR and time-of-flight (TOF) cameras produce highly sparse measurements, with outdoor point cloud densities often dropping below 5%. Existing depth completion methods either process data entirely within 2D image space—losing critical 3D spatial geometry—or attempt to process raw 3D point clouds directly, which introduces significant computational complexity and struggles with uneven point distributions over long distances.
The article demonstrates a new framework called Tri-Perspective View Decomposition (TPVD) to reconstruct high-resolution, dense depth maps while explicitly preserving 3D geometric structure. The main objective was to develop an accurate, computationally efficient architecture that overcomes the sparsity and distance-varying characteristics of raw 3D sensor data.
The authors designed a framework that converts raw 3D point clouds into three orthogonal 2D projections (top, side, and front views) and refines them using standard 2D convolutions. To preserve structural accuracy, TPVD introduces a recurrent 2D-3D-2D fusion process that projects intermediate 2D features into 3D spherical space, applies distance-aware convolutions to normalize point density across varying ranges, and projects the updated representations back to 2D. A geometric refinement network then enforces spatial consistency across all three views. The researchers evaluated the method across standard benchmarks, including the outdoor KITTI dataset and indoor NYUv2 and SUN RGBD datasets, while also introducing a new real-world mobile depth dataset (TOFDC) comprising 10,000 paired RGB-D images collected via smartphone TOF sensors.
The experimental findings show that TPVD establishes state-of-the-art performance across all benchmark tests. First, on the competitive outdoor KITTI benchmark, TPVD achieved first place across all standard evaluation metrics, outperforming the five most recent leading models by an average margin of 15.98 mm in root mean squared error. Second, on the newly created TOFDC mobile dataset, the approach reduced root mean squared error by 15.6% and relative error by 33.3% compared to the strongest 3D-assisted baseline. Third, computational efficiency evaluations demonstrated that TPVD requires 134 billion fewer floating-point operations than comparable high-performing models, resulting in faster training and a test runtime of 8.82 frames per second. Finally, the framework maintained robustness across diverse stress conditions, including depth-only inputs without camera imagery, varying levels of point cloud sparsity, and synthetic adverse weather conditions such as fog, rain, and low lighting.
These findings indicate that 3D geometric awareness can be achieved without relying on resource-intensive raw point-cloud or voxel networks. By processing spatial information across decomposed 2D views and recurrent spherical transforms, systems can achieve higher geometric fidelity at lower computational cost. For deployment teams in autonomous vehicles and edge computing devices, this translates directly to safer navigation, reduced processing latency, and lower onboard hardware overhead.
Decision-makers and engineering teams developing vision-based perception systems should consider adopting multi-view 2D decomposition techniques as an alternative to native 3D processing pipelines. Prior to full-scale production deployment, organizations should conduct on-vehicle pilot testing to validate real-time frame rates under specific hardware constraints and evaluate sensor integration across varied weather conditions.
The primary limitation of the study is that adverse weather and lighting conditions were evaluated predominantly on synthetic virtual benchmarks rather than exhaustive physical-world edge cases. Nonetheless, given the consistent outperformance across multiple established public benchmarks and the newly introduced mobile dataset, confidence in the reported accuracy and structural stability remains high.
- Paper: Dynamic Spatial Propagation Network for Depth Completion, Yuankai Lin et al. (2022). It introduces dynamic affinity learning for spatial propagation networks in depth completion, which directly establishes the foundation for the Geometric Spatial Propagation Network (GSPN) developed in this work.
- Paper: SUN RGB-D: A RGB-D scene understanding benchmark suite, Shuran Song et al. (2015). It establishes the standard SUN RGB-D benchmark dataset and evaluation protocol for RGB-D indoor scene geometry completion and understanding utilized by the source.
- Paper: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li et al. (2022). It introduces foundational multi-view and bird's-eye-view projection and spatial aggregation architectures for autonomous driving that inform multi-perspective view representations.
- Paper: Depth Map Prediction from a Single Image using a Multi-Scale Deep Network, David Eigen et al. (2014). It presents foundational multi-scale deep network architectures and training protocols for depth prediction on KITTI and NYU Depth benchmarks.
- Paper: Depth Anything V2, Lihe Yang et al. (2024). It extends metric geometric surface estimation by presenting a foundation model trained on synthetic data and scaled pseudo-labels for high-resolution dense depth recovery across complex scenes.
- Paper: Depth Pro: Sharp Monocular Metric Depth in Less Than a Second, Alexey Bochkovskiy et al. (2025). It builds on zero-shot metric depth estimation to recover sharp geometric boundaries and high-resolution depth maps without relying on sparse sensor measurements.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). It generalizes multi-view and single-view visual geometry prediction into an end-to-end feed-forward transformer framework that jointly estimates depth, poses, and dense 3D point structures.
- Paper: Symphonize 3D Semantic Scene Completion with Contextual Instance Queries, Haoyi Jiang et al. (2024). It advances 3D scene completion by integrating instance-level queries with depth rectification to resolve fine-grained geometric occupancy and semantic categories.
