Built independently by an author, for readers. Read the story and support ChapterPal

keyword

TPV fusion

TPV fusion, or tri-perspective view fusion, is a computer vision feature integration technique that combines and refines spatial representations across three mutually orthogonal two-dimensional projection planes—typically top, front, and side views—describing a three-dimensional environment. Rather than relying on computationally expensive three-dimensional voxel grids or a single planar projection such as bird's-eye view, this approach decomposes visual or geometric data into three canonical planes and merges their features through cross-view interaction or recurrent transitions between two-dimensional plane representations and three-dimensional spatial coordinates. By exchanging and consolidating geometric information across these complementary viewpoints, TPV fusion preserves fine structural detail and spatial consistency for tasks such as depth completion, semantic occupancy prediction, and 3D scene reconstruction.

1 item

Tri-Perspective view Decomposition for Geometry-Aware Depth Completion

Tri-Perspective view Decomposition for Geometry-Aware Depth Completion

Zhiqiang Yan, Yuankai Lin, Kun Wang, Yupeng Zheng, Yufei Wang, Zhenyu Zhang, Jun Li, Jian Yang

Why you should read this

Proposes a tri-perspective view decomposition framework that models structural 3D geometry across top, front, and side projections using spherical convolutions and geometric spatial propagation to achieve superior depth completion on standard benchmarks and real-world smartphone data.

Depth completion is a vital task for autonomous driving, as it involves reconstructing the precise 3D geometry of a scene from sparse and noisy depth measurements. However, most existing methods either rely only on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. To address this challenge, we introduce Tri-Perspective View Decomposition (TPVD), a novel framework that can explicitly model 3D geometry. In particular, (1) TPVD ingeniously decomposes the original point cloud into three 2D views, one of which corresponds to the sparse depth input. (2) We design TPV Fusion to update the 2D TPV features through recurrent 2D-3D-2D aggregation, where a Distance-Aware Spherical Convolution (DASC) is applied. (3) By adaptively choosing TPV affinitive neighbors, the newly proposed Geometric Spatial Propagation Network (GSPN) further improves the geometric consistency. As a result, our TPVD outperforms existing methods on KITTI, NYUv2, and SUN RGBD. Furthermore, we build a novel depth completion dataset named TOFDC, which is acquired by the time-of-flight (TOF) sensor and the color camera on smartphones. Project page.

Added

2026-09-26