CompletionFormer: Depth Completion with Convolutions and Vision Transformers
Youmin ZhangXianda GuoMatteo PoggiZheng ZhuGuan HuangStefano Mattoccia
Proposes CompletionFormer, a pyramidal depth completion architecture that integrates convolutional attention with vision transformers in a single-branch model, achieving state-of-the-art accuracy on KITTI and NYUv2 benchmarks while requiring nearly one-third the computation of pure transformer approaches.
Depth sensing technologies, such as laser-based distance sensors and visual rangefinders, are essential for autonomous driving and augmented reality. However, real-world measurements often yield sparse, incomplete, and noisy depth data due to sensor limitations, surface reflections, or long distances. Existing solutions rely heavily on standard convolutional networks or pure visual attention models, which both face operational trade-offs: convolutional networks excel at fine local boundaries but struggle with global spatial context, while pure attention-based networks capture scene-wide relationships at the expense of fine local details and high computational efficiency.
The article introduces and evaluates CompletionFormer, a deep learning architecture designed to accurately reconstruct dense depth maps from sparse inputs and visual camera images. The main objective is to demonstrate that tightly coupling convolutional layers and vision transformers within a single processing pipeline delivers superior depth reconstruction accuracy while maintaining manageable computational costs.
The authors designed a hybrid Joint Convolutional Attention and Transformer block organized in a multi-scale, single-branch pyramid framework. This design merges color images and sparse depth inputs early, extracts features across five hierarchical stages using parallel attention pathways, and refines the initial predictions using a spatial propagation network. The method was rigorously tested using benchmark datasets for outdoor autonomous driving (the KITTI depth completion benchmark) and indoor environments (the NYUv2 dataset), evaluating performance across varying levels of measurement sparsity, from 64 sensor lines down to a single scan line.
The evaluation yielded several key findings. First, CompletionFormer established state-of-the-art accuracy across both indoor and outdoor datasets, achieving the lowest overall error rates among published techniques. Second, the hybrid architecture reduced computational burden to approximately one-third of the operations required by dual-branch, pure transformer models (559.5 billion operations compared to nearly 2 trillion). Third, the performance advantage became significantly wider under extreme data sparsity; on 1-line sensor inputs, the full hybrid model reduced depth errors from 3,507 mm down to 3,250 mm compared to baseline models. Finally, the improved global context allowed the refinement module to converge effectively in only 6 iterations instead of the standard 18, reducing processing overhead.
These results show that combining local boundary precision with global spatial awareness is critical for reliable depth estimation in safety-critical systems. For industrial applications such as automated transport, high-accuracy completion from sparse inputs lowers hardware costs by enabling the use of less expensive, lower-density depth sensors without sacrificing reconstruction quality.
Decision-makers and engineering teams evaluating 3D vision systems should consider adopting hybrid convolutional-transformer backbones over single-paradigm architectures. Before deploying this architecture in real-time embedded systems, further engineering work is required to optimize its execution speed, as the current model runs at approximately 10 frames per second. The conclusions are supported by thorough comparative evaluations on standardized industry datasets, though real-time operational deployment requires testing on dedicated low-power edge hardware.
- Paper: Dynamic Spatial Propagation Network for Depth Completion, Yuankai Lin et al. (2022). Read this earlier depth-completion work to understand the spatial propagation refinement mechanism that CompletionFormer adapts and improves.
- Paper: Tri-Perspective view Decomposition for Geometry-Aware Depth Completion, Zhiqiang Yan et al. (2024). This later depth-completion framework advances the task by adding explicit 3D geometric structure to reconstruction, building on the dense depth-completion setting established by CompletionFormer.
