CompletionFormer: Depth Completion with Convolutions and Vision Transformers

Youmin ZhangXianda GuoMatteo PoggiZheng ZhuGuan HuangStefano Mattoccia

article2023CVPR198 citations

Proposes CompletionFormer, a pyramidal depth completion architecture that integrates convolutional attention with vision transformers in a single-branch model, achieving state-of-the-art accuracy on KITTI and NYUv2 benchmarks while requiring nearly one-third the computation of pure transformer approaches.

Listen

Depth sensing technologies, such as laser-based distance sensors and visual rangefinders, are essential for autonomous driving and augmented reality. However, real-world measurements often yield sparse, incomplete, and noisy depth data due to sensor limitations, surface reflections, or long distances. Existing solutions rely heavily on standard convolutional networks or pure visual attention models, which both face operational trade-offs: convolutional networks excel at fine local boundaries but struggle with global spatial context, while pure attention-based networks capture scene-wide relationships at the expense of fine local details and high computational efficiency.

The article introduces and evaluates CompletionFormer, a deep learning architecture designed to accurately reconstruct dense depth maps from sparse inputs and visual camera images. The main objective is to demonstrate that tightly coupling convolutional layers and vision transformers within a single processing pipeline delivers superior depth reconstruction accuracy while maintaining manageable computational costs.

The authors designed a hybrid Joint Convolutional Attention and Transformer block organized in a multi-scale, single-branch pyramid framework. This design merges color images and sparse depth inputs early, extracts features across five hierarchical stages using parallel attention pathways, and refines the initial predictions using a spatial propagation network. The method was rigorously tested using benchmark datasets for outdoor autonomous driving (the KITTI depth completion benchmark) and indoor environments (the NYUv2 dataset), evaluating performance across varying levels of measurement sparsity, from 64 sensor lines down to a single scan line.

The evaluation yielded several key findings. First, CompletionFormer established state-of-the-art accuracy across both indoor and outdoor datasets, achieving the lowest overall error rates among published techniques. Second, the hybrid architecture reduced computational burden to approximately one-third of the operations required by dual-branch, pure transformer models (559.5 billion operations compared to nearly 2 trillion). Third, the performance advantage became significantly wider under extreme data sparsity; on 1-line sensor inputs, the full hybrid model reduced depth errors from 3,507 mm down to 3,250 mm compared to baseline models. Finally, the improved global context allowed the refinement module to converge effectively in only 6 iterations instead of the standard 18, reducing processing overhead.

These results show that combining local boundary precision with global spatial awareness is critical for reliable depth estimation in safety-critical systems. For industrial applications such as automated transport, high-accuracy completion from sparse inputs lowers hardware costs by enabling the use of less expensive, lower-density depth sensors without sacrificing reconstruction quality.

Decision-makers and engineering teams evaluating 3D vision systems should consider adopting hybrid convolutional-transformer backbones over single-paradigm architectures. Before deploying this architecture in real-time embedded systems, further engineering work is required to optimize its execution speed, as the current model runs at approximately 10 frames per second. The conclusions are supported by thorough comparative evaluations on standardized industry datasets, though real-time operational deployment requires testing on dedicated low-power edge hardware.

Cover for CompletionFormer: Depth Completion with Convolutions and Vision Transformers

Abstract

Given sparse depths and the corresponding RGB images, depth completion aims at spatially propagating the sparse measurements throughout the whole image to get a dense depth prediction. Despite the tremendous progress of deep-learning-based depth completion methods, the locality of the convolutional layer or graph model makes it hard for the network to model the long-range relationship between pixels. While recent fully Transformer-based architecture has reported encouraging results with the global receptive field, the performance and efficiency gaps to the well-developed CNN models still exist because of its deteriorative local feature details. This paper proposes a Joint Convolutional Attention and Transformer block (JCAT), which deeply couples the convolutional attention layer and Vision Transformer into one block, as the basic unit to construct our depth completion model in a pyramidal structure. This hybrid architecture naturally benefits both the local connectivity of convolutions and the global context of the Transformer in one single model. As a result, our CompletionFormer outperforms state-of-the-art CNNs-based methods on the outdoor KITTI Depth Completion benchmark and indoor NYUv2 dataset, achieving significantly higher efficiency (nearly 1/3 FLOPs) compared to pure Transformer-based methods. Code is available at https://github.com/youmi-zym/CompletionFormer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 RGB and Depth Embedding
  • 3.2 Joint Convolutional Attention and Transformer Encoder
  • 3.3 Decoder
  • 3.4 SPN Refinement and Loss Function
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Implementation Details
  • 4.3 Evaluation Metrics
  • 4.4 Ablation Studies and Analysis
  • 4.5 Sparsity Level Analysis
  • 4.6 Comparison with SOTA Methods
  • 5 Conclusion and Limitations
  • A Appendix
  • A.1 Qualitative Results on NYUv2 Dataset
  • A.2 Model Architecture Details
  • References

Knowls

  1. Knowl 1 — CompletionFormer Architecture for Depth Completion

    model/method

    CompletionFormer is a single-branch pyramidal encoder-decoder neural network designed for depth completion from an RGB image and a sparse depth map.

    The framework consists of four primary stages:

    1. RGB and Depth Embedding: The input sparse depth map SS and RGB image II are independently mapped via separate 3×33 \times 3 convolutions into feature spaces, concatenated along the channel dimension, and fused with a convolutional layer into a shared multi-modal feature representation.
    2. Hierarchical Encoder: Stage 1 processes the fused representation using ResNet-34 BasicBlocks to output half-resolution feature map F1F_1. Stages 2 through 5 each consist of a patch embedding layer (3×33 \times 3 convolution with stride 2) followed by LiL_i repeated Joint Convolutional Attention and Transformer (JCAT) blocks, producing a multi-scale feature pyramid with resolutions {1/4,1/8,1/16,1/32}\{1/4, 1/8, 1/16, 1/32\} relative to the input resolution.
    3. Multiscale Decoder: Decoding layers upsample deeper features via transposed convolutions and apply Convolutional Block Attention Modules (CBAM) across spatial and channel dimensions while aggregating encoder features via skip connections. A prediction head fuses full-resolution features with stage-1 and raw input embeddings to predict an initial dense depth map D0D^0, an affinity matrix ww, and a confidence map.
    4. Spatial Propagation Refinement: A Non-Local Spatial Propagation Network (NLSPN) refines the initial estimate D0D^0 over KK recurrent steps using the predicted affinity matrix ww to yield the final depth prediction DKD^K.
  2. Knowl 2 — Joint Convolutional Attention and Transformer Block

    model/method

    The Joint Convolutional Attention and Transformer (JCAT) block is the basic building unit of the CompletionFormer encoder stages 2 to 5. It integrates a local convolutional attention stream and a global self-attention stream in parallel.

    Given an input feature tensor F∈RHi×Wi×CF \in \mathbb{R}^{H_i \times W_i \times C} at stage ii, the block branches into two parallel paths:

    1. Transformer Path: The input features are normalized using Layer Normalization (LN) and flattened into token vectors X∈RN×CX \in \mathbb{R}^{N \times C} with N=Hi×WiN = H_i \times W_i. Linear projections WQ,WK,WV∈RC×CW^Q, W^K, W^V \in \mathbb{R}^{C \times C} generate query QQ, key KK, and value VV matrices. Spatial-reduction attention (SRA) reduces the spatial scale of KK and VV via spatial reduction convolutions before applying multi-head self-attention: Attention(Q,K,V)=Softmax(QKTChead)V\text{Attention}(Q, K, V) = \text{Softmax}\left(\frac{QK^T}{\sqrt{C_{\text{head}}}}\right)V where CheadC_{\text{head}} is the channel dimension per attention head. The output is processed by a feed-forward network (FFN) with Layer Normalization and residual connections, then reshaped back to RHi×Wi×C\mathbb{R}^{H_i \times W_i \times C}.
    2. Convolutional Attention Path: The input features are processed by standard convolutions coupled with channel and spatial attention modules (CBAM) to extract fine-grained local connectivity and suppress noise.

    Fusion: The output representations from both paths are concatenated along the channel dimension and fused via a 3×33 \times 3 convolution layer before being passed to the next block or stage.

  3. Knowl 3 — CompletionFormer Model Configurations

    model/method

    CompletionFormer is defined across three standard architectural capacities—Tiny, Small, and Base—by varying the number of JCAT blocks per encoder stage while keeping the channel widths fixed across stages 2 through 5 at C∈{64,128,320,512}C \in \{64, 128, 320, 512\}:

    • CompletionFormer-Tiny: Layer configuration [L2,L3,L4,L5]=[2,2,2,2][L_2, L_3, L_4, L_5] = [2, 2, 2, 2], containing 41.5M parameters and requiring 191.7G FLOPs at 480×640480 \times 640 resolution.
    • CompletionFormer-Small: Layer configuration [L2,L3,L4,L5]=[3,3,6,3][L_2, L_3, L_4, L_5] = [3, 3, 6, 3], containing 78.3M parameters and requiring 231.8G FLOPs at 480×640480 \times 640 resolution (or 82.6M parameters and 429.6G FLOPs with the full decoder and NLSPN setup).
    • CompletionFormer-Base: Layer configuration [L2,L3,L4,L5]=[3,3,18,3][L_2, L_3, L_4, L_5] = [3, 3, 18, 3], containing 142.4M parameters and requiring 301.9G FLOPs at 480×640480 \times 640 resolution (146.7M parameters and 499.6G FLOPs with the full decoder).
  4. Knowl 4 — Depth Refinement Formulation and Supervised Loss Function

    equation

    CompletionFormer refines the initial depth map D0=(du,v0)∈RH×WD^0 = (d^0_{u,v}) \in \mathbb{R}^{H \times W} over KK recurrent iterations using non-local spatial propagation (NLSPN). At iteration step tt, the updated depth value du,vtd^t_{u,v} at pixel (u,v)(u,v) is computed as:

    du,vt=wu,v(0,0)du,vt−1+∑(i,j)∈Nu,vNL,i≠0,j≠0wu,v(i,j)di,jt−1d_{u,v}^t = w_{u,v}(0, 0) d_{u,v}^{t-1} + \sum_{(i,j) \in \mathcal{N}_{u,v}^{NL}, i \neq 0, j \neq 0} w_{u,v}(i, j) d_{i,j}^{t-1}

    where Nu,vNL\mathcal{N}_{u,v}^{NL} is the set of non-local spatial neighbors for pixel (u,v)(u, v), wu,v(i,j)∈(−1,1)w_{u,v}(i, j) \in (-1, 1) represents the affinity weight between reference pixel (u,v)(u, v) and neighbor pixel (i,j)(i, j) output by the decoder and modulated by an estimated confidence map, and the self-preservation weight is:

    wu,v(0,0)=1−∑(i,j)∈Nu,vNL,i≠0,j≠0wu,v(i,j)w_{u,v}(0, 0) = 1 - \sum_{(i,j) \in \mathcal{N}_{u,v}^{NL}, i \neq 0, j \neq 0} w_{u,v}(i, j)

    The network is supervised end-to-end using a joint L1L_1 and L2L_2 loss on the final refined depth map D^=DK\hat{D} = D^K over all valid ground-truth pixels VV:

    L(D^,Dgt)=1∣V∣∑v∈V(∣D^v−Dvgt∣+∣D^v−Dvgt∣2)\mathcal{L}(\hat{D}, D^{gt}) = \frac{1}{|V|} \sum_{v \in V} \left( |\hat{D}_v - D^{gt}_v| + |\hat{D}_v - D^{gt}_v|^2 \right)

    where DvgtD^{gt}_v is the ground-truth depth at pixel index vv, and ∣V∣|V| denotes the total number of valid pixels.

  5. Knowl 5 — Benchmark Depth Completion Performance on KITTI and NYUv2

    data/table

    CompletionFormer achieves state-of-the-art results on both the outdoor KITTI Depth Completion benchmark and the indoor NYUv2 dataset.

    Method KITTI DC NYUv2
    MAE ↓\downarrow iMAE ↓\downarrow iRMSE ↓\downarrow RMSE ↓\downarrow RMSE ↓\downarrow REL ↓\downarrow
    (mm) (1/km) (1/km) (mm) (m)
    CSPN 279.46 1.15 2.93 1019.64 0.117 0.016
    DeepLiDAR 226.50 1.15 2.56 758.38 0.115 0.022
    GuideNet 218.83 0.99 2.25 736.24 0.101 0.015
    NLSPN 199.59 0.84 1.99 741.68 0.092 0.012
    PENet 210.55 0.94 2.17 730.08 - -
    ACMNet 206.09 0.90 2.08 744.91 0.105 0.015
    TWISE 195.58 0.82 2.08 840.20 0.097 0.013
    RigNet 203.25 0.90 2.08 712.66 0.090 0.013
    GuideFormer 207.76 0.97 2.14 721.48 - -
    DySPN 192.71 0.82 1.88 709.12 0.090 0.012
    Ours (L1L_1) 183.88 0.80 1.89 764.87 - -
    Ours (L1+L2L_1 + L_2) 203.45 0.88 2.01 708.87 0.090 0.012

    When trained with L1L_1 loss, CompletionFormer attains the lowest MAE (183.88 mm183.88\text{ mm}) and iMAE (0.80 km−10.80\text{ km}^{-1}) on KITTI DC. When trained with joint L1+L2L_1+L_2 loss, it achieves the lowest RMSE (708.87 mm708.87\text{ mm}) among published methods on KITTI DC and matches the top performance on NYUv2 (0.090 m0.090\text{ m} RMSE, 0.0120.012 REL).

  6. Knowl 6 — Robustness Across Variable Sparse Depth Densities

    data/table

    CompletionFormer demonstrates superior resilience across severe sparsity levels in comparison to pure CNN baselines (NLSPN, DySPN) and a pure Vision Transformer variant (Ours-ViT).

    On the KITTI Depth Completion dataset (trained on 10,000 samples and evaluated on the selected validation set with LiDAR lines subsampled to 1, 4, 16, and 64 lines):

    Scanning Lines Method RMSE ↓\downarrow (mm) MAE ↓\downarrow (mm) iRMSE ↓\downarrow (1/km) iMAE ↓\downarrow (1/km)
    1 NLSPN 3507.7 1849.1 13.8 8.9
    DySPN 3625.5 1924.7 13.8 8.9
    Ours-ViT 3507.2 1807.7 12.1 7.8
    Ours 3250.2 1582.6 10.4 6.6
    4 NLSPN 2293.1 831.3 7.0 3.4
    DySPN 2285.8 834.3 6.3 3.2
    Ours-ViT 2241.2 795.9 5.8 2.9
    Ours 2150.0 740.1 5.4 2.6
    16 NLSPN 1288.9 377.2 3.4 1.4
    DySPN 1274.8 366.4 3.2 1.3
    Ours-ViT 1268.9 360.7 3.3 1.3
    Ours 1218.6 337.4 3.0 1.2
    64 NLSPN 889.4 238.8 2.6 1.0
    DySPN 878.5 228.6 2.5 1.0
    Ours-ViT 872.0 226.2 2.5 1.0
    Ours 848.7 215.9 2.5 0.9

    On the NYUv2 dataset with 0, 50, 200, and 500 depth point samples:

    Sample Count PackNet-SAN GuideNet NLSPN Ours-ViT Ours
    0 - - 0.562 0.544 0.490
    50 - - 0.223 0.218 0.208
    200 0.155 0.142 0.129 0.130 0.127
    500 0.120 0.101 0.092 0.091 0.090

    Under extreme sparsity (e.g., 1 scanning line or 0 input depth samples), the global context captured by the Transformer components provides substantial gains over pure CNN architectures, while coupling CNN local detail with Transformer global reasoning in CompletionFormer outperforms all alternatives.

  7. Knowl 7 — Ablation of CompletionFormer Architecture Components on NYUv2

    data/table

    Ablations on the NYUv2 dataset (with 500 randomly sampled depth points, input resolution 480×640480 \times 640) isolate the contribution of key structural choices in CompletionFormer:

    Configuration RMSE ↓\downarrow (mm) MAE ↓\downarrow (mm) Params ↓\downarrow (M) FLOPs ↓\downarrow (G)
    (A) Cascaded connection 91.5 35.7 82.6 429.6
    (B) Parallel connection (Ours-Small) 90.0 35.0 82.6 429.6
    (C) w/o Spatial and Channel Attention in JCAT 91.1 35.5 82.5 429.4
    (D) Dual-branch encoder 94.0 36.4 161.0 661.4
    (E) ResNet34 backbone (18 ref. iters, no decoder attn) 92.3 36.1 26.4 542.2
    (F) ResNet34 backbone (18 ref. iters, w/ decoder attn) 91.4 35.5 28.1 582.1
    (G) Swin-Tiny backbone 92.6 36.4 38.1 634.8
    (H) PVT-Large backbone 91.4 35.6 68.3 419.8
    (I) MPViT-Base backbone 91.0 35.5 83.1 1259.3
    (J) CMT-Base backbone 92.0 35.9 47.6 358.7
    (L) Ours-Small (6 NLSPN iters) 90.0 35.0 82.6 429.6
    (M) Ours-Small (CSPN++ refinement) 90.3 34.9 82.7 446.4
    (N) Ours-Tiny (6 NLSPN iters) 90.9 35.3 45.8 389.4
    (O) Ours-Base (6 NLSPN iters) 90.1 35.1 146.7 499.6

    Key takeaways from the ablation:

    1. Parallel vs. Cascaded: Parallel fusion of Transformer and CNN paths outperforms cascaded connections (90.0 mm90.0\text{ mm} vs 91.5 mm91.5\text{ mm} RMSE).
    2. Convolutional Attention: Adding spatial and channel attention in JCAT reduces RMSE from 91.1 mm91.1\text{ mm} to 90.0 mm90.0\text{ mm} with only a 0.2G0.2\text{G} FLOPs overhead.
    3. Single vs. Dual Branch: Early fusion in a single branch outperforms dual-branch encoding (90.0 mm90.0\text{ mm} at 429.6G429.6\text{G} FLOPs vs 94.0 mm94.0\text{ mm} at 661.4G661.4\text{G} FLOPs).
    4. Refinement Convergence: Due to the global context modeled by the encoder, CompletionFormer converges in only 6 NLSPN iterations (90.0 mm90.0\text{ mm} RMSE) compared to 18 iterations needed by ResNet-34 baselines (92.3 mm92.3\text{ mm} RMSE). Without any SPN refinement, Ours-Small achieves 99.2 mm99.2\text{ mm} RMSE compared to 106.5 mm106.5\text{ mm} for ResNet-34 and 100.2 mm100.2\text{ mm} for MPViT-Base.
  8. Knowl 8 — Inference Latency Limitation

    limitation

    CompletionFormer operates at an inference speed of approximately 10 frames per second (FPS). Although it substantially reduces FLOPs relative to pure dual-branch Transformer networks (such as GuideFormer, which requires near 2T FLOPs at 352×1216352 \times 1216 resolution), 10 FPS does not meet the hard real-time latency thresholds required in time-critical autonomous driving and robotics pipelines.

Coverage note — None was omitted; all key architectural components (JCAT, embedding, decoder, SPN refinement, loss functions), benchmark comparisons, sparsity studies, ablations, and stated limitations are covered.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Xinjing Cheng, Peng Wang, Chenye Guan, and Ruigang Yang. Cspn++: Learning context and resource aware convolutional spatial propagation networks for depth completion. In AAAI, volume 34, pages 10615–10622, 2020.
  3. 3.Xinjing Cheng, Peng Wang, and Ruigang Yang. Depth estimation via affinity learned with convolutional spatial propagation network. In ECCV, pages 103–119, 2018.
  4. 4.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. ArXiv preprint arXiv:2010.11929, 2020.
  5. 5.Vitor Guizilini, Rares Ambrus, Wolfram Burgard, and Adrien Gaidon. Sparse auxiliary networks for unified monocular depth prediction and completion. In CVPR, pages 11078–11088, 2021.
  6. 6.Jianyuan Guo, Kai Han, Han Wu, Yehui Tang, Xinghao Chen, Yunhe Wang, and Chang Xu. Cmt: Convolutional neural networks meet vision transformers. In CVPR, pages 12175–12185, June 2022.
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In ECCV, pages 630–645. Springer, 2016.
  8. 8.Mu Hu, Shuling Wang, Bin Li, Shiyu Ning, Li Fan, and Xiaojin Gong. Penet: Towards precise and efficient image guided depth completion. In ICRA, pages 13656–13662. IEEE, 2021.
  9. 9.Saif Imran, Xiaoming Liu, and Daniel Morris. Depth completion with twin surface extrapolation at occlusion boundaries. In CVPR, pages 2583–2592, 2021.
  10. 10.Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley. Learning to estimate hidden motions with global motion aggregation. In ICCV, pages 9772–9781, 2021.
  11. 11.Leonid Keselman, John Iselin Woodfill, Anders Grunnet-Jepsen, and Achintya Bhowmik. Intel realsense stereoscopic depth cameras. In CVPR Workshops, pages 1–10, 2017.
  12. 12.Youngwan Lee, Jonghee Kim, Jeffrey Willette, and Sung Ju Hwang. Mpvit: Multi-path vision transformer for dense prediction. In CVPR, 2022.
  13. 13.Jiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai, Ziwei Yan, Lei Yang, Jiangyu Liu, Haoqiang Fan, and Shuaicheng Liu. Practical stereo matching via cascaded recurrent network with adaptive correlation. ArXiv preprint arXiv:2203.11483, 2022.
  14. 14.Zhenyu Li, Zehui Chen, Xianming Liu, and Junjun Jiang. Depthformer: Exploiting long-range correlation and local information for accurate monocular depth estimation. ArXiv preprint arXiv:2203.14211, 2022.
  15. 15.Zhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy Ding, Francis X Creighton, Russell H Taylor, and Mathias Unberath. Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers. In ICCV, pages 6197–6206, 2021.
  16. 16.Yuankai Lin, Tao Cheng, Qi Zhong, Wending Zhou, and Hua Yang. Dynamic spatial propagation network for depth completion. In AAAI, 2022.
  17. 17.Lina Liu, Xibin Song, Xiaoyang Lyu, Junwei Diao, Mengmeng Wang, Yong Liu, and Liangjun Zhang. Fcfr-net: Feature fusion based coarse- to-fine residual learning for depth completion. In AAAI, volume 35, pages 2136–2144, 2021.
  18. 18.Sifei Liu, Shalini De Mello, Jinwei Gu, Guangyu Zhong, Ming-Hsuan Yang, and Jan Kautz. Learning affinity via spatial propagation networks. NeurIPS, 2017.
  19. 19.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, pages 10012–10022, 2021.
  20. 20.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  21. 21.Fangchang Ma, Guilherme Venturelli Cavalheiro, and Sertac Karaman. Self-supervised sparse-to-dense: Self-supervised depth completion from lidar and monocular camera. In ICRA, pages 3288–3295. IEEE, 2019.
  22. 22.Fangchang Ma and Sertac Karaman. Sparse-to-dense: Depth prediction from sparse depth samples and a single image. In ICRA, 2018.
  23. 23.Microsoft. Kinect for windows. https://developer.microsoft.com/en-us/windows/kinect/.
  24. 24.Danish Nazir, Marcus Liwicki, Didier Stricker, and Muhammad Zeshan Afzal. Semattnet: Towards attention-based semantic aware guided depth completion. arXiv preprint arXiv:2204.13635, 2022.
  25. 25.Xuran Pan, Chunjiang Ge, Rui Lu, Shiji Song, Guanfu Chen, Zeyi Huang, and Gao Huang. On the integration of self-attention and convolution. In CVPR, pages 815–825, 2022.
  26. 26.Jinsun Park, Kyungdon Joo, Zhe Hu, Chi-Kuei Liu, and In So Kweon. Non-local spatial propagation network for depth completion. In ECCV, pages 120–136. Springer, 2020.
  27. 27.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alche-Buc, E. Fox, and R. Garnett, editors, NeurIPS, pages 8024–8035. Curran Associates, Inc., 2019.
  28. 28.Zhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie, Yaowei Wang, Jianbin Jiao, and Qixiang Ye. Conformer: Local features coupling global representations for visual recognition. In ICCV, pages 367–376, October 2021.
  29. 29.Jiaxiong Qiu, Zhaopeng Cui, Yinda Zhang, Xingdi Zhang, Shuaicheng Liu, Bing Zeng, and Marc Pollefeys. Deeplidar: Deep surface normal guided depth prediction for outdoor scene from sparse lidar data and single color image. In CVPR, pages 3313–3322, 2019.
  30. 30.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. ArXiv preprint, 2021.
  31. 31.Kyeongha Rho, Jinsung Ha, and Youngjung Kim. Guideformer: Transformers for image guided depth completion. In CVPR, pages 6250–6259, 2022.
  32. 32.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, pages 4510–4520, 2018.
  33. 33.Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, pages 746–760. Springer, 2012.
  34. 34.Xiuchao Sui, Shaohua Li, Xue Geng, Yan Wu, Xinxing Xu, Yong Liu, Rick Siow Mong Goh, and Hongyuan Zhu. Craft: Cross-attentional flow transformers for robust optical flow. In CVPR, 2022.
  35. 35.Jie Tang, Fei-Peng Tian, Wei Feng, Jian Li, and Ping Tan. Learning guided convolutional network for depth completion. IEEE TIP, pages 1116–1129, 2020.
  36. 36.Jonas Uhrig, Nick Schneider, Lukas Schneider, Uwe Franke, Thomas Brox, and Andreas Geiger. Sparsity invariant cnns. In 3DV, pages 11–20. IEEE, 2017.
  37. 37.Wouter Van Gansbeke, Davy Neven, Bert De Brabandere, and Luc Van Gool. Sparse and noisy lidar completion with rgb guidance and uncertainty. In MVA, pages 1–6. IEEE, 2019.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017.
  39. 39.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In ICCV, pages 568–578, 2021.
  40. 40.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. Cbam: Convolutional block attention module. In ECCV, pages 3–19, 2018.
  41. 41.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. NeurIPS, 34, 2021.
  42. 42.Xin Xiong, Haipeng Xiong, Ke Xian, Chen Zhao, Zhiguo Cao, and Xin Li. Sparse-to-dense depth completion revisited: Sampling strategy and graph construction. In ECCV, pages 682–699. Springer, 2020.
  43. 43.Weijian Xu, Yifan Xu, Tyler Chang, and Zhuowen Tu. Coscale conv-attentional image transformers. In ICCV, pages 9981–9990, 2021.
  44. 44.Zheyuan Xu, Hongche Yin, and Jian Yao. Deformable spatial propagation networks for depth completion. In ICIP, pages 913–917. IEEE, 2020.
  45. 45.Zhiqiang Yan, Kun Wang, Xiang Li, Zhenyu Zhang, Baobei Xu, Jun Li, and Jian Yang. Rignet: Repetitive image guided network for depth completion. ECCV, 2022.
  46. 46.Yinda Zhang and Thomas Funkhouser. Deep depth completion of a single rgb-d image. In CVPR, pages 175–185, 2018.
  47. 47.Yongchi Zhang, Ping Wei, Huan Li, and Nanning Zheng. Multiscale adaptation fusion networks for depth completion. In IJCNN, pages 1–7, 2020.
  48. 48.Chaoqiang Zhao, Youmin Zhang, Matteo Poggi, Fabio Tosi, Xianda Guo, Zheng Zhu, Guan Huang, Yang Tang, and Stefano Mattoccia. Monovit: Self-supervised monocular depth estimation with a vision transformer. In 3DV, 2022.
  49. 49.Shanshan Zhao, Mingming Gong, Huan Fu, and Dacheng Tao. Adaptive context-aware multi-modal network for depth completion. TIP, pages 5264–5276, 2021.
  50. 50.Yiqi Zhong, Cho-Ying Wu, Suya You, and Ulrich Neumann. Deep rgb-d canonical correlation analysis for sparse depth completion. NeurIPS, 2019.
  51. 51.Yufan Zhu, Weisheng Dong, Leida Li, Jinjian Wu, Xin Li, and Guangming Shi. Robust depth completion with uncertainty-driven loss functions. In AAAI, 2022.

Citation

MLA
Youmin, Z., et al. “CompletionFormer: Depth Completion with Convolutions and Vision Transformers”. arXiv, 2023, http://arxiv.org/abs/2304.13030v1.
APA
Youmin, Z., Xianda, G., Matteo, P., Zheng, Z., Guan, H., & Stefano, M. (2023). CompletionFormer: Depth Completion with Convolutions and Vision Transformers. arXiv. http://arxiv.org/abs/2304.13030v1
Chicago
Youmin, Z., G. Xianda, P. Matteo, Z. Zheng, H. Guan, and M. Stefano. 2023. “CompletionFormer: Depth Completion with Convolutions and Vision Transformers”. arXiv. http://arxiv.org/abs/2304.13030v1.
Harvard
Youmin, Z. et al. (2023) “CompletionFormer: Depth Completion with Convolutions and Vision Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.13030v1.
Vancouver
1. Youmin Z, Xianda G, Matteo P, Zheng Z, Guan H, Stefano M (2023) CompletionFormer: Depth Completion with Convolutions and Vision Transformers. arXiv

BibTeX

@article{youmin2023completionformer,
  title = {CompletionFormer: Depth Completion with Convolutions and Vision Transformers},
  author = {Youmin, Zhang and Xianda, Guo and Matteo, Poggi and Zheng, Zhu and Guan, Huang and Stefano, Mattoccia},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.13030v1},
  eprint = {2304.13030}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE