Local Similarity Pattern and Cost Self-Reassembling for Deep Stereo Matching Networks

Biyang LiuHuimin YuYangqi Long

article2022AAAI85 citations

Proposes a pairwise Local Similarity Pattern to inject structural information into deep features alongside a dynamic self-reassembling refinement strategy that prevents over-smoothing in stereo matching disparity estimation.

Listen

Stereo matching is a core computer vision technology that estimates depth from paired camera images, providing a low-cost visual foundation for critical systems such as autonomous vehicles and augmented reality. While modern deep learning architectures have significantly advanced stereo matching performance, they suffer from two key shortcomings: standard convolutional features focus heavily on visual appearance while overlooking essential geometric structures, and standard post-processing refinement filters cause excessive smoothing that blurs sharp object boundaries and distorts occluded areas.

The article designs and evaluates two modular enhancements to address these flaws: the Local Similarity Pattern (LSP), which injects explicit geometric structure into feature extraction, and Cost Self-Reassembling (CSR), an adaptive refinement strategy that sharpens depth estimates by dynamically propagating reliable measurements from neighboring pixels.

To demonstrate the effectiveness and credibility of these modules, the authors integrated them into leading baseline architectures, including GwcNet and GANet-deep. They evaluated the systems across major standard benchmarks, including the large-scale synthetic SceneFlow dataset (over 39,000 image pairs), the real-world KITTI autonomous driving benchmarks, and cross-domain generalization datasets such as Middlebury and ETH3D.

The findings confirm clear, practical improvements across all tested configurations. First, integrating the full multi-scale and multi-level Local Similarity Pattern consistently boosted accuracy over standard convolutional features alone by making matching robust across different object scales. Second, applying dynamic refinement via Cost Self-Reassembling substantially reduced errors and avoided the over-smoothing seen in conventional convolutional post-processing, dropping outlier percentages significantly. Third, combining both modules on the baseline GwcNet model improved endpoint error on SceneFlow from 1.04 to 0.75 pixels (an improvement of nearly 28%) and reduced outlier rates on the KITTI 2015 validation set from 1.65% to 1.30%. Finally, the modules demonstrated strong cross-domain transferability under varying lighting conditions, with minimal parameter overhead.

These results demonstrate that blending classical geometric principles with deep learning architectures delivers sharper, more reliable depth maps. This enhancement directly improves downstream safety, reliability, and obstacle detection in autonomous navigation. The article also provides a practical alternative, Disparity Self-Reassembling (DSR), which achieves competitive accuracy gains with negligible memory overhead compared to the higher memory footprint of Cost Self-Reassembling.

For engineering and operational deployment, teams building real-time or resource-constrained embedded systems should pair lightweight baseline backbones with the Disparity Self-Reassembling variant to balance throughput, memory, and precision. Organizations focusing on high-accuracy offline 3D modeling or safety-critical perception pipelines can deploy the full Cost Self-Reassembling model. Next development steps should focus on extending this dynamic neighbor-reassembling strategy to broader pixel-level tasks such as semantic segmentation and testing performance under severe weather or extreme illumination shifts.

While confidence in the reported results is high due to comprehensive testing across standard synthetic and real-world benchmarks, operational teams should note that real-world benchmarks like KITTI provide relatively sparse ground truth for evaluation. Practitioners must carefully weigh the memory demands of the full cost-level refinement against target embedded hardware limits before deploying at scale.

arXiv: 2210.12785
Cover for Local Similarity Pattern and Cost Self-Reassembling for Deep Stereo Matching Networks

Abstract

Although convolutional neural network based stereo matching architectures have made impressive achievements, there are still some limitations: 1) Convolutional Feature (CF) tends to capture appearance information, which is inadequate for accurate matching. 2) Due to the static filters, current convolution based disparity refinement modules often produce over-smooth results. In this paper, we present two schemes to address these issues, where some traditional wisdoms are integrated. Firstly, we introduce a pairwise feature for deep stereo matching networks, named LSP (Local Similarity Pattern). Through explicitly revealing the neighbor relationships, LSP contains rich structural information, which can be leveraged to aid CF for more discriminative feature description. Secondly, we design a dynamic self-reassembling refinement strategy and apply it to the cost distribution and the disparity map respectively. The former could be equipped with the unimodal distribution constraint to alleviate the over-smoothing problem, and the latter is more practical. The effectiveness of the proposed methods is demonstrated via incorporating them into two well-known basic architectures, GwcNet and GANet-deep. Experimental results on the SceneFlow and KITTI benchmarks show that our modules significantly improve the performance of the model. Code is available at https://github.com/SpadeLiu/Lac-GwcNet.

Table of Contents

  • Introduction
  • Related Works
  • Deep Stereo Matching Networks
  • Structural Features
  • Disparity Refinement
  • Method
  • Overall Architecture
  • Local Similarity Pattern
  • Cost Self-Reassembling
  • Loss Function
  • Experiments Datasets & Evaluation Metrics
  • Implementation Details
  • Ablation Studies
  • Complexity Analysis
  • Generalization Ability
  • Benchmarks
  • Conclusions
  • References

Knowls

  1. Knowl 1 — Local Similarity Pattern feature

    model/method

    The Local Similarity Pattern (LSP) is a trainable pairwise feature designed to supplement convolutional features (CF) with local structural information. Let fC(x,y)∈RCf_C(x,y)\in\mathbb{R}^{C} be the CC-channel convolutional feature at pixel coordinates (x,y)(x,y), and let (δxk,δyk)(\delta x_k,\delta y_k) be the offset of the kk-th neighbor in a square neighborhood. LSP channel kk is

    fLk(x,y)=Φ(fC(x,y),fC(x+δxk,y+δyk)),f_L^k(x,y)=\Phi\bigl(f_C(x,y),f_C(x+\delta x_k,y+\delta y_k)\bigr),

    where fLk(x,y)f_L^k(x,y) is a scalar relationship between the central feature and its neighbor, and Φ\Phi is cosine similarity. If the neighborhood contains KK points, the LSP has KK channels, independently of the number CC of convolutional-feature channels. LSP is computed directly in a bottom-up fashion, so it does not require stacking multiple learned layers to recognize local patterns.

    The paper extends LSP in two ways. A multi-scale LSP uses dilated square neighborhoods to capture structures at different spatial scopes without the detail loss associated with pooling. A multi-level LSP computes pairwise relationships separately from convolutional features at different network layers, thereby combining low-level texture information with higher-level semantic information. Each LSP is channel-adjusted with a 1×11\times1 convolution, concatenated with the relevant CF, and passed through another 1×11\times1 fusion convolution. The resulting representation can jointly use appearance information from CF and neighbor-relationship information from LSP.

  2. Knowl 2 — Dynamic cost and disparity self-reassembling

    model/method

    The paper introduces dynamic self-reassembling refinement to replace fixed-grid convolutional disparity correction. A lightweight U-Net takes the left image as input and predicts, for every pixel (x,y)(x,y), NN two-dimensional offsets (Δxi,Δyi)(\Delta x_i,\Delta y_i) identifying content-adaptive neighbors. Cost Self-Reassembling (CSR) samples the initial cost distribution C0C_0 at those locations and averages the samples:

    Cr(x,y)=1N∑i=1NC0(x+Δxi,y+Δyi),C_r(x,y)=\frac{1}{N}\sum_{i=1}^{N}C_0(x+\Delta x_i,y+\Delta y_i),

    where CrC_r and C0C_0 are refined and initial cost distributions, respectively, each with dmax⁡d_{\max} disparity channels; NN is the number of assembled neighbors; and fractional coordinates are sampled by bilinear interpolation. The paper also predicts a modulation factor mi∈[0,1]m_i\in[0,1] for each neighbor using a sigmoid output and uses the weighted version

    Cr(x,y)=1∑i=1Nmi∑i=1Nmi C0(x+Δxi,y+Δyi).C_r(x,y)=\frac{1}{\sum_{i=1}^{N}m_i}\sum_{i=1}^{N}m_i\,C_0(x+\Delta x_i,y+\Delta y_i).

    The same operation applied after disparity regression defines Disparity Self-Reassembling (DSR):

    Dr(x,y)=1∑i=1Nmi∑i=1Nmi D0(x+Δxi,y+Δyi),D_r(x,y)=\frac{1}{\sum_{i=1}^{N}m_i}\sum_{i=1}^{N}m_i\,D_0(x+\Delta x_i,y+\Delta y_i),

    where D0D_0 and DrD_r are single-channel initial and refined disparity maps. CSR permits a unimodal-distribution constraint to be applied to the refined cost distribution through cross-entropy supervision, which the paper uses to reduce over-smoothing. DSR is more practical because it stores and reassembles scalar disparities rather than a full dmax⁡d_{\max}-dimensional cost vector, whereas CSR provides the additional distribution-level constraint.

  3. Knowl 3 — Integration into deep stereo networks

    model/method

    The proposed modules are inserted into an end-to-end stereo matching pipeline. Shared-weight feature extractors process the rectified left and right images; single-level or multi-level LSP is computed from the resulting convolutional features and fused with CF. In the authors' GwcNet-based implementation, the left and right fused features are concatenated to form the cost volume, rather than using group-wise correlation. Several 3D convolutions aggregate the cost volume, softmax produces a disparity distribution, and weighted averaging regresses the initial disparity.

    The refinement stage then applies CSR to the cost distribution before disparity regression, or applies DSR directly to the regressed disparity map. The same LSP and self-reassembling modules are also inserted into GANet-deep. Thus, the contribution is a feature representation that exposes local structural relationships and a refinement mechanism that propagates values between image-content-selected locations rather than only along a regular convolutional grid.

  4. Knowl 4 — Multi-output training objective

    equation

    The stereo model is supervised both at the disparity-distribution level and at the final disparity level. For an image with II pixels and disparity candidates d∈{0,…,dmax⁡}d\in\{0,\ldots,d_{\max}\}, the paper uses

    Lce(P^,P)=1I∑i=1I∑d=0dmax⁡−P^i(d)log⁡Pi(d),L_{\mathrm{ce}}(\widehat P,P)=\frac{1}{I}\sum_{i=1}^{I}\sum_{d=0}^{d_{\max}}-\widehat P_i(d)\log P_i(d),

    and

    Lsm(D^,D)=1I∑i=1IsmoothL1⁡(D^i,Di).L_{\mathrm{sm}}(\widehat D,D)=\frac{1}{I}\sum_{i=1}^{I}\operatorname{smoothL1}(\widehat D_i,D_i).

    Here P^i(d)\widehat P_i(d) is the predicted softmax distribution at pixel ii, Pi(d)P_i(d) is the normalized Laplacian ground-truth distribution centered at the ground-truth disparity DiD_i, D^i\widehat D_i is the predicted disparity obtained by weighted averaging over disparity candidates, and dmax⁡d_{\max} is the maximum disparity. The total objective sums the two losses over the MM supervised outputs:

    L=∑m=1Mλm(Lce(m)+μLsm(m)),L=\sum_{m=1}^{M}\lambda_m\bigl(L_{\mathrm{ce}}^{(m)}+\mu L_{\mathrm{sm}}^{(m)}\bigr),

    where λm\lambda_m weights output mm and μ\mu balances disparity regression against distribution supervision. In the GwcNet-based model, M=4M=4: the three hourglass outputs and the refinement output are supervised.

  5. Knowl 5 — Experimental protocol and evaluation measures

    experimental setup

    The main experiments use SceneFlow, KITTI 2012, and KITTI 2015. SceneFlow contains 35,454 training pairs and 4,370 test pairs with dense ground truth; it is evaluated with EPE, 1N∑i∣d^i−di∗∣\frac{1}{N}\sum_i|\widehat d_i-d_i^*|, and the percentage of pixels whose absolute disparity error exceeds 1 pixel. KITTI 2012 contains 194 training and 195 test pairs, while KITTI 2015 contains 200 training and 200 test pairs; the authors split KITTI 2015 into 160 training and 40 validation pairs for experiments. KITTI evaluation uses EPE and D1, the percentage of pixels whose error exceeds both 3 pixels and 5% of the ground-truth disparity. Middlebury 2014 and ETH3D are used only for generalization evaluation because each has fewer than 50 training pairs; their reported metrics are Bad2.0 and Bad1.0, respectively.

    The models are implemented in PyTorch and optimized with Adam using β1=0.9\beta_1=0.9 and β2=0.999\beta_2=0.999. SceneFlow training lasts 10 epochs at learning rate 0.0010.001. KITTI fine-tuning from the SceneFlow model lasts 300 epochs, using learning rate 0.0010.001 for the first 200 epochs and 0.00010.0001 thereafter; the KITTI benchmark submission is trained for 600 epochs with learning rate initialized to 0.0010.001 and divided by 10 every 200 epochs. Training crops are 512×256512\times256, the batch size is 4 on two TITAN Xp GPUs, and the maximum disparity is 192. LSP uses a 3×33\times3 neighbor window with dilation rates 1,2,4,81,2,4,8. The smooth-L1 balance is μ=0.1\mu=0.1, the refined-output weight is 11, and the remaining output weights follow the basic model.

  6. Knowl 6 — Ablation of LSP and self-reassembling refinement

    data/table

    The ablation compares a GwcNet-based baseline using feature concatenation on the SceneFlow test set and KITTI 2015 validation set. Lower values are better. CF denotes convolutional features; SS and SL denote single-scale and single-level LSP; F denotes the full multi-scale, multi-level LSP; ConvNet is the residual-prediction refinement baseline; and CSR and DSR are the proposed cost- and disparity-level refiners.

    Could not parse LaTeX table

    Every LSP variant improves the baseline, and the full multi-scale, multi-level variant is strongest among the LSP-only settings. The ConvNet residual refiner lowers EPE but increases both outlier measures, consistent with over-smoothing. DSR and CSR improve both EPE and outlier percentages, while the complete LSP-plus-CSR model achieves the best ablation result: relative to CF alone, SceneFlow EPE falls from 1.041.04 to 0.750.75 pixels and KITTI D1-all falls from 1.65%1.65\% to 1.30%1.30\%.

  7. Knowl 7 — Effect of the number of reassembled neighbors

    data/table

    The authors evaluate the number NN of assembled points in CSR on SceneFlow. Increasing NN improves EPE but increases GPU memory use, so the selected operating point balances accuracy and resource cost.

    Could not parse LaTeX table

    The model uses N=2N=2 in the reported experiments: it reduces EPE from 0.950.95 at N=0N=0 to 0.750.75 while avoiding the substantially larger memory costs of N=4N=4 and N=8N=8.

  8. Knowl 8 — Benchmark performance on SceneFlow and KITTI

    empirical result

    The proposed LaC system, comprising LSP and self-reassembling refinement, improves strong stereo baselines on the SceneFlow and KITTI benchmarks. On SceneFlow, the reported EPE values are:

    Could not parse LaTeX table

    On the KITTI online leaderboards, the reported values are percentages except for runtime in seconds. D1-bg, D1-fg, and D1-all are reported for all KITTI 2015 pixels and for non-occluded pixels; Out-Noc and Out-All are reported for KITTI 2012.

    Could not parse LaTeX table

    Relative to GwcNet, LaC lowers SceneFlow EPE by 0.230.23 pixels and KITTI 2015 D1-all by 0.340.34 percentage points. Relative to GANet, the improvements are 0.080.08 pixels and 0.140.14 percentage points, respectively. The paper notes that the smaller gain over GANet may result from GANet's guided aggregation already providing image-content-based propagation.

  9. Knowl 9 — Computational complexity and memory trade-offs

    data/table

    Complexity is measured with 960×540960\times540 input images using the GwcNet-based model. FLOPs are used as a device-independent proxy for computation time.

    Could not parse LaTeX table

    LSP adds little computational or memory overhead. DSR adds moderate parameters and FLOPs while increasing memory from 3.283.28 G to 3.453.45 G. CSR has nearly the same parameter count and FLOPs as DSR but raises memory use to 5.515.51 G because each assembled neighbor requires storing a full cost vector rather than one disparity value. This memory difference is the paper's main practical reason for recommending DSR in resource-constrained algorithms.

  10. Knowl 10 — Cross-dataset generalization

    empirical result

    To test domain generalization, the authors train all models on SceneFlow and evaluate them on the training images of Middlebury 2014 and ETH3D without dataset-specific training. The results are:

    Could not parse LaTeX table

    Both proposed components improve performance under the domain shift. The paper attributes CSR's robustness partly to its use of the input image rather than matching-network features, which may be more domain-sensitive. It attributes LSP's improvement, especially on Middlebury, to the structural relationships being useful when the stereo pair has inconsistent illumination.

Coverage note — No substantial contributed material was omitted; future-work suggestions and qualitative visualizations were not separated into knowls because their substantive claims are covered by the quantitative method and generalization results.

References

  1. 1.Chabra, R.; Straub, J.; Sweeney, C.; Newcombe, R.; and Fuchs, H. 2019. Stereodrnet: Dilated residual stereonet. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11786–11795.
  2. 2.Chang, J.-R.; and Chen, Y.-S. 2018. Pyramid stereo matching network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5410–5418.
  3. 3.Chen, C.; Chen, X.; and Cheng, H. 2019. On the over-smoothing problem of cnn based disparity estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8997–9005.
  4. 4.Chen, Z.; Sun, X.; Wang, L.; Yu, Y.; and Huang, C. 2015. A deep visual correspondence embedding model for stereo matching costs. In Proceedings of the IEEE International Conference on Computer Vision, 972–980.
  5. 5.Cheng, X.; Wang, P.; and Yang, R. 2019. Learning depth with convolutional spatial propagation network. IEEE transactions on pattern analysis and machine intelligence, 42(10): 2361–2379.
  6. 6.Cheng, X.; Zhong, Y.; Harandi, M.; Dai, Y.; Chang, X.; Li, H.; Drummond, T.; and Ge, Z. 2020. Hierarchical Neural Architecture Search for Deep Stereo Matching. Advances in Neural Information Processing Systems, 33.
  7. 7.Dai, J.; Qi, H.; Xiong, Y.; Li, Y.; Zhang, G.; Hu, H.; and Wei, Y. 2017. Deformable convolutional networks. In Proceedings of the IEEE international conference on computer vision, 764–773.
  8. 8.Duggal, S.; Wang, S.; Ma, W.-C.; Hu, R.; and Urtasun, R. 2019. Deeppruner: Learning efficient stereo matching via differentiable patchmatch. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4384–4393.
  9. 9.Gan, Y.; Xu, X.; Sun, W.; and Lin, L. 2018. Monocular depth estimation with affinity, vertical pooling, and label enhancement. In Proceedings of the European Conference on Computer Vision (ECCV), 224–239.
  10. 10.Garg, D.; Wang, Y.; Hariharan, B.; Campbell, M.; Weinberger, K. Q.; and Chao, W.-L. 2020. Wasserstein distances for stereo disparity estimation. arXiv preprint arXiv:2007.03085.
  11. 11.Geiger, A.; Lenz, P.; and Urtasun, R. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, 3354–3361. IEEE.
  12. 12.Guo, X.; Yang, K.; Yang, W.; Wang, X.; and Li, H. 2019. Group-Wise Correlation Stereo Network. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  13. 13.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 37(9): 1904–1916.
  14. 14.Hu, H.; Zhang, Z.; Xie, Z.; and Lin, S. 2019. Local relation networks for image recognition. In Proceedings of the IEEE International Conference on Computer Vision, 3464–3473.
  15. 15.Huq, S.; Koschan, A.; and Abidi, M. 2013. Occlusion filling in stereo: Theory and experiments. Computer Vision and Image Understanding, 117(6): 688–704.
  16. 16.Kendall, A.; Martirosyan, H.; Dasgupta, S.; Henry, P.; Kennedy, R.; Bachrach, A.; and Bry, A. 2017. End-to-end learning of geometry and context for deep stereo regression. In Proceedings of the IEEE International Conference on Computer Vision, 66–75.
  17. 17.Khamis, S.; Fanello, S.; Rhemann, C.; Kowdle, A.; Valentin, J.; and Izadi, S. 2018. Stereonet: Guided hierarchical refinement for real-time edge-aware depth prediction. In Proceedings of the European Conference on Computer Vision (ECCV), 573–590.
  18. 18.Liang, Z.; Feng, Y.; Guo, Y.; Liu, H.; Chen, W.; Qiao, L.; Zhou, L.; and Zhang, J. 2018. Learning for disparity estimation through feature constancy. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2811–2820.
  19. 19.Luo, W.; Schwing, A. G.; and Urtasun, R. 2016. Efficient deep learning for stereo matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5695–5703.
  20. 20.Mayer, N.; Ilg, E.; Hausser, P.; Fischer, P.; Cremers, D.; Dosovitskiy, A.; and Brox, T. 2016. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4040–4048.
  21. 21.Mei, X.; Sun, X.; Zhou, M.; Jiao, S.; Wang, H.; and Zhang, X. 2011. On building an accurate stereo matching system on graphics hardware. In 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops), 467–474. IEEE.
  22. 22.Menze, M.; and Geiger, A. 2015. Object scene flow for autonomous vehicles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3061–3070.
  23. 23.Ojala, T.; Pietikainen, M.; and Maenpaa, T. 2002. Multiresolution gray-scale and rotation invariant texture classification with local binary patterns. IEEE Transactions on pattern analysis and machine intelligence, 24(7): 971–987.
  24. 24.Ramamonjisoa, M.; Du, Y.; and Lepetit, V. 2020. Predicting sharp and accurate occlusion boundaries in monocular depth estimation using displacement fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14648–14657.
  25. 25.Ronneberger, O.; Fischer, P.; and Brox, T. 2015. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, 234–241. Springer.
  26. 26.Scharstein, D.; Hirschmuller, H.; Kitajima, Y.; Krathwohl, G.; Nesic, N.; Wang, X.; and Westling, P. 2014. High-resolution stereo datasets with subpixel-accurate ground truth. In German conference on pattern recognition, 31–42. Springer.
  27. 27.Scharstein, D.; and Szeliski, R. 2002. A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms. International Journal of Computer Vision, 47(1-3): 7–42.
  28. 28.Schops, T.; Schönberger, J. L.; Galliani, S.; Sattler, T.; Schindler, K.; Pollefeys, M.; and Geiger, A. 2017. A Multi-View Stereo Benchmark with High-Resolution Images and Multi-Camera Videos. In Conference on Computer Vision and Pattern Recognition (CVPR).
  29. 29.Shen, Z.; Dai, Y.; and Rao, Z. 2021. CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13906–13915.
  30. 30.Song, X.; Zhao, X.; Fang, L.; Hu, H.; and Yu, Y. 2020. Edgestereo: An effective multi-task learning network for stereo matching and edge detection. International Journal of Computer Vision, 1–21.
  31. 31.Tankovich, V.; Hane, C.; Zhang, Y.; Kowdle, A.; Fanello, S.; and Bouaziz, S. 2021. Hitnet: Hierarchical iterative tile refinement network for real-time stereo matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14362–14372.
  32. 32.Tosi, F.; Poggi, M.; and Mattoccia, S. 2019. Leveraging confident points for accurate depth refinement on embedded systems. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 0–0.
  33. 33.Tulyakov, S.; Ivanov, A.; and Fleuret, F. 2018. Practical deep stereo (pds): Toward applications-friendly deep stereo matching. In Advances in Neural Information Processing Systems, 5871–5881.
  34. 34.Wu, Z.; Wu, X.; Zhang, X.; Wang, S.; and Ju, L. 2019. Semantic Stereo Matching with Pyramid Cost Volumes. In Proceedings of the IEEE International Conference on Computer Vision, 7484–7493.
  35. 35.Xu, H.; and Zhang, J. 2020. AANet: Adaptive Aggregation Network for Efficient Stereo Matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1959–1968.
  36. 36.Yu, F.; and Koltun, V. 2015. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122.
  37. 37.Zabih, R.; and Woodfill, J. 1994. Non-parametric local transforms for computing visual correspondence. In European conference on computer vision, 151–158. Springer.
  38. 38.Zbontar, J.; and LeCun, Y. 2015. Computing the stereo matching cost with a convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1592–1599.
  39. 39.Zhang, F.; Prisacariu, V.; Yang, R.; and Torr, P. H. 2019. GANet: Guided Aggregation Net for End-to-end Stereo Matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 185–194.
  40. 40.Zhang, F.; Qi, X.; Yang, R.; Prisacariu, V.; Wah, B.; and Torr, P. 2020a. Domain-invariant stereo matching networks. In European Conference on Computer Vision, 420–439. Springer.
  41. 41.Zhang, Y.; Chen, Y.; Bai, X.; Yu, S.; Yu, K.; Li, Z.; and Yang, K. 2020b. Adaptive unimodal cost volume filtering for deep stereo matching. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 12926–12934.
  42. 42.Zhu, X.; Hu, H.; Lin, S.; and Dai, J. 2019. Deformable convnets v2: More deformable, better results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9308–9316.

Citation

MLA
Jiang, H., et al. “An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022”. arXiv, 2022, http://arxiv.org/abs/2210.12785v1.
APA
Jiang, H., Xu, R., & Jiang, W. (2022). An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022. arXiv. http://arxiv.org/abs/2210.12785v1
Chicago
Jiang, H., R. Xu, and W. Jiang. 2022. “An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022”. arXiv. http://arxiv.org/abs/2210.12785v1.
Harvard
Jiang, H., Xu, R. and Jiang, W. (2022) “An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.12785v1.
Vancouver
1. Jiang H, Xu R, Jiang W (2022) An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022. arXiv

BibTeX

@article{jiang2022improved,
  title = {An Improved RaftStereo Trained with A Mixed Dataset for the Robust Vision Challenge 2022},
  author = {Jiang, Hualie and Xu, Rui and Jiang, Wenjie},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.12785v1},
  eprint = {2210.12785}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF