Surface Reconstruction from Point Clouds by Learning Predictive Context Priors

Baorui MaYu-Shen LiuMatthias ZwickerZhizhong Han

article2022CVPR89 citations

Presents a surface reconstruction framework that learns test-time predictive query displacements to flexibly query pre-trained local signed distance functions, achieving accurate 3D geometry without requiring ground truth normals or distances.

Listen

Surface reconstruction from 3D point clouds is an essential task across 3D computer vision, autonomous systems, and digital modeling. Modern state-of-the-art approaches train neural networks to learn local geometric shape representations, called signed distance functions, across large datasets. However, these methods typically restrict spatial queries to fixed target coordinates during reconstruction, severely limiting the model's ability to generalize to novel, unseen geometries and complex scenes.

The article demonstrates a novel framework called Predictive Context Priors to overcome this generalization bottleneck. The primary objective is to evaluate whether predicting flexible, adjusted spatial query locations at inference time enables neural networks to match learned geometric shapes more accurately to new point clouds without requiring ground truth surface normals or signed distance annotations.

The researchers developed a two-stage approach. First, a neural network learns a local context prior across partitioned local regions from a large point cloud dataset using a pulling optimization objective, which moves sampled query points toward the true surface based on predicted distances and gradients. Second, when reconstructing a specific target point cloud at test time, the local prior is frozen, and an additional query network is trained on that specific shape to predict spatial query displacements and condition features. This procedure allows the model to flexibly search the entire learned prior space. The framework was evaluated across standard single-shape datasets, such as ShapeNet, ABC, and FAMOUS, as well as complex multi-object scene datasets, including 3D Scene and SceneNet.

The experimental findings show significant performance advantages over existing methods. Across single shapes in the ShapeNet benchmark, the proposed method achieved a Chamfer distance error roughly 60% lower than the closest competitive neural implicit method (Neural-pull) and reduced error by over 90% compared to traditional spatial grid methods. On the FAMOUS dataset, the method cut surface reconstruction error by approximately 80% compared to Neural-pull and by over 95% compared to Points2Surf. In complex scene evaluations, the framework consistently achieved superior normal consistency and the lowest average metric distance errors, such as cutting mean vertex error down to 6.33–9.28 millimeters across challenging indoor scenes. Furthermore, tests demonstrated strong robustness against substantial sensor noise and sparse sampling densities.

These results show that relaxing fixed spatial coordinate constraints allows learned geometric priors to adapt to complex and irregular structures with high precision. By eliminating the requirement for pre-computed surface normals or ground truth distance supervision, this method lowers data preparation costs and simplifies 3D processing pipelines for real-world scans.

Organizations handling high-precision 3D scanning, digital twin generation, or CAD reconstruction should evaluate the adoption of predictive query techniques within their reconstruction workflows. Implementation teams should weigh the trade-off between higher geometric accuracy and computational runtime, as the per-shape optimization during inference requires more processing time than direct feed-forward models. Future efforts should focus on optimizing this test-time inference process to accelerate reconstruction speeds for time-critical operational deployments.

The findings are supported with high confidence across standard academic benchmarks and diverse shape categories. Readers should note that performance depends on the test-time optimization phase, meaning real-time or low-latency operational environments may face computational bottlenecks until further inference speed optimizations are introduced.

Cover for Surface Reconstruction from Point Clouds by Learning Predictive Context Priors

Abstract

Surface reconstruction from point clouds is vital for 3D computer vision. State-of-the-art methods leverage large datasets to first learn local context priors that are represented as neural network-based signed distance functions (SDFs) with some parameters encoding the local contexts. To reconstruct a surface at a specific query location at inference time, these methods then match the local reconstruction target by searching for the best match in the local prior space (by optimizing the parameters encoding the local context) at the given query location. However, this requires the local context prior to generalize to a wide variety of unseen target regions, which is hard to achieve. To resolve this issue, we introduce Predictive Context Priors by learning Predictive Queries for each specific point cloud at inference time. Specifically, we first train a local context prior using a large point cloud dataset similar to previous techniques. For surface reconstruction at inference time, however, we specialize the local context prior into our Predictive Context Prior by learning Predictive Queries, which predict adjusted spatial query locations as displacements of the original locations. This leads to a global SDF that fits the specific point cloud the best. Intuitively, the query prediction enables us to flexibly search the learned local context prior over the entire prior space, rather than being restricted to the fixed query locations, and this improves the generalizability. Our method does not require ground truth signed distances, normals, or any additional procedure of signed distance fusion across overlapping regions. Our experimental results in surface reconstruction for single shapes or complex scenes show significant improvements over the state-of-the-art under widely used benchmarks. Code and data are available at https://github.com/mabaorui/PredictableContextPrior.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 4. Experiments, Analysis, and Discussion
  • 4.1. Setup
  • 4.2. Surface Reconstruction for Single Shapes
  • 4.3. Surface Reconstruction for Scenes
  • 4.4. Analysis and Discussion
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Predictive Context Prior Formulation for Neural Signed Distance Functions

    model/method

    The Predictive Context Prior framework reconstructs continuous 3D surfaces from unoriented point clouds GG by separating shape representation into pre-training a local context prior and specializing it at test time via learned predictive queries. Traditional local implicit approaches fix query points in local coordinate systems and optimize latent codes (via auto-decoding) to match the target geometry, requiring the pre-trained prior to cover every possible local geometric configuration at fixed query coordinates.

    In contrast, the Predictive Context Prior specializes a fixed, pre-trained local Signed Distance Function (SDF) FF by training a shape-specific neural query network parameterized by θ3\theta_3. For any spatial query point qg∈R3q_g \in \mathbb{R}^3 sampled in the global coordinate frame of the target point cloud GG, the query network predicts a spatial displacement Δq∈R3\Delta q \in \mathbb{R}^3 and a local condition vector fl′∈R512f'_l \in \mathbb{R}^{512}. The predicted local query is given by:

    ql′=qg+Δqq'_l = q_g + \Delta q

    The global SDF Fg(qg,G)F_g(q_g, G) is evaluated by feeding the transformed query and predicted condition into the frozen local SDF network:

    s′=Fg(qg,G)=F(ql′,fl′)s' = F_g(q_g, G) = F(q'_l, f'_l)

    This formulation enables the model to flexibly search the entire continuous prior space of the pre-trained local SDF rather than restricting the search to fixed query locations, thereby overcoming local generalization bottlenecks without requiring ground-truth signed distances, surface normals, or post-hoc blending across overlapping partitions.

  2. Knowl 2 — Local Context Prior Training Objective and Architecture

    equation

    The local context prior is learned across a training collection of shapes divided into local point cloud regions ti∈Tt_i \in T. Each local region tit_i is normalized by centering its points at the origin and scaling them linearly so the longest bounding-box dimension lies in [−0.5,0.5][-0.5, 0.5].

    A feature encoder (PointNet parameterized by θ1\theta_1) extracts a 512-dimensional local condition fl∈R512f_l \in \mathbb{R}^{512} for each region tit_i. A neural implicit network (an Occupancy Network architecture without activation on the final output layer, parameterized by θ2\theta_2) represents the local SDF s=F(ql,fl)s = F(q_l, f_l), where ql∈R3q_l \in \mathbb{R}^3 is a local spatial query.

    Without requiring ground-truth signed distances or surface normals, θ1\theta_1 and θ2\theta_2 are jointly optimized using a pulling cost CPull1C^1_{\mathrm{Pull}}:

    min⁡θ1,θ2∥nn(ql)−(ql−s∇s∥∇s∥22)∥2\min_{\theta_1, \theta_2} \left\| \mathrm{nn}(q_l) - \left( q_l - s \frac{\nabla s}{\|\nabla s\|_2^2} \right) \right\|_2

    where nn(ql)∈ti\mathrm{nn}(q_l) \in t_i denotes the Euclidean nearest neighbor of qlq_l in the point set tit_i, and ∇s=∂F∂ql\nabla s = \frac{\partial F}{\partial q_l} is the analytical gradient of the predicted signed distance with respect to the query location qlq_l.

  3. Knowl 3 — Test-Time Query Specialization Objective and Network Architecture

    equation

    At inference time, the pre-trained local SDF network parameters θ2\theta_2 are frozen. For an unnormalized target point cloud GG in the global coordinate frame, an 8-layer feed-forward neural query network parameterized by θ3\theta_3 is trained to map global query coordinates qg∈R3q_g \in \mathbb{R}^3 to the local prior domain. Each of the first 7 layers contains 512 nodes with ReLU activations, and the 8th layer has 515 nodes with linear activations, outputting the displacement Δq∈R3\Delta q \in \mathbb{R}^3 (giving ql′=qg+Δqq'_l = q_g + \Delta q) and the latent condition fl′∈R512f'_l \in \mathbb{R}^{512}.

    The query network parameters θ3\theta_3 are optimized by minimizing the test pulling cost CPull2C^2_{\mathrm{Pull}}:

    min⁡θ3∥nn(qg)−(qg−s′∇s′∥∇s′∥22)∥2\min_{\theta_3} \left\| \mathrm{nn}(q_g) - \left( q_g - s' \frac{\nabla s'}{\|\nabla s'\|_2^2} \right) \right\|_2

    where nn(qg)∈G\mathrm{nn}(q_g) \in G is the nearest neighbor to qgq_g on the target point cloud GG, s′=F(ql′,fl′)s' = F(q'_l, f'_l) is the signed distance produced by the frozen implicit network, and ∇s′=∂F∂ql′\nabla s' = \frac{\partial F}{\partial q'_l} is the gradient of FF with respect to the predicted local query ql′q'_l.

  4. Knowl 4 — Surface Reconstruction via Predictive Context Prior Optimization

    algorithm

    The complete surface reconstruction pipeline optimizes the predictive query network θ3\theta_3 on a target point cloud GG and extracts the explicit polygonal mesh via Marching Cubes.

    Input: Target point cloud G⊂R3G \subset \mathbb{R}^3, pre-trained local SDF network F(⋅;θ2)F(\cdot; \theta_2)
    Output: Triangle mesh M=(V,E)M = (V, E)
    Initialize neural query network Q(⋅;θ3)Q(\cdot; \theta_3) with parameters θ3\theta_3
    Freeze parameters θ2\theta_2
    for each optimization epoch do
        Sample NqN_q query points qgq_g around each point p∈Gp \in G using Gaussian sampling N(p,σ2)\mathcal{N}(p, \sigma^2)
        Find nearest neighbors nn(qg)∈G\mathrm{nn}(q_g) \in G for each query
        Predict local offsets and conditions: (Δq,fl′)←Q(qg;θ3)(\Delta q, f'_l) \leftarrow Q(q_g; \theta_3)
        Compute local queries: ql′←qg+Δqq'_l \leftarrow q_g + \Delta q
        Compute signed distances: s′←F(ql′,fl′;θ2)s' \leftarrow F(q'_l, f'_l; \theta_2)
        Compute gradients: ∇s′←∂F∂ql′\nabla s' \leftarrow \frac{\partial F}{\partial q'_l}
        Compute pulled query positions: qpull←qg−s′∇s′∥∇s′∥22q_{\mathrm{pull}} \leftarrow q_g - s' \frac{\nabla s'}{\|\nabla s'\|_2^2}
        Compute loss: L←∥nn(qg)−qpull∥2L \leftarrow \| \mathrm{nn}(q_g) - q_{\mathrm{pull}} \|_2
        Update θ3\theta_3 via gradient descent to minimize LL
    end for
    Define continuous global field: Fg(x)=F(x+QΔq(x),Qf(x);θ2)F_g(x) = F(x + Q_{\Delta q}(x), Q_f(x); \theta_2)
    Extract mesh M←MarchingCubes(Fg,resolution)M \leftarrow \mathrm{MarchingCubes}(F_g, \mathrm{resolution})
    return MM

    In standard execution, Nq=40N_q = 40 queries per point are sampled where σ2\sigma^2 is set to the squared distance to the 50-th nearest neighbor of pp in GG. The marching cubes grid resolution is set to 1283128^3 for single objects or 5123512^3 for complex scenes.

  5. Knowl 5 — Surface Reconstruction Accuracy on ShapeNet Benchmark

    data/table

    The Predictive Context Prior method was evaluated on a subset of the ShapeNet benchmark comprising 8 object categories (Display, Lamp, Airplane, Cabinet, Vessel, Table, Chair, Sofa) using L2 Chamfer Distance (L2L_2-CD ×100\times 100), Normal Consistency (NC), and F-score thresholds (μ=0.002\mu = 0.002 and 2μ=0.0042\mu = 0.004). Reconstructions were generated using Marching Cubes on a 1283128^3 grid and compared against 100,000 ground truth surface points.

    Method L2L_2-CD (×100)↓(\times 100) \downarrow NC ↑\uparrow F-score(μ\mu) ↑\uparrow F-score(2μ2\mu) ↑\uparrow
    PSR 0.286 0.866 0.407 0.618
    DMC 0.276 0.836 0.446 0.626
    BPA 0.116 0.923 0.769 0.885
    AtlasNet 1.368 0.695 0.062 0.158
    DeepSDF 0.766 0.884 0.468 0.647
    DGP 0.228 0.962 0.425 0.649
    MeshP 0.071 0.903 0.872 0.959
    SALD 0.062 0.666 0.429 0.950
    LIG - 0.939 0.961 0.976
    Neural-Pull 0.038 - - -
    Ours 0.0136 0.9590 0.9871 0.9899

    The Predictive Context Prior achieves the lowest Chamfer Distance and highest F-score across all categories, substantially outperforming prior implicit function baselines (e.g., DeepSDF, SALD, and Neural-Pull) without requiring ground-truth distance supervision or normal vectors.

  6. Knowl 6 — Reconstruction Robustness Under Noise and Point Density Variations

    data/table

    The model's robustness was evaluated on the ABC and FAMOUS datasets under clean point clouds, variable noise (var-noise), maximum noise (max-noise), medium noise (med-noise), sparse density (F-Sparse), and dense density (F-Dense). Metrics are reported in L2L_2-Chamfer Distance (L2L_2-CD ×100\times 100) evaluated on 10,000 sampled points.

    Dataset / Condition DSDF ATLAS PSR P2S IGR NP Ours
    ABC (clean) 8.41 4.69 2.49 1.80 0.51 0.48 0.200
    FAMOUS (clean) 10.08 4.69 1.67 1.41 1.65 0.22 0.044
    ABC var-noise 12.51 4.04 3.29 2.14 - 0.72 0.488
    ABC max-noise 11.34 4.47 3.89 2.76 - 1.24 0.571
    F-med-noise 9.89 4.54 1.80 1.51 - 0.28 0.071
    F-max-noise 13.17 4.14 3.41 2.52 - 0.31 0.298
    F-Sparse 10.41 4.91 2.17 1.93 - 0.84 0.083
    F-Dense 9.49 4.35 1.60 1.33 - 0.22 0.087
    Overall Mean (Variants) 11.73 4.30 3.10 2.23 - 0.60 0.266

    Compared to DeepSDF (DSDF), AtlasNet (ATLAS), Screened Poisson (PSR), Points2Surf (P2S), Implicit Geometric Regularization (IGR), and Neural-Pull (NP), the Predictive Context Prior demonstrates superior noise resilience and consistency across varying point densities.

  7. Knowl 7 — Surface Reconstruction Performance on Complex Indoor Scenes

    data/table

    Scene-level surface reconstruction was evaluated on five scenes from the 3D Scene benchmark (Burghers, Lounge, Copyroom, Stonewall, Totempole) and on SceneNet at varying sampling densities (20, 100, 500, and 1000 points/m2\text{m}^2). Marching Cubes was executed on a 5123512^3 grid. In 3D Scene, accuracy is measured by L2L_2-CD (×1000\times 1000), L1L_1-CD, Normal Consistency (Norm), and average vertex-to-mesh distance error (mm).

    Scene (3D Scene) Metric LIG Neural-Pull Ours
    Burghers L2L_2-CD ↓\downarrow / Error (mm) ↓\downarrow 0.839 / 28.70 1.76 / 11.23 0.267 / 9.28
    Lounge L2L_2-CD ↓\downarrow / Error (mm) ↓\downarrow 0.789 / 28.20 39.71 / 98.03 0.061 / 6.76
    Copyroom L2L_2-CD ↓\downarrow / Error (mm) ↓\downarrow 0.906 / 30.50 0.51 / 8.76 0.076 / 7.76
    Stonewall L2L_2-CD ↓\downarrow / Error (mm) ↓\downarrow 1.08 / 33.65 0.063 / 6.84 0.061 / 6.33
    Totempole L2L_2-CD ↓\downarrow / Error (mm) ↓\downarrow 1.37 / 38.90 0.19 / 10.21 0.10 / 8.36

    On the SceneNet dataset, the Predictive Context Prior achieves mean L1L_1-CD values of 0.029, 0.017, 0.017, and 0.017 across 20, 100, 500, and 1000 points/m2\text{m}^2 respectively (compared to 0.030, 0.018, 0.018, 0.018 for LIG, and 0.037, 0.047, 0.047, 0.054 for Neural-Pull), with F-scores reaching 0.831, 0.943, 0.962, and 0.965, outperforming Convolutional OccNet, LIG, and Neural-Pull.

  8. Knowl 8 — Ablation on Predictive Query Network Components and Prior Specialization

    empirical result

    Ablation experiments conducted on the ABC dataset justify the critical design choices of the Predictive Context Prior architecture (evaluated in L2L_2-CD ×1000\times 1000 and Normal Consistency):

    1. Predicted Query Offset (Δq\Delta q): Directly setting ql′=qgq'_l = q_g ("No Δq\Delta q") increases L2L_2-CD error from 2.090 to 3.13 (NC drops from 0.945 to 0.924). Predicting ql′q'_l directly as absolute coordinates ("Direct ql′q'_l") increases L2L_2-CD to 2.78 (NC 0.935). Learning residual displacement Δq\Delta q relative to qgq_g is essential for stable optimization.

    2. Predicted Condition (fl′f'_l): Removing the query-dependent condition prediction and relying solely on auto-decoding ("No fl′f'_l") increases L2L_2-CD error to 4.03. Using a static PointNet feature flf_l alongside Δq\Delta q ("No fl′+flf'_l + f_l") increases error to 4.21, proving that the condition fl′f'_l must be dynamically predicted per query point alongside ql′q'_l.

    3. Prior Freezing vs. Joint Tuning: Fitting the target point cloud without any pre-trained local prior ("No Prior") yields an L2L_2-CD error of 4.04 (NC 0.9200). Initializing with the pre-trained prior but jointly tuning θ2\theta_2 and θ3\theta_3 at test time ("Tune θ2+θ3\theta_2 + \theta_3") results in an L2L_2-CD of 3.60 (NC 0.9250). Keeping the pre-trained prior θ2\theta_2 strictly frozen while specializing θ3\theta_3 achieves the optimal L2L_2-CD of 2.09 (NC 0.9446).

  9. Knowl 9 — Impact of Local Region Partitioning and Coordinate Normalization

    empirical result

    Ablation analysis on the ABC dataset evaluates the effect of spatial partitioning resolution, boundary overlap, and normalization during local context prior learning (measured by L2L_2-CD ×1000\times 1000 and Normal Consistency):

    1. Grid Partitioning Resolution: Dividing training shapes into a 636^3 bounding-box grid achieves the best performance (L2L_2-CD 2.09, NC 0.945). Coarser partitions (030^3 unpartitioned yields 12.06; 434^3 yields 4.43) create overly complex regions that degrade prior learning. Overly fine partitions (838^3 yields 2.58) provide insufficient geometric context.

    2. Patch Overlap and Self-Prior: Overlapping adjacent regions ("Lap", L2L_2-CD 2.09) or learning the prior exclusively from local regions of the test target itself ("Self", L2L_2-CD 2.09, NC 0.942) yields nearly identical final accuracy to multi-shape training, demonstrating that flexible query prediction compensates for limited training diversity, though larger datasets accelerate optimization convergence.

    3. Geodesic Patches: Constructing local regions tit_i using geodesic neighborhoods (1000, 2000, or 4000 neighbors: G1, G2, G4) yields L2L_2-CDs of 2.9, 2.3, and 2.5, underperforming simple regular bounding-box grid partitioning.

    4. Region Normalization: Centering and scale normalization of local regions tit_i improves performance from 2.83 (no normalization) to 2.09 (both centering and scaling; centering alone yields 2.13, scaling alone yields 2.67).

  10. Knowl 10 — Computational Overhead of Test-Time Query Optimization

    limitation

    A primary limitation of the Predictive Context Prior framework is that reconstructing the surface of an unseen target shape or scene requires per-instance gradient-based optimization of the neural query network parameters θ3\theta_3 at test time. Consequently, the method incurs higher inference latency than feed-forward single-pass networks (such as Points2Surf or direct regression baselines) that produce geometry via a single forward evaluation without test-time fine-tuning.

Coverage note — Deliberately omitted the 2D circle-to-square illustrative toy example (Fig. 5) and qualitative video descriptions, as their mathematical and conceptual insights are fully captured by the Predictive Context Prior formulation and ablation knowls.

References

  1. 1.Matan Atzmon and Yaron Lipman. Sal: Sign agnostic learning of shapes from raw data. In IEEE Conference on Computer Vision and Pattern Recognition, 2020. 1, 2
  2. 2.Matan Atzmon and yaron Lipman. SALD: sign agnostic learning with derivatives. In International Conference on Learning Representations, 2021. 2, 5
  3. 3.Jan Bednarik, Shaifali Parashar, Erhan Gundogdu, and Pascal Salzmann, Mathieu andFua. Shape reconstruction by learning differentiable surface representations. In IEEE Conference on Computer Vision and Pattern Recognition, 2020. 2
  4. 4.Yizhak Ben-Shabat, Chamin Hewa Koneputugodage, and Stephen Gould. Digs : Divergence guided shape implicit neural representation for unoriented point clouds. CoRR, abs/2106.10811, 2021. 2
  5. 5.F. Bernardini, J. Mittleman, H. Rushmeier, C. Silva, and G. Taubin. The ball-pivoting algorithm for surface reconstruction. IEEE Transactions on Visualization and Computer Graphics, 5(4):349–359, 1999. 2, 5
  6. 6.Alexandre Boulch and Renaud Marlet. Poco: Point convolution for surface reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition, 2022. 2
  7. 7.Rohan Chabra, Jan Eric Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, and Richard A. Newcombe. Deep local shapes: Learning local SDF priors for detailed 3D reconstruction. In European Conference on Computer Vision, volume 12374, pages 608–625, 2020. 1, 2, 5, 6
  8. 8.Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. ShapeNet: An Information-Rich 3D Model Repository. Technical Report arXiv:1512.03012 [cs.GR], Stanford University — Princeton University — Toyota Technological Institute at Chicago, 2015. 5
  9. 9.Chao Chen, Zhizhong Han, Yu-Shen Liu, and Matthias Zwicker. Unsupervised learning of fine structure generation for 3D point clouds by 2D projection matching. In IEEE International Conference on Computer Vision, 2021. 2
  10. 10.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. IEEE Conference on Computer Vision and Pattern Recognition, 2019. 2, 5
  11. 11.Julian Chibane, Aymen Mir, and Gerard Pons-Moll. Neural unsigned distance fields for implicit function learning. arXiv, 2010.13938, 2020. 2, 5
  12. 12.Christopher Bongsoo Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3D-R2N2: A unified approach for single and multi-view 3D object reconstruction. In European Conference on Computer Vision, pages 628–644, 2016. 5
  13. 13.Thomas Davies, Derek Nowrouzezahrai, and Alec Jacobson. On the effectiveness of weight-encoded neural implicit 3d shapes. CoRR, abs/2009.09808, 2021. 2
  14. 14.Philipp Erler, Paul Guerrero, Stefan Ohrhallinger, Niloy J. Mitra, and Michael Wimmer. Points2Surf: Learning implicit surfaces from point clouds. In European Conference on Computer Vision, 2020. 1, 2, 5, 6, 8
  15. 15.Haoqiang Fan, Hao Su, and Leonidas J. Guibas. A point set generation network for 3D object reconstruction from a single image. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2463–2471, 2017. 5
  16. 16.Matheus Gadelha, Rui Wang, and Subhransu Maji. Shape reconstruction using differentiable projections and deep priors. In International Conference on Computer Vision, 2019. 2
  17. 17.Matheus Gadelha, Rui Wang, and Subhransu Maji. Deep manifold prior. CoRR, abs/2004.04242, 2020. 2
  18. 18.Kyle Genova, Forrester Cole, Avneesh Sud, Aaron Sarna, and Thomas Funkhouser. Local deep implicit functions for 3d shape. In IEEE Conference on Computer Vision and Pattern Recognition, June 2020. 2
  19. 19.Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. Implicit geometric regularization for learning shapes. arXiv, 2002.10099, 2020. 1, 2, 6
  20. 20.Thibault Groueix, Matthew Fisher, Vladimir G. Kim, Bryan C. Russell, and Mathieu Aubry. A papier-mch approach to learning 3D surface generation. In IEEE Conference on Computer Vision and Pattern Recognition, 2018. 2, 5, 6
  21. 21.Zhizhong Han, Chao Chen, Yu-Shen Liu, and Matthias Zwicker. Drwr: A differentiable renderer without rendering for unsupervised 3D structure learning from silhouette images. In International Conference on Machine Learning, 2020. 2
  22. 22.Zhizhong Han, Chao Chen, Yu-Shen Liu, and Matthias Zwicker. ShapeCaptioner: Generative caption network for 3D shapes by learning a mapping from parts detected in multiple views to sentences. In ACM International Conference on Multimedia, 2020. 2
  23. 23.Zhizhong Han, Honglei Lu, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, Junwei Han, and C.L. Philip Chen. 3D2SeqViews: Aggregating sequential views for 3D global feature learning by cnn with hierarchical attention aggregation. IEEE Transactions on Image Processing, 28(8):3986–3999, 2019. 2
  24. 24.Zhizhong Han, Guanhui Qiao, Yu-Shen Liu, and Matthias Zwicker. SeqXY2SeqZ: Structure learning for 3D shapes by sequentially predicting 1D occupancy segments from 2D coordinates. In European Conference on Computer Vision, 2020. 2
  25. 25.Zhizhong Han, Mingyang Shang, Yu-Shen Liu, and Matthias Zwicker. View Inter-Prediction GAN: Unsupervised representation learning for 3D shapes by learning global shape memories to support local view predictions. In AAAI, pages 8376–8384, 2019. 2, 4
  26. 26.Zhizhong Han, Mingyang Shang, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, Junwei Han, and C.L. Philip Chen. SeqViews2SeqLabels: Learning 3D global features via aggregating sequential views by rnn with attention. IEEE Transactions on Image Processing, 28(2):685–672, 2019. 2
  27. 27.Zhizhong Han, Mingyang Shang, Xiyang Wang, Yu-Shen Liu, and Matthias Zwicker. Y2Seq2Seq: Cross-modal representation learning for 3D shape and text by joint reconstruction and prediction of view and word sequences. In AAAI, pages 126–133, 2019. 2
  28. 28.Zhizhong Han, Xiyang Wang, Yu-Shen Liu, and Matthias Zwicker. Multi-angle point cloud-vae:unsupervised feature learning for 3D point clouds from multiple angles by joint self-reconstruction and half-to-half prediction. In IEEE International Conference on Computer Vision, 2019. 2
  29. 29.A. Handa, V. Patraucean, V. Badrinarayanan, S. Stent, and R. Cipolla. Understanding realworld indoor scenes with synthetic data. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4077–4085, 2016. 5
  30. 30.Rana Hanocka, Gal Metzer, Raja Giryes, and Daniel Cohen-Or. Point2mesh: A self-prior for deformable meshes. ACM Trans. Graph., 39(4), 2020. 2
  31. 31.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. 3
  32. 32.Tao Hu, Zhizhong Han, Abhinav Shrivastava, and Matthias Zwicker. Render4Completion: Synthesizing multi-view depth maps for 3D shape completion. ArXiv, abs/1904.08366, 2019. 2
  33. 33.Tao Hu, Zhizhong Han, and Matthias Zwicker. 3D shape completion with multi-view consistent inference. In AAAI, 2020. 2
  34. 34.Eldar Insafutdinov and Alexey Dosovitskiy. Unsupervised learning of shape and pose with differentiable point clouds. In Advances in Neural Information Processing Systems, pages 2807–2817, 2018. 2
  35. 35.Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. 2022. 2
  36. 36.Meng Jia and Matthew Kyan. Learning occupancy function from point clouds for surface reconstruction. arXiv, 2010.11378, 2020. 2
  37. 37.Chiyu Jiang, Avneesh Sud, Ameesh Makadia, Jingwei Huang, Matthias Nießner, and Thomas Funkhouser. Local implicit grid representations for 3D scenes. In IEEE Conference on Computer Vision and Pattern Recognition, 2020. 1, 2, 4, 5, 6
  38. 38.Yue Jiang, Dantong Ji, Zhizhong Han, and Matthias Zwicker. SDFDiff: Differentiable rendering of signed distance fields for 3D shape optimization. In IEEE Conference on Computer Vision and Pattern Recognition, 2020. 2
  39. 39.Michael M. Kazhdan and Hugues Hoppe. Screened poisson surface reconstruction. ACM Transactions Graphics, 32(3):29:1–29:13, 2013. 2, 5, 6
  40. 40.Sebastian Koch, Albert Matveev, Zhongshi Jiang, Francis Williams, Alexey Artemov, Evgeny Burnaev, Marc Alexa, Denis Zorin, and Daniele Panozzo. ABC: A big cad model dataset for geometric deep learning. In IEEE Conference on Computer Vision and Pattern Recognition, June 2019. 5
  41. 41.Tianyang Li, Xin Wen, Yu-Shen Liu, Hua Su, and Zhizhong Han. Learning deep implicit functions for 3D shapes with dynamic code clouds. In IEEE Conference on Computer Vision and Pattern Recognition, 2022. 2
  42. 42.Yiyi Liao, Simon Donne, and Andreas Geiger. Deep marching cubes: Learning explicit surface representations. In Conference on Computer Vision and Pattern Recognition, 2018. 5
  43. 43.Minghua Liu, Xiaoshuai Zhang, and Hao Su. Meshing point clouds with predicted intrinsic-extrinsic ratio guidance. In European Conference on Computer vision, 2020. 2, 5
  44. 44.Shi-Lin Liu, Hao-Xiang Guo, Hao Pan, Pengshuai Wang, Xin Tong, and Yang Liu. Deep implicit moving least-squares functions for 3D reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition, 2021. 2, 6
  45. 45.Xinhai Liu, Zhizhong Han, Yu-Shen Liu, and Matthias Zwicker. Point2Sequence: Learning the shape representation of 3D point clouds with an attention-based sequence to sequence network. In AAAI, pages 8778–8785, 2019. 2
  46. 46.S. Lombardi, M. R. Oswald, and M. Pollefeys. Scalable point cloud-based reconstruction with local implicit functions. In International Conference on 3D Vision, pages 997–1007, 2020. 2
  47. 47.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3D surface construction algorithm. Computer Graphics, 21(4):163–169, 1987. 1, 2, 3, 4
  48. 48.Yiming Luo, Zhenxing Mi, and Wenbing Tao. Deepdt: Learning geometry from delaunay triangulation for surface reconstruction. CoRR, abs/2101.10353, 2020. 2, 5
  49. 49.Baorui Ma, Zhizhong Han, Yu-Shen Liu, and Matthias Zwicker. Neural-pull: Learning signed distance functions from point clouds by learning to pull space onto surfaces. In International Conference on Machine Learning, 2021. 2, 3, 5, 6
  50. 50.Julien N. P. Martel, David B. Lindell, Connor Z. Lin, Eric R. Chan, Marco Monteiro, and Gordon Wetzstein. ACORN: adaptive coordinate networks for neural scene representation. CoRR, abs/2105.02788, 2021. 2
  51. 51.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3D reconstruction in function space. In IEEE Conference on Computer Vision and Pattern Recognition, 2019. 2, 4, 5
  52. 52.Zhenxing Mi, Yiming Luo, and Wenbing Tao. Ssrnet: Scalable 3D surface reconstruction network. In IEEE Conference on Computer Vision and Pattern Recognition, 2020. 2, 5
  53. 53.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. IEEE Conference on Computer Vision and Pattern Recognition, 2022. 2
  54. 54.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision, 2020. 2
  55. 55.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. arXiv:2201.05989, 2022. 2
  56. 56.Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In International Conference on Computer Vision, 2021. 2
  57. 57.Yutaka Ohtake, Alexander G. Belyaev, Marc Alexa, Greg Turk, and Hans-Peter Seidel. Multi-level partition of unity implicits. ACM Transactions on Graphics, 22(3):463–470, 2003. 6
  58. 58.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. DeepSDF: Learning continuous signed distance functions for shape representation. In IEEE Conference on Computer Vision and Pattern Recognition, 2019. 2, 4, 5, 6, 7
  59. 59.Songyou Peng, Chiyu ”Max” Jiang, Yiyi Liao, Michael Niemeyer, Marc Pollefeys, and Andreas Geiger. Shape as points: A differentiable poisson solver. In Advances in Neural Information Processing Systems, 2021. 2
  60. 60.Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. PointNet: Deep learning on point sets for 3D classification and segmentation. In IEEE Conference on Computer Vision and Pattern Recognition, 2017. 2, 3
  61. 61.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems, pages 5105–5114, 2017. 2
  62. 62.Konstantinos Rematas, Ricardo Martin-Brualla, and Vittorio Ferrari. Sharf: Shape-conditioned radiance fields from a single view. In International Conference on Machine Learning, 2021. 2
  63. 63.Darius Ruckert, Linus Franke, and Marc Stamminger. Adop: Approximate differentiable one-pixel point rendering. arXiv:2110.06635, 2021. 2
  64. 64.Sara Fridovich-Keil and Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In IEEE Conference on Computer Vision and Pattern Recognition, 2022. 2
  65. 65.Jonathan Sauder and Bjarne Sievers. Self-supervised deep learning on point clouds by reconstructing space. In Advances in Neural Information Processing Systems, pages 12962–12972. 2019. 2
  66. 66.Yawar Siddiqui, Justus Thies, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Retrievalfuse: Neural 3D scene reconstruction with a database. In International Conference on Computer Vision), 2021. 2
  67. 67.Vincent Sitzmann, Julien N.P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In Advances in Neural Information Processing Systems, 2020. 2
  68. 68.Lars Mescheder Marc Pollefeys Andreas Geiger Songyou Peng, Michael Niemeyer. Convolutional occupancy networks. In European Conference on Computer Vision, 2020. 2, 6
  69. 69.Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Neural geometric level of detail: Real-time rendering with implicit 3D shapes. In IEEE Conference on Computer Vision and Pattern Recognition, 2021. 2
  70. 70.Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. NeurIPS, 2020. 2
  71. 71.Jiapeng Tang, Jiabao Lei, Dan Xu, Feiying Ma, Kui Jia, and Lei Zhang. Sa-convonet: Sign-agnostic optimization of convolutional occupancy networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 2
  72. 72.Maxim Tatarchenko, Stephan R. Richter, Rene Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox. What do single-view 3D reconstruction networks learn? In The IEEE Conference on Computer Vision and Pattern Recognition, 2019. 5
  73. 73.Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhofer, Carsten Stoll, and Christian Theobalt. PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations. European Conference on Computer Vision, 2020. 1, 2
  74. 74.Bram Wallace and Bharath Hariharan. Few-shot generalization for single-image 3d reconstruction via priors. In IEEE International Conference on Computer Vision, pages 3817–3826. 2
  75. 75.Xin Wen, Zhizhong Han, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Yu-Shen Liu. Cycle4completion: Unpaired point cloud completion using cycle transformation with missing region coding. In IEEE Conference on Computer Vision and Pattern Recognition, 2021. 2
  76. 76.Xin Wen, Tianyang Li, Zhizhong Han, and Yu-Shen Liu. Point cloud completion by skip-attention network with hierarchical folding. In IEEE Conference on Computer Vision and Pattern Recognition, 2020. 2
  77. 77.Xin Wen, Peng Xiang, Zhizhong Han, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Yu-Shen Liu. Pmp-net: Point cloud completion by learning multi-step point moving paths. In IEEE Conference on Computer Vision and Pattern Recognition, 2021. 2
  78. 78.Xin Wen, Junsheng Zhou, Yu-Shen Liu, Hua Su, Zhen Dong, and Zhizhong Han. 3D shape reconstruction from 2D images with disentangled attribute flow. In IEEE Conference on Computer Vision and Pattern Recognition, 2022. 2
  79. 79.Francis Williams, Teseo Schneider, Claudio Silva, Denis Zorin, Joan Bruna, and Daniele Panozzo. Deep geometric prior for surface reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition, 2019. 2, 5
  80. 80.Jiajun Wu, Chengkai Zhang, Xiuming Zhang, Zhoutong Zhang, William T. Freeman, and Joshua B. Tenenbaum. Learning shape priors for single-view 3d completion and reconstruction. In European Conference on Computer Vision, pages 673–691, 2018. 2
  81. 81.Peng Xiang, Xin Wen, Yu-Shen Liu, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Zhizhong Han. Snowflakenet: Point cloud completion by snowflake point deconvolution with skip-transformer. In IEEE International Conference on Computer Vision, 2021. 2
  82. 82.Mingyue Yang, Yuxin Wen, Weikai Chen, Yongwei Chen, and Kui Jia. Deep optimized priors for 3d shape modeling and reconstruction, 2020. 2
  83. 83.Wang Yifan, Shihao Wu, Cengiz Oztireli, and Olga Sorkine-Hornung. Iso-points: Optimizing neural implicit surfaces with hybrid representations. CoRR, abs/2012.06434, 2020. 2
  84. 84.Wentao Yuan, Tejas Khot, David Held, Christoph Mertz, and Martial Hebert. PCN: Point completion network. In Proceedings of 2018 International Conference on 3D Vision, 2018. 2
  85. 85.Wenbin Zhao, Jiabao Lei, Yuxin Wen, Jianguo Zhang, and Kui Jia. Sign-agnostic implicit learning of surface self-similarities for shape modeling and reconstruction from raw point clouds. CoRR, abs/2012.07498, 2020. 2
  86. 86.Qian-Yi Zhou and Vladlen Koltun. Dense scene reconstruction with points of interest. ACM Transactions on Graphics, 32(4):112:1–112:8, 2013. 5
  87. 87.Zihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu, Hujun Bao, Zhaopeng Cui, Martin R. Oswald, and Marc Pollefeys. Nice-slam: Neural implicit scalable encoding for slam. In IEEE Conference on Computer Vision and Pattern Recognition, 2022. 2

Citation

MLA
Ma, B., et al. “Surface Reconstruction from Point Clouds by Learning Predictive Context Priors”. arXiv, 2022, http://arxiv.org/abs/2204.11015v1.
APA
Ma, B., Liu, Y.-S., Zwicker, M., & Han, Z. (2022). Surface Reconstruction from Point Clouds by Learning Predictive Context Priors. arXiv. http://arxiv.org/abs/2204.11015v1
Chicago
Ma, B., Y.-S. Liu, M. Zwicker, and Z. Han. 2022. “Surface Reconstruction from Point Clouds by Learning Predictive Context Priors”. arXiv. http://arxiv.org/abs/2204.11015v1.
Harvard
Ma, B. et al. (2022) “Surface Reconstruction from Point Clouds by Learning Predictive Context Priors”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.11015v1.
Vancouver
1. Ma B, Liu Y-S, Zwicker M, Han Z (2022) Surface Reconstruction from Point Clouds by Learning Predictive Context Priors. arXiv

BibTeX

@article{ma2022surface,
  title = {Surface Reconstruction from Point Clouds by Learning Predictive Context Priors},
  author = {Ma, Baorui and Liu, Yu-Shen and Zwicker, Matthias and Han, Zhizhong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.11015v1},
  eprint = {2204.11015}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE