PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces Using Permutohedral Lattices

Radu Alexandru RosuSven Behnke

article2023CVPR98 citations

Proposes a multi-view 3D reconstruction framework that uses permutohedral lattice hash encodings and a specialized regularization scheme to recover fine surface details like pores and wrinkles from RGB images alone while enabling real-time rendering.

Listen

Accurately reconstructing high-quality three-dimensional geometry and appearance from multi-view photographs is a core challenge across computer vision, digital mapping, and virtual graphics. While recent neural radiance field techniques render photorealistic views rapidly, they often yield inaccurate surface geometry on untextured or reflective objects. Conversely, implicit surface techniques based on signed distance functions capture better geometry but typically suffer from long computation times and overly smoothed surfaces that erase fine geometric details.

The article demonstrates an implicit surface reconstruction framework called PermutoSDF. The primary objective is to achieve fast, high-fidelity three-dimensional surface reconstruction and real-time novel-view rendering using only standard color images without requiring object masks.

The approach combines implicit surface modeling with a multi-resolution hash encoding based on a permutohedral lattice rather than traditional cubical voxels. Because simplex vertices scale linearly with dimension instead of exponentially, this lattice drastically reduces memory accesses during feature interpolation. The system separates the modeling into two dedicated networks—one for geometry and one for color—and integrates a phased regularization scheme. This scheme initially applies curvature regularization to prevent geometry artifacts in ambiguous regions and subsequently enforces a smoothness constraint on the color network to compel the underlying geometric model to capture fine surface variations.

The findings show that the proposed permutohedral lattice accelerates both model training and higher-dimensional inference compared to cubical voxel grids. In quantitative benchmarks on the DTU dataset, the method achieved lower surface error metrics than existing baselines, reducing the mean Chamfer distance to 0.68 without masks compared to 0.84 for NeuS and 1.57 for Instant Neural Graphics Primitives. The system generated higher image fidelity on novel-view synthesis, attaining an average peak signal-to-noise ratio of 33.97 decibels. Qualitative evaluations further demonstrate the framework's ability to recover subtle micro-details such as skin pores and wrinkles while enabling real-time rendering at 30 frames per second on a single commercial graphics processor after approximately 30 minutes of training.

These results establish that organizations can achieve highly detailed, production-grade 3D assets quickly using accessible hardware and standard camera inputs. By eliminating the need for manual foreground masking and cutting training time to half an hour, the framework substantially reduces labor costs and processing overhead for 3D content creation workflows.

Teams developing spatial computing or 3D mapping applications should consider transitioning from cubical voxel hashing to permutohedral lattice representations, particularly when scaling to multi-dimensional data such as spatio-temporal modeling. Practitioners looking to deploy the framework can utilize the public code base to benchmark performance against existing rendering pipelines.

Confidence in the geometric fidelity is high under controlled conditions, though minor limitations persist. The method can struggle with intricate, non-solid geometry such as individual hair strands, and real-world camera artifacts like optical defocus or imperfect color calibration can limit detail extraction relative to ideal synthetic conditions. Additionally, full direct image supervision for four-dimensional dynamic scenes remains an open area requiring further investigation.

arXiv: 2211.12562
Cover for PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces Using Permutohedral Lattices

Abstract

Neural radiance-density field methods have become increasingly popular for the task of novel-view rendering. Their recent extension to hash-based positional encoding ensures fast training and inference with visually pleasing results. However, density-based methods struggle with recovering accurate surface geometry. Hybrid methods alleviate this issue by optimizing the density based on an underlying SDF. However, current SDF methods are overly smooth and miss fine geometric details. In this work, we combine the strengths of these two lines of work in a novel hash-based implicit surface representation. We propose improvements to the two areas by replacing the voxel hash encoding with a permutohedral lattice which optimizes faster, especially for higher dimensions. We additionally propose a regularization scheme which is crucial for recovering high-frequency geometric detail. We evaluate our method on multiple datasets and show that we can recover geometric detail at the level of pores and wrinkles while using only RGB images for supervision. Furthermore, using sphere tracing we can render novel views at 30 fps on an RTX 3090. Code is publicly available at https://radualexandru.github.io/permuto_sdf

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Classical Multi-View Reconstruction
  • 2.2. NeRF Models
  • 2.3. Accelerating NeRF
  • 2.4. Implicit Representation
  • 3. Method Overview
  • 4. Permutohedral Lattice SDF Rendering
  • 4.1. Volumetric Rendering
  • 4.2. Hash Encoding with Permutohedral Lattice
  • 4.3. 4D Background Estimation
  • 5. PermutoSDF Training and Regularization
  • 5.1. SDF Regularization
  • 5.2. Color Regularization
  • 5.3. Training Schedule
  • 6. Acceleration
  • 6.1. Occupancy Grid
  • 6.2. Sphere Tracing
  • 6.3. Implementation Details
  • 7. Results
  • 7.1. DTU Data Set
  • 7.2. Multiface Data Set
  • 7.3. Rendered Head Images
  • 7.4. Performance
  • 7.5. Ablation Study
  • 8. Conclusion
  • 7.6. 4D Spatio-temporal Surface
  • References

Knowls

  1. Knowl 1 — PermutoSDF’s two-network implicit-surface representation

    model/method

    PermutoSDF reconstructs a scene from posed RGB images by representing geometry with a signed distance field (SDF) and appearance with a separate view-dependent color network. For a spatial point x∈R3x\in\mathbb{R}^3, the reconstructed surface is the zero level set

    S={x∈R3∣g(enc⁡(x;θg);Φg)=0},S=\{x\in\mathbb{R}^3\mid g(\operatorname{enc}(x;\theta_g);\Phi_g)=0\},

    where enc⁡(x;θg)\operatorname{enc}(x;\theta_g) is a learned multi-resolution permutohedral-lattice encoding, θg\theta_g contains its feature parameters, and Φg\Phi_g contains the SDF multilayer perceptron parameters. The SDF network also outputs a learned geometric feature χ\chi.

    A separate color network predicts RGB color from the color encoding hc=enc⁡(x;θc)h_c=\operatorname{enc}(x;\theta_c), viewing direction vv, the SDF normal n=∇xg(x)n=\nabla_x g(x), and the SDF feature χ\chi:

    c=c(hc,v,n,χ;Φc).c=c(h_c,v,n,\chi;\Phi_c).

    The two networks are optimized jointly through volumetric rendering, but their separate parameterizations allow geometry and appearance to receive different regularization. Training uses only posed RGB images and does not require object masks, depth, or surface point supervision.

  2. Knowl 2 — Permutohedral-lattice hash encoding and higher-dimensional representation

    model/method

    Instead of the hypercubic voxel hash encoding used by Instant Neural Graphics Primitives, PermutoSDF partitions a dd-dimensional space into uniform simplices of a permutohedral lattice. A point is assigned to its containing simplex, its barycentric coordinates are computed, and the feature vectors at the simplex vertices are interpolated. The interpolated features from multiple resolutions are concatenated into the encoding.

    A dd-dimensional permutohedral simplex has d+1d+1 vertices, whereas a hypercubic voxel cell requires 2d2^d vertices. Thus, the number of hash-map accesses grows linearly rather than exponentially with position dimensionality. The containing simplex is found in O(d2)O(d^2) time, after which barycentric interpolation is performed. Each of the LL lattice levels stores up to TT hashed feature vectors of dimension FF.

    The same representation models the NeRF++-style background in four dimensions. Foreground space is represented by a unit sphere, while an outer point is encoded as (x′,y′,z′,1/r)(x',y',z',1/r), where (x′,y′,z′)(x',y',z') is a unit direction and 1/r1/r is inverse distance. In four dimensions, each permutohedral simplex requires five vertices instead of the sixteen vertices required by a hypercubic cell.

    In the reported encoding benchmark, batches of 2192^{19} random points were processed with hash-map capacity T=218T=2^{18}, L=24L=24 levels, and F=2F=2 features per entry. The permutohedral encoding trained faster at every tested dimensionality and inferred faster than the voxel encoding for dimensions greater than two; the advantage increased with dimensionality. In two dimensions, the voxel lattice remained competitive because its four vertices versus three permutohedral vertices did not offset the cost of locating the simplex and computing barycentric coordinates. The paper also fit a four-dimensional animated surface from oriented surface-point samples and generated intermediate shapes by sweeping through the time coordinate.

  3. Knowl 3 — Unbiased SDF-based volumetric rendering

    equation

    For a camera ray with origin o∈R3o\in\mathbb{R}^3 and unit viewing direction v∈R3v\in\mathbb{R}^3, PermutoSDF uses p(t)=o+tvp(t)=o+tv for t≥0t\ge 0. Let g(p(t))g(p(t)) be the SDF value, let nn be the SDF normal, let χ\chi be the learned geometric feature, and let c(p(t),v,n,χ;Φc)c(p(t),v,n,\chi;\Phi_c) be the color-network output. The rendered pixel color is

    C^(p)=∫0+∞w(t)c(p(t),v,n,χ;Φc) dt.\widehat C(p)=\int_{0}^{+\infty}w(t)c(p(t),v,n,\chi;\Phi_c)\,dt.

    The SDF is converted into an opaque density using

    ρ(t)=max⁡(−ddtψs(g(p(t)))ψs(g(p(t))),0),ψs(z)=11+e−az,\rho(t)=\max\left(-\frac{\frac{d}{dt}\psi_s(g(p(t)))}{\psi_s(g(p(t)))} ,0\right), \qquad \psi_s(z)=\frac{1}{1+e^{-az}},

    where a>0a>0 is the sigmoid slope and ψs\psi_s is applied to a scalar SDF value zz. The transmittance and rendering weight are

    T(t)=exp⁡(−∫0tρ(u) du),w(t)=T(t)ρ(t).T(t)=\exp\left(-\int_{0}^{t}\rho(u)\,du\right), \qquad w(t)=T(t)\rho(t).

    This SDF-derived weighting is intended to be unbiased and occlusion-aware, concentrating the volume-rendering contribution around the zero level set rather than allowing an unconstrained density field to explain the images.

  4. Knowl 4 — Eikonal and tangent-plane curvature regularization

    equation

    PermutoSDF regularizes the SDF at sampled spatial points x∈R3x\in\mathbb{R}^3 with an Eikonal term and a local curvature term. For an SDF network gg, the Eikonal loss is

    Leik=∑x(∥∇xg(x)∥−1)2,\mathcal{L}_{\mathrm{eik}}=\sum_x\left(\left\|\nabla_x g(x)\right\|-1\right)^2,

    where ∇xg(x)\nabla_x g(x) is obtained by automatic differentiation. This discourages the degenerate zero-everywhere solution but does not uniquely determine geometry in reflective or textureless regions.

    To encourage smooth geometry without explicitly evaluating the full 3×33\times3 Hessian, the method compares normals at nearby points. Let n=∇xg(x)n=\nabla_x g(x), let τ\tau be a randomly sampled unit vector, let η=n×τ\eta=n\times\tau be a tangent direction, and let ϵ\epsilon be a small perturbation magnitude. Define

    xϵ=x+ϵη,nϵ=∇xg(xϵ).x_\epsilon=x+\epsilon\eta, \qquad n_\epsilon=\nabla_x g(x_\epsilon).

    The tangent-plane curvature penalty is

    Lcurv=∑x(n⋅nϵ−1)2.\mathcal{L}_{\mathrm{curv}}=\sum_x\left(n\cdot n_\epsilon-1\right)^2.

    This loss suppresses local normal changes and therefore smooths ambiguous reflective or untextured regions. Used alone, however, it can oversmooth small geometric details; the color regularization and staged training are used to counteract that tendency.

  5. Knowl 5 — Soft Lipschitz regularization of the color network

    model/method

    PermutoSDF constrains the color network so that high-frequency RGB variation cannot be represented independently of geometric variation. The color network receives the encoded position, viewing direction, SDF normal, and SDF feature; among these inputs, the normal and geometric feature are tied to the SDF, while the encoded position and network weights can otherwise store high-frequency appearance.

    A scalar function ff is kk-Lipschitz if ∥f(d)−f(e)∥≤k∥d−e∥\|f(d)-f(e)\|\le k\|d-e\| for inputs dd and ee. For a color-network layer with activation σ\sigma, input xx, weight matrix WiW_i, bias bib_i, and trainable layer bound kik_i, the method uses

    y=σ(W^ix+bi),W^i=m ⁣(Wi,softplus⁡(ki)),y=\sigma(\widehat W_i x+b_i), \qquad \widehat W_i=m\!\left(W_i,\operatorname{softplus}(k_i)\right),

    where softplus⁡(ki)=ln⁡(1+eki)\operatorname{softplus}(k_i)=\ln(1+e^{k_i}) and mm rescales every row of WiW_i so that its absolute row sum is at most softplus⁡(ki)\operatorname{softplus}(k_i). If the color MLP has ll layers, its soft Lipschitz regularizer is

    LLipschitz=∏i=1lsoftplus⁡(ki).\mathcal{L}_{\mathrm{Lipschitz}}=\prod_{i=1}^{l}\operatorname{softplus}(k_i).

    The method additionally applies weight decay of 0.010.01 to the color hash table θc\theta_c. The intended effect is that a large color change requires a corresponding change in the SDF-derived geometric inputs, encouraging fine color boundaries to be represented by fine geometry rather than by an overly smooth surface plus a highly expressive appearance field.

  6. Knowl 6 — Coarse-to-fine training schedule for recovering detail

    model/method

    PermutoSDF uses a fixed training schedule rather than learning the SDF-to-density sharpness parameter freely. In the sigmoid ψs(z)=(1+e−az)−1\psi_s(z)=(1+e^{-az})^{-1} used for rendering, the quantity 1/a1/a controls the effective standard deviation and range of SDF influence on volume rendering. The method linearly decays 1/a1/a over the first 30,00030{,}000 iterations. This was chosen because a freely learned parameter can cause large objects to dominate its gradients and make thin structures disappear.

    The SDF is initialized as the signed distance to a sphere, and the number of active hash-encoding levels is increased from coarse to fine over the first 10,00010{,}000 iterations. The first 100,000100{,}000 iterations optimize

    L=Lrgb+λ1Leik+λ2Lcurv,\mathcal{L}=\mathcal{L}_{\mathrm{rgb}}+\lambda_1\mathcal{L}_{\mathrm{eik}}+\lambda_2\mathcal{L}_{\mathrm{curv}},

    where Lrgb=∑p∥C^(p)−C(p)∥22\mathcal{L}_{\mathrm{rgb}}=\sum_p\|\widehat C(p)-C(p)\|_2^2, C(p)C(p) is the observed RGB value for pixel pp, and λ1,λ2\lambda_1,\lambda_2 are loss weights. This phase prioritizes a coherent, smooth surface. For the next 100,000100{,}000 iterations, the curvature term is removed and the color-network regularizer is added:

    L=Lrgb+λ1Leik+λ3LLipschitz,\mathcal{L}=\mathcal{L}_{\mathrm{rgb}}+\lambda_1\mathcal{L}_{\mathrm{eik}}+\lambda_3\mathcal{L}_{\mathrm{Lipschitz}},

    with λ3\lambda_3 a loss weight. This second phase transfers high-frequency image evidence back into the geometry.

  7. Knowl 7 — Occupancy-grid sampling and sphere-traced inference

    algorithm

    PermutoSDF accelerates training and rendering with a dense occupancy grid and uses the SDF directly for sphere tracing.

    The occupancy structure has resolution 1283128^3 and is stored in Morton order for fast traversal. Two grids are maintained: a full-precision grid containing the current SDF value at each voxel and a binary grid indicating whether the voxel can make a meaningful contribution to the volume-rendering integral. Unlike density grids in voxel NeRF systems, the stored scalar is an SDF value. Grid updates use the SDF-based rendering weights to identify voxels with sufficient contribution.

    For inference, the sphere-tracing procedure is:

    1. Traverse the occupancy grid along each camera ray and initialize a sample at the first occupied voxel encountered.
    2. Evaluate the SDF at each active sample.
    3. For every sample whose absolute SDF is above the convergence threshold, advance it along the ray by its SDF-estimated distance toward the surface.
    4. Repeat the marching step until every ray has converged or a predefined maximum number of sphere-tracing iterations has been reached.
    5. Evaluate the color network once at the converged surface samples and produce the rendered pixels.

    Most rays converge in approximately two to three tracing iterations. Increasing the maximum iteration count trades rendering speed for accuracy. The implementation uses custom CUDA kernels for parallel multi-resolution lattice slicing, hash-table backpropagation, spatial derivatives needed for normals, and double-backward derivatives needed by the Eikonal loss, allowing optimization through PyTorch autograd without finite differences. The complete system trains in approximately 30 minutes on an RTX 3090 and renders a 1920×10801920\times1080 image in about 3838 ms using sphere tracing.

  8. Knowl 8 — DTU geometry reconstruction accuracy

    data/table

    The DTU evaluation uses approximately 50 posed images per object, with every eighth image held out for testing and the remaining images used for training. Geometry is evaluated by Chamfer distance, for which lower values are better. All methods are trained without mask supervision except for the separate comparison condition explicitly labeled “with mask.” COLMAP uses trim=0. The values show that PermutoSDF has the lowest mean Chamfer distance in both conditions and generally preserves more fine structure in reflective and untextured regions.

    Could not parse LaTeX table
  9. Knowl 9 — Novel-view synthesis and facial-detail results

    data/table

    On the same DTU objects, novel-view synthesis is evaluated with PSNR in dB without mask supervision. PermutoSDF obtains the highest mean PSNR, 33.97 dB, exceeding NeuS at 31.99 dB, NeRF at 32.74 dB, and INGP at 33.10 dB. It is best or tied for best on most individual scans, consistent with the claim that more faithful geometry improves view generalization.

    Could not parse LaTeX table

    On the Multiface human-face dataset, PermutoSDF reconstructs finer facial detail than NeuS and reports PSNR values of 34.9234.92 versus 33.4933.49. On synthetically rendered, perfectly calibrated head images, it recovers detail at the scale of pores and wrinkles. On real Multiface captures, the method still struggles with extremely fine hair: it can create hair-like geometric streaks to explain eyebrows and beards.

  10. Knowl 10 — Ablation of initialization, smoothing, and RGB regularization

    empirical result

    An ablation on the full reconstruction pipeline measures Chamfer distance, with lower values indicating better geometry. The tested sequence starts from naive joint optimization and successively adds sphere initialization, coarse-to-fine hash-level activation, tangent-plane curvature regularization, and Lipschitz RGB regularization.

    Could not parse LaTeX table

    Sphere initialization and coarse-to-fine activation make the recovered shape smoother, especially in highly specular regions, but do not eliminate holes. Curvature regularization removes most holes while producing the worst Chamfer value in this ablation because it smooths away geometric detail. Adding the Lipschitz regularizer to the color network yields the lowest Chamfer distance, supporting the method’s claim that RGB regularization is needed to prevent the color field from explaining high-frequency image structure while the SDF becomes overly smooth.

Coverage note — No substantial contributed material was omitted; the four-dimensional point-supervised surface experiment, occupancy-grid acceleration, and optimized CUDA/autograd implementation are included in the method knowls rather than treated as separate results.

References

  1. 1.Andrew Adams, Jongmin Baek, and Myers Abraham Davis. Fast high-dimensional filtering using the permutohedral lattice. In Computer Graphics Forum, volume 29, pages 753–762. Wiley Online Library, 2010.
  2. 2.Jongmin Baek and Andrew Adams. Some useful properties of the permutohedral lattice for Gaussian filtering. Technical Report, Stanford University, 2009.
  3. 3.Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-NeRF: A multiscale representation for anti-aliasing neural radiance fields. In IEEE International Conference on Computer Vision (ICCV), 2021.
  4. 4.James Busby. 3D scan store head model. https://www.3dscanstore.com/blog/Free-3D-Head-Model. Accessed: 2022-09-30.
  5. 5.Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier. Parseval Networks: Improving robustness to adversarial examples. In 34th International Conference on Machine Learning (ICML), pages 854–863. PMLR, 2017.
  6. 6.Brian Curless and Marc Levoy. A volumetric method for building complex models from range images. In 23rd Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), pages 303–312, 1996.
  7. 7.Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5491–5500, 2022.
  8. 8.Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multiview stereopsis. IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 32(8):1362–1376, 2009.
  9. 9.Silvano Galliani, Katrin Lasinger, and Konrad Schindler. Massively parallel multiview stereopsis by surface normal diffusion. In IEEE International Conference on Computer Vision (ICCV), pages 873–881, 2015.
  10. 10.Rasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, and Henrik Aanæs. Large scale multi-view stereopsis evaluation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 406–413, 2014.
  11. 11.Yue Jiang, Dantong Ji, Zhizhong Han, and Matthias Zwicker. SDFDiff: Differentiable rendering of signed distance fields for 3D shape optimization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1251–1261, 2020.
  12. 12.Michael Kazhdan, Matthew Bolitho, and Hugues Hoppe. Poisson surface reconstruction. In 4th Eurographics Symposium on Geometry Processing (SGP), volume 7, 2006.
  13. 13.Hsueh-Ti Derek Liu, Francis Williams, Alec Jacobson, Sanja Fidler, and Or Litany. Learning smooth neural functions via Lipschitz regularization. In Special Interest Group on Computer Graphics and Interactive Techniques Conference (SIGGRAPH), pages 31:1–31:13. ACM, 2022.
  14. 14.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. In Advances in Neural Information Processing Systems 33 (NeurIPS), pages 15651–15663, 2020.
  15. 15.William E Lorensen and Harvey E Cline. Marching cubes: A high resolution 3D surface construction algorithm. ACM SIGGRAPH Computer Graphics, 21(4):163–169, 1987.
  16. 16.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision (ECCV), pages 405–421, 2020.
  17. 17.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In 6th International Conference on Learning Representations (ICLR), 2018.
  18. 18.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (SIGGRAPH), 41(4):102:1–102:15, July 2022.
  19. 19.Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon. KinectFusion: Real-time dense surface mapping and tracking. In 10th IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pages 127–136, 2011.
  20. 20.Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable volumetric rendering: Learning implicit 3D representations without 3D supervision. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3504–3515, 2020.
  21. 21.Matthias Nießner, Michael Zollhofer, Shahram Izadi, and Marc Stamminger. Real-time 3D reconstruction at scale using voxel hashing. ACM Transactions on Graphics (ToG), 32(6):1–11, 2013.
  22. 22.Michael Oechsle, Songyou Peng, and Andreas Geiger. UNISURF: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 5589–5599, 2021.
  23. 23.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural radiance fields for dynamic scenes. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10318–10327, 2021.
  24. 24.Radu Alexandru Rosu and Sven Behnke. EasyPBR: A lightweight physically-based renderer. In 16th International Conference on Computer Graphics Theory and Applications (GRAPP), 2021.
  25. 25.Johannes L. Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys. Pixelwise view selection for unstructured multi-view stereo. In European Conference on Computer Vision (ECCV), pages 501–518, 2016.
  26. 26.Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In Advances in Neural Information Processing Systems 33 (NeurIPS), pages 7462–7473, 2020.
  27. 27.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5459–5469, 2022.
  28. 28.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-NeRF: Scalable large scene neural view synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8248–8258, 2022.
  29. 29.David Terjék. Adversarial Lipschitz regularization. In 8th International Conference on Learning Representations (ICLR), 2020.
  30. 30.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. NeuS: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. In Advances in Neural Information Processing Systems 34 (NeurIPS), pages 27171–27183, 2021.
  31. 31.Thomas Whelan, Renato F. Salas-Moreno, Ben Glocker, Andrew J. Davison, and Stefan Leutenegger. ElasticFusion: Real-time dense SLAM and light source estimation. International Journal of Robotics Research (IJRR), 35(14):1697–1716, 2016.
  32. 32.Cheng-hsin Wuu, Ningyuan Zheng, Scott Ardisson, Rohan Bali, Danielle Belko, Eric Brockmeyer, Lucas Evans, Timothy Godisart, Hyowon Ha, Alexander Hypes, Taylor Koska, Steven Krenn, Stephen Lombardi, Xiaomin Luo, Kevyn McPhail, Laura Millerschoen, Michal Perdoch, Mark Pitts, Alexander Richard, Jason Saragih, Junko Saragih, Takaaki Shiratori, Tomas Simon, Matt Stewart, Autumn Trimble, Xinshuo Weng, David Whitewolf, Chenglei Wu, Shoou-I Yu, and Yaser Sheikh. Multiface: A dataset for neural face rendering. arXiv preprint arXiv:2207.11243, 2022.
  33. 33.Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. Volume rendering of neural implicit surfaces. In Advances in Neural Information Processing Systems 34 (NeurIPS), pages 4805–4815, 2021.
  34. 34.Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neural surface reconstruction by disentangling geometry and appearance. In Advances in Neural Information Processing Systems 33 (NeurIPS), pages 2492–2502, 2020.
  35. 35.Yuichi Yoshida and Takeru Miyato. Spectral norm regularization for improving the generalizability of deep learning. arXiv preprint arXiv:1705.10941, 2017.
  36. 36.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. NeRF++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020.
  37. 37.Enliang Zheng, Enrique Dunn, Vladimir Jojic, and Jan-Michael Frahm. PatchMatch based joint view selection and depthmap estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1510–1517, 2014.

Citation

MLA
Rosu, R. A., and S. Behnke. “PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces Using Permutohedral Lattices”. arXiv, 2022, http://arxiv.org/abs/2211.12562v2.
APA
Rosu, R. A., & Behnke, S. (2022). PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces using Permutohedral Lattices. arXiv. http://arxiv.org/abs/2211.12562v2
Chicago
Rosu, R. A., and S. Behnke. 2022. “PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces Using Permutohedral Lattices”. arXiv. http://arxiv.org/abs/2211.12562v2.
Harvard
Rosu, R.A. and Behnke, S. (2022) “PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces using Permutohedral Lattices”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.12562v2.
Vancouver
1. Rosu RA, Behnke S (2022) PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces using Permutohedral Lattices. arXiv

BibTeX

@article{rosu2022permutosdf,
  title = {PermutoSDF: Fast Multi-View Reconstruction with Implicit Surfaces using Permutohedral Lattices},
  author = {Rosu, Radu Alexandru and Behnke, Sven},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.12562v2},
  eprint = {2211.12562}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE