Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering

Tao LuMulin YuLinning XuYuanbo XiangliLimin WangDahua LinBo Dai

article2024CVPR983 citations

Proposes a structured 3D Gaussian Splatting framework that dynamically predicts neural Gaussian attributes from anchor points to reduce model redundancy, lower storage requirements, and improve novel view synthesis across varying viewing angles and distances.

Listen

Photo-realistic, real-time 3D scene rendering is essential for applications such as virtual reality and large-scale visual simulation. While recent breakthroughs in 3D Gaussian Splatting have achieved rapid frame rates and strong visual quality, the baseline method tends to overfit training views by generating redundant representations that ignore the underlying scene geometry. Consequently, existing models suffer from excessive storage demands and degrade when encountering significant shifts in viewing angles, distances, texture-less surfaces, or complex lighting conditions.

The article evaluates Scaffold-GS, a structured neural rendering framework designed to improve novel view synthesis while drastically reducing model size. The method constructs a sparse grid of anchor points derived from initial spatial data to establish a hierarchical, region-aware scene representation. Rather than storing millions of static elements, each anchor spawns local neural elements whose geometric and appearance properties are dynamically predicted on the fly based on the camera’s relative distance and angle within the visible field of view. The framework also introduces dynamic growing and pruning strategies to refine anchor coverage and filter out low-opacity elements.

Evaluations across 27 benchmark scenes demonstrate that Scaffold-GS matches or exceeds current state-of-the-art visual quality while maintaining real-time display speeds of 100 to 140 frames per second at 1K resolution. Crucially, the approach reduces storage footprints significantly, achieving roughly 4- to 10-fold compression ratios compared to standard 3D Gaussian Splatting across complex datasets (e.g., dropping from 676 MB to 66 MB on the Deep Blending dataset and from 1.6 GB to 203 MB on the BungeeNeRF dataset). The framework also exhibits superior robustness against visual artifacts in challenging scenarios, including specular reflections, fine-scale geometry, and unseen zoom-in distances.

These findings suggest that anchoring neural representations to regularized spatial structures offers a highly scalable and cost-effective path for deploying immersive 3D visualization pipelines on memory-constrained devices. Organizations seeking to deploy real-time digital twins or virtual spaces should consider adopting structured, view-adaptive Gaussian architectures over unstructured baselines to lower storage and streaming bandwidth without sacrificing rendering fidelity. Future efforts should focus on mitigating initialization risks in extremely sparse, texture-less environments where initial spatial points are difficult to acquire.

arXiv: 2312.00109
Cover for Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering

Abstract

Neural rendering methods have significantly advanced photo-realistic 3D scene rendering in various academic and industrial applications. The recent 3D Gaussian Splatting method has achieved the state-of-the-art rendering quality and speed combining the benefits of both primitive-based representations and volumetric representations. However, it often leads to heavily redundant Gaussians that try to fit every training view, neglecting the underlying scene geometry. Consequently, the resulting model becomes less robust to significant view changes, texture-less area and lighting effects. We introduce Scaffold-GS, which uses anchor points to distribute local 3D Gaussians, and predicts their attributes on-the-fly based on viewing direction and distance within the view frustum. Anchor growing and pruning strategies are developed based on the importance of neural Gaussians to reliably improve the scene coverage. We show that our method effectively reduces redundant Gaussians while delivering high-quality rendering. We also demonstrates an enhanced capability to accommodate scenes with varying levels-of-detail and view-dependent observations, without sacrificing the rendering speed. Project page: https://city-super.github.io/scaffold-gs/.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Methods
  • 3.1. Preliminaries
  • 3.2.1 Anchor Point Initialization
  • 3.2.2 Neural Gaussian Derivation
  • 3.3. Anchor Points Refinement
  • 3.4. Losses Design
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Results Analysis
  • 4.3. Ablation Studies
  • 4.4. Discussions and Limitations
  • 5. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Hierarchical anchor-based scene representation

    model/method

    Scaffold-GS represents a scene with a sparse set of voxel-center anchor points rather than directly storing a large, freely optimized set of 3D Gaussians. Each anchor stores a local feature, a learnable 3D scaling factor, and several learnable local offsets. For a camera, every anchor inside the view frustum spawns a fixed number kk of neural Gaussians. The position, opacity, color, scale, and rotation of each spawned Gaussian are predicted on demand from the anchor feature and the camera-to-anchor distance and direction.

    The generated Gaussians are rendered with differentiable 3D Gaussian splatting. Because only anchor points, their features and offsets, and the scene-specific MLP predictors need to be stored, the representation avoids explicitly storing the full set of view-specific Gaussian attributes. View-dependent prediction allows the same local scene structure to produce different Gaussian attributes for different viewing directions and distances, while the anchor layout preserves a spatial organization tied to the reconstructed scene.

  2. Knowl 2 — Sparse voxel anchor initialization and view-adaptive feature aggregation

    model/method

    Given a Structure-from-Motion point cloud P∈RM×3P\in\mathbb{R}^{M\times 3} and voxel size ϵ>0\epsilon>0, Scaffold-GS constructs unique voxel centers by quantizing the point cloud:

    V={round⁡(Pϵ)}ϵ,V∈RN×3.V=\left\{\operatorname{round}\left(\frac{P}{\epsilon}\right)\right\}\epsilon,\qquad V\in\mathbb{R}^{N\times 3}.

    Here round⁡\operatorname{round} is coordinate-wise rounding, duplicate quantized points are removed, and each resulting voxel center xv∈Vx_v\in V becomes an anchor. Every anchor has a local feature fv∈R32f_v\in\mathbb{R}^{32}, a 3D scale factor lv∈R3l_v\in\mathbb{R}^{3}, and kk learnable offsets Ov∈Rk×3O_v\in\mathbb{R}^{k\times 3}. The voxel grid supplies a regularized spatial scaffold over the irregular SfM points.

    To make the anchor representation both multi-resolution and view-dependent, the feature bank contains fvf_v, fv↓1f_v^{\downarrow 1}, and fv↓2f_v^{\downarrow 2}, where fv↓nf_v^{\downarrow n} is obtained by downsampling the feature channels by a factor of 2n2^n. For camera position xcx_c, the anchor-camera distance and direction are

    δvc=∥xv−xc∥2,d⃗vc=xv−xc∥xv−xc∥2.\delta_{vc}=\lVert x_v-x_c\rVert_2,\qquad \vec d_{vc}=\frac{x_v-x_c}{\lVert x_v-x_c\rVert_2}.

    A small MLP FwF_w predicts three normalized weights from (δvc,d⃗vc)(\delta_{vc},\vec d_{vc}):

    {w,w1,w2}=Softmax⁡(Fw(δvc,d⃗vc)),f^v=wfv+w1fv↓1+w2fv↓2.\{w,w_1,w_2\}=\operatorname{Softmax}\left(F_w(\delta_{vc},\vec d_{vc})\right),\qquad \hat f_v=w f_v+w_1f_v^{\downarrow 1}+w_2f_v^{\downarrow 2}.

    The aggregated feature f^v\hat f_v is used to generate the neural Gaussians visible from the current camera.

  3. Knowl 3 — On-the-fly neural Gaussian derivation and opacity filtering

    model/method

    For each visible anchor vv, Scaffold-GS spawns kk neural Gaussians. If Oi∈R3O_i\in\mathbb{R}^{3} is the iith learnable offset and lv∈R3l_v\in\mathbb{R}^{3} is the anchor scale factor, the Gaussian positions are

    μi=xv+Oi⊙lv,i∈{0,…,k−1},\mu_i=x_v+O_i\odot l_v,\qquad i\in\{0,\ldots,k-1\},

    where ⊙\odot denotes element-wise multiplication and μi∈R3\mu_i\in\mathbb{R}^{3}. Four MLPs decode the Gaussian attributes in one pass from the aggregated anchor feature f^v\hat f_v, distance δvc\delta_{vc}, and direction d⃗vc\vec d_{vc}:

    (α0,…,αk−1)=Fα(f^v,δvc,d⃗vc),(c0,…,ck−1)=Fc(f^v,δvc,d⃗vc),(q0,…,qk−1)=Fq(f^v,δvc,d⃗vc),(s0,…,sk−1)=Fs(f^v,δvc,d⃗vc).\begin{aligned} (\alpha_0,\ldots,\alpha_{k-1})&=F_\alpha(\hat f_v,\delta_{vc},\vec d_{vc}),\\ (c_0,\ldots,c_{k-1})&=F_c(\hat f_v,\delta_{vc},\vec d_{vc}),\\ (q_0,\ldots,q_{k-1})&=F_q(\hat f_v,\delta_{vc},\vec d_{vc}),\\ (s_0,\ldots,s_{k-1})&=F_s(\hat f_v,\delta_{vc},\vec d_{vc}). \end{aligned}

    For Gaussian ii, αi∈R\alpha_i\in\mathbb{R} is opacity, ci∈R3c_i\in\mathbb{R}^{3} is color, qi∈R4q_i\in\mathbb{R}^{4} is the covariance rotation quaternion, and si∈R3s_i\in\mathbb{R}^{3} is the scale vector. These attributes are predicted only for anchors inside the current view frustum. Before rasterization, the representation discards generated Gaussians whose opacity is below a fixed threshold τα\tau_\alpha, retaining only Gaussians satisfying αi≥τα\alpha_i\geq\tau_\alpha. Frustum culling and opacity-based selection reduce the number of MLP outputs and rasterized primitives without requiring all scene Gaussians to be materialized for every view.

  4. Knowl 4 — Gradient-driven anchor growing and opacity-driven pruning

    algorithm

    Scaffold-GS refines the initially voxelized anchor set using the importance of its generated neural Gaussians. During anchor updates, neural Gaussians are spatially quantized into voxels of size ϵg\epsilon_g. For each voxel, the method averages the gradients of the neural Gaussians contained in that voxel over NN training iterations, obtaining an importance measure ∇g\nabla_g. A voxel is a growing candidate when ∇g>τg\nabla_g>\tau_g and no anchor currently occupies it; a new anchor is placed at the voxel center. Candidate anchors are additionally subjected to a random elimination step to prevent rapid uncontrolled expansion.

    Growing is performed at three spatial resolutions m∈{1,2,3}m\in\{1,2,3\}, with voxel sizes and gradient thresholds

    ϵg(m)=ϵg4m−1,τg(m)=τg2m−1.\epsilon_g^{(m)}=\frac{\epsilon_g}{4^{m-1}},\qquad \tau_g^{(m)}=\tau_g2^{m-1}.

    Thus, finer voxels can add anchors for localized details while coarser voxels capture larger missing regions. The pruning operation accumulates the opacity values of all neural Gaussians generated by each anchor over the same refinement interval and removes anchors that fail to produce sufficiently opaque Gaussians.

    To avoid modifying anchors that have received too few observations, an anchor is eligible for growing only after being visited more than NrgN r_g times and eligible for pruning only after being visited more than NrpN r_p times, where rgr_g and rpr_p are observation-rate thresholds. This refinement scheme adds representation capacity in high-error regions and removes anchors that contribute little to rendering.

  5. Knowl 5 — Training objective with Gaussian volume regularization

    equation

    Scaffold-GS optimizes the anchor parameters, offsets, and MLP parameters using a rendered-image reconstruction objective combining pixelwise error, structural similarity, and a volume penalty:

    L=L1+λSSIMLSSIM+λvolLvol.\mathcal{L}=\mathcal{L}_1+\lambda_{\mathrm{SSIM}}\mathcal{L}_{\mathrm{SSIM}}+\lambda_{\mathrm{vol}}\mathcal{L}_{\mathrm{vol}}.

    Here L1\mathcal{L}_1 is the L1L_1 loss between rendered and training-view pixel colors, LSSIM\mathcal{L}_{\mathrm{SSIM}} is the structural-similarity loss, and λSSIM\lambda_{\mathrm{SSIM}} and λvol\lambda_{\mathrm{vol}} are scalar loss weights. If NngN_{\mathrm{ng}} is the number of neural Gaussians and si∈R3s_i\in\mathbb{R}^{3} is the scale vector of Gaussian ii, the volume regularizer is

    Lvol=∑i=1NngProd⁡(si),\mathcal{L}_{\mathrm{vol}}=\sum_{i=1}^{N_{\mathrm{ng}}}\operatorname{Prod}(s_i),

    where Prod⁡(si)\operatorname{Prod}(s_i) is the product of the three scale components. This term encourages small neural Gaussians and reduces unnecessary spatial overlap.

  6. Knowl 6 — Experimental protocol and implementation configuration

    experimental setup

    The evaluation covers 27 scenes from Mip-NeRF360, Tanks&Temples, DeepBlending, the synthetic Blender dataset, BungeeNeRF, and VR-NeRF. The selected scenes include nine Mip-NeRF360 scenes, two Tanks&Temples scenes, two DeepBlending scenes, six BungeeNeRF scenes, and two VR-NeRF scenes, together with the synthetic Blender scenes. The experiments report PSNR, SSIM, LPIPS, storage size in MB, and rendering speed in FPS.

    Scaffold-GS and the principal 3D-GS baseline are trained for 30,000 iterations. The number of neural Gaussians spawned per anchor is k=10k=10. Every MLP has two layers, ReLU activations, and 32 hidden units. Anchor gradients are averaged over N=100N=100 iterations. The default growing threshold is τg=64ϵ\tau_g=64\epsilon, while τg=16ϵ\tau_g=16\epsilon is used for intricate scenes and scenes dominated by texture-less regions. The observation thresholds are rg=0.4r_g=0.4 for growing and rp=0.8r_p=0.8 for pruning. An anchor is pruned when the accumulated opacity of its generated neural Gaussians is below 0.50.5 at a refinement round. The loss weights are λSSIM=0.2\lambda_{\mathrm{SSIM}}=0.2 and λvol=0.001\lambda_{\mathrm{vol}}=0.001.

  7. Knowl 7 — Novel-view quality on real-world scenes

    data/table

    The real-world evaluation compares Scaffold-GS with 3D-GS, Mip-NeRF360, iNGP, and Plenoxels using dataset-averaged PSNR and SSIM, where higher is better, and LPIPS, where lower is better. Scaffold-GS is comparable to the strongest methods on Mip-NeRF360 and improves PSNR and SSIM over 3D-GS on Tanks&Temples and DeepBlending. Its DeepBlending LPIPS is slightly worse than 3D-GS, so the improvement is not uniform across every metric.

    Could not parse LaTeX table
  8. Knowl 8 — Compact real-time rendering and multi-scale performance

    empirical result

    Scaffold-GS substantially reduces storage while retaining real-time rendering speed. On Mip-NeRF360, Tanks&Temples, and Deep Blending, it uses 4.2×, 4.7×, and 10.2× less storage than 3D-GS, respectively. Its rendering speed remains around or above 100 FPS on the reported machine. On multi-scale BungeeNeRF and VR-NeRF scenes and on the synthetic Blender dataset, Scaffold-GS also achieves higher PSNR with smaller models than 3D-GS.

    Could not parse LaTeX table
    Could not parse LaTeX table

    The multi-scale experiments show that dynamically decoding scale and other attributes from viewing distance helps avoid the blurry and needle-shaped artifacts observed with fixed Gaussian attributes. The authors also report faster convergence than 3D-GS.

  9. Knowl 9 — Filtering and anchor-refinement ablations

    empirical result

    The opacity and frustum filters primarily affect speed rather than fidelity. On DB-PLAYROOM, using both filters increases rendering speed from 84 FPS without filtering to 150 FPS, while PSNR changes from 30.4 to 31.07. On DB-DRJOHNSON, the corresponding change is from 79 to 129 FPS, with PSNR changing from 29.7 to 29.79.

    Could not parse LaTeX table

    The anchor-refinement ablation shows that growing is important for recovering details and texture-less regions, whereas pruning controls model size. On DB-PLAYROOM, disabling both operations gives 28.45 PSNR and 24 MB, adding pruning gives 29.12 PSNR and 23 MB, adding growing gives 30.54 PSNR and 71 MB, and using both gives 31.07 PSNR and 61 MB. On DB-DRJOHNSON, the corresponding settings give 28.81 PSNR and 12 MB, 28.51 PSNR and 12 MB, 29.75 PSNR and 76 MB, and 29.79 PSNR and 50 MB.

    Could not parse LaTeX table
  10. Knowl 10 — Dependence on initial scene points

    limitation

    Scaffold-GS relies on the initial SfM point cloud to place its anchor points, and the authors find that initialization is important for high-fidelity rendering. SfM initialization is practical because the point cloud is often produced during image calibration, but it can be inadequate in scenes with large texture-less regions or extremely sparse observations. The gradient-based anchor-growing strategy can repair poor initialization to some extent, yet it does not fully overcome extremely sparse starting points. The expectation that future improvements in reconstruction will progressively improve Scaffold-GS results is presented as an expectation rather than a demonstrated result.

Coverage note — The anchor-feature clustering analysis and qualitative cross-view/specular-reflection visualizations were not made separate knowls because they provide supporting qualitative evidence for feature interpretability and view adaptability rather than standalone quantitative contributions; background and 3D-GS preliminaries were also omitted because they are not Scaffold-GS contributions.

References

  1. 1.Kara-Ali Aliev, Dmitry Ulyanov, and Victor S. Lempitsky. Neural point-based graphics. In European Conference on Computer Vision, 2019.
  2. 2.Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5835–5844, 2021.
  3. 3.Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 5855–5864, 2021.
  4. 4.Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. CVPR, 2022.
  5. 5.Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 19697–19705, 2023.
  6. 6.Mario Botsch, Alexander Hornung, Matthias Zwicker, and Leif Kobbelt. High-quality surface splatting on today’s gpus. In Proceedings Eurographics/IEEE VGTC Symposium Point-Based Graphics, 2005., pages 17–141. IEEE, 2005.
  7. 7.Mario Botsch, Alexander Sorkine-Hornung, Matthias Zwicker, and Leif P. Kobbelt. High-quality surface splatting on today’s gpus. Proceedings Eurographics/IEEE VGTC Symposium Point-Based Graphics, 2005., pages 17–141, 2005.
  8. 8.Eric Chan, Connor Z. Lin, Matthew Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J. Guibas, Jonathan Tremblay, S. Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3d generative adversarial networks. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16102–16112, 2021.
  9. 9.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. ArXiv, abs/2203.09517, 2022.
  10. 10.Anpei Chen, Zexiang Xu, Xinyue Wei, Siyu Tang, Hao Su, and Andreas Geiger. Factor fields: A unified framework for neural fields and beyond. ArXiv, abs/2302.01226, 2023.
  11. 11.Zhiqin Chen, Thomas Funkhouser, Peter Hedman, and Andrea Tagliasacchi. Mobilenerf: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures. In The Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  12. 12.Christopher Bongsoo Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. ArXiv, abs/1604.00449, 2016.
  13. 13.Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In CVPR, 2022.
  14. 14.Sara Fridovich-Keil, Giacomo Meanti, Frederik Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12479–12488, 2023.
  15. 15.Kyle Genova, Forrester Cole, Avneesh Sud, Aaron Sarna, and Thomas A. Funkhouser. Local deep implicit functions for 3d shape. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4856–4865, 2019.
  16. 16.Markus Gross and Hanspeter Pfister. Point-based graphics. Elsevier, 2011.
  17. 17.Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. Deep blending for free-viewpoint image-based rendering. 37(6):257:1–257:15, 2018.
  18. 18.Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, and Gabriel Brostow. Deep blending for free-viewpoint image-based rendering. ACM Transactions on Graphics (ToG), 37(6):1–15, 2018.
  19. 19.Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron, and Paul E. Debevec. Baking neural radiance fields for real-time view synthesis. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5855–5864, 2021.
  20. 20.Eldar Insafutdinov and Alexey Dosovitskiy. Unsupervised learning of shape and pose with differentiable point clouds. In Advances in Neural Information Processing Systems (NeurIPS), 2018.
  21. 21.Abhishek Kar, Christian Hane, and Jitendra Malik. Learning a multi-view stereo machine. ArXiv, abs/1708.05375, 2017.
  22. 22.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023.
  23. 23.Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics, 36(4), 2017.
  24. 24.Georgios Kopanas, Julien Philip, Thomas Leimkuhler, and George Drettakis. Point-based neural rendering with per-view optimization. Computer Graphics Forum, 40, 2021.
  25. 25.K Krishna and M Narasimha Murty. Genetic k-means algorithm. IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), 29(3):433–439, 1999.
  26. 26.Christoph Lassner and Michael Zollhofer. Pulsar: Efficient sphere-based neural rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1440–1449, 2021.
  27. 27.Chen-Hsuan Lin, Chen Kong, and Simon Lucey. Learning efficient point cloud generation for dense 3d object reconstruction. In AAAI Conference on Artificial Intelligence (AAAI), 2018.
  28. 28.Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mixture of volumetric primitives for efficient neural rendering. ACM Transactions on Graphics (ToG), 40(4):1–13, 2021.
  29. 29.Lars M. Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozem, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4455–4465, 2018.
  30. 30.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  31. 31.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022.
  32. 32.Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Muller, and Sanja Fidler. Extracting Triangular 3D Models, Materials, and Lighting From Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8280–8290, 2022.
  33. 33.Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5569–5579, 2021.
  34. 34.Jeong Joon Park, Peter R. Florence, Julian Straub, Richard A. Newcombe, and S. Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 165–174, 2019.
  35. 35.Songyou Peng, Michael Niemeyer, Lars M. Mescheder, Marc Pollefeys, and Andreas Geiger. Convolutional occupancy networks. ArXiv, abs/2003.04618, 2020.
  36. 36.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, 2023.
  37. 37.Liu Ren, Hanspeter Pfister, and Matthias Zwicker. Object space ewa surface splatting: A hardware accelerated approach to high quality point rendering. Computer Graphics Forum, 21, 2002.
  38. 38.Miguel Sainz and Renato Pajarola. Point-based rendering techniques. Computers & Graphics, 28(6):869–879, 2004.
  39. 39.Johannes Lutz Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  40. 40.Jessica Shue, Eric Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein. 3d neural field generation using triplane diffusion. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20875–20886, 2022.
  41. 41.Vincent Sitzmann, Semon Rezchikov, William T. Freeman, Joshua B. Tenenbaum, and Fredo Durand. Light field networks: Neural scene representations with single-evaluation rendering. In Neural Information Processing Systems, 2021.
  42. 42.Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo Tourism: Exploring Photo Collections in 3D. Association for Computing Machinery, New York, NY, USA, 1 edition, 2023.
  43. 43.Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P Srinivasan, Jonathan T Barron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8248–8258, 2022.
  44. 44.Shubham Tulsiani, Tinghui Zhou, Alyosha A. Efros, and Jitendra Malik. Multi-view supervision for single-view reconstruction via differentiable ray consistency. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 209–217, 2017.
  45. 45.Haithem Turki, Deva Ramanan, and Mahadev Satyanarayanan. Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12922–12931, 2022.
  46. 46.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021.
  47. 47.Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004.
  48. 48.Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. Synsin: End-to-end view synthesis from a single image. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7465–7475, 2019.
  49. 49.Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Anyi Rao, Christian Theobalt, Bo Dai, and Dahua Lin. Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering. In The European Conference on Computer Vision (ECCV), 2022.
  50. 50.Yuanbo Xiangli, Linning Xu, Xingang Pan, Nanxuan Zhao, Bo Dai, and Dahua Lin. Assetfield: Assets mining and reconfiguration in ground feature plane representation. ArXiv, abs/2303.13953, 2023.
  51. 51.Linning Xu, Vasu Agrawal, William Laney, Tony Garcia, Aayush Bansal, Changil Kim, Samuel Rota Bulo, Lorenzo Porzi, Peter Kontschieder, Aljaz Bozic, Dahua Lin, Michael Zollhofer, and Christian Richardt. VR-NeRF: High-fidelity virtualized walkable spaces. In SIGGRAPH Asia Conference Proceedings, 2023.
  52. 52.Linning Xu, Yuanbo Xiangli, Sida Peng, Xingang Pan, Nanxuan Zhao, Christian Theobalt, Bo Dai, and Dahua Lin. Grid-guided neural radiance fields for large urban scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8296–8306, 2023.
  53. 53.Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5438–5448, 2022.
  54. 54.Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. Volume rendering of neural implicit surfaces. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  55. 55.Wang Yifan, Felice Serena, Shihao Wu, Cengiz Oztireli, and Olga Sorkine-Hornung. Differentiable surface splatting for point-based geometry processing. ACM Transactions on Graphics (TOG), 38(6):1–14, 2019.
  56. 56.Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  57. 57.Xi Zhao, Ruizhen Hu, Haisong Liu, Taku Komura, and Xinyu Yang. Localization and completion for 3d object interactions. IEEE Transactions on Visualization and Computer Graphics, 26(8):2634–2644, 2019.
  58. 58.Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. Ewa volume splatting. In Proceedings Visualization, 2001. VIS’01., pages 29–538. IEEE, 2001.

Citation

MLA
Lu, T., et al. “Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering”. arXiv, 2023, http://arxiv.org/abs/2312.00109v1.
APA
Lu, T., Yu, M., Xu, L., Xiangli, Y., Wang, L., Lin, D., & Dai, B. (2023). Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering. arXiv. http://arxiv.org/abs/2312.00109v1
Chicago
Lu, T., M. Yu, L. Xu, et al. 2023. “Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering”. arXiv. http://arxiv.org/abs/2312.00109v1.
Harvard
Lu, T. et al. (2023) “Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.00109v1.
Vancouver
1. Lu T, Yu M, Xu L, Xiangli Y, Wang L, Lin D, Dai B (2023) Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering. arXiv

BibTeX

@article{lu2023scaffold,
  title = {Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering},
  author = {Lu, Tao and Yu, Mulin and Xu, Linning and Xiangli, Yuanbo and Wang, Limin and Lin, Dahua and Dai, Bo},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.00109v1},
  eprint = {2312.00109}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE