GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

Yiwen ChenZilong ChenChi ZhangFeng WangXiaofeng YangYikai WangZhongang CaiLei YangHuaping LiuGuosheng Lin

article2024CVPR467 citations

Proposes GaussianEditor, the first 3D editing framework built on Gaussian Splatting that uses semantic tracing and hierarchical constraints to enable fast, highly localized object insertion, removal, and text-guided scene modification within minutes.

Listen

Fast and precise 3D content editing is increasingly essential for digital gaming, virtual reality, and simulation industries. While traditional 3D formats struggle to render complex scenes realistically, recent neural approaches like Neural Radiance Fields achieve high visual fidelity but suffer from slow processing speeds and imprecise localized control. The article addresses these bottlenecks by evaluating and demonstrating GaussianEditor, the first 3D editing algorithm based on Gaussian Splatting, an explicit point-based 3D scene representation that renders in real time.

The evaluated approach introduces three core mechanisms to make 3D editing swift and controllable. First, Gaussian semantic tracing projects 2D image segmentation masks into 3D space to assign dynamic semantic labels to individual points, allowing the algorithm to continuously track and modify only targeted objects as the model geometry changes. Second, Hierarchical Gaussian Splatting stabilizes the training process by grouping points into generations, applying anchor constraints to older base points while allowing newer points the flexibility needed to refine visual details. Third, dedicated 3D scene-repair workflows enable rapid object addition and removal by combining two-dimensional diffusion generation, depth alignment, and local boundary cleanup. The framework was evaluated on an RTX A6000 GPU across diverse tasks including style changes, facial edits, and object manipulation.

The findings show that GaussianEditor significantly outperforms existing neural field editing techniques across efficiency, visual fidelity, and user preference. A complete editing session with GaussianEditor takes only 5 to 10 minutes—with object removal taking roughly 2 minutes and object insertion taking 5 minutes—compared to over 30 minutes for Neural Radiance Field methods. In quantitative user studies, the proposed pipeline achieved a 72.28% preference rate over competing methods (which scored 15.45% and 12.27%), alongside superior text-image alignment metrics. Ablation tests confirmed that the generational hierarchical structure is essential to prevent uncontrolled blurring and point diffusion during AI-guided editing.

These results demonstrate that moving from implicit neural networks to explicit, semantically traced point representations substantially reduces computing time and production costs while eliminating unintended changes to surrounding background scenes. Organizations working in 3D content generation can achieve faster production turnarounds without sacrificing artistic control. Decision-makers should consider adopting Gaussian Splatting editing pipelines for interactive 3D workflows, though implementation teams should note that editing quality remains dependent on the underlying 2D AI generative models, which can struggle with highly complex or ambiguous text instructions.

arXiv: 2311.14521
Cover for GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting

Abstract

3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods, which rely on representations like meshes and point clouds, often fall short in realistically depicting complex scenes. On the other hand, methods based on implicit 3D representations, like Neural Radiance Field (NeRF), render complex scenes effectively but suffer from slow processing speeds and limited control over specific scene areas. In response to these challenges, our paper presents GaussianEditor, the first 3D editing algorithm based on Gaussian Splatting (GS), a novel 3D representation. GaussianEditor enhances precision and control in editing through our proposed Gaussian semantic tracing, which traces the editing target throughout the training process. Additionally, we propose Hierarchical Gaussian splatting (HGS) to achieve stabilized and fine results under stochastic generative guidance from 2D diffusion models. We also develop editing strategies for efficient object removal and integration, a challenging task for existing methods. Our comprehensive experiments demonstrate GaussianEditor's superior control, effective, and efficient performance, marking a significant advancement in 3D editing.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. 3D Representations
  • 2.2. 3D Editing
  • 3. Preliminary
  • 3.1. 3D Gaussian Splatting
  • 3.2. Diffusion-based Editing Guidance
  • 4. Method
  • 4.1. Gaussian Semantic Tracing
  • 4.2. Hierarchical Gaussian Splatting
  • 4.3. 3D Inpainting
  • 5. Experiments
  • 5.1. Implementation Details
  • 5.2. Qualitative Comparisons
  • 5.3. Quantitative Comparisons
  • 5.4. Ablation Study
  • 6. Conclusion
  • 7. Acknowledgements
  • References

Knowls

  1. Knowl 1 — GaussianEditor formulation for prompt-driven Gaussian Splatting edits

    model/method

    GaussianEditor treats a 3D scene as an explicit collection of Gaussian primitives. The ii-th Gaussian has parameters

    Θi={xi,si,qi,αi,ci},\Theta_i = \{x_i,s_i,q_i,\alpha_i,c_i\},

    where xi∈R3x_i\in\mathbb{R}^3 is position, si∈R3s_i\in\mathbb{R}^3 is scale, qi∈R4q_i\in\mathbb{R}^4 is a rotation quaternion, αi∈R\alpha_i\in\mathbb{R} is opacity, and cic_i contains spherical-harmonic color coefficients. Given the current scene parameters Θ\Theta, an editing prompt yy, and a rendered camera pose pp, the objective is to produce edited parameters Θy\Theta_y whose rendered appearance satisfies yy.

    GaussianEditor accepts existing 2D diffusion-based editing guidance rather than introducing a new diffusion loss. The guidance may operate on noisy renderings, as in score-distillation-style editing, or on multiview images edited by a 2D model. It is represented generically as DD, giving the editing loss

    LEdit=D(Θ;p,y).\mathcal{L}_{\mathrm{Edit}} = D(\Theta;p,y).

    The framework combines this guidance with Gaussian semantic tracing when a spatially restricted edit is required and with Hierarchical Gaussian Splatting when stable optimization under stochastic guidance is needed. A typical editing session optimizes for 500–1000 steps and takes approximately 5–10 minutes.

  2. Knowl 2 — Gaussian semantic tracing for dynamic 3D editing masks

    algorithm

    Gaussian semantic tracing assigns semantic labels directly to individual 3D Gaussians, so the edit region follows the evolving representation instead of remaining a fixed 2D or 3D mask. Each Gaussian receives a semantic-mask attribute mijm_i^j, where ii indexes a Gaussian and jj indexes a semantic category.

    The procedure renders the Gaussian scene from multiple posed viewpoints, applies a 2D segmentation model to obtain masks, and unprojects the masks to the Gaussians using inverse rendering. For pixel pp, the contribution accumulated for Gaussian ii and semantic label jj is

    wij=∑p∈Poi(p) Ti(p) Mj(p),w_i^j = \sum_{p\in\mathcal{P}} o_i(p)\,T_i(p)\,M^j(p),

    where P\mathcal{P} is the set of rendered pixels, oi(p)o_i(p) is the opacity contribution of Gaussian ii at pixel pp, Ti(p)T_i(p) is its rendering transmittance, and Mj(p)M^j(p) is the 2D segmentation-mask value for label jj. A counter records how many relevant pixel/view observations contributed to each Gaussian; the Gaussian is assigned label jj when its average accumulated weight exceeds a manually selected threshold.

    During densification, every newly created Gaussian inherits the semantic label of its parent. Consequently, the semantic mask remains synchronized with changes in Gaussian positions, appearances, and population. GaussianEditor uses these labels to update, densify, prune, or render only the Gaussians belonging to the target category. This permits edited content to expand beyond the initial projected mask while keeping unrelated scene regions unchanged, and it also allows small objects to be rendered independently for more precise 2D diffusion supervision.

  3. Knowl 3 — Hierarchical Gaussian Splatting with generation-dependent anchoring

    model/method

    Hierarchical Gaussian Splatting (HGS) organizes Gaussians by the densification round that created them. The initial Gaussians belong to generation 00; a Gaussian created during the kk-th densification round belongs to generation kk. Older generations receive stronger constraints, while newly created generations remain relatively unconstrained so that they can adapt to stochastic generative guidance and add fine details.

    At initialization, HGS stores an anchor value P^i\hat{P}_i for every Gaussian property PP of Gaussian ii. The properties are position xx, scale ss, rotation quaternion qq, opacity α\alpha, and color coefficients cc. For NN Gaussians, the anchor loss for one property is

    LanchorP=∑i=1Nλi∥Pi−P^i∥22,L_{\mathrm{anchor}}^{P} = \sum_{i=1}^{N}\lambda_i\left\|P_i-\hat{P}_i\right\|_2^2,

    where PiP_i is the current value, P^i\hat{P}_i is its stored anchor, and λi\lambda_i is larger for older generations. The anchors are updated at each densification process, and the full optimization objective is

    L=LEdit+∑P∈{x,s,q,α,c}λPLanchorP,\mathcal{L}=\mathcal{L}_{\mathrm{Edit}}+\sum_{P\in\{x,s,q,\alpha,c\}}\lambda_P L_{\mathrm{anchor}}^{P},

    where LEdit\mathcal{L}_{\mathrm{Edit}} is the diffusion-based editing loss and λP\lambda_P controls the importance of anchoring property PP. After each densification, the anchor weights of existing generations are increased, progressively reducing their mobility while leaving new Gaussians available to form details. Increasing the generation-0 weight preserves the original scene more strongly; increasing the weight for selected properties can selectively preserve geometry or color.

    HGS also replaces a manually chosen gradient threshold for densification with a percentage rule: at each densification step, only Gaussians whose 3D-position gradients lie in the top k%k\% are densified, where kk is the chosen percentage.

  4. Knowl 4 — Semantic-mask-guided object removal with boundary repair

    algorithm

    GaussianEditor removes an object by deleting the Gaussians selected by its semantic Gaussian mask, but direct deletion can expose artifacts where the object intersects the surrounding scene. The repair procedure is:

    • Delete the target Gaussians identified by Gaussian semantic tracing.
    • Apply K-nearest-neighbor search to find the remaining Gaussians nearest to the deleted set; these approximate the object–scene interface.
    • Project the interface Gaussians into multiple views.
    • Dilate the projected masks and fill holes to form a boundary mask covering the damaged interface.
    • Use a 2D image-inpainting method to generate repaired views for the boundary region.
    • Optimize the remaining Gaussian scene with mean-squared-error supervision to the repaired pixels.

    The complete removal process typically takes about two minutes. The boundary-specific mask and local repair avoid applying an unrestricted inpainting loss to the entire scene.

  5. Knowl 5 — Single-view object incorporation through 2D inpainting and coarse-to-fine 3D reconstruction

    algorithm

    For object incorporation, GaussianEditor takes an existing Gaussian scene θ\theta, a camera pose pp, the corresponding rendered image II, a user-provided 2D mask MM, and a prompt yy describing the desired object. The procedure is:

    • Generate a 2D inpainted image IyMI_y^M with a 2D diffusion inpainting model conditioned on MM and yy.
    • Segment the inserted foreground object from IyMI_y^M.
    • Convert the segmented single-view object into a coarse 3D mesh using an image-to-3D model.
    • Estimate the depth of IyMI_y^M and align it to the depth map rendered from θ\theta at pose pp using least-squares alignment. The alignment determines the inserted object's position and scale in the coordinate system of the original scene.
    • Convert the aligned coarse mesh into Gaussian primitives and refine them with HGS.
    • Concatenate the refined inserted Gaussians with the original Gaussian scene.

    The mesh generation takes approximately three minutes, Gaussian conversion and refinement take approximately two additional minutes, and concatenating the two Gaussian sets takes less than one second. The complete incorporation procedure therefore takes approximately five minutes.

  6. Knowl 6 — Experimental configuration and computational cost

    experimental setup

    GaussianEditor was implemented with an optimized differentiable Gaussian renderer and the Threestudio framework. The original Gaussian scenes were trained using the standard Gaussian-Splatting training procedure. All experiments used a single NVIDIA RTX A6000 GPU.

    For large scenes, editing camera poses were sampled from the multiview reconstruction dataset, with 24–96 poses depending on scene complexity. The 2D-to-3D semantic labeling stage typically required about one second. Prompt-based editing used 500–1000 optimization steps and took approximately 5–10 minutes, depending on scene complexity and the prompt. These settings were used for comparisons involving different diffusion guidance variants and for demonstrations of scene, face, and object edits.

  7. Knowl 7 — Quantitative superiority in user preference and CLIP directional similarity

    data/table

    The reported quantitative comparison evaluates three editing configurations using a user-study preference percentage and CLIP Directional Similarity. The two GaussianEditor variants use either delta-denoising-score guidance or Instruct-NeRF2NeRF-style image-to-image guidance; the latter is the primary GaussianEditor configuration.

    Could not parse LaTeX table

    GaussianEditor with iN2N guidance has both the highest user-study preference, 72.28%, and the highest CLIP Directional Similarity, 0.2071. The method also reduces the reported editing time from typically more than 30 minutes for Instruct-NeRF2NeRF to approximately 5–10 minutes.

  8. Knowl 8 — Qualitative control across localized and diverse edits

    empirical result

    The qualitative experiments show that GaussianEditor can perform localized appearance and semantic edits while preserving non-target regions. In edits such as converting a person into a clown, Albert Einstein, or a Modigliani-style subject, the method confines changes to the selected human body; the background and other scene content remain essentially unchanged.

    Compared with Instruct-NeRF2NeRF, GaussianEditor-iN2N produces more detailed edits and substantially better control over the editing area. Instruct-NeRF2NeRF often changes the entire scene and produces insufficient detail. GaussianEditor-DDS tends to produce oversaturated colors and less precise edits. The method also demonstrates scene-level changes such as snow or autumn appearance, object and face edits, and bear appearance changes. For face swaps, tracing the facial Gaussians keeps the edit localized; for large scenes, the method can still edit successfully without applying semantic tracing.

  9. Knowl 9 — HGS prevents uncontrolled Gaussian spreading under stochastic guidance

    empirical result

    The HGS ablation compares GaussianEditor with and without generation-dependent anchoring on an edit that makes grass appear to be on fire. Without HGS, the stochastic 2D guidance causes Gaussians to spread and densify across much of the scene, because the representation follows whole-image changes in the 2D guidance. The resulting uncontrolled densification blurs the image and damages non-target content.

    With HGS, older Gaussian generations are constrained by their anchors while newer generations remain available for adaptation. This suppresses scene-wide Gaussian mobility and confines the edit more effectively, producing a sharper and more controllable result.

  10. Knowl 10 — Dependence on the capabilities of 2D diffusion guidance

    limitation

    GaussianEditor relies on 2D diffusion models to provide effective supervision for 3D editing. As a result, prompts that current 2D diffusion models cannot interpret or guide reliably remain difficult for GaussianEditor as well. This limitation can lead to unsuccessful or low-quality 3D edits for sufficiently complex prompts.

Coverage note — No substantial contributed material was omitted; standard Gaussian-Splatting background, related work, and acknowledgements were excluded because they are not contributions of the paper.

References

  1. 1.Chong Bao, Yinda Zhang, and Bangbang et al. Yang. Sine: Semantic-driven image-based nerf editing with prior-guided editing field. In CVPR 2023, pages 20919–20929, 2023.
  2. 2.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021.
  3. 3.Jonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. CVPR, 2022.
  4. 4.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. arXiv preprint arXiv:2211.09800, 2022.
  5. 5.Yiwen Chen, Chi Zhang, Xiaofeng Yang, Zhongang Cai, Gang Yu, Lei Yang, and Guosheng Lin. It3d: Improved text-to-3d generation with explicit view synthesis. arXiv preprint arXiv:2308.11473, 2023.
  6. 6.Zhiqin Chen, Thomas Funkhouser, Peter Hedman, and Andrea Tagliasacchi. Mobilenerf: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures. arXiv preprint arXiv:2208.00277, 2022.
  7. 7.Zilong Chen, Feng Wang, and Huaping Liu. Text-to-3d using gaussian splatting. arXiv preprint arXiv:2309.16585, 2023.
  8. 8.Xinhua Cheng, Tianyu Yang, Jianan Wang, Yu Li, Lei Zhang, Jian Zhang, and Li Yuan. Progressive3d: Progressively local editing for text-to-3d content creation with complex semantic prompts. arXiv preprint arXiv:2310.11784, 2023.
  9. 9.Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Stylegan-nada: Clip-guided domain adaptation of image generators. ACM Transactions on Graphics (TOG), 41(4):1–13, 2022.
  10. 10.William Gao, Noam Aigerman, Thibault Groueix, Vladimir G Kim, and Rana Hanocka. Textdeformer: Geometry manipulation using text guidance. arXiv preprint arXiv:2304.13348, 2023.
  11. 11.Yuan-Chen Guo, Ying-Tian Liu, Chen Wang, Zi-Xin Zou, Guan Luo, Chia-Hao Chen, Yan-Pei Cao, and Song-Hai Zhang. threestudio: A unified framework for 3d content generation. https://github.com/threestudio-project/threestudio, 2023.
  12. 12.Ayaan Haque, Matthew Tancik, Alexei A Efros, Aleksander Holynski, and Angjoo Kanazawa. Instruct-nerf2nerf: Editing 3d scenes with instructions. arXiv preprint arXiv:2303.12789, 2023.
  13. 13.Peter Hedman, Pratul P. Srinivasan, Ben Mildenhall, Jonathan T. Barron, and Paul Debevec. Baking neural radiance fields for real-time view synthesis. ICCV, 2021.
  14. 14.Amir Hertz, Kfir Aberman, and Daniel Cohen-Or. Delta denoising score. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2328–2337, 2023.
  15. 15.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023.
  16. 16.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ToG, 42(4):1–14, 2023.
  17. 17.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4), 2023.
  18. 18.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  19. 19.Sosuke Kobayashi, Eiichi Matsumoto, and Vincent Sitzmann. Decomposing nerf for editing via feature field distillation. arXiv preprint arXiv:2205.15585, 2022.
  20. 20.Yuan Li, Zhi-Hao Lin, David Forsyth, Jia-Bin Huang, and Shenlong Wang. Climatenerf: Physically-based neural rendering for extreme climate synthesis. arXiv e-prints, pages arXiv–2211, 2022.
  21. 21.Zhaoshuo Li, Thomas Muller, Alex Evans, Russell H Taylor, Mathias Unberath, Ming-Yu Liu, and Chen-Hsuan Lin. Neuralangelo: High-fidelity neural surface reconstruction. In CVPR, 2023.
  22. 22.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, pages 300–309, 2023.
  23. 23.Hao-Kang Liu, I Shen, Bing-Yu Chen, et al. Nerf-in: Free-form nerf inpainting with rgb-d priors. arXiv preprint arXiv:2206.04901, 2022.
  24. 24.Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, and Bryan Russell. Editing conditional radiance fields. In ICCV 2021, pages 5773–5783, 2021.
  25. 25.Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3d: Single image to 3d using cross-domain diffusion. arXiv preprint arXiv:2310.15008, 2023.
  26. 26.Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713, 2023.
  27. 27.Aryan Mikaeili, Or Perel, Mehdi Safaee, Daniel Cohen-Or, and Ali Mahdavi-Amiri. Sked: Sketch-guided text-based 3d editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14607–14619, 2023.
  28. 28.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  29. 29.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, 2022.
  30. 30.Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Harada. Neural articulated radiance field. In ICCV 2021, pages 5762–5772, 2021.
  31. 31.Jangho Park, Gihyun Kwon, and Jong Chul Ye. Ed-nerf: Efficient text-guided editing of 3d scene using latent space nerf. arXiv preprint arXiv:2310.02712, 2023.
  32. 32.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. arXiv preprint arXiv:2106.13228, 2021.
  33. 33.Sida Peng, Yuanqing Zhang, Yinghao Xu, and et al. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR 2021, pages 9054–9063, 2021.
  34. 34.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  35. 35.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022.
  36. 36.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023.
  37. 37.Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan Barron, et al. Dreambooth3d: Subject-driven text-to-3d generation. arXiv preprint arXiv:2303.13508, 2023.
  38. 38.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision transformers for dense prediction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12179–12188, 2021.
  39. 39.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10684–10695, 2022.
  40. 40.Sara Fridovich-Keil and Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In CVPR, 2022.
  41. 41.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016.
  42. 42.Etai Sella, Gal Fiebelman, Peter Hedman, and Hadar Averbuch-Elor. Vox-e: Text-guided voxel editing of 3d objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 430–440, 2023.
  43. 43.Ruizhi Shao, Jingxiang Sun, Cheng Peng, Zerong Zheng, Boyao Zhou, Hongwen Zhang, and Yebin Liu. Control4d: Dynamic portrait editing by learning 4d gan from 2d diffusion-based editor. arXiv preprint arXiv:2305.20082, 2023.
  44. 44.Chaoyue Song, Tianyi Chen, Yiwen Chen, Jiacheng Wei, Chuan Sheng Foo, Fayao Liu, and Guosheng Lin. Moda: Modeling deformable 3d objects from casual videos. arXiv preprint arXiv:2304.08279, 2023.
  45. 45.Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023.
  46. 46.Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Clip-nerf: Text-and-image driven manipulation of neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3835–3844, 2022.
  47. 47.Can Wang, Ruixiang Jiang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. Nerf-art: Text-driven neural radiance fields stylization. IEEE Transactions on Visualization and Computer Graphics, 2023.
  48. 48.Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. arXiv preprint arXiv:2310.08528, 2023.
  49. 49.Tianhan Xu and Tatsuya Harada. Deforming radiance fields with cages. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXIII, pages 159–175. Springer, 2022.
  50. 50.Bangbang Yang, Chong Bao, and Junyi et al. Zeng. Neumesh: Learning disentangled neural mesh-based implicit field for geometry and texture editing. In ECCV 2022, pages 597–614. Springer, 2022.
  51. 51.Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction. arXiv preprint arXiv:2309.13101, 2023.
  52. 52.Zeyu Yang, Hongye Yang, Zijie Pan, Xiatian Zhu, and Li Zhang. Real-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting. arXiv preprint arXiv:2310.10642, 2023.
  53. 53.Taoran Yi, Jiemin Fang, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Qi Tian, and Xinggang Wang. Gaussiandreamer: Fast generation from text to 3d gaussian splatting with point cloud priors. arXiv preprint arXiv:2310.08529, 2023.
  54. 54.Wang Yifan, Felice Serena, Shihao Wu, Cengiz Oztireli, and Olga Sorkine-Hornung. Differentiable surface splatting for point-based geometry processing. ACM Transactions on Graphics (TOG), 38(6):1–14, 2019.
  55. 55.Yu-Jie Yuan, Yang-Tian Sun, Yu-Kun Lai, and et al. Nerf-editing: geometry editing of neural radiance fields. In CVPR 2022, pages 18353–18364, 2022.
  56. 56.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020.
  57. 57.Jingyu Zhuang, Chen Wang, Lingjie Liu, Liang Lin, and Guanbin Li. Dreameditor: Text-driven 3d scene editing with neural fields. arXiv preprint arXiv:2306.13455, 2023.
  58. 58.Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. Surface splatting. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 371–378, 2001.

Citation

MLA
Chen, Y., et al. “GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting”. arXiv, 2023, http://arxiv.org/abs/2311.14521v4.
APA
Chen, Y., Chen, Z., Zhang, C., Wang, F., Yang, X., Wang, Y., Cai, Z., Yang, L., Liu, H., & Lin, G. (2023). GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting. arXiv. http://arxiv.org/abs/2311.14521v4
Chicago
Chen, Y., Z. Chen, C. Zhang, et al. 2023. “GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting”. arXiv. http://arxiv.org/abs/2311.14521v4.
Harvard
Chen, Y. et al. (2023) “GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.14521v4.
Vancouver
1. Chen Y, Chen Z, Zhang C, Wang F, Yang X, Wang Y, Cai Z, Yang L, Liu H, Lin G (2023) GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting. arXiv

BibTeX

@article{chen2023gaussianeditor,
  title = {GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting},
  author = {Chen, Yiwen and Chen, Zilong and Zhang, Chi and Wang, Feng and Yang, Xiaofeng and Wang, Yikai and Cai, Zhongang and Yang, Lei and Liu, Huaping and Lin, Guosheng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.14521v4},
  eprint = {2311.14521}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE