GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting
Yiwen ChenZilong ChenChi ZhangFeng WangXiaofeng YangYikai WangZhongang CaiLei YangHuaping LiuGuosheng Lin
Proposes GaussianEditor, the first 3D editing framework built on Gaussian Splatting that uses semantic tracing and hierarchical constraints to enable fast, highly localized object insertion, removal, and text-guided scene modification within minutes.
Fast and precise 3D content editing is increasingly essential for digital gaming, virtual reality, and simulation industries. While traditional 3D formats struggle to render complex scenes realistically, recent neural approaches like Neural Radiance Fields achieve high visual fidelity but suffer from slow processing speeds and imprecise localized control. The article addresses these bottlenecks by evaluating and demonstrating GaussianEditor, the first 3D editing algorithm based on Gaussian Splatting, an explicit point-based 3D scene representation that renders in real time.
The evaluated approach introduces three core mechanisms to make 3D editing swift and controllable. First, Gaussian semantic tracing projects 2D image segmentation masks into 3D space to assign dynamic semantic labels to individual points, allowing the algorithm to continuously track and modify only targeted objects as the model geometry changes. Second, Hierarchical Gaussian Splatting stabilizes the training process by grouping points into generations, applying anchor constraints to older base points while allowing newer points the flexibility needed to refine visual details. Third, dedicated 3D scene-repair workflows enable rapid object addition and removal by combining two-dimensional diffusion generation, depth alignment, and local boundary cleanup. The framework was evaluated on an RTX A6000 GPU across diverse tasks including style changes, facial edits, and object manipulation.
The findings show that GaussianEditor significantly outperforms existing neural field editing techniques across efficiency, visual fidelity, and user preference. A complete editing session with GaussianEditor takes only 5 to 10 minutes—with object removal taking roughly 2 minutes and object insertion taking 5 minutes—compared to over 30 minutes for Neural Radiance Field methods. In quantitative user studies, the proposed pipeline achieved a 72.28% preference rate over competing methods (which scored 15.45% and 12.27%), alongside superior text-image alignment metrics. Ablation tests confirmed that the generational hierarchical structure is essential to prevent uncontrolled blurring and point diffusion during AI-guided editing.
These results demonstrate that moving from implicit neural networks to explicit, semantically traced point representations substantially reduces computing time and production costs while eliminating unintended changes to surrounding background scenes. Organizations working in 3D content generation can achieve faster production turnarounds without sacrificing artistic control. Decision-makers should consider adopting Gaussian Splatting editing pipelines for interactive 3D workflows, though implementation teams should note that editing quality remains dependent on the underlying 2D AI generative models, which can struggle with highly complex or ambiguous text instructions.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). Introduces the foundational 3D Gaussian Splatting representation and real-time differentiable rendering pipeline that GaussianEditor adapts and extends for 3D editing.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). Establishes the iterative dataset-update paradigm using 2D diffusion editing models for 3D radiance fields that GaussianEditor builds upon and accelerates with Gaussian splatting.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). Provides the foundational stochastic differential equation and diffusion-based image editing formulations leveraged for 2D generative guidance in 3D scene modification.
- Paper: Decomposing NeRF for Editing via Feature Field Distillation, Sosuke Kobayashi et al. (2022). Pioneers semantic feature distillation into neural fields for query-based regional decomposition, informing the semantic tracing and object localization strategies in GaussianEditor.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Establishes the volumetric neural radiance field baseline whose rendering latency and editing limitations motivate the shift to explicit Gaussian Splatting.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). Applies generative diffusion guidance directly to optimize 3D Gaussian Splatting primitives for text-to-3D asset creation.
- Paper: 2D Gaussian Splatting for Geometrically Accurate Radiance Fields, Binbin Huang et al. (2024). Advances explicit Gaussian rendering by employing 2D surface disks for geometrically accurate reconstructions, directly refining explicit splatting geometry.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). Extends controllable 3D Gaussian representations to dynamic, animatable human head avatars bound to parametric meshes.
- Paper: 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis, Zhicheng Lu et al. (2024). Builds on dynamic 3D Gaussian representations by integrating explicit local geometric structures into deformation networks.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). Generalizes multi-representation 3D asset generation by learning structured latents capable of decoding into both Gaussian splats and explicit meshes.
