Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data
Yizhao XuHongyuan ZhuCaiyun LiuTianfu WangKeyu ChenSicheng XuJiaolong YangNicholas Jing YuanQi Zhang
Presents an efficient 3D editing framework paired with a large-scale dataset, utilizing lightweight modules and annotation-free 3D masking to execute precise, text-guided modifications while preserving unchanged geometry without full-model retraining.
Generative artificial intelligence has rapidly advanced 3D content creation across gaming, virtual reality, and manufacturing, yet creating production-grade assets in a single attempt remains difficult. Users often need to make local or global modifications to existing 3D models. Existing 3D editing solutions face major bottlenecks: multi-view 2D image editing introduces geometric artifacts when projected back to 3D, native voxel manipulations struggle to modify large structures, and optimization-based methods are too slow for interactive workflows. A critical underlying barrier has been the lack of large-scale, dedicated datasets designed to train and benchmark 3D editing models.
The article demonstrates an efficient framework, named Beyond Voxel 3D Editing, designed to execute high-fidelity, text-guided 3D object modifications while strictly preserving the identity and geometry of unedited regions. The primary objective is to evaluate whether integrating lightweight, trainable adaptation modules into a frozen image-to-3D foundation model can deliver precise local and global editing without requiring expensive full-model retraining.
To accomplish this, the authors constructed Edit-3DVerse, a curated dataset of over 100,000 instruction-paired 3D editing samples derived through automated filtering, multi-modal validation, and aesthetic scoring across 31 object categories. Building on this dataset, the framework augments a pre-trained two-stage generative model (TRELLIS) with zero-initialized cross-attention and feature modulation modules—freezing 1.1 billion parameters while training the remaining components. An automated, annotation-free 3D point cloud registration technique generates spatial masks, penalizing unintended deviations in unedited areas during the training process.
The experimental findings show that the proposed method significantly outperforms existing approaches across both objective and subjective benchmarks. Quantitatively, the framework achieved an image fidelity score of 28.9 on the Fréchet Inception Distance metric, substantially outperforming the nearest baseline at 119.83. It also demonstrated superior geometric fidelity with a Chamfer Distance of 0.013, roughly 38% better than the leading alternative at 0.021, and delivered the highest text-alignment score. In user preference evaluations involving 30 participants, the method secured a 91.0% preference rate for text alignment and 88.5% for visual quality, leading all compared tools by a wide margin. Ablation analyses confirmed that removing the 3D mask loss noticeably degrades structural preservation, and executing edits across both the coarse structure and high-resolution latent stages is vital for detailed geometry.
These findings indicate that 3D generative pipelines can achieve rapid, controllable editing within seconds without risking geometric distortion or cross-view noise. By relying on parameter-efficient fine-tuning rather than full-model retraining, the framework dramatically reduces computational costs and deployment overhead, providing an interactive foundation for commercial 3D design workflows.
Organizations developing 3D generation tools should consider adopting parameter-efficient text-modulation blocks and automated masking losses instead of relying solely on multi-view lifting or manual bounding-box annotations. To support broader adoption, further engineering is recommended to streamline the two-stage generation process into a faster, unified pass.
The reported results are constrained by two primary limitations noted in the article. Because the architecture operates in a two-stage sequential flow—editing sparse structures before generating detailed surface latents—inference latency is higher than single-pass generation systems. Additionally, the visual fidelity and aesthetic style remain inherently dependent on the underlying foundation backbone, occasionally transferring stylized characteristics from the base model. Confidence in the relative performance improvements is high across the evaluated asset categories, though broader validation on diverse production domains remains an area for future work.
- Paper: GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splatting, Yiwen Chen et al. (2024). It provides essential background on localized 3D editing and semantic masking techniques for maintaining regional consistency during content manipulation.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). It introduces foundational methodologies for instruction-guided 3D scene editing through iterative multi-view 2D diffusion updates.
- Paper: Decomposing NeRF for Editing via Feature Field Distillation, Sosuke Kobayashi et al. (2022). It establishes the concept of distilling 2D visual and linguistic foundation model features into 3D representations for annotation-free decomposition and editing.
- Paper: Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields, Shijie Zhou et al. (2024). It demonstrates how to incorporate semantic feature fields into explicit 3D representations to support language-guided localized scene modifications.
- Paper: Diffusion-SDF: Text-to-Shape via Voxelized Diffusion, Muheng Li et al. (2023). It details voxelized diffusion formulations for text-to-shape synthesis, clarifying the geometric representations and limitations that BVE seeks to transcend.
- Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). It provides the foundational large-scale 3D asset repository and methodology widely adapted to train and evaluate 3D generative and editing models.
No sufficiently relevant recommendations were found.
