Learning Locally Editable Virtual Humans
Hsuan-I HoLixin XueJie SongOtmar Hilliges
Proposes a hybrid neural representation anchored to skinned mesh vertices that enables fine-grained local editing, texture painting, and generative modeling of articulate 3D human avatars.
Creating realistic, customizable three-dimensional digital humans is essential for immersive gaming, virtual reality, and metaverse environments. Traditional digital asset creation relies heavily on specialized computer graphics expertise and manual mesh editing, while modern neural avatar systems often lack fine-grained, localized control. Existing generative models struggle either because pure two-dimensional supervision entangles shape and color or because conventional body meshes cannot easily capture complex, loose clothing topologies.
The article demonstrates an end-to-end generative framework and hybrid representation that enables the creation of fully poseable, highly detailed, and locally customizable three-dimensional human avatars. It evaluates the framework’s ability to fit unseen scans, sample diverse avatars, and support cross-subject garment transfers and two-dimensional texture authoring.
The proposed approach merges parametric mesh models with implicit neural fields. The system anchors learnable local feature codebooks to the vertices of a deformable body model, providing a consistent surface topology under movement. Separate, shared neural network decoders predict surface geometry and appearance solely from local triangle-relative coordinates, preventing the model from memorizing global body positions. The authors trained this auto-decoder architecture using combined three-dimensional reconstruction and two-dimensional adversarial losses across multiple subjects, supported by CustomHumans, a new dataset of over 600 volumetric scans covering 80 individuals and 120 outfits.
The findings confirm substantial performance improvements over existing avatar methods. First, the framework achieved significantly higher geometric accuracy when fitting unseen body scans; on the standard SIZER benchmark, it attained an error of 1.364 to 1.423 millimeters, outperforming previous generative approaches like gDNA (8.006 to 8.374 millimeters) and mesh extensions like SMPL+D (2.854 to 5.192 millimeters). Second, conditioning the network on local triangle coordinates rather than global space proved essential, preventing overfitting and ensuring avatars retain consistent details during reposing. Third, decoupling geometry and texture into separate feature branches allowed seamless partial swaps—such as transferring an upper-body jacket between scans—and direct two-dimensional drawing customization without distorting underlying body shape. Finally, expanding the training dataset from 10% to 100% improved model fitting accuracy by approximately 25% and eliminated movement artifacts caused by limb self-contact.
These results provide a practical path to lower production costs and reduce content development timelines for virtual worlds. By enabling modular asset reuse and non-specialist texture editing on fully animatable characters, the method removes major bottlenecks in digital human pipelines. Unlike older mesh-based systems that over-smooth complex garments or implicit models that cannot be edited locally, this hybrid method maintains high fidelity without sacrificing rigging consistency.
Organizations developing digital human pipelines should consider adopting localized, mesh-anchored neural representations to enable modular asset authoring. Teams implementing this workflow should ensure their capture stages provide sufficiently varied poses and subjects to train robust shared decoders. Users should note that while geometric fitting achieves high precision, edge performance remains bounded by the registration accuracy of underlying base body models and the diversity of training garments.
- Paper: SMPL, Matthew Loper et al. (2015). Introduces the SMPL parametric body model that provides the underlying mesh topology and skinning foundation anchored by the neural fields in the source paper.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). Establishes pixel-aligned implicit functions for clothed 3D human reconstruction, motivating the source's hybrid combination of parametric meshes and neural implicit fields.
- Paper: Geometry-Consistent Neural Shape Representation with Implicit Displacement Fields, Yifan Wang et al. (2022). Demonstrates decomposing complex 3D shapes into base surfaces and local displacement fields, directly informing the source's use of local coordinate-conditioned geometric decoders.
- Paper: Learning Implicit Fields for Generative Shape Modeling, Zhiqin Chen et al. (2018). Pioneers the use of neural implicit field decoders for continuous 3D generative shape modeling.
- Paper: End-to-End Recovery of Human Shape and Pose, Angjoo Kanazawa et al. (2017). Introduces end-to-end parametric human mesh recovery using adversarial priors to constrain 3D body shape and pose estimations.
- Paper: AMASS: Archive of Motion Capture As Surface Shapes, Naureen Mahmood et al. (2019). Provides the foundational large-scale motion capture surface dataset (AMASS) widely used to animate and repose statistical body models.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). Extends the concept of localized, anchored neural implicit fields to full head modeling with disentangled identity and expression spaces.
- Paper: Relightable Gaussian Codec Avatars, Shunsuke Saito et al. (2024). Advances high-fidelity avatar representation by combining animatable geometric control with relightable Gaussian codec models.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). Applies instructional editing concepts to 3D neural scenes, complementing the source's focus on local and modular 3D avatar editing.
- Paper: NeUDF: Leaning Neural Unsigned Distance Fields with Volume Rendering, Yu-Tao Liu et al. (2023). Expands neural implicit surface modeling to open, non-watertight surfaces such as complex loose garments via unsigned distance fields.
