Learning Locally Editable Virtual Humans

Hsuan-I HoLixin XueJie SongOtmar Hilliges

article2023CVPR65 citations

Proposes a hybrid neural representation anchored to skinned mesh vertices that enables fine-grained local editing, texture painting, and generative modeling of articulate 3D human avatars.

Listen

Creating realistic, customizable three-dimensional digital humans is essential for immersive gaming, virtual reality, and metaverse environments. Traditional digital asset creation relies heavily on specialized computer graphics expertise and manual mesh editing, while modern neural avatar systems often lack fine-grained, localized control. Existing generative models struggle either because pure two-dimensional supervision entangles shape and color or because conventional body meshes cannot easily capture complex, loose clothing topologies.

The article demonstrates an end-to-end generative framework and hybrid representation that enables the creation of fully poseable, highly detailed, and locally customizable three-dimensional human avatars. It evaluates the framework’s ability to fit unseen scans, sample diverse avatars, and support cross-subject garment transfers and two-dimensional texture authoring.

The proposed approach merges parametric mesh models with implicit neural fields. The system anchors learnable local feature codebooks to the vertices of a deformable body model, providing a consistent surface topology under movement. Separate, shared neural network decoders predict surface geometry and appearance solely from local triangle-relative coordinates, preventing the model from memorizing global body positions. The authors trained this auto-decoder architecture using combined three-dimensional reconstruction and two-dimensional adversarial losses across multiple subjects, supported by CustomHumans, a new dataset of over 600 volumetric scans covering 80 individuals and 120 outfits.

The findings confirm substantial performance improvements over existing avatar methods. First, the framework achieved significantly higher geometric accuracy when fitting unseen body scans; on the standard SIZER benchmark, it attained an error of 1.364 to 1.423 millimeters, outperforming previous generative approaches like gDNA (8.006 to 8.374 millimeters) and mesh extensions like SMPL+D (2.854 to 5.192 millimeters). Second, conditioning the network on local triangle coordinates rather than global space proved essential, preventing overfitting and ensuring avatars retain consistent details during reposing. Third, decoupling geometry and texture into separate feature branches allowed seamless partial swaps—such as transferring an upper-body jacket between scans—and direct two-dimensional drawing customization without distorting underlying body shape. Finally, expanding the training dataset from 10% to 100% improved model fitting accuracy by approximately 25% and eliminated movement artifacts caused by limb self-contact.

These results provide a practical path to lower production costs and reduce content development timelines for virtual worlds. By enabling modular asset reuse and non-specialist texture editing on fully animatable characters, the method removes major bottlenecks in digital human pipelines. Unlike older mesh-based systems that over-smooth complex garments or implicit models that cannot be edited locally, this hybrid method maintains high fidelity without sacrificing rigging consistency.

Organizations developing digital human pipelines should consider adopting localized, mesh-anchored neural representations to enable modular asset authoring. Teams implementing this workflow should ensure their capture stages provide sufficiently varied poses and subjects to train robust shared decoders. Users should note that while geometric fitting achieves high precision, edge performance remains bounded by the registration accuracy of underlying base body models and the diversity of training garments.

arXiv: 2305.00121
Cover for Learning Locally Editable Virtual Humans

Abstract

In this paper, we propose a novel hybrid representation and end-to-end trainable network architecture to model fully editable and customizable neural avatars. At the core of our work lies a representation that combines the modeling power of neural fields with the ease of use and inherent 3D consistency of skinned meshes. To this end, we construct a trainable feature codebook to store local geometry and texture features on the vertices of a deformable body model, thus exploiting its consistent topology under articulation. This representation is then employed in a generative auto-decoder architecture that admits fitting to unseen scans and sampling of realistic avatars with varied appearances and geometries. Furthermore, our representation allows local editing by swapping local features between 3D assets. To verify our method for avatar creation and editing, we contribute a new high-quality dataset, dubbed CustomHumans, for training and evaluation. Our experiments quantitatively and qualitatively show that our method generates diverse detailed avatars and achieves better model fitting performance compared to state-of-the-art methods. Our code and dataset are available at https://ait.ethz.ch/custom-humans.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Hybrid Representation of Humans
  • 3.2. Generative Codebook Sampling
  • 3.3. Model Training
  • 3.4. Feature Editing and Avatar Customization
  • 4. Experiments
  • 4.1. Experiment Settings
  • 4.2. Customized Avatars
  • 4.3. Model Fitting Comparison
  • 4.4. Ablation Study
  • 4.5. Generalization Ability Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Locally conditioned hybrid neural-field representation

    model/method

    The avatar representation combines an explicitly deformable LBS body mesh with neural fields that encode detailed geometry and appearance. A posed coarse body mesh M\mathcal{M} is produced from pose and shape parameters (θ,β)(\boldsymbol{\theta},\boldsymbol{\beta}). The mesh has MM vertices with positions V∈RM×3\mathbf{V}\in\mathbb{R}^{M\times 3} and a fixed triangular topology shared by all subjects.

    Each mesh vertex stores two learnable local feature vectors of dimension KK: a geometry feature and a texture feature. The resulting codebook is C∈RM×2K\mathbf{C}\in\mathbb{R}^{M\times 2K}. For a query point xg∈R3\mathbf{x}_g\in\mathbb{R}^3, the closest point xc∗\mathbf{x}_c^* on the nearest mesh triangle is found by minimizing Euclidean distance. If the triangle has vertex indices (i0,i1,i2)(i_0,i_1,i_2) and barycentric coordinates (u,v,1−u−v)(u,v,1-u-v), its closest point is

    xc∗=arg⁡min⁡xc∈M∥xg−xc∥2,xc=Bu,v(V[i0,i1,i2]).\mathbf{x}_c^*=\arg\min_{\mathbf{x}_c\in\mathcal{M}}\lVert\mathbf{x}_g-\mathbf{x}_c\rVert_2,\qquad \mathbf{x}_c=\mathcal{B}_{u,v}\big(\mathbf{V}[i_0,i_1,i_2]\big).

    The query is represented locally by xl=(u,v,d)∈R3\mathbf{x}_l=(u,v,d)\in\mathbb{R}^3, where dd is the signed distance from xg\mathbf{x}_g to xc∗\mathbf{x}_c^*, together with a direction vector n\mathbf{n} from the closest surface point toward the query. The three geometry features and three texture features stored at (i0,i1,i2)(i_0,i_1,i_2) are barycentrically interpolated to obtain fs,fc∈RK\mathbf{f}_s,\mathbf{f}_c\in\mathbb{R}^K.

    Two shared MLP decoders use only local information: the geometry decoder Φ\Phi predicts the signed distance s(xg)s(\mathbf{x}_g) from (fs,xl,n)(\mathbf{f}_s,\mathbf{x}_l,\mathbf{n}), and the texture decoder Ψ\Psi predicts the RGB color r(xg)∈R3\mathbf{r}(\mathbf{x}_g)\in\mathbb{R}^3 from (fc,xl,n)(\mathbf{f}_c,\mathbf{x}_l,\mathbf{n}). Because the decoders do not receive global coordinates, the same decoders can be shared across vertices and subjects without directly memorizing subject- or vertex-specific global information.

  2. Knowl 2 — Generative sampling of geometry and texture codebooks

    model/method

    A multi-subject model stores one geometry codebook and one texture codebook per training scan in dictionaries Ds,Dc∈RN×(MK)\mathbf{D}_s,\mathbf{D}_c\in\mathbb{R}^{N\times(MK)}, where NN is the number of scans, MM is the number of body-mesh vertices, and KK is the feature dimension for one branch. The identical mesh topology across scans makes corresponding dictionary rows represent corresponding local body regions.

    To obtain a latent space from which new avatars can be sampled, the method applies PCA to the flattened codebooks. For each branch, the PCA coefficients of the NN training codebooks are modeled by a DD-dimensional Gaussian distribution, where DD is the retained PCA dimension. A random DD-dimensional coefficient vector is sampled from this distribution and multiplied by the learned PCA eigenvectors to produce a random codebook. Geometry and texture use separate dictionaries, PCA spaces, and decoders, so their random features can be sampled independently. Directly optimizing codebooks with only per-scan 3D reconstruction supervision was found insufficient to produce a useful latent space; the adversarial training described in the paper is used to regularize these random samples.

  3. Knowl 3 — Joint 3D reconstruction and 2D adversarial training

    equation

    For a sampled query point xg\mathbf{x}_g, let ss be its ground-truth signed distance, n\mathbf{n} its ground-truth unit surface normal, and r\mathbf{r} its ground-truth RGB color. The predicted signed distance and color are s(xg)s(\mathbf{x}_g) and r(xg)\mathbf{r}(\mathbf{x}_g), and ∇xgs(xg)\nabla_{\mathbf{x}_g}s(\mathbf{x}_g) is the spatial gradient of the predicted signed-distance field. The 3D supervision consists of

    Lsdf=∥s−s(xg)∥1+λn∥1−n⋅∇xgs(xg)∥1,\mathcal{L}_{\mathrm{sdf}}= \left\lVert s-s(\mathbf{x}_g)\right\rVert_1+ \lambda_n\left\lVert 1-\mathbf{n}\cdot\nabla_{\mathbf{x}_g}s(\mathbf{x}_g)\right\rVert_1, Lrgb=∥r−r(xg)∥1,L3D=λsdfLsdf+λrgbLrgb,\mathcal{L}_{\mathrm{rgb}}=\left\lVert\mathbf{r}-\mathbf{r}(\mathbf{x}_g)\right\rVert_1, \qquad \mathcal{L}_{3D}=\lambda_{\mathrm{sdf}}\mathcal{L}_{\mathrm{sdf}}+\lambda_{\mathrm{rgb}}\mathcal{L}_{\mathrm{rgb}},

    where λn\lambda_n, λsdf\lambda_{\mathrm{sdf}}, and λrgb\lambda_{\mathrm{rgb}} are loss weights. Query points are sampled in a thin shell around each training scan.

    The adversarial supervision uses real color and normal image patches obtained by rasterizing a ground-truth scan, and fake patches rendered from randomly sampled codebooks using the same virtual camera and coarse body mesh. With non-saturating logistic adversarial loss Ladv\mathcal{L}_{\mathrm{adv}}, R1 discriminator regularization LR1\mathcal{L}_{R1}, path-length regularization Lpath\mathcal{L}_{\mathrm{path}}, and feature-dictionary Gaussian regularization Lreg=∥D∥F\mathcal{L}_{\mathrm{reg}}=\lVert\mathbf{D}\rVert_F, the discriminator is optimized with

    Ldis=Ladv+λR1LR1,\mathcal{L}_{\mathrm{dis}}=\mathcal{L}_{\mathrm{adv}}+\lambda_{R1}\mathcal{L}_{R1},

    while the geometry and texture dictionaries and the shared decoders are optimized with

    L=−Ladv+L3D+λpathLpath+λregLreg.\mathcal{L}=-\mathcal{L}_{\mathrm{adv}}+\mathcal{L}_{3D}+\lambda_{\mathrm{path}}\mathcal{L}_{\mathrm{path}}+\lambda_{\mathrm{reg}}\mathcal{L}_{\mathrm{reg}}.

    Here λR1\lambda_{R1}, λpath\lambda_{\mathrm{path}}, and λreg\lambda_{\mathrm{reg}} balance the corresponding regularizers. The 3D terms preserve scan geometry and color, whereas the 2D adversarial term improves the realism and usefulness of randomly sampled codebooks.

  4. Knowl 4 — Feature inversion, local editing, texture drawing, and reposing

    model/method

    Avatar creation begins either from a codebook directly retrieved from the trained dictionaries or from a randomly sampled codebook. To fit an unseen 3D scan, the shared geometry and texture decoders are frozen and only a new codebook Cfit\mathbf{C}_{\mathrm{fit}} is optimized against the scan's signed distances and RGB colors:

    Cfit=arg⁡min⁡C(λsdfLsdf+λrgbLrgb).\mathbf{C}_{\mathrm{fit}}=\arg\min_{\mathbf{C}}\left(\lambda_{\mathrm{sdf}}\mathcal{L}_{\mathrm{sdf}}+\lambda_{\mathrm{rgb}}\mathcal{L}_{\mathrm{rgb}}\right).

    Because features are indexed by the vertices of a common body topology, a user can select a vertex subset corresponding to a clothing or body region and copy the corresponding rows of one fitted codebook into another. This transfers local geometry, texture, or both across subjects while leaving other regions unchanged.

    For 2D personalization, only the texture features of a fitted codebook are optimized using the RGB loss at the 3D points corresponding to user-edited image pixels; geometry features remain fixed. This permits logos, letters, and other drawn patterns to be transferred from edited images onto the 3D avatar.

    For reposing, the coarse LBS mesh is transformed with new pose and shape parameters (θ,β)(\boldsymbol{\theta},\boldsymbol{\beta}). The learned local geometry and texture features do not encode the original global pose, so the same local details can be applied consistently after the body is reposed.

  5. Knowl 5 — CustomHumans high-quality 3D scan dataset

    data/table

    The paper contributes CustomHumans, a dataset for training and evaluating generative 3D human avatars. It contains more than 600 high-quality scans of 80 participants wearing 120 garments in varied poses. The scans were captured with a volumetric stage containing 106 synchronized cameras: 53 RGB cameras and 53 infrared cameras. The released data include detailed meshes and accurately registered SMPL-X body models, providing the body pose and shape correspondence required by the proposed representation.

    The quantitative experiments use CustomHumans for training. THuman2.0, with approximately 500 scans of people wearing 150 garments, is used for qualitative random-sampling experiments because it provides greater texture diversity. SIZER provides 97 A-pose subjects in 22 garments and is used as an unseen test set for model fitting. Fitting quality is evaluated with Chamfer distance, normal consistency, and f-score.

  6. Knowl 6 — Unseen-scan fitting substantially outperforms baselines

    data/table

    The fitting experiment inverts unseen textured SIZER scans into the latent representation while keeping each method's remaining model parameters fixed. Chamfer distance is reported in millimeters in both prediction-to-scan and scan-to-prediction directions; lower is better. Normal consistency and f-score are dimensionless and higher is better. The proposed method retains high-frequency clothing and appearance details that fixed-resolution mesh extensions tend to smooth away.

    Could not parse LaTeX table

    The proposed method achieves the lowest Chamfer distances and the highest normal consistency and f-score among SMPL, gDNA, and SMPL+D. Qualitatively, it fits difficult loose garments such as jackets and loose T-shirts while preserving detailed surfaces, and it produces substantially more detailed textures than SMPL+D when fitting unseen textured meshes.

  7. Knowl 7 — Demonstrated avatar generation and customization capabilities

    empirical result

    Random codebooks sampled from the model trained on THuman2.0 generate plausible body geometry and appearance in arbitrary poses. The sampled geometry includes details such as garment wrinkles and facial expressions, while independently sampled texture features produce reasonable skin, hair, clothing, and trouser colors.

    Fitting codebooks to unseen scans enables cross-subject clothing transfer: upper- and lower-body local features can be copied between different avatars, including across different subjects, garments, body shapes, and poses. The transferred details remain spatially consistent after reposing.

    The same representation supports personalized texture authoring. User-drawn letters and logos fitted through the texture branch are applied seamlessly to the 3D avatar and remain attached to the clothing under changes to the SMPL-X pose parameters. These experiments demonstrate that the representation supports random avatar creation, local geometry and appearance transfer, 2D-to-3D texture editing, and pose-consistent output rather than only whole-avatar reconstruction.

  8. Knowl 8 — Local conditioning prevents decoder memorization

    empirical result

    Replacing the local query features (xl,n)(\mathbf{x}_l,\mathbf{n}) with global query coordinates xg\mathbf{x}_g gives similar reconstruction quality on the training scans, but it fails to preserve that quality when fitting unseen scans or reposing avatars. The shared decoders conditioned on global coordinates can memorize subject- and location-specific information instead of learning features transferable across vertices and subjects.

    With local triangle coordinates, signed distance, direction, and interpolated vertex features as decoder inputs, the shared decoders receive only locally reusable information. This version handles unseen body poses and out-of-distribution scans more consistently, which is necessary for local feature editing and avatar reposing.

  9. Knowl 9 — Adversarial learning and branch disentanglement are necessary for sampling

    empirical result

    Two ablations show that both components of the generative training design are needed. When the 2D adversarial loss is removed, random samples from the learned feature spaces do not produce reasonable body textures. When geometry and texture are modeled with a single decoder rather than separate branches, transferring or randomly sampling texture features can alter the intended body geometry.

    The complete model, which combines separate geometry and texture codebooks and decoders with 2D adversarial training, can apply random textures to fixed body geometries while preserving those geometries. It therefore provides the disentangled sampling behavior required for independent appearance and shape customization.

  10. Knowl 10 — More training subjects improve generalization

    data/table

    The model-fitting evaluation measures how the shared decoders generalize as the amount of CustomHumans training data changes. Here 100% corresponds to 100 training scans. Chamfer distance is reported in millimeters as scan-to-prediction / prediction-to-scan; lower is better. Normal consistency and f-score are higher-is-better metrics.

    Could not parse LaTeX table

    Increasing the number of training scans improves every reported metric, with the full training set yielding the best fitting performance and an approximately 25% accuracy improvement relative to the smallest training setting as reported by the authors. Qualitatively, more training subjects and poses reduce reposing artifacts caused by self-contact around regions such as fists and elbows, and make 2D texture fitting robust to a wider range of color distributions.

Coverage note — No substantial contributed material from the main paper was omitted; supplementary implementation details and background or related-work material were excluded because they do not constitute additional standalone contributions.

References

  1. 1.https://3dpeople.com/.
  2. 2.https://renderpeople.com/.
  3. 3.Thiemo Alldieck, Marcus Magnor, Bharat Lal Bhatnagar, Christian Theobalt, and Gerard Pons-Moll. Learning to reconstruct people in clothing from a single rgb camera. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1175–1186, 2019.
  4. 4.Thiemo Alldieck, Mihai Zanfir, and Cristian Sminchisescu. Photorealistic monocular 3d reconstruction of humans wearing clothing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  5. 5.Alexander W Bergman, Petr Kellnhofer, Yifan Wang, Eric R Chan, David B Lindell, and Gordon Wetzstein. Generative neural articulated radiance fields. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  6. 6.Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Combining implicit function learning and parametric models for 3d human reconstruction. In Proceedings of the European Conference on Computer Vision (ECCV), pages 311–329. Springer, 2020.
  7. 7.Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. Everybody dance now. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 5933–5942, 2019.
  8. 8.Xu Chen, Tianjian Jiang, Jie Song, Max Rietmann, Andreas Geiger, Michael J. Black, and Otmar Hilliges. Fast-SNARF: A fast deformer for articulated neural fields. arXiv, abs/2211.15601, 2022.
  9. 9.Xu Chen, Tianjian Jiang, Jie Song, Jinlong Yang, Michael J Black, Andreas Geiger, and Otmar Hilliges. gdna: Towards generative detailed neural avatars. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 20427–20437, 2022.
  10. 10.Xu Chen, Yufeng Zheng, Michael J Black, Otmar Hilliges, and Andreas Geiger. Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021.
  11. 11.Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Dennis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. High-quality streamable free-viewpoint video. ACM Transactions on Graphics (TOG), 34(4):1–13, 2015.
  12. 12.Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018.
  13. 13.Enric Corona, Albert Pumarola, Guillem Alenyà, Gerard Pons-Moll, and Francesc Moreno-Noguer. Smplicit: Topology-aware generative model for clothed people. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  14. 14.Boyang Deng, JP Lewis, Timothy Jeruzalski, Gerard Pons-Moll, Geoffrey Hinton, Mohammad Norouzi, and Andrea Tagliasacchi. Neural articulated shape approximation. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, August 2020.
  15. 15.Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W Taylor, and Joshua S Susskind. Unconstrained scene generation with locally conditioned radiance fields. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 14304–14313, 2021.
  16. 16.Haoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu, Bing-Cheng Chen, and Jian Yin. Fw-gan: Flow-navigated warping gan for video virtual try-on. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 1161–1170, 2019.
  17. 17.Xin Dong, Fuwei Zhao, Zhenyu Xie, Xijin Zhang, Daniel K. Du, Min Zheng, Xiang Long, Xiaodan Liang, and Jianchao Yang. Dressing in the wild by watching dance videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3480–3489, June 2022.
  18. 18.Jianglin Fu, Shikai Li, Yuming Jiang, Kwan-Yee Lin, Chen Qian, Chen Change Loy, Wayne Wu, and Ziwei Liu. Stylegan-human: A data-centric odyssey of human generation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 1–19. Springer, 2022.
  19. 19.Rony Goldenthal, David Harmon, Raanan Fattal, Michel Bercovier, and Eitan Grinspun. Efficient simulation of inextensible cloth. ACM Transactions on Graphics (TOG), pages 49–es, 2007.
  20. 20.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), pages 2672–2680, 2014.
  21. 21.Artur Grigorev, Karim Iskakov, Anastasia Ianina, Renat Bashirov, Ilya Zakharkin, Alexander Vakhitov, and Victor Lempitsky. Stylepeople: A generative model of fullbody human avatars. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5151–5160, 2021.
  22. 22.Artur Grigorev, Bernhard Thomaszewski, Michael J Black, and Otmar Hilliges. HOOD: Hierarchical graphs for generalized modelling of clothing dynamics. 2023.
  23. 23.Chen Guo, Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition. In Computer Vision and Pattern Recognition (CVPR), 2023.
  24. 24.Pat Hanrahan and Paul Haeberli. Direct wysiwyg painting and texturing on 3d shapes: (an error occurred during the printing of this article that reversed the print order of pages 118 and 119. while we have corrected the sort order of the 2 pages in the dl, the pdf did not allow us to repaginate the 2 pages.). ACM Transactions on Graphics (TOG), 24(4):215–223, 1990.
  25. 25.Fangzhou Hong, Zhaoxi Chen, Yushi Lan, Liang Pan, and Ziwei Liu. Eva3d: Compositional 3d human generation from 2d image collections. arXiv preprint arXiv:2210.04888, 2022.
  26. 26.Tero Karras, Miika Aittala, Samuli Laine, Erik Härk̈onen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. Advances in Neural Information Processing Systems (NeurIPS), 34:852–863, 2021.
  27. 27.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020.
  28. 28.Christoph Lassner, Gerard Pons-Moll, and Peter V Gehler. A generative model of people in clothing. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 853–862, 2017.
  29. 29.Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. Ai choreographer: Music conditioned 3d dance generation with aist++. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021.
  30. 30.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose control. ACM Transactions on Graphics (TOG), 40(6):1–16, 2021.
  31. 31.Wen Liu, Zhixin Piao, Jie Min, Wenhan Luo, Lin Ma, and Shenghua Gao. Liquid warping gan: A unified framework for human motion imitation, appearance transfer and novel view synthesis. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 5904–5913, 2019.
  32. 32.Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1096–1104, 2016.
  33. 33.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (TOG), 34(6):248:1–248:16, October 2015.
  34. 34.Qianli Ma, Shunsuke Saito, Jinlong Yang, Siyu Tang, and Michael J. Black. SCALE: Modeling clothed humans with a surface codec of articulated local elements. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 16082–16093, June 2021.
  35. 35.Qianli Ma, Jinlong Yang, Michael J. Black, and Siyu Tang. Neural point-based shape modeling of humans in challenging clothing. In 2022 International Conference on 3D Vision (3DV), September 2022.
  36. 36.Qianli Ma, Jinlong Yang, Anurag Ranjan, Sergi Pujades, Gerard Pons-Moll, Siyu Tang, and Michael J. Black. Learning to Dress 3D People in Generative Clothing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  37. 37.Qianli Ma, Jinlong Yang, Siyu Tang, and Michael J. Black. The power of points for modeling humans in clothing. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), October 2021.
  38. 38.Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for gans do actually converge? In International Conference on Machine Learning (ICML), pages 3481–3490. PMLR, 2018.
  39. 39.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4460–4470, 2019.
  40. 40.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, July 2022.
  41. 41.Rahul Narain, Armin Samii, and James F O’brien. Adaptive anisotropic remeshing for cloth simulation. ACM Transactions on Graphics (TOG), 31(6):1–10, 2012.
  42. 42.Atsuhiro Noguchi, Xiao Sun, Stephen Lin, and Tatsuya Harada. Unsupervised learning of efficient geometry-aware neural articulated representations. In Proceedings of the European Conference on Computer Vision (ECCV), 2022.
  43. 43.Ahmed AA Osman, Timo Bolkart, and Michael J Black. Star: Sparse trained articulated human body regressor. In Proceedings of the European Conference on Computer Vision (ECCV), pages 598–613. Springer, 2020.
  44. 44.Pablo Palafox, Aljaz Božić, Justus Thies, Matthias Nießner, and Angela Dai. Npms: Neural parametric models for 3d deformable shapes. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2021.
  45. 45.Pablo Palafox, Nikolaos Sarafianos, Tony Tung, and Angela Dai. Spams: Structured implicit parametric models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  46. 46.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 165–174, 2019.
  47. 47.Priyanka Patel, Chun-Hao P Huang, Joachim Tesch, David T Hoffmann, Shashank Tripathi, and Michael J Black. Agora: Avatars in geography optimized for regression analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 13468–13478, 2021.
  48. 48.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 10975–10985, 2019.
  49. 49.Hans Køhling Pedersen. Decorating implicit surfaces. In ACM Transactions on Graphics (TOG), pages 291–300, 1995.
  50. 50.Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 9054–9063, 2021.
  51. 51.Sergey Prokudin, Michael J Black, and Javier Romero. Smplpix: Neural avatars from 3d human models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1810–1819, 2021.
  52. 52.Daniel Rebain, Mark Matthews, Kwang Moo Yi, Dmitry Lagun, and Andrea Tagliasacchi. Lolnerf: Learn from one look. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1558–1567, 2022.
  53. 53.Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito, Chenglei Wu, Tomas Simon, Shih-En Wei, Kaiwen Guo, Zhe Cao, Fabian Prada, Jason Saragih, et al. Drivable volumetric avatars using texel-aligned features. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–9, 2022.
  54. 54.Daniel Roich, Ron Mokady, Amit H. Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real images. ACM Transactions on Graphics (TOG), 42(1), aug 2022.
  55. 55.Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bodies together. ACM Transactions on Graphics (TOG), 36(6), Nov. 2017.
  56. 56.Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), October 2019.
  57. 57.Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 84–93, 2020.
  58. 58.Shunsuke Saito, Jinlong Yang, Qianli Ma, and Michael J. Black. SCANimate: Weakly supervised learning of skinned clothed avatar networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2021.
  59. 59.Kaiyue Shen, Chen Guo, Manuel Kaufmann, Juan Zarate, Julien Valentin, Jie Song, and Otmar Hilliges. X-avatar: Expressive human avatars. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  60. 60.Yichun Shi, Divyansh Aggarwal, and Anil K Jain. Lifting 2d stylegan for 3d-aware face generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 6258–6266, 2021.
  61. 61.Zifan Shi, Sida Peng, Yinghao Xu, Yiyi Liao, and Yujun Shen. Deep generative models on 3d representations: A survey. arXiv preprint arXiv:2210.15663, 2022.
  62. 62.Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Neural geometric level of detail: Real-time rendering with implicit 3D shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  63. 63.Ayush Tewari, Ohad Fried, Justus Thies, Vincent Sitzmann, Stephen Lombardi, Kalyan Sunkavalli, Ricardo Martin-Brualla, Tomas Simon, Jason Saragih, Matthias Nießner, et al. State of the art on neural rendering. In Annual Conference of the European Association for Computer Graphics (EUROGRAPHICS), volume 39, pages 701–727. Wiley Online Library, 2020.
  64. 64.Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, W Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, et al. Advances in neural rendering. In Annual Conference of the European Association for Computer Graphics (EUROGRAPHICS), volume 41, pages 703–735. Wiley Online Library, 2022.
  65. 65.Justus Thies, Michael Zollhofer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 38(4):1–12, 2019.
  66. 66.Garvita Tiwari, Bharat Lal Bhatnagar, Tony Tung, and Gerard Pons-Moll. Sizer: A dataset and model for parsing 3d clothing and learning size sensitive 3d clothing. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, August 2020.
  67. 67.Gul Varol, Javier Romero, Xavier Martin, Naureen Mahmood, Michael J Black, Ivan Laptev, and Cordelia Schmid. Learning from synthetic humans. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 109–117, 2017.
  68. 68.Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 10039–10049, 2021.
  69. 69.Weihao Xia, Yulun Zhang, Yujiu Yang, Jing-Hao Xue, Bolei Zhou, and Ming-Hsuan Yang. Gan inversion: A survey. TPAMI, 2022.
  70. 70.Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. Neural fields in visual computing and beyond. In Annual Conference of the European Association for Computer Graphics (EUROGRAPHICS), volume 41, pages 641–676. Wiley Online Library, 2022.
  71. 71.Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, and Xiaodan Liang. Towards scalable unpaired virtual try-on via patch-routed spatially-adaptive gan. Advances in Neural Information Processing Systems (NeurIPS), 34:2598–2610, 2021.
  72. 72.Bangbang Yang, Chong Bao, Junyi Zeng, Hujun Bao, Yinda Zhang, Zhaopeng Cui, and Guofeng Zhang. Neumesh: Learning disentangled neural mesh-based implicit field for geometry and texture editing. In Proceedings of the European Conference on Computer Vision (ECCV), pages 597–614. Springer, 2022.
  73. 73.Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qionghai Dai, and Yebin Liu. Function4d: Real-time human volumetric capture from very sparse consumer rgbd sensors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2021.
  74. 74.Zhixuan Yu, Jae Shin Yoon, In Kyu Lee, Prashanth Venkatesh, Jaesik Park, Jihun Yu, and Hyun Soo Park. Humbi: A large multiview dataset of human body expressions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2990–3000, 2020.
  75. 75.Jianfeng Zhang, Zihang Jiang, Dingdong Yang, Hongyi Xu, Yichun Shi, Guoxian Song, Zhongcong Xu, Xinchao Wang, and Jiashi Feng. Avatargen: A 3d generative model for animatable human avatars. In Arxiv, 2022.
  76. 76.Yufeng Zheng, Victoria Fernandez Abrevaya, Marcel C. Bühler, Xu Chen, Michael J. Black, and Otmar Hilliges. I M Avatar: Implicit morphable head avatars from videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.

Citation

MLA
Ho, H.-I., et al. “Learning Locally Editable Virtual Humans”. arXiv, 2023, http://arxiv.org/abs/2305.00121v1.
APA
Ho, H.-I., Xue, L., Song, J., & Hilliges, O. (2023). Learning Locally Editable Virtual Humans. arXiv. http://arxiv.org/abs/2305.00121v1
Chicago
Ho, H.-I., L. Xue, J. Song, and O. Hilliges. 2023. “Learning Locally Editable Virtual Humans”. arXiv. http://arxiv.org/abs/2305.00121v1.
Harvard
Ho, H.-I. et al. (2023) “Learning Locally Editable Virtual Humans”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.00121v1.
Vancouver
1. Ho H-I, Xue L, Song J, Hilliges O (2023) Learning Locally Editable Virtual Humans. arXiv

BibTeX

@article{ho2023learning,
  title = {Learning Locally Editable Virtual Humans},
  author = {Ho, Hsuan-I and Xue, Lixin and Song, Jie and Hilliges, Otmar},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.00121v1},
  eprint = {2305.00121}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE