ICON: Implicit Clothed humans Obtained from Normals
Yuliang XiuJinlong YangDimitrios TzionasMichael J. Black
Reconstructs detailed 3D clothed humans from single unconstrained images by coupling body-guided surface normal estimation with local implicit function regression, enabling animatable avatar creation from in-the-wild video.
Generating realistic and animatable three-dimensional digital humans from everyday two-dimensional images is a key technological pillar for virtual reality, telepresence, digital entertainment, and interactive media. Traditional workflows require expensive specialized multi-camera scanning rigs and extensive manual artist intervention, which severely limits scalability. While recent computer vision techniques attempt to reconstruct 3D humans from single images using deep implicit functions, they rely heavily on global image encoders. Consequently, existing methods break down when encountering challenging, unconstrained human poses or cropped photographs, frequently outputting severe visual defects such as missing limbs, detached body parts, and unrealistic surface noise.
The article demonstrates a deep-learning framework named ICON (Implicit Clothed humans Obtained from Normals) that robustly reconstructs highly detailed 3D clothed humans from single unconstrained color images and creates animatable avatars from video sequences. The authors evaluate this framework against existing state-of-the-art systems to test geometric accuracy, generalization to unseen poses, and overall data efficiency.
The evaluation compares ICON against leading baseline models across standard benchmark datasets (AGORA and CAPE) and a curated collection of 200 real-world, in-the-wild images featuring complex movements such as dance, sports, and martial arts. The approach combines a statistical body model prior with an implicit surface regressor guided entirely by local geometric features rather than global encoders. The system incorporates an iterative feedback loop during inference that alternates between refining the underlying body fit and improving front and back surface normal predictions. The team evaluated geometric reconstruction errors using 3D surface distance metrics and conducted a perceptual user study to quantify visual realism.
The findings show that ICON significantly outperforms existing methods across both standard and out-of-distribution pose benchmarks. When tested on complex non-fashion poses, ICON reduces surface reconstruction errors substantially compared to competing implicit-function baselines. Furthermore, the architecture demonstrates remarkable training-data efficiency, matching or exceeding state-of-the-art accuracy when trained on only 12 percent of the standard dataset size. In human perceptual evaluations on challenging real-world photos, human raters favored ICON reconstructions over top competing methods by a wide margin, choosing alternative baselines in only 22 to 31 percent of pairwise comparisons. By feeding per-frame reconstructions into an avatar creation pipeline, the method successfully produces animatable digital avatars featuring realistic, pose-dependent clothing deformations directly from monocular video.
These results demonstrate that switching from global feature representations to pose-agnostic local features effectively solves severe limb-detachment artifacts and eliminates reliance on massive 3D scanning datasets. For organizations developing virtual human pipelines, this provides a scalable, lower-cost path to automated avatar creation without requiring dedicated capture hardware. However, deploying such accessible digital human generation introduces societal and corporate risks surrounding full-body deepfakes and unauthorized digital likeness replication, underscoring the necessity of clear governance and technical licensing controls.
Organizations looking to adopt automated avatar pipelines should consider deploying local-feature implicit architectures to process unconstrained video and photography. Practitioners should integrate visibility-aware weighting when converting monocular video frames into animatable avatars to account for occluded viewpoints. Future development should focus on expanding avatar generation capabilities to build large-scale diverse datasets, while establishing organizational compliance policies for synthetic media.
Confidence in these findings is high for standard poses and close-fitting clothing, supported by rigorous quantitative ablations and perceptual validation. Readers should exercise caution in specific operational edge cases: the framework struggles with loose clothing that departs widely from the body (such as long dresses), severe failures in the initial statistical body fit, and extreme camera perspective distortions that deviate from orthographic assumptions.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). PIFu introduced pixel-aligned implicit functions for digitizing clothed humans from monocular images, establishing the core implicit formulation that ICON aims to make robust under unconstrained poses.
- Paper: SMPL, Matthew Loper et al. (2015). SMPL provides the foundational statistical parametric 3D body model whose geometry and normals guide ICON's local feature conditioning and iterative refinement.
- Paper: Expressive Body Capture: 3D Hands, Face, and Body From a Single Image, Georgios Pavlakos et al. (2019). SMPL-X extends the SMPL model to expressively capture body, face, and hands, serving as the explicit parametric prior directly utilized by ICON during surface reconstruction.
- Paper: End-to-End Recovery of Human Shape and Pose, Angjoo Kanazawa et al. (2017). Human Mesh Recovery pioneered end-to-end regression of parametric SMPL meshes from single RGB images via iterative feedback, informing the body initialization and alignment mechanisms in ICON.
- Paper: DensePose: Dense Human Pose Estimation in the Wild, Rıza Alp Güler et al. (2018). DensePose established dense pixel-to-surface mapping between 2D images and 3D body models, inspiring body-conditioned surface feature extraction.
- Paper: Learning Implicit Fields for Generative Shape Modeling, Zhiqin Chen et al. (2018). IM-NET demonstrates continuous 3D shape generation via implicit occupancy fields, providing the mathematical foundation for implicit surface reconstruction methods.
- Paper: Keep It SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image, Federica Bogo et al. (2016). SMPLify introduced optimization-based fitting of SMPL bodies to 2D image observations, laying the groundwork for the iterative mesh refinement loops employed by ICON.
- Paper: Learning Locally Editable Virtual Humans, Hsuan-I Ho et al. (2023). This work builds on local implicit surface modeling anchored to deformable parametric meshes to generate poseable and locally editable 3D human avatars.
- Paper: Geometry-Consistent Neural Shape Representation with Implicit Displacement Fields, Yifan Wang et al. (2022). Implicit Displacement Fields advance surface-normal conditioned implicit geometry representations by decomposing complex shapes into smooth base meshes and high-frequency normal displacements.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). GaussianAvatars extends parametric mesh-guided neural reconstruction by binding discrete 3D Gaussian primitives to local mesh coordinate frames for animatable avatar generation.
