ICON: Implicit Clothed humans Obtained from Normals

Yuliang XiuJinlong YangDimitrios TzionasMichael J. Black

article2022CVPR446 citations

Reconstructs detailed 3D clothed humans from single unconstrained images by coupling body-guided surface normal estimation with local implicit function regression, enabling animatable avatar creation from in-the-wild video.

Listen

Generating realistic and animatable three-dimensional digital humans from everyday two-dimensional images is a key technological pillar for virtual reality, telepresence, digital entertainment, and interactive media. Traditional workflows require expensive specialized multi-camera scanning rigs and extensive manual artist intervention, which severely limits scalability. While recent computer vision techniques attempt to reconstruct 3D humans from single images using deep implicit functions, they rely heavily on global image encoders. Consequently, existing methods break down when encountering challenging, unconstrained human poses or cropped photographs, frequently outputting severe visual defects such as missing limbs, detached body parts, and unrealistic surface noise.

The article demonstrates a deep-learning framework named ICON (Implicit Clothed humans Obtained from Normals) that robustly reconstructs highly detailed 3D clothed humans from single unconstrained color images and creates animatable avatars from video sequences. The authors evaluate this framework against existing state-of-the-art systems to test geometric accuracy, generalization to unseen poses, and overall data efficiency.

The evaluation compares ICON against leading baseline models across standard benchmark datasets (AGORA and CAPE) and a curated collection of 200 real-world, in-the-wild images featuring complex movements such as dance, sports, and martial arts. The approach combines a statistical body model prior with an implicit surface regressor guided entirely by local geometric features rather than global encoders. The system incorporates an iterative feedback loop during inference that alternates between refining the underlying body fit and improving front and back surface normal predictions. The team evaluated geometric reconstruction errors using 3D surface distance metrics and conducted a perceptual user study to quantify visual realism.

The findings show that ICON significantly outperforms existing methods across both standard and out-of-distribution pose benchmarks. When tested on complex non-fashion poses, ICON reduces surface reconstruction errors substantially compared to competing implicit-function baselines. Furthermore, the architecture demonstrates remarkable training-data efficiency, matching or exceeding state-of-the-art accuracy when trained on only 12 percent of the standard dataset size. In human perceptual evaluations on challenging real-world photos, human raters favored ICON reconstructions over top competing methods by a wide margin, choosing alternative baselines in only 22 to 31 percent of pairwise comparisons. By feeding per-frame reconstructions into an avatar creation pipeline, the method successfully produces animatable digital avatars featuring realistic, pose-dependent clothing deformations directly from monocular video.

These results demonstrate that switching from global feature representations to pose-agnostic local features effectively solves severe limb-detachment artifacts and eliminates reliance on massive 3D scanning datasets. For organizations developing virtual human pipelines, this provides a scalable, lower-cost path to automated avatar creation without requiring dedicated capture hardware. However, deploying such accessible digital human generation introduces societal and corporate risks surrounding full-body deepfakes and unauthorized digital likeness replication, underscoring the necessity of clear governance and technical licensing controls.

Organizations looking to adopt automated avatar pipelines should consider deploying local-feature implicit architectures to process unconstrained video and photography. Practitioners should integrate visibility-aware weighting when converting monocular video frames into animatable avatars to account for occluded viewpoints. Future development should focus on expanding avatar generation capabilities to build large-scale diverse datasets, while establishing organizational compliance policies for synthetic media.

Confidence in these findings is high for standard poses and close-fitting clothing, supported by rigorous quantitative ablations and perceptual validation. Readers should exercise caution in specific operational edge cases: the framework struggles with loose clothing that departs widely from the body (such as long dresses), severe failures in the initial statistical body fit, and extreme camera perspective distortions that deviate from orthographic assumptions.

Cover for ICON: Implicit Clothed humans Obtained from Normals

Abstract

Current methods for learning realistic and animatable 3D clothed avatars need either posed 3D scans or 2D images with carefully controlled user poses. In contrast, our goal is to learn an avatar from only 2D images of people in unconstrained poses. Given a set of images, our method estimates a detailed 3D surface from each image and then combines these into an animatable avatar. Implicit functions are well suited to the first task, as they can capture details like hair and clothes. Current methods, however, are not robust to varied human poses and often produce 3D surfaces with broken or disembodied limbs, missing details, or non-human shapes. The problem is that these methods use global feature encoders that are sensitive to global pose. To address this, we propose ICON (“Implicit Clothed humans Obtained from Normals”), which, instead, uses local features. ICON has two main modules, both of which exploit the SMPL(-X) body model. First, ICON infers detailed clothed-human normals (front/back) conditioned on the SMPL(-X) normals. Second, a visibility-aware implicit surface regressor produces an iso-surface of a human occupancy field. Importantly, at inference time, a feedback loop alternates between refining the SMPL(-X) mesh using the inferred clothed normals and then refining the normals. Given multiple reconstructed frames of a subject in varied poses, we use a modified version of SCANimate to produce an animatable avatar from them. Evaluation on the AGORA and CAPE datasets shows that ICON outperforms the state of the art in reconstruction, even with heavily limited training data. Additionally, it is much more robust to out-of-distribution samples, e.g., in-the-wild poses/images and out-of-frame cropping. ICON takes a step towards robust 3D clothed human reconstruction from in-the-wild images. This enables avatar creation directly from video with personalized pose-dependent cloth deformation. Models and code are available for research at https://icon.is.tue.mpg.de.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Body-guided normal prediction
  • 3.2 Local-feature based implicit 3D reconstruction
  • 4 Experiments
  • 4.1 Baseline models
  • 4.2 Datasets
  • 4.3 Evaluation
  • 5 Applications
  • 5.1 Reconstruction from in-the-wild images
  • 5.2 Animatable avatar creation from video
  • 6 Conclusion
  • A Method & Experiment Details
  • A.1 Dataset ()
  • A.2 Refining SMPL ()
  • A.3 Perceptual study ()
  • A.4 Implementation details ()
  • B More Quantitative Results ()
  • C More Qualitative Results ()
  • References

Knowls

  1. Knowl 1 — Local Feature Representation for Implicit 3D Clothed Human Reconstruction

    model/method

    ICON reconstructs 3D clothed human geometry from a single RGB image by querying an implicit surface function parameterized by local, pose-agnostic features instead of global image or volumetric encodings.

    Let P∈R3P \in \mathbb{R}^3 denote a 3D query point, M∈RN×3M \in \mathbb{R}^{N \times 3} denote an estimated parametric SMPL body mesh with N=6,890N = 6{,}890 vertices, and Pb∈MP^b \in M denote the closest point on the SMPL body surface to PP. Let π(P)\pi(P) denote the 2D perspective or weak-perspective projection of PP onto the image plane. The local feature vector FPF_P assigned to PP is defined as:

    FP=[Fs(P),Fnb(P),Fnc(P)]F_P = [F_s(P), F_n^b(P), F_n^c(P)]

    where:

    • Fs(P)∈RF_s(P) \in \mathbb{R} is the signed distance from PP to the closest body point Pb∈MP^b \in M.
    • Fnb(P)∈R3F_n^b(P) \in \mathbb{R}^3 is the barycentric surface normal of the SMPL mesh at PbP^b.
    • Fnc(P)∈R3F_n^c(P) \in \mathbb{R}^3 is a clothed-body surface normal vector sampled from the predicted front or back clothed-body normal maps N^c={N^frontc,N^backc}\hat{\mathcal{N}}^c = \{\hat{\mathcal{N}}^c_\text{front}, \hat{\mathcal{N}}^c_\text{back}\}, determined by the visibility of PbP^b from the camera viewpoint:

    Fnc(P)={N^frontc(π(P)),if Pb is visible to the cameraN^backc(π(P)),otherwiseF_n^c(P) = \begin{cases} \hat{\mathcal{N}}^c_\text{front}(\pi(P)), & \text{if } P^b \text{ is visible to the camera} \\ \hat{\mathcal{N}}^c_\text{back}(\pi(P)), & \text{otherwise} \end{cases}

    The combined feature vector FPF_P is fed into an implicit function IF\mathcal{IF}, parameterized as a Multi-Layer Perceptron (MLP), to predict the continuous human occupancy field o^(P)∈[0,1]\hat{o}(P) \in [0, 1] at query point PP. The network is supervised using a mean squared error loss against ground-truth occupancy labels o(P)o(P). Meshes are extracted from the continuous occupancy field using fast surface localization and marching cubes. Because FPF_P depends only on local relative distances and surface orientations rather than global body coordinates or global image feature maps, it is invariant to global pose changes and avoids learning spurious correlations between body pose and clothing surface details.

  2. Knowl 2 — SMPL-Guided Clothed-Body Normal Prediction

    model/method

    To infer full 360∘360^\circ detailed surface normals from a single RGB image II containing a segmented human, ICON conditions 2D normal estimation on rendered front and back surface normals of an underlying SMPL body model M(β,θ)M(\beta, \theta).

    Given the estimated SMPL shape parameters β∈R10\beta \in \mathbb{R}^{10} and pose parameters θ∈R3×K\theta \in \mathbb{R}^{3 \times K} (with K=24K=24 joints), a differentiable renderer renders normal maps of MM from opposite viewpoints under a weak-perspective camera model (scale s∈Rs \in \mathbb{R}, translation t∈R3t \in \mathbb{R}^3), yielding the front (observable) and back (occluded) SMPL-body normal maps:

    Nb={Nfrontb,Nbackb}\mathcal{N}^b = \{\mathcal{N}^b_\text{front}, \mathcal{N}^b_\text{back}\}

    Two normal prediction networks with separate parameters, GN={GfrontN,GbackN}G^N = \{G^N_\text{front}, G^N_\text{back}\}, infer the detailed front and back clothed-body normal maps N^c={N^frontc,N^backc}\hat{\mathcal{N}}^c = \{\hat{\mathcal{N}}^c_\text{front}, \hat{\mathcal{N}}^c_\text{back}\} conditioned on Nb\mathcal{N}^b and the input image II:

    GN(Nb,I)→N^cG^N(\mathcal{N}^b, I) \to \hat{\mathcal{N}}^c

    The normal networks are trained by minimizing:

    LN=Lpixel+λVGGLVGG\mathcal{L}_N = \mathcal{L}_\text{pixel} + \lambda_\text{VGG}\mathcal{L}_\text{VGG}

    where Lpixel=∑v∈{front,back}∣Nvc−N^vc∣\mathcal{L}_\text{pixel} = \sum_{v \in \{\text{front}, \text{back}\}} |\mathcal{N}^c_v - \hat{\mathcal{N}}^c_v| is the L1L_1 loss between ground-truth clothed normals Nvc\mathcal{N}^c_v and predicted normals N^vc\hat{\mathcal{N}}^c_v, and LVGG\mathcal{L}_\text{VGG} is a perceptual loss weighted by hyperparameter λVGG\lambda_\text{VGG} that recovers high-frequency clothing wrinkles and folds.

  3. Knowl 3 — Inference-Time Feedback Optimization for SMPL Body Mesh and Surface Normals

    algorithm

    At inference time, monocular human pose and shape estimators often produce imperfectly aligned SMPL body fits. ICON addresses this by running an iterative feedback loop that optimizes the SMPL parameters against the predicted clothed normal maps, and subsequently updates the predicted normal maps using the refined SMPL body.

    Input: Segmented color image II, foreground mask ScS^c, initial SMPL parameters β,θ,t\beta, \theta, t, normal prediction networks GNG^N, differentiable renderer DRDR, weighting hyperparameter λNdiff\lambda_{N_\text{diff}}
    Output: Refined SMPL mesh MM, refined clothed normal maps N^c\hat{\mathcal{N}}^c, and 3D reconstructed clothed human mesh
    Render initial body normals Nb=DR(M(β,θ,t))\mathcal{N}^b = DR(M(\beta, \theta, t))
    Predict initial clothed normals N^c=GN(Nb,I)\hat{\mathcal{N}}^c = G^N(\mathcal{N}^b, I)
    repeat until convergence or maximum iterations:
        Render current SMPL body normals Nb=DR(M(β,θ,t))\mathcal{N}^b = DR(M(\beta, \theta, t))
        Render current SMPL body silhouette mask SbS^b
        Compute normal difference loss LNdiff=∣Nb−N^c∣\mathcal{L}_{N_\text{diff}} = |\mathcal{N}^b - \hat{\mathcal{N}}^c|
        Compute silhouette difference loss LSdiff=∣Sb−Sc∣\mathcal{L}_{S_\text{diff}} = |S^b - S^c|
        Compute total SMPL loss LSMPL=λNdiffLNdiff+LSdiff\mathcal{L}_\text{SMPL} = \lambda_{N_\text{diff}} \mathcal{L}_{N_\text{diff}} + \mathcal{L}_{S_\text{diff}}
        Update SMPL parameters β,θ,t\beta, \theta, t via gradient descent on LSMPL\mathcal{L}_\text{SMPL}
        Re-render updated SMPL body normals Nb=DR(M(β,θ,t))\mathcal{N}^b = DR(M(\beta, \theta, t))
        Re-infer refined clothed-body normals N^c=GN(Nb,I)\hat{\mathcal{N}}^c = G^N(\mathcal{N}^b, I)
    Extract local point features FP=[Fs(P),Fnb(P),Fnc(P)]F_P = [F_s(P), F_n^b(P), F_n^c(P)] for 3D query points using refined MM and N^c\hat{\mathcal{N}}^c
    Infer 3D occupancy field using implicit function IF(FP)\mathcal{IF}(F_P)
    Extract 3D surface mesh using marching cubes
    return 3D surface mesh and refined parameters
  4. Knowl 4 — Quantitative Clothed Human Reconstruction Performance on AGORA and CAPE

    data/table

    ICON is evaluated on 3D clothed human reconstruction across the in-distribution AGORA-50 dataset and the unseen CAPE test dataset, split into fashion poses (CAPE-FP) and non-fashion / out-of-distribution poses (CAPE-NFP). Evaluation metrics are Chamfer distance (cm), point-to-surface distance (P2S, cm), and rendered surface normal difference (Normals, L2L_2 error).

    Method AGORA-50 CAPE-FP CAPE-NFP CAPE (All)
    Chamfer ↓\downarrow P2S ↓\downarrow Normals ↓\downarrow Chamfer ↓\downarrow P2S ↓\downarrow Normals ↓\downarrow Chamfer ↓\downarrow P2S ↓\downarrow Normals ↓\downarrow Chamfer ↓\downarrow P2S ↓\downarrow Normals ↓\downarrow
    PIFu 3.453 3.660 0.094 2.823 2.796 0.100 4.029 4.195 0.124 3.627 3.729 0.116
    PIFuHD 3.119 3.333 0.085 2.302 2.335 0.090 3.704 3.517 0.123 3.237 3.123 0.112
    PaMIR 2.035 1.873 0.079 1.936 1.263 0.078 2.216 1.611 0.093 2.122 1.495 0.088
    SMPL-X GT 1.518 1.985 0.072 1.335 1.259 0.085 1.070 1.058 0.068 1.158 1.125 0.074
    PIFu* 2.688 2.573 0.097 2.100 2.093 0.091 2.973 2.940 0.111 2.682 2.658 0.104
    PaMIR* 1.401 1.500 0.063 1.225 1.206 0.055 1.413 1.321 0.063 1.350 1.283 0.060
    ICON 1.204 1.584 0.060 1.233 1.170 0.072 1.096 1.013 0.063 1.142 1.065 0.066

    While PaMIR* performs well on in-distribution poses (AGORA-50 and CAPE-FP), its error increases substantially on out-of-distribution complex poses (CAPE-NFP) due to its global voxel feature encoder. ICON achieves the lowest Chamfer error (1.142 cm) and P2S error (1.065 cm) on CAPE overall, generalizing across out-of-distribution poses due to its local, pose-agnostic features.

  5. Knowl 5 — Training Data Sample Efficiency of Local Point Features

    empirical result

    Because ICON extracts features strictly defined in local body-centric coordinates rather than global image or 3D voxel representations, it avoids learning spurious correlations between global poses and clothing geometry. Consequently, ICON achieves state-of-the-art reconstruction accuracy with a fraction of the training data required by baseline architectures.

    When trained on varying fractions of the 450 standard Renderpeople training scans:

    • At a 1/8x data scale (56 scans), ICON achieves a Chamfer distance of 1.336 cm1.336\text{ cm}, which already outperforms the full-dataset (1x scale, 450 scans) performance of PIFu (2.682 cm2.682\text{ cm}) and matches that of PaMIR (1.350 cm1.350\text{ cm}).
    • At 1/4x (112 scans), 1/2x (225 scans), and 1x (450 scans), ICON achieves Chamfer distances of 1.266 cm1.266\text{ cm}, 1.219 cm1.219\text{ cm}, and 1.142 cm1.142\text{ cm}, respectively.
    • At 8x data scale (3,709 scans from AGORA and THuman), ICON reaches a Chamfer distance of 1.036 cm1.036\text{ cm}, consistently maintaining a performance gap over PIFu (1.760 cm1.760\text{ cm}) and PaMIR (1.095 cm1.095\text{ cm}).
  6. Knowl 6 — Animatable Clothed Avatar Creation from Monocular Video via Visibility-Aware SCANimate

    model/method

    ICON reconstructs detailed 3D clothed human meshes for individual frames of an unconstrained monocular video of a person in motion. To turn these per-frame meshes into a coherent, pose-driven animatable clothed avatar, the meshes are processed using a modified formulation of SCANimate.

    Unlike 3D scanner meshes captured by multi-view camera rigs, monocular reconstructions from ICON have high geometric fidelity and detail on body regions visible to the camera, but lower fidelity on occluded regions. To account for this asymmetry, the optimization loss in SCANimate is reformulated to downweight implicit surface alignment errors in occluded body regions based on camera viewpoint visibility. Training SCANimate on these visibility-weighted per-frame ICON meshes yields an avatar that reproduces pose-dependent clothing deformations.

  7. Knowl 7 — Robustness of ICON Reconstruction to SMPL-X Pose and Shape Fitting Noise

    empirical result

    When evaluated with artificially perturbed SMPL-X body model fits rather than ground-truth body fits on the CAPE dataset:

    • Unrefined ICON conditioned on noisy SMPL-X fits yields a Chamfer distance of 1.417 cm1.417\text{ cm}, a P2S error of 1.436 cm1.436\text{ cm}, and a normal difference of 0.0820.082.
    • Enabling ICON's inference-time body refinement (ICON + BR) reduces the Chamfer distance to 1.339 cm1.339\text{ cm}, P2S error to 1.378 cm1.378\text{ cm}, and normal difference to 0.0720.072.

    With body refinement active, ICON conditioned on noisy SMPL-X fits achieves performance comparable to PaMIR* conditioned on ground-truth SMPL-X fits (Chamfer distance of 1.350 cm1.350\text{ cm}, P2S error of 1.283 cm1.283\text{ cm}, and normal difference of 0.0600.060).

  8. Knowl 8 — Perceptual Evaluation on In-the-Wild Clothed Human Reconstructions

    empirical result

    To assess reconstruction realism on unconstrained real-world imagery, a two-alternative forced-choice perceptual study was conducted using 200 in-the-wild images from Pinterest featuring challenging, dynamic human poses (e.g., parkour, dance, martial arts, sports). Participants were shown the original image alongside 3D renderings from ICON and a baseline method, and asked to select the 3D shape that best represents the human in the image.

    Baseline models were trained on 3,709 scans from AGORA and THuman (with pre-trained weights used for PIFuHD). The preference rates for baseline models against ICON were:

    • PIFu*: 30.9%30.9\% user preference (p=1.35×10−33p = 1.35 \times 10^{-33} against equal performance)
    • PIFuHD: 22.3%22.3\% user preference (p=1.08×10−48p = 1.08 \times 10^{-48})
    • PaMIR*: 26.6%26.6\% user preference (p=3.60×10−54p = 3.60 \times 10^{-54})

    In all cases, ICON was judged significantly more realistic and faithful to the input images than competing approaches.

  9. Knowl 9 — Limitations of ICON

    limitation

    ICON exhibits three main limitations:

    1. Loose and non-body-conforming clothing: Because query point features are regularized by proximity and surface normals of the underlying parametric SMPL body model, garments that stand significantly far off the body (such as wide skirts, dresses, and loose cloaks) cannot be accurately represented.
    2. Catastrophic parametric body fitting failure: While the inference feedback loop provides robustness against moderate body pose and shape misalignments, severe initialization failures by the upstream human pose and shape regressor cause reconstruction failure.
    3. Perspective distortion: Because ICON is trained with an orthographic / weak-perspective camera model, single-view inputs captured with strong perspective distortion can lead to distorted or anatomically implausible reconstructions, such as asymmetric or elongated limbs.

Coverage note — None was omitted; all key architectural components, equations, algorithms, empirical evaluations (AGORA, CAPE, data efficiency, perceptual study), avatar creation applications, and limitations are covered.

References

  1. 1.RenderPeople. renderpeople.com, 2018. 2, 5
  2. 2.Thiemo Alldieck, Marcus A. Magnor, Bharat Lal Bhatnagar, Christian Theobalt, and Gerard Pons-Moll. Learning to reconstruct people in clothing from a single RGB camera. In Computer Vision and Pattern Recognition (CVPR), pages 1175–1186, 2019. 3
  3. 3.Thiemo Alldieck, Marcus A. Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Detailed human avatars from monocular video. In International Conference on 3D Vision (3DV), pages 98–109, 2018. 3
  4. 4.Thiemo Alldieck, Marcus A. Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruction of 3D people models. In Computer Vision and Pattern Recognition (CVPR), pages 8387–8397, 2018. 1, 3
  5. 5.Thiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, and Marcus A. Magnor. Tex2Shape: Detailed full human body geometry from a single image. In International Conference on Computer Vision (ICCV), pages 2293–2303, 2019. 1, 3
  6. 6.Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Combining implicit function learning and parametric models for 3D human reconstruction. In European Conference on Computer Vision (ECCV), volume 12347, pages 311–329, 2020. 3
  7. 7.Bharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. LoopReg: Self-supervised learning of implicit surface correspondences, pose and shape for 3D human mesh registration. In Conference on Neural Information Processing Systems (NeurIPS), 2020. 3
  8. 8.Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, and Gerard Pons-Moll. Multi-Garment Net: Learning to dress 3D people from images. In International Conference on Computer Vision (ICCV), pages 5419–5429, 2019. 3
  9. 9.Aljaz Bozic, Pablo R. Palafox, Michael Zollhöfer, Justus Thies, Angela Dai, and Matthias Nießner. Neural deformation graphs for globally-consistent non-rigid reconstruction. In Computer Vision and Pattern Recognition (CVPR), pages 1450–1459, 2021. 3
  10. 10.Xu Chen, Tianjian Jiang, Jie Song, Jinlong Yang, Michael J. Black, Andreas Geiger, and Otmar Hilliges. gDNA: Towards generative detailed neural avatars. In Computer Vision and Pattern Recognition (CVPR), 2022. 7
  11. 11.Xu Chen, Yufeng Zheng, Michael J. Black, Otmar Hilliges, and Andreas Geiger. SNARF: Differentiable forward skinning for animating non-rigid neural implicit shapes. In International Conference on Computer Vision (ICCV), pages 11594–11604, 2021. 3
  12. 12.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In Computer Vision and Pattern Recognition (CVPR), pages 5939–5948, 2019. 3
  13. 13.Vasileios Choutas, Lea Müller, Chun-Hao P. Huang, Siyu Tang, Dimitrios Tzionas, and Michael J. Black. Accurate 3D body shape regression via linguistic attributes and anthropometric measurements. In Computer Vision and Pattern Recognition (CVPR), 2022. 1, 3
  14. 14.Boyang Deng, JP Lewis, Timothy Jeruzalski, Gerard Pons-Moll, Geoffrey Hinton, Mohammad Norouzi, and Andrea Tagliasacchi. Neural Articulated Shape Approximation. In European Conference on Computer Vision (ECCV), volume 12352, pages 612–628, 2020. 3
  15. 15.Zijian Dong, Chen Guo, Jie Song, Xu Chen, Andreas Geiger, and Otmar Hilliges. PINA: Learning a personalized implicit neural avatar from a single RGB-D video sequence. In Computer Vision and Pattern Recognition (CVPR), 2022. 3
  16. 16.Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, and Michael J. Black. Collaborative regression of expressive bodies using moderation. In International Conference on 3D Vision (3DV), pages 792–804, 2021. 1
  17. 17.Tong He, John P. Collomosse, Hailin Jin, and Stefano Soatto. Geo-PIFu: Geometry and pixel aligned implicit functions for single-view human reconstruction. In Conference on Neural Information Processing Systems (NeurIPS), 2020. 2, 3
  18. 18.Tong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, and Tony Tung. ARCH++: Animation-ready clothed human reconstruction revisited. In International Conference on Computer Vision (ICCV), pages 11046–11056, 2021. 3
  19. 19.Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li, and Tony Tung. ARCH: Animatable reconstruction of clothed humans. In Computer Vision and Pattern Recognition (CVPR), pages 3090–3099, 2020. 2, 3, 5
  20. 20.Aaron S. Jackson, Chris Manafas, and Stefan Roth Georgios Tzimiropoulos. 3D human body reconstruction from a single image via volumetric regression. In European Conference on Computer Vision Workshops (ECCVw), volume 11132, pages 64–77, 2018. 6
  21. 21.Yasamin Jafarian and Hyun Soo Park. Learning high fidelity depths of dressed humans by watching social media dance videos. In Computer Vision and Pattern Recognition (CVPR), pages 12753–12762, June 2021. 3
  22. 22.Boyi Jiang, Juyong Zhang, Yang Hong, Jinhao Luo, Ligang Liu, and Hujun Bao. BCNet: Learning body and cloth shape from a single image. In European Conference on Computer Vision (ECCV), pages 18–35, 2020. 3
  23. 23.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision (ECCV), volume 9906, pages 694–711, 2016. 4
  24. 24.Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Total capture: A 3D deformation model for tracking faces, hands, and bodies. In Computer Vision and Pattern Recognition (CVPR), pages 8320–8329, 2018. 1, 2
  25. 25.Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In Computer Vision and Pattern Recognition (CVPR), pages 7122–7131, 2018. 3
  26. 26.Muhammed Kocabas, Nikos Athanasiou, and Michael J. Black. VIBE: Video inference for human body pose and shape estimation. In Computer Vision and Pattern Recognition (CVPR), pages 5252–5262, 2020. 3
  27. 27.Muhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, and Michael J. Black. PARE: Part attention regressor for 3D human body estimation. In International Conference on Computer Vision (ICCV), pages 11127–11137, 2021. 2
  28. 28.Nikos Kolotouros, Georgios Pavlakos, Michael J. Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In International Conference on Computer Vision (ICCV), pages 2252–2261, 2019. 1, 3
  29. 29.Verica Lazova, Eldar Insafutdinov, and Gerard Pons-Moll. 360-Degree textures of people in clothing from a single image. In International Conference on 3D Vision (3DV), pages 643–653, 2019. 3
  30. 30.Ruilong Li, Yuliang Xiu, Shunsuke Saito, Zeng Huang, Kyle Olszewski, and Hao Li. Monocular real-time volumetric performance capture. In European Conference on Computer Vision (ECCV), volume 12368, pages 49–67, 2020. 3, 5
  31. 31.Zhe Li, Tao Yu, Chuanyu Pan, Zerong Zheng, and Yebin Liu. Robust 3D self-portraits in Seconds. In Computer Vision and Pattern Recognition (CVPR), pages 1341–1350, 2020. 3
  32. 32.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. Transactions on Graphics (TOG), 34(6):248:1–248:16, 2015. 1, 2, 3
  33. 33.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3D surface construction algorithm. International Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 21(4):163–169, 1987. 5
  34. 34.Qianli Ma, Shunsuke Saito, Jinlong Yang, Siyu Tang, and Michael J. Black. SCALE: Modeling clothed humans with a surface codec of articulated local elements. In Computer Vision and Pattern Recognition (CVPR), pages 16082–16093, 2021. 3
  35. 35.Qianli Ma, Jinlong Yang, Anurag Ranjan, Sergi Pujades, Gerard Pons-Moll, Siyu Tang, and Michael J. Black. Learning to dress 3D people in generative clothing. In Computer Vision and Pattern Recognition (CVPR), pages 6468–6477, 2020. 2, 5
  36. 36.Qianli Ma, Jinlong Yang, Siyu Tang, and Michael J. Black. The power of points for modeling humans in clothing. In International Conference on Computer Vision (ICCV), pages 10974–10984, 2021. 3
  37. 37.Lars M. Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3D reconstruction in function space. In Computer Vision and Pattern Recognition (CVPR), pages 4460–4470, 2019. 3
  38. 38.Jeong Joon Park, Peter Florence, Julian Straub, Richard A. Newcombe, and Steven Lovegrove. DeepSDF: Learning continuous signed distance functions for shape representation. In Computer Vision and Pattern Recognition (CVPR), pages 165–174, 2019. 3
  39. 39.Priyanka Patel, Chun-Hao P. Huang, Joachim Tesch, David T. Hoffmann, Shashank Tripathi, and Michael J. Black. AGORA: Avatars in geography optimized for regression analysis. In Computer Vision and Pattern Recognition (CVPR), pages 13468–13478, 2021. 2, 5, 6, 7
  40. 40.Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Computer Vision and Pattern Recognition (CVPR), pages 10975–10985, 2019. 1, 2, 3
  41. 41.PIFuHD code on GitHub. https://github.com/facebookresearch/pifuhd, 2020. 3
  42. 42.Gerard Pons-Moll, Sergi Pujades, Sonny Hu, and Michael J. Black. ClothCap: seamless 4D clothing capture and retargeting. Transactions on Graphics (TOG), 36(4):73:1–73:15, 2017. 3, 5
  43. 43.Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. Accelerating 3D deep learning with PyTorch3D. arXiv:2007.08501, 2020. 3
  44. 44.Rembg: A tool to remove images background. https://github.com/danielgatis/rembg, 2022. 4
  45. 45.Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bodies together. Transactions on Graphics (TOG), 36(6):245:1–245:17, 2017. 1, 2
  46. 46.Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Hao Li, and Angjoo Kanazawa. PIFu: Pixel-aligned implicit function for high-resolution clothed human digitization. In International Conference on Computer Vision (ICCV), pages 2304–2314, 2019. 2, 3, 5, 6, 7
  47. 47.Shunsuke Saito, Tomas Simon, Jason M. Saragih, and Hanbyul Joo. PIFuHD: Multi-level pixel-aligned implicit function for high-resolution 3D human digitization. In Computer Vision and Pattern Recognition (CVPR), pages 81–90, 2020. 2, 3, 5, 6, 7
  48. 48.Shunsuke Saito, Jinlong Yang, Qianli Ma, and Michael J. Black. SCANimate: Weakly supervised learning of skinned clothed avatar networks. In Computer Vision and Pattern Recognition (CVPR), pages 2886–2897, 2021. 2, 3, 7
  49. 49.David Smith, Matthew Loper, Xiaochen Hu, Paris Mavroidis, and Javier Romero. FACSIMILE: Fast and accurate scans from an image in less than a second. In International Conference on Computer Vision (ICCV), pages 5330–5339, 2019. 3
  50. 50.Yu Sun, Wu Liu, Qian Bao, Yili Fu, Tao Mei, and Michael J. Black. Putting people in their place: Monocular regression of 3D people in depth. In Computer Vision and Pattern Recognition (CVPR), 2022. 3
  51. 51.Sicong Tang, Feitong Tan, Kelvin Cheng, Zhaoyang Li, Siyu Zhu, and Ping Tan. A neural network for detailed human depth estimation from a single image. In International Conference on Computer Vision (ICCV), pages 7750–7759, 2019. 3
  52. 52.Garvita Tiwari, Nikolaos Sarafianos, Tony Tung, and Gerard Pons-Moll. Neural-GIF: Neural generalized implicit functions for animating people in clothing. In International Conference on Computer Vision (ICCV), pages 11708–11718, 2021. 3
  53. 53.Twindom. twindom.com, 2018. 5
  54. 54.Shaofei Wang, Marko Mihajlovic, Qianli Ma, Andreas Geiger, and Siyu Tang. MetaAvatar: Learning animatable clothed human models from few depth images. In Conference on Neural Information Processing Systems (NeurIPS), 2021. 3
  55. 55.Donglai Xiang, Fabian Prada, Chenglei Wu, and Jessica K. Hodgins. MonoClothCap: Towards temporally coherent clothing capture from monocular RGB video. In International Conference on 3D Vision (3DV), pages 322–332, 2020. 3
  56. 56.Hongyi Xu, Eduard Gabriel Bazavan, Andrei Zanfir, William T. Freeman, Rahul Sukthankar, and Cristian Sminchisescu. GHUM & GHUML: Generative 3D human shape and articulated pose models. In Computer Vision and Pattern Recognition (CVPR), pages 6183–6192, 2020. 1, 2
  57. 57.Ze Yang, Shenlong Wang, Sivabalan Manivasagam, Zeng Huang, Wei-Chiu Ma, Xinchen Yan, Ersin Yumer, and Raquel Urtasun. S3: Neural shape, skeleton, and skinning fields for 3D human modeling. In Computer Vision and Pattern Recognition (CVPR), pages 13284–13293, 2021. 2, 3
  58. 58.Hongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas, Mohamed Hassan, Siyu Tang, Justus Thies, and Michael J. Black. Human-aware object placement for visual environment reconstruction. In Computer Vision and Pattern Recognition (CVPR), 2022. 3
  59. 59.Chao Zhang, Sergi Pujades, Michael J. Black, and Gerard Pons-Moll. Detailed, accurate, human shape estimation from clothed 3D scan sequences. In Computer Vision and Pattern Recognition (CVPR), pages 5484–5493, 2017. 5
  60. 60.Hongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang, Yebin Liu, Limin Wang, and Zhenan Sun. Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop. In International Conference on Computer Vision (ICCV), pages 11446–11456, 2021. 3
  61. 61.Yang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu, Zerong Zheng, Qionghai Dai, and Yebin Liu. DeepMultiCap: Performance capture of multiple characters using sparse multiview cameras. In International Conference on Computer Vision (ICCV), pages 6239–6249, 2021. 3
  62. 62.Zerong Zheng, Tao Yu, Yebin Liu, and Qionghai Dai. PaMIR: Parametric model-conditioned implicit representation for image-based human reconstruction. Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2021. 2, 3, 5, 6
  63. 63.Zerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai, and Yebin Liu. DeepHuman: 3D human reconstruction from a single image. In International Conference on Computer Vision (ICCV), pages 7738–7748, 2019. 5, 6, 7
  64. 64.Hao Zhu, Xinxin Zuo, Sen Wang, Xun Cao, and Ruigang Yang. Detailed human shape estimation from a single image by hierarchical mesh deformation. In Computer Vision and Pattern Recognition (CVPR), pages 4491–4500, 2019. 3

Citation

MLA
Xiu, Y., et al. “ICON: Implicit Clothed Humans Obtained from Normals”. arXiv, 2021, http://arxiv.org/abs/2112.09127v2.
APA
Xiu, Y., Yang, J., Tzionas, D., & Black, M. J. (2021). ICON: Implicit Clothed humans Obtained from Normals. arXiv. http://arxiv.org/abs/2112.09127v2
Chicago
Xiu, Y., J. Yang, D. Tzionas, and M. J. Black. 2021. “ICON: Implicit Clothed Humans Obtained from Normals”. arXiv. http://arxiv.org/abs/2112.09127v2.
Harvard
Xiu, Y. et al. (2021) “ICON: Implicit Clothed humans Obtained from Normals”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2112.09127v2.
Vancouver
1. Xiu Y, Yang J, Tzionas D, Black MJ (2021) ICON: Implicit Clothed humans Obtained from Normals. arXiv

BibTeX

@article{xiu2021icon,
  title = {ICON: Implicit Clothed humans Obtained from Normals},
  author = {Xiu, Yuliang and Yang, Jinlong and Tzionas, Dimitrios and Black, Michael J.},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2112.09127v2},
  eprint = {2112.09127}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE