PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization

Shunsuke SaitoZeng HuangRyota NatsumeShigeo MorishimaAngjoo KanazawaHao Li

article2019ICCV1,448 citations

Proposes a pixel-aligned implicit function representation that reconstructs complete, high-resolution 3D clothed humans and surface textures directly from single 2D photographs, overcoming the memory limits of traditional voxel-based methods.

Listen

High-resolution 3D digitization of clothed humans typically requires expensive multi-camera studios or specialized scanning equipment. Existing single-image 3D deep learning approaches often rely on memory-heavy volumetric grids or coarse body templates that fail to recover complex clothing shapes, loose hairstyles, and realistic color textures.

The article evaluates a new deep learning framework, the Pixel-aligned Implicit Function, designed to infer both high-resolution 3D surface geometry and complete 360-degree color textures from a single photograph or sparse multi-view images.

The framework combines a fully convolutional image encoder with a continuous implicit function that classifies whether any point along a camera ray is inside or outside the object's surface. By tying local pixel features directly to 3D spatial coordinates rather than using discrete grid cells, the system operates with high memory efficiency and handles arbitrary clothing shapes. The model was trained and evaluated on photogrammetry datasets containing hundreds of high-quality 3D scans and benchmarked against real-world photographs.

The evaluation yielded several key findings. First, the framework achieved state-of-the-art accuracy in single-image reconstruction, cutting point-to-surface errors by more than half compared to voxel-based and template-free baselines on benchmark datasets. Second, the method successfully reconstructs intricate surface details such as clothing wrinkles and high heels while synthesizing plausible geometry and color for unseen regions like the subject's back. Third, the system naturally scales across varying inputs: when three calibrated views were provided, reconstruction errors decreased significantly, outperforming prior multi-view and video-based methods.

These results demonstrate that high-quality, fully textured 3D digital humans can be generated directly from standard RGB imagery without specialized capture hardware. This significantly lowers production costs and deployment complexity for immersive virtual reality, visual effects, and digital avatar creation.

Organizations developing 3D capture workflows should consider adopting continuous implicit representations over legacy voxel or rigid template pipelines. Before production deployment, technical teams should implement robust preprocessing steps to estimate scale factors and evaluate generative enhancement techniques for ultra-high-resolution texture detail.

Confidence in these findings is high for fully visible, upright subjects, as supported by consistent quantitative benchmarks. However, stakeholders should exercise caution regarding real-world scenes with partial body occlusions, foreground clutter, or unknown camera scaling, which remain active areas for further refinement.

Cover for PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization

Abstract

We introduce Pixel-aligned Implicit Function (PIFu), a highly effective implicit representation that locally aligns pixels of 2D images with the global context of their corresponding 3D object. Using PIFu, we propose an end-to-end deep learning method for digitizing highly detailed clothed humans that can infer both 3D surface and texture from a single image, and optionally, multiple input images. Highly intricate shapes, such as hairstyles, clothing, as well as their variations and deformations can be digitized in a unified way. Compared to existing representations used for 3D deep learning, PIFu can produce high-resolution surfaces including largely unseen regions such as the back of a person. In particular, it is memory efficient unlike the voxel representation, can handle arbitrary topology, and the resulting surface is spatially aligned with the input image. Furthermore, while previous techniques are designed to process either a single image or multiple views, PIFu extends naturally to arbitrary number of views. We demonstrate high-resolution and robust reconstructions on real world images from the DeepFashion dataset, which contains a variety of challenging clothing types. Our method achieves state-of-the-art performance on a public benchmark and outperforms the prior work for clothed human digitization from a single image.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 PIFu: Pixel-Aligned Implicit Function
  • 3.1 Single-view Surface Reconstruction
  • 3.2 Texture Inference
  • 3.3 Multi-View Stereo
  • 4 Experiments
  • 4.1 Quantitative Results
  • 4.2 Qualitative Results
  • 5 Discussion
  • References

Knowls

  1. Knowl 1 — Pixel-Aligned Implicit Function (PIFu) Formulation

    model/method

    The Pixel-aligned Implicit Function (PIFu) defines a continuous 3D implicit surface or property field conditioned on local 2D image features. Given a 3D coordinate X∈R3X \in \mathbb{R}^3, its 2D projection on the image plane is x=π(X)x = \pi(X) under a camera model (such as weak-perspective projection), and its depth along the camera ray is z(X)z(X). A fully convolutional image encoder gg maps the input image II to a spatial feature map. The pixel-aligned feature vector at continuous 2D location xx is obtained via bilinear interpolation:

    F(x)=g(I(x))F(x) = g(I(x))

    An implicit function ff, parameterized by a multi-layer perceptron (MLP), maps the concatenated pixel-aligned feature vector and the continuous depth coordinate to the target implicit value ss:

    f(F(x),z(X))=swhere s∈Rf(F(x), z(X)) = s \quad \text{where } s \in \mathbb{R}

    For 3D surface geometry, ss represents the continuous inside/outside probability field whose level set defines the 3D surface mesh. Because the feature vector F(x)F(x) is extracted locally at the projected pixel coordinates rather than from a global image pooling operation, the representation preserves fine-grained spatial details (such as clothing wrinkles and hair geometry) while maintaining memory efficiency and topology-free reconstruction.

  2. Knowl 2 — Single-View Surface Reconstruction and Loss Formulation

    model/method

    For 3D surface geometry reconstruction, the ground truth 3D shape is represented by an indicator occupancy field fv∗(X)f_v^*(X) defined over points X∈R3X \in \mathbb{R}^3:

    fv∗(X)={1,if X is inside the mesh surface0,otherwisef_v^*(X) = \begin{cases} 1, & \text{if } X \text{ is inside the mesh surface} \\ 0, & \text{otherwise} \end{cases}

    The surface reconstruction network consists of an image encoder gg producing geometric feature maps FV(x)F_V(x) and an implicit function fvf_v. Given nn sampled 3D points {Xi}i=1n\{X_i\}_{i=1}^n with 2D projections xi=π(Xi)x_i = \pi(X_i) and depths z(Xi)z(X_i), the network parameters are trained end-to-end by minimizing the mean squared error loss LV\mathcal{L}_V:

    LV=1n∑i=1n∣fv(FV(xi),z(Xi))−fv∗(Xi)∣2\mathcal{L}_V = \frac{1}{n} \sum_{i=1}^n \left| f_v(F_V(x_i), z(X_i)) - f_v^*(X_i) \right|^2

    During inference, the 3D volume is densely sampled on a uniform grid, evaluating the predicted occupancy probabilities fv(FV(x),z(X))f_v(F_V(x), z(X)). The explicit 3D mesh surface is extracted as the 0.5 iso-surface of the predicted occupancy field using the Marching Cubes algorithm.

  3. Knowl 3 — Surface-Adaptive and Uniform Point Sampling Strategy

    model/method

    Training the implicit occupancy field requires sampling 3D query points X∈R3X \in \mathbb{R}^3 and evaluating ground truth inside/outside status via ray tracing against watertight meshes. The sampling distribution combines surface-adaptive sampling and uniform bounding box sampling in a 16:116:1 ratio:

    1. Surface-Adaptive Sampling: Points are randomly sampled directly on the ground truth surface mesh and perturbed by adding 3D offset vectors sampled from an isotropic Gaussian distribution N(0,σ2I3)\mathcal{N}(0, \sigma^2 I_3) with standard deviation σ=5.0 cm\sigma = 5.0\text{ cm}. This concentrates sampling near the iso-surface boundary to capture sharp geometric details.
    2. Uniform Sampling: Points are uniformly sampled inside the 3D bounding box enclosing the object to provide negative supervision and prevent the network from predicting spurious geometry far from the surface.
  4. Knowl 4 — Surface-Conditioned Texture Inference (Tex-PIFu)

    model/method

    Tex-PIFu predicts RGB color values on 3D surfaces with arbitrary topology and self-occlusions by defining the implicit function output as a 3-channel color field fcf_c. To prevent severe overfitting and decouple appearance prediction from shape learning, the texture encoder is conditioned on the geometric image feature FV(x)F_V(x) learned during surface reconstruction.

    Query points are sampled on the ground truth surface mesh Xi∈ΩX_i \in \Omega with surface normal NiN_i and perturbed along the normal direction by an offset ϵ∼N(0,d2)\epsilon \sim \mathcal{N}(0, d^2) with standard deviation d=1.0 cmd = 1.0\text{ cm}:

    Xi′=Xi+ϵ⋅NiX'_i = X_i + \epsilon \cdot N_i

    Let xi′=π(Xi′)x'_i = \pi(X'_i) and Xi,z′=z(Xi′)X'_{i, z} = z(X'_i). The texture implicit network fcf_c is trained using the average L1L_1 loss over nn sampled perturbed points:

    LC=1n∑i=1n∣fc(FC(xi′,FV),Xi,z′)−C(Xi)∣\mathcal{L}_C = \frac{1}{n} \sum_{i=1}^n \left| f_c(F_C(x'_i, F_V), X'_{i, z}) - C(X_i) \right|

    where C(Xi)∈[0,1]3C(X_i) \in [0, 1]^3 is the ground truth RGB color at the surface point XiX_i, and FC(xi′,FV)F_C(x'_i, F_V) is the feature vector from the texture image encoder conditioned on the surface reconstruction feature map.

  5. Knowl 5 — Multi-View PIFu via Latent Feature Aggregation

    model/method

    PIFu extends to an arbitrary number of calibrated views V≥1V \ge 1 by decomposing the implicit function ff into a feature embedding network f1f_1 and a multi-view reasoning network f2f_2 (f=f2∘f1f = f_2 \circ f_1):

    1. Feature Embedding (f1f_1): For each view k∈{1,…,V}k \in \{1, \dots, V\}, the shared 3D point XX in world coordinates is projected to image coordinates xk=πk(X)x_k = \pi_k(X) with depth zk(X)z_k(X). f1f_1 maps the view-specific pixel feature Fk(xk)F_k(x_k) and depth zk(X)z_k(X) to an intermediate latent feature embedding Φk∈Rne\Phi_k \in \mathbb{R}^{n_e}: Φk=f1(Fk(xk),zk(X))\Phi_k = f_1(F_k(x_k), z_k(X))
    2. Feature Fusion: Latent embeddings from all available views are aggregated via average pooling: Φˉ=1V∑k=1VΦk\bar{\Phi} = \frac{1}{V} \sum_{k=1}^V \Phi_k
    3. Multi-View Reasoning (f2f_2): The reasoning network maps the pooled embedding to the target implicit field ss (occupancy probability for geometry or RGB color for texture): s=f2(Φˉ)s = f_2(\bar{\Phi})

    This additive pooling formulation naturally accommodates varying numbers of input viewpoints without architectural changes and collapses to single-view inference when V=1V=1.

  6. Knowl 6 — High-Fidelity Clothed Human Dataset and PRT Rendering Setup

    experimental setup

    The training dataset consists of 491 high-resolution photogrammetry 3D scans of clothed humans from RenderPeople (each mesh comprising approximately 100,000 triangles), split into 442 training subjects and 49 test subjects.

    To simulate realistic illumination and global light transport effects (such as ambient occlusion and self-shadowing) on synthetic training images, a Precomputed Radiance Transfer (PRT) method is employed. Surface visibility is precomputed using spherical harmonics. Rendering utilizes 163 second-order spherical harmonics environment maps of indoor scenes from HDRI Haven with random rotations around the vertical (yy) axis.

    Each subject is centered in a weak-perspective camera frame and rendered at 512×512512 \times 512 resolution across 360 yaw angles (1∘1^\circ increments), yielding 442×360=159,120442 \times 360 = 159,120 training images. For evaluation, 49 test subjects from RenderPeople and 5 subjects from the BUFF dataset are rendered from 4 orthogonal viewpoints (0∘,90∘,180∘,270∘0^\circ, 90^\circ, 180^\circ, 270^\circ yaw).

  7. Knowl 7 — PIFu Network Architecture and Training Parameters

    model/method

    The single-view PIFu framework uses the following architectures and training protocols:

    • Surface Geometry: The image encoder is a Stacked Hourglass network with Group Normalization, outputting a 256-dimensional feature map FV(x)∈R256F_V(x) \in \mathbb{R}^{256}. The implicit MLP has layer widths (257,1024,512,256,128,1)(257, 1024, 512, 256, 128, 1) with LeakyReLU activations and a final sigmoid output. Skip connections feed the input vector (FV(x),z)(F_V(x), z) into each intermediate MLP layer. Training uses RMSProp with initial learning rate 1×10−31 \times 10^{-3} (decayed by 0.1 at epoch 10), batch size 3, 5,000 sampled points per object, and runs for 12 epochs (4 days on a single NVIDIA GTX 1080Ti).
    • Texture Inference: The image encoder is a CycleGAN generator backbone with 6 residual blocks outputting FC(x)∈R256F_C(x) \in \mathbb{R}^{256}. The Tex-PIFu MLP takes 513 input channels (concatenating FC(x)∈R256F_C(x) \in \mathbb{R}^{256}, FV(x)∈R256F_V(x) \in \mathbb{R}^{256}, and z∈Rz \in \mathbb{R}) with layer widths (513,1024,512,256,128,3)(513, 1024, 512, 256, 128, 3) and a final tanh activation. Training uses Adam with learning rate 1×10−31 \times 10^{-3}, batch size 5, 10,000 sampled points per object, and runs for 6 epochs (2 days on a single 1080Ti).
    • Multi-View PIFu: The multi-view model uses the 4th layer output of the MLP (width 256) as the embedding Φk\Phi_k, pooling embeddings across views before feeding into the remaining MLP layers. It is fine-tuned from the single-view weights with learning rate 1×10−41 \times 10^{-4} for 2 epochs (1 day on a single 1080Ti).
  8. Knowl 8 — Single-View Human Digitization Quantitative Benchmark

    data/table

    Reconstruction accuracy is evaluated using three metrics: normal reprojection error (L2L_2 error between rendered normal maps of reconstructed and ground truth surfaces in the input view), Point-to-Surface (P2S) distance (average Euclidean distance from reconstructed vertices to ground truth in cm), and Chamfer distance (in cm).

    RenderPeople Buff
    Methods Normal P2S Chamfer Normal P2S Chamfer
    BodyNet 0.262 5.72 5.64 0.308 4.94 4.52
    SiCloPe 0.216 3.81 4.02 0.222 4.06 3.99
    IM-GAN 0.258 2.87 3.14 0.337 5.11 5.32
    VRN 0.116 1.42 1.56 0.130 2.33 2.48
    Ours (PIFu) 0.084 1.52 1.50 0.0928 1.15 1.14

    PIFu achieves the lowest normal reprojection error across both datasets (0.084 and 0.0928) due to pixel alignment preserving fine surface details. It outperforms global implicit representations (IM-GAN) and voxel regression (VRN) in geometric generalization on unseen datasets (BUFF).

  9. Knowl 9 — Multi-View Clothed Human Reconstruction Quantitative Evaluation

    data/table

    Multi-view PIFu trained and evaluated with 3 input views is quantitatively compared against learned multi-view stereo (LSM) and Deep Visual Hull (Huang et al., which corresponds to ablated PIFu without depth conditioning zz). In addition, 3-view PIFu is compared against a video-based parametric fitting method (Alldieck et al. 2018) using a dense 360-degree video sequence on the BUFF dataset.

    RenderPeople Buff
    Methods (3 views) Normal P2S Chamfer Normal P2S Chamfer
    LSM 0.251 4.40 3.93 0.272 3.58 3.30
    Deep V-Hull 0.093 0.639 0.632 0.119 0.698 0.709
    Ours (PIFu, 3 views) 0.094 0.554 0.567 0.107 0.665 0.641
    Methods Normal P2S Chamfer
    Alldieck et al. 18 (Dense Video) 0.127 0.820 0.795
    Ours (PIFu, 3 views) 0.107 0.665 0.641

    Multi-view PIFu using 3 calibrated views achieves lower P2S and Chamfer errors than voxel stereo (LSM) and outperforms dense monocular video fitting while operating on sparse views.

  10. Knowl 10 — Ablation on Sampling Strategies and Image Encoder Architectures

    data/table

    Ablation experiments evaluate the impact of 3D point sampling strategies and image encoder architectures on surface reconstruction quality across RenderPeople and BUFF datasets.

    RenderPeople Buff
    Sampling Method Normal P2S Chamfer Normal P2S Chamfer
    Uniform 0.119 5.07 4.23 0.132 5.98 4.53
    σ=3 cm\sigma = 3\text{ cm} 0.104 2.03 1.62 0.114 6.15 3.81
    σ=5 cm\sigma = 5\text{ cm} 0.105 1.73 1.55 0.115 1.54 1.41
    σ=15 cm\sigma = 15\text{ cm} 0.100 1.49 1.43 0.105 1.37 1.26
    σ=5 cm\sigma = 5\text{ cm} + Uniform (16:1) 0.084 1.52 1.50 0.092 1.15 1.14
    Image Encoder Normal P2S Chamfer Normal P2S Chamfer
    VGG16 0.125 3.02 2.25 0.144 4.65 3.08
    ResNet34 0.097 1.49 1.43 0.099 1.68 1.50
    Stacked Hourglass (HG) 0.084 1.52 1.50 0.092 1.15 1.14

    Uniform sampling alone produces oversmoothed boundaries, while pure surface-adaptive sampling without uniform samples causes spurious surface artifacts outside the subject. Combining σ=5 cm\sigma = 5\text{ cm} adaptive sampling with uniform sampling achieves the best balance. For image encoders, Stacked Hourglass achieves superior out-of-domain generalization on BUFF compared to ResNet34 and VGG16.

  11. Knowl 11 — Assumptions and Limitations of PIFu

    limitation

    The standard PIFu model possesses several explicit operational constraints and failure modes:

    1. Foreground Isolation and Occlusion: The model requires the human subject to be fully visible and segmented from the background (e.g., using off-the-shelf segmentation and GrabCut). It cannot reconstruct subjects occluded by other people or external scene objects.
    2. Scale Ambiguity: Reconstruction operates in normalized pixel coordinates with pre-aligned subject scale; predicting metric absolute scale from a single image remains unresolved.
    3. Generalization to Multi-Category Objects: When trained on diverse class-agnostic 3D objects (e.g., ShapeNet), PIFu struggles to infer globally coherent 3D shapes from purely local pixel-aligned features without explicit global shape conditioning.

Coverage note — None was omitted; all key contributions—including mathematical formulation, surface and texture network pipelines, multi-view feature aggregation, PRT synthetic data rendering, empirical quantitative benchmarks, ablations, and stated limitations—are fully covered.

References

  1. 1.Thiemo Alldieck, Marcus Magnor, Bharat Lal Bhatnagar, Christian Theobalt, and Gerard Pons-Moll. Learning to reconstruct people in clothing from a single RGB camera. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1175–1186, 2019.
  2. 2.Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Detailed human avatars from monocular video. In International Conference on 3D Vision, pages 98–109, 2018.
  3. 3.Thiemo Alldieck, Marcus A Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruction of 3d people models. In IEEE Conference on Computer Vision and Pattern Recognition, pages 8387–8397, 2018.
  4. 4.Dragomir Anguelov, Praveen Srinivasan, Daphne Koller, Sebastian Thrun, Jim Rodgers, and James Davis. SCAPE: shape completion and animation of people. ACM Transactions on Graphics, 24(3):408–416, 2005.
  5. 5.Alexandru O Balan, Leonid Sigal, Michael J Black, James E Davis, and Horst W Haussecker. Detailed human shape and pose from images. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1–8, 2007.
  6. 6.Aayush Bansal, Xinlei Chen, Bryan Russell, Abhinav Gupta, and Deva Ramanan. Pixelnet: Representation of the pixels, by the pixels, and for the pixels. arXiv:1702.06506, 2017.
  7. 7.Gavin Barill, Neil Dickson, Ryan Schmidt, David I.W. Levin, and Alec Jacobson. Fast winding numbers for soups and clouds. ACM Transactions on Graphics, 37(4):43, 2018.
  8. 8.Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J Black. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In European Conference on Computer Vision, pages 561–578, 2016.
  9. 9.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015.
  10. 10.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In European Conference on Computer Vision, pages 801–818, 2018.
  11. 11.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In IEEE Conference on Computer Vision and Pattern Recognition, pages 5939–5948, 2019.
  12. 12.Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European Conference on Computer Vision, pages 628–644, 2016.
  13. 13.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  14. 14.Carlos Hernández Esteban and Francis Schmitt. Silhouette and stereo fusion for 3d object modeling. Computer Vision and Image Understanding, 96(3):367–392, 2004.
  15. 15.Yasutaka Furukawa and Jean Ponce. Carved visual hulls for image-based modeling. In European Conference on Computer Vision, pages 564–577, 2006.
  16. 16.Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multiview stereopsis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(8):1362–1376, 2010.
  17. 17.Juergen Gall, Carsten Stoll, Edilson De Aguiar, Christian Theobalt, Bodo Rosenhahn, and Hans-Peter Seidel. Motion capture using joint skeleton tracking and surface estimation. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1746–1753, 2009.
  18. 18.Andrew Gilbert, Marco Volino, John Collomosse, and Adrian Hilton. Volumetric performance capture from minimal camera viewpoints. In European Conference on Computer Vision, pages 566–581, 2018.
  19. 19.Thibault Groueix, Matthew Fisher, Vladimir G Kim, Bryan C Russell, and Mathieu Aubry. Atlasnet: A papier-mâché approach to learning 3d surface generation. In IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  20. 20.Peng Guan, Alexander Weiss, Alexandru O Balan, and Michael J Black. Estimating human shape and pose from a single image. In IEEE International Conference on Computer Vision, pages 1381–1388, 2009.
  21. 21.Rıza Alp Güler, Natalia Neverova, and Iasonas Kokkinos. Densepose: Dense human pose estimation in the wild. In IEEE Conference on Computer Vision and Pattern Recognition, pages 7297–7306, 2018.
  22. 22.Christian Häne, Shubham Tulsiani, and Jitendra Malik. Hierarchical surface prediction for 3d object reconstruction. In arXiv preprint arXiv:1704.00710. 2017.
  23. 23.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  24. 24.Haibin Huang, Evangelos Kalogerakis, Siddhartha Chaudhuri, Duygu Ceylan, Vladimir G Kim, and Ersin Yumer. Learning local shape descriptors from part correspondences with multiview convolutional networks. ACM Transactions on Graphics, 37(1):6, 2018.
  25. 25.Yinghao Huang, Federica Bogo, Christoph Lassner, Angjoo Kanazawa, Peter V Gehler, Javier Romero, Ijaz Akhter, and Michael J Black. Towards accurate marker-less human shape and pose estimation over time. In International Conference on 3D Vision, pages 421–430, 2017.
  26. 26.Zeng Huang, Tianye Li, Weikai Chen, Yajie Zhao, Jun Xing, Chloe LeGendre, Linjie Luo, Chongyang Ma, and Hao Li. Deep volumetric video from very sparse multi-view performance capture. In European Conference on Computer Vision, pages 336–354, 2018.
  27. 27.Aaron S Jackson, Chris Manafas, and Georgios Tzimiropoulos. 3D Human Body Reconstruction from a Single Image via Volumetric Regression. In ECCV Workshop Proceedings, PeopleCap 2018, pages 0–0, 2018.
  28. 28.Mengqi Ji, Juergen Gall, Haitian Zheng, Yebin Liu, and Lu Fang. Surfacenet: An end-to-end 3d neural network for multiview stereopsis. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2307–2315, 2017.
  29. 29.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision, pages 694–711. Springer, 2016.
  30. 30.Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In IEEE Conference on Computer Vision and Pattern Recognition, pages 7122–7131, 2018.
  31. 31.Angjoo Kanazawa, Shubham Tulsiani, Alexei A. Efros, and Jitendra Malik. Learning category-specific mesh reconstruction from image collections. In European Conference on Computer Vision, pages 371–386, 2018.
  32. 32.Abhishek Kar, Christian Häne, and Jitendra Malik. Learning a multi-view stereo machine. In Advances in Neural Information Processing Systems, pages 364–375, 2017.
  33. 33.Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J Black, and Peter V Gehler. Unite the people: Closing the loop between 3d and 2d human representations. In IEEE Conference on Computer Vision and Pattern Recognition, pages 6050–6059, 2017.
  34. 34.Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In IEEE Conference on Computer Vision and Pattern Recognition, pages 1096–1104, 2016.
  35. 35.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. SMPL: A skinned multi-person linear model. ACM Transactions on Graphics, 34(6):248, 2015.
  36. 36.William E Lorensen and Harvey E Cline. Marching cubes: A high resolution 3d surface construction algorithm. In ACM siggraph computer graphics, volume 21, pages 163–169. ACM, 1987.
  37. 37.Wojciech Matusik, Chris Buehler, Ramesh Raskar, Steven J Gortler, and Leonard McMillan. Image-based visual hulls. In ACM SIGGRAPH, pages 369–374, 2000.
  38. 38.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. arXiv preprint arXiv:1812.03828, 2018.
  39. 39.Ryota Natsume, Shunsuke Saito, Zeng Huang, Weikai Chen, Chongyang Ma, Hao Li, and Shigeo Morishima. Siclope: Silhouette-based clothed people. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4480–4490, 2019.
  40. 40.Natalia Neverova, Riza Alp Guler, and Iasonas Kokkinos. Dense pose transfer. In European Conference on Computer Vision, pages 123–138, 2018.
  41. 41.Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. In European Conference on Computer Vision, pages 483–499, 2016.
  42. 42.Mohamed Omran, Christoph Lassner, Gerard Pons-Moll, Peter V. Gehler, and Bernt Schiele. Neural body fitting: Unifying deep learning and model-based human pose and shape estimation. In International Conference on 3D Vision, pages 484–494, 2018.
  43. 43.Eunbyung Park, Jimei Yang, Ersin Yumer, Duygu Ceylan, and Alexander C. Berg. Transformation-grounded image generation network for novel 3d view synthesis. In IEEE Conference on Computer Vision and Pattern Recognition, pages 3500–3509, 2017.
  44. 44.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. arXiv preprint arXiv:1901.05103, 2019.
  45. 45.Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros. Context encoders: Feature learning by inpainting. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2536–2544, 2016.
  46. 46.Georgios Pavlakos, Luyang Zhu, Xiaowei Zhou, and Kostas Daniilidis. Learning to estimate 3D human pose and shape from a single color image. In IEEE Conference on Computer Vision and Pattern Recognition, pages 459–468, 2018.
  47. 47.Gerard Pons-Moll, Sergi Pujades, Sonny Hu, and Michael J Black. Clothcap: Seamless 4d clothing capture and retargeting. ACM Transactions on Graphics, 36(4):73, 2017.
  48. 48.Renderpeople, 2018. https://renderpeople.com/3d-people.
  49. 49.Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. Grabcut: Interactive foreground extraction using iterated graph cuts. ACM Transactions on Graphics, 23(3):309–314, 2004.
  50. 50.Stan Sclaroff and Alex Pentland. Generalized implicit functions for computer graphics, volume 25. ACM, 1991.
  51. 51.Evan Shelhamer, Jonathan Long, and Trevor Darrell. Fully convolutional networks for semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 39(4):640–651, 2017.
  52. 52.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  53. 53.Peter-Pike Sloan, Jan Kautz, and John Snyder. Precomputed radiance transfer for real-time rendering in dynamic, low-frequency lighting environments. In ACM Transactions on Graphics, volume 21, pages 527–536, 2002.
  54. 54.Cristian Sminchisescu and Alexandru Telea. Human pose estimation from silhouettes. a consistent approach using distance level sets. In International Conference on Computer Graphics, Visualization and Computer Vision, volume 10, 2002.
  55. 55.Jonathan Starck and Adrian Hilton. Surface capture for performance-based animation. IEEE Computer Graphics and Applications, 27(3):21–31, 2007.
  56. 56.Shao-Hua Sun, Minyoung Huh, Yuan-Hong Liao, Ning Zhang, and Joseph J Lim. Multi-view to novel view: Synthesizing novel views with self-learned confidence. In European Conference on Computer Vision, pages 155–171, 2018.
  57. 57.Jonathan J Tompson, Arjun Jain, Yann LeCun, and Christoph Bregler. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in neural information processing systems, pages 1799–1807, 2014.
  58. 58.Shubham Tulsiani, Tinghui Zhou, Alexei A. Efros, and Jitendra Malik. Multi-view supervision for single-view reconstruction via differentiable ray consistency. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2626–2634, 2017.
  59. 59.Gül Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid. BodyNet: Volumetric inference of 3D human body shapes. In European Conference on Computer Vision, pages 20–36, 2018.
  60. 60.Gül Varol, Javier Romero, Xavier Martin, Naureen Mahmood, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning from synthetic humans. In IEEE Conference on Computer Vision and Pattern Recognition, pages 109–117, 2017.
  61. 61.Daniel Vlasic, Ilya Baran, Wojciech Matusik, and Jovan Popović. Articulated mesh animation from multi-view silhouettes. ACM Transactions on Graphics, 27(3):97, 2008.
  62. 62.Daniel Vlasic, Pieter Peers, Ilya Baran, Paul Debevec, Jovan Popović, Szymon Rusinkiewicz, and Wojciech Matusik. Dynamic shape capture using multi-view photometric stereo. ACM Transactions on Graphics, 28(5):174, 2009.
  63. 63.Ingo Wald, Sven Woop, Carsten Benthin, Gregory S Johnson, and Manfred Ernst. Embree: a kernel framework for efficient cpu ray tracing. ACM Transactions on Graphics, 33(4):143, 2014.
  64. 64.Weiyue Wang, Xu Qiangeng, Duygu Ceylan, Radomir Mech, and Ulrich Neumann. Disn: Deep implicit surface network for high-quality single-view 3d reconstruction. arXiv preprint arXiv:1905.10711, 2019.
  65. 65.Michael Waschbüsch, Stephan Würmlin, Daniel Cotting, Filip Sadlo, and Markus Gross. Scalable 3D video of dynamic scenes. The Visual Computer, 21(8):629–638, 2005.
  66. 66.Shih-En Wei, Varun Ramakrishna, Takeo Kanade, and Yaser Sheikh. Convolutional pose machines. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4724–4732, 2016.
  67. 67.Chung-Yi Weng, Brian Curless, and Ira Kemelmacher-Shlizerman. Photo wake-up: 3d character animation from a single photo. arXiv preprint arXiv:1812.02246, 2018.
  68. 68.Chenglei Wu, Kiran Varanasi, and Christian Theobalt. Full body performance capture under uncontrolled and varying illumination: A shading-based approach. European Conference on Computer Vision, pages 757–770, 2012.
  69. 69.Yuxin Wu and Kaiming He. Group normalization. In European Conference on Computer Vision, pages 3–19, 2018.
  70. 70.Jinlong Yang, Jean-Sébastien Franco, Franck Hétroy-Wheeler, and Stefanie Wuhrer. Estimation of human body shape in motion with wide clothing. In European Conference on Computer Vision, pages 439–454, 2016.
  71. 71.Chao Zhang, Sergi Pujades, Michael Black, and Gerard Pons-Moll. Detailed, accurate, human shape estimation from clothed 3D scan sequences. In IEEE Conference on Computer Vision and Pattern Recognition, pages 4191–4200, 2017.
  72. 72.Shizhe Zhou, Hongbo Fu, Ligang Liu, Daniel Cohen-Or, and Xiaoguang Han. Parametric reshaping of human bodies in images. In ACM Transactions on Graphics, page 126, 2010.
  73. 73.Tinghui Zhou, Shubham Tulsiani, Weilun Sun, Jitendra Malik, and Alexei A Efros. View synthesis by appearance flow. In European Conference on Computer Vision, pages 286–301, 2016.
  74. 74.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In IEEE International Conference on Computer Vision, pages 2223–2232, 2017.
  75. 75.C Lawrence Zitnick, Sing Bing Kang, Matthew Uyttendaele, Simon Winder, and Richard Szeliski. High-quality video view interpolation using a layered representation. ACM Transactions on Graphics, 23(3):600–608, 2004.

Citation

MLA
Saito, S., et al. “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization”. The IEEE International Conference on Computer Vision (ICCV), 2019, Pp. 2304-2314, 2019, http://arxiv.org/abs/1905.05172v3.
APA
Saito, S., Huang, Z., Natsume, R., Morishima, S., Kanazawa, A., & Li, H. (2019). PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization. The IEEE International Conference on Computer Vision (ICCV), 2019, Pp. 2304-2314. http://arxiv.org/abs/1905.05172v3
Chicago
Saito, S., Z. Huang, R. Natsume, S. Morishima, A. Kanazawa, and H. Li. 2019. “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization”. The IEEE International Conference on Computer Vision (ICCV), 2019, Pp. 2304-2314. http://arxiv.org/abs/1905.05172v3.
Harvard
Saito, S. et al. (2019) “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization”, The IEEE International Conference on Computer Vision (ICCV), 2019, pp. 2304-2314 [Preprint]. Available at: http://arxiv.org/abs/1905.05172v3.
Vancouver
1. Saito S, Huang Z, Natsume R, Morishima S, Kanazawa A, Li H (2019) PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization. The IEEE International Conference on Computer Vision (ICCV), 2019, pp. 2304-2314

BibTeX

@article{saito2019pifu,
  title = {PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization},
  author = {Saito, Shunsuke and Huang, Zeng and Natsume, Ryota and Morishima, Shigeo and Kanazawa, Angjoo and Li, Hao},
  year = {2019},
  journal = {The IEEE International Conference on Computer Vision (ICCV), 2019, pp. 2304-2314},
  url = {http://arxiv.org/abs/1905.05172v3},
  eprint = {1905.05172}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE