Learning Neural Parametric Head Models

Simon GiebenhainTobias KirschsteinMarkos GeorgopoulosMartin RünzLourdes AgapitoMatthias Nießner

article2023CVPR87 citations

Proposes a neural parametric head model using local signed distance fields and forward deformation fields trained on high-resolution full-head scans to achieve fine geometry reconstruction with disentangled identity and expression controls.

Listen

Digital representations of complete human heads are essential for applications across virtual reality, gaming, telepresence, and digital avatars. Traditional 3D morphable face models rely on fixed-topology mesh templates and linear dimensionality reduction methods such as Principal Component Analysis. While these conventional methods provide good regularization against noisy inputs, their rigid structural assumptions prevent them from capturing fine local facial details, natural variations in head shape, and diverse hair styles.

The article develops and demonstrates a novel neural parametric head model (NPHM) based on hybrid neural fields. The primary objective is to accurately reconstruct full 3D head geometry and complex facial expressions from sparse, noisy depth data by learning separate, disentangled representations for personal identity and facial movements.

To build and validate this system, the authors captured a high-resolution 3D head dataset comprising over 3,700 scans across 203 distinct individuals performing 23 facial expressions. Each scan averaged roughly 1.5 million vertices and 3.5 million triangles. The approach uses an implicit neural field to represent the canonical geometry of an identity as a signed distance field, broken down into an ensemble of smaller local networks anchored to specific facial keypoints with shared weights across symmetrical facial regions. A separate forward deformation network then models facial expressions. When deployed at inference time, the model fits to sparse point clouds—such as single depth frames from consumer-grade sensors—by optimizing the underlying identity and expression codes.

The evaluation produced several clear findings. First, the proposed framework significantly outperformed traditional mesh-based models such as FLAME and Basel Face Models, achieving lower reconstruction error (an L1-Chamfer distance of 0.00182 versus 0.00640 for FLAME during identity fitting) and higher geometric accuracy (F-Score of 0.954 versus 0.530). Second, it outperformed competing neural implicit baselines like ImFace and Neural Parametric Models across both identity and expression tasks. Third, ablation analyses confirmed that decomposing the face into localized regions and sharing symmetric weights directly improves geometric accuracy and surface normal consistency.

These results demonstrate that moving from fixed mesh templates to localized neural fields eliminates key structural bottlenecks in digital human modeling. For engineering and product teams, this approach enables high-fidelity 3D avatar generation and tracking directly from consumer depth sensors without requiring dense multi-camera capture rigs or manual clean-up. Additionally, the forward deformation design allows faster animation compared to backward deformation architectures.

Organizations developing 3D facial capture and avatar pipelines should consider adopting local implicit neural representations over legacy mesh templates for tracking and geometry reconstruction tasks. Future development should focus on integrating photometric texture pipelines, as the current model focuses exclusively on 3D geometry rather than visual color appearance. Furthermore, while the model captures full head shapes with short hair, extending the dataset to accommodate long, loose hair styles remains necessary to broaden real-world applicability.

Cover for Learning Neural Parametric Head Models

Abstract

We propose a novel 3D morphable model for complete human heads based on hybrid neural fields. At the core of our model lies a neural parametric representation that disentangles identity and expressions in disjoint latent spaces. To this end, we capture a person’s identity in a canonical space as a signed distance field (SDF), and model facial expressions with a neural deformation field. In addition, our representation achieves high-fidelity local detail by introducing an ensemble of local fields centered around facial anchor points. To facilitate generalization, we train our model on a newly-captured dataset of over 3700 head scans from 203 different identities using a custom high-end 3D scanning setup. Our dataset significantly exceeds comparable existing datasets, both with respect to quality and completeness of geometry, averaging around 3.5M mesh faces per scan1. Finally, we demonstrate that our approach outperforms state-of-the-art methods in terms of fitting error and reconstruction quality.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Dataset Acquisition
  • 3.1. Capture Setup
  • 3.2. Registration Pipeline
  • 3.2.1 Rigid Alignment
  • 3.2.2 Non-Rigid Registration
  • 4. Neural Parametric Head Models
  • 4.1. Identity Representation
  • 4.1.1 Local Decomposition
  • 4.1.2 Global Blending
  • 4.2. Expression Representation
  • 4.3. Training Strategy
  • 5. Results
  • 5.1. Identity Reconstruction
  • 5.2. Expression Reconstruction
  • 5.3. Ablations
  • 5.4. Limitations
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Canonical Identity Representation via Local MLP Ensemble and Symmetry

    model/method

    Neural Parametric Head Models (NPHM) represent human head identity as a continuous Signed Distance Field (SDF) in a canonical open-mouth space. To capture localized geometric details and reduce optimization complexity, the canonical field is partitioned into K=2Ksymm+KmiddleK = 2K_{\text{symm}} + K_{\text{middle}} semantic regions centered at anchor points a∈RK×3\mathbf{a} \in \mathbb{R}^{K \times 3}, where KmiddleK_{\text{middle}} anchors lie on the facial symmetry axis M\mathcal{M}, and 2Ksymm2K_{\text{symm}} anchors form symmetric pairs between left-side regions S\mathcal{S} and right-side regions S∗\mathcal{S}^*.

    Each region is parameterized by an identity-specific local latent code zkid∈Rdloc\mathbf{z}_k^{\text{id}} \in \mathbb{R}^{d_{\text{loc}}} and conditioned on a shared global identity latent code zglobid∈Rdglob\mathbf{z}_{\text{glob}}^{\text{id}} \in \mathbb{R}^{d_{\text{glob}}}. The kk-th local field for a 3D query point x∈R3x \in \mathbb{R}^3 is predicted by an MLP:

    fk(x,zglobid,zkid)=MLP⁡θk([x−ak,zglobid,zkid])f_k(x, \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_k^{\text{id}}) = \operatorname{MLP}_{\theta_k}([x - \mathbf{a}_k, \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_k^{\text{id}}])

    where [ ⋅ ][\,\cdot\,] denotes vector concatenation. Facial bilateral symmetry is enforced by sharing network parameters between each region k∈Sk \in \mathcal{S} on the left side and its symmetric counterpart k∗∈S∗k^* \in \mathcal{S}^* on the right side while mirroring spatial coordinates along the symmetry axis:

    fk∗(x,zglobid,zk∗id)=fk(flip⁡(x−ak∗),zglobid,zk∗id)f_{k^*}(x, \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_{k^*}^{\text{id}}) = f_k(\operatorname{flip}(x - \mathbf{a}_{k^*}), \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_{k^*}^{\text{id}})

    Anchor coordinates a\mathbf{a} are predicted from the global latent code using a dedicated network a=MLP⁡pos(zglobid)\mathbf{a} = \operatorname{MLP}_{\text{pos}}(\mathbf{z}_{\text{glob}}^{\text{id}}).

  2. Knowl 2 — Global Blending of Local SDFs with Background Field

    model/method

    To combine the ensemble of local SDF networks into a continuous global SDF Fid(x)\mathscr{F}_{\text{id}}(x), NPHM computes a spatially normalized weighted sum of local fields fkf_k and an additional global background field f0f_0 centered at a0=0∈R3\mathbf{a}_0 = \mathbf{0} \in \mathbb{R}^3:

    Fid(x)=∑k=0Kwk(x,ak)fk(x,zglobid,zkid)\mathscr{F}_{\text{id}}(x) = \sum_{k=0}^{K} w_k(x, \mathbf{a}_k) f_k(x, \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_k^{\text{id}})

    The background network f0(x,zglobid,z0id)=MLP⁡0(x,zglobid,z0id)f_0(x, \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_0^{\text{id}}) = \operatorname{MLP}_0(x, \mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_0^{\text{id}}) operates in global coordinates to represent geometry far away from any facial anchor. The blending weights wk(x,ak)w_k(x, \mathbf{a}_k) are computed using normalized isotropic Gaussian kernels with standard deviation σ\sigma and a constant baseline response cc for the background field:

    wk∗(x,ak)={exp⁡(−∥x−ak∥22σ),if k>0c,if k=0w_k^*(x, \mathbf{a}_k) = \begin{cases} \exp\left(-\frac{\|x - \mathbf{a}_k\|_2}{2\sigma}\right), & \text{if } k > 0 \\ c, & \text{if } k = 0 \end{cases}

    wk(x,ak)=wk∗(x,ak)∑k′=0Kwk′∗(x,ak′)w_k(x, \mathbf{a}_k) = \frac{w_k^*(x, \mathbf{a}_k)}{\sum_{k'=0}^K w_{k'}^*(x, \mathbf{a}_{k'})}

  3. Knowl 3 — Forward Neural Deformation Field with Identity Bottleneck

    model/method

    NPHM models facial expressions as a forward deformation field Fex\mathscr{F}_{\text{ex}} mapping canonical 3D coordinates x∈R3x \in \mathbb{R}^3 to 3D deformation vectors in posed space. The deformation MLP is conditioned on a latent expression code zex∈Rdex\mathbf{z}^{\text{ex}} \in \mathbb{R}^{d_{\text{ex}}} and an identity embedding Z^id∈Rdid-ex\hat{\mathbf{Z}}^{\text{id}} \in \mathbb{R}^{d_{\text{id-ex}}}:

    Fex(x,zex,Z^id):Rdex+did-ex→R3\mathscr{F}_{\text{ex}}(x, \mathbf{z}^{\text{ex}}, \hat{\mathbf{Z}}^{\text{id}}): \mathbb{R}^{d_{\text{ex}} + d_{\text{id-ex}}} \rightarrow \mathbb{R}^3

    To prevent the expression code from encoding identity-specific geometric variations while retaining the necessary conditioning for identity-dependent deformations, an information bottleneck is applied by projecting the concatenated global/local identity latent vectors and anchor positions through a single learned linear layer WW:

    Z^id=W[zglobid,z0id,…,zKid,a1,…,aK]\hat{\mathbf{Z}}^{\text{id}} = W [\mathbf{z}_{\text{glob}}^{\text{id}}, \mathbf{z}_0^{\text{id}}, \dots, \mathbf{z}_K^{\text{id}}, \mathbf{a}_1, \dots, \mathbf{a}_K]

  4. Knowl 4 — Auto-Decoder Training Objective for Disentangled Canonical and Deformation Fields

    model/method

    NPHM trains identity and expression spaces sequentially in an auto-decoder framework.

    1. Identity Training: For each subject j∈Jj \in J, latent codes Zjid={zglob,jid,z0,jid,…,zK,jid}\mathbf{Z}_j^{\text{id}} = \{\mathbf{z}_{\text{glob},j}^{\text{id}}, \mathbf{z}_{0,j}^{\text{id}}, \dots, \mathbf{z}_{K,j}^{\text{id}}\} and network parameters {θpos,θ0,…,θK}\{\theta_{\text{pos}}, \theta_0, \dots, \theta_K\} are optimized by minimizing:

    Lid=∑j∈J[LIGR+λa∥a^j−aj∥22+λsyLsy+λregid∥Zjid∥22]\mathcal{L}_{\text{id}} = \sum_{j \in J} \left[ \mathcal{L}_{\text{IGR}} + \lambda_a \|\hat{\mathbf{a}}_j - \mathbf{a}_j\|_2^2 + \lambda_{\text{sy}} \mathcal{L}_{\text{sy}} + \lambda_{\text{reg}}^{\text{id}} \|\mathbf{Z}_j^{\text{id}}\|_2^2 \right]

    where LIGR\mathcal{L}_{\text{IGR}} enforces zero SDF on surface samples with Eikonal gradient regularization, a^j\hat{\mathbf{a}}_j denotes anchor positions from registrations, and symmetry regularization is defined as:

    Lsy=∑k∈S∥zk,jid−zk∗,jid∥22\mathcal{L}_{\text{sy}} = \sum_{k \in \mathcal{S}} \|\mathbf{z}_{k,j}^{\text{id}} - \mathbf{z}_{k^*,j}^{\text{id}}\|_2^2

    1. Expression Training: With the identity model frozen, deformation network parameters θex\theta_{\text{ex}}, projection matrix WW, and per-expression latent codes {zj,lex}\{\mathbf{z}_{j,l}^{\text{ex}}\} are optimized on sampled canonical points x∈Xj,lx \in X_{j,l} with ground-truth displacement vectors δ(x)j,l\delta(x)_{j,l} via:

    Lex=∑j∈J,l∈L∑x∈Xj,l∥Fex(x,zj,lex,Z^jid)−δ(x)j,l∥22+λregex∥zj,lex∥22\mathcal{L}_{\text{ex}} = \sum_{j \in J, l \in L} \sum_{x \in X_{j,l}} \|\mathscr{F}_{\text{ex}}(x, \mathbf{z}_{j,l}^{\text{ex}}, \hat{\mathbf{Z}}_j^{\text{id}}) - \delta(x)_{j,l}\|_2^2 + \lambda_{\text{reg}}^{\text{ex}} \|\mathbf{z}_{j,l}^{\text{ex}}\|_2^2

  5. Knowl 5 — Two-Stage Scan Registration Pipeline

    model/method

    To build ground truth training pairs for canonical identity and non-rigid forward deformations, raw 3D scans are registered through a two-stage process:

    1. Rigid Alignment and FLAME Initialization: 2D facial landmarks are detected on frontal mesh renderings using MediaPipe and back-projected to the 3D scan, yielding 48 iBUG68 landmarks. An initial Umeyama similarity transform aligns them with the FLAME template face. Across the 23 expression scans {Sj}j=123\{\mathcal{S}_j\}_{j=1}^{23} of a subject, parameters Φj\Phi_j (identity zid∈R100\mathbf{z}^{\text{id}} \in \mathbb{R}^{100}, expressions zjex\mathbf{z}_j^{\text{ex}}, jaw poses θj\theta_j, shared scale s∈Rs \in \mathbb{R}, and per-scan rigid corrections Rj,tjR_j, t_j) are jointly solved by minimizing:

    arg⁡min⁡Φ1,…,Φ23∑j=123[λl∥Lj−L^j∥1+d(VΦj,Sj)+R(Φj)]\arg\min_{\Phi_1, \dots, \Phi_{23}} \sum_{j=1}^{23} \left[ \lambda_l \|L_j - \hat{L}_j\|_1 + d(V_{\Phi_j}, \mathcal{S}_j) + \mathcal{R}(\Phi_j) \right]

    where Lj∈R68×3L_j \in \mathbb{R}^{68 \times 3} and L^j\hat{L}_j denote scan and model landmarks, d(VΦj,Sj)d(V_{\Phi_j}, \mathcal{S}_j) is point-to-plane distance, and R(Φj)\mathcal{R}(\Phi_j) regularizes FLAME parameters. Hair regions are masked out using FaRL segmentation.

    1. Fine Tuning via As-Rigid-As-Possible (ARAP) Optimization: The facial mesh resolution is upsampled by a factor of 16. Vertex-specific offsets {δv}v∈V\{\delta_v\}_{v \in V} and rotations {Rv}v∈V\{R_v\}_{v \in V} are optimized per scan using L-BFGS:

    arg⁡min⁡{δv},{Rv}∑v∈V[d(v+δv,S)+∑u∈Nv∥Rv(v−u)−((v+δv)−(u+δu))∥22]\arg\min_{\{\delta_v\}, \{R_v\}} \sum_{v \in V} \left[ d(v + \delta_v, \mathcal{S}) + \sum_{u \in \mathcal{N}_v} \|R_v(v - u) - ((v + \delta_v) - (u + \delta_u))\|_2^2 \right]

    where Nv\mathcal{N}_v denotes the neighboring vertices of vv.

  6. Knowl 6 — High-Fidelity 3D Head Scanning Dataset

    experimental setup

    The captured dataset consists of 3720 high-fidelity 3D head scans from 203 subjects (144 male, 59 female; 29% female). The acquisition rig uses two Artec Eva structured light 3D scanners rotating 360∘360^\circ around the subject's head via a robotic actuator. Each full scan takes 6 seconds at 16 frames per second. Individual frames are aligned and fused into a single mesh containing approximately 1.5 million vertices and 3.5 million triangles per scan.

    Participants perform 23 distinct facial expressions based on the Facial Action Coding System (FACS) from FaceWarehouse, including a canonical neutral expression with the mouth open to prevent self-intersection and topological merging. Subjects do not wear bathing caps, enabling the capture of natural head and hair geometry.

  7. Knowl 7 — Identity Reconstruction Benchmark from Single-View Depth

    data/table

    Identity reconstruction performance is evaluated on 18 unseen test identities (6 female, 12 male) by fitting each model's identity space to a single neutral depth map containing 5000 randomly sampled points. Quality is measured by L1L_1-Chamfer distance (10−210^{-2}), normal consistency (N. C., cosine similarity), and F-Score at a 1.5 mm threshold.

    Method L1L_1-Chamfer ↓\downarrow N. C. ↑\uparrow [email protected] ↑\uparrow
    BFM 1.341×10−21.341\times 10^{-2} 0.936 0.319
    FLAME 0.640×10−20.640\times 10^{-2} 0.931 0.530
    Global PCA 0.563×10−20.563\times 10^{-2} 0.954 0.571
    Local PCA 0.416×10−20.416\times 10^{-2} 0.960 0.756
    ImFace 0.404×10−20.404\times 10^{-2} 0.954 0.832
    ImFace* 0.312×10−20.312\times 10^{-2} 0.971 0.883
    NPM 0.200×10−20.200\times 10^{-2} 0.975 0.947
    Ours (NPHM) 0.182×10−20.182\times 10^{-2} 0.978 0.954

    *ImFace* denotes training on the authors' dataset with ImFace preprocessing. NPHM outperforms mesh-based and neural implicit baselines across all metrics.

  8. Knowl 8 — Multi-Expression Reconstruction Benchmark

    data/table

    Expression reconstruction is evaluated by fitting models across 23 distinct depth map observations per subject for 18 unseen identities, solving for a shared identity code and per-scan expression codes. For forward neural deformation models (NPM and NPHM), the identity code is fixed from neutral fitting and expression codes are recovered using iterative root finding.

    Method L1L_1-Chamfer ↓\downarrow N. C. ↑\uparrow [email protected] ↑\uparrow
    BFM 1.271×10−21.271\times 10^{-2} 0.937 0.508
    FLAME 0.679×10−20.679\times 10^{-2} 0.924 0.351
    Global PCA 0.515×10−20.515\times 10^{-2} 0.956 0.606
    Local PCA 0.535×10−20.535\times 10^{-2} 0.950 0.641
    ImFace 0.369×10−20.369\times 10^{-2} 0.959 0.824
    ImFace* 0.321×10−20.321\times 10^{-2} 0.971 0.879
    NPM 0.299×10−20.299\times 10^{-2} 0.962 0.891
    Ours (NPHM) 0.272×10−20.272\times 10^{-2} 0.969 0.913

    *ImFace* is trained on the authors' dataset. NPHM achieves the lowest Chamfer error (0.272×10−20.272\times 10^{-2}) and highest [email protected] (0.913).

  9. Knowl 9 — Ablation on Anchor Count and Symmetry Constraints

    data/table

    An ablation study isolates the contributions of the number of local anchor regions KK and bilateral weight sharing on neutral identity reconstruction from a single depth map.

    Method L1L_1-Chamfer ↓\downarrow N. C. ↑\uparrow [email protected] ↑\uparrow
    NPM (effectively K=1K=1) 0.254 0.972 0.906
    K=12K=12, w/ sy. 0.289 0.966 0.876
    K=26K=26, w/ sy. 0.237 0.971 0.913
    K=39K=39, w/o sy. 0.230 0.974 0.917
    Ours (K=39K=39, w/ sy.) 0.206 0.976 0.938

    Increasing the number of anchor points up to K=39K=39 combined with weight-sharing for symmetric anchors yields the highest accuracy (L1L_1-Chamfer of 0.206 and F-Score of 0.938).

  10. Knowl 10 — Limitations of NPHM

    limitation

    Neural Parametric Head Models exhibit two primary limitations:

    1. Absence of Photometric/Appearance Modeling: NPHM models geometry exclusively via signed distance and deformation fields, without a texture or appearance field, preventing direct end-to-end fitting to single RGB images using dense photometric losses.
    2. Hair Style Constraints: While NPHM models full head shape without requiring skull caps, it does not capture dynamic, open, or long unconstrained hair, limiting the geometric diversity of represented hairstyles.

Coverage note — No substantial contributed material from the main paper was omitted. Baseline optimization setups detailed in the supplementary material were summarized within the respective empirical benchmark knowls.

References

  1. 1.Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pages 187–194, 1999. 2, 6, 7, 8
  2. 2.Timo Bolkart and Stefanie Wuhrer. A groupwise multilinear correspondence optimization for 3d faces. In Proceedings of the IEEE international conference on computer vision, pages 3604–3612, 2015. 2
  3. 3.James Booth, Epameinondas Antonakos, Stylianos Ploumpis, George Trigeorgis, Yannis Panagakis, and Stefanos Zafeiriou. 3d face morphable models” in-the-wild”. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 48–57, 2017. 2
  4. 4.James Booth, Anastasios Roussos, Stefanos Zafeiriou, Allan Ponniah, and David Dunaway. A 3d morphable model learnt from 10,000 faces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5543–5552, 2016. 2
  5. 5.Giorgos Bouritsas, Sergiy Bokhnyak, Stylianos Ploumpis, Michael Bronstein, and Stefanos Zafeiriou. Neural 3d morphable models: Spiral convolutional networks for 3d shape representation learning and generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7213–7222, 2019. 2
  6. 6.Alan Brunton, Timo Bolkart, and Stefanie Wuhrer. Multilinear wavelets: A statistical shape space for human faces. In European Conference on Computer Vision, pages 297–312. Springer, 2014. 2
  7. 7.Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou. Facewarehouse: A 3d facial expression database for visual computing. 20(3):413–425, mar 2014. 3
  8. 8.Rohan Chabra, Jan E. Lenssen, Eddy Ilg, Tanner Schmidt, Julian Straub, Steven Lovegrove, and Richard Newcombe. Deep local shapes: Learning local sdf priors for detailed 3d reconstruction. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX, page 608–625, Berlin, Heidelberg, 2020. Springer-Verlag. 2
  9. 9.Xu Chen, Tianjian Jiang, Jie Song, Jinlong Yang, Michael J. Black, Andreas Geiger, and Otmar Hilliges. gdna: Towards generative detailed neural avatars. CoRR, abs/2201.04123, 2022. 2, 5
  10. 10.Xu Chen, Yufeng Zheng, Michael J Black, Otmar Hilliges, and Andreas Geiger. Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes. In International Conference on Computer Vision (ICCV), 2021. 2, 7
  11. 11.Zhixiang Chen and Tae-Kyun Kim. Learning feature aggregation for deep 3d morphable models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13164–13173, 2021. 2
  12. 12.Kyle Genova, Forrester Cole, Avneesh Sud, Aaron Sarna, and Thomas Funkhouser. Local deep implicit functions for 3d shape. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4857–4866, 2020. 2, 5
  13. 13.Simon Giebenhain and Bastian Goldluecke. Air-nets: An attention-based framework for locally conditioned implicit representations. In 2021 International Conference on 3D Vision (3DV). IEEE, 2021. 2
  14. 14.Shunwang Gong, Lei Chen, Michael Bronstein, and Stefanos Zafeiriou. Spiralnet++: A fast and highly efficient mesh convolution operator. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 2
  15. 15.Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. Implicit geometric regularization for learning shapes. arXiv preprint arXiv:2002.10099, 2020. 6
  16. 16.Ruilong Li, Karl Bladin, Yajie Zhao, Chinmay Chinara, Owen Ingraham, Pengda Xiang, Xinglei Ren, Pratusha Prasad, Bipin Kishore, Jun Xing, and Hao Li. Learning formation of physically-based face attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2
  17. 17.Ruilong Li, Julian Tanke, Minh Vo, Michael Zollhofer, Jurgen Gall, Angjoo Kanazawa, and Christoph Lassner. Tava: Template-free animatable volumetric actors. arXiv preprint arXiv:2206.08929, 2022. 2
  18. 18.Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6):194:1–194:17, 2017. 2, 3, 6, 7, 8
  19. 19.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose control. ACM Transactions on Graphics (TOG), 40(6):1–16, 2021. 2
  20. 20.Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Yong, Juhyun Lee, Wan-Teh Chang, Wei Hua, Manfred Georg, and Matthias Grundmann. Mediapipe: A framework for perceiving and processing reality. In Third Workshop on Computer Vision for AR/VR at IEEE Computer Vision and Pattern Recognition (CVPR) 2019, 2019. 3
  21. 21.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4460–4470, 2019. 2
  22. 22.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. 2
  23. 23.Thomas Neumann, Kiran Varanasi, Stephan Wenger, Markus Wacker, Marcus Magnor, and Christian Theobalt. Sparse localized deformation components. ACM Trans. Graph., 32(6), nov 2013. 2
  24. 24.Pablo Palafox, Aljaž Božič, Justus Thies, Matthias Nießner, and Angela Dai. Npms: Neural parametric models for 3d deformable shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 12695–12705, October 2021. 2, 5, 6, 7, 8
  25. 25.Pablo Palafox, Nikolaos Sarafianos, Tony Tung, and Angela Dai. Spams: Structured implicit parametric models. CVPR, 2022. 2, 5
  26. 26.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 165–174, 2019. 2
  27. 27.Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5865–5874, 2021. 2
  28. 28.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M. Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. ACM Trans. Graph., 40(6), dec 2021. 2
  29. 29.Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami Romdhani, and Thomas Vetter. A 3d face model for pose and illumination invariant face recognition. In 2009 sixth IEEE international conference on advanced video and signal based surveillance, pages 296–301. Ieee, 2009. 2, 6, 7
  30. 30.Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, and Andreas Geiger. Convolutional occupancy networks. In European Conference on Computer Vision, pages 523–540. Springer, 2020. 2
  31. 31.Stylianos Ploumpis, Haoyang Wang, Nick Pears, William AP Smith, and Stefanos Zafeiriou. Combining 3d morphable models: A large scale face-and-head model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10934–10943, 2019. 2
  32. 32.Eduard Ramon, Gil Triginer, Janna Escur, Albert Pumarola, Jaime Garcia, Xavier Giro-i Nieto, and Francesc Moreno-Noguer. H3d-net: Few-shot high-fidelity 3d head reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5620–5629, 2021. 2
  33. 33.Christos Sagonas, Georgios Tzimiropoulos, Stefanos Zafeiriou, and Maja Pantic. 300 faces in-the-wild challenge: The first facial landmark localization challenge. In Proceedings of the IEEE international conference on computer vision workshops, pages 397–403, 2013. 3
  34. 34.Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. Advances in Neural Information Processing Systems, 33:7462–7473, 2020. 6
  35. 35.Janujah Sivanandan, Eugene Liscio, and P Eng. Assessing structured light 3d scanning using artec eva for injury documentation during autopsy. J Assoc Crime Scene Reconstr, 21:5–14, 2017. 3
  36. 36.Olga Sorkine and Marc Alexa. As-Rigid-As-Possible Surface Modeling. In Alexander Belyaev and Michael Garland, editors, Geometry Processing. The Eurographics Association, 2007. 4
  37. 37.Luan Tran, Feng Liu, and Xiaoming Liu. Towards high-fidelity nonlinear 3d face morphable model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1126–1135, 2019. 2
  38. 38.Luan Tran and Xiaoming Liu. Nonlinear 3d face morphable model. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7346–7355, 2018. 2
  39. 39.Luan Tran and Xiaoming Liu. On learning 3d face morphable model from in-the-wild images. IEEE transactions on pattern analysis and machine intelligence, 43(1):157–171, 2019. 2
  40. 40.Shinji Umeyama. Least-squares estimation of transformation parameters between two point patterns. IEEE Transactions on Pattern Analysis & Machine Intelligence, 13(04):376–380, 1991. 4
  41. 41.Daoye Wang, Prashanth Chandran, Gaspard Zoss, Derek Bradley, and Paulo Gotardo. Morf: Morphable radiance fields for multiview neural head modeling. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–9, 2022. 2
  42. 42.Lizhen Wang, Zhiyua Chen, Tao Yu, Chenguang Ma, Liang Li, and Yebin Liu. Faceverse: a fine-grained and detail-controllable 3d face morphable model from a hybrid dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR2022), June 2022. 2, 3
  43. 43.Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao. Facescape: A large-scale high quality 3d face dataset and detailed riggable 3d face prediction. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2, 3, 6
  44. 44.Tarun Yenamandra, Ayush Tewari, Florian Bernard, Hans-Peter Seidel, Mohamed Elgharib, Daniel Cremers, and Christian Theobalt. i3dmm: Deep implicit 3d morphable model of human heads. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12803–12813, 2021. 2
  45. 45.Mingwu Zheng, Hongyu Yang, Di Huang, and Liming Chen. Imface: A nonlinear 3d morphable face model with implicit neural representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 2, 6, 7, 8
  46. 46.Yufeng Zheng, Victoria Fernandez Abrevaya, Xu Chen, Marcel C. Bühler, Michael J. Black, and Otmar Hilliges. I M avatar: Implicit morphable head avatars from videos. CoRR, abs/2112.07471, 2021. 2
  47. 47.Yinglin Zheng, Hao Yang, Ting Zhang, Jianmin Bao, Dongdong Chen, Yangyu Huang, Lu Yuan, Dong Chen, Ming Zeng, and Fang Wen. General facial representation learning in a visual-linguistic manner. arXiv preprint arXiv:2112.03109, 2021. 4
  48. 48.Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yandong Guo, and Yebin Liu. Structured local radiance fields for human avatar modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022. 2, 5
  49. 49.Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt. State of the art on monocular 3d face reconstruction, tracking, and applications. In Computer graphics forum, volume 37, pages 523–550. Wiley Online Library, 2018. 1

Citation

MLA
Giebenhain, S., et al. “Learning Neural Parametric Head Models”. arXiv, 2022, http://arxiv.org/abs/2212.02761v2.
APA
Giebenhain, S., Kirschstein, T., Georgopoulos, M., Rünz, M., Agapito, L., & Nießner, M. (2022). Learning Neural Parametric Head Models. arXiv. http://arxiv.org/abs/2212.02761v2
Chicago
Giebenhain, S., T. Kirschstein, M. Georgopoulos, M. Rünz, L. Agapito, and M. Nießner. 2022. “Learning Neural Parametric Head Models”. arXiv. http://arxiv.org/abs/2212.02761v2.
Harvard
Giebenhain, S. et al. (2022) “Learning Neural Parametric Head Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.02761v2.
Vancouver
1. Giebenhain S, Kirschstein T, Georgopoulos M, Rünz M, Agapito L, Nießner M (2022) Learning Neural Parametric Head Models. arXiv

BibTeX

@article{giebenhain2022learning,
  title = {Learning Neural Parametric Head Models},
  author = {Giebenhain, Simon and Kirschstein, Tobias and Georgopoulos, Markos and Rünz, Martin and Agapito, Lourdes and Nießner, Matthias},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.02761v2},
  eprint = {2212.02761}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/