ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations

Mingwu ZhengHongyu YangDi HuangLiming Chen

article2022CVPR85 citations

Proposes a nonlinear 3D morphable face model based on implicit neural representations that explicitly disentangles identity and expression deformation fields, enabling high-fidelity face reconstruction and synthesis directly from non-watertight surfaces.

Listen

Precise digital 3D face modeling is critical for applications in computer vision, biometric security, graphics, and medical imaging. However, standard 3D morphable face models struggle to capture fine details and exaggerated facial expressions because they rely on linear statistical assumptions and discrete representations like point clouds or meshes. While continuous implicit neural representations offer a continuous alternative, they typically require closed, watertight 3D geometries—a condition rarely met by standard facial scans—and struggle to isolate complex facial expressions from individual identities.

The article aims to design, evaluate, and demonstrate ImFace, a nonlinear 3D morphable face model based on implicit neural representations. ImFace creates a continuous shape space that separates identity variations from facial expressions while operating directly on open, non-watertight facial surfaces.

To achieve this, the authors implemented two distinct deformation fields to decouple identity and expression variations relative to a neutral template face. They introduced an adaptive blending framework, known as a Neural Blend-Field, which decomposes the face into five semantic regions and merges local implicit functions to capture subtle surface details. To bypass the watertight data requirement, the authors developed a preprocessing pipeline that cleans internal cavities and applies triangulation to construct pseudo-watertight surfaces. They trained and tested the framework using high-resolution scans from the FaceScape dataset, utilizing 5,323 training scans from 355 individuals and evaluating accuracy on 200 unseen scans from 10 individuals.

The experimental findings show that ImFace substantially improves geometric fidelity over current state-of-the-art models. Quantitatively, ImFace achieved a Chamfer reconstruction error of 0.625 mm, representing a reduction of roughly 33% to 62% in geometric error compared to established alternatives (0.929 mm for FaceScape, 0.971 mm for FLAME, and 1.635 mm for i3DMM). Under a strict geometric accuracy metric (F-score at 0.001), ImFace reached 91.11%, outperforming alternative models that scored between 42.26% and 67.09%. Furthermore, it accomplished this with a compact 256-dimensional embedding, establishing reliable point-to-point correspondences across diverse subjects and expressions without requiring manual registration or dense expression labels during training.

These results demonstrate that implicit neural representations can capture fine, non-rigid anatomical deformations—such as frowns and pouts—more efficiently than traditional discrete models. By decoupling identity from expression and eliminating the need for watertight meshes, ImFace reduces data preparation overhead and improves reconstruction quality. This offers strong performance advantages for automated 3D avatar generation, biometric matching, and digital clinical planning.

Organizations developing 3D facial technology should evaluate continuous implicit representations as an alternative to linear mesh pipelines. For deployment in production or consumer applications, teams should develop ethical safeguards and access controls to mitigate privacy invasion and identity spoofing risks associated with high-fidelity digital replicas. Future technical development should focus on integrating realistic surface appearance, such as diffuse and specular reflectance, to match the geometric fidelity.

Confidence in the geometric modeling performance is high based on quantitative benchmarks. However, readers should note that the current model evaluates only facial geometry without photographic surface textures, and minor alignment inaccuracies can still occur around dynamic areas like mouth corners during extreme motion.

arXiv: 2203.14510
Cover for ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations

Abstract

Precise representations of 3D faces are beneficial to various computer vision and graphics applications. Due to the data discretization and model linearity, however, it remains challenging to capture accurate identity and expression clues in current studies. This paper presents a novel 3D morphable face model, namely ImFace, to learn a nonlinear and continuous space with implicit neural representations. It builds two explicitly disentangled deformation fields to model complex shapes associated with identities and expressions, respectively, and designs an improved learning strategy to extend embeddings of expressions to allow more diverse changes. We further introduce a Neural Blend-Field to learn sophisticated details by adaptively blending a series of local fields. In addition to ImFace, an effective preprocessing pipeline is proposed to address the issue of watertight input requirement in implicit representations, enabling them to work with common facial surfaces for the first time. Extensive experiments are performed to demonstrate the superiority of ImFace.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Disentangled INRs Network
  • 3.2. Neural Blend-Field
  • 3.3. Improved Expression Embedding Learning
  • 3.4. Loss Functions
  • 3.5. Data Preprocessing
  • 4. Experiments
  • 4.1. Reconstruction
  • 4.2. Correspondence
  • 4.3. Ablation Study
  • 5. Discussion
  • 6. Conclusion
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — ImFace nonlinear implicit 3D morphable model

    model/method

    ImFace is a nonlinear 3D morphable face model represented by a continuous signed-distance function rather than a mesh, voxel grid, or point cloud. A query point is evaluated together with an expression embedding and an identity embedding. The model factorizes facial variation into an expression deformation field, an identity deformation field, and a shared template signed-distance field. The resulting continuous representation supports arbitrary-resolution surface extraction, fine-grained nonlinear deformations, and learned dense correspondence without requiring pre-registered training meshes.

  2. Knowl 2 — Disentangled deformation-field architecture

    equation

    For a query point p∈R3p\in\mathbb{R}^{3}, expression code zexp∈Rdexpz_{\mathrm{exp}}\in\mathbb{R}^{d_{\mathrm{exp}}}, and identity code zid∈Rdidz_{\mathrm{id}}\in\mathbb{R}^{d_{\mathrm{id}}}, ImFace first maps the observed face to a person-specific neutral canonical space and then maps that canonical space to a shared template space. With KK facial landmarks, the network computes

    l=η(zexp,zid)∈RK×3,p′=E(p,zexp,l)∈R3,l′=η′(zid)∈RK×3,(p′′,δ)=I(p′,zid,l′)∈R3×R,s0=T(p′′,l′′)∈R,f(p,zexp,zid)=s0+δ.\begin{aligned} l &= \eta(z_{\mathrm{exp}},z_{\mathrm{id}})\in\mathbb{R}^{K\times 3}, \\ p' &= E(p,z_{\mathrm{exp}},l)\in\mathbb{R}^{3}, \\ l' &= \eta'(z_{\mathrm{id}})\in\mathbb{R}^{K\times 3}, \\ (p'',\delta) &= I(p',z_{\mathrm{id}},l')\in\mathbb{R}^{3}\times\mathbb{R}, \\ s_{0} &= T(p'',l'')\in\mathbb{R}, \\ f(p,z_{\mathrm{exp}},z_{\mathrm{id}}) &= s_{0}+\delta . \end{aligned}

    Here EE is the expression deformation field, II is the identity deformation field, TT is the template signed-distance field, η\eta and η′\eta' generate landmarks for the observed and canonical faces, respectively, and l′′∈RK×3l''\in\mathbb{R}^{K\times 3} contains landmarks on the shared template obtained by averaging the training faces. The intermediate point p′p' represents the corresponding point on a neutral face of the same person, while p′′p'' represents the point in the shared template space. The residual δ\delta corrects the template SDF when preprocessing produces imperfect point correspondences; the final signed distance is f=s0+δf=s_{0}+\delta.

  3. Knowl 3 — Neural Blend-Field for local facial detail

    model/method

    Each of the expression, identity, and template subnetworks uses a Neural Blend-Field instead of a single global implicit function. For an input coordinate x∈R3x\in\mathbb{R}^{3} and field output qq, the blended field is

    q=ψ(x)=∑n=1Kwn(x) ψn(x−ln),q=\psi(x)=\sum_{n=1}^{K}w_{n}(x)\,\psi_{n}(x-l_{n}),

    where ln∈R3l_{n}\in\mathbb{R}^{3} is the center of local region nn, ψn\psi_{n} is its local implicit field, and the weights satisfy wn(x)≥0w_{n}(x)\geq 0 and ∑n=1Kwn(x)=1\sum_{n=1}^{K}w_{n}(x)=1. The implementation uses K=5K=5 regions centered at the outer eye corners, mouth corners, and nose tip. Each local field is a small MLP with sinusoidal activations and sinusoidal positional encoding of x−lnx-l_{n}, while a three-layer Fusion Network takes the absolute coordinate xx and predicts the blend weights through a softmax. This lets the model represent local high-frequency geometry and deformations while using fewer parameters than an equivalently sized global network.

    For deformation fields, the local output is an SE(3)SE(3) transformation represented by a rotation vector ω∈R3\omega\in\mathbb{R}^{3} and translation parameter u∈R3u\in\mathbb{R}^{3}. The transformed coordinate is x′=Rx+tx'=Rx+t, with θ=∥ω∥\theta=\lVert\omega\rVert, skew-symmetric matrix [ω]×[\omega]_{\times}, and

    R=I3+sin⁡θθ[ω]×+1−cos⁡θθ2[ω]×2,R=I_{3}+\frac{\sin\theta}{\theta}[\omega]_{\times}+\frac{1-\cos\theta}{\theta^{2}}[\omega]_{\times}^{2}, t=(I3+1−cos⁡θθ2[ω]×+θ−sin⁡θθ3[ω]×2)u, t=\left(I_{3}+\frac{1-\cos\theta}{\theta^{2}}[\omega]_{\times}+\frac{\theta-\sin\theta}{\theta^{3}}[\omega]_{\times}^{2}\right)u,

    where I3I_{3} is the 3×33\times3 identity matrix. Hyper Networks take the latent expression or identity code and generate instance-specific parameters for the local MLPs in EE and II, increasing the variety represented by the latent spaces. The SE(3)SE(3) formulation is used because it can represent mandibular rotations and is more robust to pose perturbations than a pure translation field.

  4. Knowl 4 — Per-scan expression embedding extension

    model/method

    Instead of assigning one expression embedding to every expression category, ImFace assigns a distinct expression embedding to each non-neutral training scan. This enlarges the expression latent space so that the expression deformation field can encode individual-specific and fine-grained variations rather than only an average category shape.

    To prevent identity information from being absorbed into the enlarged expression space, the expression deformation is suppressed for neutral scans. For every point pnup_{\mathrm{nu}} on a neutral face, the training constraint is

    E(pnu,zexp,l)=pnu.E(p_{\mathrm{nu}},z_{\mathrm{exp}},l)=p_{\mathrm{nu}}.

    Consequently, the identity field and template field learn the shape variation present in neutral faces, while the expression field is trained to model non-neutral deformation. The strategy requires only neutral-versus-non-neutral labels and does not require dense expression annotations.

  5. Knowl 5 — Pseudo-watertight preprocessing for open facial surfaces

    algorithm

    ImFace converts ordinary non-watertight facial meshes into pseudo-watertight surfaces so that a differentiable signed-distance function can be learned.

    Input: a facial mesh with facial landmarks.

    Output: sampled training triples (p,nˉ,sˉ)(p,\bar n,\bar s) consisting of a query point, its signed-distance gradient, and its signed-distance value.

    Input facial mesh and facial landmarks
    Rigidly align the face frontally using the landmarks
    Normalize the face using a 10 cm scale
    Place the coordinate origin 4 cm behind the nose tip
    Define a sphere of radius 10 cm centered at the origin
    Crop mesh triangles outside the sphere
    Use ray-triangle intersections to remove hidden surfaces, including nasal and oral cavities
    Project the remaining mesh to the x-y plane
    Apply Delaunay triangulation to obtain an oriented pseudo-watertight mesh
    Compute distances from sampled points to the nearest surface
    Assign negative sign to points behind the facial surface along the positive z-direction convention
    Uniformly sample 250,000 points on the facial surface and 15,000 points in the sphere
    Compute the signed-distance gradient at every sampled point
    Return the triples (query point, gradient, signed distance)

    The sign is determined from the angle between the vector to the nearest surface and the positive zz direction, with coordinates behind the facial surface assigned negative values. The procedure avoids the discontinuous gradient at the boundary of an unsigned-distance function while allowing implicit networks to process common open facial scans.

  6. Knowl 6 — Training and latent-code fitting objectives

    equation

    For training face ii, let Ωi\Omega_i be its sampled query points, sˉ\bar s the ground-truth signed distance, nˉ\bar n the ground-truth SDF gradient, ff the complete ImFace SDF, and λ1,…,λ7\lambda_1,\ldots,\lambda_7 nonnegative loss weights. The model uses SDF reconstruction, gradient alignment, Eikonal regularization, latent-code regularization, landmark supervision, landmark consistency, and residual suppression:

    Lsdfi=λ1∑p∈Ωi∣f(p)−sˉ∣+λ2∑p∈Ωi(1−⟨∇f(p),nˉ⟩),\mathcal{L}^{i}_{\mathrm{sdf}}=\lambda_{1}\sum_{p\in\Omega_i}|f(p)-\bar s|+\lambda_{2}\sum_{p\in\Omega_i}\left(1-\langle\nabla f(p),\bar n\rangle\right), Leiki=λ3∑p∈Ωi(∣∥∇f(p)∥2−1∣+∣∥∇T(I(E(p,zexp)))(p)∥2−1∣),\mathcal{L}^{i}_{\mathrm{eik}}=\lambda_{3}\sum_{p\in\Omega_i}\left(|\lVert\nabla f(p)\rVert_{2}-1|+|\lVert\nabla T(I(E(p,z_{\mathrm{exp}})))(p)\rVert_{2}-1|\right), Lembi=λ4(∥zexp∥22+∥zid∥22),\mathcal{L}^{i}_{\mathrm{emb}}=\lambda_{4}\left(\lVert z_{\mathrm{exp}}\rVert_{2}^{2}+\lVert z_{\mathrm{id}}\rVert_{2}^{2}\right), Llmkgi=λ5∑n=1K(∣ln−lˉni∣+∣ln′−lˉn′∣),\mathcal{L}^{i}_{\mathrm{lmkg}}=\lambda_{5}\sum_{n=1}^{K}\left(|l_{n}-\bar l^{i}_{n}|+|l'_{n}-\bar l'_{n}|\right), Llmkci=λ6∑n=1K(∣E(ln)−lˉn′∣+∣I(E(ln))−ln′′∣),\mathcal{L}^{i}_{\mathrm{lmkc}}=\lambda_{6}\sum_{n=1}^{K}\left(|E(l_{n})-\bar l'_{n}|+|I(E(l_{n}))-l''_{n}|\right), Lresi=λ7∑p∈Ωi∣δ(p)∣,\mathcal{L}^{i}_{\mathrm{res}}=\lambda_{7}\sum_{p\in\Omega_i}|\delta(p)|, L=∑i(Lsdfi+Leiki+Lembi+Llmkgi+Llmkci+Lresi).\mathcal{L}=\sum_i\left(\mathcal{L}^{i}_{\mathrm{sdf}}+\mathcal{L}^{i}_{\mathrm{eik}}+\mathcal{L}^{i}_{\mathrm{emb}}+\mathcal{L}^{i}_{\mathrm{lmkg}}+\mathcal{L}^{i}_{\mathrm{lmkc}}+\mathcal{L}^{i}_{\mathrm{res}}\right).

    Here lnl_n and ln′l'_n are predicted observed-face and canonical-face landmarks, lˉni\bar l^{i}_n and lˉn′\bar l'_n are their ground-truth positions, and ln′′l''_n is the corresponding template landmark. The Eikonal term constrains spatial SDF gradients in both observation and canonical/template spaces, while the residual penalty prevents δ\delta from storing excessive template geometry. At test time, the network parameters are fixed and each face is reconstructed by optimizing only its expression and identity codes with Lsdf+Leik+Lemb\mathcal{L}_{\mathrm{sdf}}+\mathcal{L}_{\mathrm{eik}}+\mathcal{L}_{\mathrm{emb}}.

  7. Knowl 7 — Training dataset and implementation configuration

    experimental setup

    The experiments use FaceScape, which contains 938 individuals and 20 expression types; the publicly available portion contains 365 individuals. ImFace is trained on 5,323 scans from 355 people covering 15 expressions and tested on 200 scans from 10 different people covering 20 expressions.

    Each local Mini-Net is an MLP with three hidden layers, 32 hidden features per layer, and sine activations. Each Hyper Net is a three-layer ReLU MLP with hidden dimensionality 64. The two Landmark-Nets use three fully connected layers of width 128. The complete model is trained end-to-end with Adam for 1,500 epochs, starting at learning rate 0.00010.0001; after epoch 200, the learning rate is multiplied by 0.950.95 every 10 epochs. Training uses minibatches of 72 on four NVIDIA RTX 3090 GPUs and takes approximately two days. Optimizing embeddings for 200 test scans takes approximately four hours on one GPU.

  8. Knowl 8 — Reconstruction accuracy against existing 3D face models

    data/table

    ImFace is evaluated by fitting complete test scans and is compared with i3DMM, FLAME, and FaceScape. All methods use the same preprocessed facial region for metric computation. The reported dimensionality is the total latent dimensionality; ImFace and i3DMM use 128-dimensional identity and expression codes, FLAME uses 300 identity plus 100 expression parameters, and FaceScape uses 300 identity plus 52 expression parameters. i3DMM is retrained on the ImFace training set, while the FaceScape test scans are included in the released FaceScape training data.

    Model Dim. Chamfer (mm) ↓\downarrow [email protected] ↑\uparrow
    i3DMM 256 1.635 42.26
    FLAME 400 0.971 64.73
    FaceScape 352 0.929 67.09
    ImFace 256 0.625 91.11

    ImFace obtains the lowest symmetric Chamfer distance and the highest strict-threshold F-score. Qualitatively, it reconstructs identity more accurately than the alternatives, preserves subtle nonlinear muscle deformations such as frowns and pouts, and handles expressions not seen during learning with fewer latent parameters. i3DMM produces artifacts on complicated deformations, FLAME tends toward stiff expressions, and FaceScape remains less precise for expression morphs despite its high-quality scans.

  9. Knowl 9 — Automatically learned cross-expression and cross-identity correspondence

    empirical result

    ImFace provides dense correspondence by fitting the expression and identity codes of two faces, deforming densely sampled points from each face into the shared template space, and matching points there by nearest-neighbor search. This procedure does not require an externally supplied point-to-point registration for the input faces.

    Visual correspondence tests show that a source face can be transferred consistently across multiple expressions and identities, with facial color patterns remaining aligned over the corresponding regions. Small internal texture dispersions occasionally appear around mouth corners because those regions undergo especially large expression-dependent shape changes; nevertheless, the learned correspondences are generally coherent across both identity and expression variation.

  10. Knowl 10 — Ablation evidence for the three core design choices

    data/table

    The ablations remove disentangled deformation fields, the Neural Blend-Field, or the extended expression-embedding strategy while keeping the remaining system fixed. Without disentanglement, identity and expression codes are concatenated and supplied to one universal deformation field. Without blending, the local fields are replaced by vanilla MLPs with the same parameter count. Without extension, the number of expression embeddings is restricted to the number of expression categories.

    Variant Chamfer (mm) ↓\downarrow [email protected] ↑\uparrow
    ImFace without disentanglement 0.772 82.70
    ImFace without Neural Blend-Field 0.767 82.37
    ImFace without expression extension 0.705 86.98
    Full ImFace 0.625 91.11

    The single-field variant produces chaotic reconstructions, especially for large expressions. Replacing the Neural Blend-Field causes visible blurring and weakens high-frequency detail. Restricting expression codes to category-level embeddings yields averaged expressions and can prevent exaggerated expressions such as strong mouth stretching from converging to plausible shapes. The full model is best under both metrics.

Coverage note — The stated limitation that ImFace models geometry but not realistic diffuse or specular facial texture, along with the paper's societal-impact discussion, was omitted because it is not a contributed method, result, or experimentally validated component.

References

  1. 1.Victoria Fernández Abrevaya, Stefanie Wuhrer, and Edmond Boyer. Spatiotemporal modeling for efficient registration of dynamic 3d faces. In 3DV, 2018.
  2. 2.Oswald Aldrian and William AP Smith. Inverse rendering of faces with a 3d morphable model. IEEE TPAMI, 35(5):1080–1093, 2012.
  3. 3.Thiemo Alldieck, Hongyi Xu, and Cristian Sminchisescu. Imghum: Implicit generative models of 3d human shape and articulated pose. In ICCV, 2021.
  4. 4.Brian Amberg, Sami Romdhani, and Thomas Vetter. Optimal step nonrigid icp algorithms for surface registration. In CVPR, 2007.
  5. 5.Timur Bagautdinov, Chenglei Wu, Jason Saragih, Pascal Fua, and Yaser Sheikh. Modeling facial geometry using compositional vaes. In CVPR, 2018.
  6. 6.Mehdi Bahri, Eimear O’Sullivan, Shunwang Gong, Feng Liu, Xiaoming Liu, Michael M Bronstein, and Stefanos Zafeiriou. Shape my face: registering 3d face scans by surface-to-surface translation. IJCV, 129(9):2680–2713, 2021.
  7. 7.Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In SIGGRAPH, 1999.
  8. 8.Volker Blanz and Thomas Vetter. Face recognition based on fitting a 3d morphable model. IEEE TPAMI, 25(9):1063–1074, 2003.
  9. 9.Timo Bolkart and Stefanie Wuhrer. A groupwise multilinear correspondence optimization for 3d faces. In ICCV, 2015.
  10. 10.James Booth, Anastasios Roussos, Allan Ponniah, David Dunaway, and Stefanos Zafeiriou. Large scale 3d morphable models. IJCV, 126(2):233–254, 2018.
  11. 11.Giorgos Bouritsas, Sergiy Bokhnyak, Stylianos Ploumpis, Michael Bronstein, and Stefanos Zafeiriou. Neural 3d morphable models: Spiral convolutional networks for 3d shape representation learning and generation. In ICCV, 2019.
  12. 12.Alan Brunton, Timo Bolkart, and Stefanie Wuhrer. Multilinear wavelets: A statistical shape space for human faces. In ECCV, 2014.
  13. 13.Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou. Facewarehouse: A 3d facial expression database for visual computing. IEEE TVCG, 20(3):413–425, 2013.
  14. 14.Xu Chen, Yufeng Zheng, Michael J. Black, Otmar Hilliges, and Andreas Geiger. Snarf: Differentiable forward skinning for animating non-rigid neural implicit shapes. In ICCV, 2021.
  15. 15.Zhixiang Chen and Tae-Kyun Kim. Learning feature aggregation for deep 3d morphable models. In CVPR, 2021.
  16. 16.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In CVPR, 2019.
  17. 17.Shiyang Cheng, Michael M. Bronstein, Yuxiang Zhou, Irene Kotsia, Maja Pantic, and Stefanos Zafeiriou. Meshgan: Non-linear 3d morphable models of faces. CoRR, abs/1903.10384, 2019.
  18. 18.Julian Chibane, Aymen Mir, and Gerard Pons-Moll. Neural unsigned distance fields for implicit function learning. In NeurIPS, 2020.
  19. 19.Enric Corona, Albert Pumarola, Guillem Alenya, Gerard Pons-Moll, and Francesc Moreno-Noguer. Smplicit: Topology-aware generative model for clothed people. In CVPR, 2021.
  20. 20.Darren Cosker, Eva Krumhuber, and Adrian Hilton. A facs valid 3d dynamic action unit database with applications to 3d dynamic morphable facial modeling. In ICCV, 2011.
  21. 21.Yu Deng, Jiaolong Yang, and Xin Tong. Deformed implicit field: Modeling 3d shapes with learned dense correspondence. In CVPR, 2021.
  22. 22.Bernhard Egger, William AP Smith, Ayush Tewari, Stefanie Wuhrer, Michael Zollhoefer, Thabo Beeler, Florian Bernard, Timo Bolkart, Adam Kortylewski, Sami Romdhani, et al. 3d morphable face models—past, present, and future. ACM TOG, 39(5):1–38, 2020.
  23. 23.Kyle Genova, Forrester Cole, Avneesh Sud, Aaron Sarna, and Thomas Funkhouser. Local deep implicit functions for 3d shape. In CVPR, 2020.
  24. 24.Kyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna, William T Freeman, and Thomas Funkhouser. Learning shape templates with structured implicit functions. In ICCV, 2019.
  25. 25.Syed Zulqarnain Gilani, Ajmal Mian, Faisal Shafait, and Ian Reid. Dense 3d face correspondence. IEEE TPAMI, 40(7):1584–1598, 2017.
  26. 26.Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. Implicit geometric regularization for learning shapes. In ICML. 2020.
  27. 27.Guosheng Hu, Fei Yan, Chi-Ho Chan, Weihong Deng, William Christmas, Josef Kittler, and Neil M Robertson. Face recognition using a unified 3d morphable model. In ECCV, 2016.
  28. 28.Moritz Ibing, Isaak Lim, and Leif Kobbelt. 3d shape generation with grid-based implicit functions. In CVPR, 2021.
  29. 29.Der-Tsai Lee and Bruce J Schachter. Two algorithms for constructing a delaunay triangulation. International Journal of Computer & Information Sciences, 9(3):219–242, 1980.
  30. 30.John P Lewis, Matt Cordner, and Nickson Fong. Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation. In SIGGRAPH, 2000.
  31. 31.Tianye Li, Timo Bolkart, Michael J Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4d scans. ACM TOG, 36(6):194–1, 2017.
  32. 32.Yaron Lipman. Phase transitions, distance functions, and implicit neural representations. In ICML, 2021.
  33. 33.Feng Liu and Xiaoming Liu. Learning implicit functions for topology-varying dense 3d shape correspondence. In NeurIPS, 2020.
  34. 34.Feng Liu, Luan Tran, and Xiaoming Liu. 3d face modeling from diverse raw scan data. In ICCV, 2019.
  35. 35.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In CVPR, 2019.
  36. 36.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  37. 37.Tomas Moller and Ben Trumbore. Fast, minimum storage ray-triangle intersection. Journal of graphics tools, 2(1):21–28, 1997.
  38. 38.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In CVPR, 2019.
  39. 39.Ankur Patel and William AP Smith. 3d morphable face models revisited. In CVPR, 2009.
  40. 40.Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami Romdhani, and Thomas Vetter. A 3d face model for pose and illumination invariant face recognition. In AVSS, 2009.
  41. 41.Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Animatable neural radiance fields for modeling dynamic human bodies. In ICCV, 2021.
  42. 42.Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, and Andreas Geiger. Convolutional occupancy networks. In ECCV, 2020.
  43. 43.Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR, 2021.
  44. 44.Mallikarjun B R, Ayush Tewari, Hans-Peter Seidel, Mohamed Elgharib, and Christian Theobalt. Learning complete 3d morphable face models from images and videos. In CVPR, 2021.
  45. 45.Eduard Ramon, Gil Triginer, Janna Escur, Albert Pumarola, Jaime Garcia, Xavier Giro-i Nieto, and Francesc Moreno-Noguer. H3d-net: Few-shot high-fidelity 3d head reconstruction. In ICCV, 2021.
  46. 46.Anurag Ranjan, Timo Bolkart, Soubhik Sanyal, and Michael J Black. Generating 3d faces using convolutional mesh autoencoders. In ECCV, 2018.
  47. 47.Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In ICCV, 2019.
  48. 48.Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In CVPR, 2020.
  49. 49.Vincent Sitzmann, Julien N.P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. Implicit neural representations with periodic activation functions. In NeurIPS, 2020.
  50. 50.Femke CR Staal, Allan JT Ponniah, Freida Angullia, Clifford Ruff, Maarten J Koudstaal, and David Dunaway. Describing crouzon and pfeiffer syndrome based on principal component analysis. Journal of Cranio-Maxillofacial Surgery, 43(4):528–536, 2015.
  51. 51.Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Neural geometric level of detail: Real-time rendering with implicit 3d shapes. In CVPR, 2021.
  52. 52.Jia-Heng Tang, Weikai Chen, Jie Yang, Bo Wang, Songrun Liu, Bo Yang, and Lin Gao. Octfield: Hierarchical implicit functions for 3d modeling. In NeurIPS, 2021.
  53. 53.Luan Tran, Feng Liu, and Xiaoming Liu. Towards high-fidelity nonlinear 3d face morphable model. In CVPR, 2019.
  54. 54.Luan Tran and Xiaoming Liu. Nonlinear 3d face morphable model. In CVPR, 2018.
  55. 55.Daniel Vlasic, Matthew Brand, Hanspeter Pfister, and Jovan Popovic. Face transfer with multilinear models. In ACM SIGGRAPH Courses. 2006.
  56. 56.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. In NeurIPS, 2021.
  57. 57.Qiangeng Xu, Weiyue Wang, Duygu Ceylan, Radomir Mech, and Ulrich Neumann. Disn: Deep implicit surface network for high-quality single-view 3d reconstruction. In NeurIPS, 2019.
  58. 58.Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao. Facescape: a large-scale high quality 3d face dataset and detailed riggable 3d face prediction. In CVPR, 2020.
  59. 59.Tarun Yenamandra, Ayush Tewari, Florian Bernard, Hans-Peter Seidel, Mohamed Elgharib, Daniel Cremers, and Christian Theobalt. i3dmm: Deep implicit 3d morphable model of human heads. In CVPR, 2021.
  60. 60.Li Yi, Vladimir G Kim, Duygu Ceylan, I-Chao Shen, Mengyan Yan, Hao Su, Cewu Lu, Qixing Huang, Alla Sheffer, and Leonidas Guibas. A scalable active framework for region annotation in 3d shape collections. ACM TOG, 35(6):1–12, 2016.
  61. 61.Jingyang Zhang, Yao Yao, and Long Quan. Learning signed distance field for multi-view surface reconstruction. In ICCV, 2021.
  62. 62.Zerong Zheng, Tao Yu, Qionghai Dai, and Yebin Liu. Deep implicit templates for 3d shape representation. In CVPR, 2021.

Citation

MLA
Zheng, M., et al. “ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations”. arXiv, 2022, http://arxiv.org/abs/2203.14510v2.
APA
Zheng, M., Yang, H., Huang, D., & Chen, L. (2022). ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations. arXiv. http://arxiv.org/abs/2203.14510v2
Chicago
Zheng, M., H. Yang, D. Huang, and L. Chen. 2022. “ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations”. arXiv. http://arxiv.org/abs/2203.14510v2.
Harvard
Zheng, M. et al. (2022) “ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14510v2.
Vancouver
1. Zheng M, Yang H, Huang D, Chen L (2022) ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations. arXiv

BibTeX

@article{zheng2022imface,
  title = {ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations},
  author = {Zheng, Mingwu and Yang, Hongyu and Huang, Di and Chen, Liming},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14510v2},
  eprint = {2203.14510}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE