ImFace: A Nonlinear 3D Morphable Face Model with Implicit Neural Representations
Mingwu ZhengHongyu YangDi HuangLiming Chen
Proposes a nonlinear 3D morphable face model based on implicit neural representations that explicitly disentangles identity and expression deformation fields, enabling high-fidelity face reconstruction and synthesis directly from non-watertight surfaces.
Precise digital 3D face modeling is critical for applications in computer vision, biometric security, graphics, and medical imaging. However, standard 3D morphable face models struggle to capture fine details and exaggerated facial expressions because they rely on linear statistical assumptions and discrete representations like point clouds or meshes. While continuous implicit neural representations offer a continuous alternative, they typically require closed, watertight 3D geometries—a condition rarely met by standard facial scans—and struggle to isolate complex facial expressions from individual identities.
The article aims to design, evaluate, and demonstrate ImFace, a nonlinear 3D morphable face model based on implicit neural representations. ImFace creates a continuous shape space that separates identity variations from facial expressions while operating directly on open, non-watertight facial surfaces.
To achieve this, the authors implemented two distinct deformation fields to decouple identity and expression variations relative to a neutral template face. They introduced an adaptive blending framework, known as a Neural Blend-Field, which decomposes the face into five semantic regions and merges local implicit functions to capture subtle surface details. To bypass the watertight data requirement, the authors developed a preprocessing pipeline that cleans internal cavities and applies triangulation to construct pseudo-watertight surfaces. They trained and tested the framework using high-resolution scans from the FaceScape dataset, utilizing 5,323 training scans from 355 individuals and evaluating accuracy on 200 unseen scans from 10 individuals.
The experimental findings show that ImFace substantially improves geometric fidelity over current state-of-the-art models. Quantitatively, ImFace achieved a Chamfer reconstruction error of 0.625 mm, representing a reduction of roughly 33% to 62% in geometric error compared to established alternatives (0.929 mm for FaceScape, 0.971 mm for FLAME, and 1.635 mm for i3DMM). Under a strict geometric accuracy metric (F-score at 0.001), ImFace reached 91.11%, outperforming alternative models that scored between 42.26% and 67.09%. Furthermore, it accomplished this with a compact 256-dimensional embedding, establishing reliable point-to-point correspondences across diverse subjects and expressions without requiring manual registration or dense expression labels during training.
These results demonstrate that implicit neural representations can capture fine, non-rigid anatomical deformations—such as frowns and pouts—more efficiently than traditional discrete models. By decoupling identity from expression and eliminating the need for watertight meshes, ImFace reduces data preparation overhead and improves reconstruction quality. This offers strong performance advantages for automated 3D avatar generation, biometric matching, and digital clinical planning.
Organizations developing 3D facial technology should evaluate continuous implicit representations as an alternative to linear mesh pipelines. For deployment in production or consumer applications, teams should develop ethical safeguards and access controls to mitigate privacy invasion and identity spoofing risks associated with high-fidelity digital replicas. Future technical development should focus on integrating realistic surface appearance, such as diffuse and specular reflectance, to match the geometric fidelity.
Confidence in the geometric modeling performance is high based on quantitative benchmarks. However, readers should note that the current model evaluates only facial geometry without photographic surface textures, and minor alignment inaccuracies can still occur around dynamic areas like mouth corners during extreme motion.
- Paper: A Morphable Model For The Synthesis Of 3D Faces, Volker Blanz et al. (1999). This seminal work introduced 3D morphable face models using linear combinations of shape and texture, establishing the foundational paradigm that ImFace seeks to overcome through nonlinear implicit representations.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). FLAME presents the standard parametric mesh formulation that explicitly decouples identity shape and facial expressions, serving as the direct linear baseline and motivation for ImFace's disentangled neural deformation fields.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). DeepSDF established continuous signed distance functions conditioned on latent codes for 3D shape representation, providing the foundational implicit neural architecture utilized by ImFace.
- Paper: Learning Implicit Fields for Generative Shape Modeling, Zhiqin Chen et al. (2018). IM-NET demonstrates generative 3D shape modeling with continuous implicit fields, supplying essential context on learning continuous spatial decoders.
- Paper: Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains, Matthew Tancik et al. (2020). This work analyzes the spectral bias of coordinate-based MLPs and shows how Fourier feature mappings enable implicit neural networks to represent fine, high-frequency spatial details.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Nerfies establishes the methodology of modeling non-rigid scene deformation by mapping observation spaces into a canonical static template via continuous deformation fields.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). NPHM extends implicit parametric face modeling to full-head 3D reconstructions from sparse depth data by using local neural fields anchored to keypoints.
- Paper: FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models, Shivangi Aneja et al. (2024). FaceTalk builds upon the latent expression spaces of neural parametric head models like ImFace to drive realistic 3D volumetric facial animations directly from speech.
- Paper: Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos, Ziqian Bai et al. (2023). This method leverages parametric facial priors and implicit neural fields to synthesize dynamic, personalized volumetric head avatars from in-the-wild monocular videos.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). GaussianAvatars advances animatable neural avatars by attaching controllable 3D Gaussian primitives to parametric face models for real-time expressive rendering.
- Paper: Learning Locally Editable Virtual Humans, Hsuan-I Ho et al. (2023). This paper extends hybrid parametric-implicit neural representations to full virtual human bodies, enabling local region editing and fine-grained geometric control.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). Real3D-Portrait adapts 3D facial modeling techniques into a one-shot audio-driven synthesis pipeline that reconstructs and animates 3D talking portraits from a single image.
