Learning Neural Parametric Head Models
Simon GiebenhainTobias KirschsteinMarkos GeorgopoulosMartin RünzLourdes AgapitoMatthias Nießner
Proposes a neural parametric head model using local signed distance fields and forward deformation fields trained on high-resolution full-head scans to achieve fine geometry reconstruction with disentangled identity and expression controls.
Digital representations of complete human heads are essential for applications across virtual reality, gaming, telepresence, and digital avatars. Traditional 3D morphable face models rely on fixed-topology mesh templates and linear dimensionality reduction methods such as Principal Component Analysis. While these conventional methods provide good regularization against noisy inputs, their rigid structural assumptions prevent them from capturing fine local facial details, natural variations in head shape, and diverse hair styles.
The article develops and demonstrates a novel neural parametric head model (NPHM) based on hybrid neural fields. The primary objective is to accurately reconstruct full 3D head geometry and complex facial expressions from sparse, noisy depth data by learning separate, disentangled representations for personal identity and facial movements.
To build and validate this system, the authors captured a high-resolution 3D head dataset comprising over 3,700 scans across 203 distinct individuals performing 23 facial expressions. Each scan averaged roughly 1.5 million vertices and 3.5 million triangles. The approach uses an implicit neural field to represent the canonical geometry of an identity as a signed distance field, broken down into an ensemble of smaller local networks anchored to specific facial keypoints with shared weights across symmetrical facial regions. A separate forward deformation network then models facial expressions. When deployed at inference time, the model fits to sparse point clouds—such as single depth frames from consumer-grade sensors—by optimizing the underlying identity and expression codes.
The evaluation produced several clear findings. First, the proposed framework significantly outperformed traditional mesh-based models such as FLAME and Basel Face Models, achieving lower reconstruction error (an L1-Chamfer distance of 0.00182 versus 0.00640 for FLAME during identity fitting) and higher geometric accuracy (F-Score of 0.954 versus 0.530). Second, it outperformed competing neural implicit baselines like ImFace and Neural Parametric Models across both identity and expression tasks. Third, ablation analyses confirmed that decomposing the face into localized regions and sharing symmetric weights directly improves geometric accuracy and surface normal consistency.
These results demonstrate that moving from fixed mesh templates to localized neural fields eliminates key structural bottlenecks in digital human modeling. For engineering and product teams, this approach enables high-fidelity 3D avatar generation and tracking directly from consumer depth sensors without requiring dense multi-camera capture rigs or manual clean-up. Additionally, the forward deformation design allows faster animation compared to backward deformation architectures.
Organizations developing 3D facial capture and avatar pipelines should consider adopting local implicit neural representations over legacy mesh templates for tracking and geometry reconstruction tasks. Future development should focus on integrating photometric texture pipelines, as the current model focuses exclusively on 3D geometry rather than visual color appearance. Furthermore, while the model captures full head shapes with short hair, extending the dataset to accommodate long, loose hair styles remains necessary to broaden real-world applicability.
- Paper: A Morphable Model For The Synthesis Of 3D Faces, Volker Blanz et al. (1999). It introduces the fundamental 3D morphable model framework for linearly disentangling identity and facial variations from 3D scans.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). It establishes the FLAME parametric head model that decouples identity shape, pose, and facial expressions, serving as a primary baseline and conceptual precursor.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). It provides the foundational deep learning formulation for representing continuous 3D shapes via auto-decoded signed distance functions (SDFs).
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). It introduces non-rigid continuous neural deformation fields to map dynamic observations to a canonical template space.
- Paper: NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction, Peng Wang et al. (2021). It develops the neural implicit surface reconstruction framework using signed distance fields with volume rendering.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). It pioneers pixel-aligned implicit functions for continuous 3D human digitizations with fine localized geometric details.
- Paper: Learning Implicit Fields for Generative Shape Modeling, Zhiqin Chen et al. (2018). It establishes generative implicit field decoding (IM-NET) for continuous 3D shape modeling without topological restrictions.
- Paper: Face2Face: Real-Time Face Capture and Reenactment of RGB Videos, Justus Thies et al. (2016). It presents foundational methods for tracking and transferring facial expressions by optimizing parametric facial models.
- Paper: LRM: Large Reconstruction Model for Single Image to 3D, Yicong Hong et al. (2024). It scales up feedforward neural 3D reconstruction from single images to general objects using large-scale transformer architectures and triplane neural fields.
