RigNeRF: Fully Controllable Neural 3D Portraits
ShahRukh AtharZexiang XuKalyan SunkavalliEli ShechtmanZhixin Shu
Combines 3D morphable face models with neural radiance fields to enable precise control over rigid head poses and non-rigid facial expressions using only a short casual smartphone video.
Photo-realistic editing of human portraits with full control over head orientation, facial expressions, and camera viewpoints is a critical capability for immersive media, virtual reality, and digital video production. Standard neural radiance fields synthesize high-quality 3D views of static environments but cannot edit dynamic objects. Meanwhile, existing dynamic face-modeling methods struggle to render novel camera angles, distort full 3D backgrounds, or lose natural head rigidity during reanimation.
The article demonstrates RigNeRF, a volumetric neural rendering system that enables simultaneous, explicit control over head pose, facial expression, and novel camera viewpoints for an entire portrait scene. The primary objective is to learn these capabilities directly from a single, short video captured on a standard consumer device without requiring specialized multi-camera rigs.
To achieve this, the approach combines deformable neural radiance fields with a 3D morphable face model (3DMM). The 3DMM provides a structural prior that defines coarse rigid and non-rigid head movements, while a neural network predicts residual offsets to capture fine details omitted by standard face meshes, such as hair, teeth, and glasses. The model was evaluated using smartphone video captures of multiple subjects enacting varied expressions and head motions across 40- to 70-second sequences (approximately 1,200 to 2,100 frames).
The evaluations show that RigNeRF consistently outperforms leading dynamic rendering baselines—including HyperNeRF, NerFACE, and First Order Motion Model—across visual quality, perceptual distance, and facial reconstruction accuracy. Specifically, RigNeRF achieved higher peak signal-to-noise ratios (ranging from 27.0 to 29.55 dB) and significantly lower face reconstruction errors across all test subjects. Naively applying expression and pose conditioning to dynamic fields caused unnatural warping and visual distortions, whereas RigNeRF’s deformation prior maintained structural rigidity while generalizing to unobserved poses and expressions.
These findings indicate that 3D geometric priors provide an effective inductive bias for neural volumetric rendering, lowering production barriers and capture costs for high-fidelity avatar generation. By enabling high-accuracy reanimation and free-viewpoint rendering from casual smartphone video, the system demonstrates strong potential for low-cost, high-performance content creation pipelines in entertainment and telepresence.
Organizations evaluating this technology should focus on piloting the system in controlled video production or avatar animation workflows to validate performance across broader operational use cases. Because photorealistic face reanimation introduces risks related to synthetic media misuse, deployment should incorporate authentication protocols, watermarking, or other media compliance measures.
Key limitations include that the current implementation is subject-specific, requiring an individual network to be trained for each new subject and scene. Rendering fidelity also depends heavily on precise camera pose tracking during the capture phase. Confidence in the evaluated metrics is strong for standard monocular captures, though users should exercise caution when extrapolating results to scenes with extreme lighting changes or rapid, uncalibrated camera movements.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Introduces Neural Radiance Fields (NeRF), establishing the foundational coordinate-based volumetric rendering framework that RigNeRF builds upon.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Introduces deformable neural radiance fields for monocular video using canonical space deformation fields, which directly informs RigNeRF's dynamic deformation modeling.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). Provides a core methodology for mapping dynamic scenes to a canonical NeRF via time-conditioned deformation networks.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). Presents the statistical 3D morphable head model (FLAME) that underpins the 3D parametric geometric priors used by RigNeRF to guide controllable deformations.
- Paper: Face2Face: Real-Time Face Capture and Reenactment of RGB Videos, Justus Thies et al. (2016). Demonstrates parametric 3D face model tracking and real-time expression transfer from standard RGB video, providing key context for RGB-driven reenactment.
- Paper: First Order Motion Model for Image Animation, Aliaksandr Siarohin et al. (2019). Serves as an essential 2D image animation baseline compared directly against RigNeRF in dynamic portrait reanimation.
- Paper: Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos, Ziqian Bai et al. (2023). Extends 3DMM-guided neural volumetric head avatars by predicting localized 3DMM-anchored feature fields to better synthesize out-of-distribution expressions from monocular video.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). Advances controllable head avatar synthesis by binding 3D Gaussian splatting primitives directly to parametric face meshes for faster, highly detailed rendering.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). Generalizes 3D portrait reanimation to a one-shot setup driven by audio or video without requiring subject-specific NeRF training.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). Builds upon implicit neural representations for heads by learning a complete neural parametric head model that overcomes the expressive limitations of classical linear 3DMMs.
- Paper: Relightable Gaussian Codec Avatars, Shunsuke Saito et al. (2024). Expands animatable mesh-anchored neural avatars to support real-time relighting and sub-millimeter rendering under complex illumination.
- Paper: Learning Locally Editable Virtual Humans, Hsuan-I Ho et al. (2023). Applies mesh-anchored local neural fields beyond portrait heads to achieve controllable and editable full-body digital human avatars.
