GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians
Shenhan QianTobias KirschsteinLiam SchoneveldDavide DavoliSimon GiebenhainMatthias Nießner
Proposes a dynamic head avatar representation that rigs 3D Gaussian splats to a FLAME mesh using a binding inheritance strategy, enabling real-time, photorealistic facial animation and novel expression transfer.
Creating realistic and animatable digital head avatars is essential for emerging visual technologies in gaming, film production, virtual telepresence, and augmented reality. However, existing methods struggle to balance fine visual fidelity with flexible control. Traditional approaches based on neural radiance fields or dynamic point representations often fail to animate novel facial expressions accurately, produce noticeable visual artifacts, or require expensive processing that prevents efficient rendering.
To address this challenge, the article introduces GaussianAvatars, a framework designed to reconstruct photorealistic head avatars from multi-view video that are fully controllable across arbitrary camera viewpoints, head poses, and facial expressions. The core approach attaches discrete 3D Gaussian splats directly to the local coordinate frames of triangles in a parametric face mesh. By transforming splats from local to global space during animation and simultaneously refining the underlying face model parameters, the method allows the visual primitives to compensate for geometric inaccuracies while preserving explicit animation control. A binding inheritance mechanism ensures that newly added or pruned splats retain their connection to the mesh, while spatial and scaling regularizations prevent visual distortions during motion.
The evaluation demonstrates that GaussianAvatars significantly improves rendering fidelity and animation transfer over previous leading methods. In novel-view rendering tasks across multiple subjects, the framework achieved a peak signal-to-noise ratio of 31.6 dB and a perceptual error score (LPIPS) of 0.065, outperforming existing baselines by a substantial margin. In self-reenactment and cross-identity driving tests, the method produced visibly sharper details, accurately capturing complex dynamics such as eye blinks, mouth interiors, eye reflections, and facial wrinkles without the jitter or tearing seen in prior techniques. Ablation analyses confirmed that binding inheritance and regularization are indispensable for maintaining structural stability and avoiding severe visual spikes.
These findings indicate that directly rigging 3D Gaussian primitives to parametric geometric models provides a robust, high-performance path for real-time digital avatar production. Organizations operating in virtual production and telepresence can achieve higher visual quality with reduced manual correction. However, adopting photorealistic avatar synthesis presents compliance, legal, and security risks regarding identity theft, unauthorized likeness manipulation, and misleading deepfake media. Stakeholders must pair deployment with strict data governance, identity consent protocols, and detection mechanisms.
For practical implementation, teams exploring digital human pipelines should evaluate rigged Gaussian splatting as a baseline for head animation while planning additional technical development. Current limitations include the inability to relight avatars dynamically, as appearance and lighting are baked into the captured radiance field, as well as unconstrained motion in areas outside the face mesh, such as loose hair and accessories. Further engineering should integrate specialized hair modeling and separate material properties from illumination to achieve production-ready versatility.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). FLAME supplies the articulated parametric face model for identity, head pose, and expression that GaussianAvatars uses to rig and animate its Gaussian primitives.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). 3D Gaussian Splatting establishes the explicit Gaussian representation, adaptive densification, and real-time differentiable rendering that GaussianAvatars adapts to animated heads.
- Paper: RigNeRF: Fully Controllable Neural 3D Portraits, ShahRukh Athar et al. (2022). RigNeRF provides the preceding framework for coupling neural radiance representations with controllable head pose and facial expression from portrait video.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). Learning Neural Parametric Head Models develops disentangled identity and expression geometry that clarifies GaussianAvatars' use of a deformable head prior.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). D-NeRF introduces canonical-space deformation for dynamic view synthesis, a conceptual precursor to transforming appearance primitives under facial motion.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). 4D Gaussian Splatting shows how a canonical Gaussian field can be deformed over time, providing the dynamic-representation context for rigged facial Gaussians.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Nerfies establishes deformation-field reconstruction of nonrigid subjects from video, framing the motion and regularization challenges that GaussianAvatars addresses with mesh-bound splats.
- Paper: Relightable Gaussian Codec Avatars, Shunsuke Saito et al. (2024). Relightable Gaussian Codec Avatars extends rigged Gaussian head avatars with explicit material and illumination modeling, addressing GaussianAvatars' baked-lighting limitation.
- Paper: 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis, Zhicheng Lu et al. (2024). 3D Geometry-aware Deformable Gaussian Splatting generalizes Gaussian motion modeling with local geometric features and continuous rotations for more structurally consistent dynamic rendering.
