GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians

Shenhan QianTobias KirschsteinLiam SchoneveldDavide DavoliSimon GiebenhainMatthias Nießner

article2024CVPR358 citationsHighlight

Proposes a dynamic head avatar representation that rigs 3D Gaussian splats to a FLAME mesh using a binding inheritance strategy, enabling real-time, photorealistic facial animation and novel expression transfer.

Listen

Creating realistic and animatable digital head avatars is essential for emerging visual technologies in gaming, film production, virtual telepresence, and augmented reality. However, existing methods struggle to balance fine visual fidelity with flexible control. Traditional approaches based on neural radiance fields or dynamic point representations often fail to animate novel facial expressions accurately, produce noticeable visual artifacts, or require expensive processing that prevents efficient rendering.

To address this challenge, the article introduces GaussianAvatars, a framework designed to reconstruct photorealistic head avatars from multi-view video that are fully controllable across arbitrary camera viewpoints, head poses, and facial expressions. The core approach attaches discrete 3D Gaussian splats directly to the local coordinate frames of triangles in a parametric face mesh. By transforming splats from local to global space during animation and simultaneously refining the underlying face model parameters, the method allows the visual primitives to compensate for geometric inaccuracies while preserving explicit animation control. A binding inheritance mechanism ensures that newly added or pruned splats retain their connection to the mesh, while spatial and scaling regularizations prevent visual distortions during motion.

The evaluation demonstrates that GaussianAvatars significantly improves rendering fidelity and animation transfer over previous leading methods. In novel-view rendering tasks across multiple subjects, the framework achieved a peak signal-to-noise ratio of 31.6 dB and a perceptual error score (LPIPS) of 0.065, outperforming existing baselines by a substantial margin. In self-reenactment and cross-identity driving tests, the method produced visibly sharper details, accurately capturing complex dynamics such as eye blinks, mouth interiors, eye reflections, and facial wrinkles without the jitter or tearing seen in prior techniques. Ablation analyses confirmed that binding inheritance and regularization are indispensable for maintaining structural stability and avoiding severe visual spikes.

These findings indicate that directly rigging 3D Gaussian primitives to parametric geometric models provides a robust, high-performance path for real-time digital avatar production. Organizations operating in virtual production and telepresence can achieve higher visual quality with reduced manual correction. However, adopting photorealistic avatar synthesis presents compliance, legal, and security risks regarding identity theft, unauthorized likeness manipulation, and misleading deepfake media. Stakeholders must pair deployment with strict data governance, identity consent protocols, and detection mechanisms.

For practical implementation, teams exploring digital human pipelines should evaluate rigged Gaussian splatting as a baseline for head animation while planning additional technical development. Current limitations include the inability to relight avatars dynamically, as appearance and lighting are baked into the captured radiance field, as well as unconstrained motion in areas outside the face mesh, such as loose hair and accessories. Further engineering should integrate specialized hair modeling and separate material properties from illumination to achieve production-ready versatility.

arXiv: 2312.02069
  • Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). FLAME supplies the articulated parametric face model for identity, head pose, and expression that GaussianAvatars uses to rig and animate its Gaussian primitives.
  • Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). 3D Gaussian Splatting establishes the explicit Gaussian representation, adaptive densification, and real-time differentiable rendering that GaussianAvatars adapts to animated heads.
  • Paper: RigNeRF: Fully Controllable Neural 3D Portraits, ShahRukh Athar et al. (2022). RigNeRF provides the preceding framework for coupling neural radiance representations with controllable head pose and facial expression from portrait video.
  • Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). Learning Neural Parametric Head Models develops disentangled identity and expression geometry that clarifies GaussianAvatars' use of a deformable head prior.
  • Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). D-NeRF introduces canonical-space deformation for dynamic view synthesis, a conceptual precursor to transforming appearance primitives under facial motion.
  • Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). 4D Gaussian Splatting shows how a canonical Gaussian field can be deformed over time, providing the dynamic-representation context for rigged facial Gaussians.
  • Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Nerfies establishes deformation-field reconstruction of nonrigid subjects from video, framing the motion and regularization challenges that GaussianAvatars addresses with mesh-bound splats.
  • Paper: Relightable Gaussian Codec Avatars, Shunsuke Saito et al. (2024). Relightable Gaussian Codec Avatars extends rigged Gaussian head avatars with explicit material and illumination modeling, addressing GaussianAvatars' baked-lighting limitation.
  • Paper: 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis, Zhicheng Lu et al. (2024). 3D Geometry-aware Deformable Gaussian Splatting generalizes Gaussian motion modeling with local geometric features and continuous rotations for more structurally consistent dynamic rendering.
Cover for GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Radiance Field Reconstruction
  • 2.2. Human Head Reconstruction and Animation
  • 3. Method
  • 3.1. Preliminary
  • 3.2. 3D Gaussian Rigging
  • 3.3. Binding Inheritance
  • 3.4. Optimization and Regularization
  • 4. Experiments
  • 4.1. Setup
  • 4.2. Head Avatar Reconstruction and Animation
  • 4.3. Ablation Study
  • 5. Limitations and Potential Negative Impacts
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — GaussianAvatars representation and controllability

    model/method

    GaussianAvatars represents a human-head avatar as anisotropic 3D Gaussian splats attached to triangles of a FLAME parametric face mesh. Each Gaussian stores a local position, local rotation, anisotropic local scale, opacity, spherical-harmonic color coefficients, and the index of its parent FLAME triangle. The initial representation contains one Gaussian at the center of every mesh triangle, but optimization and adaptive density control can add additional splats.

    When FLAME pose or expression parameters change, each Gaussian is transformed with its parent triangle, giving direct control over head pose and facial expression while retaining the expressive appearance of 3D Gaussian Splatting. The Gaussian layer can move away from the mesh surface to represent details that FLAME cannot reproduce accurately, including hair strands, teeth, wrinkles, and other fine-scale appearance. Gaussian parameters and FLAME parameters are jointly optimized from multi-view video in an end-to-end reconstruction process.

  2. Knowl 2 — Triangle-local rigging transformation

    equation

    For every FLAME triangle, GaussianAvatars constructs a local coordinate frame. Let T∈R3\mathbf{T}\in\mathbb{R}^3 be the triangle centroid, let R∈R3×3\mathbf{R}\in\mathbb{R}^{3\times 3} be the rotation matrix whose columns are an edge direction, the triangle normal, and their cross product, and let k>0k>0 be a scalar describing the triangle's metric scale. A Gaussian attached to this triangle has local position μ∈R3\boldsymbol{\mu}\in\mathbb{R}^3, local rotation r∈R3×3\mathbf{r}\in\mathbb{R}^{3\times 3}, and local scale s∈R3\mathbf{s}\in\mathbb{R}^3. Its global parameters at rendering time are

    r′=Rr,μ′=kRμ+T,s′=ks.\mathbf{r}'=\mathbf{R}\mathbf{r},\qquad \boldsymbol{\mu}'=k\mathbf{R}\boldsymbol{\mu}+\mathbf{T},\qquad \mathbf{s}'=k\mathbf{s}.

    The local position, rotation, and scale are initialized to μ=0\boldsymbol{\mu}=\mathbf{0}, r=I\mathbf{r}=\mathbf{I}, and s=(1,1,1)\mathbf{s}=(1,1,1), respectively. The triangle-scale factor makes local parameters relative to the parent triangle: a Gaussian attached to a smaller triangle takes smaller metric-space steps under the same local update than a Gaussian attached to a larger triangle.

  3. Knowl 3 — Binding inheritance during adaptive density control

    algorithm

    GaussianAvatars preserves mesh control while using the adaptive density operations of 3D Gaussian Splatting.

    Input: Gaussian splats with parent-triangle indices, local Gaussian parameters, view-space positional gradients, and opacities.

    Output: A densified and pruned set of Gaussians in which every splat remains attached to a FLAME triangle.

    For each Gaussian with a large view-space positional gradient, split it into two smaller Gaussians when the Gaussian is large, or clone it when the Gaussian is small. Perform these operations in the local coordinate system and initialize every new Gaussian close to the triggering Gaussian. Assign every new Gaussian the triggering Gaussian's parent-triangle index.

    Periodically reset opacities that are close to zero and remove Gaussians whose opacity falls below the pruning threshold. Track the number of Gaussians attached to every FLAME triangle and do not prune the last Gaussian attached to any triangle.

    This inheritance rule allows multiple splats to model regions such as curved hair while preventing pruning from eliminating all splats on frequently occluded regions such as eyeballs.

  4. Knowl 4 — Animation-aware optimization objective

    equation

    GaussianAvatars supervises differentiable renders with an RGB loss combining an L1L_1 image loss and a D-SSIM loss:

    Lrgb=(1−λ)L1+λLD-SSIM,λ=0.2.\mathcal{L}_{\mathrm{rgb}}=(1-\lambda)\mathcal{L}_1+\lambda\mathcal{L}_{\mathrm{D\text{-}SSIM}},\qquad \lambda=0.2.

    To keep each Gaussian geometrically associated with its parent triangle during animation, the method adds thresholded penalties on the local position μ\boldsymbol{\mu} and local scale s\mathbf{s}. For a vector x\mathbf{x}, define the elementwise threshold-excess operator Eϵ(x)=max⁡(∣x∣−ϵ,0)E_\epsilon(\mathbf{x})=\max(|\mathbf{x}|-\epsilon,0). The regularizers are

    Lposition=∥Eϵposition(μ)∥2,ϵposition=1,\mathcal{L}_{\mathrm{position}}=\left\|E_{\epsilon_{\mathrm{position}}}(\boldsymbol{\mu})\right\|_2, \qquad \epsilon_{\mathrm{position}}=1, Lscaling=∥Eϵscaling(s)∥2,ϵscaling=0.6.\mathcal{L}_{\mathrm{scaling}}=\left\|E_{\epsilon_{\mathrm{scaling}}}(\mathbf{s})\right\|_2, \qquad \epsilon_{\mathrm{scaling}}=0.6.

    The final objective is

    L=Lrgb+λpositionLposition+λscalingLscaling,λposition=0.01,λscaling=1.\mathcal{L}=\mathcal{L}_{\mathrm{rgb}}+\lambda_{\mathrm{position}}\mathcal{L}_{\mathrm{position}}+\lambda_{\mathrm{scaling}}\mathcal{L}_{\mathrm{scaling}}, \qquad \lambda_{\mathrm{position}}=0.01,\quad \lambda_{\mathrm{scaling}}=1.

    The position threshold permits small deviations from the parent triangle, while the scale threshold prevents Gaussians from becoming excessively large relative to the triangle. The regularizers are applied only to visible splats, so occluded regions are not unnecessarily altered when they do not contribute to the current color loss.

  5. Knowl 5 — Joint Gaussian and FLAME optimization procedure

    algorithm

    GaussianAvatars first obtains per-frame FLAME parameters with a photometric head tracker using the multi-view images and known camera parameters. The optimized FLAME quantities include shape β\boldsymbol{\beta}, translation t\mathbf{t}, joint pose θ\boldsymbol{\theta}, expression ψ\boldsymbol{\psi}, and canonical-space vertex offsets Δv\Delta\mathbf{v}.

    Use Adam for 600,000 iterations. Optimize Gaussian parameters together with per-frame FLAME translation, joint rotation, and expression. The learning rates for Gaussian local position and local scale are 5×10−35\times10^{-3} and 1.7×10−21.7\times10^{-2}; the learning rates for FLAME translation, joint rotation, and expression are 10−610^{-6}, 10−510^{-5}, and 10−310^{-3}, respectively. Use the standard 3D Gaussian Splatting learning rates for the remaining Gaussian parameters.

    Exponentially decay the Gaussian-position learning rate so that it reaches 0.010.01 times its initial value at the final iteration. Run adaptive density control with binding inheritance every 2,000 iterations from iteration 10,000 through the end of training, and reset Gaussian opacities every 60,000 iterations.

  6. Knowl 6 — Multi-view evaluation protocol and data

    experimental setup

    The evaluation uses the NeRSemble multi-view head-video dataset with 9 subjects. Each subject has 16 cameras covering the front and sides, 11 video sequences, and images downsampled to 802×550802\times550 pixels. Participants perform prescribed expressions or emotions in 10 sequences and perform freely in an eleventh sequence.

    Three evaluation settings are used. In novel-view synthesis, an avatar is driven by poses and expressions from training sequences and rendered from a held-out camera. In self-reenactment, an avatar is driven by unseen poses and expressions from a held-out sequence of the same subject and rendered from all 16 cameras. In cross-identity reenactment, the pose and expression parameters tracked from one subject drive the avatar reconstructed for another subject.

    For quantitative evaluation, each method is trained on 9 of the 10 prescribed sequences and 15 of the 16 cameras. The free-performance sequence is used to assess cross-identity reenactment. GaussianAvatars is compared with AvatarMAV, PointAvatar, and INSTA using PSNR, SSIM, and LPIPS.

  7. Knowl 7 — Quantitative improvement over competing head-avatar methods

    data/table

    The comparison below evaluates novel-view synthesis and self-reenactment on the NeRSemble protocol. Higher PSNR and SSIM are better, while lower LPIPS is better. GaussianAvatars is best on all three novel-view metrics and achieves the best SSIM and LPIPS in self-reenactment; INSTA has slightly higher self-reenactment PSNR.

    Could not parse LaTeX table

    The results demonstrate a large novel-view advantage for GaussianAvatars, especially in perceptual similarity. Its self-reenactment LPIPS of 0.0760.076 is lower than AvatarMAV's 0.1680.168, PointAvatar's 0.1020.102, and INSTA's 0.1100.110, indicating sharper and more perceptually faithful animated renderings.

  8. Knowl 8 — Cross-identity expression transfer

    empirical result

    In cross-identity reenactment, GaussianAvatars drives a target avatar with tracked FLAME pose and expression parameters from a different source actor. The resulting renderings reproduce eye blinks, mouth movements, wrinkles, and other complex facial dynamics vividly while maintaining high visual quality.

    The reported qualitative comparison shows that INSTA can produce aliasing when motion leaves the occupancy grid optimized for its training sequences, PointAvatar does not maintain a deformation space consistently aligned with FLAME and therefore transfers motion imprecisely, and AvatarMAV degrades substantially because it lacks sufficiently strong deformation priors. GaussianAvatars remains controllable because every splat retains a fixed triangle correspondence throughout animation.

  9. Knowl 9 — Ablation evidence for the principal components

    data/table

    The ablation study evaluates GaussianAvatars on subject #304. It removes adaptive density control with binding inheritance, scale regularization or its tolerance, position regularization or its tolerance, and FLAME fine-tuning. Higher PSNR and SSIM are better, while lower LPIPS is better.

    Could not parse LaTeX table

    Removing adaptive density control causes a large perceptual-quality loss because the representation cannot add enough splats and existing splats become overly large. Removing scale regularization produces spike artifacts, while removing its tolerance causes the strongest metric degradation because splats shrink excessively. Removing position regularization can improve held-out-view pixel metrics by allowing overfitting, but produces cracks and floating blobs under novel expressions and poses. Removing FLAME fine-tuning reduces novel-view quality and increases self-reenactment perceptual error.

  10. Knowl 10 — Technical and societal limitations

    limitation

    GaussianAvatars directly models radiance with 3D Gaussian splats without separating material properties from illumination, so relighting is not feasible within the presented approach. The method also lacks direct control over hair and accessories that are not represented by the FLAME face model; explicitly modeling these parts is identified as future work.

    The paper additionally notes risks associated with photorealistic avatar generation, including unauthorized manipulation of a person's likeness, deceptive deepfake content, misinformation, defamation, identity theft, impersonation, and fraud. These risks limit the acceptable deployment of the technology even though they do not arise from the reconstruction objective itself.

Coverage note — Only standard 3D Gaussian Splatting preliminaries and related-work material were omitted because they are not GaussianAvatars' own contribution; the contributed representation, rigging, density control, optimization, evaluation, ablations, and limitations are covered.

References

  1. 1.Benjamin Attal, Jia-Bin Huang, Christian Richardt, Michael Zollhoefer, Johannes Kopf, Matthew O’Toole, and Changil Kim. Hyperreel: High-fidelity 6-dof video with ray-conditioned sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16610–16620, 2023. 2
  2. 2.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5855–5864, 2021. 1
  3. 3.Volker Blanz and Thomas Vetter. Face recognition based on fitting a 3d morphable model. IEEE Transactions on pattern analysis and machine intelligence, 25(9):1063–1074, 2003. 2
  4. 4.Ang Cao and Justin Johnson. Hexplane: A fast representation for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 130–141, 2023. 2
  5. 5.Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. Everybody dance now. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5933–5942, 2019. 2
  6. 6.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In European Conference on Computer Vision, pages 333–350. Springer, 2022. 1, 2
  7. 7.Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. K-planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12479–12488, 2023. 2
  8. 8.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8649–8658, 2021. 2
  9. 9.Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5712–5721, 2021. 2
  10. 10.Xuan Gao, Chenglai Zhong, Jun Xiang, Yang Hong, Yudong Guo, and Juyong Zhang. Reconstructing personalized semantic facial nerf models from monocular video. ACM Transactions on Graphics (Proceedings of SIGGRAPH Asia), 41(6), 2022. 2
  11. 11.Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. Neural head avatars from monocular rgb videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18653–18664, 2022. 2
  12. 12.Yang Hong, Bo Peng, Haiyao Xiao, Ligang Liu, and Juyong Zhang. Headnerf: A real-time nerf-based parametric head model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20374–20384, 2022. 2
  13. 13.Mustafa Is¸ık, Martin Runz, Markos Georgopoulos, Taras ¨Khakhulin, Jonathan Starck, Lourdes Agapito, and Matthias Nießner. Humanrf: High-fidelity neural radiance fields for humans in motion. ACM Transactions on Graphics (TOG), 42(4):1–12, 2023. 1
  14. 14.Rohit Jena, Ganesh Subramanian Iyer, Siddharth Choudhary, Brandon Smith, Pratik Chaudhari, and James Gee. Splatarmor: Articulated gaussian splatting for animatable humans from monocular rgb videos. arXiv preprint arXiv:2311.10812, 2023. 2
  15. 15.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, ¨and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023. 1, 2, 3, 4, 5
  16. 16.Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Perez, Christian ´Richardt, Michael Zollhofer, and Christian Theobalt. Deep video portraits. ACM transactions on graphics (TOG), 37(4):1–14, 2018. 2
  17. 17.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  18. 18.Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view radiance field reconstruction of human heads. ACM Trans. Graph., 42(4), 2023. 1, 5
  19. 19.Muhammed Kocabas, Jen-Hao Rick Chang, James Gabriel, Oncel Tuzel, and Anurag Ranjan. Hugs: Human gaussian splats. arXiv preprint arXiv:2311.17910, 2023. 2
  20. 20.Jiahui Lei, Yufu Wang, Georgios Pavlakos, Lingjie Liu, and Kostas Daniilidis. Gart: Gaussian articulated template models. arXiv preprint arXiv:2311.16099, 2023. 2
  21. 21.Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6):194:1–194:17, 2017. 2, 3, 4, 5
  22. 22.Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard Newcombe, et al. Neural 3d video synthesis from multi-view video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5521–5531, 2022. 1
  23. 23.Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling. arXiv, 2023. 2
  24. 24.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose control. ACM transactions on graphics (TOG), 40(6):1–16, 2021. 2
  25. 25.Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mixture of volumetric primitives for efficient neural rendering. ACM Transactions on Graphics (ToG), 40(4):1–13, 2021. 2
  26. 26.Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multi-person linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023. 2
  27. 27.Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713, 2023. 2
  28. 28.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. 1, 2
  29. 29.Thomas Muller, Alex Evans, Christoph Schied, and Alexan- ¨der Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022. 1, 2, 7
  30. 30.Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5865–5874, 2021. 1, 2
  31. 31.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. arXiv preprint arXiv:2106.13228, 2021. 1, 2
  32. 32.Pascal Paysan, Reinhard Knothe, Brian Amberg, Sami Romdhani, and Thomas Vetter. A 3d face model for pose and illumination invariant face recognition. In 2009 sixth IEEE international conference on advanced video and signal based surveillance, pages 296–301. Ieee, 2009. 2
  33. 33.Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Animatable neural radiance fields for modeling dynamic human bodies. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14314–14323, 2021. 2
  34. 34.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10318–10327, 2021. 2
  35. 35.Radu Alexandru Rosu, Shunsuke Saito, Ziyan Wang, Chenglei Wu, Sven Behnke, and Giljoo Nam. Neural strands: Learning hair geometry and appearance from multi-view images. In European Conference on Computer Vision, pages 73–89. Springer, 2022. 8
  36. 36.Liangchen Song, Anpei Chen, Zhong Li, Zhang Chen, Lele Chen, Junsong Yuan, Yi Xu, and Andreas Geiger. Nerfplayer: A streamable dynamic scene representation with decomposed neural radiance fields. IEEE Transactions on Visualization and Computer Graphics, 29(5):2732–2742, 2023. 2
  37. 37.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5459–5469, 2022. 2
  38. 38.Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (ToG), 36(4):1–13, 2017. 2
  39. 39.Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2387–2395, 2016. 2, 3
  40. 40.Justus Thies, Michael Zollhofer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (TOG), 38(4):1–12, 2019. 2
  41. 41.Ziyan Wang, Giljoo Nam, Tuur Stuyck, Stephen Lombardi, Chen Cao, Jason Saragih, Michael Zollhofer, Jessica Hodgins, and Christoph Lassner. Neuwigs: A neural dynamic model for volumetric hair capture and animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8641–8651, 2023. 8
  42. 42.Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9421–9431, 2021. 2
  43. 43.Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5438–5448, 2022. 2
  44. 44.Yuelang Xu, Lizhen Wang, Xiaochen Zhao, Hongwen Zhang, and Yebin Liu. Avatarmav: Fast 3d head avatar reconstruction using motion-aware neural voxels. In ACM SIGGRAPH 2023 Conference Proceedings, 2023. 2, 5, 7
  45. 45.Yuelang Xu, Hongwen Zhang, Lizhen Wang, Xiaochen Zhao, Huang Han, Qi Guojun, and Yebin Liu. Latentavatar: Learning latent expression code for expressive neural head avatar. In ACM SIGGRAPH 2023 Conference Proceedings, 2023. 2
  46. 46.Keyang Ye, Tianjia Shao, and Kun Zhou. Animatable 3d gaussians for high-fidelity synthesis of human motions. arXiv preprint arXiv:2311.13404, 2023. 2
  47. 47.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4578–4587, 2021. 2
  48. 48.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 5
  49. 49.Yufeng Zheng, Victoria Fernandez Abrevaya, Marcel C ´Buhler, Xu Chen, Michael J Black, and Otmar Hilliges. Im avatar: Implicit morphable head avatars from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13545–13555, 2022. 2
  50. 50.Yufeng Zheng, Wang Yifan, Gordon Wetzstein, Michael J Black, and Otmar Hilliges. Pointavatar: Deformable point-based head avatars from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21057–21067, 2023. 2, 5, 7
  51. 51.Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito, Michael Zollhofer, Justus Thies, and Javier Romero. Drivable 3d gaussian avatars. 2023. 2
  52. 52.Wojciech Zielonka, Timo Bolkart, and Justus Thies. Instant volumetric head avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4574–4584, 2023. 2, 5, 7

Citation

MLA
Qian, S., et al. “GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians”. arXiv, 2023, http://arxiv.org/abs/2312.02069v2.
APA
Qian, S., Kirschstein, T., Schoneveld, L., Davoli, D., Giebenhain, S., & Nießner, M. (2023). GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians. arXiv. http://arxiv.org/abs/2312.02069v2
Chicago
Qian, S., T. Kirschstein, L. Schoneveld, D. Davoli, S. Giebenhain, and M. Nießner. 2023. “GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians”. arXiv. http://arxiv.org/abs/2312.02069v2.
Harvard
Qian, S. et al. (2023) “GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.02069v2.
Vancouver
1. Qian S, Kirschstein T, Schoneveld L, Davoli D, Giebenhain S, Nießner M (2023) GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians. arXiv

BibTeX

@article{qian2023gaussianavatars,
  title = {GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians},
  author = {Qian, Shenhan and Kirschstein, Tobias and Schoneveld, Liam and Davoli, Davide and Giebenhain, Simon and Nießner, Matthias},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.02069v2},
  eprint = {2312.02069}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE