Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos

Ziqian BaiFeitong TanZeng HuangKripasindhu SarkarDanhang TangDi QiuAbhimitra MekaRuofei DuMingsong DouSergio Orts-Escolano

article2023CVPR65 citations

Presents a hybrid pipeline that reconstructs photorealistic, user-controllable 3D head avatars from casual monocular RGB videos by anchoring CNN-predicted local UV features onto a parametric face mesh to guide a neural radiance field.

Listen

Generating realistic and controllable 3D human head avatars is essential for emerging digital communication, virtual reality, gaming, and visual effects. However, existing high-fidelity systems typically rely on expensive multi-camera rigs or specialized depth sensors, creating high barriers to broader adoption. Previous attempts to build controllable avatars from standard, casual smartphone or webcam video struggle with a fundamental trade-off: traditional surface models fail to capture dynamic features like hair, wrinkles, and accessories, while newer neural rendering methods often produce over-smoothed results, blurry artifacts, or distorted facial geometry during novel expressions.

The article demonstrates a novel framework that learns a personalized, photorealistic, and fully controllable 3D volumetric head avatar using only a single, short monocular video clip of one to two minutes. The primary objective is to achieve fine-grained control over head poses and facial expressions while preserving subject-specific details across novel camera viewpoints.

To achieve this, the approach anchors a neural radiance field to the geometry of a 3D Morphable Model (a standard parametric face mesh). Instead of relying on global networks that struggle to capture fine details, the method predicts spatially local, expression-dependent features anchored directly to the face mesh vertices. These dynamic features are generated by rasterizing vertex displacements into a 2D surface map and processing them with a convolutional neural network. Volumetric color and density at any point in 3D space are then decoded locally by interpolating adjacent vertex features, with per-frame error correction used during training to account for tracking inaccuracies.

The evaluation yields several key findings across quantitative metrics and visual assessments. First, the proposed displacement-driven model consistently outperforms current state-of-the-art monocular avatar methods across standard perceptual and image quality benchmarks (LPIPS, SSIM, and PSNR). Second, using convolutional networks over surface vertex displacements provides critical spatial context, significantly outperforming standard multilayer perceptrons in synthesizing fine details like teeth, cheek contours, and reflections on eyeglasses. Third, the system demonstrates superior generalization and robustness when rendering unseen expressions or extrapolated, extreme facial movements. Finally, when evaluated on training data restricted to only 50% of the video clip, the method degrades far less than alternative designs, confirming that high-quality avatars can be generated from very brief user captures.

These findings prove that high-end 3D avatars can be democratized using standard consumer hardware without requiring specialized studio environments, significantly lowering production costs and deployment friction for digital human applications. However, organizations evaluating this technology should account for key operational trade-offs and limitations. Training remains subject-specific and rendering relies on volumetric neural fields, which are computationally intensive and not yet optimized for low-latency, real-time edge environments. Additionally, because the architecture is anchored to a facial parametric model, it cannot synthesize structures absent from the underlying mesh, such as the tongue or the torso.

Next steps for development include exploring model compression and acceleration to enable real-time mobile rendering, extending the geometric anchor to support upper-body and full-body avatars, and implementing proactive governance measures. Given the realism of the synthesized heads, deploying teams should integrate cryptographic signatures or digital watermarking to prevent potential deepfake misuse and ensure content authenticity.

arXiv: 2304.01436
  • Paper: RigNeRF: Fully Controllable Neural 3D Portraits, ShahRukh Athar et al. (2022). RigNeRF introduces deformation fields guided by 3D morphable models to control neural radiance fields from monocular portrait video, establishing the core hybrid methodology that the source paper builds upon.
  • Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). FLAME provides the articulated 3D morphable head model that establishes the geometric prior and parametric facial control utilized by the source paper.
  • Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Nerfies establishes foundational techniques for optimizing continuous deformation fields to reconstruct dynamic and non-rigid human subjects from monocular video.
  • Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). D-NeRF introduces the canonical space and time-conditioned deformation field framework for dynamic neural radiance fields that underpins volumetric avatar reconstruction.
  • Paper: A Morphable Model For The Synthesis Of 3D Faces, Volker Blanz et al. (1999). This seminal paper introduces 3D Morphable Models (3DMM), defining the parametric representation of facial geometry and texture that the source relies on for expression control.
  • Paper: FENeRF: Face Editing in Neural Radiance Fields, Jingxiang Sun et al. (2022). FENeRF establishes methods for decoupling facial geometry and semantics in 3D-aware neural radiance fields to enable local editing.
Cover for Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos

Abstract

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracking of a 3DMM with a neural radiance field to achieve fine-grained control and photorealism. To reduce over-smoothing and improve out-of-model expressions synthesis, we propose to predict local features anchored on the 3DMM geometry. These learnt features are driven by 3DMM deformation and interpolated in 3D space to yield the volumetric radiance at a designated query point. We further show that using a Convolutional Neural Network in the UV space is critical in incorporating spatial context and producing representative local features. Extensive experiments show that we are able to reconstruct high-quality avatars, with more accurate expression-dependent details, good generalization to out-of-training expressions, and quantitatively superior renderings compared to other state-of-the-art approaches.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Method
  • 3.1. Avatar Representation
  • 3.2. Predicting Expression-Dependent Features
  • 3.3. Training Schema
  • 4. Experiments
  • 4.1. Datasets and Metrics
  • 4.2. Comparisons with State-of-Art
  • 4.3. Driving the Avatar
  • 4.4. Ablation Study: Expression Features
  • 4.5. Robustness to Expression Extrapolation
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — 3DMM-Anchored Neural Radiance Field Architecture

    model/method

    A 3D head avatar is represented by anchoring a dynamic neural radiance field to the deformable mesh surface of a 3D Morphable Model (such as FLAME). Let V_i = V(eta, oldsymbol{ heta}_i, oldsymbol{oldsymbol{ heta}}_i) \subset \mathbb{R}^3 denote the 3DMM mesh vertices deformed by identity shape β\beta, expression parameters ψi\boldsymbol{\psi}_i, and pose parameters θi\boldsymbol{\theta}_i (including neck, jaw, and eye rotations) at frame ii. Each mesh vertex vijv_i^j is assigned a dynamic feature vector zijz_i^j.

    To compute the volume density σi(q)\sigma_i(q) and RGB color ci(q,di)c_i(q, d_i) at an arbitrary 3D query point q∈R3q \in \mathbb{R}^3 for ray direction di∈S2d_i \in \mathbb{S}^2, the kk-nearest neighbor (kk-NN) vertices Nkq\mathcal{N}_k^q of qq on the mesh ViV_i are identified. The local features are transformed and aggregated using two Multi-Layer Perceptrons (MLPs), F0F_0 and F1F_1, weighted by inverse Euclidean distance:

    z^ij=F0(vij−q,zij)\hat{z}_i^j = F_0(v_i^j - q, z_i^j)

    z^i=∑j∈Nkqwjz^ij,where wj=dj∑m∈Nkqdm,dj=1∥vij−q∥22\hat{z}_i = \sum_{j \in \mathcal{N}_k^q} w^j \hat{z}_i^j, \quad \text{where } w^j = \frac{d^j}{\sum_{m \in \mathcal{N}_k^q} d^m}, \quad d^j = \frac{1}{\|v_i^j - q\|_2^2}

    ci(q,di),σi(q)=F1(z^i,di)c_i(q, d_i), \sigma_i(q) = F_1(\hat{z}_i, d_i)

    Given a camera ray r(t)=o+tdir(t) = o + t d_i with near bound tnt_n and far bound tft_f, the rendered pixel color Ci(r)C_i(r) is obtained through numerical volume rendering:

    Ci(r)=∫tntfT(t)σi(r(t))ci(r(t),di) dt,where T(t)=exp⁡(−∫tntσi(r(s)) ds)C_i(r) = \int_{t_n}^{t_f} T(t) \sigma_i(r(t)) c_i(r(t), d_i) \, dt, \quad \text{where } T(t) = \exp\left(-\int_{t_n}^t \sigma_i(r(s)) \, ds\right)

  2. Knowl 2 — UV-Space Expression-Dependent Feature Learning

    model/method

    To capture fine-grained, dynamic expression variations and overcome the capacity limitations of global MLPs, dynamic vertex features zijz_i^j are synthesized via a 2D Convolutional Neural Network operating in the 3DMM texture atlas space (UV space).

    1. Vertex Displacement Formulation (Ours-D): Per-vertex 3D displacement vectors are computed relative to the neutral mesh configuration:

    Di=Vi(ψi,θi)−Vneutral(0,0)D_i = V_i(\boldsymbol{\psi}_i, \boldsymbol{\theta}_i) - V_{\text{neutral}}(0, 0)

    where Vi(ψi,θi)V_i(\boldsymbol{\psi}_i, \boldsymbol{\theta}_i) is the posed/expressed mesh and Vneutral(0,0)V_{\text{neutral}}(0, 0) is the neutral template mesh. The displacement vectors DiD_i are rasterized into a 2D UV-space map and processed through a 2D convolutional U-Net FD\mathcal{F}_D. The resulting 2D feature map is sampled back onto the mesh vertices ViV_i using UV coordinates to yield the dynamic feature vectors {zij}\{z_i^j\}.

    1. 1D Parameter Decoding Variant (Ours-C): In an alternative formulation, a transposed convolutional decoder processes the concatenated 1D parameter vectors [ψi,θi][\boldsymbol{\psi}_i, \boldsymbol{\theta}_i] directly into the 2D UV feature map.

    Conditioning the network on explicit per-vertex displacement fields in UV space (Ours-D) provides spatial convolutional context between adjacent facial regions, improving robustness to extreme or out-of-training expressions.

  3. Knowl 3 — Error-Correction Warping Field for Monocular Avatars

    model/method

    To compensate for tracking noise, non-rigid hair motion, and registration misalignments between the parametric 3DMM mesh and the ground-truth video frames, an error-correction warping field is optimized during training.

    For frame ii, an MLP FE\mathcal{F}_E takes the spatial query point q∈R3q \in \mathbb{R}^3 and an optimizable per-frame latent code eie_i to output a spatial transformation Ti(q)T_i(q):

    q′=Ti(q)=FE(q,ei)q' = T_i(q) = \mathcal{F}_E(q, e_i)

    The warped query point q′q' is used in place of qq to find the kk-NN mesh vertices and evaluate color and density. This warping module is active exclusively during training to absorb unmodeled per-frame tracking discrepancies and is disabled at inference time to ensure strict 3DMM-driven controllability.

  4. Knowl 4 — Optimization Objective and Regularization Schema

    equation

    The complete avatar model parameters—including the UV feature extraction network, the local MLPs F0F_0 and F1F_1, and the deformation field FE\mathcal{F}_E with latent codes eie_i—are trained end-to-end on monocular RGB frames using the total loss function:

    L=Lrgb+λelasticLelastic+λmagLmag\mathcal{L} = \mathcal{L}_{rgb} + \lambda_{\text{elastic}} \mathcal{L}_{\text{elastic}} + \lambda_{\text{mag}} \mathcal{L}_{\text{mag}}

    where:

    Lrgb=∑i∑r∥Ci(r)−Ii(r)∥22\mathcal{L}_{rgb} = \sum_{i} \sum_{r} \|C_i(r) - I_i(r)\|_2^2

    penalizes the ℓ2\ell_2 photometric difference between rendered ray colors Ci(r)C_i(r) and ground-truth pixel values Ii(r)I_i(r);

    Lmag=∑q∥q−T(q)∥22\mathcal{L}_{\text{mag}} = \sum_{q} \|q - T(q)\|_2^2

    constrains the magnitude of the predicted error-correction warping field T(q)T(q); and Lelastic\mathcal{L}_{\text{elastic}} penalizes non-rigid deformations in the warping field.

    The hyperparameter weights are set to λmag=10−2\lambda_{\text{mag}} = 10^{-2}, while λelastic\lambda_{\text{elastic}} is initialized to 10−410^{-4} and decayed to 10−510^{-5} after 155,000 optimization iterations.

  5. Knowl 5 — Quantitative Avatar Reconstruction and Novel View Benchmark Comparison

    data/table

    Models are evaluated on monocular RGB video test splits using Learned Perceptual Image Patch Similarity (LPIPS, lower is better), Structural Similarity Index Measure (SSIM, higher is better), and Peak Signal-to-Noise Ratio (PSNR in dB, higher is better). The comparison includes 2D motion transfer methods (TPSMM, FOMM) and 3D avatar approaches (NHA, IMAvatar, NerFACE), evaluated across five subjects.

    Method Subject 0 Subject 1 Subject 2 Subject 3 Subject 4
    LPIPS / SSIM / PSNR LPIPS / SSIM / PSNR LPIPS / SSIM / PSNR LPIPS / SSIM / PSNR LPIPS / SSIM / PSNR
    TPSMM 0.192 / 0.852 / 22.60 0.205 / 0.830 / 16.38 0.216 / 0.782 / 18.40 0.222 / 0.799 / 20.28 0.156 / 0.913 / 21.29
    FOMM 0.171 / 0.841 / 22.93 0.179 / 0.827 / 16.02 0.202 / 0.777 / 18.98 0.186 / 0.798 / 22.28 0.122 / 0.915 / 23.94
    NHA 0.165 / 0.836 / 20.20 0.166 / 0.840 / 15.48 0.178 / 0.809 / 17.99 0.153 / 0.798 / 21.31 0.091 / 0.926 / 23.78
    IMAvatar 0.207 / 0.852 / 21.26 0.187 / 0.848 / 15.98 0.265 / 0.729 / 15.80 0.214 / 0.782 / 20.37 0.142 / 0.897 / 20.63
    NerFACE 0.205 / 0.817 / 20.06 0.182 / 0.833 / 15.78 0.188 / 0.793 / 19.41 0.229 / 0.747 / 18.16 0.093 / 0.938 / 25.57
    Ours-D 0.144 / 0.864 / 21.92 0.152 / 0.855 / 16.23 0.141 / 0.841 / 20.42 0.156 / 0.833 / 23.05 0.075 / 0.944 / 25.71

    Ours-D consistently achieves the lowest perceptual distortion (LPIPS) across all subjects, outperforming both global MLP neural radiance field methods (NerFACE, IMAvatar) and neural mesh texture methods (NHA).

  6. Knowl 6 — Ablation Study on Expression Feature Encoding and Data Sparsity

    data/table

    Ablations evaluate different feature conditioning mechanisms on LPIPS perceptual error (lower is better) across four subjects, comparing full training sequences against reduced training sets consisting only of the first 50% of frames.

    Configuration Subject 0 Subject 1 Subject 2 Subject 3
    Full Training Data
    Static Features 0.1559 0.1586 0.1552 0.1688
    3DMM Codes 0.1599 0.1620 0.1746 0.1738
    3DMM Codes MLP 0.1568 0.1551 0.1505 0.1686
    Ours-C 0.1417 0.1457 0.1383 0.1550
    Ours-D 0.1439 0.1523 0.1415 0.1559
    First 50% Training Data
    Ours-C 0.2038 (+0.0621) 0.1483 (+0.0026) 0.1580 (+0.0197) 0.1566 (+0.0016)
    Ours-D 0.1711 (+0.0272) 0.1511 (-0.0012) 0.1516 (+0.0101) 0.1558 (-0.0001)

    Key observations:

    1. UV-CNN architectures (Ours-C, Ours-D) outperform global/per-vertex MLP decoders (Static Features, 3DMM Codes, 3DMM Codes MLP) because CNN operations capture spatial context between neighboring vertices.
    2. When training data is halved, Ours-D (UV displacement input) generalizes to unseen expressions with significantly smaller degradation (+0.0272+0.0272 LPIPS on Subject 0) compared to Ours-C (+0.0621+0.0621 LPIPS).
  7. Knowl 7 — Limitations of 3DMM-Anchored Neural Head Avatars

    limitation

    The 3DMM-anchored neural head avatar approach has the following limitations:

    1. 3DMM Representation Bounds: The avatar is strictly driven by the expression parameter space of the underlying linear 3DMM (FLAME). Structures that are not represented in the parametric model (such as the tongue and inner mouth geometry) cannot be synthesized or controlled.
    2. NeRF Computational Overhead: The pipeline inherits volumetric neural radiance field training and rendering overhead, resulting in time-consuming subject-specific optimization and non-real-time volumetric raymarching rendering.
    3. Domain Restriction: The model reconstructs only portrait head avatars and does not extend to upper-body, clothing, or full-body motion capture.

Coverage note — None was omitted; all key contributions—including geometry-anchored NeRF formulation, UV displacement feature learning, error-correction deformation, loss functions, quantitative benchmarks, ablation studies, and limitations—are fully covered.

References

  1. 1.MetaHuman - Unreal Engine. https : / / www . unrealengine.com/en- US/metahuman, 2. Accessed: 2022-10-17. 1
  2. 2.Shruti Agarwal, Hany Farid, Tarek El-Gaaly, and Ser-Nam Lim. Detecting deep-fake videos from appearance and behavior. In 2020 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–6. IEEE, 2020. 8
  3. 3.ShahRukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman, and Zhixin Shu. RigNeRF: Fully Controllable Neural 3D Portraits. pages 20364–20373, 2022. 2, 3, 4
  4. 4.Ziqian Bai, Zhaopeng Cui, Xiaoming Liu, and Ping Tan. Riggable 3D Face Reconstruction via In-Network Optimization. pages 6216–6225, June 2021. 2
  5. 5.Ziqian Bai, Zhaopeng Cui, Jamal Ahmed Rahim, Xiaoming Liu, and Ping Tan. Deep Facial Non-Rigid Multi-View Stereo. pages 5850–5860, 2020. 2
  6. 6.Thabo Beeler, Fabian Hahn, Derek Bradley, Bernd Bickel, Paul Beardsley, Craig Gotsman, Robert W Sumner, and Markus Gross. High-Quality Passive Facial Performance Capture Using Anchor Frames. In ACM SIGGRAPH 2011 Papers, pages 1–10. 2011. 1
  7. 7.Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques, pages 187–194, 1999. 2
  8. 8.Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhoefer, Shun-Suke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, and Jason Saragih. Authentic Volumetric Avatars from a Phone Scan. ACM Transactions on Graphics (TOG), 41(4), 2022. 1
  9. 9.Chen Cao, Hongzhi Wu, Yanlin Weng, Tianjia Shao, and Kun Zhou. Real-Time Facial Animation With Image-Based Dynamic Avatars. 35(4), 2016. 2
  10. 10.Bindita Chaudhuri, Noranart Vesdapunt, Linda Shapiro, and Baoyuan Wang. Personalized Face Modeling for Improved Face Reconstruction and Motion Retargeting. In ECCV 2020: 16th European Conference on Computer Vision, pages 142–160, 2020. 2
  11. 11.Chong Bao and Bangbang Yang, Zeng Junyi, Bao Hujun, Zhang Yinda, Cui Zhaopeng, and Zhang Guofeng. Neumesh: Learning disentangled neural mesh-based implicit field for geometry and texture editing. In European Conference on Computer Vision (ECCV), 2022. 3
  12. 12.Ruofei Du, Ming Chuang, Wayne Chang, Hugues Hoppe, and Amitabh Varshney. Montage4D: Real-time Seamless Fusion and Stylization of Multiview Video Textures. Journal of Computer Graphics Techniques, 8(1):1–34, Jan. 2019. 1
  13. 13.Ruofei Du, David Li, and Amitabh Varshney. Geollery: A Mixed Reality Social Media Platform. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, number 685 in CHI. ACM, May 2019. 1
  14. 14.Bernhard Egger, William AP Smith, Ayush Tewari, Stefanie Wuhrer, Michael Zollhoefer, Thabo Beeler, Florian Bernard, Timo Bolkart, Adam Kortylewski, Sami Romdhani, et al. 3D Morphable Face Models—Past, Present, and Future. 39(5):1–38, 2020. 2
  15. 15.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction. pages 8649–8658, 2021. 2, 3, 4, 5, 6, 8
  16. 16.Pablo Garrido, Michael Zollhofer, Dan Casas, Levi Valgaerts, Kiran Varanasi, Patrick Perez, and Christian Theobalt. Reconstruction of Personalized 3D Face Rigs From Monocular Video. ACM Transactions on Graphics (TOG), 35(3):28, 2016. 2
  17. 17.Philip-William Grassal, Malte Prinzler, Titus Leistner, Carsten Rother, Matthias Nießner, and Justus Thies. Neural Head Avatars From Monocular RGB Videos. pages 18653–18664, 2022. 2, 5, 6, 8
  18. 18.Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio OrtsEscolano, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor, Paul Debevec, and Shahram Izadi. The Relightables: Volumetric Performance Capture of Humans With Realistic Relighting. ACM Transactions on Graphics, Nov. 2019. 1
  19. 19.Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio OrtsEscolano, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor, Paul Debevec, and Shahram Izadi. The relightables: Volumetric performance capture of humans with realistic relighting. ACM Trans. Graph., 38(6), nov 2019. 3
  20. 20.Zhenyi He, Ruofei Du, and Ken Perlin. CollaboVR: A Reconfigurable Framework for Multi-user to Communicate in Virtual Reality. In 2020 IEEE International Symposium on Mixed and Augmented Reality, ISMAR, pages 542–554. IEEE, Nov. 2020. 1
  21. 21.Liwen Hu, Shunsuke Saito, Lingyu Wei, Koki Nagano, Jaewoo Seo, Jens Fursund, Iman Sadeghi, Carrie Sun, YenChun Chen, and Hao Li. Avatar Digitization From a Single Image for Real-Time Rendering. 36(6):1–14, 2017. 2
  22. 22.Alexandru Eugen Ichim, Sofien Bouaziz, and Mark Pauly. Dynamic 3D Avatar Creation From Hand-Held Video Input. 34(4):1–14, 2015. 2
  23. 23.Wei Jiang, Kwang Moo Yi, Golnoosh Samei, Oncel Tuzel, and Anurag Ranjan. NeuMan: Neural Human Radiance Field From a Single Video. 2022. 4
  24. 24.Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a Model of Facial Shape and Expression From 4D Scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6):194:1–194:17, 2017. 2, 3
  25. 25.Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3207–3216, 2020. 8
  26. 26.Siyou Lin, Hongwen Zhang, Zerong Zheng, Ruizhi Shao, and Yebin Liu. Learning implicit templates for point-based clothed human modeling. In ECCV, 2022. 3
  27. 27.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. Advances in Neural Information Processing Systems, 33:15651–15663, 2020. 3
  28. 28.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural Actor: Neural Free-view Synthesis of Human Actors with Pose Control. ACM Transactions on Graphics (TOG), 40(6):1–16, 2021. 3
  29. 29.Abhimitra Meka, Rohit Pandey, Christian Haene, Sergio Orts-Escolano, Peter Barnum, Philip David-Son, Daniel Erickson, Yinda Zhang, Jonathan Taylor, Sofien Bouaziz, et al. Deep Relightable Textures: Volumetric Performance Capture With Neural Rendering. ACM Transactions on Graphics (TOG), 39(6):1–21, 2020. 1
  30. 30.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes As Neural Radiance Fields for View Synthesis. Communications of the ACM, 65(1):99–106, 2021. 2, 4, 8
  31. 31.Sergio Orts-Escolano, Christoph Rhemann, Sean Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles Loop, Qin Cai, Philip Chou, Sarah Mennicken, Julien Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, and Shahram Izadi. Holoportation: Virtual 3D Teleportation in Real-Time. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology (UIST). ACM, Oct. 2016. 1
  32. 32.Rohit Pandey, Sergio Orts Escolano, Chloe Legendre, Christian Haene, Sofien Bouaziz, Christoph Rhemann, Paul Debevec, and Sean Fanello. Total relighting: learning to relight portraits for background replacement. ACM Transactions on Graphics (TOG), 40(4):1–21, 2021. 3, 5
  33. 33.Keunhong Park, Utkarsh Sinha, Jonathan T Barron, Sofien Bouaziz, Dan B Goldman, Steven M Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable Neural Radiance Fields. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5865–5874, 2021. 5
  34. 34.Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural Body: Implicit Neural Representations With Structured Latent Codes for Novel View Synthesis of Dynamic Humans. pages 9054–9063, 2021. 3, 4
  35. 35.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics: A large-scale video dataset for forgery detection in human faces. arXiv preprint arXiv:1803.09179, 2018. 8
  36. 36.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1–11, 2019. 8
  37. 37.Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Hao Li, and Angjoo Kanazawa. PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, Oct. 2019. 1
  38. 38.Aliaksandr Siarohin, Stephane Lathuilı`ere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. Animating arbitrary objects via deep motion transfer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2377–2386, 2019. 3
  39. 39.Aliaksandr Siarohin, Stephane Lathuilı`ere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First Order Motion Model for Image Animation. In Advances in Neural Information Processing Systems (NeurIPS), volume 32, 2019. 3, 5, 6
  40. 40.Ayush Tewari, Florian Bernard, Pablo Garrido, Gaurav Bharaj, Mohamed Elgharib, Hans-Peter Seidel, Patrick Perez, Michael Zollhofer, and Christian Theobalt. FML: Face Model Learning From Videos. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10812–10822, 2019. 2
  41. 41.Ruben Tolosana, Ruben Vera-Rodriguez, Julian Fierrez, Aythami Morales, and Javier Ortega-Garcia. Deepfakes and beyond: A survey of face manipulation and fake detection. Information Fusion, 64:131–148, 2020. 8
  42. 42.Zach Waggoner. My Avatar, My Self: Identity in Video RolePlaying Games. McFarland, 2009. 1
  43. 43.Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Jan Kautz, and Bryan Catanzaro. Few-Shot Video-to-Video Synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 3
  44. 44.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image Quality Assessment: From Error Visibility to Structural Similarity. 13(4):600–612, 2004. 5
  45. 45.Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Humannerf: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16210–16220, 2022. 4
  46. 46.Olivia Wiles, A Koepke, and Andrew Zisserman. X2face: A network for controlling face generation using images, audio, and pose codes. In Proceedings of the European conference on Computer Vision (ECCV), pages 670–686, 2018. 3
  47. 47.Bangbang Yang, Chong Bao, Junyi Zeng, Hujun Bao, Yinda Zhang, Zhaopeng Cui, and Guofeng Zhang. Neumesh: Learning disentangled neural mesh-based implicit field for geometry and texture editing. In European Conference on Computer Vision, pages 597–614. Springer, 2022. 4
  48. 48.Haotian Yang, Hao Zhu, Yanru Wang, Mingkai Huang, Qiu Shen, Ruigang Yang, and Xun Cao. FaceScape: A LargeScale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction. pages 601–610, 2020. 2
  49. 49.Egor Zakharov, Aleksei Ivakhnenko, Aliaksandra Shysheya, and Victor Lempitsky. Fast bi-layer neural synthesis of oneshot realistic head avatars. In European Conference on Computer Vision, pages 524–540. Springer, 2020. 3
  50. 50.Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. Few-Shot Adversarial Learning of Realistic Neural Talking Head Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9459–9468, 2019. 3
  51. 51.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The Unreasonable Effectiveness of Deep Features As a Perceptual Metric. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 586–595, 2018. 5
  52. 52.Jian Zhao and Hui Zhang. Thin-plate spline motion model for image animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3657–3666, 2022. 5, 6
  53. 53.Yufeng Zheng, Victoria Fernández Abrevaya, Marcel C Bühler, Xu Chen, Michael J Black, and Otmar Hilliges. Im avatar: Implicit morphable head avatars from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13545–13555, 2022. 3, 4, 5, 6, 8
  54. 54.Zerong Zheng, Han Huang, Tao Yu, Hongwen Zhang, Yandong Guo, and Yebin Liu. Structured Local Radiance Fields for Human Avatar Modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15893–15903, 2022. 3, 4, 7
  55. 55.Michael Zollhöfer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick Pérez, Marc Stamminger, Matthias Nießner, and Christian Theobalt. State of the Art on Monocular 3D Face Reconstruction, Tracking, and Applications. volume 37, pages 523–550, 2018. 2

Citation

MLA
Bai, Z., et al. “Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos”. arXiv, 2023, http://arxiv.org/abs/2304.01436v1.
APA
Bai, Z., Tan, F., Huang, Z., Sarkar, K., Tang, D., Qiu, D., Meka, A., Du, R., Dou, M., Orts-Escolano, S., Pandey, R., Tan, P., Beeler, T., Fanello, S., & Zhang, Y. (2023). Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos. arXiv. http://arxiv.org/abs/2304.01436v1
Chicago
Bai, Z., F. Tan, Z. Huang, et al. 2023. “Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos”. arXiv. http://arxiv.org/abs/2304.01436v1.
Harvard
Bai, Z. et al. (2023) “Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.01436v1.
Vancouver
1. Bai Z, Tan F, Huang Z, et al (2023) Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos. arXiv

BibTeX

@article{bai2023learning,
  title = {Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos},
  author = {Bai, Ziqian and Tan, Feitong and Huang, Zeng and Sarkar, Kripasindhu and Tang, Danhang and Qiu, Di and Meka, Abhimitra and Du, Ruofei and Dou, Mingsong and Orts-Escolano, Sergio and Pandey, Rohit and Tan, Ping and Beeler, Thabo and Fanello, Sean and Zhang, Yinda},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.01436v1},
  eprint = {2304.01436}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE