RigNeRF: Fully Controllable Neural 3D Portraits

ShahRukh AtharZexiang XuKalyan SunkavalliEli ShechtmanZhixin Shu

article2022CVPR180 citations

Combines 3D morphable face models with neural radiance fields to enable precise control over rigid head poses and non-rigid facial expressions using only a short casual smartphone video.

Listen

Photo-realistic editing of human portraits with full control over head orientation, facial expressions, and camera viewpoints is a critical capability for immersive media, virtual reality, and digital video production. Standard neural radiance fields synthesize high-quality 3D views of static environments but cannot edit dynamic objects. Meanwhile, existing dynamic face-modeling methods struggle to render novel camera angles, distort full 3D backgrounds, or lose natural head rigidity during reanimation.

The article demonstrates RigNeRF, a volumetric neural rendering system that enables simultaneous, explicit control over head pose, facial expression, and novel camera viewpoints for an entire portrait scene. The primary objective is to learn these capabilities directly from a single, short video captured on a standard consumer device without requiring specialized multi-camera rigs.

To achieve this, the approach combines deformable neural radiance fields with a 3D morphable face model (3DMM). The 3DMM provides a structural prior that defines coarse rigid and non-rigid head movements, while a neural network predicts residual offsets to capture fine details omitted by standard face meshes, such as hair, teeth, and glasses. The model was evaluated using smartphone video captures of multiple subjects enacting varied expressions and head motions across 40- to 70-second sequences (approximately 1,200 to 2,100 frames).

The evaluations show that RigNeRF consistently outperforms leading dynamic rendering baselines—including HyperNeRF, NerFACE, and First Order Motion Model—across visual quality, perceptual distance, and facial reconstruction accuracy. Specifically, RigNeRF achieved higher peak signal-to-noise ratios (ranging from 27.0 to 29.55 dB) and significantly lower face reconstruction errors across all test subjects. Naively applying expression and pose conditioning to dynamic fields caused unnatural warping and visual distortions, whereas RigNeRF’s deformation prior maintained structural rigidity while generalizing to unobserved poses and expressions.

These findings indicate that 3D geometric priors provide an effective inductive bias for neural volumetric rendering, lowering production barriers and capture costs for high-fidelity avatar generation. By enabling high-accuracy reanimation and free-viewpoint rendering from casual smartphone video, the system demonstrates strong potential for low-cost, high-performance content creation pipelines in entertainment and telepresence.

Organizations evaluating this technology should focus on piloting the system in controlled video production or avatar animation workflows to validate performance across broader operational use cases. Because photorealistic face reanimation introduces risks related to synthetic media misuse, deployment should incorporate authentication protocols, watermarking, or other media compliance measures.

Key limitations include that the current implementation is subject-specific, requiring an individual network to be trained for each new subject and scene. Rendering fidelity also depends heavily on precise camera pose tracking during the capture phase. Confidence in the evaluated metrics is strong for standard monocular captures, though users should exercise caution when extrapolating results to scenes with extreme lighting changes or rapid, uncalibrated camera movements.

arXiv: 2206.06481
Cover for RigNeRF: Fully Controllable Neural 3D Portraits

Abstract

Volumetric neural rendering methods, such as neural radiance fields (NeRFs), have enabled photo-realistic novel view synthesis. However, in their standard form, NeRFs do not support the editing of objects, such as a human head, within a scene. In this work, we propose RigNeRF, a system that goes beyond just novel view synthesis and enables full control of head pose and facial expressions learned from a single portrait video. We model changes in head pose and facial expressions using a deformation field that is guided by a 3D morphable face model (3DMM). The 3DMM effectively acts as a prior for RigNeRF that learns to predict only residuals to the 3DMM deformations and allows us to render novel (rigid) poses and (non-rigid) expressions that were not present in the input sequence. Using only a smartphone-captured short video of a subject for training, we demonstrate the effectiveness of our method on free view synthesis of a portrait scene with explicit head pose and expression controls.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. RigNeRF
  • 3.1. Deformable Neural Radiance Fields
  • 3.2. A 3DMM-guided deformation field
  • 3.3. 3DMM-conditioned Appearance
  • 4. Results
  • 4.1. Evaluation on Test Data
  • 4.2. Reanimation with pose and expression control
  • 4.3. Comparison with HyperNeRF+E/P
  • 5. Limitations and Conclusion
  • 6. Acknowledgements
  • References

Knowls

  1. Knowl 1 — RigNeRF 3DMM-Guided Deformable Neural Radiance Field Architecture

    model/method

    RigNeRF models fully controllable dynamic 3D portrait scenes (head pose, facial expressions, and camera viewpoint) using a continuous volumetric representation trained from a monocular smartphone video. The system comprises two main components:

    1. Deformation Module: Maps sampled 3D points x=(x,y,z)∈R3\mathbf{x} = (x, y, z) \in \mathbb{R}^3 along camera rays from an observation space (at frame ii with head pose βi,pose\beta_{i,\text{pose}} and expression βi,exp\beta_{i,\text{exp}}) to a canonical space xcan=(x′,y′,z′)∈R3\mathbf{x}_{\text{can}} = (x', y', z') \in \mathbb{R}^3, where the head is in a standardized zero head-pose and neutral facial expression configuration.
    2. Appearance/Radiance Field Module: A Multi-Layer Perceptron (MLP) FF that operates in the canonical space to predict volume density σ\sigma and view-dependent color c=(r,g,b)\mathbf{c} = (r, g, b) conditioned on viewing direction, appearance embeddings, deformation feature vectors, and 3DMM expression/pose parameters.

    By leveraging a 3D Morphable Face Model (3DMM) prior, rigid head movements and non-rigid facial deformations are explicitly guided, while an MLP learns corrective residual displacements to capture dynamic facial details, hair, glasses, and accessories.

  2. Knowl 2 — 3DMM Spatial Deformation Field Formulation

    equation

    To define a continuous 3D deformation prior across all points in space from discrete 3D Morphable Model (3DMM) mesh vertices, RigNeRF derives a dense 3DMM deformation field. For expression parameters βexp\beta_{\text{exp}} and head-pose parameters βpose\beta_{\text{pose}}, the 3DMM deformation at an arbitrary 3D spatial coordinate x=(x,y,z)∈R3\mathbf{x} = (x, y, z) \in \mathbb{R}^3 is defined as:

    3DMMDef(x,βexp,βpose)=3DMMDef(x^,βexp,βpose)exp⁡(DistToMesh(x))\text{3DMMDef}(\mathbf{x}, \beta_{\text{exp}}, \beta_{\text{pose}}) = \frac{\text{3DMMDef}(\hat{\mathbf{x}}, \beta_{\text{exp}}, \beta_{\text{pose}})}{\exp(\text{DistToMesh}(\mathbf{x}))}

    where x^=(x^,y^,z^)∈R3\hat{\mathbf{x}} = (\hat{x}, \hat{y}, \hat{z}) \in \mathbb{R}^3 is the closest point to x\mathbf{x} on the articulated face mesh, and DistToMesh(x)=∥x−x^∥2\text{DistToMesh}(\mathbf{x}) = \|\mathbf{x} - \hat{\mathbf{x}}\|_2 denotes the Euclidean distance from x\mathbf{x} to the mesh surface.

    The mesh surface deformation 3DMMDef(x^,βexp,βpose)\text{3DMMDef}(\hat{\mathbf{x}}, \beta_{\text{exp}}, \beta_{\text{pose}}) is given by the difference between the coordinate of x^\hat{\mathbf{x}} in the canonical space (zero head pose and neutral expression in FLAME) and its articulated configuration:

    3DMMDef(x^,βexp,βpose)=x^FLAME(0,0)−x^FLAME(βexp,βpose)\text{3DMMDef}(\hat{\mathbf{x}}, \beta_{\text{exp}}, \beta_{\text{pose}}) = \hat{\mathbf{x}}_{\text{FLAME}(\mathbf{0}, \mathbf{0})} - \hat{\mathbf{x}}_{\text{FLAME}(\beta_{\text{exp}}, \beta_{\text{pose}})}

    This inverse exponential distance weighting attenuates the 3DMM prior smoothly away from the head mesh into empty space and the background.

  3. Knowl 3 — Residual-Refined Deformation Field with Low-Frequency 3DMM Conditioning

    equation

    The complete deformation field D^(x)\hat{D}(\mathbf{x}) mapping a point x\mathbf{x} in the observation frame to the canonical space point xcan\mathbf{x}_{\text{can}} is defined as the sum of the coarse 3DMM deformation field and a corrective residual field predicted by a deformation MLP DD:

    D^(x)=3DMMDef(x,βi,exp,βi,pose)+D(γa(x),γb(3DMMDef(x,βi,exp,βi,pose)),ωi)\hat{D}(\mathbf{x}) = \text{3DMMDef}(\mathbf{x}, \beta_{i,\text{exp}}, \beta_{i,\text{pose}}) + D(\pmb{\gamma}_a(\mathbf{x}), \pmb{\gamma}_b(\text{3DMMDef}(\mathbf{x}, \beta_{i,\text{exp}}, \beta_{i,\text{pose}})), \omega_i)

    xcan=x+D^(x)\mathbf{x}_{\text{can}} = \mathbf{x} + \hat{D}(\mathbf{x})

    where:

    • βi,exp\beta_{i,\text{exp}} and βi,pose\beta_{i,\text{pose}} are the 3DMM expression and pose parameters of frame ii.
    • ωi\omega_i is a per-frame latent deformation code used to account for residual motions not captured by 3DMM parameters.
    • γa(⋅)\pmb{\gamma}_a(\cdot) and γb(⋅)\pmb{\gamma}_b(\cdot) are positional encodings. For the 3-dimensional 3DMM displacement vector 3DMMDef(x,… )\text{3DMMDef}(\mathbf{x}, \dots), using b=2b = 2 frequency bands in γb\pmb{\gamma}_b yields optimal performance.

    Conditioning the deformation MLP DD on the positional encoding of the 3D displacement vector 3DMMDef\text{3DMMDef} rather than directly on the raw 59-dimensional expression and pose parameter vector prevents overfitting and improves generalization to novel expressions and poses.

  4. Knowl 4 — 3DMM-Conditioned Radiance and Appearance Field Formulation

    equation

    To render fine-grained expression- and pose-dependent details (such as teeth and wrinkles) in the canonical volume, the color MLP FF is conditioned on the canonical coordinate, viewing direction, per-frame appearance embedding, internal deformation features, and 3DMM expression and pose vectors:

    (c(x,d),σ(x))=F(γc(xcan),γd(d),ϕi,DF,i(xcan),βi,exp,βi,pose)(\mathbf{c}(\mathbf{x}, \mathbf{d}), \sigma(\mathbf{x})) = F(\pmb{\gamma}_c(\mathbf{x}_{\text{can}}), \pmb{\gamma}_d(\mathbf{d}), \phi_i, D_{F,i}(\mathbf{x}_{\text{can}}), \beta_{i,\text{exp}}, \beta_{i,\text{pose}})

    where:

    • xcan∈R3\mathbf{x}_{\text{can}} \in \mathbb{R}^3 is the canonicalized 3D coordinate obtained via deformation.
    • d∈S2\mathbf{d} \in \mathbb{S}^2 is the viewing ray direction.
    • γc\pmb{\gamma}_c and γd\pmb{\gamma}_d are standard sinusoidal positional encodings applied to xcan\mathbf{x}_{\text{can}} and d\mathbf{d}.
    • ϕi\phi_i is a per-frame latent appearance code for frame ii.
    • DF,i(xcan)D_{F,i}(\mathbf{x}_{\text{can}}) denotes activation feature representations extracted from the penultimate layer of the deformation MLP DD.
    • βi,exp\beta_{i,\text{exp}} and βi,pose\beta_{i,\text{pose}} are the frame's expression and pose coefficients.
    • c∈[0,1]3\mathbf{c} \in [0, 1]^3 is the predicted RGB color and σ≥0\sigma \ge 0 is the predicted volume density.
  5. Knowl 5 — Latent Deformation Optimization for Unseen Test Frame Evaluation

    algorithm

    To evaluate RigNeRF or related deformable NeRF baselines on held-out validation/test frames without utilizing the arbitrary deformation code of the first training frame, only the per-frame latent deformation code ω\omega is optimized while all network weights and appearance parameters remain frozen.

    Input: Test image ground-truth pixel values C^{GT}, camera ray parameters {x, d}, 3DMM parameters {\beta_{i,exp}, \beta_{i,pose}}, frozen appearance code \phi_0, frozen radiance field parameters \theta
    Output: Optimized validation deformation code \omega_v
    Initialize deformation code \omega
    for epoch = 1 to 200 do
        for each sampled pixel p with ray direction d and coordinates x do
            Predict color C_p(\omega; x, d, \theta, \phi_0, \beta_{i,exp}, \beta_{i,pose}) via volumetric rendering
        Compute loss L(\omega) = \sum_p ||C_p(\omega; x, d, \theta, \phi_0, \beta_{i,exp}, \beta_{i,pose}) - C_p^{GT}||_2^2
        Update \omega using gradient descent on \nabla_\omega L(\omega)
    end for
    return \omega_v = \omega

    Optimizing ω\omega for 200 epochs reaches the loss plateau, isolating the model's geometric deformation quality from mismatched per-frame latent codes.

  6. Knowl 6 — Monocular Portrait Video Capture and Preprocessing Protocol

    experimental setup

    RigNeRF models are trained on monocular smartphone video sequences captured using an iPhone XR or iPhone 12:

    • Capture Protocol: Video length is between 40 to 70 seconds (~1200 to 2100 frames). The capture is divided into two phases:
      1. Expression/Speech Phase: The subject performs a wide range of facial expressions and speech while attempting to keep their head stationary as the camera is panned around them.
      2. Head Rotation Phase: The camera is held fixed at head level while the subject rotates their head across various angles while performing varied facial expressions.
    • Preprocessing: Camera intrinsic and extrinsic parameters are estimated across the sequence using COLMAP. Per-frame FLAME 3DMM expression and shape parameters are extracted with DECA and refined using dense 3D landmark fitting alongside COLMAP camera poses.
    • Training Details: Video frames are downsampled to 256×256256 \times 256 resolution. Training employs coarse-to-fine positional encoding scheduling and vertex deformation regularization on the deformation MLP D(x,ωi)D(\mathbf{x}, \omega_i).
  7. Knowl 7 — Quantitative Evaluation of Novel View and Expression Synthesis on Held-Out Test Data

    data/table

    RigNeRF was evaluated on held-out test frames across four human subjects and compared against HyperNeRF (a dynamic NeRF baseline lacking explicit controls), NerFACE (a dynamic NeRF for facial avatars assuming static camera/background), and First Order Motion Model (FOMM, a 2D image animation baseline). Performance is measured by Peak Signal-to-Noise Ratio (PSNR, higher is better), Learned Perceptual Image Patch Similarity (LPIPS, lower is better), and Mean Squared Error localized to the face region (FaceMSE, lower is better).

    Subject 1 Subject 2 Subject 3 Subject 4
    Models PSNR ↑\uparrow LPIPS ↓\downarrow FaceMSE ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow FaceMSE ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow FaceMSE ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow FaceMSE ↓\downarrow
    RigNeRF (Ours) 29.55 0.136 9.6e-5 29.36 0.102 1e-4 28.39 0.109 8e-5 27.00 0.092 2.3e-4
    HyperNeRF 24.58 0.220 8.14e-4 22.55 0.1546 9.48e-4 19.29 0.260 2.74e-3 21.19 0.182 1.58e-3
    NerFACE 24.20 0.217 7.84e-4 24.57 0.1740 6.70e-4 28.00 0.1292 1.2e-4 28.47 0.134 2.7e-4
    FOMM 11.45 0.432 7.65e-3 12.70 0.5820 6.31e-3 10.17 0.6010 1.7e-2 11.17 0.529 6.8e-3

    RigNeRF consistently achieves lower LPIPS and lower FaceMSE across all four subjects, demonstrating superior reconstruction fidelity in the facial region compared to methods that lack deformation fields (NerFACE), lack explicit parametric controls (HyperNeRF), or lack 3D geometric consistency (FOMM).

  8. Knowl 8 — HyperNeRF+E/P Baseline Architecture and Rigidity Degradation

    model/method

    To evaluate the specific role of the 3DMM prior against learning deformations purely via an MLP conditioned on pose and expression parameters, an augmented baseline named HyperNeRF+E/P is constructed. The model modifies the HyperNeRF forward pass by appending expression and pose parameters {βi,exp,βi,pose}\{\beta_{i,\text{exp}}, \beta_{i,\text{pose}}\} to the inputs of both the deformation field DD and the radiance field FF:

    xcan=D(γa(x),βi,exp,βi,pose,ωi)\mathbf{x}_{\text{can}} = D(\pmb{\gamma}_a(\mathbf{x}), \beta_{i,\text{exp}}, \beta_{i,\text{pose}}, \omega_i)

    w=H(γ1(x),ωi)\mathbf{w} = H(\pmb{\gamma}_1(\mathbf{x}), \omega_i)

    (c(x,d),σ(x))=F(γc(xcan),γd(d),ϕi,w,βi,exp,βi,pose)(\mathbf{c}(\mathbf{x}, \mathbf{d}), \sigma(\mathbf{x})) = F(\pmb{\gamma}_c(\mathbf{x}_{\text{can}}), \pmb{\gamma}_d(\mathbf{d}), \phi_i, \mathbf{w}, \beta_{i,\text{exp}}, \beta_{i,\text{pose}})

    where HH is the ambient MLP and w\mathbf{w} represents ambient slicing coordinates.

    Without the explicit 3DMM geometric prior, HyperNeRF+E/P suffers from loss of head rigidity during reanimation under novel head poses and exhibits severe facial distortion artefacts due to overfitting on the high-dimensional conditioning space.

  9. Knowl 9 — Quantitative Comparison Between RigNeRF and HyperNeRF+E/P

    data/table

    A quantitative ablation comparing RigNeRF against the HyperNeRF+E/P baseline (HyperNeRF+Exp) on test data for Subject 1 and Subject 2:

    Subject 1 Subject 2
    Models PSNR ↑\uparrow LPIPS ↓\downarrow FaceMSE ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow FaceMSE ↓\downarrow
    RigNeRF (Ours) 29.55 0.136 9.6e-5 29.36 0.102 1e-4
    HyperNeRF+Exp 31.30 0.161 1.3e-4 30.00 0.116 1.9e-4

    While HyperNeRF+Exp achieves higher PSNR due to background/ambient coordinate capacity, RigNeRF achieves lower LPIPS and lower FaceMSE, verifying that the 3DMM-guided deformation field produces perceptually superior and geometrically more accurate reconstructions of the human face.

  10. Knowl 10 — Practical Limitations of RigNeRF

    limitation

    RigNeRF has three primary operational limitations:

    1. Subject-Specific Training: A separate network must be trained from scratch for each portrait scene; the weights do not generalize across different individuals without retraining.
    2. Training Video Length and Variety: The model requires 40-70 seconds of video that deliberately captures diverse combinations of camera views, expressions, and head poses.
    3. Dependence on Camera and 3DMM Tracking: Reconstruction and rendering quality are bounded by the accuracy of camera calibration (COLMAP) and 3DMM landmark/parameter extraction (DECA).

Coverage note — None was omitted; all key equations, architectural details, baselines (HyperNeRF+E/P), training setups, benchmark evaluation data, and stated limitations are fully covered.

References

  1. 1.ShahRukh Athar, Albert Pumarola, Francesc Moreno-Noguer, and Dimitris Samaras. Facedet3d: Facial expressions with 3d geometric detail prediction. arXiv preprint arXiv:2012.07999, 2020. 3
  2. 2.S Athar, Z Shu, and D Samaras. Self-supervised deformation modeling for facial expression editing. 2020. 3
  3. 3.Mojtaba Bemana, Karol Myszkowski, Hans-Peter Seidel, and Tobias Ritschel. X-fields: Implicit neural view-, light-and time-image interpolation. 2020. 2
  4. 4.Volker Blanz, Thomas Vetter, et al. A morphable model for the synthesis of 3d faces. 1999. 1, 2, 3
  5. 5.Aggelina Chatziagapi, ShahRukh Athar, Francesc Moreno-Noguer, and Dimitris Samaras. Sider: Single-image neural optimization for facial geometric detail recovery. arXiv preprint arXiv:2108.05465, 2021. 3
  6. 6.Julian Chibane, Aayush Bansal, Verica Lazova, and Gerard Pons-Moll. Stereo radiance fields (srf): Learning view synthesis for sparse views of novel scenes. In CVPR, 2021. 2
  7. 7.Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR, 2018. 3
  8. 8.Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha. Stargan v2: Diverse image synthesis for multiple domains. In CVPR, 2020. 3
  9. 9.Yu Deng, Jiaolong Yang, Dong Chen, Fang Wen, and Xin Tong. Disentangled and controllable face image generation via 3d imitative-contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5154–5163, 2020. 3
  10. 10.M. Doukas, Mohammad Rami Koujan, V. Sharmanska, A. Roussos, and S. Zafeiriou. Head2head++: Deep facial attributes re-targeting. IEEE Transactions on Biometrics, Behavior, and Identity Science, 3:31–43, 2021. 3
  11. 11.Yao Feng, Haiwen Feng, Michael J. Black, and Timo Bolkart. Learning an animatable detailed 3D face model from in-the-wild images. volume 40, 2021. 3, 5, 7
  12. 12.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In CVPR, June 2021. 3, 5, 6, 7, 8
  13. 13.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction, 2020. 2, 3, 6
  14. 14.Chen Gao, Yichang Shih, Wei-Sheng Lai, Chia-Kai Liang, and Jia-Bin Huang. Portrait neural radiance fields from a single image. arXiv preprint arXiv:2012.05903, 2020. 2
  15. 15.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In NeurIPS, 2014. 2
  16. 16.Jianzhu Guo, Xiangyu Zhu, Yang Yang, Fan Yang, Zhen Lei, and Stan Z Li. Towards fast, accurate and stable 3d dense face alignment. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 3, 5, 7
  17. 17.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In CVPR, 2017. 2
  18. 18.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019. 2
  19. 19.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019. 2
  20. 20.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In CVPR, 2020. 2
  21. 21.H. Kim, P. Garrido, A. Tewari, Weipeng Xu, Justus Thies, M. Nießner, Patrick Perez, C. Richardt, M. Zollhofer, and C. Theobalt. Deep video portraits. ACM Transactions on Graphics (TOG), 37:1 – 14, 2018. 2
  22. 22.Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Perez, Christian Richardt, Michael Zollhofer, and Christian Theobalt. Deep video portraits. ACM TOG, 2018. 3
  23. 23.M. Koujan, M. Doukas, A. Roussos, and S. Zafeiriou. Head2head: Video-based neural head synthesis. In 2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020) (FG), pages 319–326, Los Alamitos, CA, USA, may 2020. IEEE Computer Society. 3
  24. 24.Marek Kowalski, Stephan J Garbin, Virginia Estellers, Tadas Baltrusaitis, Matthew Johnson, and Jamie Shotton. Config: Controllable neural face image generation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16, pages 299–315. Springer, 2020. 3
  25. 25.Christoph Lassner and Michael ZollhA˜ ¶fer. Pulsar: Efficient sphere-based neural rendering. In CVPR, 2021. 2
  26. 26.Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6), 2017. 3
  27. 27.Tianye Li, Mira Slavcheva, Michael Zollhoefer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, and Zhaoyang Lv. Neural 3d video synthesis, 2021. 2
  28. 28.Zhengqi Li, Simon Niklaus, Noah Snavely, and Oliver Wang. Neural scene flow fields for space-time view synthesis of dynamic scenes. In CVPR, 2021. 2
  29. 29.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. 2020. 2
  30. 30.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose control. ACM TOG, 2021. 3
  31. 31.Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mixture of volumetric primitives for efficient neural rendering, 2021. 2
  32. 32.Ricardo Martin-Brualla, Noha Radwan, Mehdi SM Sajjadi, Jonathan T Barron, Alexey Dosovitskiy, and Daniel Duckworth. Nerf in the wild: Neural radiance fields for unconstrained photo collections. arXiv:2008.02268, 2020. 2
  33. 33.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. 2020. 2, 3
  34. 34.Michael Oechsle, Songyou Peng, and Andreas Geiger. Unisurf: Unifying neural implicit surfaces and radiance fields for multi-view reconstruction. arXiv preprint arXiv:2104.10078, 2021. 2
  35. 35.Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Steven M. Seitz, and Ricardo Martin-Brualla. Nerfies: Deformable neural radiance fields. ICCV, 2021. 2, 3, 5
  36. 36.Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M. Seitz. Hypernerf: A higher-dimensional representation for topologically varying neural radiance fields. arXiv preprint arXiv:2106.13228, 2021. 2, 3, 5, 6, 7, 8
  37. 37.Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Animatable neural radiance fields for modeling dynamic human bodies. In ICCV, 2021. 3
  38. 38.Albert Pumarola, Antonio Agudo, Aleix M Martinez, Alberto Sanfeliu, and Francesc Moreno-Noguer. Ganimation: One-shot anatomically consistent facial animation. International Journal of Computer Vision, 128(3):698–713, 2020. 3
  39. 39.Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural Radiance Fields for Dynamic Scenes. In CVPR, 2021. 2, 3
  40. 40.Gernot Riegler and Vladlen Koltun. Stable view synthesis. In CVPR, 2021. 2
  41. 41.Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In CVPR, 2016. 5
  42. 42.Zhixin Shu, Mihir Sahasrabudhe, Riza Alp Guler, Dimitris Samaras, Nikos Paragios, and Iasonas Kokkinos. Deforming autoencoders: Unsupervised disentangling of shape and appearance. In ECCV, 2018. 3
  43. 43.Z. Shu, E. Yumer, S. Hadap, K. Sunkavalli, E. Shechtman, and D. Samaras. Neural face editing with intrinsic image disentangling. In CVPR, 2017. 3
  44. 44.Aliaksandr Siarohin, Stephane Lathuiliere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First order motion model for image animation. 2019. 5, 6, 7
  45. 45.Vincent Sitzmann, Michael Zollhofer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. 2019. 2
  46. 46.Ayush Tewari, Mohamed Elgharib, Florian Bernard, Hans-Peter Seidel, Patrick Perez, Michael Zollhofer, and Christian Theobalt. Pie: Portrait image embedding for semantic control. ACM Transactions on Graphics (TOG), 39(6):1–14, 2020. 3
  47. 47.Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj, Florian Bernard, Hans-Peter Seidel, Patrick Perez, Michael Zollhofer, and Christian Theobalt. Stylerig: Rigging stylegan for 3d control over portrait images, cvpr 2020. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, june 2020. 2, 3
  48. 48.Justus Thies, M. Zollhofer, and M. Nießner. Deferred neural rendering. ACM Transactions on Graphics (TOG), 2019. 2
  49. 49.Suttisak Wizadwongsa, Pakkapon Phongthawee, Jiraphon Yenphraphai, and Supasorn Suwajanakorn. Nex: Real-time view synthesis with neural basis expansion. In CVPR, 2021. 2
  50. 50.Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. Space-time neural irradiance fields for free-viewpoint video. In CVPR, 2021. 2
  51. 51.Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neural surface reconstruction by disentangling geometry and appearance. NIPS, 33, 2020. 2
  52. 52.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv:2010.07492, 2020. 2
  53. 53.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv:2010.07492, 2020. 2
  54. 54.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 8
  55. 55.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017. 2

Citation

MLA
Athar, S., et al. “RigNeRF: Fully Controllable Neural 3D Portraits”. arXiv, 2022, http://arxiv.org/abs/2206.06481v1.
APA
Athar, S., Xu, Z., Sunkavalli, K., Shechtman, E., & Shu, Z. (2022). RigNeRF: Fully Controllable Neural 3D Portraits. arXiv. http://arxiv.org/abs/2206.06481v1
Chicago
Athar, S., Z. Xu, K. Sunkavalli, E. Shechtman, and Z. Shu. 2022. “RigNeRF: Fully Controllable Neural 3D Portraits”. arXiv. http://arxiv.org/abs/2206.06481v1.
Harvard
Athar, S. et al. (2022) “RigNeRF: Fully Controllable Neural 3D Portraits”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.06481v1.
Vancouver
1. Athar S, Xu Z, Sunkavalli K, Shechtman E, Shu Z (2022) RigNeRF: Fully Controllable Neural 3D Portraits. arXiv

BibTeX

@article{athar2022rignerf,
  title = {RigNeRF: Fully Controllable Neural 3D Portraits},
  author = {Athar, ShahRukh and Xu, Zexiang and Sunkavalli, Kalyan and Shechtman, Eli and Shu, Zhixin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.06481v1},
  eprint = {2206.06481}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE