Nerfies: Deformable Neural Radiance Fields

Keunhong ParkUtkarsh SinhaJonathan T. BarronSofien BouazizDan B GoldmanSteven M. SeitzRicardo Martin-Brualla

article2020ICCV1,540 citations

Proposes an extension to neural radiance fields that integrates a continuous deformation field with elastic regularization, enabling photorealistic free-viewpoint rendering of non-rigid, deforming subjects captured casually on mobile phones.

Listen

High-quality three-dimensional scanning of humans has traditionally required expensive specialized studio laboratories with multiple synchronized cameras and controlled lighting rigs. Standard reconstruction algorithms fail when applied to casual hand-held mobile phone videos because subjects inevitably move, and fine features such as hair, glasses, and jewelry break core assumptions about rigid geometry. The article addresses this accessibility barrier by introducing a method to generate photorealistic, deformable three-dimensional models—termed "nerfies"—from casual mobile phone videos without requiring domain-specific templates or specialized multi-camera hardware.

The main objective of the article is to demonstrate and evaluate a system that reconstructs non-rigidly deforming scenes by extending Neural Radiance Fields (NeRF) with an optimized continuous volumetric deformation field. To achieve this, the authors decompose a dynamic scene into a canonical static three-dimensional template and a per-observation deformation field modeled by a neural network. To prevent distortions and stabilize training, they introduce three core innovations: a rigid motion parameterization, an elastic energy regularization that penalizes deviations from local rigidity, and a coarse-to-fine optimization strategy that gradually introduces high-frequency details. The method was evaluated using monocular RGB image sequences captured on mobile phones, with a dual-phone synchronized rig used to collect ground-truth novel viewpoints for benchmarking.

The core evaluation demonstrates several key findings. First, the proposed approach consistently outperforms baseline methods in perceptual visual quality across both quasi-static and dynamic scenes, achieving substantially better perceptual error scores on unseen viewpoints. Second, the elastic regularization successfully eliminates geometric distortion artifacts in under-constrained capture scenarios, such as when a user sweeps the camera predominantly across one side of their face. Third, the coarse-to-fine optimization scheme proves essential for resolving complex motions; it avoids suboptimal registrations while capturing subtle expressions like smiles alongside large head turns. Finally, the framework generalizes beyond human portraits to dynamic objects and full-body scans without modifying the underlying architecture.

These findings indicate that accessible, consumer-grade hardware can replace specialized laboratory rigs for high-fidelity three-dimensional human modeling, significantly reducing capture costs and technical friction. For product teams and technical leaders, this approach unlocks practical pipelines for virtual avatars, telepresence, and interactive media directly from standard smartphone cameras. Furthermore, the volumetric elastic prior and coarse-to-fine annealing establish an effective blueprint for regularizing under-constrained coordinate-based neural representations.

Organizations seeking to implement this technology should pilot casual video capture workflows while accounting for current computational requirements, which involve intensive training times across multiple high-performance graphics processors. Next development efforts should focus on optimizing training and inference speeds, handling topological changes like opening and closing mouths, and developing robust handling for rapid subject motions and orientation flips.

Cover for Nerfies: Deformable Neural Radiance Fields

Abstract

We present the first method capable of photorealistically reconstructing deformable scenes using photos/videos captured casually from mobile phones. Our approach augments neural radiance fields (NeRF) by optimizing an additional continuous volumetric deformation field that warps each observed point into a canonical 5D NeRF. We observe that these NeRF-like deformation fields are prone to local minima, and propose a coarse-to-fine optimization method for coordinate-based models that allows for more robust optimization. By adapting principles from geometry processing and physical simulation to NeRF-like models, we propose an elastic regularization of the deformation field that further improves robustness. We show that our method can turn casually captured selfie photos/videos into deformable NeRF models that allow for photorealistic renderings of the subject from arbitrary viewpoints, which we dub "nerfies." We evaluate our method by collecting time-synchronized data using a rig with two mobile phones, yielding train/validation images of the same pose at different viewpoints. We show that our method faithfully reconstructs non-rigidly deforming scenes and reproduces unseen views with high fidelity.

Citation

MLA
Park, K., et al. “Nerfies: Deformable Neural Radiance Fields”. arXiv, 2020, http://arxiv.org/abs/2011.12948v5.
APA
Park, K., Sinha, U., Barron, J. T., Bouaziz, S., Goldman, D. B., Seitz, S. M., & Martin-Brualla, R. (2020). Nerfies: Deformable Neural Radiance Fields. arXiv. http://arxiv.org/abs/2011.12948v5
Chicago
Park, K., U. Sinha, J. T. Barron, et al. 2020. “Nerfies: Deformable Neural Radiance Fields”. arXiv. http://arxiv.org/abs/2011.12948v5.
Harvard
Park, K. et al. (2020) “Nerfies: Deformable Neural Radiance Fields”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2011.12948v5.
Vancouver
1. Park K, Sinha U, Barron JT, Bouaziz S, Goldman DB, Seitz SM, Martin-Brualla R (2020) Nerfies: Deformable Neural Radiance Fields. arXiv

BibTeX

@article{park2020nerfies,
  title = {Nerfies: Deformable Neural Radiance Fields},
  author = {Park, Keunhong and Sinha, Utkarsh and Barron, Jonathan T. and Bouaziz, Sofien and Goldman, Dan B and Seitz, Steven M. and Martin-Brualla, Ricardo},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2011.12948v5},
  eprint = {2011.12948}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE