Nerfies: Deformable Neural Radiance Fields
Keunhong ParkUtkarsh SinhaJonathan T. BarronSofien BouazizDan B GoldmanSteven M. SeitzRicardo Martin-Brualla
Proposes an extension to neural radiance fields that integrates a continuous deformation field with elastic regularization, enabling photorealistic free-viewpoint rendering of non-rigid, deforming subjects captured casually on mobile phones.
High-quality three-dimensional scanning of humans has traditionally required expensive specialized studio laboratories with multiple synchronized cameras and controlled lighting rigs. Standard reconstruction algorithms fail when applied to casual hand-held mobile phone videos because subjects inevitably move, and fine features such as hair, glasses, and jewelry break core assumptions about rigid geometry. The article addresses this accessibility barrier by introducing a method to generate photorealistic, deformable three-dimensional models—termed "nerfies"—from casual mobile phone videos without requiring domain-specific templates or specialized multi-camera hardware.
The main objective of the article is to demonstrate and evaluate a system that reconstructs non-rigidly deforming scenes by extending Neural Radiance Fields (NeRF) with an optimized continuous volumetric deformation field. To achieve this, the authors decompose a dynamic scene into a canonical static three-dimensional template and a per-observation deformation field modeled by a neural network. To prevent distortions and stabilize training, they introduce three core innovations: a rigid motion parameterization, an elastic energy regularization that penalizes deviations from local rigidity, and a coarse-to-fine optimization strategy that gradually introduces high-frequency details. The method was evaluated using monocular RGB image sequences captured on mobile phones, with a dual-phone synchronized rig used to collect ground-truth novel viewpoints for benchmarking.
The core evaluation demonstrates several key findings. First, the proposed approach consistently outperforms baseline methods in perceptual visual quality across both quasi-static and dynamic scenes, achieving substantially better perceptual error scores on unseen viewpoints. Second, the elastic regularization successfully eliminates geometric distortion artifacts in under-constrained capture scenarios, such as when a user sweeps the camera predominantly across one side of their face. Third, the coarse-to-fine optimization scheme proves essential for resolving complex motions; it avoids suboptimal registrations while capturing subtle expressions like smiles alongside large head turns. Finally, the framework generalizes beyond human portraits to dynamic objects and full-body scans without modifying the underlying architecture.
These findings indicate that accessible, consumer-grade hardware can replace specialized laboratory rigs for high-fidelity three-dimensional human modeling, significantly reducing capture costs and technical friction. For product teams and technical leaders, this approach unlocks practical pipelines for virtual avatars, telepresence, and interactive media directly from standard smartphone cameras. Furthermore, the volumetric elastic prior and coarse-to-fine annealing establish an effective blueprint for regularizing under-constrained coordinate-based neural representations.
Organizations seeking to implement this technology should pilot casual video capture workflows while accounting for current computational requirements, which involve intensive training times across multiple high-performance graphics processors. Next development efforts should focus on optimizing training and inference speeds, handling topological changes like opening and closing mouths, and developing robust handling for rapid subject motions and orientation flips.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Nerfies directly builds upon Neural Radiance Fields (NeRF) by augmenting its continuous 5D scene representation with a volumetric deformation field for non-rigid scenes.
- Paper: Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains, Matthew Tancik et al. (2020). This work establishes the theoretical foundation and positional/Fourier encoding strategies that coordinate-based neural representations, including Nerfies' canonical and deformation fields, rely on to fit high-frequency geometry.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). DeepSDF introduces continuous neural implicit fields conditioned on latent codes, providing foundational architectural concepts for continuous 3D field optimization.
- Paper: D-NeRF: neural radiance fields for dynamic scenes, Albert Pumarola et al. (2021). D-NeRF develops a concurrent and complementary formulation for dynamic neural radiance fields using canonical space mappings driven by time.
- Paper: 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering, Guanjun Wu et al. (2023). 4D Gaussian Splatting extends dynamic scene reconstruction to explicit Gaussian primitives coupled with neural deformation fields to achieve real-time rendering speeds.
- Paper: NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections, Ricardo Martin-Brualla et al. (2021). NeRF in the Wild tackles unconstrained in-the-wild captures by learning transient elements and appearance embeddings alongside static neural radiance fields.
- Paper: Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields, Jonathan T. Barron et al. (2021). Mip-NeRF introduces integrated positional encodings along cone frustums to resolve multiscale aliasing issues present in baseline coordinate networks.
- Paper: Instant neural graphics primitives with a multiresolution hash encoding, Thomas Müller et al. (2022). Instant NGP introduces multiresolution hash encodings that drastically accelerate training and inference for neural graphics primitives such as radiance and deformation fields.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). Instruct-NeRF2NeRF advances neural scene manipulation by integrating 2D diffusion models to apply instruction-guided edits directly across 3D radiance fields.
