FENeRF: Face Editing in Neural Radiance Fields
Jingxiang SunXuan WangYong ZhangXiaoyu LiQi ZhangYebin LiuJue Wang
Proposes a 3D-aware face generator that couples neural radiance fields with decoupled semantic and texture latent spaces, enabling view-consistent portrait synthesis alongside precise local attribute editing trained solely on monocular image-mask pairs.
Synthesizing photo-realistic human portraits with computer graphics has rapidly evolved, yet existing methods force a compromise between editing capability and 3D geometric realism. Standard 2D image generators allow flexible local attribute editing but produce severe visual distortions when rendering the subject from different camera angles. Conversely, recent 3D-aware methods maintain strict view consistency across angles but lack the ability to support interactive, local adjustments to specific facial features. The article introduces FENeRF (Face Editing in Neural Radiance Fields), an approach that achieves both strict multi-view consistency and user-friendly, local semantic editing of portrait images.
The authors designed a generative 3D representation that separates shape and appearance using two independent control codes while sharing an underlying geometric volume. Crucially, the system is trained exclusively on standard 2D portrait images paired with 2D semantic attribute masks, eliminating the need for expensive multi-view photography or 3D scan data. To capture fine facial details without corrupting the geometry, a learnable coordinate embedding is incorporated into the color generation branch, and dual discriminators enforce realism and alignment between images and semantic labels. The model was evaluated on the benchmark CelebAMask-HQ and FFHQ portrait datasets.
The experiments show that FENeRF establishes a new state-of-the-art in portrait generation quality, outperforming leading 3D-aware baselines. On the CelebAMask-HQ dataset, FENeRF improved the standard image fidelity score (Frechet Inception Distance) to 12.1 compared to 14.7 for pi-GAN and 16.2 for Giraffe, alongside a substantial reduction in kernel inception error. Furthermore, joint learning of semantic masks and textures significantly refined underlying 3D geometry, eliminating visual artifacts and surface distortions present in earlier models. When inverting real portraits into the model's control space, the system converged to an average intersection-over-union segmentation accuracy above 0.7 within 200 iterations, enabling reliable free-view reconstruction, style mixing, and local modifications such as altering hairstyles or reshaping facial features without distorting adjacent areas.
These findings demonstrate that high-fidelity 3D facial modeling and fine-grained editing can be achieved without costly 3D data collection. This substantially lowers the technical barrier for creating interactive digital avatars, virtual production assets, and graphic editing tools. However, the technology introduces security risks regarding digital identity manipulation, such as creating convincing synthetic video avatars that could deceive facial recognition or liveliness detection systems. Deploying organizations should anticipate these risks by strengthening synthetic media detection protocols.
Before implementing this technology in production environments, teams should address its current operational constraints. The underlying volumetric rendering is computationally intensive, limiting generated image resolutions to lower dimensions and preventing real-time editing due to iterative optimization speeds. Stakeholders should pilot this approach in non-real-time pipelines while future engineering work focuses on accelerating rendering speeds and scaling output to high-definition resolutions.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Introduces Neural Radiance Fields and differentiable volumetric rendering, establishing the core 3D representation that FENeRF adapts for view-consistent face generation.
- Paper: GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields, Michael Niemeyer et al. (2021). Pioneers the integration of neural feature fields with GAN-based adversarial training to achieve compositional, 3D-aware image generation.
- Paper: Analyzing and Improving the Image Quality of StyleGAN, Tero Karras et al. (2020). Provides the foundational generative architecture and latent space modeling used in modern GAN inversion and high-fidelity facial synthesis.
- Paper: Semantic Image Synthesis With Spatially-Adaptive Normalization, Taesung Park et al. (2019). Formulates spatially-adaptive normalization for semantic mask conditioning, underpinning FENeRF's approach to semantic-guided facial generation and editing.
- Paper: Generative Visual Manipulation on the Natural Image Manifold, Jun-Yan Zhu et al. (2016). Presents the foundational principles of interactive generative image manipulation via manifold projection and latent space inversion.
- Paper: A Morphable Model For The Synthesis Of 3D Faces, Volker Blanz et al. (1999). Establishes the classical paradigm of decoupling facial identity shape and texture in parametric 3D face representations.
- Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). Advances 3D-aware GANs by introducing efficient triplane representations for high-resolution geometry and texture synthesis, scaling beyond earlier neural radiance field generators.
- Paper: Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions, Ayaan Haque et al. (2023). Extends 3D neural scene editing beyond semantic mask inversion to interactive, language-guided instruction editing across arbitrary radiance fields.
- Paper: Learning Locally Editable Virtual Humans, Hsuan-I Ho et al. (2023). Builds on localized neural rendering and editable generative fields to enable fine-grained, localized customization of full 3D human bodies.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). Transitions 3D head avatar representation from volumetric radiance fields to controllable 3D Gaussian splatting for faster rendering and expressive animation.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). Applies decoupled 3D representation techniques to one-shot talking portrait animation, controlling geometry and expression dynamically from audio and video.
