FENeRF: Face Editing in Neural Radiance Fields

Jingxiang SunXuan WangYong ZhangXiaoyu LiQi ZhangYebin LiuJue Wang

article2022CVPR174 citations

Proposes a 3D-aware face generator that couples neural radiance fields with decoupled semantic and texture latent spaces, enabling view-consistent portrait synthesis alongside precise local attribute editing trained solely on monocular image-mask pairs.

Listen

Synthesizing photo-realistic human portraits with computer graphics has rapidly evolved, yet existing methods force a compromise between editing capability and 3D geometric realism. Standard 2D image generators allow flexible local attribute editing but produce severe visual distortions when rendering the subject from different camera angles. Conversely, recent 3D-aware methods maintain strict view consistency across angles but lack the ability to support interactive, local adjustments to specific facial features. The article introduces FENeRF (Face Editing in Neural Radiance Fields), an approach that achieves both strict multi-view consistency and user-friendly, local semantic editing of portrait images.

The authors designed a generative 3D representation that separates shape and appearance using two independent control codes while sharing an underlying geometric volume. Crucially, the system is trained exclusively on standard 2D portrait images paired with 2D semantic attribute masks, eliminating the need for expensive multi-view photography or 3D scan data. To capture fine facial details without corrupting the geometry, a learnable coordinate embedding is incorporated into the color generation branch, and dual discriminators enforce realism and alignment between images and semantic labels. The model was evaluated on the benchmark CelebAMask-HQ and FFHQ portrait datasets.

The experiments show that FENeRF establishes a new state-of-the-art in portrait generation quality, outperforming leading 3D-aware baselines. On the CelebAMask-HQ dataset, FENeRF improved the standard image fidelity score (Frechet Inception Distance) to 12.1 compared to 14.7 for pi-GAN and 16.2 for Giraffe, alongside a substantial reduction in kernel inception error. Furthermore, joint learning of semantic masks and textures significantly refined underlying 3D geometry, eliminating visual artifacts and surface distortions present in earlier models. When inverting real portraits into the model's control space, the system converged to an average intersection-over-union segmentation accuracy above 0.7 within 200 iterations, enabling reliable free-view reconstruction, style mixing, and local modifications such as altering hairstyles or reshaping facial features without distorting adjacent areas.

These findings demonstrate that high-fidelity 3D facial modeling and fine-grained editing can be achieved without costly 3D data collection. This substantially lowers the technical barrier for creating interactive digital avatars, virtual production assets, and graphic editing tools. However, the technology introduces security risks regarding digital identity manipulation, such as creating convincing synthetic video avatars that could deceive facial recognition or liveliness detection systems. Deploying organizations should anticipate these risks by strengthening synthetic media detection protocols.

Before implementing this technology in production environments, teams should address its current operational constraints. The underlying volumetric rendering is computationally intensive, limiting generated image resolutions to lower dimensions and preventing real-time editing due to iterative optimization speeds. Stakeholders should pilot this approach in non-real-time pipelines while future engineering work focuses on accelerating rendering speeds and scaling output to high-definition resolutions.

arXiv: 2111.15490
Cover for FENeRF: Face Editing in Neural Radiance Fields

Abstract

Previous portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome these limitations, we propose FENeRF, a 3D-aware generator that can produce view-consistent and locally-editable portrait images. Our method uses two decoupled latent codes to generate corresponding facial semantics and texture in a spatial-aligned 3D volume with shared geometry. Benefiting from such underlying 3D representation, FENeRF can jointly render the boundary-aligned image and semantic mask and use the semantic mask to edit the 3D volume via GAN inversion. We further show such 3D representation can be learned from widely available monocular image and semantic mask pairs. Moreover, we reveal that joint learning semantics and texture helps to generate finer geometry. Our experiments demonstrate that FENeRF outperforms state-of-the-art methods in various face editing tasks. Code is available at https://github.com/MrTornado24/FENeRF.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Locally Editable NeRF Generator
  • 3.2 Discriminators
  • 3.3 Training
  • 4 Experiments
  • 4.1 Comparisons
  • 4.2 Applications
  • 4.3 Ablation Studies
  • 5 Limitations
  • 6 Conclusion
  • 7 Potential Social Impact
  • References

Knowls

  1. Knowl 1 — Decoupled Semantic and Radiance Field Generator Architecture in FENeRF

    model/method

    FENeRF (Face Editing in Neural Radiance Fields) generates 3D-aware portrait images along with spatially aligned semantic segmentations by conditioning a shared neural implicit volume on decoupled latent codes for shape and texture. The generator GG maps a 3D spatial coordinate x=(x,y,z)∈R3\mathbf{x} = (x, y, z) \in \mathbb{R}^3, a 2D viewing direction d=(θ,ϕ)\mathbf{d} = (\theta, \phi), a shape latent code zs∼N(0,I)\mathbf{z}_s \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), a texture latent code zt∼N(0,I)\mathbf{z}_t \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), and a local coordinate feature embedding ecoordx\mathbf{e}_{\text{coord}}^\mathbf{x} to view-invariant volume density σ∈R+\sigma \in \mathbb{R}^+, view-invariant semantic label logits s∈Rk\mathbf{s} \in \mathbb{R}^k (for kk semantic classes), and view-dependent color c∈R3\mathbf{c} \in \mathbb{R}^3:

    G(x,d,zs,zt,ecoordx)↦(σ,c,s)G(\mathbf{x}, \mathbf{d}, \mathbf{z}_s, \mathbf{z}_t, \mathbf{e}_{\text{coord}}^\mathbf{x}) \mapsto (\sigma, \mathbf{c}, \mathbf{s})

    The generator is parameterized as a Multi-Layer Perceptron (MLP) employing sinusoidal representation networks (SIREN) modulated via Feature-wise Linear Modulation (FiLM). Mapping networks transform sampled latent codes zs\mathbf{z}_s and zt\mathbf{z}_t into intermediate latent representations in space W\mathcal{W} that produce layer-wise modulation frequencies γi\boldsymbol{\gamma}_i and phase shifts βi\boldsymbol{\beta}_i. For the ii-th network layer with input xi∈RMi\mathbf{x}_i \in \mathbb{R}^{M_i}, weight matrix Wi∈RNi×Mi\mathbf{W}_i \in \mathbb{R}^{N_i \times M_i}, and bias bi∈RMi\mathbf{b}_i \in \mathbb{R}^{M_i}, the modulated sine activation ϕi(xi)∈RNi\phi_i(\mathbf{x}_i) \in \mathbb{R}^{N_i} is given by:

    ϕi(xi)=sin⁡(γi⊙(Wixi+bi)+βi)\phi_i(\mathbf{x}_i) = \sin\left(\boldsymbol{\gamma}_i \odot (\mathbf{W}_i \mathbf{x}_i + \mathbf{b}_i) + \boldsymbol{\beta}_i\right)

    The complete network stack maps spatial coordinates through nn layers:

    Φ(x)=Wn(ϕn−1∘ϕn−2∘⋯∘ϕ0)(x)+bn\Phi(\mathbf{x}) = \mathbf{W}_n (\phi_{n-1} \circ \phi_{n-2} \circ \dots \circ \phi_0)(\mathbf{x}) + \mathbf{b}_n

    To ensure geometric alignment across shape, semantics, and appearance, the network uses a shared intermediate feature backbone that branches into three heads: a density head predicting σ\sigma conditioned on zs\mathbf{z}_s, a semantic head predicting category logits s\mathbf{s} conditioned on zs\mathbf{z}_s, and an appearance head predicting view-dependent color c\mathbf{c} conditioned on zt\mathbf{z}_t, viewing direction d\mathbf{d}, and positional embedding ecoordx\mathbf{e}_{\text{coord}}^\mathbf{x}.

  2. Knowl 2 — Learnable 3D Positional Feature Grid Injection

    model/method

    To overcome the blurriness and detail loss inherent to pure SIREN-based generative radiance fields, FENeRF incorporates a learnable 3D feature grid. For any queried 3D point x=(x,y,z)\mathbf{x} = (x, y, z), a local coordinate feature embedding ecoordx\mathbf{e}_{\text{coord}}^\mathbf{x} is sampled from this 3D grid via tricubic (bi-cubic across grid slices) interpolation.

    The sampled feature vector ecoordx\mathbf{e}_{\text{coord}}^\mathbf{x} is injected specifically as an auxiliary input into the color prediction branch of the generator alongside viewing direction d=(θ,ϕ)\mathbf{d} = (\theta, \phi), rather than being concatenated to the initial spatial coordinates x\mathbf{x} at the input layer. Injecting ecoordx\mathbf{e}_{\text{coord}}^\mathbf{x} exclusively into the color branch preserves high-frequency appearance details (such as facial wrinkles, eyes, and teeth) while preventing high-frequency noise from perturbing the density volume σ\sigma and semantic field s\mathbf{s}, which ensures smooth underlying 3D geometry.

  3. Knowl 3 — Neural Volume Rendering for Joint Radiance and Semantic Maps

    equation

    Given a camera ray r(t)=o+td\mathbf{r}(t) = \mathbf{o} + t\mathbf{d} with near bound tnt_n, far bound tft_f, origin o\mathbf{o}, and unit viewing direction d\mathbf{d}, FENeRF simultaneously synthesizes the accumulated pixel color C(r)∈R3\mathbf{C}(\mathbf{r}) \in \mathbb{R}^3 and the accumulated semantic class probability vector S(r)∈Rk\mathbf{S}(\mathbf{r}) \in \mathbb{R}^k via continuous volume rendering integrals:

    C(r)=∫tntfT(t) σ(r(t)) c(r(t),d) dt\mathbf{C}(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma(\mathbf{r}(t))\,\mathbf{c}(\mathbf{r}(t), \mathbf{d})\,dt

    S(r)=∫tntfT(t) σ(r(t)) s(r(t),d) dt\mathbf{S}(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma(\mathbf{r}(t))\,\mathbf{s}(\mathbf{r}(t), \mathbf{d})\,dt

    where the volume transmittance T(t)T(t) represents the probability that the ray traverses from tnt_n to tt without hitting any particles:

    T(t)=exp⁡(−∫tntσ(r(u)) du)T(t) = \exp\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(u))\,du\right)

    Here σ(r(t))∈R+\sigma(\mathbf{r}(t)) \in \mathbb{R}^+ denotes the volume density, c(r(t),d)∈R3\mathbf{c}(\mathbf{r}(t), \mathbf{d}) \in \mathbb{R}^3 is the emitted radiance at point r(t)\mathbf{r}(t) along direction d\mathbf{d}, and s(r(t),d)∈Rk\mathbf{s}(\mathbf{r}(t), \mathbf{d}) \in \mathbb{R}^k denotes the vector of kk semantic label logits. In practice, these integrals are discretized using stratified sampling along the ray. Because both rendering integrals share the identical volume density σ(r(t))\sigma(\mathbf{r}(t)) along the ray, the resulting 2D RGB portrait image and 2D semantic segmentation mask are inherently boundary-aligned in 3D space.

  4. Knowl 4 — Dual Discriminator Training Objectives and Camera Pose Regularization

    equation

    FENeRF is trained using unpaired monocular 2D images and semantic masks under an adversarial formulation supervised by two convolutional neural network discriminators: an image discriminator DcD_c and a joint semantic-image discriminator DsD_s.

    Let xc=C(r)\mathbf{x}_c = \mathbf{C}(\mathbf{r}) and xs=S(r)\mathbf{x}_s = \mathbf{S}(\mathbf{r}) denote the generated color image and semantic mask rendered from sampled latent codes zs,zt∼N(0,I)\mathbf{z}_s, \mathbf{z}_t \sim \mathcal{N}(\mathbf{0}, \mathbf{I}) and camera pose ξ∼pξ\xi \sim p_\xi. Let I∼pi\mathbf{I} \sim p_i and L∼pl\mathbf{L} \sim p_l denote real portrait images and real semantic masks from the training dataset. The non-saturating GAN loss with R1R_1 gradient penalty is parameterized using f(t)=−log⁡(1+exp⁡(−t))f(t) = -\log(1 + \exp(-t)) with penalty coefficients λc=λs=λp=10\lambda_c = \lambda_s = \lambda_p = 10.

    The objective for image discriminator DcD_c is:

    LDc=Ezs,zt∼N,ξ∼pξ[f(Dc(xc))]+EI∼pi[f(−Dc(I))+λc∥∇Dc(I)∥2]\mathcal{L}_{D_c} = \mathbb{E}_{\mathbf{z}_s, \mathbf{z}_t \sim \mathcal{N}, \xi \sim p_\xi} \left[ f(D_c(\mathbf{x}_c)) \right] + \mathbb{E}_{\mathbf{I} \sim p_i} \left[ f(-D_c(\mathbf{I})) + \lambda_c \|\nabla D_c(\mathbf{I})\|^2 \right]

    The objective for semantic discriminator DsD_s is:

    LDs=Ezs,zt∼N,ξ∼pξ[f(Ds(xs,xc))]+EI∼pi,L∼pl[f(−Ds(L,I))+λs∥∇Ds(L,I)∥2]\mathcal{L}_{D_s} = \mathbb{E}_{\mathbf{z}_s, \mathbf{z}_t \sim \mathcal{N}, \xi \sim p_\xi} \left[ f(D_s(\mathbf{x}_s, \mathbf{x}_c)) \right] + \mathbb{E}_{\mathbf{I} \sim p_i, \mathbf{L} \sim p_l} \left[ f(-D_s(\mathbf{L}, \mathbf{I})) + \lambda_s \|\nabla D_s(\mathbf{L}, \mathbf{I})\|^2 \right]

    The objective for generator GG is:

    LG=Ezs,zt∼N,ξ∼pξ[f(Dc(xc))]+Ezs,zt∼N,ξ∼pξ[f(Ds(xs,xc))]+λp∥ξ^−ξ∥\mathcal{L}_G = \mathbb{E}_{\mathbf{z}_s, \mathbf{z}_t \sim \mathcal{N}, \xi \sim p_\xi} \left[ f(D_c(\mathbf{x}_c)) \right] + \mathbb{E}_{\mathbf{z}_s, \mathbf{z}_t \sim \mathcal{N}, \xi \sim p_\xi} \left[ f(D_s(\mathbf{x}_s, \mathbf{x}_c)) \right] + \lambda_p \|\hat{\xi} - \xi\|

    where ξ^\hat{\xi} is the camera pose predicted by two auxiliary output channels appended to DcD_c. The pose correction loss λp∥ξ^−ξ∥\lambda_p \|\hat{\xi} - \xi\| enforces canonical 3D alignment and prevents pose drift. During optimization of LG\mathcal{L}_G, gradients backpropagating from DsD_s into the color branch are blocked to prevent texture details from being smoothed into flat semantic regions.

  5. Knowl 5 — Semantic-Guided 3D Face Attribute Editing via GAN Inversion

    algorithm

    FENeRF enables interactive, 3D view-consistent facial attribute editing (e.g., hair style, eyes, nose shape) by decoupling shape and appearance and performing semantic inversion into the shape latent space.

    Input: Real portrait image I\mathbf{I}, novel camera pose ξtarget\xi_{\text{target}}, edited 2D semantic mask Ledit\mathbf{L}_{\text{edit}}, trained generator GG, image discriminator DcD_c
    Output: Edited free-view 3D portrait images C(r;ξtarget)\mathbf{C}(\mathbf{r}; \xi_{\text{target}})
    Step 1: Estimate the initial camera pose ξinit\xi_{\text{init}} of image I\mathbf{I} using the pose prediction head of DcD_c.
    Step 2: Optimize latent codes zs\mathbf{z}_s and zt\mathbf{z}_t via GAN inversion to minimize reconstruction error between rendered image C(r;zs,zt,ξinit)\mathbf{C}(\mathbf{r}; \mathbf{z}_s, \mathbf{z}_t, \xi_{\text{init}}) and target image I\mathbf{I}, yielding inverted codes zs∗\mathbf{z}_s^* and zt∗\mathbf{z}_t^*.
    Step 3: Render the corresponding boundary-aligned 2D semantic mask S(r;zs∗,ξinit)\mathbf{S}(\mathbf{r}; \mathbf{z}_s^*, \xi_{\text{init}}).
    Step 4: Interactively modify the rendered 2D semantic mask to produce target layout Ledit\mathbf{L}_{\text{edit}}.
    Step 5: Fix the appearance latent code zt∗\mathbf{z}_t^* and optimize only the shape latent code zs\mathbf{z}_s to match Ledit\mathbf{L}_{\text{edit}} by minimizing the semantic loss between rendered mask S(r;zs,ξinit)\mathbf{S}(\mathbf{r}; \mathbf{z}_s, \xi_{\text{init}}) and Ledit\mathbf{L}_{\text{edit}}, yielding edited shape code zsedit\mathbf{z}_s^{\text{edit}}.
    Step 6: Render the edited 3D scene from any desired novel camera pose ξtarget\xi_{\text{target}} using G(⋅,⋅,zsedit,zt∗,ecoord)G(\cdot, \cdot, \mathbf{z}_s^{\text{edit}}, \mathbf{z}_t^*, \mathbf{e}_{\text{coord}}).
    return Free-view edited portrait C(r;zsedit,zt∗,ξtarget)\mathbf{C}(\mathbf{r}; \mathbf{z}_s^{\text{edit}}, \mathbf{z}_t^*, \xi_{\text{target}})

    Because zs\mathbf{z}_s controls both density σ\sigma and semantics s\mathbf{s} in 3D space, optimizing zs\mathbf{z}_s deforms 3D geometry locally while keeping unedited regions and texture consistency intact across all viewpoints.

  6. Knowl 6 — Quantitative Image Synthesis Quality Comparison on CelebAMask-HQ and FFHQ

    data/table

    The image synthesis quality of FENeRF was quantitatively evaluated against state-of-the-art 3D-aware GAN baselines (GRAF, π\pi-GAN, and GIRAFFE) on CelebAMask-HQ (30,000 images) and FFHQ (70,000 images) at 128×128128 \times 128 resolution. Metrics include Fréchet Inception Distance (FID, lower is better) and Kernel Inception Distance (KID×103\text{KID} \times 10^3, lower is better) computed on 2,048 randomly sampled images.

    Method FID ↓\downarrow KID (×103\times 10^3) ↓\downarrow
    CelebA-HQ FFHQ CelebA-HQ FFHQ
    GRAF 34.7 66.5 15.6 49.3
    π\pi-GAN 14.7 40.3 3.9 23.5
    Giraffe 16.2 31.9 9.1 32.7
    FENeRF 12.1 28.2 1.6 17.3

    FENeRF achieves the best FID and KID across both datasets, outperforming π\pi-GAN by 2.6 FID points on CelebA-HQ and 12.1 FID points on FFHQ. This improvement stems from two mechanisms: joint training with semantic fields which stabilizes optimization and guides 3D convergence, and the learnable coordinate feature grid which introduces high-frequency texture details.

  7. Knowl 7 — Geometry Regularization via Joint Semantic and Texture Volume Learning

    empirical result

    In 3D-aware GANs such as π\pi-GAN, unsupervised optimization of density from 2D RGB images alone frequently leads to noisy facial boundaries, floating artifacts, and inaccurate background separation.

    FENeRF reveals a bidirectional synergy between semantic rendering and appearance rendering:

    1. Training the generator with semantic mask supervision alongside RGB images provides strong geometric constraints that force the volume density σ\sigma to form smooth, accurate 3D facial surfaces without requiring 3D scans, multi-view data, or explicit geometric regularization.
    2. Conversely, training with image rendering is essential for semantic field convergence: an ablated model trained to render semantic maps alone fails to segment facial regions cleanly or establish reliable 3D geometry, whereas joint learning with RGB supervision allows both modalities to align accurately in 3D space.
  8. Knowl 8 — Ablation on Positional Feature Embedding and Injection Strategy

    empirical result

    Ablation experiments on the role and injection location of the learnable 3D coordinate feature embedding grid (ecoord\mathbf{e}_{\text{coord}}) demonstrate the following properties:

    1. Without ecoord\mathbf{e}_{\text{coord}}: The generator produces blurry image details in high-frequency regions such as teeth, eyes, and hair.
    2. Injecting ecoord\mathbf{e}_{\text{coord}} at the MLP input: When ecoord\mathbf{e}_{\text{coord}} is concatenated with spatial coordinates (x,y,z)(x, y, z) at the input layer of the generator, high-frequency signals enter the shared backbone and perturb the density volume σ\sigma, creating surface noise, visual artifacts on synthesized images, and distorted semantic segmentation maps.
    3. Injecting ecoord\mathbf{e}_{\text{coord}} into the color branch: Concatenating the tricubically interpolated feature vector ecoordx\mathbf{e}_{\text{coord}}^\mathbf{x} exclusively into the color prediction MLP head preserves fine-grained texture details without introducing geometric artifacts or degrading the smoothness of the underlying 3D density and semantic fields.
  9. Knowl 9 — View Consistency Comparison in 3D Inversion and Semantic Rendering Against Baselines

    empirical result

    When evaluated on latent space inversion and novel view synthesis, FENeRF demonstrates superior 3D view consistency compared to 2D GAN inversion methods (InterFaceGAN and e4e) and semantic-field methods (SofGAN):

    1. Comparison with 2D GAN Inversion: InterFaceGAN exhibits severe visual artifacts such as texture sticking to the image plane during camera rotation, while e4e fails to preserve facial identity and geometric structure across viewpoint changes. FENeRF projects input 2D portraits into latent codes controlling a continuous 3D volume, maintaining strict identity and viewpoint consistency under rotation.
    2. Comparison with SofGAN: SofGAN relies on a semantic occupancy field trained on multi-view scans and synthesizes RGB images via a 2D image generator conditioned on rendered semantic maps. SofGAN produces incorrect semantic labels and visible artifacts at extreme camera poses, and exhibits viewpoint inconsistency (e.g., eye gaze direction changing upon zoom). In contrast, FENeRF's joint 3D volume rendering guarantees strict pixel-level view consistency without requiring 3D scans or multi-view training data.
  10. Knowl 10 — Computational and Latent Optimization Limitations of FENeRF

    limitation

    FENeRF exhibits two primary operational limitations:

    1. Computational rendering overhead: Due to the dense ray marching and numerical volume integration required by neural radiance fields, generating high-definition (HD) portrait images is computationally intensive, restricting training and real-time inference resolutions (e.g., 128×128128 \times 128).
    2. Inversion latency: Performing semantic or appearance editing on real portrait images relies on iterative gradient-based GAN inversion to project target images and masks into the latent spaces (zs,zt)(\mathbf{z}_s, \mathbf{z}_t), which requires several hundred iterations and prevents real-time interactive free-view portrait manipulation.

Coverage note — None was omitted; all key contributions—the decoupled architecture, feature grid, rendering equations, dual discriminator losses, inversion algorithm, quantitative benchmarks, ablations, baseline comparisons, and limitations—are covered.

References

  1. 1.Mikołaj Bińkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans. arXiv preprint arXiv:1801.01401, 2018.
  2. 2.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In arXiv, 2021.
  3. 3.Eric R Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5799–5809, 2021.
  4. 4.Anpei Chen, Ruiyang Liu, Ling Xie, Zhang Chen, Hao Su, and Jingyi Yu. Sofgan: A portrait image generator with dynamic styling. ACM Transactions on Graphics (TOG), 41(1):1–26, 2022.
  5. 5.Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14124–14133, 2021.
  6. 6.Shu-Yu Chen, Feng-Lin Liu, Yu-Kun Lai, Paul L. Rosin, Chunpeng Li, Hongbo Fu, and Lin Gao. Deepfaceediting: Deep generation of face images from sketches. ACM Transactions on Graphics (TOG), 40(4):90:1–90:15, 2021.
  7. 7.Shu-Yu Chen, Wanchao Su, Lin Gao, Shihong Xia, and Hongbo Fu. Deepfacedrawing: Deep generation of face images from sketches. ACM Transactions on Graphics (TOG), 39(4):72–1, 2020.
  8. 8.Forrester Cole, Kyle Genova, Avneesh Sud, Daniel Vlasic, and Zhoutong Zhang. Differentiable surface rendering via non-differentiable sampling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6088–6097, 2021.
  9. 9.Edo Collins, Raja Bala, Bob Price, and Sabine Süsstrunk. Editing in style: Uncovering the local semantics of gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5771–5780, 2020.
  10. 10.Yu Deng, Jiaolong Yang, Jianfeng Xiang, and Xin Tong. Gram: Generative radiance manifolds for 3d-aware image generation. In arXiv, 2021.
  11. 11.William Fedus, Ian Goodfellow, and Andrew M Dai. Maskgan: better text generation via filling in the . arXiv preprint arXiv:1801.07736, 2018.
  12. 12.Matheus Gadelha, Subhransu Maji, and Rui Wang. 3d shape induction from 2d views of multiple objects. In 2017 International Conference on 3D Vision (3DV), pages 402–411. IEEE, 2017.
  13. 13.Stephan J Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien Valentin. Fastnerf: High-fidelity neural rendering at 200fps. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14346–14355, 2021.
  14. 14.Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis, 2021.
  15. 15.Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5784–5794, 2021.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  17. 17.Peter Hedman, Pratul P Srinivasan, Ben Mildenhall, Jonathan T Barron, and Paul Debevec. Baking neural radiance fields for real-time view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5875–5884, 2021.
  18. 18.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  19. 19.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017.
  20. 20.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations, 2018.
  21. 21.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019.
  22. 22.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8110–8119, 2020.
  23. 23.Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5549–5558, 2020.
  24. 24.Thomas Leimkühler and George Drettakis. Freestylegan: Free-view editable portrait rendering with the camera manifold. ACM Transactions on Graphics (SIGGRAPH Asia), 40(6), 2021.
  25. 25.Yuhang Li, Xuejin Chen, Binxin Yang, Zihan Chen, Zhihua Cheng, and Zheng-Jun Zha. Deepfacepencil: Creating face images from freehand sketches. In Proceedings of the 28th ACM International Conference on Multimedia, pages 991–999, 2020.
  26. 26.Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. Barf: Bundle-adjusting neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5741–5751, 2021.
  27. 27.Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose control. ACM Transactions on Graphics (TOG), 40(6):1–16, 2021.
  28. 28.Rosanne Liu, Joel Lehman, Piero Molino, Felipe Petroski Such, Eric Frank, Alex Sergeev, and Jason Yosinski. An intriguing failing of convolutional neural networks and the coordconv solution. Advances in neural information processing systems, 31, 2018.
  29. 29.Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, and Bryan Russell. Editing conditional radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5773–5783, 2021.
  30. 30.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, pages 3730–3738, 2015.
  31. 31.Quan Meng, Anpei Chen, Haimin Luo, Minye Wu, Hao Su, Lan Xu, Xuming He, and Jingyi Yu. Gnerf: Gan-based neural radiance field without posed camera. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6351–6361, 2021.
  32. 32.Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for gans do actually converge? In International conference on machine learning, pages 3481–3490. PMLR, 2018.
  33. 33.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4460–4470, 2019.
  34. 34.Mateusz Michalkiewicz, Jhony K Pontes, Dominic Jack, Mahsa Baktashmotlagh, and Anders Eriksson. Implicit surface representations as layers in neural networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4743–4752, 2019.
  35. 35.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision, pages 405–421. Springer, 2020.
  36. 36.Thu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt, and Yong-Liang Yang. Hologan: Unsupervised learning of 3d representations from natural images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7588–7597, 2019.
  37. 37.Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11453–11464, 2021.
  38. 38.Roy Or-El, Xuan Luo, Mengyi Shan, Eli Shechtman, Jeong Joon Park, and Ira Kemelmacher-Shlizerman. StyleSDF: High-Resolution 3D-Consistent Image and Geometry Generation. arXiv preprint arXiv:2112.11427, 2021.
  39. 39.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 165–174, 2019.
  40. 40.Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2337–2346, 2019.
  41. 41.Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Animatable neural radiance fields for modeling dynamic human bodies. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14314–14323, 2021.
  42. 42.Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  43. 43.Christian Reiser, Songyou Peng, Yiyi Liao, and Andreas Geiger. Kilonerf: Speeding up neural radiance fields with thousands of tiny mlps. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14335–14345, 2021.
  44. 44.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. Advances in Neural Information Processing Systems, 33:20154–20166, 2020.
  45. 45.Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE transactions on pattern analysis and machine intelligence, 2020.
  46. 46.Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj, Florian Bernard, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, and Christian Theobalt. Stylerig: Rigging stylegan for 3d control over portrait images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6142–6151, 2020.
  47. 47.Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. Designing an encoder for stylegan image manipulation. ACM Transactions on Graphics (TOG), 40(4):1–14, 2021.
  48. 48.Tengfei Wang, Yong Zhang, Yanbo Fan, Jue Wang, and Qifeng Chen. High-fidelity gan inversion for image attribute editing. arXiv preprint arXiv:2109.06590, 2021.
  49. 49.Alex Yu, Ruilong Li, Matthew Tancik, Hao Li, Ren Ng, and Angjoo Kanazawa. Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5752–5761, 2021.
  50. 50.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021.
  51. 51.Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang. Bisenet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 325–341, 2018.
  52. 52.Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and Andrew J Davison. In-place scene labelling and understanding with implicit scene representation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15838–15847, 2021.
  53. 53.Peng Zhou, Lingxi Xie, Bingbing Ni, and Qi Tian. Cips-3d: A 3d-aware generator of gans based on conditionally-independent pixel synthesis. arXiv preprint arXiv:2110.09788, 2021.
  54. 54.Jun-Yan Zhu, Zhoutong Zhang, Chengkai Zhang, Jiajun Wu, Antonio Torralba, Josh Tenenbaum, and Bill Freeman. Visual object networks: Image generation with disentangled 3d representations. Advances in neural information processing systems, 31, 2018.
  55. 55.Peihao Zhu, Rameen Abdal, Yipeng Qin, and Peter Wonka. Sean: Image synthesis with semantic region-adaptive normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5104–5113, 2020.

Citation

MLA
Sun, J., et al. “FENeRF: Face Editing in Neural Radiance Fields”. arXiv, 2021, http://arxiv.org/abs/2111.15490v2.
APA
Sun, J., Wang, X., Zhang, Y., Li, X., Zhang, Q., Liu, Y., & Wang, J. (2021). FENeRF: Face Editing in Neural Radiance Fields. arXiv. http://arxiv.org/abs/2111.15490v2
Chicago
Sun, J., X. Wang, Y. Zhang, et al. 2021. “FENeRF: Face Editing in Neural Radiance Fields”. arXiv. http://arxiv.org/abs/2111.15490v2.
Harvard
Sun, J. et al. (2021) “FENeRF: Face Editing in Neural Radiance Fields”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.15490v2.
Vancouver
1. Sun J, Wang X, Zhang Y, Li X, Zhang Q, Liu Y, Wang J (2021) FENeRF: Face Editing in Neural Radiance Fields. arXiv

BibTeX

@article{sun2021fenerf,
  title = {FENeRF: Face Editing in Neural Radiance Fields},
  author = {Sun, Jingxiang and Wang, Xuan and Zhang, Yong and Li, Xiaoyu and Zhang, Qi and Liu, Yebin and Wang, Jue},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.15490v2},
  eprint = {2111.15490}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE