Relightable Gaussian Codec Avatars

Shunsuke SaitoGabriel SchwartzTomas SimonJunxuan LiGiljoo Nam

article2024CVPR139 citations

Combines animatable 3D Gaussian splatting with learnable radiance transfer to achieve real-time, all-frequency relighting of dynamic head avatars featuring sub-millimeter geometric details and explicit gaze control.

Listen

Generating realistic, animatable digital human avatars that can adapt seamlessly to any lighting environment is essential for telecommunication, virtual reality, and interactive gaming. However, current systems struggle to produce photorealistic results in real time. Human heads comprise varied, complex materials—such as translucent skin, multi-layered reflective eyes, and fine hair fibers—which traditionally require prohibitively expensive computation to render under changing illumination conditions, often blurring fine features like individual hair strands.

The article demonstrates a novel framework called Relightable Gaussian Codec Avatars, which creates dynamic, highly detailed, and relightable digital head models capable of running in real time. The primary objective is to evaluate whether combining 3D Gaussian point-based geometry with a learned radiance transfer appearance model can accurately reproduce complex global illumination and sharp reflections across all head materials simultaneously.

To accomplish this, the authors captured multi-view performance data from human subjects across approximately 144,000 frames using a synchronized 110-camera, 460-LED light-stage rig. The framework employs 3D Gaussian Splatting anchored to a template mesh to efficiently render intricate sub-millimeter geometry, a neural radiance transfer function decomposing diffuse and specular light components via spherical harmonics and spherical Gaussians, and an explicit dual-sphere geometric eye model to control gaze and realistic cornea reflections.

The key findings confirm substantial performance improvements across multiple metrics. First, the proposed method consistently outperformed existing real-time baselines—such as volumetric primitives and linear neural networks—achieving superior visual quality with peak signal-to-noise ratios exceeding 36 dB on held-out expressions. Second, the 3D Gaussian geometry captured sub-millimeter details like pores and distinct hair strands that previous volumetric methods blurred. Third, the spherical Gaussian specular model achieved sharp, all-frequency environmental reflections without the computational bottlenecks or visual flickering seen in prior spherical harmonic formulations. Finally, the explicit eye model successfully disentangled eye gaze control while rendering realistic cornea reflections.

These results demonstrate that photorealistic, fully relightable avatars can run efficiently on consumer-grade virtual reality headsets without requiring specialized rendering workarounds or expensive per-frame offline ray tracing. By maintaining the mathematical linearity of light transport, the model renders both point-source lights and complex environment maps instantaneously, significantly lowering runtime performance costs.

Organizations developing virtual presence or interactive graphics systems should consider adopting 3D Gaussian splatting paired with learned radiance transfer for avatar rendering pipelines. However, decision-makers should note that current avatar construction relies on highly controlled light-stage data capture and multi-view coarse tracking pipelines. Future efforts should focus on enabling end-to-end training that reduces preprocessing dependencies and exploring per-pixel fragment shader offloading to scale rendering capacity when animating multiple avatars simultaneously.

arXiv: 2312.03704
Cover for Relightable Gaussian Codec Avatars

Abstract

The fidelity of relighting is bounded by both geometry and appearance representations. For geometry, both mesh and volumetric approaches have difficulty modeling intricate structures like 3D hair geometry. For appearance, existing relighting models are limited in fidelity and often too slow to render in real-time with high-resolution continuous environments. In this work, we present Relightable Gaussian Codec Avatars, a method to build high-fidelity relightable head avatars that can be animated to generate novel expressions. Our geometry model based on 3D Gaussians can capture 3D-consistent sub-millimeter details such as hair strands and pores on dynamic face sequences. To support diverse materials of human heads such as the eyes, skin, and hair in a unified manner, we present a novel relightable appearance model based on learnable radiance transfer. Together with global illumination-aware spherical harmonics for the diffuse components, we achieve real-time relighting with all-frequency reflections using spherical Gaussians. This appearance model can be efficiently relit under both point light and continuous illumination. We further improve the fidelity of eye reflections and enable explicit gaze control by introducing relightable explicit eye models. Our method outperforms existing approaches without compromising real-time performance. We also demonstrate real-time relighting of avatars on a tethered consumer VR headset, showcasing the efficiency and fidelity of our avatars.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Data Acquisition
  • 3.2. Geometry: 3D Gaussian Avatars
  • 3.3. Appearance: Learned Radiance Transfer
  • 3.4. Relightable Explicit Eye Model
  • 3.5. Training
  • 4. Experiments
  • 4.1. Qualitative Results
  • 4.2. Discussion
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Drivable 3D Gaussian Head Avatar Geometry Representation

    model/method

    The dynamic head geometry is represented as a collection of MM 3D anisotropic Gaussians, where each Gaussian gk={tk,Rk,sk,ok,ck}g_k = \{\mathbf{t}_k, \mathbf{R}_k, \mathbf{s}_k, o_k, \mathbf{c}_k\} is defined by its spatial translation center tk∈R3\mathbf{t}_k \in \mathbb{R}^3, rotation matrix Rk∈SO(3)\mathbf{R}_k \in \mathrm{SO}(3) parameterized as a quaternion, per-axis scale vector sk∈R3\mathbf{s}_k \in \mathbb{R}^3, opacity ok∈Ro_k \in \mathbb{R}, and base color ck∈R3\mathbf{c}_k \in \mathbb{R}^3.

    To drive facial animations while maintaining temporal coherence, each 3D Gaussian is mapped to an individual texel on a shared 2D UV texture parameterization of a template head mesh. A Conditional Variational Auto-Encoder (CVAE) models the facial expression distribution. An encoder E\mathcal{E} takes tracked coarse mesh vertices V\mathbf{V} and an unwrapped average texture T\mathbf{T} to output parameters μe,σe=E(V,T;Θe)\boldsymbol{\mu}_e, \boldsymbol{\sigma}_e = \mathcal{E}(\mathbf{V}, \mathbf{T}; \Theta_e). A latent expression code z∈R256\mathbf{z} \in \mathbb{R}^{256} is sampled via z∼N(μe,σe)\mathbf{z} \sim \mathcal{N}(\boldsymbol{\mu}_e, \boldsymbol{\sigma}_e).

    A geometry decoder Dg\mathcal{D}_g predicts Gaussian transformations conditioned on z\mathbf{z} and left/right eye gaze directions e{l,r}∈R3\mathbf{e}_{\{l,r\}} \in \mathbb{R}^3:

    {δtk,Rk,sk,ok}k=1M=Dg(z,e{l,r};Θg)\{\delta \mathbf{t}_k, \mathbf{R}_k, \mathbf{s}_k, o_k\}_{k=1}^{M} = \mathcal{D}_g(\mathbf{z}, \mathbf{e}_{\{l,r\}}; \Theta_g)

    To avoid poor local minima under large head and face motions, Gaussian centers are guided by coarse mesh deformation V′=Dv(z;Θv)\mathbf{V}' = \mathcal{D}_v(\mathbf{z}; \Theta_v), where Dv\mathcal{D}_v is a vertex decoder. The final translation is computed as tk=t^k+δtk\mathbf{t}_k = \hat{\mathbf{t}}_k + \delta \mathbf{t}_k, where t^k\hat{\mathbf{t}}_k is the barycentrically interpolated position on the deformed mesh V′\mathbf{V}' corresponding to the Gaussian's UV coordinate.

  2. Knowl 2 — Diffuse Radiance Transfer via Hybrid-Order Spherical Harmonics

    model/method

    The outgoing color ck\mathbf{c}_k of each 3D Gaussian is decomposed into a view-independent diffuse term ckdiffuse\mathbf{c}_k^{\text{diffuse}} and a view-dependent specular term ckspecular(ωo)\mathbf{c}_k^{\text{specular}}(\boldsymbol{\omega}_o), such that ck=ckdiffuse+ckspecular(ωo)\mathbf{c}_k = \mathbf{c}_k^{\text{diffuse}} + \mathbf{c}_k^{\text{specular}}(\boldsymbol{\omega}_o), where ωo∈S2\boldsymbol{\omega}_o \in \mathbb{S}^2 is the viewing direction.

    The diffuse color contribution incorporates global light transport effects (subsurface scattering, multi-bounce scattering, and ambient occlusion) through a learned intrinsic radiance transfer function dk(⋅)\mathbf{d}_k(\cdot) integrated against incident illumination L(⋅)\mathbf{L}(\cdot) over the sphere S2\mathbb{S}^2:

    ckdiffuse=ρk⊙∫S2L(ωi)⊙dk(ωi)dωi=ρk⊙∑i=1(n+1)2Li⊙dki\mathbf{c}_k^{\text{diffuse}} = \boldsymbol{\rho}_k \odot \int_{\mathbb{S}^2} \mathbf{L}(\boldsymbol{\omega}_i) \odot \mathbf{d}_k(\boldsymbol{\omega}_i) \mathrm{d}\boldsymbol{\omega}_i = \boldsymbol{\rho}_k \odot \sum_{i=1}^{(n+1)^2} \mathbf{L}_i \odot \mathbf{d}_k^i

    where Li\mathbf{L}_i and dki∈R3\mathbf{d}_k^i \in \mathbb{R}^3 are the nn-th order spherical harmonics (SH) expansion coefficients of the incident lighting and the transfer function, respectively, and ρk∈R3\boldsymbol{\rho}_k \in \mathbb{R}^3 is a statically defined learnable per-Gaussian diffuse albedo ensuring temporal consistency across frames.

    To balance memory consumption with the need to capture higher-frequency shadow boundaries, the intrinsic transfer representation splits the SH bands: RGB color SH coefficients dkc\mathbf{d}_k^c are decoded up to the 3rd order (n=3n=3, 16 coefficients per color channel), while monochrome SH coefficients dkm\mathbf{d}_k^m are decoded from the 4th to the 8th order (n=8n=8, 65 scalar coefficients).

  3. Knowl 3 — Specular Radiance Transfer via Normalized Angle-Based Spherical Gaussians

    model/method

    To render sharp, all-frequency specular reflections in real-time, the view-dependent specular color ckspecular(ωo)\mathbf{c}_k^{\text{specular}}(\boldsymbol{\omega}_o) of each 3D Gaussian is modeled using a normalized angle-based Spherical Gaussian (SG) angular basis GsG_s:

    Gs(p;q,σ)=Ce−12(arccos⁡(p⋅q)σ)2G_s(\mathbf{p}; \mathbf{q}, \sigma) = C e^{-\frac{1}{2}\left(\frac{\arccos(\mathbf{p} \cdot \mathbf{q})}{\sigma}\right)^2}

    where q∈S2\mathbf{q} \in \mathbb{S}^2 is the lobe center axis, p∈S2\mathbf{p} \in \mathbb{S}^2 is the evaluation direction, σ∈R+\sigma \in \mathbb{R}^+ is the standard deviation of angular decay (roughness), and C=1/(2π2/3σ)C = 1/(\sqrt{2\pi^{2/3}}\sigma) is a normalization constant that preserves the Gaussian integral over the unit sphere.

    The specular lobe axis qk\mathbf{q}_k is aligned with the mirror reflection vector computed at the Gaussian center:

    qk=2(ωko⋅nk)nk−ωko\mathbf{q}_k = 2(\boldsymbol{\omega}_k^o \cdot \mathbf{n}_k)\mathbf{n}_k - \boldsymbol{\omega}_k^o

    where ωko\boldsymbol{\omega}_k^o is the Gaussian-centric view direction and nk\mathbf{n}_k is the Gaussian's surface normal. The integrated specular color is:

    ckspecular(ωo)=vk(ωo)∫S2L(ωi)Gs(ωi;qk,σk)dωi\mathbf{c}_k^{\text{specular}}(\boldsymbol{\omega}_o) = v_k(\boldsymbol{\omega}_o) \int_{\mathbb{S}^2} \mathbf{L}(\boldsymbol{\omega}_i) G_s(\boldsymbol{\omega}_i; \mathbf{q}_k, \sigma_k) \mathrm{d}\boldsymbol{\omega}_i

    where vk(ωo)∈(0,1)v_k(\boldsymbol{\omega}_o) \in (0, 1) is a decoded learnable view-dependent visibility factor that accounts for Fresnel reflectance, geometric shadowing, and occlusion without explicit ray tracing. Under continuous illumination, this integral is evaluated in constant real time with a single mipmap texture lookup on prefiltered environment maps.

  4. Knowl 4 — View-Conditioned Normals for Unified Surface and Fiber Specular Reflection

    model/method

    Human heads contain both 2D surface structures (skin) and thin 1D cylindrical fibers (hair strands). Fiber reflections do not adhere to fixed surface reflection equations when the viewpoint rotates around the fiber tangent axis. To model both skin and hair within a single framework, a view-conditioned decoder Dcv\mathcal{D}_{cv} predicts a view-dependent normal residual δnk\delta \mathbf{n}_k alongside the view-dependent visibility vkv_k:

    {δnk,vk}k=1M=Dcv(z,e{l,r},ωo;Θcv)\{\delta \mathbf{n}_k, v_k\}_{k=1}^M = \mathcal{D}_{cv}(\mathbf{z}, \mathbf{e}_{\{l,r\}}, \boldsymbol{\omega}_o; \Theta_{cv})

    where z\mathbf{z} is the expression latent code, e{l,r}\mathbf{e}_{\{l,r\}} are gaze vectors, ωo\boldsymbol{\omega}_o is the viewing direction, and Θcv\Theta_{cv} are network parameters. The effective surface normal nk\mathbf{n}_k is computed by perturbing the barycentrically interpolated base mesh normal n^k\hat{\mathbf{n}}_k:

    nk=n^k+δnk∥n^k+δnk∥\mathbf{n}_k = \frac{\hat{\mathbf{n}}_k + \delta \mathbf{n}_k}{\|\hat{\mathbf{n}}_k + \delta \mathbf{n}_k\|}

    For flat skin regions, the learned network sets δnk\delta \mathbf{n}_k such that nk\mathbf{n}_k remains static under view angle shifts. For hair fibers, nk\mathbf{n}_k dynamically rotates around the hair tangent axis in response to view changes ωo\boldsymbol{\omega}_o, enabling accurate anisotropic-like specular highlights without explicitly reconstructing micro-geometry or classifying material categories.

  5. Knowl 5 — Relightable Explicit Eye Model with Refraction Adaptation

    model/method

    To reproduce sharp corneal reflections and enable explicit gaze control, each eye is modeled as a blend of two spheres parameterized by E={re,rc,d,ce}E = \{r_e, r_c, d, \mathbf{c}_e\}, where rer_e is eyeball radius, rcr_c is cornea radius, dd is the optical axis offset from eyeball center to cornea center, and ce\mathbf{c}_e is eyeball center in canonical head space.

    To avoid training collapse caused by saturated, sub-pixel point light glints on the reflective cornea, two constraints are imposed on the MeM_e eye Gaussians:

    1. Gaussian spatial translations tk\mathbf{t}_k are frozen directly onto the parametric eyeball surface mesh.
    2. Gaussian normals nk\mathbf{n}_k are locked to the analytical eyeball surface normals.

    To account for view-dependent optical refraction of the iris through the cornea without ray tracing, eye diffuse albedo is conditioned on view direction via specialized eye decoders Dei\mathcal{D}_{ei} (view-independent) and Dev\mathcal{D}_{ev} (view-dependent):

    {Rk,sk,ok,dkc,dkm,σk}k=1Me=Dei(e,hp,hr;Θei)\{\mathbf{R}_k, \mathbf{s}_k, o_k, \mathbf{d}_k^c, \mathbf{d}_k^m, \sigma_k\}_{k=1}^{M_e} = \mathcal{D}_{ei}(\mathbf{e}, \mathbf{h}_p, \mathbf{h}_r; \Theta_{ei})

    {ρk,vk}k=1Me=Dev(e,hp,hr,ωo;Θev)\{\boldsymbol{\rho}_k, v_k\}_{k=1}^{M_e} = \mathcal{D}_{ev}(\mathbf{e}, \mathbf{h}_p, \mathbf{h}_r, \boldsymbol{\omega}_o; \Theta_{ev})

    where e\mathbf{e} is eye gaze, hp∈R3\mathbf{h}_p \in \mathbb{R}^3 and hr∈SO(3)\mathbf{h}_r \in \mathrm{SO}(3) are relative head position and rotation compensating for tracking misalignment, and Θei,Θev\Theta_{ei}, \Theta_{ev} are trainable decoder weights.

  6. Knowl 6 — Time-Multiplexed Illumination Capture Setup

    experimental setup

    Training data is acquired in a multi-camera light-stage system comprising 110 calibrated and synchronized cameras capturing at 4096×26684096 \times 2668 resolution and 90 Hz, surrounding the subject alongside 460 controllable white LED lights. Each participant performs diverse facial expressions, speech sentences, and gaze sequences across approximately 144,000 frames.

    To decouple continuous facial tracking from relighting supervision, time-multiplexed lighting is used across frame capture:

    • Every third frame is illuminated with uniform full-on lighting, providing stable image features to track topologically consistent coarse 3D facial meshes, head poses, and gaze orientations.
    • The remaining two frames are illuminated by either grouped light sources or randomized sparse subsets of 5 individual LED point lights to supervise the radiance transfer models under isolated directional illuminations.

    Tracked coarse meshes, head poses, gaze directions, and unwrapped textures from full-on frames are linearly interpolated onto adjacent partially lit frames during training.

  7. Knowl 7 — Avatar Training Objective and Regularization Losses

    equation

    The trainable avatar parameters (network weights Θ\Theta, static albedos ρ\boldsymbol{\rho}, and eyeball parameters E{l,r}E_{\{l,r\}}) are jointly optimized using the composite loss function:

    L=Lrec+Lreg+λklLkl\mathcal{L} = \mathcal{L}_{\text{rec}} + \mathcal{L}_{\text{reg}} + \lambda_{\text{kl}}\mathcal{L}_{\text{kl}}

    where Lkl\mathcal{L}_{\text{kl}} is the Kullback-Leibler divergence on the CVAE latent space with weight λkl=2.0×10−3\lambda_{\text{kl}} = 2.0 \times 10^{-3}.

    The reconstruction loss Lrec\mathcal{L}_{\text{rec}} combines pixel-level L1L_1 color distance, structural dissimilarity (D-SSIM), and an L2L_2 geometric loss on coarse mesh vertices V′\mathbf{V}':

    Lrec=λl1Ll1+λssimLssim+λgeoLgeo\mathcal{L}_{\text{rec}} = \lambda_{\text{l1}} \mathcal{L}_{\text{l1}} + \lambda_{\text{ssim}} \mathcal{L}_{\text{ssim}} + \lambda_{\text{geo}} \mathcal{L}_{\text{geo}}

    with loss weights λgeo=10\lambda_{\text{geo}} = 10, λl1=10\lambda_{\text{l1}} = 10, and λssim=0.2\lambda_{\text{ssim}} = 0.2.

    The regularization loss Lreg\mathcal{L}_{\text{reg}} constrains Gaussian scales, eliminates negative SH colors, and prevents eye transparency:

    Lreg=λsLs+λc−Lc−+λesLes+λevLev+λeoLeo\mathcal{L}_{\text{reg}} = \lambda_s \mathcal{L}_s + \lambda_{c-} \mathcal{L}_{c-} + \lambda_{es} \mathcal{L}_{es} + \lambda_{ev} \mathcal{L}_{ev} + \lambda_{eo} \mathcal{L}_{eo}

    with weights λs=λc−=λes=1.0×10−2\lambda_s = \lambda_{c-} = \lambda_{es} = 1.0 \times 10^{-2} and λeo=λev=1.0×10−4\lambda_{eo} = \lambda_{ev} = 1.0 \times 10^{-4}, where the individual component terms are defined per-axis scale ss, diffuse color ckdiffuse\mathbf{c}_k^{\text{diffuse}}, opacity oko_k, and visibility vkv_k as:

    Ls=mean(ls),ls={1/max⁡(s,10−7)if s<0.1(s−10.0)2if s>10.00otherwise\mathcal{L}_s = \mathrm{mean}(l_s), \quad l_s = \begin{cases} 1/\max(s, 10^{-7}) & \text{if } s < 0.1 \\ (s - 10.0)^2 & \text{if } s > 10.0 \\ 0 & \text{otherwise} \end{cases}

    Lc−=mean(min⁡(ckdiffuse,0)2)\mathcal{L}_{c-} = \mathrm{mean}(\min(\mathbf{c}_k^{\text{diffuse}}, 0)^2)

    Les=mean(max⁡(s−0.1,0)2),Leo=mean((1−ok)2),Lev=mean((1−vk)2)\mathcal{L}_{es} = \mathrm{mean}(\max(s - 0.1, 0)^2), \quad \mathcal{L}_{eo} = \mathrm{mean}((1 - o_k)^2), \quad \mathcal{L}_{ev} = \mathrm{mean}((1 - v_k)^2)

  8. Knowl 8 — Quantitative Evaluation across Geometry and Appearance Representations

    data/table

    The Relightable Gaussian Codec Avatar was evaluated across combinations of geometry models (Proposed 3D Gaussians with Explicit Eye Model [EEM], 3D Gaussians without EEM, and Mixture of Volumetric Primitives [MVP]) and relightable appearance models (Proposed SH+SG radiance transfer, EyeNeRF SH-based appearance, and Linear light transport network). Evaluation was performed across ~9,000 conversational expression frames and ~100 disgust frames on held-out motion segments (Table 1) and 10 held-out frontally-biased lighting patterns (~1,800 frames, Table 2). Metrics are PSNR, SSIM, and LPIPS computed on face-masked regions:

    Table 1: Held-out Motion Segments
    Geometry Appearance PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow
    Ours w/ EEM EyeNeRF 34.550 0.939 0.115
    Ours w/ EEM Ours 36.501 0.943 0.110
    Ours (w/o EEM) EyeNeRF 35.110 0.938 0.113
    Ours (w/o EEM) Linear 33.831 0.936 0.184
    Ours (w/o EEM) Ours 36.529 0.943 0.111
    MVP EyeNeRF 27.594 0.922 0.151
    MVP Linear 36.294 0.942 0.140
    MVP Ours 35.789 0.943 0.134
    Table 2: Held-out Novel Illumination Patterns
    Geometry Appearance PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow
    Ours w/ EEM EyeNeRF 30.7976 0.828 0.162
    Ours w/ EEM Ours 34.042 0.858 0.148
    Ours (w/o EEM) EyeNeRF 30.836 0.815 0.163
    Ours (w/o EEM) Linear 32.829 0.870 0.202
    Ours (w/o EEM) Ours 33.845 0.831 0.148
    MVP EyeNeRF 28.030 0.812 0.210
    MVP Linear 33.444 0.726 0.192
    MVP Ours 33.778 0.877 0.168

    The combination of 3D Gaussian geometry with the proposed learned radiance transfer model yields top performance, with 3D Gaussians reconstructing thin hair strands and skin pore details superior to voxel primitives (MVP), and the SG formulation outperforming the band-limited SH specular model of EyeNeRF.

  9. Knowl 9 — Self-Supervised Intrinsic Reflectance and Shading Decomposition

    empirical result

    Without explicit direct supervision of surface normals, albedo, or reflectance material properties, training the 3D Gaussian avatar solely on multiview point-light images automatically produces a 3D-consistent intrinsic reflectance decomposition. The reconstructed avatar decomposes rendering into:

    1. Static diffuse albedo maps ρk\boldsymbol{\rho}_k across the face and dynamic albedos for the eyes.
    2. Diffuse shading maps resulting from the spherical harmonics radiance transfer integral over incident illumination, encoding subsurface scattering and global light transport.
    3. Specular reflection fields driven by normalized spherical Gaussian lobes scaled by view-dependent visibility factors vk(ωo)v_k(\boldsymbol{\omega}_o).
    4. Surface normal vector fields nk\mathbf{n}_k that correspond to fine geometric features (facial pores, skin micro-geometry, and hair fiber orientations).
  10. Knowl 10 — System Limitations and Scalability Constraints

    limitation

    The Relightable Gaussian Codec Avatar framework exhibits three primary operational limitations:

    1. Tracking Preprocessing Dependency: The framework relies on topologically consistent coarse mesh tracking and explicit eye gaze estimations as preprocessing steps; errors or failures in upstream tracking can propagate into avatar artifacts.
    2. Controlled Illumination Dependency: Training requires known calibrated light stage lighting with time-multiplexed illumination, preventing direct out-of-the-box application to in-the-wild video captures with uncalibrated ambient lighting.
    3. Multi-Avatar Scalability: Radiance transfer calculations (diffuse SH dot products and specular SG evaluations) are executed per individual 3D Gaussian rather than in screen-space fragment shaders, causing rendering compute to scale linearly with the number of avatars rendered simultaneously.

Coverage note — None omitted; all core contributions including drivable 3D Gaussian geometry, hybrid diffuse-specular radiance transfer, explicit relightable eye modeling, loss formulations, experimental benchmark data, and stated limitations are represented.

References

  1. 1.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan: How to embed images into the stylegan latent space? In Proceedings of the IEEE/CVF international conference on computer vision, pages 4432–4441, 2019.
  2. 2.Thabo Beeler, Bernd Bickel, Paul Beardsley, Bob Sumner, and Markus Gross. High-quality single-shot capture of facial geometry. In ACM SIGGRAPH 2010 papers, pages 1–9. 2010.
  3. 3.Thabo Beeler, Bernd Bickel, Gioacchino Noris, Paul Beardsley, Steve Marschner, Robert W Sumner, and Markus Gross. Coupled 3d reconstruction of sparse facial hair and skin. ACM Transactions on Graphics (ToG), 31(4):1–10, 2012.
  4. 4.Thabo Beeler, Fabian Hahn, Derek Bradley, Bernd Bickel, Paul A Beardsley, Craig Gotsman, Robert W Sumner, and Markus H Gross. High-quality passive facial performance capture using anchor frames. ACM Trans. Graph., 30(4):75, 2011.
  5. 5.Pascal Bérard, Derek Bradley, Markus Gross, and Thabo Beeler. Lightweight eye capture using a parametric model. ACM Transactions on Graphics (TOG), 35(4):1–12, 2016.
  6. 6.Sai Bi, Stephen Lombardi, Shunsuke Saito, Tomas Simon, Shih-En Wei, Kevyn Mcphail, Ravi Ramamoorthi, Yaser Sheikh, and Jason Saragih. Deep relightable appearance models for animatable faces. ACM Transactions on Graphics (TOG), 40(4):1–15, 2021.
  7. 7.Sai Bi, Zexiang Xu, Kalyan Sunkavalli, Miloš Hašan, Yannick Hold-Geoffroy, David Kriegman, and Ravi Ramamoorthi. Deep reflectance volumes: Relightable reconstructions from multi-view photometric images. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16, pages 294–311. Springer, 2020.
  8. 8.Timo Bolkart, Tianye Li, and Michael J Black. Instant multiview head capture through learnable registration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 768–779, 2023.
  9. 9.Derek Bradley, Wolfgang Heidrich, Tiberiu Popa, and Alla Sheffer. High resolution passive facial performance capture. In ACM SIGGRAPH 2010 papers, pages 1–10. 2010.
  10. 10.Paul Debevec, Tim Hawkins, Chris Tchou, Haarm-Pieter Duiker, Westley Sarokin, and Mark Sagar. Acquiring the reflectance field of a human face. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 145–156, 2000.
  11. 11.Boyang Deng, Yifan Wang, and Gordon Wetzstein. Lumigan: Unconditional generation of relightable 3d human faces. arXiv preprint arXiv:2304.13153, 2023.
  12. 12.Yasutaka Furukawa and Jean Ponce. Accurate, dense, and robust multiview stereopsis. IEEE transactions on pattern analysis and machine intelligence, 32(8):1362–1376, 2009.
  13. 13.Duan Gao, Guojun Chen, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. Deferred neural lighting: free-viewpoint relighting from unstructured photographs. ACM Transactions on Graphics (TOG), 39(6):1–15, 2020.
  14. 14.Stephan J Garbin, Marek Kowalski, Virginia Estellers, Stanislaw Szymanowicz, Shideh Rezaeifar, Jingjing Shen, Matthew Johnson, and Julien Valentin. Voltemorph: Real-time, controllable and generalisable animation of volumetric representations. arXiv preprint arXiv:2208.00949, 2022.
  15. 15.Pablo Garrido, Michael Zollhöfer, Chenglei Wu, Derek Bradley, Patrick Pérez, Thabo Beeler, and Christian Theobalt. Corrective 3d reconstruction of lips from monocular video. ACM Trans. Graph., 35(6):219–1, 2016.
  16. 16.Abhijeet Ghosh, Graham Fyffe, Borom Tunwattanapong, Jay Busch, Xueming Yu, and Paul Debevec. Multiview face capture using polarized spherical gradient illumination. ACM Transactions on Graphics (TOG), 30(6):1–10, 2011.
  17. 17.Paul Green, Jan Kautz, Wojciech Matusik, and Frédo Durand. View-dependent precomputed light transport using nonlinear gaussian function approximations. In Proceedings of the 2006 symposium on Interactive 3D graphics and games, pages 7–14, 2006.
  18. 18.Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts-Escolano, Rohit Pandey, Jason Dourgarian, et al. The relightables: Volumetric performance capture of humans with realistic relighting. ACM Transactions on Graphics (ToG), 38(6):1–19, 2019.
  19. 19.Liwen Hu, Chongyang Ma, Linjie Luo, and Hao Li. Robust hair capture using simulated examples. ACM Transactions on Graphics (TOG), 33(4):1–10, 2014.
  20. 20.Shun Iwase, Shunsuke Saito, Tomas Simon, Stephen Lombardi, Timur Bagautdinov, Rohan Joshi, Fabian Prada, Takaaki Shiratori, Yaser Sheikh, and Jason Saragih. Relightablehands: Efficient neural relighting of articulated hand models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16663–16673, 2023.
  21. 21.James T Kajiya. The rendering equation. In Proceedings of the 13th annual conference on Computer graphics and interactive techniques, pages 143–150, 1986.
  22. 22.James T Kajiya and Timothy L Kay. Rendering fur with three dimensional textures. ACM Siggraph Computer Graphics, 23(3):271–280, 1989.
  23. 23.Jan Kautz, Pere-Pau Vázquez, Wolfgang Heidrich, and Hans-Peter Seidel. A unified approach to prefiltered environment maps. In Rendering Techniques 2000: Proceedings of the Eurographics Workshop in Brno, Czech Republic, June 26–28, 2000 11, pages 185–196. Springer, 2000.
  24. 24.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023.
  25. 25.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  26. 26.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  27. 27.Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view radiance field reconstruction of human heads. arXiv preprint arXiv:2305.03027, 2023.
  28. 28.Georgios Kopanas, Thomas Leimkühler, Gilles Rainer, Clément Jambon, and George Drettakis. Neural point catacaustics for novel-view synthesis of reflections. ACM Transactions on Graphics (TOG), 41(6):1–15, 2022.
  29. 29.Georgios Kopanas, Julien Philip, Thomas Leimkühler, and George Drettakis. Point-based neural rendering with per-view optimization. In Computer Graphics Forum, volume 40, pages 29–43. Wiley Online Library, 2021.
  30. 30.Mathieu Lamarre, John P Lewis, and Etienne Danvoye. Face stabilization by mode pursuit for avatar construction. In 2018 International Conference on Image and Vision Computing New Zealand (IVCNZ), pages 1–6. IEEE, 2018.
  31. 31.Alexandros Lattas, Stylianos Moschoglou, Baris Gecer, Stylianos Ploumpis, Vasileios Triantafyllou, Abhijeet Ghosh, and Stefanos Zafeiriou. Avatarme: Realistically renderable 3d facial reconstruction” in-the-wild”. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 760–769, 2020.
  32. 32.Alexandros Lattas, Stylianos Moschoglou, Stylianos Ploumpis, Baris Gecer, Abhijeet Ghosh, and Stefanos Zafeiriou. Avatarme++: Facial shape and brdf inference with photorealistic rendering-aware gans. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12):9269–9284, 2021.
  33. 33.Gengyan Li, Abhimitra Meka, Franziska Mueller, Marcel C Buehler, Otmar Hilliges, and Thabo Beeler. Eyenerf: a hybrid representation for photorealistic synthesis, animation and relighting of human eyes. ACM Transactions on Graphics (TOG), 41(4):1–16, 2022.
  34. 34.Jiaman Li, Zhengfei Kuang, Yajie Zhao, Mingming He, Kalle Bladin, and Hao Li. Dynamic facial asset and rig generation from a single scan. ACM Transactions on Graphics (TOG), 39:1 – 18, 2020.
  35. 35.Junxuan Li, Shunsuke Saito, Tomas Simon, Stephen Lombardi, Hongdong Li, and Jason Saragih. Megane: Morphable eyeglass and avatar network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12769–12779, 2023.
  36. 36.Ruilong Li, Kalle Bladin, Yajie Zhao, Chinmay Chinara, Owen Ingraham, Pengda Xiang, Xinglei Ren, Pratusha Bhuvana Prasad, Bipin Kishore, Jun Xing, and Hao Li. Learning formation of physically-based face attributes. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3407–3416, 2020.
  37. 37.Tianye Li, Shichen Liu, Timo Bolkart, Jiayi Liu, Hao Li, and Yajie Zhao. Topologically consistent multi-view face inference using volumetric sampling. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3824–3834, 2021.
  38. 38.Shichen Liu, Yunxuan Cai, Haiwei Chen, Yichao Zhou, and Yajie Zhao. Rapid face asset acquisition with recurrent feature alignment. ACM Transactions on Graphics (TOG), 41(6):1–17, 2022.
  39. 39.Stephen Lombardi, Jason Saragih, Tomas Simon, and Yaser Sheikh. Deep appearance models for face rendering. ACM Transactions on Graphics (ToG), 37(4):1–13, 2018.
  40. 40.Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural volumes: Learning dynamic renderable volumes from images. arXiv preprint arXiv:1906.07751, 2019.
  41. 41.Stephen Lombardi, Tomas Simon, Gabriel Schwartz, Michael Zollhoefer, Yaser Sheikh, and Jason Saragih. Mixture of volumetric primitives for efficient neural rendering. ACM Transactions on Graphics (ToG), 40(4):1–13, 2021.
  42. 42.Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. arXiv preprint arXiv:2308.09713, 2023.
  43. 43.Linjie Luo, Hao Li, and Szymon Rusinkiewicz. Structure-aware hair capture. ACM Transactions on Graphics (TOG), 32(4):1–12, 2013.
  44. 44.Shugao Ma, Tomas Simon, Jason Saragih, Dawei Wang, Yuecheng Li, Fernando De La Torre, and Yaser Sheikh. Pixel codec avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 64–73, 2021.
  45. 45.Wan-Chun Ma, Tim Hawkins, Pieter Peers, Charles-Felix Chabert, Malte Weiss, Paul E Debevec, et al. Rapid acquisition of specular and diffuse normal maps from polarized spherical gradient illumination. Rendering Techniques, 2007(9):10, 2007.
  46. 46.Stephen R Marschner, Henrik Wann Jensen, Mike Cammarano, Steve Worley, and Pat Hanrahan. Light scattering from human hair fibers. ACM Transactions on Graphics (TOG), 22(3):780–791, 2003.
  47. 47.David K McAllister, Anselmo Lastra, and Wolfgang Heidrich. Efficient rendering of spatial bi-directional reflectance distribution functions. In Proceedings of the ACM SIGGRAPH/EUROGRAPHICS conference on Graphics hardware, pages 79–88, 2002.
  48. 48.Abhimitra Meka, Christian Haene, Rohit Pandey, Michael Zollhöfer, Sean Fanello, Graham Fyffe, Adarsh Kowdle, Xueming Yu, Jay Busch, Jason Dourgarian, et al. Deep reflectance fields: high-quality facial reflectance field inference from color gradient illumination. ACM Transactions on Graphics (TOG), 38(4):1–12, 2019.
  49. 49.Abhimitra Meka, Rohit Pandey, Christian Haene, Sergio Orts-Escolano, Peter Barnum, Philip David-Son, Daniel Erickson, Yinda Zhang, Jonathan Taylor, Sofien Bouaziz, et al. Deep relightable textures: volumetric performance capture with neural rendering. ACM Transactions on Graphics (TOG), 39(6):1–21, 2020.
  50. 50.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  51. 51.Erick Miller and Dmitriy Pinskiy. Realistic eye motion using procedural geometric methods. In SIGGRAPH 2009: Talks, pages 1–1. 2009.
  52. 52.Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Müller, and Sanja Fidler. Extracting triangular 3d models, materials, and lighting from images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8280–8290, 2022.
  53. 53.Koki Nagano, Graham Fyffe, Oleg Alexander, Jernej Barbic, Hao Li, Abhijeet Ghosh, and Paul E Debevec. Skin microstructure deformation with displacement map convolution. ACM Trans. Graph., 34(4):109–1, 2015.
  54. 54.Giljoo Nam, Joo Ho Lee, Diego Gutierrez, and Min H Kim. Practical svbrdf acquisition of 3d objects with unstructured flash photography. ACM Transactions on Graphics (TOG), 37(6):1–12, 2018.
  55. 55.Giljoo Nam, Chenglei Wu, Min H Kim, and Yaser Sheikh. Strand-accurate multi-view hair capture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 155–164, 2019.
  56. 56.Ren Ng, Ravi Ramamoorthi, and Pat Hanrahan. All-frequency shadows using non-linear wavelet lighting approximation. In ACM SIGGRAPH 2003 Papers, pages 376–381. 2003.
  57. 57.Sergio Orts-Escolano, Christoph Rhemann, Sean Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim, Philip L Davidson, Sameh Khamis, Mingsong Dou, et al. Holoportation: Virtual 3d teleportation in real-time. In Proceedings of the 29th annual symposium on user interface software and technology, pages 741–754, 2016.
  58. 58.Rohit Pandey, Sergio Orts Escolano, Chloe Legendre, Christian Haene, Sofien Bouaziz, Christoph Rhemann, Paul Debevec, and Sean Fanello. Total relighting: learning to relight portraits for background replacement. ACM Transactions on Graphics (TOG), 40(4):1–21, 2021.
  59. 59.Foivos Paraperas Papantoniou, Alexandros Lattas, Stylianos Moschoglou, and Stefanos Zafeiriou. Relightify: Relightable 3d faces from a single image via diffusion models. arXiv preprint arXiv:2305.06077, 2023.
  60. 60.Sylvain Paris, Will Chang, Oleg I Kozhushnyan, Wojciech Jarosz, Wojciech Matusik, Matthias Zwicker, and Frédo Durand. Hair photobooth: geometric and photometric acquisition of real hairstyles. ACM Trans. Graph., 27(3):30, 2008.
  61. 61.Frederic I Parke and Keith Waters. Computer facial animation. CRC press, 2008.
  62. 62.Pieter Peers, Naoki Tamura, Wojciech Matusik, and Paul Debevec. Post-production facial performance relighting using reflectance transfer. ACM Transactions on Graphics (TOG), 26(3):52–es, 2007.
  63. 63.Gilles Rainer, Adrien Bousseau, Tobias Ritschel, and George Drettakis. Neural precomputed radiance transfer. In Computer Graphics Forum, volume 41, pages 365–378. Wiley Online Library, 2022.
  64. 64.Ravi Ramamoorthi and Pat Hanrahan. An efficient representation for irradiance environment maps. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 497–500, 2001.
  65. 65.Anurag Ranjan, Kwang Moo Yi, Jen-Hao Rick Chang, and Oncel Tuzel. Facelit: Neural 3d relightable faces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8619–8628, 2023.
  66. 66.Kripasindhu Sarkar, Marcel C. Buehler, Gengyan Li, Daoye Wang, Delio Vicini, Jerémy Riviere, Yinda Zhang, Sergio Orts-Escolano, Paulo Gotardo, Thabo Beeler, and Abhimitra Meka. Litnerf: Intrinsic radiance decomposition for high-quality view synthesis and relighting of faces. In ACM SIGGRAPH Asia 2023, 2023.
  67. 67.Gabriel Schwartz, Shih-En Wei, Te-Li Wang, Stephen Lombardi, Tomas Simon, Jason Saragih, and Yaser Sheikh. The eyes have it: An integrated eye and face model for photorealistic facial animation. ACM Transactions on Graphics (TOG), 39(4):91–1, 2020.
  68. 68.Mike Seymour, Chris Evans, and Kim Libreri. Meet mike: epic avatars. In ACM SIGGRAPH 2017 VR Village, pages 1–2. 2017.
  69. 69.Peter-Pike Sloan, Jan Kautz, and John Snyder. Precomputed radiance transfer for real-time rendering in dynamic, low-frequency lightingevironments. ACM Trans. Graph., 21(3):527–536, jul 2002.
  70. 70.Tiancheng Sun, Jonathan T Barron, Yun-Ta Tsai, Zexiang Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay Busch, Paul Debevec, and Ravi Ramamoorthi. Single image portrait relighting. ACM Transactions on Graphics (TOG), 38(4):1–12, 2019.
  71. 71.Ayush Tewari, Mohamed Elgharib, Gaurav Bharaj, Florian Bernard, Hans-Peter Seidel, Patrick Pérez, Michael Zollhöfer, and Christian Theobalt. Stylerig: Rigging stylegan for 3d control over portrait images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6142–6151, 2020.
  72. 72.Justus Thies, Michael Zollhöfer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (TOG), 38(4):1–12, 2019.
  73. 73.Yu-Ting Tsai and Zen-Chung Shih. All-frequency precomputed radiance transfer using spherical radial basis functions and clustered tensor approximation. ACM Transactions on graphics (TOG), 25(3):967–976, 2006.
  74. 74.Jiaping Wang, Peiran Ren, Minmin Gong, John Snyder, and Baining Guo. All-frequency rendering of dynamic, spatially-varying reflectance. In ACM SIGGRAPH Asia 2009 papers, pages 1–10. 2009.
  75. 75.Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4):600–612, 2004.
  76. 76.Zhibo Wang, Xin Yu, Ming Lu, Quan Wang, Chen Qian, and Feng Xu. Single image portrait relighting via explicit multiple reflectance channel modeling. ACM Transactions on Graphics (TOG), 39(6):1–13, 2020.
  77. 77.Shih-En Wei, Jason Saragih, Tomas Simon, Adam W Harley, Stephen Lombardi, Michal Perdoch, Alexander Hypes, Dawei Wang, Hernan Badino, and Yaser Sheikh. Vr facial animation via multiview image translation. ACM Transactions on Graphics (TOG), 38(4):1–16, 2019.
  78. 78.Tim Weyrich, Wojciech Matusik, Hanspeter Pfister, Bernd Bickel, Craig Donner, Chien Tu, Janet McAndless, Jinho Lee, Addy Ngan, Henrik Wann Jensen, et al. Analysis of human faces using a measurement-based skin reflectance model. ACM Transactions on Graphics (ToG), 25(3):1013–1024, 2006.
  79. 79.Chenglei Wu, Derek Bradley, Pablo Garrido, Michael Zollhöfer, Christian Theobalt, Markus H Gross, and Thabo Beeler. Model-based teeth reconstruction. ACM Trans. Graph., 35(6):220–1, 2016.
  80. 80.Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. arXiv preprint arXiv:2310.08528, 2023.
  81. 81.Kun Xu, Wei-Lun Sun, Zhao Dong, Dan-Yong Zhao, Run-Dong Wu, and Shi-Min Hu. Anisotropic spherical gaussians. ACM Transactions on Graphics (TOG), 32(6):1–11, 2013.
  82. 82.Yuelang Xu, Hongwen Zhang, Lizhen Wang, Xiaochen Zhao, Han Huang, Guojun Qi, and Yebin Liu. Latentavatar: Learning latent expression code for expressive neural head avatar. arXiv preprint arXiv:2305.01190, 2023.
  83. 83.Yingyan Xu, Gaspard Zoss, Prashanth Chandran, Markus Gross, Derek Bradley, and Paulo Gotardo. Renerf: Relightable neural radiance fields with nearfield lighting. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22581–22591, 2023.
  84. 84.Zexiang Xu, Kalyan Sunkavalli, Sunil Hadap, and Ravi Ramamoorthi. Deep image-based relighting from optimal sparse samples. ACM Transactions on Graphics (ToG), 37(4):1–13, 2018.
  85. 85.Zilin Xu, Zheng Zeng, Lifan Wu, Lu Wang, and Ling-Qi Yan. Lightweight neural basis functions for all-frequency shading. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9, 2022.
  86. 86.Shugo Yamaguchi, Shunsuke Saito, Koki Nagano, Yajie Zhao, Weikai Chen, Kyle Olszewski, Shigeo Morishima, and Hao Li. High-fidelity facial reflectance and geometry inference from an unconstrained image. ACM Transactions on Graphics (TOG), 37(4):1–14, 2018.
  87. 87.Haotian Yang, Mingwu Zheng, Wanquan Feng, Haibin Huang, Yu-Kun Lai, Pengfei Wan, Zhongyuan Wang, and Chongyang Ma. Towards practical capture of high-fidelity relightable avatars. In SIGGRAPH Asia 2023 Conference Proceedings, 2023.
  88. 88.Yu-Ying Yeh, Koki Nagano, Sameh Khamis, Jan Kautz, Ming-Yu Liu, and Ting-Chun Wang. Learning to relight portrait images via a virtual light stage and synthetic-to-real adaptation. ACM Transactions on Graphics (TOG), 41(6):1–21, 2022.
  89. 89.Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5453–5462, 2021.
  90. 90.Li Zhang, Noah Snavely, Brian Curless, and Steven M Seitz. Spacetime faces: high resolution capture for modeling and animation. In ACM SIGGRAPH 2004 Papers, pages 548–558. 2004.
  91. 91.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018.
  92. 92.Xiuming Zhang, Sean Fanello, Yun-Ta Tsai, Tiancheng Sun, Tianfan Xue, Rohit Pandey, Sergio Orts-Escolano, Philip Davidson, Christoph Rhemann, Paul Debevec, et al. Neural light transport for relighting and view synthesis. ACM Transactions on Graphics (TOG), 40(1):1–17, 2021.
  93. 93.Xiuming Zhang, Pratul P Srinivasan, Boyang Deng, Paul Debevec, William T Freeman, and Jonathan T Barron. Nerfactor: Neural factorization of shape and reflectance under an unknown illumination. ACM Transactions on Graphics (ToG), 40(6):1–18, 2021.
  94. 94.Xiaochen Zhao, Lizhen Wang, Jingxiang Sun, Hongwen Zhang, Jinli Suo, and Yebin Liu. Havatar: High-fidelity head avatar via facial model conditioned neural radiance field. ACM Transactions on Graphics, 2023.
  95. 95.Yufeng Zheng, Wang Yifan, Gordon Wetzstein, Michael J Black, and Otmar Hilliges. Pointavatar: Deformable point-based head avatars from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21057–21067, 2023.
  96. 96.Wojciech Zielonka, Timo Bolkart, and Justus Thies. Instant volumetric head avatars. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4574–4584, 2023.
  97. 97.Matthias Zwicker, Hanspeter Pfister, Jeroen Van Baar, and Markus Gross. Ewa splatting. IEEE Transactions on Visualization and Computer Graphics, 8(3):223–238, 2002.

Citation

MLA
Saito, S., et al. “Relightable Gaussian Codec Avatars”. arXiv, 2023, http://arxiv.org/abs/2312.03704v2.
APA
Saito, S., Schwartz, G., Simon, T., Li, J., & Nam, G. (2023). Relightable Gaussian Codec Avatars. arXiv. http://arxiv.org/abs/2312.03704v2
Chicago
Saito, S., G. Schwartz, T. Simon, J. Li, and G. Nam. 2023. “Relightable Gaussian Codec Avatars”. arXiv. http://arxiv.org/abs/2312.03704v2.
Harvard
Saito, S. et al. (2023) “Relightable Gaussian Codec Avatars”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.03704v2.
Vancouver
1. Saito S, Schwartz G, Simon T, Li J, Nam G (2023) Relightable Gaussian Codec Avatars. arXiv

BibTeX

@article{saito2023relightable,
  title = {Relightable Gaussian Codec Avatars},
  author = {Saito, Shunsuke and Schwartz, Gabriel and Simon, Tomas and Li, Junxuan and Nam, Giljoo},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.03704v2},
  eprint = {2312.03704}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE