Relightable Gaussian Codec Avatars
Shunsuke SaitoGabriel SchwartzTomas SimonJunxuan LiGiljoo Nam
Combines animatable 3D Gaussian splatting with learnable radiance transfer to achieve real-time, all-frequency relighting of dynamic head avatars featuring sub-millimeter geometric details and explicit gaze control.
Generating realistic, animatable digital human avatars that can adapt seamlessly to any lighting environment is essential for telecommunication, virtual reality, and interactive gaming. However, current systems struggle to produce photorealistic results in real time. Human heads comprise varied, complex materials—such as translucent skin, multi-layered reflective eyes, and fine hair fibers—which traditionally require prohibitively expensive computation to render under changing illumination conditions, often blurring fine features like individual hair strands.
The article demonstrates a novel framework called Relightable Gaussian Codec Avatars, which creates dynamic, highly detailed, and relightable digital head models capable of running in real time. The primary objective is to evaluate whether combining 3D Gaussian point-based geometry with a learned radiance transfer appearance model can accurately reproduce complex global illumination and sharp reflections across all head materials simultaneously.
To accomplish this, the authors captured multi-view performance data from human subjects across approximately 144,000 frames using a synchronized 110-camera, 460-LED light-stage rig. The framework employs 3D Gaussian Splatting anchored to a template mesh to efficiently render intricate sub-millimeter geometry, a neural radiance transfer function decomposing diffuse and specular light components via spherical harmonics and spherical Gaussians, and an explicit dual-sphere geometric eye model to control gaze and realistic cornea reflections.
The key findings confirm substantial performance improvements across multiple metrics. First, the proposed method consistently outperformed existing real-time baselines—such as volumetric primitives and linear neural networks—achieving superior visual quality with peak signal-to-noise ratios exceeding 36 dB on held-out expressions. Second, the 3D Gaussian geometry captured sub-millimeter details like pores and distinct hair strands that previous volumetric methods blurred. Third, the spherical Gaussian specular model achieved sharp, all-frequency environmental reflections without the computational bottlenecks or visual flickering seen in prior spherical harmonic formulations. Finally, the explicit eye model successfully disentangled eye gaze control while rendering realistic cornea reflections.
These results demonstrate that photorealistic, fully relightable avatars can run efficiently on consumer-grade virtual reality headsets without requiring specialized rendering workarounds or expensive per-frame offline ray tracing. By maintaining the mathematical linearity of light transport, the model renders both point-source lights and complex environment maps instantaneously, significantly lowering runtime performance costs.
Organizations developing virtual presence or interactive graphics systems should consider adopting 3D Gaussian splatting paired with learned radiance transfer for avatar rendering pipelines. However, decision-makers should note that current avatar construction relies on highly controlled light-stage data capture and multi-view coarse tracking pipelines. Future efforts should focus on enabling end-to-end training that reduces preprocessing dependencies and exploring per-pixel fragment shader offloading to scale rendering capacity when animating multiple avatars simultaneously.
- Paper: 3D Gaussian Splatting for Real-Time Radiance Field Rendering, Bernhard Kerbl et al. (2023). This foundational work introduces 3D Gaussian Splatting and real-time differentiable rasterization, which forms the core geometric representation utilized in Relightable Gaussian Codec Avatars.
- Paper: Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields, Dor Verbin et al. (2022). This paper establishes principled decomposition and reflection modeling using directional representations for view-dependent specular highlights that directly influenced neural radiance transfer formulation.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). This work introduces the FLAME parametric head model, providing the foundational underlying mesh rigging and face parameterization on which avatar Gaussians are anchored.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). This study demonstrates learning disentangled parametric head representations for identity and expression tracking, a core requirement for driving controllable codec avatars.
- Paper: Rendering synthetic objects into real scenes: bridging traditional and image-based graphics with global illumination and high dynamic range photography, P. Debevec (1998). This classic text develops the principles of image-based relighting and environment map radiance transport that underpin real-time relightable avatar pipelines.
- Paper: Models of light reflection for computer synthesized pictures, J. Blinn (1977). This seminal graphics paper formalizes microfacet specular reflection models, serving as the physical foundation for the spherical Gaussian specular radiance transfer functions.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). This work directly extends the concept of binding 3D Gaussians to parametric face meshes to reconstruct and animate photorealistic, controllable head avatars from multi-view video.
- Paper: FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models, Shivangi Aneja et al. (2024). This paper applies generative diffusion modeling to drive neural parametric head motion directly from audio signals, building on advanced animatable head models.
- Paper: 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis, Zhicheng Lu et al. (2024). This research continues the development of deformable Gaussian splatting by incorporating explicit 3D geometric structure and continuous rotations to improve dynamic view synthesis.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). This framework generalizes talking avatar synthesis to one-shot single-image scenarios by learning expressive motion adapters and full portrait composition.
