FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models
Shivangi AnejaJustus ThiesAngela DaiMatthias Nießner
Presents FaceTalk, a latent diffusion framework that generates expressive, temporally consistent 3D head animations from audio signals by synthesizing motion within the compact expression space of volumetric Neural Parametric Head Models.
Creating realistic, speech-driven 3D animations of human faces is essential for digital media, video games, virtual avatars, and automated assistants. Conventional 3D animation systems typically rely on linear blendshape templates, which produce static meshes that struggle to capture dynamic, fine-scale facial details such as skin creasing, wrinkles, eye blinks, and diverse hairstyles. Furthermore, existing methods often generate rigid upper-face expressions and suffer from temporal unnaturalness or jitter. The article presents FaceTalk, the first audio-driven generative framework that uses a transformer-based latent diffusion model to synthesize high-fidelity, temporally coherent 3D volumetric head animations directly from speech signals.
To achieve realistic motion without the constraints of fixed-topology meshes, the approach couples input audio embeddings extracted using a pretrained speech model with the latent expression space of Neural Parametric Head Models (NPHMs)—a volumetric representation capable of handling complex geometry and expressions. Because paired datasets of audio and volumetric facial expressions did not exist, the researchers built a training dataset of 1,000 sequences by fitting multi-view video recordings to the NPHM space using geometric and temporal regularization. The core architecture uses a multi-head transformer decoder to progressively denoise expression sequences conditioned on audio, applying an expression-audio alignment mask, feature modulation, data augmentation to encourage expressive diversity, and post-generation Gaussian smoothing to eliminate structural head wobble.
FaceTalk significantly outperformed existing state-of-the-art animation techniques across standard objective and subjective benchmarks. In quantitative evaluations, FaceTalk reduced the Fréchet Inception Distance on mouth regions by more than 75% compared to the strongest baselines (achieving a score of 40.69 compared to over 200 for competing methods) and improved lip-sync error and general visual quality metrics. In a perceptual user study with 40 participants across 15 unseen audio clips, over 70% to 75% of respondents preferred FaceTalk over competing baseline models across overall animation quality, lip synchronization, and facial realism. The ablation analyses confirmed that the diffusion formulation, feature-wise modulation, and explicit audio-expression alignment were crucial for preventing expression collapse and maintaining phoneme-level synchronization.
These findings demonstrate that volumetric latent diffusion provides a scalable, identity-agnostic mechanism to automate high-fidelity 3D facial motion without manual artist intervention. The model accurately transfers speech-driven expressions across distinct facial identities while allowing adjustable expression intensity via classifier-free guidance. For production and content-creation pipelines, this shift reduces the time and cost required to generate expressive 3D character performances while delivering substantially higher realism than legacy morphable models.
Organizations planning to adopt this workflow should focus near-term investments on non-real-time production pipelines, such as offline rendering, cinematic content creation, and pre-recorded digital media. Further engineering is required before deploying the system in real-time or low-latency interactive applications, as the iterative diffusion denoising steps introduce computational latency. Future development should incorporate efficient sampling acceleration and expand the generative architecture to jointly synthesize diverse facial identities aligned directly with audio-inferred characteristics.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). It introduces Neural Parametric Head Models (NPHM), providing the foundational 3D volumetric head representation and expression latent space that FaceTalk directly builds upon.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). It establishes generative diffusion modeling tailored for continuous 3D human motion, laying the conceptual basis for FaceTalk's motion diffusion architecture.
- Paper: Executing your Commands via Motion Diffusion in Latent Space, Xin Chen et al. (2023). It introduces latent-space motion diffusion models, providing the core methodological blueprint for performing conditional diffusion synthesis within compressed parametric motion spaces.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). It defines the widely used FLAME statistical head model for disentangling identity, pose, and expression, contextualizing the parametric head modeling paradigms extended by NPHM and FaceTalk.
- Paper: Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation, Lingting Zhu et al. (2023). It demonstrates audio-conditioned diffusion models for synthesizing coherent human communicative motion, providing foundational techniques for cross-modal speech-to-motion generation.
- Paper: GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians, Shenhan Qian et al. (2024). It extends 3D head avatar modeling by rigging discrete 3D Gaussian primitives onto parametric face representations for photorealistic rendering and novel expression animation.
- Paper: Relightable Gaussian Codec Avatars, Shunsuke Saito et al. (2024). It advances dynamic head avatar synthesis by introducing relightable 3D Gaussian codec representations capable of rendering sub-millimeter dynamic facial details under varying illumination.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). It applies audio-to-motion generative modeling to complete one-shot 3D talking portrait synthesis incorporating head, torso, and background elements.
