FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models

Shivangi AnejaJustus ThiesAngela DaiMatthias Nießner

article2024CVPR72 citations

Presents FaceTalk, a latent diffusion framework that generates expressive, temporally consistent 3D head animations from audio signals by synthesizing motion within the compact expression space of volumetric Neural Parametric Head Models.

Listen

Creating realistic, speech-driven 3D animations of human faces is essential for digital media, video games, virtual avatars, and automated assistants. Conventional 3D animation systems typically rely on linear blendshape templates, which produce static meshes that struggle to capture dynamic, fine-scale facial details such as skin creasing, wrinkles, eye blinks, and diverse hairstyles. Furthermore, existing methods often generate rigid upper-face expressions and suffer from temporal unnaturalness or jitter. The article presents FaceTalk, the first audio-driven generative framework that uses a transformer-based latent diffusion model to synthesize high-fidelity, temporally coherent 3D volumetric head animations directly from speech signals.

To achieve realistic motion without the constraints of fixed-topology meshes, the approach couples input audio embeddings extracted using a pretrained speech model with the latent expression space of Neural Parametric Head Models (NPHMs)—a volumetric representation capable of handling complex geometry and expressions. Because paired datasets of audio and volumetric facial expressions did not exist, the researchers built a training dataset of 1,000 sequences by fitting multi-view video recordings to the NPHM space using geometric and temporal regularization. The core architecture uses a multi-head transformer decoder to progressively denoise expression sequences conditioned on audio, applying an expression-audio alignment mask, feature modulation, data augmentation to encourage expressive diversity, and post-generation Gaussian smoothing to eliminate structural head wobble.

FaceTalk significantly outperformed existing state-of-the-art animation techniques across standard objective and subjective benchmarks. In quantitative evaluations, FaceTalk reduced the Fréchet Inception Distance on mouth regions by more than 75% compared to the strongest baselines (achieving a score of 40.69 compared to over 200 for competing methods) and improved lip-sync error and general visual quality metrics. In a perceptual user study with 40 participants across 15 unseen audio clips, over 70% to 75% of respondents preferred FaceTalk over competing baseline models across overall animation quality, lip synchronization, and facial realism. The ablation analyses confirmed that the diffusion formulation, feature-wise modulation, and explicit audio-expression alignment were crucial for preventing expression collapse and maintaining phoneme-level synchronization.

These findings demonstrate that volumetric latent diffusion provides a scalable, identity-agnostic mechanism to automate high-fidelity 3D facial motion without manual artist intervention. The model accurately transfers speech-driven expressions across distinct facial identities while allowing adjustable expression intensity via classifier-free guidance. For production and content-creation pipelines, this shift reduces the time and cost required to generate expressive 3D character performances while delivering substantially higher realism than legacy morphable models.

Organizations planning to adopt this workflow should focus near-term investments on non-real-time production pipelines, such as offline rendering, cinematic content creation, and pre-recorded digital media. Further engineering is required before deploying the system in real-time or low-latency interactive applications, as the iterative diffusion denoising steps introduce computational latency. Future development should incorporate efficient sampling acceleration and expand the generative architecture to jointly synthesize diverse facial identities aligned directly with audio-inferred characteristics.

  • Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). It introduces Neural Parametric Head Models (NPHM), providing the foundational 3D volumetric head representation and expression latent space that FaceTalk directly builds upon.
  • Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). It establishes generative diffusion modeling tailored for continuous 3D human motion, laying the conceptual basis for FaceTalk's motion diffusion architecture.
  • Paper: Executing your Commands via Motion Diffusion in Latent Space, Xin Chen et al. (2023). It introduces latent-space motion diffusion models, providing the core methodological blueprint for performing conditional diffusion synthesis within compressed parametric motion spaces.
  • Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). It defines the widely used FLAME statistical head model for disentangling identity, pose, and expression, contextualizing the parametric head modeling paradigms extended by NPHM and FaceTalk.
  • Paper: Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation, Lingting Zhu et al. (2023). It demonstrates audio-conditioned diffusion models for synthesizing coherent human communicative motion, providing foundational techniques for cross-modal speech-to-motion generation.
Cover for FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models

Abstract

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears, and finer-scale eye movements, we propose to couple speech signal with the latent space of neural parametric head models to create high-fidelity, temporally coherent motion sequences. We propose a new latent diffusion model for this task, operating in the expression space of neural parametric head models, to synthesize audio-driven realistic head sequences. In the absence of a dataset with corresponding NPHM expressions to audio, we optimize for these correspondences to produce a dataset of temporally-optimized NPHM expressions fit to audio-video recordings of people talking. To the best of our knowledge, this is the first work to propose a generative approach for realistic and high-quality motion synthesis of volumetric human heads, representing a significant advancement in the field of audio-driven 3D animation. Notably, our approach stands out in its ability to generate plausible motion sequences that can produce high-fidelity head animation coupled with the NPHM shape space. Our experimental results substantiate the effectiveness of FaceTalk, consistently achieving superior and visually natural motion, encompassing diverse facial expressions and styles, outperforming existing methods by 75% in perceptual user study evaluation.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Method
  • 5. Dataset Creation
  • 6. Results
  • 7. Conclusion
  • 8. Acknowledgments
  • References

Knowls

  1. Knowl 1 — FaceTalk latent diffusion for volumetric head animation

    model/method

    FaceTalk generates an entire sequence of high-fidelity 3D talking-head motions from a speech waveform. The model predicts a sequence of NPHM expression codes, {θ^expi}i=1N∈RN×200\{\hat{\boldsymbol{\theta}}_{\mathrm{exp}}^i\}_{i=1}^{N}\in\mathbb{R}^{N\times 200}, conditioned on aligned speech features, and combines these codes with a separately selected identity code. The predicted expressions are decoded by a frozen Neural Parametric Head Model (NPHM) into volumetric facial deformations and signed-distance fields, from which meshes are extracted with Marching Cubes.

    Unlike autoregressive speech-animation systems, FaceTalk denoises the complete expression sequence jointly. This allows it to model temporally coherent mouth motion together with stochastic upper-face behavior, including eye blinks, wrinkles, frowning, and other fine-scale facial expressions. Because identity and expression are represented separately, the same generated expression sequence can be applied to different NPHM identities.

  2. Knowl 2 — Neural parametric head representation used by FaceTalk

    definition

    NPHM represents a head with a volumetric identity code and a volumetric expression code. Its pretrained neural function is

    F(x,θid,θexp)→s,R3×R1344×R200→R,\mathcal{F}(\boldsymbol{x},\boldsymbol{\theta}_{\mathrm{id}},\boldsymbol{\theta}_{\mathrm{exp}})\rightarrow s, \qquad \mathbb{R}^{3}\times\mathbb{R}^{1344}\times\mathbb{R}^{200}\rightarrow\mathbb{R},

    where x∈R3\boldsymbol{x}\in\mathbb{R}^{3} is a query point in canonical 3D space, θid∈R1344\boldsymbol{\theta}_{\mathrm{id}}\in\mathbb{R}^{1344} is an identity latent code, θexp∈R200\boldsymbol{\theta}_{\mathrm{exp}}\in\mathbb{R}^{200} is an expression latent code, and s∈Rs\in\mathbb{R} is the predicted signed-distance-field value. A frozen identity network represents the subject’s overall geometry, including hair and ears, while a frozen expression network represents facial motion and produces expression deformations such as jaw motion, wrinkles, and eye blinks.

    The head surface is the zero-level isosurface of the predicted signed-distance field. FaceTalk extracts this surface with Marching Cubes, allowing the topology and fine geometric details of the generated head to come from the NPHM shape space rather than from a fixed-topology blendshape mesh.

  3. Knowl 3 — Audio-conditioned transformer expression decoder

    model/method

    FaceTalk encodes a raw speech waveform with a frozen Wave2Vec 2.0 encoder. Temporal convolution layers first produce audio features at the waveform sampling rate of 16 kHz16\,\mathrm{kHz}. Frequency interpolation resamples them to the expression-sequence rate of 24 Hz24\,\mathrm{Hz}, after which a stacked transformer encoder and a linear projection produce aligned conditioning features A1:N={ai}i=1N\boldsymbol{A}^{1:N}=\{\boldsymbol{a}_i\}_{i=1}^{N}.

    The expression decoder receives noisy NPHM expression codes, projects them into a transformer latent space, and processes them with stacked transformer decoder blocks. Each block contains multi-head self-attention, multi-head cross-attention to A1:N\boldsymbol{A}^{1:N}, and a feed-forward layer. A sinusoidal embedding of the diffusion timestamp is passed through a three-layer MLP, and FiLM layers inject this timestamp information between the attention and feed-forward components. The decoder then linearly projects its output back to the 200-dimensional NPHM expression space.

    The self-attention uses a binary look-ahead mask T∈{True,False}N×N\mathcal{T}\in\{\mathrm{True},\mathrm{False}\}^{N\times N} defined by

    Tij={True,i≤j,False,i>j,\mathcal{T}_{ij}=\begin{cases} \mathrm{True}, & i\leq j,\\ \mathrm{False}, & i>j, \end{cases}

    for sequence positions i,j∈{1,…,N}i,j\in\{1,\ldots,N\}. Audio-expression cross-attention uses an alignment mask M\mathcal{M} with Mij=True\mathcal{M}_{ij}=\mathrm{True} only when i=ji=j, so the audio feature at a timestep attends to the expression feature at the corresponding timestep. This explicit alignment is used to preserve phonetic synchronization and also permits inference on audio sequences longer than the training clips.

  4. Knowl 4 — Diffusion training, expression augmentation, and guidance

    algorithm

    FaceTalk trains a conditional diffusion model to denoise complete NPHM expression sequences. For a clean sequence x={θexpi}i=1N\boldsymbol{x}=\{\boldsymbol{\theta}_{\mathrm{exp}}^i\}_{i=1}^{N}, an integer diffusion timestamp tt is sampled uniformly from {0,…,T}\{0,\ldots,T\} and Gaussian noise is added according to

    q(xt∣x0)=N(αˉt x0,(1−αˉt)I),q(\boldsymbol{x}_t\mid\boldsymbol{x}_0) =\mathcal{N}\left(\sqrt{\bar{\alpha}_t}\,\boldsymbol{x}_0,(1-\bar{\alpha}_t)\boldsymbol{I}\right),

    where I\boldsymbol{I} is the identity covariance matrix and αˉt\bar{\alpha}_t follows the cosine noise schedule. The decoder Gϕ\mathcal{G}_{\boldsymbol{\phi}} is trained to predict the clean sequence from the noisy sequence, timestamp, and aligned audio conditioning A1:N\boldsymbol{A}^{1:N} by minimizing

    Lϕ=Ex,t[∥x−Gϕ(xt,t,A1:N)∥22].\mathcal{L}_{\boldsymbol{\phi}}= \mathbb{E}_{\boldsymbol{x},t}\left[ \left\|\boldsymbol{x}-\mathcal{G}_{\boldsymbol{\phi}}(\boldsymbol{x}_t,t,\boldsymbol{A}^{1:N})\right\|_2^2 \right].

    To prevent the speech signal from determining a single expression style, the training data are augmented by sampling r∼Uniform(a,b)r\sim\mathrm{Uniform}(a,b) and replacing each expression sequence with rxr\boldsymbol{x}. The paper specifies the modulation interval symbolically as [a,b][a,b] and uses uniform expression-code modulation in the implementation.

    Classifier-free guidance is trained by replacing the audio condition with a null condition with probability 25%25\%. During inference, the conditional and unconditional denoiser outputs are combined as

    xtguided=w xtcond+(1−w) xtuncond,\boldsymbol{x}_t^{\mathrm{guided}} =w\,\boldsymbol{x}_t^{\mathrm{cond}} +(1-w)\,\boldsymbol{x}_t^{\mathrm{uncond}},

    where ww is the guidance strength; choosing w>1w>1 amplifies the influence of the speech condition. Sampling starts from Gaussian expression noise at timestamp TT and iteratively denoises until timestamp 00.

  5. Knowl 5 — Spatial smoothing and NPHM mesh extraction

    model/method

    After diffusion sampling, FaceTalk decodes each predicted expression code with the frozen NPHM expression network to obtain spatial expression deformations. To suppress unwanted head and neck wobble, it attenuates deformations away from a control center C=[cx,cy,cz]\boldsymbol{C}=[c_x,c_y,c_z] located in the mouth region of canonical space.

    For a grid point pxyz=[px,py,pz]\boldsymbol{p}_{xyz}=[p_x,p_y,p_z] and Gaussian standard deviations Σ=[σx,σy,σz]\boldsymbol{\Sigma}=[\sigma_x,\sigma_y,\sigma_z], the normalized distance from the control center is

    dxyz=(px−cx)2σx2+(py−cy)2σy2+(pz−cz)2σz2.d_{xyz}=\sqrt{\frac{(p_x-c_x)^2}{\sigma_x^2}+\frac{(p_y-c_y)^2}{\sigma_y^2}+\frac{(p_z-c_z)^2}{\sigma_z^2}}.

    The corresponding unnormalized Gaussian weight is

    wxyz=12πσxσyσzexp⁡(−12dxyz2),w_{xyz}=\frac{1}{2\pi\sigma_x\sigma_y\sigma_z} \exp\left(-\frac{1}{2}d_{xyz}^{2}\right),

    and each weight is min-max normalized over all M=∣X∣∣Y∣∣Z∣M=|X||Y||Z| grid points:

    w~xyz=wxyz−min⁡(w1:M)max⁡(w1:M)−min⁡(w1:M).\tilde{w}_{xyz}=\frac{w_{xyz}-\min(w^{1:M})} {\max(w^{1:M})-\min(w^{1:M})}.

    The spatially varying weights multiply the decoded expression deformations, preserving detailed mouth motion while reducing deformation in distant regions. The smoothed deformations, the identity code, and the predicted expression codes are passed through the frozen NPHM identity network to obtain a signed-distance field for every frame; Marching Cubes then extracts the final mesh sequence.

  6. Knowl 6 — Temporally optimized audio-paired NPHM dataset

    model/method

    Because no dataset directly pairs audio with NPHM expression codes, FaceTalk creates supervision from multi-view NeRSemble recordings of talking subjects captured with 16 cameras and synchronized audio. The subjects already lie in the NPHM identity space, so only their framewise expression codes need to be estimated.

    For each sequence, COLMAP produces a point cloud Pi∈RK×3P_i\in\mathbb{R}^{K\times 3} for frame ii, where KK is the number of points. With the subject identity code θid\boldsymbol{\theta}_{\mathrm{id}} fixed, the method repeatedly samples 5000 points SiS_i from PiP_i, feeds SiS_i, θid\boldsymbol{\theta}_{\mathrm{id}}, and a learnable expression code θexpi\boldsymbol{\theta}_{\mathrm{exp}}^i to the frozen NPHM expression network, and adds the resulting deformation δi\boldsymbol{\delta}_i to obtain deformed points Di=Si+δiD_i=S_i+\boldsymbol{\delta}_i. The NPHM identity network evaluates the signed-distance field on these deformed points.

    The expression codes are optimized with

    Ltotal=λsdfLsdf+λtempLtemp+λregLreg,\mathcal{L}_{\mathrm{total}}= \lambda_{\mathrm{sdf}}\mathcal{L}_{\mathrm{sdf}}+ \lambda_{\mathrm{temp}}\mathcal{L}_{\mathrm{temp}}+ \lambda_{\mathrm{reg}}\mathcal{L}_{\mathrm{reg}},

    where Lsdf\mathcal{L}_{\mathrm{sdf}} is the mean absolute signed-distance value on the sampled points, Ltemp=∑i∥θexpi+1−θexpi∥Huber\mathcal{L}_{\mathrm{temp}}=\sum_i\|\boldsymbol{\theta}_{\mathrm{exp}}^{i+1}-\boldsymbol{\theta}_{\mathrm{exp}}^i\|_{\mathrm{Huber}} penalizes temporal changes, and Lreg=∑i∥θexpi∥2\mathcal{L}_{\mathrm{reg}}=\sum_i\|\boldsymbol{\theta}_{\mathrm{exp}}^i\|_2 keeps the codes near the learned NPHM expression distribution. Optimization is performed in sliding windows of n=10n=10 frames with a two-frame overlap, rather than independently per frame, to reduce flicker. The dataset contains 1000 optimized audio-expression sequences.

  7. Knowl 7 — Training and evaluation protocol

    experimental setup

    For diffusion training, sequences are randomly clipped to 2 seconds, corresponding to 48 frames at the 24-Hz expression rate. FaceTalk is optimized with Adam using learning rate 0.00010.0001, 1000 diffusion noising/denoising timestamps, and a cosine noise schedule. The NPHM expression-code augmentation uses uniform modulation, and the pretrained Wave2Vec 2.0 and NPHM components remain frozen.

    For dataset creation, expression codes are optimized for 500 iterations in each 10-frame window. The step size is 0.0010.001 through iteration 300 and 0.00010.0001 thereafter. The loss weights are λsdf=10\lambda_{\mathrm{sdf}}=10, λtemp=0.1\lambda_{\mathrm{temp}}=0.1, and λreg=0.0025\lambda_{\mathrm{reg}}=0.0025.

    Evaluation uses a 100-sequence hybrid test set containing 25 approximately 2–3-second clips from Vocaset, 25 approximately 2–3-second clips from the FaceTalk dataset, and 50 approximately 5–7-second clips from LJSpeech. Test identities are unseen during training. Lip synchronization is measured with LSE-D because NPHM meshes can have varying topology, making lip-vertex error unsuitable. Diversity is evaluated with FID, KID, and a diversity score; perceptual quality is evaluated with FIQA and VQA. Comparisons include VOCA, MeshTalk, FaceFormer, CodeTalker, Imitator, EmoTalk, and EMOTE.

  8. Knowl 8 — Quantitative comparison with speech-driven baselines

    data/table

    The table compares FaceTalk with template-based, personalized, and emotion-aware speech-driven facial-animation methods on the hybrid unseen-identity test set. FID and KID are computed on mouth-region crops, while LSE-D measures lip synchronization. Lower LSE-D, FID, and KID are better; higher FIQA and VQA are better. FaceTalk has the best value on every reported metric, especially reducing FID and KID relative to the baselines, which indicates improved mouth-motion realism and diversity while retaining audio synchronization.

    Could not parse LaTeX table
  9. Knowl 9 — Perceptual preference and generative behavior

    empirical result

    In a user study with 40 participants and 15 unseen audio clips, FaceTalk was preferred over Imitator and CodeTalker for overall animation quality, lip synchronization, and facial realism. The preference percentages were:

    Could not parse LaTeX table

    The generated expression codes transfer across identities with substantially different head geometry, hair, and facial structure. For a fixed identity and audio clip, stochastic generation produces different speaking styles, including different mouth-opening intensities and variations in eye blinks and frowning. Qualitative comparisons also show more detailed nasolabial creasing and more accurate phonetic lip articulation than the template-based baselines.

  10. Knowl 10 — Ablation of FaceTalk components

    data/table

    The ablation evaluates the contribution of expression augmentation, spatial facial smoothing, FiLM timestamp conditioning, expression-audio alignment, and diffusion training. LSE-D is lower when lip synchronization is better, while the diversity score is higher when multiple expression styles are produced. The full system gives the best LSE-D and maintains high diversity. Removing expression augmentation nearly eliminates diversity; removing facial smoothing preserves synchronization but reduces temporal visual stability; removing FiLM worsens synchronization; removing the alignment mask causes the model to ignore audio; and replacing diffusion with the ablated alternative produces zero measured diversity.

    Could not parse LaTeX table

    The visual ablations further show that omitting FiLM produces uncanny expressions and inaccurate lip synchronization, whereas omitting the alignment mask produces nearly constant expressions independent of the speech.

  11. Knowl 11 — Stated limitations of FaceTalk

    limitation

    FaceTalk requires multiple diffusion denoising steps at inference time, which limits its suitability for real-time animation. The paper identifies efficient diffusion sampling as a possible way to reduce this cost.

    The current system synthesizes only NPHM expression codes while requiring an identity code to be supplied separately. It therefore does not generate complete facial identity and expression jointly. The authors identify audio-conditioned generation of diverse identities, potentially including gender inferred from speech, as an unresolved direction for holistic 3D facial animation.

Coverage note — No substantial contributed material was deliberately omitted; background, related work, acknowledgements, and references were excluded.

References

  1. 1.Chaitanya Ahuja and Louis-Philippe Morency. Language2pose: Natural language grounded pose forecasting, 2019. 1
  2. 2.Simon Alexanderson, Rajmund Nagy, Jonas Beskow, and Gustav Eje Henter. Listen, denoise, action! audio-driven motion synthesis with diffusion models. ACM Trans. Graph., 42 (4):44:1–44:20, 2023. 1, 2
  3. 3.Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gul Varol. TEACH: Temporal Action Compositions for 3D Humans. In International Conference on 3D Vision (3DV), 2022. 1
  4. 4.Nikos Athanasiou, Mathis Petrovich, Michael J. Black, and Gul Varol. SINC: Spatial composition of 3D human motions for simultaneous action generation. In International Conference on Computer Vision (ICCV), 2023. 1
  5. 5.Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh, and Sonal Gupta. Make-an-animation: Large-scale text-conditional 3d human motion generation, 2023. 2
  6. 6.Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020. 2, 3, 4
  7. 7.Volker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 157–164. 2023. 1
  8. 8.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
  9. 9.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In arXiv, 2021. 2
  10. 10.Lele Chen, Zhiheng Li, Ross K. Maddox, Zhiyao Duan, and Chenliang Xu. Lip movements generation at a glance, 2018. 2
  11. 11.Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7832–7841, 2019.
  12. 12.Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu. Talking-head generation with rhythmic head motion. arXiv preprint arXiv:2007.08547, 2020. 2
  13. 13.Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. You said that?, 2017. 2
  14. 14.Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, and Michael Black. Capture, learning, and synthesis of 3D speaking styles. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pages 10101–10111, 2019. 2, 6, 7
  15. 15.Radek Daneˇcek, Kiran Chhatre, Shashank Tripathi, Yandong Wen, Michael Black, and Timo Bolkart. Emotional speech-driven animation with content-emotion disentanglement. In SIGGRAPH Asia 2023 Conference Papers, New York, NY, USA, 2023. Association for Computing Machinery. 2, 7
  16. 16.Dipanjan Das, Sandika Biswas, Sanjana Sinha, and Brojeshwar Bhowmick. Speech-driven facial animation using cascaded gans for learning of motion and texture. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXX, page 408–424, Berlin, Heidelberg, 2020. Springer-Verlag. 2
  17. 17.Prafulla Dhariwal and Alex Nichol. Diffusion models beat gans on image synthesis, 2021. 2
  18. 18.Rohan Dhesikan and Vignesh Rajmohan. Sketching the future (stf): Applying conditional control techniques to text-to-video models, 2023. 2
  19. 19.Michail Christos Doukas, Stefanos Zafeiriou, and Viktoriia Sharmanska. Headgan: One-shot neural head synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 14398–14407, 2021. 2
  20. 20.Fanda Fan, Chaoxu Guo, Litong Gong, Biao Wang, Tiezheng Ge, Yuning Jiang, Chunjie Luo, and Jianfeng Zhan. Hierarchical masked 3d diffusion model for video outpainting, 2023. 2
  21. 21.Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. Faceformer: Speech-driven 3d facial animation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 7
  22. 22.Guy Gafni, Justus Thies, Michael Zollhofer, and Matthias Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8649–8658, 2021. 2
  23. 23.Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Runz, Lourdes Agapito, and Matthias Nießner. Learning neural parametric head models. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 3, 4
  24. 24.Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Runz, Lourdes Agapito, and Matthias Nießner. MonoNPHM: Dynamic head reconstruction from monocular videos. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2024. 2, 3
  25. 25.Chuan Guo, Xinxin Zuo, Sen Wang, Shihao Zou, Qingyao Sun, Annan Deng, Minglun Gong, and Li Cheng. Action2motion: Conditioned generation of 3d human motions. In Proceedings of the 28th ACM International Conference on Multimedia, page 2021–2029, New York, NY, USA, 2020. Association for Computing Machinery. 1
  26. 26.Yudong Guo, Keyu Chen, Sen Liang, Yongjin Liu, Hujun Bao, and Juyong Zhang. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 2
  27. 27.Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. Space: Speech-driven portrait animation with controllable expression, 2022. 2
  28. 28.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 5
  29. 29.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. arXiv preprint arxiv:2006.11239, 2020. 2
  30. 30.Rongjie Huang, Max W. Y. Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao. Fastdiff: A fast conditional diffusion model for high-quality speech synthesis, 2022. 2
  31. 31.Keith Ito and Linda Johnson. The lj speech dataset. https://keithito.com/LJ-Speech-Dataset/, 2017. 7
  32. 32.Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. Audio-driven emotional video portraits. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
  33. 33.Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao. Eamm: One-shot emotional talking face via audio-based emotion-aware motion model. In ACM SIGGRAPH 2022 Conference Proceedings, 2022. 2
  34. 34.Prajwal K R, Rudrabha Mukhopadhyay, Jerin Philip, Abhishek Jha, Vinay Namboodiri, and C V Jawahar. Towards automatic face-to-face translation. In Proceedings of the 27th ACM International Conference on Multimedia, page 1428–1436, New York, NY, USA, 2019. Association for Computing Machinery. 2
  35. 35.Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Trans. Graph., 36 (4), 2017. 2
  36. 36.Jihoon Kim, Jiseob Kim, and Sungjoon Choi. FLAME: Free-form language-based motion synthesis & editing. arXiv preprint arXiv:2209.00349, 2022. 1, 2
  37. 37.Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, and Matthias Nießner. Nersemble: Multi-view radiance field reconstruction of human heads. ACM Trans. Graph., 42(4), 2023. 2, 5, 6
  38. 38.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. Diffwave: A versatile diffusion model for audio synthesis, 2021. 2
  39. 39.Max WY Lam, Jun Wang, Dan Su, and Dong Yu. Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis. In International Conference on Learning Representations, 2022. 2
  40. 40.Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6):194:1–194:17, 2017. 1
  41. 41.Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou. Semantic-aware implicit neural audio-driven video portrait generation. arXiv preprint arXiv:2201.07786, 2022. 2
  42. 42.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3d surface construction algorithm. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques, page 163–169, New York, NY, USA, 1987. Association for Computing Machinery. 3, 5, 6
  43. 43.Fu-Zhao Ou, Xingyu Chen, Ruixin Zhang, Yuge Huang, Shaoxin Li, Jilin Li, Yong Li, Liujuan Cao, and Yuan-Gen Wang. SDD-FIQA: Unsupervised face image quality assessment with similarity distribution distance. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 6
  44. 44.Ziqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu, Xiangyu Zhu, Jun He, Hongyan Liu, and Zhaoxin Fan. Emotalk: Speech-driven emotional disentanglement for 3d face animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 20687–20697, 2023. 2, 7
  45. 45.Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer, 2017. 3, 4
  46. 46.Mathis Petrovich, Michael J. Black, and Gul Varol. Action-conditioned 3D human motion synthesis with transformer VAE. In International Conference on Computer Vision (ICCV), 2021. 1
  47. 47.Mathis Petrovich, Michael J. Black, and Gul Varol. TEMOS: Generating diverse human motions from textual descriptions. In European Conference on Computer Vision (ECCV), 2022. 1
  48. 48.K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C.V. Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM International Conference on Multimedia, page 484–492, New York, NY, USA, 2020. Association for Computing Machinery. 2, 6
  49. 49.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022. 2
  50. 50.Zhiyuan Ren, Zhihong Pan, Xin Zhou, and Le Kang. Diffusion motion: Generate text-guided 3d human motion by diffusion model, 2023. 6
  51. 51.Alexander Richard, Michael Zollhofer, Yandong Wen, Fernando de la Torre, and Yaser Sheikh. Meshtalk: 3d face animation from speech using cross-modality disentanglement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1173–1182, 2021. 2, 7
  52. 52.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  53. 53.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding, 2022. 2
  54. 54.Johannes Lutz Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 5
  55. 55.Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu. Learning dynamic facial radiance fields for few-shot talking head synthesis. In ECCV, 2022. 2
  56. 56.Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li, Zheng Zhu, Jie Zhou, and Jiwen Lu. Difftalk: Crafting diffusion models for generalized audio-driven portraits animation. In CVPR, 2023. 2
  57. 57.Linsen Song, Wayne Wu, Chen Qian, Ran He, and Chen Change Loy. Everybody’s talkin’: Let me talk as you want, 2020. 2
  58. 58.Stefan Stan, Kazi Injamamul Haque, and Zerrin Yumak. Facediffuser: Speech-driven 3d facial animation synthesis using diffusion. In Proceedings of the 16th ACM SIGGRAPH Conference on Motion, Interaction and Games, New York, NY, USA, 2023. Association for Computing Machinery. 2
  59. 59.Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zikeba, Stavros Petridis, and Maja Pantic. Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation. In https://arxiv.org/abs/2301.03396, 2023. 2
  60. 60.Zhiyao Sun, Tian Lv, Sheng Ye, Matthieu Gaetan Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, and Yong jin Liu. Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models, 2023. 2
  61. 61.Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing obama: Learning lip sync from audio. ACM Trans. Graph., 2017. 2
  62. 62.Anni Tang, Tianyu He, Xu Tan, Jun Ling, Runnan Li, Sheng Zhao, Li Song, and Jiang Bian. Memories are one-to-many mapping alleviators in talking face generation. arXiv preprint arXiv:2212.05005, 2022. 2
  63. 63.Jiapeng Tang, Angela Dai, Yinyu Nie, Lev Markhasin, Justus Thies, and Matthias Niessner. Dphms: Diffusion parametric head models for depth-based tracking. 2024. 2
  64. 64.Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffusion model. In The Eleventh International Conference on Learning Representations, 2023. 1, 2
  65. 65.Balamurugan Thambiraja, Ikhsanul Habibie, Sadegh Aliakbarian, Darren Cosker, Christian Theobalt, and Justus Thies. Imitator: Personalized speech-driven 3d facial animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 20621–20631, 2023. 2, 7
  66. 66.Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner. Neural voice puppetry: Audio-driven facial reenactment. ECCV 2020, 2020. 2
  67. 67.Jonathan Tseng, Rodrigo Castellon, and C Karen Liu. Edge: Editable dance generation from music. arXiv preprint arXiv:2211.10658, 2022. 1, 2
  68. 68.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023. 3, 4
  69. 69.Vikram Voleti, Alexia Jolicoeur-Martineau, and Christopher Pal. Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation. In (NeurIPS) Advances in Neural Information Processing Systems, 2022. 2
  70. 70.Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. End-to-end speech-driven facial animation with temporal gans, 2018. 2
  71. 71.Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. Realistic speech-driven facial animation with gans, 2019. 2
  72. 72.Daniel Watson, Jonathan Ho, Mohammad Norouzi, and William Chan. Learning to efficiently sample from diffusion probabilistic models, 2021. 8
  73. 73.O. Wiles, A.S. Koepke, and A. Zisserman. X2face: A network for controlling face generation by using images, audio, and pose codes. In European Conference on Computer Vision, 2018. 2
  74. 74.Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin. Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In International Conference on Computer Vision (ICCV), 2023. 6
  75. 75.Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang, and Tien-Tsin Wong. Codetalker: Speech-driven 3d facial animation with discrete motion prior. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12780–12790, 2023. 2, 7
  76. 76.Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai, and Xiaokang Yang. Dfa-nerf: Personalized talking head generation via disentangled face attributes neural rendering. arXiv preprint arXiv:2201.00791, 2022. 2
  77. 77.Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Shaoli Huang, Yong Zhang, Hongwei Zhao, Hongtao Lu, and Xi Shen. T2m-gpt: Generating human motion from textual descriptions with discrete representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 1
  78. 78.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023. 2
  79. 79.Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. Talking face generation by adversarially disentangled audio-visual representation. In AAAI Conference on Artificial Intelligence (AAAI), 2019. 2
  80. 80.Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. Makelttalk: Speaker-aware talking-head animation. ACM Trans. Graph., 39(6), 2020. 2
  81. 81.Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian, Ziwei Liu, and Lequan Yu. Taming diffusion models for audio-driven co-speech gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10544–10553, 2023. 2

Citation

MLA
Aneja, S., et al. “FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models”. CVPR 2024, 2023, http://arxiv.org/abs/2312.08459v2.
APA
Aneja, S., Thies, J., Dai, A., & Nießner, M. (2023). FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models. CVPR 2024. http://arxiv.org/abs/2312.08459v2
Chicago
Aneja, S., J. Thies, A. Dai, and M. Nießner. 2023. “FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models”. CVPR 2024. http://arxiv.org/abs/2312.08459v2.
Harvard
Aneja, S. et al. (2023) “FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models”, CVPR 2024 [Preprint]. Available at: http://arxiv.org/abs/2312.08459v2.
Vancouver
1. Aneja S, Thies J, Dai A, Nießner M (2023) FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models. CVPR 2024

BibTeX

@article{aneja2023facetalk,
  title = {FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models},
  author = {Aneja, Shivangi and Thies, Justus and Dai, Angela and Nießner, Matthias},
  year = {2023},
  journal = {CVPR 2024},
  url = {http://arxiv.org/abs/2312.08459v2},
  eprint = {2312.08459}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE