Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Zhenhui YeTianyun ZhongYi RenJiaqi YangWeichuang LiJiawei HuangZiyue JiangJinzheng HeRongjie HuangJinglin Liu

article2024ICLR124 citations

Presents a one-shot 3D talking portrait framework that distills generative 3D priors and models full head-torso dynamics to synthesize realistic, audio- and video-driven avatar videos from a single unseen image.

Listen

Generating realistic talking portrait videos from a single still image is an important capability for interactive digital media, video conferencing, customer service, and virtual avatars. However, existing methods face major challenges: two-dimensional approaches suffer from severe visual distortion during large head movements, while three-dimensional methods either require hours of individual training per identity or struggle with facial animation stability and identity preservation. Furthermore, previous systems typically focus solely on head generation, neglecting natural torso movements and realistic background integration.

The article introduces and evaluates Real3D-Portrait, a one-shot framework designed to generate realistic, three-dimensional talking portrait videos from a single unseen reference image, driven either by an input video or an audio track.

The approach uses a modular, multi-stage architecture trained across large-scale video and synthetic datasets. First, an image-to-plane model pre-trained on multi-view synthetic data reconstructs an accurate canonical three-dimensional representation from a single photo. A lightweight motion adapter then animates facial expressions using standardized coordinate codes without distorting underlying geometry. To ensure natural full-frame composition, a dedicated super-resolution model independently handles the head, warps the torso based on keypoint movement, inpaints the background, and blends them using an occlusion-aware mask. Finally, a generalized audio-to-motion model translates speech audio into precise facial motion parameters with controllable eye blinking and mouth amplitude.

Evaluation shows that the system achieves state-of-the-art results across both video- and audio-driven benchmarks. In video-driven reenactment, Real3D-Portrait attained the highest image fidelity (FID score of 37.50 for same-identity and 42.37 for cross-identity) and superior identity similarity compared to previous one-shot methods. In audio-driven generation, the framework demonstrated lip-synchronization scores (Sync score of 6.565) and visual quality comparable to person-specific models that require lengthy per-user training. Human perceptual studies confirmed that viewers rated Real3D-Portrait higher in identity preservation, visual smoothness, and lip synchronization than competing one-shot tools.

These findings indicate that high-quality, three-dimensional talking portraits no longer require costly per-individual model fine-tuning or compromise full-body visual coherence. This significantly lowers computational costs and production turnaround times for deploying personalized digital avatars at scale. Because the system cleanly separates head, torso, and background elements, it also offers practical flexibility, such as customizable backgrounds.

Organizations looking to implement digital avatar systems should adopt modular, one-shot three-dimensional pipelines over rigid, identity-specific approaches. To manage ethical and security risks associated with deepfake technologies, deployers should incorporate watermarking and tracking safeguards. Future development should focus on integrating advanced neural inpainting for background synthesis and gathering broader training data for extreme side-view head poses. While confidence in the evaluated head poses and standard portrait angles is high, caution is advised when deploying the model in scenarios with extreme rotational head movement.

Cover for Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Abstract

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talking face animation. Besides, while the existing works mainly focus on synthesizing the head part, it is also vital to generate natural torso and background segments to obtain a realistic talking portrait video. To address these limitations, we present Real3D-Potrait, a framework that (1) improves the one-shot 3D reconstruction power with a large image-to-plane model that distills 3D prior knowledge from a 3D face generative model; (2) facilitates accurate motion-conditioned animation with an efficient motion adapter; (3) synthesizes realistic video with natural torso movement and switchable background using a head-torso-background super-resolution model; and (4) supports one-shot audio-driven talking face generation with a generalizable audio-to-motion model. Extensive experiments show that Real3D-Portrait generalizes well to unseen identities and generates more realistic talking portrait videos compared to previous methods. Video samples and source code are available at this https URL .

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Real3D-Portrait
  • 3.1 Image-to-Plane model for 3D Face Reconstruction
  • 3.2 Motion Adapter for 3D Face Animation
  • 3.3 Head-Torso-Background Super-Resolution Model
  • 3.4 Generic Audio-to-Motion Model
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Quantitative Evaluation
  • 4.3 Qualitative Evaluation
  • 4.4 Ablation Studies
  • 5 Conclusion
  • 6 Acknowledgments
  • 7 Ethics Impacts
  • References
  • A Comparsion between Different Methods
  • B Additional Network and Training Details
  • B.1 Pretraining Image-to-Plane Model
  • B.2 Obtaining PNCC from 3DMM Coefficients
  • B.3 Warping-based Torso Branch
  • B.4 Background Branch with KNN-based Inpainting
  • B.5 Obtaining the Head and Torso Occlusion Mask
  • B.6 Detailed Structure of Audio-to-Motion Model
  • C Detailed Model Configuration
  • C.1 Model Configuration
  • C.2 Training Details.
  • D Additional Experiments
  • D.1 Evaluation Details
  • D.2 User Study Setting
  • D.3 Additional Qualitative Results
  • D.4 Additional Ablation Studies
  • D.5 Additional quantitative comparison with recent 2D baselines
  • E Limitations and Future Work

Knowls

  1. Knowl 1 — Real3D-Portrait Architecture and Pipeline for One-Shot 3D Talking Portrait Synthesis

    model/method

    Real3D-Portrait is a one-shot 3D talking portrait generation framework that reconstructs a 3D avatar from an unseen single reference image IsrcI_{\text{src}} and animates it using either driving audio or a driving motion sequence to render a 512×512512 \times 512 talking portrait video with natural torso motion and switchable background.

    The framework consists of four sequential core components:

    1. Image-to-Plane (I2P) Model: A feed-forward network combining a Vision Transformer (ViT) branch and a VGG branch that maps the source image IsrcI_{\text{src}} directly into a canonical 3D tri-plane representation Pcano∈R256×256×96P_{\text{cano}} \in \mathbb{R}^{256 \times 256 \times 96}, pre-trained on multi-view face data distilled from a 3D GAN.
    2. Facial Motion Adapter (MA): A lightweight SegFormer-based network that animates the head by predicting a residual motion diff-plane Pdiff∈R256×256×96P_{\text{diff}} \in \mathbb{R}^{256 \times 256 \times 96} from driving and source Projected Normalized Coordinate Codes (PNCC), editing the canonical tri-plane to Pcano+PdiffP_{\text{cano}} + P_{\text{diff}}.
    3. Head-Torso-Background Super-Resolution (HTB-SR) Model: A multi-branch rendering network that separately processes volume-rendered head features, 2D warping-based torso features driven by 3DMM keypoints, and in उन्हें (KNN-inpainted) background features, fusing them with an occlusion-aware alpha-blending mechanism.
    4. Generic Audio-to-Motion (A2M) Model: A flow-enhanced Variational Auto-Encoder (VAE) that transforms HuBERT speech features, pitch contour, eye blink, and mouth amplitude into 3DMM expression parameter sequences for audio-driven animation.
  2. Knowl 2 — Hybrid Image-to-Plane Model Architecture and 3D Prior Pre-training

    model/method

    The Image-to-Plane (I2P) model reconstructs a canonical tri-plane representation PcanoP_{\text{cano}} from a single 512×512×3512 \times 512 \times 3 reference face image IsrcI_{\text{src}} through a single feed-forward pass using a hybrid dual-branch architecture:

    • ViT Branch: A sequence of SegFormer blocks utilizing self-attention across image patches to execute the spatial coordinate transformation from 2D pixel space to the 3D canonical coordinate space.
    • VGG Branch: A 10-layer convolutional network operating without normalization layers to preserve high-frequency texture details and retain identity-specific appearance bias across all layers.

    The outputs of both branches are concatenated and passed through 4 convolutional layers to produce the final canonical tri-plane of dimensions 256×256×32×3256 \times 256 \times 32 \times 3.

    To overcome the lack of multi-view coverage in talking-head video datasets, the I2P model is pre-trained on synthetic multi-view image pairs generated on-the-fly by an EG3D generator with relaxed camera radius constraints (sampled between distances 2.42.4 and 5.05.0). Given reference camera crefc_{\text{ref}} and target camera cmvc_{\text{mv}}, the reconstructed tri-plane is rendered via volume rendering to yield multi-view prediction I^mv\hat{I}_{\text{mv}}, supervised by the loss: Lpretrain=MSE(Imv,I^mv)+LVGG19(Imv,I^mv)+LVGGFace(Imv,I^mv)+LDualAdv(I^mv_raw,I^mv),\mathcal{L}_{\text{pretrain}} = \text{MSE}(I_{\text{mv}}, \hat{I}_{\text{mv}}) + \mathcal{L}_{\text{VGG19}}(I_{\text{mv}}, \hat{I}_{\text{mv}}) + \mathcal{L}_{\text{VGGFace}}(I_{\text{mv}}, \hat{I}_{\text{mv}}) + \mathcal{L}_{\text{DualAdv}}(\hat{I}_{\text{mv\_raw}}, \hat{I}_{\text{mv}}), where MSE\text{MSE} is the mean squared error, LVGG19\mathcal{L}_{\text{VGG19}} and LVGGFace\mathcal{L}_{\text{VGGFace}} denote perceptual losses on VGG19 and VGGFace feature spaces, and LDualAdv\mathcal{L}_{\text{DualAdv}} is the dual discriminator adversarial loss operating jointly on the raw volume-rendered image I^mv_raw\hat{I}_{\text{mv\_raw}} (128×128128 \times 128) and super-resolved image I^mv\hat{I}_{\text{mv}} (512×512512 \times 512).

  3. Knowl 3 — PNCC-Conditioned Facial Motion Adapter with Laplacian Regularization

    model/method

    The Motion Adapter (MA) animates the reconstructed 3D head representation by predicting a residual motion diff-plane PdiffP_{\text{diff}} that edits the canonical tri-plane PcanoP_{\text{cano}} via element-wise addition: Idrv=SR(VR(Pcano+Pdiff,cam)),where Pcano=I2P(Isrc),  Pdiff=MA(PNCCdrv,PNCCsrc),I_{\text{drv}} = \text{SR}(\text{VR}(P_{\text{cano}} + P_{\text{diff}}, \text{cam})), \quad \text{where } P_{\text{cano}} = \text{I2P}(I_{\text{src}}), \; P_{\text{diff}} = \text{MA}(\text{PNCC}_{\text{drv}}, \text{PNCC}_{\text{src}}), where VR\text{VR} is the NeRF volume renderer, SR\text{SR} is the super-resolution module, cam\text{cam} is the target camera pose, and PNCC\text{PNCC} is the Projected Normalized Coordinate Code.

    The PNCC is an appearance- and pose-agnostic 2D feature map encoding geometry and facial expression. Given 3DMM identity parameters i∈R80i \in \mathbb{R}^{80} and expression parameters e∈R64e \in \mathbb{R}^{64}: PNCC=Z-Buffer(Vertex3D(i,e),NCC),s.t. Vertex3D(i,e)=Vertex‾3D+Bidi+Bexpe,\text{PNCC} = \text{Z-Buffer}(\text{Vertex}_{\text{3D}}(i, e), \text{NCC}), \quad \text{s.t. } \text{Vertex}_{\text{3D}}(i, e) = \overline{\text{Vertex}}_{\text{3D}} + B_{\text{id}} i + B_{\text{exp}} e, where Vertex‾3D\overline{\text{Vertex}}_{\text{3D}}, BidB_{\text{id}}, and BexpB_{\text{exp}} are the mean shape, identity basis, and expression basis of the BFM 2009 morphable model, and NCC\text{NCC} is the normalized coordinate colormap rasterized at canonical pose.

    The Motion Adapter is implemented as a 4-block shallow SegFormer taking the channel concatenation of PNCCsrc\text{PNCC}_{\text{src}} and PNCCdrv\text{PNCC}_{\text{drv}} (512×512×6512 \times 512 \times 6). During training on video frame pairs (Isrc,Itgt)(I_{\text{src}}, I_{\text{tgt}}), the total objective is: L=∥Itgt−Itgt′∥1+LVGGs(Itgt,Itgt′)+LDualAdv(Itgt′)+LLap,\mathcal{L} = \|I_{\text{tgt}} - I'_{\text{tgt}}\|_1 + \mathcal{L}_{\text{VGGs}}(I_{\text{tgt}}, I'_{\text{tgt}}) + \mathcal{L}_{\text{DualAdv}}(I'_{\text{tgt}}) + \mathcal{L}_{\text{Lap}}, where LLap\mathcal{L}_{\text{Lap}} is a temporal Laplacian loss penalizing high-frequency jitter across consecutive frames: LLap=∥MA(PNCCt)−0.5×(MA(PNCCt−1)+MA(PNCCt+1))∥22.\mathcal{L}_{\text{Lap}} = \|\text{MA}(\text{PNCC}_t) - 0.5 \times (\text{MA}(\text{PNCC}_{t-1}) + \text{MA}(\text{PNCC}_{t+1}))\|_2^2.

  4. Knowl 4 — Head-Torso-Background Super-Resolution Model and Occlusion-Aware Alpha Blending

    model/method

    The Head-Torso-Background Super-Resolution (HTB-SR) model individually represents head, torso, and background segments and integrates them to produce full 512×512512 \times 512 portraits.

    1. Head SR Branch: Processes low-resolution (128×128128 \times 128) volume-rendered head features.
    2. Torso Branch: Employs a 2D warping renderer driven by 68 3DMM keypoints extracted from the 3D mesh vertices: KP=IdxKP(R⋅Vertex3D+t),\text{KP} = \text{IdxKP}(R \cdot \text{Vertex}_{\text{3D}} + t), Ftorso=DBD(TAE(Itorso),DME(KPsrc,KPtgt)),F_{\text{torso}} = \text{DBD}(\text{TAE}(I_{\text{torso}}), \text{DME}(\text{KP}_{\text{src}}, \text{KP}_{\text{tgt}})), where R,tR, t are the camera rotation and translation, TAE\text{TAE} is the torso appearance encoder, DME\text{DME} is the dense motion estimator, and DBD\text{DBD} is the deformation-based decoder.
    3. Background Branch: Inpaints foreground-occluded regions of the source background using K-nearest-neighbor (KNN) pixel filling, followed by a 3-layer convolutional feature extractor to produce FbgF_{\text{bg}}.
    4. Occlusion-Aware Alpha-Blending Fusion: Combines feature maps according to structural depth order: F=(Fhead⋅Mhead+Ftorso⋅(1−Mhead))⋅Mperson+Fbg⋅(1−Mperson),F = (F_{\text{head}} \cdot M_{\text{head}} + F_{\text{torso}} \cdot (1 - M_{\text{head}})) \cdot M_{\text{person}} + F_{\text{bg}} \cdot (1 - M_{\text{person}}), where Mperson=Mhead∨MtorsoM_{\text{person}} = M_{\text{head}} \lor M_{\text{torso}}, with MtorsoM_{\text{torso}} predicted by DME\text{DME}, and MheadM_{\text{head}} computed via NeRF ray volume rendering density integration: Mhead=∫tntfσ(r(t))exp⁡(−∫tntσ(r(s)) ds)dt.M_{\text{head}} = \int_{t_n}^{t_f} \sigma(r(t)) \exp\left(-\int_{t_n}^t \sigma(r(s)) \, ds\right) dt.
  5. Knowl 5 — Flow-Enhanced Variational Audio-to-Motion Model

    model/method

    The Audio-to-Motion (A2M) model converts speech audio into 3DMM expression parameter trajectories edrv∈RT×64e_{\text{drv}} \in \mathbb{R}^{T \times 64} for identity-agnostic, controllable talking face synthesis.

    • Architecture: A Variational Auto-Encoder (VAE) with an 8-layer WavNet encoder, a 4-layer WavNet decoder, and a 4-layer Flow-based prior module.
    • Input Signals: HuBERT speech representation features, pitch contour, eye blink control signal, and mouth amplitude control scalar.
    • Training Objective: LA2M=LKL+LExpRecon+LLdmRecon+LExpLap,\mathcal{L}_{\text{A2M}} = \mathcal{L}_{\text{KL}} + \mathcal{L}_{\text{ExpRecon}} + \mathcal{L}_{\text{LdmRecon}} + \mathcal{L}_{\text{ExpLap}}, where LKL\mathcal{L}_{\text{KL}} is the KL-divergence of the VAE latent space, LExpRecon\mathcal{L}_{\text{ExpRecon}} is the L2L_2 reconstruction error on the 64-dimensional expression code ee, LLdmRecon\mathcal{L}_{\text{LdmRecon}} is an auxiliary L2L_2 error on 468 3D facial landmarks derived from the 3DMM face mesh, and LExpLap\mathcal{L}_{\text{ExpLap}} is a Laplacian loss on the predicted temporal sequence e1:Te_{1:T} to ensure smooth transitions without expression jitter.
  6. Knowl 6 — Quantitative Evaluation on Video-Driven Talking Portrait Reenactment

    data/table

    Video-driven reenactment performance was evaluated on the CelebV-HQ dataset under both same-identity (100 test video pairs) and cross-identity (100 source identities animated by driving videos) conditions. Evaluation metrics comprise L1L_1 distance, PSNR, SSIM, LPIPS, Fréchet Inception Distance (FID), Cosine Similarity of ArcFace identity embeddings (CSIM), Average Expression Distance (AED), Average Pose Distance (APD), and Average Keypoint Distance (AKD).

    Same-Identity Reenactment
    Methods L1 ↓\downarrow PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow FID ↓\downarrow CSIM ↑\uparrow AED ↓\downarrow APD ↓\downarrow AKD ↓\downarrow
    Face-vid2vid 0.078 16.47 0.779 0.184 42.96 0.808 0.116 0.023 3.176
    OTAvatar 0.106 13.17 0.672 0.232 65.37 0.568 0.170 0.040 5.891
    HiDe-NeRF 0.084 15.92 0.752 0.189 50.04 0.753 0.129 0.021 3.531
    Real3D-Portrait (Ours) 0.067 18.95 0.801 0.171 37.50 0.821 0.111 0.018 2.829
    Cross-Identity Reenactment
    Methods – CSIM ↑\uparrow FID ↓\downarrow AED ↓\downarrow APD ↓\downarrow
    Face-vid2vid – 0.726 45.18 0.144 0.029
    OTAvatar – 0.544 64.28 0.195 0.046
    HiDe-NeRF – 0.699 53.28 0.161 0.025
    Real3D-Portrait (Ours) – 0.758 42.37 0.138 0.022

    Real3D-Portrait outperforms 2D warping (Face-vid2vid) and one-shot 3D baselines (OTAvatar, HiDe-NeRF) across visual reconstruction quality (PSNR 18.95, FID 37.50/42.37), source identity preservation (CSIM 0.821/0.758), and motion transfer precision (AED 0.111/0.138, APD 0.018/0.022).

  7. Knowl 7 — Quantitative Evaluation on Audio-Driven Talking Face Generation

    data/table

    Audio-driven talking face generation was evaluated on 100 test samples generated from 10 unseen identities paired with 10 multilingual audio clips. Metrics include identity similarity (CSIM), image generation quality (FID), Average Expression Distance (AED), and SyncNet audio-lip synchronization confidence score (Sync).

    Methods CSIM ↑\uparrow FID ↓\downarrow AED ↓\downarrow Sync ↑\uparrow
    MakeItTalk 0.715 52.65 0.213 3.286
    PC-AVS 0.327 82.02 0.162 6.483
    RAD-NeRF 0.784 39.45 0.197 3.779
    Real3D-Portrait (Ours AD) 0.763 43.02 0.146 6.565

    Real3D-Portrait substantially outperforms one-shot 2D audio-driven systems (MakeItTalk and PC-AVS) across all dimensions. When compared to RAD-NeRF—a person-specific 3D NeRF over-fitted to a 3-minute video of the target identity—Real3D-Portrait achieves higher audio-lip synchronization (Sync score 6.565 vs. 3.779, AED 0.146 vs. 0.197) and comparable identity preservation and visual fidelity in a zero-shot setting.

  8. Knowl 8 — Subjective Mean Opinion Score User Study

    data/table

    A user study was performed with 20 participants evaluating 100 synthesized video clips (10 identities ×\times 10 audio/video inputs) across all baseline methods. Ratings were collected on a 1-to-5 Mean Opinion Score (MOS) scale (with 95% confidence intervals) across three criteria: Identity Preservation, Visual Quality (fidelity and temporal smoothness), and Audio-Lip Synchronization.

    Methods Modality ID Preserving ↑\uparrow Visual Quality ↑\uparrow Lip Sync ↑\uparrow
    Face-vid2vid Video 3.79 ±\pm 0.34 3.73 ±\pm 0.25 3.97 ±\pm 0.20
    OTAvatar Video 3.29 ±\pm 0.24 3.28 ±\pm 0.29 3.80 ±\pm 0.28
    HiDe-NeRF Video 3.74 ±\pm 0.31 3.45 ±\pm 0.29 3.54 ±\pm 0.32
    Real3D-Portrait (Ours VD) Video 4.08 ±\pm 0.31 4.16 ±\pm 0.23 4.13 ±\pm 0.29
    MakeItTalk Audio 3.38 ±\pm 0.44 3.34 ±\pm 0.34 2.96 ±\pm 0.37
    PC-AVS Audio 3.21 ±\pm 0.36 3.40 ±\pm 0.37 4.04 ±\pm 0.31
    RAD-NeRF Audio 4.12 ±\pm 0.26 4.25 ±\pm 0.24 3.18 ±\pm 0.52
    Real3D-Portrait (Ours AD) Audio 4.05 ±\pm 0.27 4.14 ±\pm 0.29 4.08 ±\pm 0.25

    Real3D-Portrait achieved higher perceived visual quality and identity preservation than all one-shot video- and audio-driven baselines, while achieving lip-sync ratings superior to the person-specific RAD-NeRF.

  9. Knowl 9 — Ablation Analysis of Real3D-Portrait Architecture Components

    data/table

    Ablation experiments evaluated on cross-identity reenactment on CelebV-HQ illustrate the impact of each design decision across identity preservation (CSIM), visual quality (FID), expression distance (AED), and pose distance (APD).

    Setting CSIM ↑\uparrow FID ↓\downarrow AED ↓\downarrow APD ↓\downarrow
    w/o pre-train 0.487 65.32 0.181 0.031
    w/o finetune 0.683 49.21 0.233 0.027
    w/ 40M params 0.725 45.48 0.140 0.026
    w/ 200M params 0.754 43.15 0.143 0.023
    w/o Lap loss 0.748 42.66 0.158 0.024
    w/ unsup. KP. 0.746 44.86 0.138 0.023
    w/ concat 0.737 46.38 0.144 0.025
    w/o inpaint 0.744 43.95 0.140 0.022
    Full Real3D-Portrait (VD) 0.758 42.37 0.138 0.022

    Key takeaways:

    • Removing EG3D multi-view pre-training (w/o pre-train) causes severe drops across CSIM (0.487) and FID (65.32), confirming its necessity for 3D prior acquisition.
    • Omitting video fine-tuning (w/o finetune) degrades expression accuracy (AED 0.233 vs. 0.138).
    • 87M parameters strikes the best trade-off compared to smaller (40M) and larger (200M) backbones.
    • Replacing alpha-blending fusion with naive concatenation (w/ concat) leads to boundary artifacts and higher FID (46.38).
    • Unsupervised keypoints (w/ unsup. KP.) and removing background inpainting (w/o inpaint) degrade stability and visual quality.
  10. Knowl 10 — Limitations in Extreme Head Poses and Torso Disocclusion

    limitation

    Real3D-Portrait exhibits two primary limitations:

    1. Extreme Head Pose Degradation: Due to the scarcity of extreme pose angles (such as full profile/side views) in the video training distribution, the model suffers visual degradation and quality loss when driving portraits to large side poses.
    2. Torso Disocclusion Artifacts: Because the background branch utilizes a simple K-nearest-neighbor (KNN) inpainting heuristic rather than deep inpainting networks, disoccluded background regions can display blurring or unnatural textures under large torso movements.

Coverage note — None was omitted; all key architectural components, training losses, quantitative experiments, subjective user studies, ablation settings, and limitations are fully covered.

References

  1. 1.V Blanz and T Vetter. A morphable model for the synthesis of 3d faces. In 26th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH 1999), pp. 187–194. ACM Press, 1999.
  2. 2.Adrian Bulat and Georgios Tzimiropoulos. How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks). In International Conference on Computer Vision, 2017.
  3. 3.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J. Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3d generative adversarial networks. In CVPR, pp. 16123–16133, June 2022.
  4. 4.Lele Chen, Guofeng Cui, Ziyi Kou, Haitian Zheng, and Chenliang Xu. What comprises a good talking-head video generation?: A survey and benchmark. arXiv preprint arXiv:2005.03201, 2020.
  5. 5.Joon Son Chung, Amir Jamaludin, and Andrew Zisserman. You said that? arXiv preprint arXiv:1705.02966, 2017.
  6. 6.Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. Voxceleb2: Deep speaker recognition. arXiv preprint arXiv:1806.05622, 2018.
  7. 7.Radek Daněček, Michael J Black, and Timo Bolkart. Emoca: Emotion driven monocular face capture and animation. In CVPR, pp. 20311–20322, 2022.
  8. 8.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4690–4699, 2019a.
  9. 9.Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. In CVPRW, 2019b.
  10. 10.Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pp. 0–0, 2019c.
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  12. 12.Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In ICCV, pp. 5784–5794, 2021.
  13. 13.Fa-Ting Hong and Dan Xu. Implicit identity representation conditioned memory compensation network for talking head video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 23062–23072, 2023.
  14. 14.Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. Depth-aware generative adversarial network for talking head video generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 3397–3406, 2022a.
  15. 15.Yang Hong, Bo Peng, Haiyao Xiao, Ligang Liu, and Juyong Zhang. Headnerf: A real-time nerf-based parametric head model. In CVPR, pp. 20374–20384, June 2022b.
  16. 16.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3451–3460, 2021.
  17. 17.Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren. Prodiff: Progressive fast diffusion model for high-quality text-to-speech. In Proceedings of the 30th ACM International Conference on Multimedia, pp. 2595–2605, 2022.
  18. 18.Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. Audiogpt: Understanding and generating speech, music, sound, and talking head. arXiv preprint arXiv:2304.12995, 2023a.
  19. 19.Rongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Linjun Li, Zhenhui Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiang Yin, et al. Av-transpeech: Audio-visual robust speech-to-speech translation. arXiv preprint arXiv:2305.15403, 2023b.
  20. 20.Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Niessner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Christian Theobalt. Deep video portraits. ACM transactions on graphics (TOG), 37(4):1–14, 2018.
  21. 21.Hyeongwoo Kim, Mohamed Elgharib, Michael Zollhöfer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt. Neural style-preserving visual dubbing. ACM Transactions on Graphics (TOG), 38(6):1–13, 2019.
  22. 22.Shaoxu Li. Ophavatars: One-shot photo-realistic head avatars. arXiv preprint arXiv:2307.09153, 2023.
  23. 23.Weichuang Li, Longhao Zhang, Dong Wang, Bin Zhao, Zhigang Wang, Mulin Chen, Bang Zhang, Zhongjian Wang, Liefeng Bo, and Xuelong Li. One-shot high-fidelity talking-head synthesis with deformable neural radiance field. In CVPR, pp. 17969–17978, 2023a.
  24. 24.Xueting Li, Shalini De Mello, Sifei Liu, Koki Nagano, Umar Iqbal, and Jan Kautz. Generalizable one-shot neural head avatar. arXiv preprint arXiv:2306.08768, 2023b.
  25. 25.Yuanxun Lu, Jinxiang Chai, and Xun Cao. Live speech portraits: real-time photorealistic talking-head animation. ACM Transactions on Graphics, 40(6):1–17, 2021.
  26. 26.Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan, Esha Uboweja, Michael Hays, Fan Zhang, Chuo-Ling Chang, Ming Guang Yong, Juhyun Lee, et al. Mediapipe: A framework for building perception pipelines. arXiv preprint arXiv:1906.08172, 2019.
  27. 27.Zhiyuan Ma, Xiangyu Zhu, Guo-Jun Qi, Zhen Lei, and Lei Zhang. Otavatar: One-shot talking face avatar with controllable tri-plane rendering. In CVPR, pp. 16901–16910, 2023.
  28. 28.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, pp. 405–421. Springer, 2020.
  29. 29.Youxin Pang, Yong Zhang, Weize Quan, Yanbo Fan, Xiaodong Cun, Ying Shan, and Dong-ming Yan. Dpe: Disentanglement of pose and expression for general video portrait editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 427–436, 2023.
  30. 30.Omkar Parkhi, Andrea Vedaldi, and Andrew Zisserman. Deep face recognition. In BMVC 2015-Proceedings of the British Machine Vision Conference 2015. British Machine Vision Association, 2015.
  31. 31.P. Paysan, R. Knothe, B. Amberg, S. Romdhani, and T. Vetter. A 3d face model for pose and illumination invariant face recognition. Proceedings of the 6th IEEE International Conference on Advanced Video and Signal based Surveillance (AVSS) for Security, Safety and Monitoring in Smart Environments, 2009.
  32. 32.Bui Tuong Phong. Illumination for computer generated pictures. In Seminal graphics: pioneering efforts that shaped the field, pp. 95–101. 1998.
  33. 33.KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In ACM MM, pp. 484–492, 2020.
  34. 34.Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real images. ACM Transactions on graphics (TOG), 42(1):1–13, 2022.
  35. 35.Shuai Shen, Wanhua Li, Zheng Zhu, Yueqi Duan, Jie Zhou, and Jiwen Lu. Learning dynamic facial radiance fields for few-shot talking head synthesis. In ECCV, pp. 666–682. Springer, 2022.
  36. 36.Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. First order motion model for image animation. NIPS, 32, 2019.
  37. 37.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  38. 38.Jingxiang Sun, Xuan Wang, Lizhen Wang, Xiaoyu Li, Yong Zhang, Hongwen Zhang, and Yebin Liu. Next3d: Generative neural texture rasterization for 3d-aware head avatars. In CVPR, 2023.
  39. 39.Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. arXiv preprint arXiv:2109.07161, 2021.
  40. 40.Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou, Xiaokang Chen, Dongliang He, Tianshu Hu, Jingtuo Liu, Gang Zeng, and Jingdong Wang. Real-time neural radiance talking portrait synthesis via audio-spatial decomposition. arXiv preprint arXiv:2211.12368, 2022.
  41. 41.Alex Trevithick, Matthew Chan, Michael Stengel, Eric Chan, Chao Liu, Zhiding Yu, Sameh Khamis, Manmohan Chandraker, Ravi Ramamoorthi, and Koki Nagano. Real-time radiance fields for single-image portrait view synthesis. ACM Transactions on Graphics (TOG), 42(4):1–15, 2023.
  42. 42.Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In CVPR, pp. 10039–10049, 2021.
  43. 43.Yaohui Wang, Di Yang, Francois Bremond, and Antitza Dantcheva. Latent image animator: Learning to animate images via latent space navigation. arXiv preprint arXiv:2203.09043, 2022.
  44. 44.Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, and Qingshan Deng. Imitating arbitrary talking style for realistic audio-driven talking face synthesis. In Proceedings of the 29th ACM International Conference on Multimedia, pp. 1478–1486, 2021.
  45. 45.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. NIPS, 34:12077–12090, 2021.
  46. 46.Zhenhui Ye, Jinzheng He, Ziyue Jiang, Rongjie Huang, Jiawei Huang, Jinglin Liu, Yi Ren, Xiang Yin, Zejun Ma, and Zhou Zhao. Geneface++: Generalized and stable real-time audio-driven 3d talking face generation. arXiv preprint arXiv:2305.00787, 2023a.
  47. 47.Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu, Jinzheng He, and Zhou Zhao. Geneface: Generalized and high-fidelity audio-driven 3d talking face synthesis. In ICLR, 2023b.
  48. 48.Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu, Chen Zhang, Xiang Yin, Zejun Ma, and Zhou Zhao. Ada-tta: Towards adaptive high-quality text-to-talking avatar synthesis. arXiv preprint arXiv:2306.03504, 2023c.
  49. 49.Zipeng Ye, Mengfei Xia, Ran Yi, Juyong Zhang, Yu-Kun Lai, Xuwei Huang, Guoxin Zhang, and Yong-jin Liu. Audio-driven talking face video generation with dynamic convolution kernels. IEEE Transactions on Multimedia, 2022.
  50. 50.Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu. Audio-driven talking face video generation with learning-based personalized head pose. arXiv preprint arXiv:2002.10137, 2020.
  51. 51.Fei Yin, Yong Zhang, Xuan Wang, Tengfei Wang, Xiaoyu Li, Yuan Gong, Yanbo Fan, Xiaodong Cun, Ying Shan, Cengiz Oztireli, et al. 3d gan inversion with facial symmetry prior. In CVPR, pp. 342–351, 2023.
  52. 52.Wangbo Yu, Yanbo Fan, Yong Zhang, Xuan Wang, Fei Yin, Yunpeng Bai, Yan-Pei Cao, Ying Shan, Yang Wu, Zhongqian Sun, et al. Nofa: Nerf-based one-shot facial avatar reconstruction. In SIGGRAPH, pp. 1–12, 2023.
  53. 53.Jian Zhao and Hui Zhang. Thin-plate spline motion model for image animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3657–3666, 2022.
  54. 54.Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In CVPR, pp. 4176–4186, 2021.
  55. 55.Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. Makelttalk: speaker-aware talking-head animation. ACM Transactions on Graphics (TOG), 39(6): 1–15, 2020.
  56. 56.Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. CelebV-HQ: A large-scale video facial attributes dataset. In ECCV, 2022.
  57. 57.Xiangyu Zhu, Zhen Lei, Xiaoming Liu, Hailin Shi, and Stan Z Li. Face alignment across large poses: A 3d solution. In CVPR, pp. 146–155, 2016.

Citation

MLA
Ye, Z., et al. “Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis”. arXiv, 2024, http://arxiv.org/abs/2401.08503v3.
APA
Ye, Z., Zhong, T., Ren, Y., Yang, J., Li, W., Huang, J., Jiang, Z., He, J., Huang, R., Liu, J., Zhang, C., Yin, X., Ma, Z., & Zhao, Z. (2024). Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis. arXiv. http://arxiv.org/abs/2401.08503v3
Chicago
Ye, Z., T. Zhong, Y. Ren, et al. 2024. “Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis”. arXiv. http://arxiv.org/abs/2401.08503v3.
Harvard
Ye, Z. et al. (2024) “Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.08503v3.
Vancouver
1. Ye Z, Zhong T, Ren Y, et al (2024) Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis. arXiv

BibTeX

@article{ye2024real3d,
  title = {Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis},
  author = {Ye, Zhenhui and Zhong, Tianyun and Ren, Yi and Yang, Jiaqi and Li, Weichuang and Huang, Jiawei and Jiang, Ziyue and He, Jinzheng and Huang, Rongjie and Liu, Jinglin and Zhang, Chen and Yin, Xiang and Ma, Zejun and Zhao, Zhou},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.08503v3},
  eprint = {2401.08503}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors