Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
Zhenhui YeTianyun ZhongYi RenJiaqi YangWeichuang LiJiawei HuangZiyue JiangJinzheng HeRongjie HuangJinglin Liu
Presents a one-shot 3D talking portrait framework that distills generative 3D priors and models full head-torso dynamics to synthesize realistic, audio- and video-driven avatar videos from a single unseen image.
Generating realistic talking portrait videos from a single still image is an important capability for interactive digital media, video conferencing, customer service, and virtual avatars. However, existing methods face major challenges: two-dimensional approaches suffer from severe visual distortion during large head movements, while three-dimensional methods either require hours of individual training per identity or struggle with facial animation stability and identity preservation. Furthermore, previous systems typically focus solely on head generation, neglecting natural torso movements and realistic background integration.
The article introduces and evaluates Real3D-Portrait, a one-shot framework designed to generate realistic, three-dimensional talking portrait videos from a single unseen reference image, driven either by an input video or an audio track.
The approach uses a modular, multi-stage architecture trained across large-scale video and synthetic datasets. First, an image-to-plane model pre-trained on multi-view synthetic data reconstructs an accurate canonical three-dimensional representation from a single photo. A lightweight motion adapter then animates facial expressions using standardized coordinate codes without distorting underlying geometry. To ensure natural full-frame composition, a dedicated super-resolution model independently handles the head, warps the torso based on keypoint movement, inpaints the background, and blends them using an occlusion-aware mask. Finally, a generalized audio-to-motion model translates speech audio into precise facial motion parameters with controllable eye blinking and mouth amplitude.
Evaluation shows that the system achieves state-of-the-art results across both video- and audio-driven benchmarks. In video-driven reenactment, Real3D-Portrait attained the highest image fidelity (FID score of 37.50 for same-identity and 42.37 for cross-identity) and superior identity similarity compared to previous one-shot methods. In audio-driven generation, the framework demonstrated lip-synchronization scores (Sync score of 6.565) and visual quality comparable to person-specific models that require lengthy per-user training. Human perceptual studies confirmed that viewers rated Real3D-Portrait higher in identity preservation, visual smoothness, and lip synchronization than competing one-shot tools.
These findings indicate that high-quality, three-dimensional talking portraits no longer require costly per-individual model fine-tuning or compromise full-body visual coherence. This significantly lowers computational costs and production turnaround times for deploying personalized digital avatars at scale. Because the system cleanly separates head, torso, and background elements, it also offers practical flexibility, such as customizable backgrounds.
Organizations looking to implement digital avatar systems should adopt modular, one-shot three-dimensional pipelines over rigid, identity-specific approaches. To manage ethical and security risks associated with deepfake technologies, deployers should incorporate watermarking and tracking safeguards. Future development should focus on integrating advanced neural inpainting for background synthesis and gathering broader training data for extreme side-view head poses. While confidence in the evaluated head poses and standard portrait angles is high, caution is advised when deploying the model in scenarios with extreme rotational head movement.
- Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). EG3D establishes the foundational tri-plane neural volume rendering representation that enables efficient, high-resolution 3D-aware face synthesis.
- Paper: Learning a model of facial shape and expression from 4D scans, Tianye Li et al. (2017). FLAME introduces the parametric head model that provides the decoupled identity, expression, and head-pose parameterization used to drive 3D portrait animations.
- Paper: First Order Motion Model for Image Animation, Aliaksandr Siarohin et al. (2019). This work establishes the keypoint-based warping and occlusion-aware masking formulation used for animating and composing 2D/3D human body parts.
- Paper: StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based Generator, Jiazhi Guan et al. (2023). StyleSync details key mechanisms for speech-driven lower-face lip synchronization and style-based identity preservation.
- Paper: Nerfies: Deformable Neural Radiance Fields, Keunhong Park et al. (2020). Nerfies introduces the canonical-to-deformed field formulation essential for animating dynamic and non-rigid 3D neural radiance fields.
- Paper: Learning Neural Parametric Head Models, Simon Giebenhain et al. (2023). NPHM demonstrates learning disentangled identity and expression representations in neural parametric head fields for high-fidelity 3D facial reconstruction.
- Paper: LRM: Large Reconstruction Model for Single Image to 3D, Yicong Hong et al. (2024). LRM provides the core feed-forward transformer architecture for directly reconstructing 3D tri-plane NeRF representations from a single reference image.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Zero-1-to-3 develops zero-shot single-image novel-view synthesis priors that underpin feedforward multi-view 3D portrait reconstruction.
- Paper: VideoBooth: Diffusion-based Video Generation with Image Prompts, Yuming Jiang et al. (2024). VideoBooth extends single-image conditioning principles to general video diffusion by injecting fine-grained visual attention features across frames without inference fine-tuning.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). UniReal generalizes identity-consistent image manipulation and dynamic portrait synthesis into a unified diffusion transformer framework trained on real-world video dynamics.
- Paper: AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. (2024). AnimateDiff applies generalized motion modeling to animate personalized visual assets without requiring subject-specific retraining.
- Paper: DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation, Hong Chen et al. (2024). DisenBooth extends identity-preserving visual conditioning by disentangling subject identity from background and pose features during generative tuning.
