StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based Generator

Jiazhi GuanZhanwang ZhangHang ZhouTianshu HuKaisiyuan WangDongliang HeHaocheng FengJingtuo LiuErrui DingZiwei Liu

article2023CVPR112 citations

Proposes a style-based lip synchronization framework that uses mask-guided spatial encoding and few-shot generator refinement to accurately match speech audio while preserving identity and personalized talking styles across diverse target videos.

Listen

Generating realistic lip-synchronized video from arbitrary audio tracks is vital for digital human production, film dubbing, and entertainment media. Current solutions struggle to balance visual quality and generalization. Existing techniques either regenerate the entire head—causing unstable backgrounds and facial distortion—or require extensive subject-specific video data and fragile structural representations that limit practical use in seamless video editing.

The article evaluates StyleSync, a machine learning framework designed to modify lower-face video regions to match driving audio while keeping surrounding scenes untouched. The main objective is to demonstrate high-fidelity, generalized lip synchronization on single reference images alongside efficient few-shot personalization that preserves unique individual speaking styles.

To achieve this, the authors modified a style-based image generation network by encoding facial structural features into noise space while injecting speech and visual context through modulated convolution layers. The system was trained in a self-reconstruction format using video datasets from LRW and VoxCeleb2, evaluated on 256x256 resolution frames, and benchmarked against leading methods using standard visual quality metrics, landmark distance measurements, alignment scores, and a 15-participant user study.

The evaluation yielded several key findings. First, the generalized framework outperformed existing methods in visual clarity, scoring 0.85 in structural similarity on the LRW benchmark compared to 0.80 for the strongest baseline, while maintaining precise audio-visual alignment. Second, applying personalized optimization using less than 10 seconds of target video further improved visual realism and identity preservation, reducing mouth landmark distance error by roughly 27% on LRW. Third, human evaluators strongly favored the framework, giving it an average rating of 4.52 out of 5 for generation quality and 4.06 for video realness, markedly higher than competing approaches.

These results indicate that high-fidelity lip dubbing can be deployed rapidly without long capture sessions or expensive per-person training pipelines. Incorporating few-shot personalization offers media workflows a practical method to preserve speaker identity and speaking nuances at low operational overhead, eliminating common visual artifacts and background drift found in full-head generation methods.

Organizations exploring automated dubbing and visual media production should consider adopting modular, style-based architectures that restrict synthesis to masked facial regions. Given that personalization can cause minor reductions in metric-based mouth opening due to target-specific speaking quirks, teams should conduct visual audits alongside automated scoring to assess dubbing realism.

Certain constraints should be noted. Because the model relies on a fixed lower-face mask, it cannot adjust overall head pose or upper-face emotional expressions, and faces with unusually large jaw structures may exceed the masked boundary. Furthermore, the risk of deepfake misuse necessitates controlled distribution, warranting restricted access policies when deploying core models.

arXiv: 2305.05445
Cover for StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based Generator

Abstract

Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model’s generalization ability. Previous studies either require long-term data for training or produce a similar movement pattern on all subjects with low quality. In this paper, we propose StyleSync, an effective framework that enables high-fidelity lip synchronization. We identify that a style-based generator would sufficiently enable such a charming property on both one-shot and few-shot scenarios. Specifically, we design a mask-guided spatial information encoding module that preserves the details of the given face. The mouth shapes are accurately modified by audio through modulated convolutions. Moreover, our design also enables personalized lip-sync by introducing style space and generator refinement on only limited frames. Thus the identity and talking style of a target person could be accurately preserved. Extensive experiments demonstrate the effectiveness of our method in producing high-fidelity results on a variety of scenes. Resources can be found at https://hangz-nju-cuhk.github.io/projects/StyleSync.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Audio-Driven Facial Animation
  • 2.2. Style-based Generator for Faces
  • 3. Methodology
  • 3.1. Modifying Style-based Generator for Lip Sync
  • 3.2. Backbone Training Objectives
  • 3.3. Personalized Optimization
  • 4. Experiments
  • 4.1. Quantitative Evaluation
  • 4.2. Qualitative Evaluation
  • 5. Conclusion and Discussions
  • References

Knowls

  1. Knowl 1 — StyleSync Architecture and Mask-Guided Spatial Information Encoding

    model/method

    StyleSync adapts a StyleGAN2 generator for arbitrary-subject and personalized audio-driven lip synchronization by decomposing input information into spatial contextual noise maps and style-modulated convolutions.

    Given a target video frame ItI_t with lower-face mask MM (yielding masked frame Itm=(1−M)∗ItI_t^m = (1 - M) * I_t) and a random reference face frame IrefI_{ref} from the same video sequence, the 6-channel concatenation [Itm,Iref][I_t^m, I_{ref}] is processed by a visual encoder EfaceE_{face} to produce multi-scale feature maps F={F1,…,FL}F = \{F_1, \dots, F_L\}, where LL is the total number of generative layers (L=14L = 14).

    To prevent the reference frame's mouth structure from overriding the target speech dynamics while preserving high-resolution background and facial context, a mask-guided filtering is applied to the high-resolution layers:

    Fl′=(1−M)∗Fl,for l>⌊L2⌋F'_l = (\mathbf{1} - M) * F_l, \quad \text{for } l > \left\lfloor \frac{L}{2} \right\rfloor

    where 1\mathbf{1} is an all-ones matrix matching the spatial dimensions of MM. For l≤⌊L/2⌋l \le \lfloor L/2 \rfloor, Fl′=FlF'_l = F_l. These modified feature maps N={F1′,…,FL′}N = \{F'_1, \dots, F'_L\} are injected directly into the generator layers as spatial noise inputs.

    Simultaneously, speech audio features faf_a extracted by an audio encoder EaE_a from spectrogram input ata_t are concatenated with the bottleneck face embedding fff_f of EfaceE_{face} to form the base style vector w=concat(fa,ff)∈Ww = \text{concat}(f_a, f_f) \in \mathcal{W}. The style vector ww modulates the generator's convolution weights across all LL layers to synthesize the reconstructed frame It′=G(N,W)I'_t = G(N, W).

  2. Knowl 2 — Few-Shot Personalized Optimization Procedure in StyleSync

    model/method

    To adapt a generalized StyleSync model to the unique talking style and facial identity of a specific target subject using short video footage (under one minute, and as short as 10 seconds), StyleSync introduces a personalized optimization procedure over the extended style space W+\mathcal{W}^+ and the generator weights.

    Given a template video Vˉ={Iˉ1,…,IˉT}\bar{V} = \{\bar{I}_1, \dots, \bar{I}_T\} and audio aˉ={aˉ1,…,aˉT}\bar{a} = \{\bar{a}_1, \dots, \bar{a}_T\}, pairs of target frames Iˉt\bar{I}_t and reference frames Iˉref\bar{I}_{ref} are sampled:

    1. Fixed Encoders: The audio encoder EaE_a (which extracts person-agnostic phoneme dynamics) and the face encoder EfaceE_{face} (which extracts visual identity priors) are kept frozen to prevent overfitting.
    2. Style Space Optimization: Layer-specific style displacements ΔW={Δw1,…,ΔwL}=MLPΔw(ff)\Delta W = \{\Delta w_1, \dots, \Delta w_L\} = \text{MLP}_{\Delta w}(f_f) are learned via a multi-layer perceptron MLPΔw\text{MLP}_{\Delta w} taking the bottleneck face feature fff_f as input, expanding the style vector to W={w+Δw1,…,w+ΔwL}∈W+W = \{w + \Delta w_1, \dots, w + \Delta w_L\} \in \mathcal{W}^+.
    3. Generator Tuning: The parameters PP of the generator network GG are allowed to undergo bounded parameter shifts ΔP\Delta P.

    Training is conducted for 5 epochs with a 1:1 mixture of personalized video data and general training data to maintain synthesis stability.

  3. Knowl 3 — Backbone Training Objectives of StyleSync

    equation

    The generalized StyleSync backbone is trained end-to-end using a composite loss function balancing pixel reconstruction, perceptual similarity, adversarial realism, and audio-visual synchronization:

    Lg=Ladv+λrLrec+λsLsync\mathcal{L}_g = \mathcal{L}_{adv} + \lambda_r \mathcal{L}_{rec} + \lambda_s \mathcal{L}_{sync}

    where λr=10\lambda_r = 10 and λs=1\lambda_s = 1. The constituent loss terms are:

    1. Reconstruction Loss: Combines ℓ1\ell_1 pixel error and multi-layer VGG-19 perceptual error between the predicted frame It′=G(N,W)I'_t = G(N, W) and the unmasked ground-truth frame ItI_t:

    Lrec=∥It′−It∥1+∑m=1Nvgg∥VGGm(It′)−VGGm(It)∥1\mathcal{L}_{rec} = \|I'_t - I_t\|_1 + \sum_{m=1}^{N_{vgg}} \|\text{VGG}_m(I'_t) - \text{VGG}_m(I_t)\|_1

    1. Adversarial Loss: Supervised by a discriminator DD initialized from pre-trained StyleGAN2:

    Ladv=min⁡Gmax⁡D(EIt[log⁡D(It)]+EIt′[log⁡(1−D(It′))])\mathcal{L}_{adv} = \min_G \max_D \left( \mathbb{E}_{I_t}[\log D(I_t)] + \mathbb{E}_{I'_t}[\log(1 - D(I'_t))] \right)

    1. Lip-Sync Loss: Supervised by a pretrained SyncNet consisting of visual encoder SvS_v and audio encoder SaS_a, evaluated over sequences of 5 consecutive predicted frames It:t+4′I'_{t:t+4} and corresponding audio spectrograms at:t+4a_{t:t+4} via cosine similarity:

    Lsync=−Sv(It:t+4′)⊤⋅Sa(at:t+4)∥Sv(It:t+4′)∥2∥Sa(at:t+4)∥2\mathcal{L}_{sync} = - \frac{S_v(I'_{t:t+4})^\top \cdot S_a(a_{t:t+4})}{\|S_v(I'_{t:t+4})\|_2 \|S_a(a_{t:t+4})\|_2}

  4. Knowl 4 — Personalized Optimization Loss Objective

    equation

    During the personalized optimization stage of StyleSync, the generalized objective Lg\mathcal{L}_g is augmented with ℓ2\ell_2 regularization penalties on the learned W+\mathcal{W}^+ offsets ΔW\Delta W and generator parameter shifts ΔP\Delta P:

    Lp=Lg+λp(∑i,j∣ΔWij∣22+∑m,n∣ΔPnm∣22)\mathcal{L}_p = \mathcal{L}_g + \lambda_p \left( \sum_{i,j} |\Delta \mathbf{W}_i^j|_2^2 + \sum_{m,n} |\Delta \mathbf{P}_n^m|_2^2 \right)

    where:

    • Lg=Ladv+λrLrec+λsLsync\mathcal{L}_g = \mathcal{L}_{adv} + \lambda_r \mathcal{L}_{rec} + \lambda_s \mathcal{L}_{sync} is the generalized task loss,
    • ΔWij\Delta \mathbf{W}_i^j denotes the jj-th scalar parameter of the ii-th layer style displacement vector Δwi∈ΔW\Delta w_i \in \Delta W,
    • ΔPnm\Delta \mathbf{P}_n^m denotes the nn-th scalar weight parameter displacement in the mm-th tensor of generator parameters PP,
    • λp=1\lambda_p = 1 balances the regularizers against the reconstruction and synchronization objectives to ensure displacements remain locally bounded.
  5. Knowl 5 — Quantitative Evaluation on LRW and VoxCeleb2 Benchmarks

    data/table

    The quantitative performance of generalized StyleSync (Ours-G) and personalized StyleSync (Ours-P) was benchmarked against leading lip-sync and talking-head models on the test splits of the Lip Reading in the Wild (LRW) and VoxCeleb2 datasets. Visual generation quality was measured by SSIM and PSNR; mouth landmark synchronization accuracy was measured by Lip Landmark Distance (LMD, lower is better); audio-visual sync confidence was measured by SyncNet confidence score (SyncconfSync_{conf}); and identity preservation was measured by ArcFace cosine distance (DIDD_{ID}).

    Method LRW VoxCeleb2
    SSIM ↑\uparrow PSNR ↑\uparrow LMD ↓\downarrow Syncconf↑\text{Sync}_{conf} \uparrow DID↑D_{ID} \uparrow SSIM ↑\uparrow PSNR ↑\uparrow LMD ↓\downarrow Syncconf↑\text{Sync}_{conf} \uparrow DID↑D_{ID} \uparrow
    MakeitTalk 0.69 29.83 2.75 3.88 0.79 0.63 28.38 6.94 2.15 0.71
    PC-AVS 0.79 30.26 1.84 7.19 0.82 0.71 29.53 2.75 8.16 0.74
    Wav2Lip 0.79 30.54 1.28 7.39 0.90 0.80 30.53 1.92 8.90 0.90
    Wav2Lip-H 0.80 31.38 1.20 7.19 0.88 0.81 30.53 1.87 8.35 0.90
    GT 1.00 100.00 0.00 7.65 1.00 1.00 100.00 0.00 7.71 1.00
    Ours-G 0.85 31.78 1.18 7.25 0.89 0.79 31.00 1.47 8.25 0.90
    Ours-P 0.88 32.66 0.86 6.35 0.93 0.82 31.54 1.15 7.26 0.93

    Key observations:

    1. Ours-G outperforms prior methods in image reconstruction quality (SSIM 0.85 vs. 0.80 for Wav2Lip-H on LRW; PSNR 31.78 vs. 31.38).
    2. Personalized optimization (Ours-P) substantially improves landmark accuracy (LMD 0.86 on LRW, 1.15 on VoxCeleb2) and identity preservation (DIDD_{ID} reaches 0.93 on both).
    3. SyncconfSync_{conf} drops slightly under Ours-P because personalized speaking styles involve smaller, idiosyncratic mouth openings that deviate from the generalized SyncNet training distribution, even while matching the true target motion more closely.
  6. Knowl 6 — Mean Opinion Score (MOS) User Study

    data/table

    A user study with 15 participants evaluated 63 randomly selected test videos from LRW and VoxCeleb2 across three perceptual dimensions using a Mean Opinion Score (MOS) protocol (ratings from 1 to 5, where 5 is best).

    Metric \ Approach MakeitTalk PC-AVS Wav2Lip Wav2Lip-H Ours-G
    Lip-Sync Quality 2.06 3.00 3.49 3.67 4.24
    Generation Quality 2.63 2.16 1.87 3.42 4.52
    Video Realness 1.89 2.16 2.17 2.98 4.06

    StyleSync (Ours-G) achieves the highest scores across all subjective metrics, surpassing the high-resolution baseline Wav2Lip-H by +0.57 in Lip-Sync Quality, +1.10 in Generation Quality, and +1.08 in Video Realness.

  7. Knowl 7 — Ablation Analysis of StyleSync Components

    data/table

    Ablation experiments conducted on the VoxCeleb2 dataset evaluated the isolated contributions of mask-guided spatial encoding, the SyncNet loss, and W+\mathcal{W}^+ displacement optimization:

    Method SSIM ↑\uparrow PSNR ↑\uparrow LMD ↓\downarrow Syncconf↑\text{Sync}_{conf} \uparrow DID↑D_{ID} \uparrow
    w/o mask 0.81 31.20 1.80 7.89 0.90
    w/o sync 0.80 31.34 1.55 7.41 0.90
    Ours-G 0.79 31.00 1.47 8.25 0.90
    P w/o ΔW\Delta W 0.82 31.51 1.32 7.31 0.92
    Ours-P 0.82 31.54 1.15 7.26 0.93

    Findings:

    • Spatial Masking (w/o mask vs. Ours-G): Omitting the mask operation Fl′=(1−M)∗FlF'_l = (\mathbf{1}-M)*F_l degrades LMD from 1.47 to 1.80 and SyncconfSync_{conf} from 8.25 to 7.89, confirming that unmasked spatial noise injects unwanted mouth shape features from the reference frame.
    • Lip-Sync Loss (w/o sync vs. Ours-G): Omitting Lsync\mathcal{L}_{sync} degrades LMD to 1.55 and SyncconfSync_{conf} to 7.41, causing temporal lip motion inconsistencies.
    • Style Space Offsets (P w/o ΔW\Delta W vs. Ours-P): Eliminating ΔW\Delta W during personalized optimization worsens LMD (1.32 vs. 1.15) and identity cosine distance DIDD_{ID} (0.92 vs. 0.93), showing that W+\mathcal{W}^+ displacement is necessary to fit individual speaking styles.
  8. Knowl 8 — Experimental and Training Setup of StyleSync

    experimental setup

    Video and audio processing along with network training follow standardized protocols:

    • Video Processing: All video sequences are sampled at 25 fps. Faces are aligned via eye landmarks and cropped to 256×256256 \times 256 pixels. A U-shaped binary mask MM erases the mouth, cheek, and jaw areas in the target frame.
    • Audio Processing: Driving audio waveforms are converted to spectrograms matching the temporal window of the corresponding video frames.
    • Backbone Architecture: The generator contains L=14L = 14 style-convolution layers based on StyleGAN2. The discriminator is initialized from pre-trained StyleGAN2 weights. Hyperparameters are set to λr=10\lambda_r = 10, λs=1\lambda_s = 1, and λp=1\lambda_p = 1.
    • Personalized Fine-Tuning: For subject-specific adaptation, models are fine-tuned for 5 epochs on less than 1 minute (typically 10 seconds) of target footage mixed in a 1:1 ratio with general training data.
  9. Knowl 9 — Structural Limitations of StyleSync

    limitation

    The StyleSync framework exhibits two primary structural limitations:

    1. Fixed Head Pose and Expression Outside Mask: Because StyleSync operates by inpainting a lip-synced mouth directly into an existing target video frame under a predefined lower-face mask, it cannot modify the subject's overall head pose, upper-face expressions, or head dynamics to reflect emotional content present in the driving audio.
    2. Mask Boundary Excursions: In extreme cases where a target subject has an unusually large jaw or makes wide jaw movements that exceed the spatial extent of the fixed U-shaped mask MM, visual boundary artifacts can appear where the synthesized area meets the unmasked background.

Coverage note — None was omitted; all key architectural components, training objectives, personalization algorithms, quantitative benchmarks, user studies, ablation experiments, and stated limitations are fully represented.

References

  1. 1.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan: How to embed images into the stylegan latent space? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4432–4441, 2019.
  2. 2.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan++: How to edit the embedded images? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8296–8305, 2020.
  3. 3.Yuval Alaluf, Or Patashnik, and Daniel Cohen-Or. Restyle: A residual-based stylegan encoder via iterative refinement. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6711–6720, 2021.
  4. 4.Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, and Amit Bermano. Hyperstyle: Stylegan inversion with hypernetworks for real image editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18511–18521, 2022.
  5. 5.Egor Burkov, Igor Pasechnik, Artur Grigorev, and Victor Lempitsky. Neural head reenactment with latent pose descriptors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13786–13795, 2020.
  6. 6.Lele Chen, Guofeng Cui, Ziyi Kou, Haitian Zheng, and Chenliang Xu. What comprises a good talking-head video generation?: A survey and benchmark. arXiv preprint arXiv:2005.03201, 2020.
  7. 7.Lele Chen, Guofeng Cui, Celong Liu, Zhong Li, Ziyi Kou, Yi Xu, and Chenliang Xu. Talking-head generation with rhythmic head motion. In European Conference on Computer Vision, pages 35–51. Springer, 2020.
  8. 8.Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. Lip movements generation at a glance. In Proceedings of the European Conference on Computer Vision (ECCV), pages 520–535, 2018.
  9. 9.Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7832–7841, 2019.
  10. 10.J. S. Chung, A. Nagrani, and A. Zisserman. Voxceleb2: Deep speaker recognition. In INTERSPEECH, 2018.
  11. 11.Joon Son Chung and Andrew Zisserman. Lip reading in the wild. In Asian conference on computer vision, pages 87–103. Springer, 2016.
  12. 12.Joon Son Chung and Andrew Zisserman. Out of time: automated lip sync in the wild. In ACCV, 2016.
  13. 13.Yu Deng, Jiaolong Yang, Sicheng Xu, Dong Chen, Yunde Jia, and Xin Tong. Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 0–0, 2019.
  14. 14.Bo Fan, Lijuan Wang, Frank K Soong, and Lei Xie. Photoreal talking head with deep bidirectional lstm. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4884–4888. IEEE, 2015.
  15. 15.Yingruo Fan, Zhaojiang Lin, Jun Saito, Wenping Wang, and Taku Komura. Faceformer: Speech-driven 3d facial animation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18770–18780, 2022.
  16. 16.Yudong Guo, Keyu Chen, Sen Liang, Yongjin Liu, Hujun Bao, and Juyong Zhang. Ad-nerf: Audio driven neural radiance fields for talking head synthesis. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  17. 17.Amir Jamaludin, Joon Son Chung, and Andrew Zisserman. You said that?: Synthesising talking faces from audio. International Journal of Computer Vision, 127(11):1767–1779, 2019.
  18. 18.Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao. Eamm: One-shot emotional talking face via audio-based emotion-aware motion model. SIGGRAPH, 2022.
  19. 19.Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. Audio-driven emotional video portraits. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14080–14089, 2021.
  20. 20.Tero Karras, Timo Aila, Samuli Laine, Antti Herva, and Jaakko Lehtinen. Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics (TOG), 36(4):1–12, 2017.
  21. 21.Tero Karras, Miika Aittala, Samuli Laine, Erik H¨ark ¨onen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. In Proc. NeurIPS, 2021.
  22. 22.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019.
  23. 23.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of StyleGAN. In Proc. CVPR, 2020.
  24. 24.Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, and Chris Bregler. Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2755–2764, 2021.
  25. 25.Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. Expressive talking head generation with granular audio-visual control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3387–3396, 2022.
  26. 26.Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou. Semantic-aware implicit neural audio-driven video portrait generation. ECCV, 2022.
  27. 27.Yuanxun Lu, Jinxiang Chai, and Xun Cao. Live speech portraits: real-time photorealistic talking-head animation. ACM Transactions on Graphics (TOG), 40(6):1–17, 2021.
  28. 28.Yifeng Ma, Suzhen Wang, Zhipeng Hu, Changjie Fan, Tangjie Lv, Yu Ding, Zhidong Deng, and Xin Yu. Styletalk: One-shot talking head generation with controllable speaking styles. AAAI, 2023.
  29. 29.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision, pages 405–421. Springer, 2020.
  30. 30.A. Nagrani, J. S. Chung, and A. Zisserman. Voxceleb: a large-scale speaker identification dataset. In INTERSPEECH, 2017.
  31. 31.Se Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi, and Yong Man Ro. Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory. In AAAI Conference on Artificial Intelligence. Association for the Advancement of Artificial Intelligence, 2022.
  32. 32.KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the 28th ACM International Conference on Multimedia, pages 484–492, 2020.
  33. 33.Alexander Richard, Michael Z¨ollhofer, Yandong Wen, Fernando de la Torre, and Yaser Sheikh. Meshtalk: 3d face animation from speech using cross-modality disentanglement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  34. 34.Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. Encoding in style: a stylegan encoder for image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2287–2296, 2021.
  35. 35.Daniel Roich, Ron Mokady, Amit H Bermano, and Daniel Cohen-Or. Pivotal tuning for latent-based editing of real images. ACM Trans. Graph., 2021.
  36. 36.Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. TPAMI, 2020.
  37. 37.Yujun Shen and Bolei Zhou. Closed-form factorization of latent semantics in gans. In CVPR, 2021.
  38. 38.Linsen Song, Wayne Wu, Chen Qian, Ran He, and Chen Change Loy. Everybody’s talkin’: Let me talk as you want. arXiv preprint arXiv:2001.05201, 2020.
  39. 39.Yang Song, Jingwen Zhu, Dawei Li, Xiaolong Wang, and Hairong Qi. Talking face generation by conditional recurrent adversarial network. arXiv preprint arXiv:1804.04786, 2018.
  40. 40.Yasheng Sun, Hang Zhou, Ziwei Liu, and Hideki Koike. Speech2talking-face: Inferring and driving a face with synchronized audio-visual representation. In IJCAI, 2021.
  41. 41.Yasheng Sun, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Zhibin Hong, Jingtuo Liu, Errui Ding, Jingdong Wang, Ziwei Liu, and Koike Hideki. Masked lip-sync prediction by audio-visual contextual exploitation in transformers. In SIGGRAPH Asia 2022 Conference Papers, pages 1–9, 2022.
  42. 42.Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (ToG), 36(4):1–13, 2017.
  43. 43.Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou, Xiaokang Chen, Dongliang He, Tianshu Hu, Jingtuo Liu, Gang Zeng, and Jingdong Wang. Real-time neural radiance talking portrait synthesis via audio-spatial decomposition. arXiv preprint arXiv:2211.12368, 2022.
  44. 44.Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner. Neural voice puppetry: Audio-driven facial reenactment. In European Conference on Computer Vision, pages 716–731. Springer, 2020.
  45. 45.Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. Designing an encoder for stylegan image manipulation. ACM Transactions on Graphics (TOG), 2021.
  46. 46.Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. Audio2head: Audio-driven one-shot talking-head generation with natural head motion. arXiv preprint arXiv:2107.09293, 2021.
  47. 47.Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. High-resolution image synthesis and semantic manipulation with conditional gans. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8798–8807, 2018.
  48. 48.Xintao Wang, Yu Li, Honglun Zhang, and Ying Shan. Towards real-world blind face restoration with generative facial prior. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  49. 49.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004.
  50. 50.Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou, Chao Duan, and Qingshan Deng. Imitating arbitrary talking style for realistic audio-driven talking face synthesis. In Proceedings of the 29th ACM International Conference on Multimedia, pages 1478–1486, 2021.
  51. 51.Chao Xu, Jiangning Zhang, Miao Hua, Qian He, Zili Yi, and Yong Liu. Region-aware face swapping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7632–7641, 2022.
  52. 52.Zhiliang Xu, Hang Zhou, Zhibin Hong, Ziwei Liu, Jiaming Liu, Zhizhi Guo, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. Styleswap: Style-based generator empowers robust face swapping. In European Conference on Computer Vision, pages 661–677. Springer, 2022.
  53. 53.Tao Yang, Peiran Ren, Xuansong Xie, and Lei Zhang. Gan prior embedded network for blind face restoration in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 672–681, 2021.
  54. 54.Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, and Yong-Jin Liu. Audio-driven talking face video generation with learning-based personalized head pose. arXiv e-prints, pages arXiv–2002, 2020.
  55. 55.Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang. Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan. ECCV, 2022.
  56. 56.Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. Few-shot adversarial learning of realistic neural talking head models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9459–9468, 2019.
  57. 57.Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo. Facial: Synthesizing dynamic talking face with implicit attribute learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 3867–3876, 2021.
  58. 58.Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. Talking face generation by adversarially disentangled audio-visual representation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 9299–9306, 2019.
  59. 59.Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4176–4186, 2021.
  60. 60.Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. Makelttalk: speaker-aware talking-head animation. ACM Transactions on Graphics (TOG), 39(6):1–15, 2020.
  61. 61.Yang Zhou, Zhan Xu, Chris Landreth, Evangelos Kalogerakis, Subhransu Maji, and Karan Singh. Visemenet: Audio-driven animator-centric speech animation. ACM Transactions on Graphics (TOG), 37(4):1–10, 2018.

Citation

MLA
Guan, J., et al. “StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator”. arXiv, 2023, http://arxiv.org/abs/2305.05445v1.
APA
Guan, J., Zhang, Z., Zhou, H., Hu, T., Wang, K., He, D., Feng, H., Liu, J., Ding, E., Liu, Z., & Wang, J. (2023). StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator. arXiv. http://arxiv.org/abs/2305.05445v1
Chicago
Guan, J., Z. Zhang, H. Zhou, et al. 2023. “StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator”. arXiv. http://arxiv.org/abs/2305.05445v1.
Harvard
Guan, J. et al. (2023) “StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.05445v1.
Vancouver
1. Guan J, Zhang Z, Zhou H, et al (2023) StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator. arXiv

BibTeX

@article{guan2023stylesync,
  title = {StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator},
  author = {Guan, Jiazhi and Zhang, Zhanwang and Zhou, Hang and Hu, Tianshu and Wang, Kaisiyuan and He, Dongliang and Feng, Haocheng and Liu, Jingtuo and Ding, Errui and Liu, Ziwei and Wang, Jingdong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.05445v1},
  eprint = {2305.05445}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE