Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
Sicheng XuYunhao DengShoukang HuYichuan WangYizhong ZhangZhanglu ChenJiaolong YangBaining Guo
Develops a real-time talking portrait video generator that combines a reference-guided causal VAE with an autoregressive Rectified Flow Transformer to produce high-fidelity streaming video directly from speech audio.
Generating realistic talking portrait videos from speech audio is critical for interactive applications such as digital assistants, education, and social companionship. However, existing high-fidelity solutions rely on heavy video diffusion foundation models that require several minutes to generate a few seconds of video. This high computational latency makes real-time, streaming deployment infeasible on standard hardware.
The article demonstrates a real-time, streamable video generation framework capable of synthesizing lifelike, arbitrary-length talking portrait videos from speech audio and one or more reference images.
To achieve this, the authors designed a causal video variational autoencoder (VAE) paired with a blockwise autoregressive transformer generator based on rectified flow matching. The framework achieves high computational efficiency by using reference image guidance inside the VAE decoder—allowing the system to focus on motion rather than static appearance—and extending causal residual auto-encoding across spatial and temporal dimensions. The generator uses key–value caching to sequentially synthesize compact video representations block-by-block. The system was trained on approximately 330 hours of video comprising about 10,000 unique identities and evaluated on two benchmark datasets across visual quality, audio synchronization, and inference latency.
The findings show that the proposed framework achieves an overall video compression ratio of 768, which is 10 to 15 times higher than the latent compression used in standard video diffusion models. The complete pipeline achieves a throughput of 42.3 frames per second at 512x512 resolution on a single modern graphics processing unit, running more than 25 times faster than leading diffusion-based talking portrait baselines. Despite its compact footprint, the model achieved the highest lip-sync accuracy and head-pose alignment across evaluated benchmarks while matching or outperforming foundation-scale baselines in visual quality metrics. Furthermore, incorporating reference image guidance into the autoencoder improved reconstruction peak signal-to-noise ratio by up to 6.7 decibels, with additional reference images providing progressive gains in fidelity and stability.
These results demonstrate that interactive, real-time portrait video streaming does not require computationally prohibitive foundation models. By shifting the generative burden from static appearance reconstruction to dynamic motion modeling, the framework significantly cuts computational costs and eliminates streaming latency bottlenecks. This makes real-time video deployment technically viable for consumer-facing interactive products on a single graphics processing unit.
Organizations developing interactive avatars should adopt reference-guided deep compression and causal blockwise autoregression as practical design patterns to achieve low-latency video streaming. When deploying the system, practitioners should utilize multiple reference images where available to maximize visual stability and reconstruction quality, while carefully balancing audio-visual guidance scales during generation.
Current limitations include the absence of fine-grained manual controls over specific attributes such as explicit eye gaze direction, as well as an inability to model hand or large-scale full-body gestures due to the portrait-focused training scope. There is strong confidence in the reported speed and fidelity for talking portrait generation within the evaluated operational domain, though further data collection across broader motion ranges is necessary before expanding beyond torso-level portraits.
- Paper: CV-VAE: A Compatible Video VAE for Latent Generative Video Models, Sijie Zhao et al. (2024). Reading CV-VAE provides the essential foundation for spatial-temporal video variational autoencoding and compression upon which the source builds its reference-guided causal VAE framework.
- Paper: StyleTalk: One-Shot Talking Head Generation with Controllable Speaking Styles, Yifeng Ma et al. (2023). Understanding StyleTalk introduces the core concepts and challenges of driving expressive, one-shot talking head generation from reference media and speech audio.
- Paper: Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis, Zhenhui Ye et al. (2024). Reading Real3D-Portrait establishes the baseline paradigms for single-image audio-driven portrait animation and motion adaptation.
- Paper: StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text, Roberto Henschel et al. (2025). This paper establishes the principles of autoregressive segment conditioning and chunked generation necessary to maintain consistency in continuous, streamable video diffusion.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). This work introduces the foundational two-stage paradigm of compressed latent autoencoding combined with autoregressive transformer modeling for video generation.
No sufficiently relevant recommendations were found.
