Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

Sicheng XuYunhao DengShoukang HuYichuan WangYizhong ZhangZhanglu ChenJiaolong YangBaining Guo

article2026arXiv1 citations

Develops a real-time talking portrait video generator that combines a reference-guided causal VAE with an autoregressive Rectified Flow Transformer to produce high-fidelity streaming video directly from speech audio.

Listen

Generating realistic talking portrait videos from speech audio is critical for interactive applications such as digital assistants, education, and social companionship. However, existing high-fidelity solutions rely on heavy video diffusion foundation models that require several minutes to generate a few seconds of video. This high computational latency makes real-time, streaming deployment infeasible on standard hardware.

The article demonstrates a real-time, streamable video generation framework capable of synthesizing lifelike, arbitrary-length talking portrait videos from speech audio and one or more reference images.

To achieve this, the authors designed a causal video variational autoencoder (VAE) paired with a blockwise autoregressive transformer generator based on rectified flow matching. The framework achieves high computational efficiency by using reference image guidance inside the VAE decoder—allowing the system to focus on motion rather than static appearance—and extending causal residual auto-encoding across spatial and temporal dimensions. The generator uses key–value caching to sequentially synthesize compact video representations block-by-block. The system was trained on approximately 330 hours of video comprising about 10,000 unique identities and evaluated on two benchmark datasets across visual quality, audio synchronization, and inference latency.

The findings show that the proposed framework achieves an overall video compression ratio of 768, which is 10 to 15 times higher than the latent compression used in standard video diffusion models. The complete pipeline achieves a throughput of 42.3 frames per second at 512x512 resolution on a single modern graphics processing unit, running more than 25 times faster than leading diffusion-based talking portrait baselines. Despite its compact footprint, the model achieved the highest lip-sync accuracy and head-pose alignment across evaluated benchmarks while matching or outperforming foundation-scale baselines in visual quality metrics. Furthermore, incorporating reference image guidance into the autoencoder improved reconstruction peak signal-to-noise ratio by up to 6.7 decibels, with additional reference images providing progressive gains in fidelity and stability.

These results demonstrate that interactive, real-time portrait video streaming does not require computationally prohibitive foundation models. By shifting the generative burden from static appearance reconstruction to dynamic motion modeling, the framework significantly cuts computational costs and eliminates streaming latency bottlenecks. This makes real-time video deployment technically viable for consumer-facing interactive products on a single graphics processing unit.

Organizations developing interactive avatars should adopt reference-guided deep compression and causal blockwise autoregression as practical design patterns to achieve low-latency video streaming. When deploying the system, practitioners should utilize multiple reference images where available to maximize visual stability and reconstruction quality, while carefully balancing audio-visual guidance scales during generation.

Current limitations include the absence of fine-grained manual controls over specific attributes such as explicit eye gaze direction, as well as an inability to model hand or large-scale full-body gestures due to the portrait-focused training scope. There is strong confidence in the reported speed and fidelity for talking portrait generation within the evaluated operational domain, though further data collection across broader motion ranges is necessary before expanding beyond torso-level portraits.

No sufficiently relevant recommendations were found.

Cover for Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

Abstract

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a framework for streamable talking portrait video generation conditioned on speech audio and reference images. Designed meticulously for streaming scenarios, it features a causal video VAE for deep latent compression and an autoregressive latent denoising model. Our causal VAE integrates a variable number of reference images as guidance, allowing the network to focus on dynamic information rather than static appearance, thereby enhancing compression efficacy and reconstruction quality. Additionally, we extend the residual auto-encoding paradigm to improve spatial-temporal causality handling in our VAE. The generator is based on a Rectified Flow Transformer architecture and produces video latents in a blockwise auto-regressive manner. Our method enables the real-time generation of high-quality talking portrait videos, achieving speeds significantly faster than baseline models. Furthermore, comprehensive experiments demonstrate that it is on par with or even outperforms these large models in realism, vividness, and video quality.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Reference-Guided Deep Video Compression
  • 3.2 Blockwise AutoRegressive Latent Generation
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Talking Portrait Generation Results
  • 4.3 Comparison with Prior Methods
  • 4.3.1 Audio-Driven Talking Portrait Generation
  • 4.4 Ablation Study
  • 5 Conclusion
  • References
  • A More Results and Comparisons
  • B More Implementation Details
  • C Limitations
  • D Ethics Consideration

Knowls

  1. Knowl 1 — Reference-Guided Causal Video VAE Architecture and Deep Compression

    model/method

    The causal video Variational Autoencoder (VAE) compresses an input portrait video y∈R3×(T+1)×H×Wy \in \mathbb{R}^{3 \times (T+1) \times H \times W} into a compact latent tensor z∈RCz×Tz×Hz×Wzz \in \mathbb{R}^{C_z \times T_z \times H_z \times W_z}, where Cz=64C_z = 64, Tz=(T/4+1)T_z = (T/4 + 1), Hz=H/64H_z = H/64, and Wz=W/64W_z = W/64. This configuration achieves an overall video compression factor of 768.

    The autoencoder employs a two-stage symmetric hierarchical structure:

    1. Encoder (E1,E2E_1, E_2): E1E_1 applies spatial downsampling by a factor of 16 and temporal downsampling by a factor of 4 using causal 3D convolutions with RMSNorm layers, producing a 1024-dimensional feature representation. E2E_2 performs an additional 4×4\times spatial downsampling with two residual blocks to yield the latent zz.

    2. Reference-Guided Decoder (D1,Dref,D2D_1, D_{\text{ref}}, D_2): D1D_1 upsamples zz spatially to feature map fz∈RCf×Tz×Hf×Wff_z \in \mathbb{R}^{C_f \times T_z \times H_f \times W_f}. A set of MM reference portrait images r∈R3×M×H×Wr \in \mathbb{R}^{3 \times M \times H \times W} is passed through E1E_1 to extract mid-level appearance feature maps fc∈RCf×M×Hf×Wff_c \in \mathbb{R}^{C_f \times M \times H_f \times W_f}. The fusion network DrefD_{\text{ref}}, consisting of 6 Transformer blocks with 16 attention heads and head dimension 64, applies frame-wise self-attention to fzf_z (preserving temporal causality) followed by cross-attention between fzf_z (as queries) and fcf_c (as keys and values) to inject appearance details. The fused representation is decoded to RGB video frames by D2D_2 via spatial-temporal upsampling.

    The VAE is optimized with the reconstruction and perceptual loss: LVAE=Ey^[λ1∥y^−y∥1+λ2LPIPS(y^,y)]+λKLDKL(q(z∣y)∥p(z))\mathcal{L}_{\text{VAE}} = \mathbb{E}_{\hat{y}} \left[ \lambda_1 \|\hat{y} - y\|_1 + \lambda_2 \text{LPIPS}(\hat{y}, y) \right] + \lambda_{\text{KL}} D_{\text{KL}}(q(z|y) \parallel p(z)) where y^\hat{y} is the decoded reconstruction and yy is the ground-truth video.

  2. Knowl 2 — Causal Residual Video Auto-Encoding (CR-VA) Mechanism

    model/method

    Causal Residual Video Auto-Encoding (CR-VA) decouples spatial and temporal resolution transitions into separate sequential steps while enforcing strict temporal causality via a split-first strategy.

    Downsampling Operation

    Given input tensor x∈RB×C×T×H×Wx \in \mathbb{R}^{B \times C \times T \times H \times W}:

    1. Temporal Step (Split-First): The first frame x0=x[:,:,:1]x_0 = x[:, :, :1] is isolated. The remaining frames xt=x[:,:,1:]x_t = x[:, :, 1:] are processed along the temporal axis. In the parameter-free residual shortcut, PixelUnshuffle3d folds temporal slices into the channel dimension with factor (kt,1,1)(k_t, 1, 1), followed by channel averaging to maintain dimension CC. In the main branch, a causal 3D convolution CausalConv3d reduces channels before PixelUnshuffle3d. x0x_0 and the processed xtx_t are concatenated along the temporal dimension.
    2. Spatial Step: PixelUnshuffle3d folds spatial dimensions with factor (1,ks,ks)(1, k_s, k_s), followed by channel averaging to target dimension CoutC_{\text{out}} on the shortcut, and CausalConv3d on the main branch.

    Upsampling Operation

    1. Spatial Step: The shortcut applies channel replication (repeat_interleave) followed by PixelShuffle3d with factor (1,ks,ks)(1, k_s, k_s). The main branch uses CausalConv3d followed by PixelShuffle3d.
    2. Temporal Step (Split-First): The first frame x0x_0 is preserved directly. Slices xt=x[:,:,1:]x_t = x[:, :, 1:] are replicated along channels and expanded along time using PixelShuffle3d with factor (kt,1,1)(k_t, 1, 1), then concatenated with x0x_0.
  3. Knowl 3 — Blockwise Autoregressive Rectified Flow Transformer Generator

    model/method

    The video generator GG models the conditional velocity field of video latents zz given reference image latents zrz_r and audio features aa. The architecture is a 24-block Transformer with 12 attention heads and head dimension 128 (totaling 1B parameters), using 3D Rotary Positional Embeddings (3D RoPE) and Adaptive Layer Normalization (AdaLN) for timestep conditioning.

    Audio and Condition Alignment

    Audio features extracted by a pretrained audio encoder are temporally compressed by a factor of 4 via a trainable MLP embedder to match the temporal rate of the video latents Tz=T/4+1T_z = T/4 + 1. The compressed audio a′a' is broadcast spatially across Hz×WzH_z \times W_z and channel-wise concatenated with the noisy video latent ztz^t.

    Blockwise Causal Attention

    The generation target sequence is partitioned into non-overlapping temporal blocks of size k=4k=4 latent frames (equivalent to 16 video frames). Attention within each block is fully bidirectional, while inter-block attention is causally constrained to attend only to past blocks via a blockwise causal mask: Ai,j={1if ⌊i/k⌋≥⌊j/k⌋0otherwiseA_{i, j} = \begin{cases} 1 & \text{if } \lfloor i / k \rfloor \ge \lfloor j / k \rfloor \\ 0 & \text{otherwise} \end{cases} where ii and jj denote temporal latent token indices. During autoregressive streaming inference, past key-value (KV) states are cached and reused.

  4. Knowl 4 — Dual Condition Classifier-Free Guidance Formulation

    equation

    Classifier-free guidance (CFG) is applied simultaneously to the speech audio condition a′a' and reference image latents zrz_r during Rectified Flow sampling according to:

    Gcfg=λa(Gfull−Ga′=∅)+λr(Gfull−Gzr=∅)−Gfull+Ga′=∅+Gzr=∅G_{\text{cfg}} = \lambda_a \left( G_{\text{full}} - G_{a'=\emptyset} \right) + \lambda_r \left( G_{\text{full}} - G_{z_r=\emptyset} \right) - G_{\text{full}} + G_{a'=\emptyset} + G_{z_r=\emptyset}

    which can equivalently be expressed as:

    Gcfg=(1+λa+λr)Gfull−λaGa′=∅−λrGzr=∅G_{\text{cfg}} = (1 + \lambda_a + \lambda_r) G_{\text{full}} - \lambda_a G_{a'=\emptyset} - \lambda_r G_{z_r=\emptyset}

    where:

    • Gfull=G(zt,zr,a′,t)G_{\text{full}} = G(z^t, z_r, a', t) is the velocity prediction conditioned on both reference latents zrz_r and audio embeddings a′a'.
    • Ga′=∅=G(zt,zr,∅,t)G_{a'=\emptyset} = G(z^t, z_r, \emptyset, t) is the prediction with the audio embedding replaced by a learned null audio embedding token.
    • Gzr=∅=G(zt,∅,a′,t)G_{z_r=\emptyset} = G(z^t, \emptyset, a', t) is the prediction with the reference latents replaced by a learned null visual embedding token.
    • λa\lambda_a is the audio guidance scale (default value λa=2\lambda_a = 2).
    • λr\lambda_r is the reference guidance scale (default value λr=2\lambda_r = 2).
  5. Knowl 5 — Training Algorithm for Blockwise Autoregressive Latent Generator with Teacher Forcing

    algorithm

    The training procedure for the Rectified Flow Transformer generator GG uses teacher forcing with Gaussian noise context augmentation to bridge the train-inference exposure gap.

    Input: Video latents zz, reference latents zrz_r, audio features aa, generator parameters GG
    Output: Trained generator network GG
    for each training iteration do
        Sample a latent segment of N=32N = 32 frames: zwin⊂zz_{\text{win}} \subset z
        Sample M∈{1,2,3}M \in \{1, 2, 3\} reference frames: zref⊂zrz_{\text{ref}} \subset z_r
        if first frame of video is in zwinz_{\text{win}} then
            Add learnable mask token to the first frame
        end if
        
        Sample flow timestep t∼LogitNormal(0,1)t \sim \text{LogitNormal}(0, 1) and Gaussian noise ϵt∼N(0,I)\epsilon^t \sim \mathcal{N}(0, I)
        Compute linear interpolation: zt=t⋅ϵt+(1−t)⋅zwinz^t = t \cdot \epsilon^t + (1 - t) \cdot z_{\text{win}}
        Compute target flow velocity: vt=ϵt−zwinv^t = \epsilon^t - z_{\text{win}}
        
        Sample noise level for context t′∼Uniform[0,0.7]t' \sim \text{Uniform}[0, 0.7] and noise ϵt′∼N(0,I)\epsilon^{t'} \sim \mathcal{N}(0, I)
        Augment ground-truth context: zaug=t′⋅ϵt′+(1−t′)⋅zwinz_{\text{aug}} = t' \cdot \epsilon^{t'} + (1 - t') \cdot z_{\text{win}}
        
        Compute temporally aligned and spatially broadcast audio features a′a'
        With probability 0.1, randomly drop zrefz_{\text{ref}} or a′a' to null tokens
        
        Construct network input sequence xinput=[zref,zaug,zt⊕a′]x_{\text{input}} = [z_{\text{ref}}, z_{\text{aug}}, z^t \oplus a']
        Predict velocity: v^=G(xinput,t)\hat{v} = G(x_{\text{input}}, t) using causal blockwise self-attention
        Compute loss: L=∥vt−v^∥22\mathcal{L} = \|v^t - \hat{v}\|^2_2
        Update GG via gradient descent
    end for
  6. Knowl 6 — Streaming Talking Portrait Video Generation Algorithm

    algorithm

    The streaming inference procedure generates arbitrary-length talking portrait video frames in an online, blockwise autoregressive manner.

    Input: Reference frames IrI_r, Generator GG, Decoder DD, streaming audio AA
    Output: Continuous streaming video frames yiy_i
    Precompute visual features: fr=E1(Ir)f_r = E_1(I_r), zr=E(Ir)z_r = E(I_r)
    Initialize key-value cache with reference keys/values: KVlist=[GKV(zr)]\text{KV}_{\text{list}} = [G_{\text{KV}}(z_r)]
    Set block index i=0i = 0
    while streaming audio AA is active do
        Wait until the next chunk AiA_i (16 audio frames) is received
        Compute temporally aligned audio embedding ai′a'_i from AiA_i
        Initialize latent noise ziT∼N(0,I)z^{T}_i \sim \mathcal{N}(0, I) for current block of 4 latent frames
        if i==0i == 0 then
            Add learnable mask token to first frame
        end if
        
        for denoising step j=Tj = T down to 11 with timestep tjt_j do
            Predict velocity and compute block KV: vij,KVij=G(zitj,tj,ai′,KVlist)v_i^j, \text{KV}_i^j = G(z_i^{t_j}, t_j, a'_i, \text{KV}_{\text{list}})
            Update latent via ODE solver: zitj−1=SolveODE(zitj,vij,tj)z_i^{t_{j-1}} = \text{SolveODE}(z_i^{t_j}, v_i^j, t_j)
        end for
        
        Obtain clean latent block zi0=zit0z^0_i = z_i^{t_0}
        Append clean KV state KVi0\text{KV}_i^0 to KVlist\text{KV}_{\text{list}}
        
        if length of KVlist≥maximum KV length per window\text{KV}_{\text{list}} \ge \text{maximum KV length per window} then
            Reset KVlist=[GKV(zr)]\text{KV}_{\text{list}} = [G_{\text{KV}}(z_r)]
            Re-feed the last latent block zi0z^0_i through GG to seed KVlist\text{KV}_{\text{list}} for the next window
        end if
        
        Decode 16 video frames: yi=D(zi0,fr)y_i = D(z^0_i, f_r)
        Yield output frames yiy_i
        i=i+1i = i + 1
    end while
  7. Knowl 7 — Real-Time Streaming Latency and Throughput Analysis

    empirical result

    On a single NVIDIA H100 GPU at 512×512512 \times 512 video resolution, the full inference pipeline executes in real time without post-training distillation:

    • Audio Encoding and Denoising Generation (tGt_G): Extracting audio embeddings and running the 1B-parameter Rectified Flow Transformer generator GG for 12 denoising steps across a block of 4 latent frames (which maps to 16 video frames) takes tG=0.288 st_G = 0.288\text{ s} under the longest key-value cache context.
    • VAE Decoding (tDt_D): Upsampling and reconstructing the 16 video frames with reference feature guidance in decoder DD takes tD=0.090 st_D = 0.090\text{ s}.
    • Total Latency and Throughput: Total per-block latency is tG+tD=0.378 st_G + t_D = 0.378\text{ s}, yielding a streaming throughput of 16/0.378=42.3 FPS16 / 0.378 = 42.3\text{ FPS}. This is over 25×25\times faster than baseline diffusion-based portrait video generation models (e.g., Hallo3 at 0.27 FPS, FantasyTalking at 0.36 FPS, Hallo2 at 1.2 FPS, EchoMIMIC at 1.4 FPS, Sonic at 1.7 FPS).
  8. Knowl 8 — Quantitative Evaluation on HDTF and PortraitOneMin Benchmarks

    data/table

    The performance of the real-time streamable model was evaluated at 512×512512 \times 512 resolution against prior offline diffusion-based portrait generation methods on the HDTF benchmark (123 segments, 66 unseen identities) and the PortraitOneMin benchmark (32 one-minute clips, 16 unseen identities). Metrics include SyncNet confidence score (SC↑S_C \uparrow), SyncNet feature distance (SD↓S_D \downarrow), Cross-Attention Audio-Pose alignment (CAPP↑CAPP \uparrow), 25-frame Fréchet Video Distance (FVD25↓\text{FVD}_{25} \downarrow), and inference speed in frames per second (FPS ↑\uparrow) on an H100 GPU.

    HDTF PortraitOneMin Speed
    Method SC↑S_C \uparrow SD↓S_D \downarrow CAPP↑CAPP \uparrow FVD25↓\text{FVD}_{25} \downarrow SC↑S_C \uparrow SD↓S_D \downarrow CAPP↑CAPP \uparrow FVD25↓\text{FVD}_{25} \downarrow FPS ↑\uparrow
    EchoMIMIC 5.291 9.557 0.341 143.620 4.863 9.594 0.208 177.141 1.4
    EchoMIMIC-Distilled 5.513 9.350 0.348 174.061 5.531 9.187 0.201 201.899 13.3
    Hallo 7.457 7.841 0.242 90.880 6.721 8.275 0.210 179.072 1.2
    Hallo2 7.547 7.819 0.247 88.170 6.829 8.244 0.196 162.672 1.2
    Hallo3 7.256 8.596 0.337 76.430 6.836 8.724 0.301 175.340 0.27
    Sonic 8.799 6.602 0.689 43.920 8.185 7.031 0.598 95.047 1.7
    FantasyTalking 4.167 11.144 0.407 89.726 — — — — 0.36
    Ours (M=1M=1) 8.943 6.286 0.699 62.300 8.537 6.619 0.648 91.964 42.3
    Ours (M=2M=2) 9.056 6.175 0.739 49.400 8.438 6.688 0.677 81.517 42.3
    Ours (M=3M=3) 8.998 6.226 0.739 43.270 8.546 6.648 0.659 73.693 42.3

    In the single-reference setting (M=1M=1), the proposed model achieves superior lip-sync (SC=8.943,SD=6.286S_C=8.943, S_D=6.286) and pose-audio alignment (CAPP=0.699CAPP=0.699) compared to all baseline methods, while running at 42.3 FPS. Adding additional reference frames (M=2,3M=2, 3) steadily decreases video distortion (FVD25\text{FVD}_{25} drops from 62.300 to 43.270 on HDTF, and from 91.964 to 73.693 on PortraitOneMin).

  9. Knowl 9 — Ablation on Reference Guidance and CR-VA for Video VAE Reconstruction Fidelity

    data/table

    Ablation experiments evaluated the impact of varying the number of reference images M∈{0,1,2,3}M \in \{0, 1, 2, 3\} and incorporating Causal Residual Video Auto-encoding (CR-VA) on reconstruction quality for VoxCeleb2 (1K clips) and HDTF test sets at 512×512512 \times 512 resolution. ΔPSNR\Delta\text{PSNR} denotes the PSNR improvement relative to the unguided baseline (M=0M=0).

    VoxCeleb2 HDTF
    Configuration L1↓L_1 \downarrow PSNR↑\text{PSNR} \uparrow ΔPSNR\Delta\text{PSNR} LPIPS↓\text{LPIPS} \downarrow L1↓L_1 \downarrow PSNR↑\text{PSNR} \uparrow ΔPSNR\Delta\text{PSNR} LPIPS↓\text{LPIPS} \downarrow
    w/o CR-VA
    M=0M=0 (No ref.) 0.020 29.071 — 0.088 0.021 28.306 — 0.087
    M=1M=1 0.014 31.676 +2.605 0.051 0.013 32.068 +3.762 0.040
    M=2M=2 0.013 32.305 +3.234 0.043 0.012 32.663 +4.357 0.035
    M=3M=3 0.012 32.766 +3.695 0.039 0.012 33.149 +4.843 0.032
    w. CR-VA (Ours)
    M=0M=0 (No ref.) 0.018 29.604 — 0.087 0.020 28.678 — 0.086
    M=1M=1 0.013 32.354 +2.750 0.045 0.012 33.469 +4.791 0.032
    M=2M=2 0.012 33.281 +3.677 0.036 0.010 34.510 +5.832 0.027
    M=3M=3 0.011 33.979 +4.375 0.031 0.010 35.374 +6.696 0.023

    Key takeaways from the ablation data:

    1. Adding a single reference image (M=1M=1) improves reconstruction PSNR by +2.750 dB+2.750\text{ dB} on VoxCeleb2 and +4.791 dB+4.791\text{ dB} on HDTF when using CR-VA.
    2. CR-VA acts synergistically with reference guidance, widening the total PSNR gain from M=0M=0 to M=3M=3 to +4.375 dB+4.375\text{ dB} on VoxCeleb2 (vs. +3.695 dB+3.695\text{ dB} without CR-VA) and +6.696 dB+6.696\text{ dB} on HDTF (vs. +4.843 dB+4.843\text{ dB} without CR-VA).
  10. Knowl 10 — Limitations in Fine-Grained Attribute Control and Upper-Body Dynamics

    limitation

    The proposed talking portrait video generation framework has two primary limitations:

    1. Lack of Fine-Grained Attribute Controllability: The latent denoising generator does not accept explicit conditioning signals for fine-grained motion attributes, such as independent control over eye gaze direction, blink frequency, or precise 3D head pose angles.
    2. Absence of Hand and Large-Scale Body Movements: The generation capability is restricted to head, facial, and torso movements; it does not model hand gestures or full upper-body dynamics due to dataset constraints.

Coverage note — No substantial contributed material was omitted. All key models, equations, training/inference algorithms, benchmark tables, ablation studies, latency breakdowns, and stated limitations are fully represented.

References

  1. 1.Gpt-4o system card. 2024. 2, 8
  2. 2.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 2, 8
  3. 3.Matthew Brand. Voice puppetry. In Proceedings of the 26th annual conference on Computer graphics and interactive techniques, pages 21–28, 1999. 3
  4. 4.Christoph Bregler, Michele Covell, and Malcolm Slaney. Video rewrite: Driving visual speech with audio. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 715–722. 2023. 3
  5. 5.Boyuan Chen, Diego Martí Monsó, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion. Advances in Neural Information Processing Systems, 37:24081–24125, 2024. 3
  6. 6.Junyu Chen, Han Cai, Junsong Chen, Enze Xie, Shang Yang, Haotian Tang, Muyang Li, Yao Lu, and Song Han. Deep compression autoencoder for efficient high-resolution diffusion models. arXiv preprint arXiv:2410.10733, 2024. 2, 3, 4
  7. 7.Junyu Chen, Wenkun He, Yuchao Gu, Yuyang Zhao, Jincheng Yu, Junsong Chen, Dongyun Zou, Yujun Lin, Zhekai Zhang, Muyang Li, et al. Dc-videogen: Efficient video generation with deep compression video autoencoder. arXiv preprint arXiv:2509.25182, 2025. 3
  8. 8.Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu. Lip movements generation at a glance. In European Conference on Computer Vision, pages 520–535, 2018. 2, 3
  9. 9.Zhiyuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li, and Chenguang Ma. Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2403–2410, 2025. 3, 8
  10. 10.Joon Son Chung and Andrew Zisserman. Out of time: automated lip sync in the wild. In Asian Conference on Computer Vision Workshops, pages 251–263. Springer, 2017. 8
  11. 11.Joon Son Chung, Arsha Nagrani, and Andrew Zisserman. Voxceleb2: Deep speaker recognition. arXiv preprint arXiv:1806.05622, 2018. 7
  12. 12.Jiahao Cui, Hui Li, Yao Yao, Hao Zhu, Hanlin Shang, Kaihui Cheng, Hang Zhou, Siyu Zhu, and Jingdong Wang. Hallo2: Long-duration and high-resolution audio-driven portrait image animation. In The Thirteenth International Conference on Learning Representations. 8
  13. 13.Jiahao Cui, Hui Li, Yun Zhan, Hanlin Shang, Kaihui Cheng, Yuqi Ma, Shan Mu, Hang Zhou, Jingdong Wang, and Siyu Zhu. Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 21086–21095, 2025. 2, 3, 8
  14. 14.Nikita Drobyshev, Antoni Bigata Casademunt, Konstantinos Vougioukas, Zoe Landgraf, Stavros Petridis, and Maja Pantic. Emoportraits: Emotion-enhanced multimodal one-shot head avatars. arXiv preprint arXiv:2404.19110, 2024. 3
  15. 15.Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, 2024. 2, 7
  16. 16.Alisa Fortin, Guillaume Vernade, Kat Kampf, and Ammaar Reshi. Introducing gemini 2.5 flash image: Our state-of-the-art image model. https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/, 2025. Google Developer Blog. 1, 8
  17. 17.Jianzhu Guo, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang. Liveportrait: Efficient portrait animation with stitching and retargeting control. arXiv preprint arXiv:2407.03168, 2024. 2
  18. 18.Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. Space: Speech-driven portrait animation with controllable expression. In Proceedings of the ieee/cvf international conference on computer vision, pages 20914–20923, 2023. 3
  19. 19.Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, et al. Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103, 2024. 3
  20. 20.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. 13
  21. 21.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 3
  22. 22.Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train–test gap in autoregressive video diffusion. arXiv preprint arXiv:2506.08009, 2025. 3
  23. 23.Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu, Wayne Wu, Feng Xu, and Xun Cao. Eamm: One-shot emotional talking face via audio-based emotion-aware motion model. In ACM SIGGRAPH 2022 conference proceedings, pages 1–10, 2022. 3
  24. 24.Xiaozhong Ji, Xiaobin Hu, Zhihong Xu, Junwei Zhu, Chuming Lin, Qingdong He, Jiangning Zhang, Donghao Luo, Yi Chen, Qin Lin, Qinglin Lu, and Chengjie Wang. Sonic: Shifting focus to global audio perception in portrait animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 193–203, 2025. 2, 3, 8
  25. 25.Diederik P Kingma, Max Welling, et al. Auto-encoding variational bayes, 2013. 5
  26. 26.Ruineng Li, Daitao Xing, Huiming Sun, Yuanzhou Ha, Jinglin Shen, and Chiuman Ho. Tokenmotion: Decoupled motion control via token disentanglement for human-centric video generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 1951–1961, 2025. 3
  27. 27.Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, and Jingdong Wang. Expressive talking head generation with granular audio-visual control. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3387–3396, 2022. 3
  28. 28.Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng, and Chao Liang. Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models. arXiv preprint arXiv:2502.01061, 2025. 2, 3
  29. 29.Shanchuan Lin, Xin Xia, Yuxi Ren, Ceyuan Yang, Xuefeng Xiao, and Lu Jiang. Diffusion adversarial post-training for one-step video generation. arXiv preprint arXiv:2501.08316, 2025. 3
  30. 30.Shanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang, Yuxi Ren, Xin Xia, Yang Zhao, Xuefeng Xiao, and Lu Jiang. Autoregressive adversarial post-training for real-time interactive video generation. arXiv preprint arXiv:2506.09350, 2025. 3
  31. 31.Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. 2, 3, 5
  32. 32.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022. 3, 5
  33. 33.Chetwin Low and Weimin Wang. Talkingmachines: Real-time audio-driven facetime-style video via autoregressive diffusion models. arXiv preprint arXiv:2506.03099, 2025. 3
  34. 34.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023. 5
  35. 35.Xiangyu Peng, Zangwei Zheng, Chenhui Shen, Tom Young, Xinying Guo, Binluo Wang, Hang Xu, Hongxin Liu, Mingyan Jiang, Wenjun Li, et al. Open-sora 2.0: Training a commercial-level video generation model in $200 k. arXiv preprint arXiv:2503.09642, 2025. 3
  36. 36.KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In ACM International Conference on Multimedia, pages 484–492, 2020. 2, 3, 5, 15
  37. 37.Jinwei Qi, Chaonan Ji, Sheng Xu, Peng Zhang, Bang Zhang, and Liefeng Bo. Chatanyone: Stylized real-time portrait video generation with hierarchical motion diffusion model. arXiv preprint arXiv:2503.21144, 2025. 3
  38. 38.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 5, 8
  39. 39.Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high-resolution image synthesis with latent adversarial diffusion distillation. In SIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 3
  40. 40.Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial diffusion distillation. In European Conference on Computer Vision, pages 87–103. Springer, 2024. 3
  41. 41.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 3
  42. 42.Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, and Maja Pantic. Diffused heads: Diffusion models beat gans on talking-face generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5091–5100, 2024. 2
  43. 43.Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics, 36 (4):1–13, 2017. 3
  44. 44.Hansi Teng, Hongyu Jia, Lei Sun, Lingzhi Li, Maolin Li, Mingqiu Tang, Shuai Han, Tianning Zhang, WQ Zhang, Weifeng Luo, et al. Magi-1: Autoregressive video generation at scale. arXiv preprint arXiv:2505.13211, 2025. 3
  45. 45.Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions. In European Conference on Computer Vision, pages 244–260. Springer, 2024. 2, 3
  46. 46.Rui Tian, Qi Dai, Jianmin Bao, Kai Qiu, Yifan Yang, Chong Luo, Zuxuan Wu, and Yu-Gang Jiang. Reducio! generating 1k video within 16 seconds using extremely compressed motion latents. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19237–19247, 2025. 3
  47. 47.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. 2019. 8
  48. 48.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Conditional image generation with pixelcnn decoders. Advances in neural information processing systems, 29, 2016. 3
  49. 49.Aäron Van Den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In International conference on machine learning, pages 1747–1756. PMLR, 2016. 3
  50. 50.Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, Pingyu Wu, Ruihang Chu, Ruili Feng, Shiwei Zhang, Siyang Sun, Tao Fang, Tianxing Wang, Tianyi Gui, Tingyu Weng, Tong Shen, Wei Lin, Wei Wang, Wei Wang, Wenmeng Zhou, Wente Wang, Wenting Shen, Wenyuan Yu, Xianzhong Shi, Xiaoming Huang, Xin Xu, Yan Kou, Yangyu Lv, Yifei Li, Yijing Liu, Yiming Wang, Yingya Zhang, Yitong Huang, Yong Li, You Wu, Yu Liu, Yulin Pan, Yun Zheng, Yuntao Hong, Yupeng Shi, Yutong Feng, Zeyinzi Jiang, Zhen Han, Zhi-Fan Wu, and Ziyu Liu. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. 3, 4, 5, 7, 8
  51. 51.Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum, and Baoyuan Wang. Progressive disentangled representation learning for fine-grained controllable talking head synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17979–17989, 2023. 2, 3
  52. 52.Jiahao Wang, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, and Jieping Ye. Echoshot: Multi-shot portrait video generation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. 3
  53. 53.Mengchao Wang, Qiang Wang, Fan Jiang, Yaqi Fan, Yunpeng Zhang, Yonggang Qi, Kun Zhao, and Mu Xu. Fantasytalking: Realistic talking portrait generation via coherent motion synthesis. In Proceedings of the 33rd ACM International Conference on Multimedia, pages 9891–9900, 2025. 2, 3, 8
  54. 54.Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. Audio2head: Audio-driven one-shot talking-head generation with natural head motion. In International Joint Conference on Artificial Intelligence, 2021. 2, 3
  55. 55.Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10039–10049, 2021. 2
  56. 56.Huawei Wei, Zejun Yang, and Zhisheng Wang. Aniportrait: Audio-driven synthesis of photorealistic portrait animation. arXiv preprint arXiv:2403.17694, 2024. 2, 3
  57. 57.Ronald J Williams and David Zipser. A learning algorithm for continually running fully recurrent neural networks. Neural computation, 1(2):270–280, 1989. 5
  58. 58.Mingwang Xu, Hui Li, Qingkun Su, Hanlin Shang, Liwei Zhang, Ce Liu, Jingdong Wang, Yao Yao, and Siyu Zhu. Hallo: Hierarchical audio-driven visual synthesis for portrait image animation. arXiv preprint arXiv:2406.08801, 2024. 2, 8
  59. 59.Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Baining Guo. Vasa-1: Lifelike audio-driven talking faces generated in real time. Advances in Neural Information Processing Systems, 37:660–684, 2024. 2, 3, 8, 13
  60. 60.Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021. 3
  61. 61.Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 3, 5, 8
  62. 62.Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang, Weichuang Li, Jiawei Huang, Ziyue Jiang, Jinzheng He, Rongjie Huang, Jinglin Liu, et al. Real3d-portrait: One-shot realistic 3d talking portrait synthesis. arXiv preprint arXiv:2401.08503, 2024. 3
  63. 63.Fei Yin, Yong Zhang, Xiaodong Cun, Mingdeng Cao, Yanbo Fan, Xuan Wang, Qingyan Bai, Baoyuan Wu, Jue Wang, and Yujiu Yang. Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan. In European Conference on Computer Vision, pages 85–101, 2022. 3
  64. 64.Tianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang, Eli Shechtman, Fredo Durand, and William T Freeman. Improved distribution matching distillation for fast image synthesis. In NeurIPS, 2024. 2, 3
  65. 65.Tianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman, Fredo Durand, William T Freeman, and Taesung Park. One-step diffusion with distribution matching distillation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6613–6623, 2024. 3
  66. 66.Tianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman, Frédo Durand, William T Freeman, and Taesung Park. One-step diffusion with distribution matching distillation. In CVPR, 2024. 2
  67. 67.Tianwei Yin, Qiang Zhang, Richard Zhang, William T Freeman, Fredo Durand, Eli Shechtman, and Xun Huang. From slow bidirectional to fast autoregressive video diffusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 22963–22974, 2025. 3
  68. 68.Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. Advances in Neural Information Processing Systems, 37:128940–128966, 2024. 3
  69. 69.Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, and Anima Anandkumar. Efficient video diffusion models via content-frame motion-latent decomposition. arXiv preprint arXiv:2403.14148, 2024. 3
  70. 70.Zhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang, Finn Wong, and Baoyuan Wang. Talking head generation with probabilistic audio-to-visual diffusion priors. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7645–7655, 2023. 3
  71. 71.Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyang Ge, Yujun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. Identity-preserving text-to-video generation by frequency decomposition. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 12978–12988, 2025. 3
  72. 72.Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2019. 4
  73. 73.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 5
  74. 74.Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang, Xi Shen, Yu Guo, Ying Shan, and Fei Wang. Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8652–8661, 2023. 2, 3, 13
  75. 75.Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3661–3670, 2021. 7
  76. 76.Deyu Zhou, Quan Sun, Yuang Peng, Kun Yan, Runpei Dong, Duomin Wang, Zheng Ge, Nan Duan, and Xiangyu Zhang. Taming teacher forcing for masked autoregressive video generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 7374–7384, 2025. 2
  77. 77.Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE/CVF Conference on computer Vision and Pattern Recognition, pages 4176–4186, 2021. 3
  78. 78.Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. Makelttalk: speaker-aware talking-head animation. ACM Transactions On Graphics (TOG), 39(6):1–15, 2020. 3

Citation

MLA
Xu, S., et al. “Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs”. arXiv, 2026, http://arxiv.org/abs/2606.01620v1.
APA
Xu, S., Deng, Y., Hu, S., Wang, Y., Zhang, Y., Chen, Z., Yang, J., & Guo, B. (2026). Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs. arXiv. http://arxiv.org/abs/2606.01620v1
Chicago
Xu, S., Y. Deng, S. Hu, et al. 2026. “Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs”. arXiv. http://arxiv.org/abs/2606.01620v1.
Harvard
Xu, S. et al. (2026) “Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.01620v1.
Vancouver
1. Xu S, Deng Y, Hu S, Wang Y, Zhang Y, Chen Z, Yang J, Guo B (2026) Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs. arXiv

BibTeX

@article{xu2026real,
  title = {Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs},
  author = {Xu, Sicheng and Deng, Yu and Hu, Shoukang and Wang, Yichuan and Zhang, Yizhong and Chen, Zhan and Yang, Jiaolong and Guo, Baining},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.01620v1},
  eprint = {2606.01620}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/