CV-VAE: A Compatible Video VAE for Latent Generative Video Models
Sijie ZhaoYong ZhangXiaodong CunShaoshu YangMuyao NiuXiaoyu LiWenbo HuYing Shan
Proposes a compatible 3D video VAE using latent space regularization that aligns with pretrained 2D image models like Stable Diffusion, enabling existing diffusion models to generate four times more frames with smooth spatio-temporal compression and minimal finetuning.
Generating high-quality, continuous video requires compressing visual data across both space and time to keep computational costs manageable. Most current video generation systems rely on two-dimensional image autoencoders adapted from text-to-image systems, such as Stable Diffusion. Because these systems only compress space, they handle time by sampling frames at fixed intervals, which discards motion dynamics and produces choppy, low-frame-rate videos. Meanwhile, independently training three-dimensional video autoencoders creates an incompatible representation space, requiring massive computing power and lengthy retraining to integrate with existing image and video foundations.
The article introduces and evaluates a compatible video variational autoencoder, termed CV-VAE, designed to achieve true spatial and temporal compression while remaining directly compatible with the latent representations of established text-to-image models. The main objective is to demonstrate that aligning representation spaces allows existing models to generate longer, smoother videos with minimal or no additional training.
To evaluate this approach, the authors developed a regularized training framework and an efficient network architecture. The architecture expands existing two-dimensional image components into three dimensions while keeping half of the internal layers in two dimensions to reduce complexity. The model was trained on large image and video datasets (including LAION-COCO, Unsplash, and WebVid-10M) using sixteen advanced graphical processing units. Performance was evaluated across standard image benchmarks (COCO2017) and video benchmarks (WebVid, UCF-101, and MSR-VTT) measuring reconstruction fidelity, generation quality, motion consistency, and computational efficiency.
The experimental findings demonstrate significant performance and efficiency gains. First, the autoencoder achieves a four-fold temporal compression ratio while maintaining reconstruction quality on par with or superior to existing image and video autoencoders. Second, integrating the architecture into pretrained models like Stable Diffusion allows direct image generation without fine-tuning, matching baseline visual quality. Third, when integrated into video generators such as Stable Video Diffusion and VideoCrafter2, the model expands output lengths by four times (for instance, converting 25-frame latents into 97-frame videos) with noticeably smoother motion after fine-tuning only a tiny fraction of model parameters. Fourth, ablations revealed that regularizing latent space alignments using the two-dimensional decoder alongside random frame mapping produced the highest reconstruction accuracy. Finally, the hybrid two-dimensional and three-dimensional architecture cut model parameters and computational complexity by approximately 30% relative to a full three-dimensional design without sacrificing visual quality.
These findings indicate that generative video systems can achieve superior frame rates and visual continuity without the prohibitive cost of training large diffusion backbones from scratch. By preserving compatibility with existing image models, organizations can upgrade video generation pipelines and frame-interpolation capabilities at substantially lower training budgets and shortened development timelines.
Teams developing video generation workflows should consider adopting compatible temporal autoencoders to expand video length and smoothness efficiently. For existing deployments, fine-tuning only the output layers of current models provides an immediate path to generating four times as many frames. Decision-makers should also account for the societal risks of high-fidelity synthetic media generation and establish appropriate usage safeguards.
The approach exhibits some limitations. Reconstruction fidelity remains bound by the channel dimensions of the underlying image autoencoder; using four latent channels constrains video fidelity compared to higher-dimensional representations. Furthermore, while the reported benchmarks show robust improvements across multiple open datasets, the evaluations omitted statistical significance bounds. Overall confidence in the technical feasibility of compatible spatio-temporal compression is high, though scaling to newer, higher-channel foundation models will require further validation.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It introduces Stable Video Diffusion, a primary baseline and foundation model whose latent representation and output length CV-VAE directly aims to extend and improve.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It details how 2D latent diffusion models are adapted to video by inserting temporal layers, establishing the latent framework and architectural conventions that CV-VAE optimizes.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It lays the foundational methodology for extending 2D image architectures to 3D spatio-temporal video diffusion models.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). It demonstrates the foundational paradigm of compressing video spatiotemporally into discrete latent codes via 3D autoencoders prior to generative modeling.
- Paper: Tutorial on Variational Autoencoders, Carl Doersch (2016). It provides the essential theoretical foundation and mathematical formulation of variational autoencoders upon which CV-VAE builds its compression framework.
- Paper: Towards Accurate Generative Models of Video: A New Metric & Challenges, Thomas Unterthiner et al. (2018). It introduces the Fréchet Video Distance (FVD) metric, which serves as a core benchmark evaluation tool throughout CV-VAE's experiments.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). It builds directly upon the paradigm of 3D variational autoencoders to achieve unified spatial and temporal compression within large-scale diffusion transformer architectures.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). It extends latent video diffusion transformers by designing spacetime motion representations for precise trajectory-guided video generation.
- Paper: Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution, Shangchen Zhou et al. (2024). It leverages temporal adaptation of pretrained latent decoders to maintain temporal consistency and visual fidelity in video upscaling pipelines.
- Paper: StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text, Roberto Henschel et al. (2025). It develops advanced memory and blending strategies to extend short-form latent video generation models into continuous, long-duration video streams.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). It exploits the temporal attention mechanisms of pretrained video models to transfer dynamic motion trajectories in a training-free manner.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). It advances temporal conditioning in video diffusion architectures by establishing per-frame noise assignments and flexible history guidance.
