CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Sijie ZhaoYong ZhangXiaodong CunShaoshu YangMuyao NiuXiaoyu LiWenbo HuYing Shan

article2024NeurIPS80 citations

Proposes a compatible 3D video VAE using latent space regularization that aligns with pretrained 2D image models like Stable Diffusion, enabling existing diffusion models to generate four times more frames with smooth spatio-temporal compression and minimal finetuning.

Listen

Generating high-quality, continuous video requires compressing visual data across both space and time to keep computational costs manageable. Most current video generation systems rely on two-dimensional image autoencoders adapted from text-to-image systems, such as Stable Diffusion. Because these systems only compress space, they handle time by sampling frames at fixed intervals, which discards motion dynamics and produces choppy, low-frame-rate videos. Meanwhile, independently training three-dimensional video autoencoders creates an incompatible representation space, requiring massive computing power and lengthy retraining to integrate with existing image and video foundations.

The article introduces and evaluates a compatible video variational autoencoder, termed CV-VAE, designed to achieve true spatial and temporal compression while remaining directly compatible with the latent representations of established text-to-image models. The main objective is to demonstrate that aligning representation spaces allows existing models to generate longer, smoother videos with minimal or no additional training.

To evaluate this approach, the authors developed a regularized training framework and an efficient network architecture. The architecture expands existing two-dimensional image components into three dimensions while keeping half of the internal layers in two dimensions to reduce complexity. The model was trained on large image and video datasets (including LAION-COCO, Unsplash, and WebVid-10M) using sixteen advanced graphical processing units. Performance was evaluated across standard image benchmarks (COCO2017) and video benchmarks (WebVid, UCF-101, and MSR-VTT) measuring reconstruction fidelity, generation quality, motion consistency, and computational efficiency.

The experimental findings demonstrate significant performance and efficiency gains. First, the autoencoder achieves a four-fold temporal compression ratio while maintaining reconstruction quality on par with or superior to existing image and video autoencoders. Second, integrating the architecture into pretrained models like Stable Diffusion allows direct image generation without fine-tuning, matching baseline visual quality. Third, when integrated into video generators such as Stable Video Diffusion and VideoCrafter2, the model expands output lengths by four times (for instance, converting 25-frame latents into 97-frame videos) with noticeably smoother motion after fine-tuning only a tiny fraction of model parameters. Fourth, ablations revealed that regularizing latent space alignments using the two-dimensional decoder alongside random frame mapping produced the highest reconstruction accuracy. Finally, the hybrid two-dimensional and three-dimensional architecture cut model parameters and computational complexity by approximately 30% relative to a full three-dimensional design without sacrificing visual quality.

These findings indicate that generative video systems can achieve superior frame rates and visual continuity without the prohibitive cost of training large diffusion backbones from scratch. By preserving compatibility with existing image models, organizations can upgrade video generation pipelines and frame-interpolation capabilities at substantially lower training budgets and shortened development timelines.

Teams developing video generation workflows should consider adopting compatible temporal autoencoders to expand video length and smoothness efficiently. For existing deployments, fine-tuning only the output layers of current models provides an immediate path to generating four times as many frames. Decision-makers should also account for the societal risks of high-fidelity synthetic media generation and establish appropriate usage safeguards.

The approach exhibits some limitations. Reconstruction fidelity remains bound by the channel dimensions of the underlying image autoencoder; using four latent channels constrains video fidelity compared to higher-dimensional representations. Furthermore, while the reported benchmarks show robust improvements across multiple open datasets, the evaluations omitted statistical significance bounds. Overall confidence in the technical feasibility of compatible spatio-temporal compression is high, though scaling to newer, higher-channel foundation models will require further validation.

Cover for CV-VAE: A Compatible Video VAE for Latent Generative Video Models

Abstract

Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI’s SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e.g., image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. To improve the training efficiency, we also design a novel architecture for the video VAE. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE.

Citation

MLA
Zhao, S., et al. “CV-VAE: A Compatible Video VAE for Latent Generative Video Models”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 12847–71, https://proceedings.neurips.cc/paper_files/paper/2024/file/1787533e171dcc8549cc2eb5a4840eec-Paper-Conference.pdf.
APA
Zhao, S., Zhang, Y., Cun, X., Yang, S., Niu, M., Li, X., Hu, W., & Shan, Y. (2024). CV-VAE: A Compatible Video VAE for Latent Generative Video Models. Advances in Neural Information Processing Systems, 37, 12847–12871. https://proceedings.neurips.cc/paper_files/paper/2024/file/1787533e171dcc8549cc2eb5a4840eec-Paper-Conference.pdf
Chicago
Zhao, S., Y. Zhang, X. Cun, et al. 2024. “CV-VAE: A Compatible Video VAE for Latent Generative Video Models”. Advances in Neural Information Processing Systems 37: 12847–71. https://proceedings.neurips.cc/paper_files/paper/2024/file/1787533e171dcc8549cc2eb5a4840eec-Paper-Conference.pdf.
Harvard
Zhao, S. et al. (2024) “CV-VAE: A Compatible Video VAE for Latent Generative Video Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 12847–12871. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/1787533e171dcc8549cc2eb5a4840eec-Paper-Conference.pdf.
Vancouver
1. Zhao S, Zhang Y, Cun X, Yang S, Niu M, Li X, Hu W, Shan Y (2024) CV-VAE: A Compatible Video VAE for Latent Generative Video Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 12847–12871

BibTeX

@inproceedings{zhao2024vae,
  title = {CV-VAE: A Compatible Video VAE for Latent Generative Video Models},
  author = {Zhao, Sijie and Zhang, Yong and Cun, Xiaodong and Yang, Shaoshu and Niu, Muyao and Li, Xiaoyu and Hu, Wenbo and Shan, Ying},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {12847-12871},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/1787533e171dcc8549cc2eb5a4840eec-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors