Hierarchical Spatio-temporal Decoupling for Text-to- Video Generation

Zhiwu QingShiwei ZhangJiayu WangXiang WangYujie WeiYingya ZhangChangxin GaoNong Sang

article2024CVPR62 citations

Proposes HiGen, a diffusion-based framework that decouples spatial reasoning from temporal dynamics at both the structural and content levels to generate realistic, temporally stable videos from text prompts.

Listen

Generating realistic and dynamic videos from text prompts holds substantial value for digital media, gaming, and entertainment. However, current automated systems struggle to produce video clips that simultaneously offer sharp image quality and lively, coherent movement. Existing approaches generally try to generate spatial imagery and temporal dynamics all at once, which introduces excessive mathematical complexity. Consequently, current models tend to generate either high-quality images with almost no movement or dynamic scenes with severe visual degradation and instability.

The article introduces and evaluates "HiGen," an artificial intelligence method designed to resolve this trade-off by decoupling video generation into separate spatial and temporal steps across two distinct levels: structural architecture and content guidance.

The researchers developed a hierarchical framework utilizing an established image synthesis model as its base. Structurally, the system performs spatial reasoning first to establish a high-quality visual anchor, followed by temporal reasoning to generate movement across frames. At the content level, the method calculates explicit motion cues from adjacent pixel differences and appearance cues using a semantic vision model, allowing independent control over movement speed and scene variations. The system was trained using an image-video joint strategy on 17 million video-text pairs and roughly 60 million image-text pairs, and its performance was evaluated against benchmark datasets and existing tools.

The evaluation revealed several key findings. First, the hierarchical decoupling strategy significantly outperformed existing models on benchmark tests, achieving a substantial reduction in video distortion metrics (a Fréchet Video Distance score of 406 compared to 550 for the baseline). Second, human evaluations showed strong preference for the new approach, most notably in temporal movement quality, where it scored 74.0%—surpassing the nearest open-source alternative by 18.8 percentage points. Third, structural decoupling without content-level guidance severely degraded frame-to-frame consistency, confirming that both decoupling layers are necessary for stable video output. Finally, explicit motion and appearance factors proved highly effective for manual control during video generation, whereas standard playback frame-rate adjustments had minimal impact on temporal dynamics.

These findings demonstrate that separating visual content from time-based motion reduces computational complexity and makes automated video synthesis more practical and controllable. By enabling users to independently tune motion speed and appearance shifts, this approach provides a viable foundation for reliable commercial video production tools. Importantly, the analysis shows that prioritizing extreme frame-to-frame consistency often results in static, unengaging videos, indicating that controlled variability is essential for natural motion.

Organizations developing or deploying automated video generation systems should adopt decoupled architectures to improve output quality and operational control. Future development should focus on addressing remaining technical boundaries. Specifically, the system's ability to render fine object details still lags behind dedicated single-image generators, and synthesizing complex human or animal movements that adhere strictly to common-sense physical actions remains challenging during large motions. Addressing these issues will require further research into training data curation and specialized model architectures.

Cover for Hierarchical Spatio-temporal Decoupling for Text-to- Video Generation

Abstract

Despite diffusion models having shown powerful abilities to generate photorealistic images, generating videos that are realistic and diverse still remains in its infancy. One of the key reasons is that current methods intertwine spatial content and temporal dynamics together, leading to a notably increased complexity of text-to-video generation (T2V). In this work, we propose HiGen, a diffusion model-based method that improves performance by decoupling the spatial and temporal factors of videos from two perspectives, i.e., structure level and content level. At the structure level, we decompose the T2V task into two steps, including spatial reasoning and temporal reasoning, using a unified denoiser. Specifically, we generate spatially coherent priors using text during spatial reasoning and then generate temporally coherent motions from these priors during temporal reasoning. At the content level, we extract two subtle cues from the content of the input video that can express motion and appearance changes, respectively. These two cues then guide the model's training for generating videos, enabling flexible content variations and enhancing temporal stability. Through the decoupled paradigm, HiGen can effectively reduce the complexity of this task and generate realistic videos with semantics accuracy and motion stability. Extensive experiments demonstrate the superior performance of HiGen over the state-of-the-art T2V methods. We have released our source code and models.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Approach
  • 3.1. Preliminaries
  • 3.2. Structure-level Decoupling
  • 3.3. Content-level Decoupling
  • 3.4. Training and Inference
  • 4. Experiments
  • 4.1. Implementation Details
  • 4.2. Ablation Studies
  • 4.3. Comparison with State-of-the-art
  • 5. Discussions
  • References

Knowls

  1. Knowl 1 — Hierarchical Spatio-Temporal Decoupling Architecture in HiGen

    model/method

    HiGen is a diffusion-based text-to-video (T2V) generation method that decomposes the complex distribution of video data by decoupling spatial content and temporal dynamics at two hierarchical levels:

    1. Structure-level decoupling: Decomposes video synthesis into spatial reasoning and temporal reasoning using a unified 3D-UNet denoiser. Spatial reasoning utilizes a text-to-image (T2I) prior to synthesize a static, semantically coherent spatial latent prior. Temporal reasoning then conditions on this prior to synthesize temporally coherent inter-frame dynamics.
    2. Content-level decoupling: Decomposes video content into two independent control signals extracted directly from the video data: a frame-difference motion factor that measures displacement velocity, and a semantic-similarity appearance factor that measures global appearance change across time.

    Through this hierarchical decoupling, the model reduces the joint spatio-temporal optimization complexity, avoiding the trade-off between static high-quality frames and dynamic low-quality videos.

  2. Knowl 2 — Structure-Level Decoupling via Shared-Weight 3D-UNet

    model/method

    Structure-level decoupling executes video generation in two stages within a single 3D-UNet architecture:

    • Spatial Reasoning: The temporal convolution and temporal self-attention layers of the 3D-UNet are disabled. The network uses only its spatial layers (pre-trained on text-to-image generation) to perform TT denoising steps conditioned on the input text prompt yy. This produces a clean spatial latent prior z0s∈RCin×H×W\mathbf{z}_0^\text{s} \in \mathbb{R}^{C_\text{in} \times H \times W}. The prior remains in the latent space and is not decoded to RGB pixels by the VAE decoder D\mathcal{D}.
    • Temporal Reasoning: The spatial prior z0s\mathbf{z}_0^\text{s} is passed through a zero-initialized convolutional stem ConvStemt(⋅)\text{ConvStem}_t(\cdot), which has an identical architecture to the spatial stem ConvStems(⋅)\text{ConvStem}_s(\cdot). The output feature map is repeated FF times along the temporal axis and added directly to the noisy video latent representation zt∈RF×Cin×H×W\mathbf{z}_t \in \mathbb{R}^{F \times C_\text{in} \times H \times W}.
    • Weight Sharing: Both stages share the exact same spatial layers, transferring pre-trained spatial generative knowledge to temporal reasoning while allowing the temporal layers to focus exclusively on inter-frame transitions.
  3. Knowl 3 — Motion Guidance Extraction and Conditioning

    model/method

    To quantify frame-to-frame motion magnitude independently of video frame rate (FPS), HiGen calculates adjacent frame differences in latent space. For a latent video sequence z0=[z01,…,z0F]\mathbf{z}_0 = [\mathbf{z}_0^1, \dots, \mathbf{z}_0^F] of FF frames, the motion factor γfm\gamma_f^\text{m} between consecutive frames ff and f+1f+1 is defined as:

    γfm=∥z0f−z0f+1∥2\gamma_f^\text{m} = \|\mathbf{z}_0^f - \mathbf{z}_0^{f+1}\|_2

    For an FF-frame sequence, this yields an (F−1)(F-1)-dimensional motion vector r~m=[γ1m,…,γF−1m]∈RF−1\tilde{\mathbf{r}}^\text{m} = [\gamma_1^\text{m}, \dots, \gamma_{F-1}^\text{m}] \in \mathbb{R}^{F-1}.

    To condition the 3D-UNet on motion:

    1. Each scalar γfm\gamma_f^\text{m} is rounded to the nearest integer.
    2. The vector is mapped into a CC-dimensional embedding using sinusoidal positional encoding Sin(⋅)\text{Sin}(\cdot) and a zero-initialized Multi-Layer Perceptron (MLP\text{MLP}).
    3. A linear interpolation function Interpolate(⋅)\text{Interpolate}(\cdot) resamples the sequence from length F−1F-1 to FF, yielding the motion guidance tensor rm\mathbf{r}^\text{m}:

    rm=Interpolate(MLP(Sin(Round(r~m))))∈RF×C\mathbf{r}^\text{m} = \text{Interpolate}(\text{MLP}(\text{Sin}(\text{Round}(\tilde{\mathbf{r}}^\text{m})))) \in \mathbb{R}^{F \times C}

    1. rm\mathbf{r}^\text{m} is added to the diffusion time-step embedding vector and integrated into each residual block of the 3D-UNet.
  4. Knowl 4 — Appearance Guidance Extraction and Semantic Variation Matrix

    model/method

    To capture semantic and appearance variations across frames without coupling to pixel-level motion, HiGen constructs an appearance variation matrix using a pre-trained visual semantic backbone Ω(⋅)\Omega(\cdot) (specifically DINO ViT):

    1. For an RGB video x0=[x01,…,x0F]\mathbf{x}_0 = [\mathbf{x}_0^1, \dots, \mathbf{x}_0^F], normalized semantic features are extracted:

    g=Norm(Ω(x0))∈RF×D\mathbf{g} = \text{Norm}(\Omega(\mathbf{x}_0)) \in \mathbb{R}^{F \times D}

    1. The full pairwise inter-frame cosine similarity matrix r~a∈RF×F\tilde{\mathbf{r}}^\text{a} \in \mathbb{R}^{F \times F} is computed:

    r~a=g⊗T(g)\tilde{\mathbf{r}}^\text{a} = \mathbf{g} \otimes \mathcal{T}(\mathbf{g})

    where ⊗\otimes denotes matrix multiplication and T(⋅)\mathcal{T}(\cdot) denotes the transpose operation.

    1. The appearance guidance ra\mathbf{r}^\text{a} is obtained using a zero-initialized Multi-Layer Perceptron:

    ra=MLP(r~a)∈RF×C\mathbf{r}^\text{a} = \text{MLP}(\tilde{\mathbf{r}}^\text{a}) \in \mathbb{R}^{F \times C}

    and injected into each residual block alongside the diffusion time-step embedding.

    1. The global appearance factor γa\gamma^\text{a} across the entire clip is defined by:

    γa=1−r~0,F−1a\gamma^\text{a} = 1 - \tilde{\mathbf{r}}^\text{a}_{0, F-1}

    where r~0,F−1a\tilde{\mathbf{r}}^\text{a}_{0, F-1} represents the cosine similarity between the first frame (f=0f=0) and the last frame (f=F−1f=F-1). Higher γa\gamma^\text{a} values indicate larger semantic transitions over the sequence.

  5. Knowl 5 — Controllable HiGen Video Inference Algorithm

    algorithm

    During inference, HiGen accepts user-specified scalars for the motion factor γm\gamma^\text{m} and appearance factor γa\gamma^\text{a} to dynamically control scene dynamics and visual variation:

    Input: Text prompt yy, diffusion steps TT, frame count FF, target motion factor γm∈[300,600]\gamma^\text{m} \in [300, 600], target appearance factor γa∈[0,1.0]\gamma^\text{a} \in [0, 1.0]
    Output: Synthesized video latent sequence z0∈RF×Cin×H×W\mathbf{z}_0 \in \mathbb{R}^{F \times C_\text{in} \times H \times W}
    # Stage 1: Spatial Reasoning
    z0s←SampleSpatialPrior(y,T)\mathbf{z}_0^\text{s} \leftarrow \text{SampleSpatialPrior}(y, T) # Denoise using spatial layers only
    # Stage 2: Guidance Construction
    r~m←[γm,γm,…,γm]∈RF−1\tilde{\mathbf{r}}^\text{m} \leftarrow [\gamma^\text{m}, \gamma^\text{m}, \dots, \gamma^\text{m}] \in \mathbb{R}^{F-1}
    rm←Interpolate(MLP(Sin(Round(r~m))))∈RF×C\mathbf{r}^\text{m} \leftarrow \text{Interpolate}(\text{MLP}(\text{Sin}(\text{Round}(\tilde{\mathbf{r}}^\text{m})))) \in \mathbb{R}^{F \times C}
    k←−γaF−1k \leftarrow -\frac{\gamma^\text{a}}{F-1}
    for i=0i = 0 to F−1F-1 do
        for j=0j = 0 to F−1F-1 do
            r~i,ja←∣i−j∣⋅k+1\tilde{\mathbf{r}}^\text{a}_{i, j} \leftarrow |i - j| \cdot k + 1
        end for
    end for
    ra←MLP(r~a)∈RF×C\mathbf{r}^\text{a} \leftarrow \text{MLP}(\tilde{\mathbf{r}}^\text{a}) \in \mathbb{R}^{F \times C}
    # Stage 3: Temporal Reasoning
    zT∼N(0,I)∈RF×Cin×H×W\mathbf{z}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I}) \in \mathbb{R}^{F \times C_\text{in} \times H \times W}
    zprior←RepeatF(ConvStemt(z0s))\mathbf{z}_\text{prior} \leftarrow \text{Repeat}_F(\text{ConvStem}_t(\mathbf{z}_0^\text{s}))
    for t=Tt = T down to 1 do
        zt′←zt+zprior\mathbf{z}_t^\prime \leftarrow \mathbf{z}_t + \mathbf{z}_\text{prior}
        ϵθ←UNet3D(zt′,t,y,rm,ra)\mathbf{\epsilon}_\theta \leftarrow \text{UNet3D}(\mathbf{z}_t^\prime, t, y, \mathbf{r}^\text{m}, \mathbf{r}^\text{a})
        zt−1←DenoiseStep(zt,ϵθ,t)\mathbf{z}_{t-1} \leftarrow \text{DenoiseStep}(\mathbf{z}_t, \mathbf{\epsilon}_\theta, t)
    end for
    return z0\mathbf{z}_0
  6. Knowl 6 — Image-Video Joint Training Protocol for HiGen

    experimental setup

    HiGen fine-tunes a pre-trained Stable Diffusion model using an asynchronous image-video joint training strategy on 8 NVIDIA A100 GPUs:

    • GPU Partitioning: Two GPUs (25%) are dedicated to image fine-tuning (spatial reasoning) with a total batch size of 512 images. Six GPUs (75%) are dedicated to video fine-tuning (temporal reasoning) with a total batch size of 72 video clips of length F=32F = 32 frames at 448×256448 \times 256 resolution.
    • Optimization: The network is trained for 25,000 iterations using the AdamW optimizer with a learning rate of 5×10−55 \times 10^{-5} and weight decay of 0.
    • Gradient Decay: For the spatial parameters on image GPUs, a gradient decay factor of 0.2 is applied to prevent catastrophic forgetting of pre-trained image generative knowledge. On video GPUs, all network parameters (spatial and temporal) are updated.
    • Training Prior Approximation: During video fine-tuning, the spatial prior z0s\mathbf{z}_0^\text{s} is set to the latent embedding of the input video's middle frame (f=⌊F/2⌋f = \lfloor F/2 \rfloor) for computational efficiency.
  7. Knowl 7 — Quantitative Text-to-Video Evaluation on MSR-VTT

    data/table

    Performance comparison on the MSR-VTT dataset measuring spatial frame quality via Fréchet Inception Distance (FID), temporal fidelity via Fréchet Video Distance (FVD), and text-video semantic alignment via CLIP Similarity (CLIPSIM):

    Method FID ↓\downarrow FVD ↓\downarrow CLIPSIM ↑\uparrow
    CogVideo (English) 23.59 1294 0.2631
    Latent-Shift 15.23 - 0.2773
    Make-A-Video 13.17 - 0.3049
    Video LDM - - 0.2929
    MagicVideo - 998 -
    VideoComposer 10.77 580 0.2932
    ModelScopeT2V 11.09 550 0.2930
    PYoCo 9.73 - -
    HiGen 8.60 406 0.2947

    HiGen achieves the lowest FID (8.60) and lowest FVD (406) among all compared methods, reducing FVD by 144 points (26.2%) relative to ModelScopeT2V while maintaining competitive CLIP similarity.

  8. Knowl 8 — Ablation of Structure-Level and Content-Level Decoupling Components

    data/table

    Ablation study evaluating the isolated and joint contributions of Structure-Level (SL) and Content-Level (CL) decoupling on Temporal Consistency (average inter-frame CLIP cosine similarity) and CLIPSIM:

    Configuration SL CL Temporal Consistency ↑\uparrow CLIPSIM ↑\uparrow
    ModelScopeT2V 0.931 0.292
    w/ SL, w/o CL ✓ 0.889 0.313
    HiGen (w/ SL, w/ CL) ✓ ✓ 0.944 0.318

    Adding structure-level decoupling alone improves spatial alignment (CLIPSIM rises from 0.292 to 0.313) but reduces inter-frame temporal consistency (drops from 0.931 to 0.889). Adding content-level decoupling recovers temporal stability (0.944) while further improving semantic fidelity (0.318).

  9. Knowl 9 — Semantic Model Selection for Decoupled Appearance Guidance

    empirical result

    To ensure independence between motion and appearance guidance, the Pearson Correlation Coefficient (PCC, rr) between the motion factor γm\gamma^\text{m} and appearance factor γa\gamma^\text{a} was computed across 8,000 random video first-and-last frame pairs using two semantic models Ω(⋅)\Omega(\cdot):

    • DINO: Achieved a Pearson correlation of r=0.40r = 0.40 with a uniform scatter distribution across the factor space.
    • CLIP: Achieved a Pearson correlation of r=0.43r = 0.43 with a less uniform distribution.

    The lower correlation and broader distribution indicate that self-supervised vision transformer features (DINO) are more selectively sensitive to appearance changes without confounding pixel-level motion, making DINO the preferred semantic backbone for appearance guidance.

  10. Knowl 10 — Limitations of HiGen

    limitation

    HiGen exhibits two key limitations:

    1. Fine-Grained Detail Generation: Constrained by training dataset quality and computation resources, the synthesized frame detail lags behind state-of-the-art pure text-to-image synthesis models.
    2. Commonsense Physical Dynamics: Modeling complex human and animal actions adhering strictly to real-world physics and commonsense remains challenging, particularly in scenes with substantial motion magnitude.

Coverage note — None was omitted; all key architectural components, mathematical formulations, training algorithms, ablations, benchmark results, and stated limitations were converted into knowls.

References

  1. 1.Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin. Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv preprint arXiv:2304.08477, 2023. 3, 8
  2. 2.Yogesh Balaji, Martin Renqiang Min, Bing Bai, Rama Chellappa, and Hans Peter Graf. Conditional gan with discriminative filter generation for text-to-video synthesis. In IJCAI, page 2, 2019. 3
  3. 3.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. corr, vol. abs/2211.01324 (2022), 2022. 2
  4. 4.Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Text2live: Text-driven layered image and video editing. In ECCV, pages 707–723. Springer, 2022. 3
  5. 5.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In CVPR, pages 22563–22575, 2023. 1, 3, 8
  6. 6.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, pages 9650–9660, 2021. 4, 6, 7
  7. 7.Duygu Ceylan, Chun-Hao P Huang, and Niloy J Mitra. Pix2video: Video editing using image diffusion. In ICCV, pages 23206–23217, 2023. 3
  8. 8.Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter1: Open diffusion models for high-quality video generation, 2023. 1, 7, 8
  9. 9.Zhongjie Duan, Lizhou You, Chengyu Wang, Cen Chen, Ziheng Wu, Weining Qian, Jun Huang, Fei Chao, and Rongrong Ji. Diffsynth: Latent in-iteration deflickering for realistic video synthesis. arXiv preprint arXiv:2308.03463, 2023. 3
  10. 10.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In ICCV, pages 7346–7356, 2023. 1, 2, 4, 5, 8
  11. 11.Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. Convolutional two-stream network fusion for video action recognition. In CVPR, pages 1933–1941, 2016. 2
  12. 12.Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming-Yu Liu, and Yogesh Balaji. Preserve your own correlation: A noise prior for video diffusion models. In ICCV, pages 22930–22941, 2023. 3, 8
  13. 13.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In CVPR, pages 10696–10706, 2022. 2
  14. 14.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 3
  15. 15.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity video generation with arbitrary lengths. arXiv preprint arXiv:2211.13221, 2022. 3
  16. 16.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. NeurIPS, 33:6840–6851, 2020. 2, 4
  17. 17.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 1, 3, 5
  18. 18.Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. JMLR, 23 (1):2249–2281, 2022. 2
  19. 19.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arxiv e-prints, page. arXiv preprint arXiv:2204.03458, 3, 2022. 4
  20. 20.Susung Hong, Junyoung Seo, Sunghwan Hong, Heeseong Shin, and Seungryong Kim. Large language models are frame-level directors for zero-shot text-to-video generation. arXiv preprint arXiv:2305.14330, 2023. 3
  21. 21.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 3
  22. 22.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In ICLR, 2023. 8
  23. 23.Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. Free-bloom: Zero-shot text-to-video generator with llm director and ldm animator. arXiv preprint arXiv:2309.14494, 2023. 3
  24. 24.Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou. Composer: Creative and controllable image synthesis with composable conditions. arXiv preprint arXiv:2302.09778, 2023. 3
  25. 25.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv preprint arXiv:2303.13439, 2023. 3, 7, 8
  26. 26.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 3
  27. 27.Zhifeng Kong and Wei Ping. On fast sampling of diffusion probabilistic models. arXiv preprint arXiv:2106.00132, 2021. 3
  28. 28.Xin Li, Wenqing Chu, Ye Wu, Weihang Yuan, Fanglong Liu, Qi Zhang, Fu Li, Haocheng Feng, Errui Ding, and Jingdong Wang. Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation. arXiv preprint arXiv:2309.00398, 2023. 3
  29. 29.Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell, and Boyi Li. Llm-grounded video diffusion models. arXiv preprint arXiv:2309.17444, 2023. 3
  30. 30.Binhui Liu, Xin Liu, Anbo Dai, Zhiyong Zeng, Zhen Cui, and Jian Yang. Dual-stream diffusion net for text-to-video generation. arXiv preprint arXiv:2308.08316, 2023. 3
  31. 31.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 5
  32. 32.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. In CVPR, pages 10209–10218, 2023. 2, 3
  33. 33.Eyal Molad, Eliahu Horwitz, Dani Valevski, Alex Rav Acha, Yossi Matias, Yael Pritch, Yaniv Leviathan, and Yedid Hoshen. Dreamix: Video diffusion models are general video editors. arXiv preprint arXiv:2302.01329, 2023. 3
  34. 34.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 2
  35. 35.Maxime Oquab, Timothee Darcet, Theo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 4, 6, 7
  36. 36.Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in gan evaluation. In CVPR, pages 11410–11420, 2022. 5
  37. 37.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 2, 8
  38. 38.Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. Fatezero: Fusing attentions for zero-shot text-based video editing. arXiv preprint arXiv:2303.09535, 2023. 3
  39. 39.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021. 4, 6, 7
  40. 40.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 2
  41. 41.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10684–10695, 2022. 2, 3, 4, 5
  42. 42.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In CVPR, pages 22500–22510, 2023. 3
  43. 43.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 35:36479–36494, 2022. 2
  44. 44.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. 5
  45. 45.Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Mostgan-v: Video generation with temporal motion styles. In CVPR, pages 5652–5661, 2023. 3
  46. 46.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 3, 5, 8
  47. 47.Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In CVPR, pages 3626–3636, 2022. 3
  48. 48.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 2
  49. 49.Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N Metaxas, and Sergey Tulyakov. A good image generator is what you need for high-resolution video synthesis. arXiv preprint arXiv:2104.15069, 2021. 3
  50. 50.Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 5
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 30, 2017. 4
  52. 52.Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics. NeurIPS, 29, 2016. 3
  53. 53.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 1, 2, 3, 4, 5, 7, 8
  54. 54.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks: Towards good practices for deep action recognition. In ECCV, pages 20–36. Springer, 2016. 2, 4
  55. 55.Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation. arXiv preprint arXiv:2305.10874, 2023. 1, 3
  56. 56.Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. Videocomposer: Compositional video synthesis with motion controllability. arXiv preprint arXiv:2306.02018, 2023. 4, 8
  57. 57.Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023. 3
  58. 58.Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021. 5
  59. 59.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In ICCV, pages 7623–7633, 2023. 3
  60. 60.Zhen Xing, Qi Dai, Han Hu, Zuxuan Wu, and Yu-Gang Jiang. Simda: Simple diffusion adapter for efficient video generation. arXiv preprint arXiv:2308.09710, 2023. 3
  61. 61.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, pages 5288–5296, 2016. 2, 5, 8
  62. 62.Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. Generating videos with dynamics-aware implicit generative adversarial networks. arXiv preprint arXiv:2202.10571, 2022. 3
  63. 63.David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. arXiv preprint arXiv:2309.15818, 2023. 3
  64. 64.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In ICCV, pages 3836–3847, 2023. 3
  65. 65.Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:2311.04145, 2023. 8
  66. 66.Long Zhao, Xi Peng, Yu Tian, Mubbasir Kapadia, and Dimitris Metaxas. Learning to forecast and refine residual motion for image-to-video generation. In ECCV, pages 387–403, 2018. 3
  67. 67.Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jiawei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. Motiondirector: Motion customization of text-to-video diffusion models. arXiv preprint arXiv:2310.08465, 2023. 3
  68. 68.Yuan Zhi, Zhan Tong, Limin Wang, and Gangshan Wu. Mgsampler: An explainable sampling strategy for video action recognition. In ICCV, pages 1513–1522, 2021. 4
  69. 69.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022. 3, 4, 5, 8

Citation

MLA
Qing, Z., et al. “Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation”. arXiv, 2023, http://arxiv.org/abs/2312.04483v1.
APA
Qing, Z., Zhang, S., Wang, J., Wang, X., Wei, Y., Zhang, Y., Gao, C., & Sang, N. (2023). Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation. arXiv. http://arxiv.org/abs/2312.04483v1
Chicago
Qing, Z., S. Zhang, J. Wang, et al. 2023. “Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation”. arXiv. http://arxiv.org/abs/2312.04483v1.
Harvard
Qing, Z. et al. (2023) “Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.04483v1.
Vancouver
1. Qing Z, Zhang S, Wang J, Wang X, Wei Y, Zhang Y, Gao C, Sang N (2023) Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation. arXiv

BibTeX

@article{qing2023hierarchical,
  title = {Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation},
  author = {Qing, Zhiwu and Zhang, Shiwei and Wang, Jiayu and Wang, Xiang and Wei, Yujie and Zhang, Yingya and Gao, Changxin and Sang, Nong},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.04483v1},
  eprint = {2312.04483}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE