StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Roberto HenschelLevon KhachatryanHayk PoghosyanDaniil HayrapetyanVahram TadevosyanZhangyang WangShant NavasardyanHumphrey Shi

article2025CVPR261 citations

Proposes an autoregressive text-to-video framework that combines short- and long-term memory modules with randomized blending to generate temporally consistent, high-motion videos of extended length without visual stagnation or frame degradation.

Listen

Text-to-video diffusion models have rapidly advanced, enabling the automated creation of short video clips from descriptive prompts. However, expanding these tools to generate longer videos—such as sequences extending to several minutes—remains a major operational bottleneck for real-world applications like advertising and digital storytelling. Existing long-form generation methods typically suffer from abrupt scene transitions, severe visual quality degradation over time, or motion stagnation where the scene freezes into a static image.

The article introduces and evaluates StreamingT2V, an autoregressive framework designed to synthesize consistent, high-dynamic, and extendable long videos from text descriptions without accumulating errors or visual artifacts.

The researchers developed an architecture comprising three integrated components. First, a conditional attention module manages short-term memory by extracting features from the preceding video segment and injecting them via attention mechanisms to ensure smooth frame transitions. Second, an appearance preservation module maintains long-term memory by extracting high-level object and scene details from an initial anchor frame, preventing the system from drifting away from the original subject. Third, a randomized blending technique enables a standard high-resolution video enhancement model to upscale overlapping video chunks smoothly across indefinite durations without requiring new model training.

The experimental evaluation demonstrated that StreamingT2V substantially outperforms existing open-source baselines on 240-frame video generations across 50 diverse test prompts. In motion-aware consistency evaluations, the framework achieved an error score roughly 28% lower than the second-best competitor, confirming both natural movement and continuity. Competing models either suffered from severe motion freezing or exhibited frequent artificial scene cuts—some producing over 100 times more abrupt cuts than StreamingT2V. In text-alignment metrics, StreamingT2V achieved the highest score among all evaluated approaches, demonstrating stable fidelity over time without visual degradation.

These findings indicate that incorporating dedicated short-term and long-term memory mechanisms solves the core technical challenges of autoregressive video generation. By avoiding the extreme compute costs of training massive end-to-end long-video models and repurposing existing short-video enhancement tools without additional training, this framework provides a highly cost-effective path for scalable content generation workflows.

For practical implementation, organizations looking to build long-form video capabilities should adopt memory-conditioned autoregressive pipelines rather than naive frame-by-frame extension. The authors note that the architecture can readily generalize to newer diffusion transformer backbones. Moving forward, engineering teams should conduct broader pilot deployments across varied visual styles and larger prompt libraries to evaluate edge cases, while researchers focus on validating performance on diverse model architectures.

arXiv: 2403.14773
Cover for StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text

Abstract

Text-to-video diffusion models enable the generation of high-quality videos that follow text instructions, simplifying the process of producing diverse and individual content. Current methods excel in generating short videos (up to 16s), but produce hard-cuts when naively extended to long video synthesis. To overcome these limitations, we present StreamingT2V, an autoregressive method that generates long videos of up to 2 minutes or longer with seamless transitions. The key components are: (i) a short-term memory block called conditional attention module (CAM), which conditions the current generation on the features extracted from the preceding chunk via an attentional mech-

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Method
  • 4.1. Conditional Attention Module
  • 4.2. Appearance Preservation Module
  • 4.3. Auto-regressive Video Enhancement
  • 5. Experiments
  • 5.1. Metrics
  • 5.2. Comparison with Baselines
  • 6. Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Three-Stage Pipeline for Consistent Long Video Synthesis in StreamingT2V

    model/method

    StreamingT2V generates extended, temporally consistent videos (e.g., 240 to 1200+ frames) from text descriptions through a three-stage autoregressive framework:

    1. Initialization Stage: An off-the-shelf text-to-video latent diffusion model (such as ModelScope) generates an initial base video chunk V1V_1 of F=16F = 16 frames at resolution 256×256256 \times 256.

    2. Streaming T2V Stage: Subsequent 16-frame chunks V2,V3,…V_2, V_3, \dots are generated autoregressively at 256×256256 \times 256 resolution. To maintain smooth motion and continuity across chunk boundaries without motion stagnation, a Conditional Attention Module (CAM) injects short-term memory from the last Fcond=8F_{\text{cond}} = 8 frames of the preceding chunk via cross-attention. Simultaneously, an Appearance Preservation Module (APM) injects long-term memory derived from a fixed anchor frame of the first chunk to maintain scene identity and object features without accumulated drift.

    3. Streaming Refinement Stage: The entire low-resolution long video is divided into overlapping chunks of F=24F = 24 frames (with overlap O=8\mathcal{O} = 8) and autoregressively upscaled and enhanced to high resolution (720×720720 \times 720 or 1280×7201280 \times 720) using an image-conditioned refiner video diffusion model (such as MS-Vid2Vid-XL). Smooth chunk enhancement is achieved using shared noise sampling and a randomized blending mechanism.

  2. Knowl 2 — Conditional Attention Module for Short-Term Memory

    model/method

    The Conditional Attention Module (CAM) provides short-term memory conditioning for autoregressive video generation. Rather than concatenating conditional frames directly into the diffusion model's latent input—which often leads to severe motion stagnation or hard cuts—CAM extracts feature maps from the preceding chunk's final FcondF_{\text{cond}} frames (typically Fcond=8F_{\text{cond}} = 8) and injects them via temporal multi-head attention into the UNet skip connections of the base video latent diffusion model (Video-LDM).

    CAM extracts features using a frame-wise image encoder Econd\mathcal{E}_{\text{cond}} followed by the initial encoder layers of the Video-LDM UNet up to its middle block. Let xCAM∈R(b⋅w⋅h)×Fcond×cx_{\text{CAM}} \in \mathbb{R}^{(b \cdot w \cdot h) \times F_{\text{cond}} \times c} denote the extracted CAM feature tensor, where bb is the batch size, w×hw \times h is the spatial feature resolution, and cc is the channel dimension. For each UNet skip-connection tensor xSC∈Rb×F×h×w×cx_{\text{SC}} \in \mathbb{R}^{b \times F \times h \times w \times c}, spatial-temporal group normalization and an input linear projection PinP_{\text{in}} are applied, followed by reshaping to xSC′∈R(b⋅w⋅h)×F×cx'_{\text{SC}} \in \mathbb{R}^{(b \cdot w \cdot h) \times F \times c}.

    CAM then performs Temporal Multi-Head Attention (T-MHA) per spatial position, treating the base model's skip features as queries and CAM features as keys and values:

    Q=PQ(xSC′),K=PK(xCAM),V=PV(xCAM)Q = P_Q(x'_{\text{SC}}), \quad K = P_K(x_{\text{CAM}}), \quad V = P_V(x_{\text{CAM}})

    xSC′′=T-MHA(Q,K,V)x''_{\text{SC}} = \text{T-MHA}(Q, K, V)

    where PQ,PK,PVP_Q, P_K, P_V are learnable linear projection matrices. The resulting features are projected through an output linear layer PoutP_{\text{out}}, reshaped by R\mathcal{R} back to dimensions b×F×h×w×cb \times F \times h \times w \times c, and added residually to the skip connection:

    xSC′′′=xSC+R(Pout(xSC′′))x'''_{\text{SC}} = x_{\text{SC}} + \mathcal{R}(P_{\text{out}}(x''_{\text{SC}}))

    PoutP_{\text{out}} is initialized to zero so that CAM initially produces an identity transformation on the base model's skip connections, stabilizing early training.

  3. Knowl 3 — Appearance Preservation Module for Long-Term Memory

    model/method

    Autoregressive video generation methods that only condition on the immediately preceding chunk tend to suffer from identity drift and forgetting of global scene details over long time horizons. The Appearance Preservation Module (APM) preserves global identity by conditioning every chunk generation on a fixed anchor frame from the first chunk.

    APM extracts high-level semantic features using a frozen CLIP image encoder on the anchor frame. The single CLIP image embedding token is expanded into k=16k = 16 tokens using a multi-layer perceptron (MLP). These kk image tokens are concatenated along the sequence dimension with the CLIP text tokens of the generation prompt τ\tau and passed through a 1D convolution and LayerNorm projection block, yielding a combined conditioning representation xmixed∈Rb×77×1024x_{\text{mixed}} \in \mathbb{R}^{b \times 77 \times 1024}, where bb is the batch size.

    To balance anchor frame visual guidance with text prompt fidelity, the conditioning representation xcrossx_{\text{cross}} supplied to the key and value projections of cross-attention layer ll of the Video-LDM UNet is computed as:

    xcross=SiLU(αl)xmixed+xtextx_{\text{cross}} = \text{SiLU}(\alpha_l) x_{\text{mixed}} + x_{\text{text}}

    where xtextx_{\text{text}} is the standard CLIP text encoding of the prompt, SiLU\text{SiLU} is the Sigmoid Linear Unit activation function, and αl∈R\alpha_l \in \mathbb{R} is a learnable scalar parameter per cross-attention layer ll, initialized to 0.

  4. Knowl 4 — Autoregressive Video Enhancement with Shared Noise and Randomized Blending

    algorithm

    To enhance long videos to high resolution (e.g., 1280×7201280 \times 720) without seam artifacts between sequentially processed chunks, StreamingT2V uses SDEdit on an image-conditioned video diffusion model (Refiner Video-LDM) combined with shared initial noise and step-wise randomized latent blending across overlapping frames.

    Input: Low-resolution video chunks V1,…,VmV_1, \dots, V_m, each with F=24F = 24 frames and an overlap of O=8\mathcal{O} = 8 frames between Vi−1V_{i-1} and ViV_i; Refiner Video-LDM ϵθ\epsilon_\theta; intermediate diffusion step T′<TT' < T.
    Output: Enhanced high-resolution long video.
    for each chunk Vi,i=1,…,mV_i, i = 1, \dots, m do
        Upscale ViV_i to target resolution using bilinear interpolation and encode to latent code x0(i)x_0(i)
        if i==1i == 1 then
            Sample standard Gaussian noise ϵ1∼N(0,I)\epsilon_1 \sim \mathcal{N}(0, I) of shape F×h×w×cF \times h \times w \times c
        else
            Sample new noise ϵ^i∼N(0,I)\hat{\epsilon}_i \sim \mathcal{N}(0, I) of shape (F−O)×h×w×c(F - \mathcal{O}) \times h \times w \times c
            Construct shared noise ϵi=concat([ϵi−1(F−O):F,ϵ^i],dim=0)\epsilon_i = \text{concat}([\epsilon_{i-1}^{(F-\mathcal{O}):F}, \hat{\epsilon}_i], \text{dim}=0)
        end if
        Add noise for T′T' forward steps to obtain noisy latent xT′(i)x_{T'}(i)
    end for
    for diffusion step t=T′,T′−1,…,1t = T', T'-1, \dots, 1 do
        for each chunk i=1,…,mi = 1, \dots, m do
            Compute denoised latent xt−1(i)x_{t-1}(i) using Refiner Video-LDM ϵθ\epsilon_\theta
        end for
        for each pair of adjacent chunks (Vi−1,Vi),i=2,…,m(V_{i-1}, V_i), i = 2, \dots, m do
            Let xL=xt−1(i−1)x_L = x_{t-1}(i-1) and xR=xt−1(i)x_R = x_{t-1}(i)
            Sample cutoff index fthr∼Uniform({0,…,O})f_{\text{thr}} \sim \text{Uniform}(\{0, \dots, \mathcal{O}\})
            Merge overlapping latents: xLR=concat([xL1:F−fthr,xRfthr+1:F],dim=0)x_{LR} = \text{concat}([x_L^{1:F-f_{\text{thr}}}, x_R^{f_{\text{thr}}+1:F}], \text{dim}=0)
            Update long video latent representation at overlapping indices with xLRx_{LR}
        end for
    end for
    Decode latents frame-wise using the VQ-GAN decoder D\mathcal{D} to yield the final enhanced video.

    Under this randomized blending rule, for frame f∈{1,…,O}f \in \{1, \dots, \mathcal{O}\} within the overlap region, the latent code from the preceding chunk Vi−1V_{i-1} is selected with probability 1−fO+11 - \frac{f}{\mathcal{O} + 1}, providing a seamless probabilistic transition between consecutive chunks.

  5. Knowl 5 — Quantitative Evaluation of Open-Source Text-to-Long-Video Generators

    data/table

    The quantitative performance of open-source text-to-long-video generation methods was evaluated on a benchmark of 50 diverse text prompts covering various actions, objects, and scenes. All methods generated 240-frame videos starting from an identical 16-frame initial chunk generated by Video-LDM at 720×720720 \times 720 resolution. Three metrics were evaluated: Motion Aware Warp Error (MAWE, assessing both motion magnitude and optical flow consistency; lower is better), SCuts (average number of scene cuts detected by PySceneDetect AdaptiveDetector; lower is better), and CLIP text-image similarity score (higher is better).

    Method ↓\downarrow MAWE ↓\downarrow SCuts ↑\uparrow CLIP
    SparseCtrl 6069.7 5.48 29.32
    I2VGenXL 2846.4 0.40 27.28
    DynamiCrafterXL 176.7 1.30 27.79
    SEINE 718.9 0.28 30.13
    SVD 857.2 1.10 23.95
    FreeNoise 1298.4 0.00 31.55
    OpenSora 1165.7 0.16 31.54
    OpenSoraPlan 72.9 0.24 29.34
    StreamingT2V (Ours) 52.3 0.04 31.73

    StreamingT2V achieves the lowest MAWE (52.3, nearly 30% lower than the second-best method, OpenSoraPlan at 72.9) and the highest CLIP score (31.73). While FreeNoise achieves an SCuts score of 0.00, it generates near-static videos with frozen camera and object motion. StreamingT2V achieves the lowest SCuts score (0.04) among all dynamic video generators.

  6. Knowl 6 — Generalization of StreamingT2V to Diffusion Transformer (DiT) Architectures

    model/method

    The StreamingT2V framework generalizes directly to Diffusion Transformer (DiT) backbones (such as OpenSora) without architectural incompatibility:

    1. CAM Adaptation: In DiT architectures, the Conditional Attention Module (CAM) is integrated by enabling the final 14 transformer blocks of the DiT model to attend to the previous chunk's extracted feature representations using CAM's temporal multi-head attention mechanism.
    2. APM Adaptation: The Appearance Preservation Module (APM) injects anchor frame representations into the cross-attention layers of the DiT blocks using the gated text-image mixture formulation.

    Visual inspection confirms that applying CAM and APM to DiT-based models maintains long-term object consistency and smooth chunk transitions in extended video generation.

  7. Knowl 7 — SCuts and Motion Aware Warp Error (MAWE) for Long Video Benchmarking

    definition

    Two quantitative metrics are defined to assess temporal consistency and motion quality in text-to-long-video synthesis:

    1. SCuts: Measures hard chunk boundaries and sudden transition artifacts by counting the total number of scene cuts detected in a video sequence using PySceneDetect's AdaptiveDetector with default parameters.

    2. Motion Aware Warp Error (MAWE): Evaluates temporal consistency by computing optical flow-based warping errors across consecutive frames while simultaneously accounting for the total magnitude of motion. A video with frozen content achieves a low raw warp error despite failing to animate; MAWE jointly penalizes both warping artifacts and motion stagnation, achieving its lowest values only when a video exhibits both high optical flow consistency and meaningful, dynamic motion.

Coverage note — None was omitted; all key architectural modules (CAM, APM, Refiner with randomized blending), evaluation metrics (MAWE, SCuts), DiT generalization, and quantitative benchmarks from the paper are fully covered.

References

  1. 1.Pyscenedetect. https://www.scenedetect.com/. Accessed: 2024-03-03. 6
  2. 2.Isaac Amidror. Scattered data interpolation methods for electronic imaging systems: a survey. Journal of electronic imaging, 11(2):157–176, 2002. 6
  3. 3.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 2, 3, 6, 7
  4. 4.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023. 2, 3, 4
  5. 5.Xi Chen, Zhiheng Liu, Mengting Chen, Yutong Feng, Yu Liu, Yujun Shen, and Hengshuang Zhao. Livephoto: Real image animation with text-guided motion control. arXiv preprint arXiv:2312.02928, 2023. 3
  6. 6.Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Seine: Short-to-long video diffusion model for generative transition and prediction. In The Twelfth International Conference on Learning Representations, 2023. 2, 3, 6, 7
  7. 7.Zuozhuo Dai, Zhenghao Zhang, Yao Yao, Bingxue Qiu, Siyu Zhu, Long Qin, and Weizhi Wang. Animateanything: Finegrained open domain image animation with motion guidance, 2023. 2, 3
  8. 8.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021. 3
  9. 9.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7346–7356, 2023. 3
  10. 10.Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu video: Factorizing text-to-video generation by explicit image conditioning. arXiv preprint arXiv:2311.10709, 2023. 2, 4
  11. 11.Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Sparsectrl: Adding sparse controls to text-to-video diffusion models. arXiv preprint arXiv:2311.16933, 2023. 2, 3, 5, 6, 7
  12. 12.Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized textto-image diffusion models without specific tuning. In The Twelfth International Conference on Learning Representations, 2023. 2
  13. 13.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. 2, 3
  14. 14.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 3, 4
  15. 15.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022. 2, 3
  16. 16.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Textto-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15954–15964, 2023. 2, 3
  17. 17.Jihwan Kim, Junoh Kang, Jinyoung Choi, and Bohyung Han. Fifo-diffusion: Generating infinite videos from text without training. In NeurIPS, 2024. 3
  18. 18.Xin Li, Wenqing Chu, Ye Wu, Weihang Yuan, Fanglong Liu, Qi Zhang, Fu Li, Haocheng Feng, Errui Ding, and Jingdong Wang. Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation. arXiv preprint arXiv:2309.00398, 2023. 2
  19. 19.Fuchen Long, Zhaofan Qiu, Ting Yao, and Tao Mei. Videodrafter: Content-consistent multi-scene video generation with llm. arXiv preprint arXiv:2401.01256, 2024. 3
  20. 20.Chenlin Meng, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022. 2, 6
  21. 21.Gyeongrok Oh, Jaehwan Jeong, Sieun Kim, Wonmin Byeon, Jinkyu Kim, Sungwoong Kim, Hyeokmin Kwon, and Sangpil Kim. Mtvg: Multi-text video generation with text-to-video models. arXiv preprint arXiv:2312.04086, 2023. 2, 3
  22. 22.PKU-Yuan-Lab and Tuzhan-AI. Open-sora-plan, 2024. 2, 3, 6, 7
  23. 23.Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. Freenoise: Tuning-free longer video diffusion via noise rescheduling. In The Twelfth International Conference on Learning Representations, 2024. 3, 6, 7
  24. 24.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 2, 3, 5, 7
  25. 25.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2, 4
  26. 26.Weiming Ren, Harry Yang, Ge Zhang, Cong Wei, Xinrun Du, Stephen Huang, and Wenhu Chen. Consisti2v: Enhancing visual consistency for image-to-video generation. arXiv preprint arXiv:2402.04324, 2024. 3
  27. 27.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 2, 4
  28. 28.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015. 4
  29. 29.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. In The Eleventh International Conference on Learning Representations, 2022. 2, 3, 4
  30. 30.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256–2265. PMLR, 2015. 2
  31. 31.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2020. 2
  32. 32.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017. 3
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 4
  34. 34.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual descriptions. In International Conference on Learning Representations, 2022. 2
  35. 35.Fu-Yun Wang, Wenshuo Chen, Guanglu Song, Han-Jia Ye, Yu Liu, and Hongsheng Li. Gen-l-video: Multi-text to long video generation via temporal co-denoising. arXiv preprint arXiv:2305.18264, 2023. 3
  36. 36.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 2, 4
  37. 37.Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. Videocomposer: Compositional video synthesis with motion controllability. Advances in Neural Information Processing Systems, 36, 2024. 2, 3, 6
  38. 38.Wenming Weng, Ruoyu Feng, Yanhui Wang, Qi Dai, Chunyu Wang, Dacheng Yin, Zhiyuan Zhao, Kai Qiu, Jianmin Bao, Yuhui Yuan, Chong Luo, Yueyi Zhang, and Zhiwei Xiong. Art•v: Auto-regressive text-to-video generation with diffusion models. arXiv preprint arXiv:2311.18834, 2023. 3
  39. 39.Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Xintao Wang, Tien-Tsin Wong, and Ying Shan. Dynamicrafter: Animating open-domain images with video diffusion priors. arXiv preprint arXiv:2310.12190, 2023. 2, 3, 6, 7
  40. 40.Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. Make pixels dance: Highdynamic video generation. arXiv:2311.10982, 2023. 3
  41. 41.David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou. Show-1: Marrying pixel and latent diffusion models for text-to-video generation. arXiv preprint arXiv:2309.15818, 2023. 2
  42. 42.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 3, 4, 5
  43. 43.Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qing, Xiang Wang, Deli Zhao, and Jingren Zhou. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. 2023. 2, 3, 4, 6, 7
  44. 44.Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, 2024. 2, 3, 6, 7, 8
  45. 45.Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou. Storydiffusion: Consistent self-attention for long-range image and video generation. 2024. 3

Citation

MLA
Henschel, R., et al. “StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text”. arXiv, 2024, http://arxiv.org/abs/2403.14773v2.
APA
Henschel, R., Khachatryan, L., Poghosyan, H., Hayrapetyan, D., Tadevosyan, V., Wang, Z., Navasardyan, S., & Shi, H. (2024). StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text. arXiv. http://arxiv.org/abs/2403.14773v2
Chicago
Henschel, R., L. Khachatryan, H. Poghosyan, et al. 2024. “StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text”. arXiv. http://arxiv.org/abs/2403.14773v2.
Harvard
Henschel, R. et al. (2024) “StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.14773v2.
Vancouver
1. Henschel R, Khachatryan L, Poghosyan H, Hayrapetyan D, Tadevosyan V, Wang Z, Navasardyan S, Shi H (2024) StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text. arXiv

BibTeX

@article{henschel2024streamingt2v,
  title = {StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text},
  author = {Henschel, Roberto and Khachatryan, Levon and Poghosyan, Hayk and Hayrapetyan, Daniil and Tadevosyan, Vahram and Wang, Zhangyang and Navasardyan, Shant and Shi, Humphrey},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.14773v2},
  eprint = {2403.14773}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE