StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
Roberto HenschelLevon KhachatryanHayk PoghosyanDaniil HayrapetyanVahram TadevosyanZhangyang WangShant NavasardyanHumphrey Shi
Proposes an autoregressive text-to-video framework that combines short- and long-term memory modules with randomized blending to generate temporally consistent, high-motion videos of extended length without visual stagnation or frame degradation.
Text-to-video diffusion models have rapidly advanced, enabling the automated creation of short video clips from descriptive prompts. However, expanding these tools to generate longer videos—such as sequences extending to several minutes—remains a major operational bottleneck for real-world applications like advertising and digital storytelling. Existing long-form generation methods typically suffer from abrupt scene transitions, severe visual quality degradation over time, or motion stagnation where the scene freezes into a static image.
The article introduces and evaluates StreamingT2V, an autoregressive framework designed to synthesize consistent, high-dynamic, and extendable long videos from text descriptions without accumulating errors or visual artifacts.
The researchers developed an architecture comprising three integrated components. First, a conditional attention module manages short-term memory by extracting features from the preceding video segment and injecting them via attention mechanisms to ensure smooth frame transitions. Second, an appearance preservation module maintains long-term memory by extracting high-level object and scene details from an initial anchor frame, preventing the system from drifting away from the original subject. Third, a randomized blending technique enables a standard high-resolution video enhancement model to upscale overlapping video chunks smoothly across indefinite durations without requiring new model training.
The experimental evaluation demonstrated that StreamingT2V substantially outperforms existing open-source baselines on 240-frame video generations across 50 diverse test prompts. In motion-aware consistency evaluations, the framework achieved an error score roughly 28% lower than the second-best competitor, confirming both natural movement and continuity. Competing models either suffered from severe motion freezing or exhibited frequent artificial scene cuts—some producing over 100 times more abrupt cuts than StreamingT2V. In text-alignment metrics, StreamingT2V achieved the highest score among all evaluated approaches, demonstrating stable fidelity over time without visual degradation.
These findings indicate that incorporating dedicated short-term and long-term memory mechanisms solves the core technical challenges of autoregressive video generation. By avoiding the extreme compute costs of training massive end-to-end long-video models and repurposing existing short-video enhancement tools without additional training, this framework provides a highly cost-effective path for scalable content generation workflows.
For practical implementation, organizations looking to build long-form video capabilities should adopt memory-conditioned autoregressive pipelines rather than naive frame-by-frame extension. The authors note that the architecture can readily generalize to newer diffusion transformer backbones. Moving forward, engineering teams should conduct broader pilot deployments across varied visual styles and larger prompt libraries to evaluate edge cases, while researchers focus on validating performance on diverse model architectures.
- Paper: FIFO-Diffusion: Generating Infinite Videos from Text without Training, Jihwan Kim et al. (2024). This paper presents FIFO-Diffusion, introducing diagonal denoising for infinite video generation, establishing a foundational baseline for non-training long-video synthesis that StreamingT2V contrasts with and improves upon.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). This work establishes standard latent video diffusion architectures and temporal layer extensions that serve as common backbones for modern text-to-video generation pipelines.
- Paper: AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. (2024). This paper introduces modular motion adapters for diffusion backbones, providing key architectural concepts for inserting temporal attention into pre-trained image diffusion models.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). This foundational paper presents temporal latent diffusion modeling for high-resolution video synthesis, demonstrating the core spatial-temporal decoupling adapted by subsequent long-video generators.
- Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). This paper explores autoregressive and subset-conditioned diffusion models for generating long video sequences, articulating the challenges of temporal error accumulation and context forgetting.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). This study introduces VBench, the comprehensive multi-dimensional benchmark suite widely used to quantify temporal consistency, motion dynamism, and visual quality in video generation models.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). This work details cascaded video diffusion models and spatial-temporal super-resolution stages, illustrating standard video enhancement architectures repurposed in modular generation systems.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This pioneering paper introduces 3D U-Net diffusion architectures and reconstruction-guided sampling for video synthesis, laying the theoretical foundation for video diffusion.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). This paper develops History Guidance and Diffusion Forcing Transformers to condition autoregressive video generation on flexible subsets of past frames, directly addressing rollout stability in extended-horizon video synthesis.
- Paper: Wan: Open and Advanced Large-Scale Video Generative Models, Ang Wang et al. (2025). This work introduces the large-scale Wan diffusion transformer foundation models, incorporating advanced flow matching and streaming video generation mechanisms.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). This paper extends diffusion transformer video generation to long-horizon trajectories by explicitly fusing user-specified motion paths into the architecture.
- Paper: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation, Kaiyue Sun 0001 et al. (2025). This study introduces T2V-CompBench to evaluate complex spatio-temporal compositional dynamics and dynamic attribute consistency over time across text-to-video models.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). This paper presents a training-free framework that extracts and guides temporal attention to clone complex motion dynamics into text-to-video diffusion pipelines.
- Paper: Learning Temporally Consistent Video Depth from Video Diffusion Priors, Jiahao Shao et al. (2025). This work leverages pre-trained video diffusion priors and context sliding-window inference to maintain temporal consistency across continuous multi-clip downstream tasks like video depth estimation.
