FIFO-Diffusion: Generating Infinite Videos from Text without Training
Jihwan KimJunoh KangJinyoung ChoiBohyung Han
Proposes FIFO-Diffusion, a training-free inference technique that enables pretrained diffusion models to generate infinitely long videos with constant memory usage by processing a first-in-first-out frame queue via diagonal denoising.
Generating extended video content using generative artificial intelligence has historically been constrained by heavy computational demands and limited temporal coherence. While diffusion models excel at producing short video clips of around 16 frames, extending them to generate long videos typically results in severe quality decay, unnatural motion discontinuities, or excessive memory requirements. Existing workarounds, such as chunked autoregressive generation that stitches together short segments, often fail to maintain consistent narrative context across scene transitions and suffer from cumulative visual errors.
The article addresses this limitation by introducing FIFO-Diffusion, a novel inference framework designed to generate arbitrarily long videos from text prompts without requiring additional model training or fine-tuning. The primary objective is to demonstrate that standard video diffusion models pretrained only on short clips can be adapted at test time to produce seamless, infinitely long videos while maintaining constant memory overhead and visual fidelity.
The approach operates via a first-in-first-out queue mechanism using diagonal denoising. Instead of processing entire video chunks at uniform noise levels, the system maintains a sequence of consecutive frames with progressively increasing noise levels. At each step, the fully refined frame at the head of the queue is finalized and output, while a new random noise frame enters the tail. To resolve the discrepancy between training conditions (uniform noise across frames) and inference conditions (varying noise across frames), the method incorporates latent partitioning to narrow noise differences and lookahead denoising to allow noisier frames to accurately reference cleaner preceding frames. The authors evaluated the framework across multiple open-source baseline models, benchmark datasets such as UCF-101, and extensive user studies.
The findings show that FIFO-Diffusion successfully produces videos exceeding 10,000 frames without perceptual degradation or loss of semantic coherence. Quantitatively, the method achieved state-of-the-art video quality and consistency scores on standard benchmarks, outperforming competing chunked autoregressive models while requiring only 64 inference steps compared to 400 steps in prior methods. In human evaluations covering 70 participants across 111 rating sets, users demonstrated a strong preference for FIFO-Diffusion over competing training-free baselines, particularly regarding motion plausibility and dynamics. Crucially, the method maintained a strictly constant memory footprint of approximately 11.2 to 13.5 gigabytes regardless of whether the output was 128 or 512 frames, whereas baseline methods quickly encountered out-of-memory errors.
These results carry significant operational implications for media production, simulation, and generative AI infrastructure. By eliminating the need to retrain large diffusion architectures for long video outputs, organizations can substantially reduce training compute costs and deployment timelines. Furthermore, the ability to partition latents allows parallel execution across multiple graphical processing units (GPUs), bringing per-frame inference latency down from 12.37 seconds on a single GPU to 1.84 seconds on an eight-GPU cluster.
Decision-makers considering long video generation pipelines can explore FIFO-Diffusion as an efficient, plug-and-play inference upgrade for existing diffusion architectures. However, while latent partitioning substantially reduces the training-inference gap, a subtle structural distribution gap remains because base models were originally trained on uniform noise levels. For future development, aligning the training phase itself with the diagonal denoising paradigm offers a promising pathway to enhance output fidelity and stability further.
- Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). This foundational work establishes flexible conditioning and sampling strategies for long video generation using diffusion models, providing the core context for subsequent infinite-video inference methods like FIFO-Diffusion.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This paper introduces video diffusion models and reconstruction-guided autoregressive sampling for longer sequences, which FIFO-Diffusion builds upon to enable continuous queue-based diagonal denoising.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It introduces latent diffusion models adapted for video generation via interleaved temporal layers, serving as a primary baseline architecture that FIFO-Diffusion leverages at inference time.
- Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). It establishes key cascaded diffusion concepts and temporal modeling architectures for text-conditional video synthesis that motivate training-free infinite generation.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It provides a widely adopted state-of-the-art latent video diffusion foundation model architecture that is directly relevant to pretrained video generation pipelines.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). It formalizes deterministic non-Markovian sampling for diffusion models, which underlies the multi-step trajectory manipulations used in diagonal denoising.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). It provides the foundational mathematical formulations and U-Net-based reverse denoising principles that all video diffusion methods extend.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). This work directly addresses long-horizon video rollout stability and per-frame independent noise assignments, extending concepts related to forward referencing and history conditioning in video diffusion.
- Paper: Learning Temporally Consistent Video Depth from Video Diffusion Priors, Jiahao Shao et al. (2025). This work adapts per-frame independent noise scheduling and sliding-window multi-clip inference principles to achieve temporally consistent long-video depth estimation.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). This paper advances extended-horizon video synthesis by incorporating explicit trajectory and motion controls into scalable diffusion transformer architectures.
