FIFO-Diffusion: Generating Infinite Videos from Text without Training

Jihwan KimJunoh KangJinyoung ChoiBohyung Han

article2024NeurIPS154 citations

Proposes FIFO-Diffusion, a training-free inference technique that enables pretrained diffusion models to generate infinitely long videos with constant memory usage by processing a first-in-first-out frame queue via diagonal denoising.

Listen

Generating extended video content using generative artificial intelligence has historically been constrained by heavy computational demands and limited temporal coherence. While diffusion models excel at producing short video clips of around 16 frames, extending them to generate long videos typically results in severe quality decay, unnatural motion discontinuities, or excessive memory requirements. Existing workarounds, such as chunked autoregressive generation that stitches together short segments, often fail to maintain consistent narrative context across scene transitions and suffer from cumulative visual errors.

The article addresses this limitation by introducing FIFO-Diffusion, a novel inference framework designed to generate arbitrarily long videos from text prompts without requiring additional model training or fine-tuning. The primary objective is to demonstrate that standard video diffusion models pretrained only on short clips can be adapted at test time to produce seamless, infinitely long videos while maintaining constant memory overhead and visual fidelity.

The approach operates via a first-in-first-out queue mechanism using diagonal denoising. Instead of processing entire video chunks at uniform noise levels, the system maintains a sequence of consecutive frames with progressively increasing noise levels. At each step, the fully refined frame at the head of the queue is finalized and output, while a new random noise frame enters the tail. To resolve the discrepancy between training conditions (uniform noise across frames) and inference conditions (varying noise across frames), the method incorporates latent partitioning to narrow noise differences and lookahead denoising to allow noisier frames to accurately reference cleaner preceding frames. The authors evaluated the framework across multiple open-source baseline models, benchmark datasets such as UCF-101, and extensive user studies.

The findings show that FIFO-Diffusion successfully produces videos exceeding 10,000 frames without perceptual degradation or loss of semantic coherence. Quantitatively, the method achieved state-of-the-art video quality and consistency scores on standard benchmarks, outperforming competing chunked autoregressive models while requiring only 64 inference steps compared to 400 steps in prior methods. In human evaluations covering 70 participants across 111 rating sets, users demonstrated a strong preference for FIFO-Diffusion over competing training-free baselines, particularly regarding motion plausibility and dynamics. Crucially, the method maintained a strictly constant memory footprint of approximately 11.2 to 13.5 gigabytes regardless of whether the output was 128 or 512 frames, whereas baseline methods quickly encountered out-of-memory errors.

These results carry significant operational implications for media production, simulation, and generative AI infrastructure. By eliminating the need to retrain large diffusion architectures for long video outputs, organizations can substantially reduce training compute costs and deployment timelines. Furthermore, the ability to partition latents allows parallel execution across multiple graphical processing units (GPUs), bringing per-frame inference latency down from 12.37 seconds on a single GPU to 1.84 seconds on an eight-GPU cluster.

Decision-makers considering long video generation pipelines can explore FIFO-Diffusion as an efficient, plug-and-play inference upgrade for existing diffusion architectures. However, while latent partitioning substantially reduces the training-inference gap, a subtle structural distribution gap remains because base models were originally trained on uniform noise levels. For future development, aligning the training phase itself with the diagonal denoising paradigm offers a promising pathway to enhance output fidelity and stability further.

  • Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). This foundational work establishes flexible conditioning and sampling strategies for long video generation using diffusion models, providing the core context for subsequent infinite-video inference methods like FIFO-Diffusion.
  • Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This paper introduces video diffusion models and reconstruction-guided autoregressive sampling for longer sequences, which FIFO-Diffusion builds upon to enable continuous queue-based diagonal denoising.
  • Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It introduces latent diffusion models adapted for video generation via interleaved temporal layers, serving as a primary baseline architecture that FIFO-Diffusion leverages at inference time.
  • Paper: Imagen Video: High Definition Video Generation with Diffusion Models, Jonathan Ho et al. (2022). It establishes key cascaded diffusion concepts and temporal modeling architectures for text-conditional video synthesis that motivate training-free infinite generation.
  • Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It provides a widely adopted state-of-the-art latent video diffusion foundation model architecture that is directly relevant to pretrained video generation pipelines.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). It formalizes deterministic non-Markovian sampling for diffusion models, which underlies the multi-step trajectory manipulations used in diagonal denoising.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). It provides the foundational mathematical formulations and U-Net-based reverse denoising principles that all video diffusion methods extend.
Cover for FIFO-Diffusion: Generating Infinite Videos from Text without Training

Abstract

We propose a novel inference technique based on a pretrained diffusion model for text-conditional video generation. Our approach, called FIFO-Diffusion, is conceptually capable of generating infinitely long videos without additional training. This is achieved by iteratively performing diagonal denoising, which simultaneously processes a series of consecutive frames with increasing noise levels in a queue; our method dequeues a fully denoised frame at the head while enqueuing a new random noise frame at the tail. However, diagonal denoising is a double-edged sword as the frames near the tail can take advantage of cleaner frames by forward reference but such a strategy induces the discrepancy between training and inference. Hence, we introduce latent partitioning to reduce the training-inference gap and lookahead denoising to leverage the benefit of forward referencing. Practically, FIFO-Diffusion consumes a constant amount of memory regardless of the target video length given a baseline model, while well-suited for parallel inference on multiple GPUs. We have demonstrated the promising results and effectiveness of the proposed methods on existing text-to-video generation baselines. Generated video examples and source codes are available at our project page^1.

Table of Contents

  • 1 Introduction
  • 2 Text-to-Video Diffusion Models
  • 3 FIFO-Diffusion
  • 3.1 Diagonal denoising
  • 3.2 Latent partitioning
  • 3.3 Lookahead denoising
  • 4 Experiment
  • 4.1 Implementation details
  • 4.2 Qualitative results
  • 4.3 Quantitative results
  • 4.4 Computational cost
  • 4.5 Ablation study
  • 5 Related work
  • 5.1 Video diffusion models
  • 5.2 Long video generation
  • 5.3 Diffusion models with latents of different noise levels
  • 6 Conclusion
  • Acknowledgements
  • References
  • Appendix
  • A Details for Lemma 3.2 and Theorem 3.3
  • A.1 Proof of Lemma 3.2
  • A.2 Justification on (Hypothesis 2) of Theorem 3.3
  • B Implementation details
  • B.1 Details for user study
  • C Algorithms of FIFO-Diffusion
  • D Qualitative results of FIFO-Diffusion
  • D.1 VideoCrafter2
  • D.2 VideoCrafter1
  • D.3 zeroscope
  • D.4 Open-Sora Plan
  • E Multi-prompts generation for FIFO-Diffusion
  • E.1 Method
  • E.2 Qualitative results
  • F Qualitative comparisons with other long video generation methods
  • G Motion evaluation
  • H Ablation study
  • I Potential Broader Impact
  • NeurIPS Paper Checklist

Knowls

  1. Knowl 1 — Diagonal Denoising Framework for Text-to-Video Diffusion Models

    model/method

    Diagonal denoising is a training-free inference framework that enables a text-to-video diffusion model trained on short clips of length ff to synthesize arbitrarily long video sequences of NN frames (N≫fN \gg f) with constant memory consumption.

    In standard video diffusion generation, a sequence of ff frames is initialized at the maximum noise level and denoised simultaneously across synchronized timesteps τt\tau_t. In diagonal denoising, the diffusion process operates on a first-in, first-out (FIFO) queue QQ holding ff consecutive latent frames with strictly increasing noise levels:

    [zτ11;zτ22;… ;zτff][z^1_{\tau_1}; z^2_{\tau_2}; \dots; z^f_{\tau_f}]

    where 0=τ0<τ1<⋯<τf=T0 = \tau_0 < \tau_1 < \dots < \tau_f = T defines the discrete diffusion timestep schedule. At each generation step, the model applies a multi-frame denoising operator Φ\Phi conditioned on text prompt cc and the vector of timesteps:

    [zτ01;zτ12;… ;zτf−1f]=Φ([zτ11;zτ22;… ;zτff],[τ1;τ2;… ;τf],c;ϵθ)[z^1_{\tau_0}; z^2_{\tau_1}; \dots; z^f_{\tau_{f-1}}] = \Phi([z^1_{\tau_1}; z^2_{\tau_2}; \dots; z^f_{\tau_f}], [\tau_1; \tau_2; \dots; \tau_f], c; \epsilon_\theta)

    Upon completing the denoising step:

    1. The head frame zτ01z^1_{\tau_0}, having reached noise level τ0=0\tau_0 = 0 (completely denoised), is popped from the queue and decoded via the latent decoder Dec(⋅)\text{Dec}(\cdot) into an RGB video frame.
    2. A fresh random Gaussian noise latent zτfi+f∼N(0,I)z^{i+f}_{\tau_f} \sim \mathcal{N}(0, I) at the maximum noise level τf\tau_f is enqueued at the tail.

    By progressing frame by frame with a temporal stride of 1, each frame is iteratively denoised over ff steps while conditioning on preceding frames at cleaner noise levels, preventing chunk boundary discontinuities without requiring additional training.

  2. Knowl 2 — Latent Partitioning in FIFO-Diffusion

    model/method

    Latent partitioning is a technique designed to mitigate the training-inference gap of diagonal denoising and enable multi-GPU parallel processing. Pretrained video diffusion models are typically trained on video clips where all frames share identical noise levels, whereas diagonal denoising inputs frames across the full noise schedule τ1\tau_1 to τf\tau_f.

    To narrow the range of noise levels processed simultaneously, latent partitioning extends the FIFO queue length by a factor of nn (n≥2n \ge 2) from ff to nfn f frames, containing diagonal latents [zτ11;… ;zτnfnf][z^1_{\tau_1}; \dots; z^{nf}_{\tau_{nf}}] over a finer nfn f-step schedule 0=τ0<τ1<⋯<τnf=T0 = \tau_0 < \tau_1 < \dots < \tau_{nf} = T. The queue is divided into nn contiguous blocks of equal size ff:

    Q=[Q0;Q1;… ;Qn−1]Q = [Q_0; Q_1; \dots; Q_{n-1}]

    where each block Qk=[zτkf+1kf+1;… ;zτ(k+1)f(k+1)f]Q_k = [z^{kf+1}_{\tau_{kf+1}}; \dots; z^{(k+1)f}_{\tau_{(k+1)f}}] contains latents spanning the narrower timestep sub-sequence τk=[τkf+1;… ;τ(k+1)f]\tau_k = [\tau_{kf+1}; \dots; \tau_{(k+1)f}].

    Each block is denoised independently using the base model:

    Qk←Φ(Qk,τk,c;ϵθ)for k=0,…,n−1Q_k \leftarrow \Phi(Q_k, \tau_k, c; \epsilon_\theta) \quad \text{for } k = 0, \dots, n-1

    Latent partitioning yields three properties:

    1. It reduces the maximum noise level difference seen by the network from ∣στnf−στ1∣|\sigma_{\tau_{nf}} - \sigma_{\tau_1}| to max⁡k∣στ(k+1)f−στkf+1∣\max_k |\sigma_{\tau_{(k+1)f}} - \sigma_{\tau_{kf+1}}|, bounding the prediction error.
    2. The nn blocks can be computed concurrently across multiple GPUs, significantly accelerating inference.
    3. It increases the total number of denoising steps per frame from ff to nfn f, reducing ODE discretization error.
  3. Knowl 3 — Lookahead Denoising Mechanism

    model/method

    Lookahead denoising leverages forward referencing—the observation that noisier frames achieve more accurate score predictions when conditioned on cleaner preceding frames—to improve diagonal denoising quality and reduce temporal flickering.

    In standard latent partitioning with nn blocks of size ff, frames near the head of a block reference fewer preceding cleaner frames. Lookahead denoising introduces a sliding denoising window with a stride of f′=⌊f/2⌋f' = \lfloor f/2 \rfloor, processing 2n2n overlapping windows but updating only the benefited trailing half (the last f′f' frames) of each window.

    For k=0,…,2n−1k = 0, \dots, 2n - 1, a window QkQ_k of length ff covering timesteps τk=[τkf′+1;… ;τkf′+f]\tau_k = [\tau_{k f' + 1}; \dots; \tau_{k f' + f}] is passed to the denoiser Φ\Phi, and only the output indices from f′+1f' + 1 to ff are retained:

    Qkf′+1:f←Φ(Qk,τk,c;ϵθ)f′+1:fQ_k^{f'+1:f} \leftarrow \Phi(Q_k, \tau_k, c; \epsilon_\theta)^{f'+1:f}

    This guarantees that every latent frame updated in the queue is predicted while attending to at least f′f' cleaner preceding frames in the context window. To provide initial forward context for the earliest frames, the queue is padded at the front with f′f' dummy replicas of the cleanest latent. While this requires evaluating 2n2n blocks per step, the computations are mutually parallelizable across GPUs.

  4. Knowl 4 — Noise Prediction Error Bound for Diagonal Denoising

    theoretical result

    Let a video diffusion latent trajectory be defined by ztvdm:=[zt1;… ;ztf]z^{\text{vdm}}_t := [z^1_t; \dots; z^f_t], where ztiz^i_t represents the latent of the ii-th frame at diffusion time tt with noise scale σt=ct\sigma_t = c t (cc constant), satisfying the probability flow ODE:

    dztvdm=c⋅ϵ(ztvdm,t⋅1)dt\mathrm{d}z^{\text{vdm}}_t = c \cdot \epsilon(z^{\text{vdm}}_t, t \cdot \mathbf{1}) \mathrm{d}t

    where 1=[1;… ;1]∈Rf\mathbf{1} = [1; \dots; 1] \in \mathbb{R}^f and ϵ(⋅)\epsilon(\cdot) is the scaled score function −σ∇zlog⁡p(⋅)-\sigma \nabla_z \log p(\cdot).

    Assume:

    1. The scaled score function ϵ(⋅)\epsilon(\cdot) is bounded: ∥ϵ(⋅)∥≤M\|\epsilon(\cdot)\| \le M for some constant M>0M > 0.
    2. The learned diffusion network ϵθ(⋅)\epsilon_\theta(\cdot) is KK-Lipschitz continuous.

    Then, for diagonal latents zdiag=[zτ11;… ;zτff]z^{\text{diag}} = [z^1_{\tau_1}; \dots; z^f_{\tau_f}] with timesteps τdiag=[τ1;… ;τf]\tau^{\text{diag}} = [\tau_1; \dots; \tau_f] (0<τ1<⋯<τf≤T0 < \tau_1 < \dots < \tau_f \le T), the error in the noise prediction for the ii-th frame compared to the ground-truth score satisfies:

    ∥ϵθ(zdiag,τdiag)i−ϵ(zτivdm,τi⋅1)i∥=∥ϵθ(zτivdm,τi⋅1)i−ϵ(zτivdm,τi⋅1)i∥+O(∣στf−στ1∣)\|\epsilon_\theta(z^{\text{diag}}, \tau^{\text{diag}})^i - \epsilon(z^{\text{vdm}}_{\tau_i}, \tau_i \cdot \mathbf{1})^i\| = \|\epsilon_\theta(z^{\text{vdm}}_{\tau_i}, \tau_i \cdot \mathbf{1})^i - \epsilon(z^{\text{vdm}}_{\tau_i}, \tau_i \cdot \mathbf{1})^i\| + \mathcal{O}(|\sigma_{\tau_f} - \sigma_{\tau_1}|)

    where ϵθ(⋅)i\epsilon_\theta(\cdot)^i and ϵ(⋅)i\epsilon(\cdot)^i denote the ii-th frame components. The discrepancy between diagonal denoising and standard uniform-timestep denoising is directly bounded by the maximum noise scale difference across the input latents ∣στf−στ1∣|\sigma_{\tau_f} - \sigma_{\tau_1}|, formally justifying the reduction of noise spreads via latent partitioning.

  5. Knowl 5 — FIFO-Diffusion Inference with Lookahead Denoising and Latent Partitioning

    algorithm

    The complete inference procedure generates an NN-frame video from a pretrained clip denoiser Φ(⋅)\Phi(\cdot), latent decoder Dec(⋅)\text{Dec}(\cdot), text conditioning cc, clip window size ff, and partition count nn (n≥2n \ge 2). The stride is set to f′=⌊f/2⌋f' = \lfloor f/2 \rfloor, and total diffusion steps per frame are S=nfS = n f with schedule τ=[τ1,…,τnf]\tau = [\tau_1, \dots, \tau_{nf}].

    Input: Target frames NN, clip size ff, partitions nn, denoiser Φ\Phi, decoder Dec\text{Dec}, schedule τ=[τ1,…,τnf]\tau = [\tau_1, \dots, \tau_{nf}], prompt cc, initial diagonal queue Q=[zτ11,…,zτnfnf]Q = [z^1_{\tau_1}, \dots, z^{nf}_{\tau_{nf}}]
    Output: Generated video frames vv
    f′←⌊f/2⌋f' \leftarrow \lfloor f / 2 \rfloor
    v←[]v \leftarrow []
    \tau_{\text{pad}} \leftarrow [\tau_1, \dots, \tau_1 \text{ (f' times)}, \tau_1, \dots, \tau_{nf}]
    Q \leftarrow [z^1_{\tau_1}, \dots, z^1_{\tau_1} \text{ (f' times)}, z^1_{\tau_1}, \dots, z^{nf}_{\tau_{nf}}]
    for i=1i = 1 to NN do
        zτ1i←Q[f′+1]z^i_{\tau_1} \leftarrow Q[f' + 1]
        for k=0k = 0 to 2n−12n - 1 do in parallel
            τk←τpad[kf′+1:kf′+2f′]\tau_k \leftarrow \tau_{\text{pad}}[k f' + 1 : k f' + 2 f']
            Qk←Q[kf′+1:kf′+2f′]Q_k \leftarrow Q[k f' + 1 : k f' + 2 f']
            Q~kf′+1:f←Φ(Qk,τk,c;ϵθ)[f′+1:f]\tilde{Q}_k^{f'+1:f} \leftarrow \Phi(Q_k, \tau_k, c; \epsilon_\theta)[f' + 1 : f]
        end for
        zτ0i←Q~0f′+1z^i_{\tau_0} \leftarrow \tilde{Q}_0^{f'+1}
        v.append(Dec(zτ0i))v.\text{append}(\text{Dec}(z^i_{\tau_0}))
        Q[f′+1]←zτ1iQ[f' + 1] \leftarrow z^i_{\tau_1}
        Q←[Q[1:f′],Q~0f′+1:f,…,Q~2n−1f′+1:f]Q \leftarrow [Q[1 : f'], \tilde{Q}_0^{f'+1:f}, \dots, \tilde{Q}_{2n-1}^{f'+1:f}]
        Q.dequeue()Q.\text{dequeue}()
        Sample zτnfi+nf∼N(0,I)z^{i+nf}_{\tau_{nf}} \sim \mathcal{N}(0, I)
        Q.enqueue(zτnfi+nf)Q.\text{enqueue}(z^{i+nf}_{\tau_{nf}})
    end for
    return vv

    Dummy prefix padding of size f′f' ensures valid lookahead reference context for the earliest partition blocks.

  6. Knowl 6 — Initial Diagonal Latent Construction Algorithm

    algorithm

    To start diagonal denoising without requiring pregenerated video clips or specialized training, the diagonal latent queue [zτ11;… ;zτff][z^1_{\tau_1}; \dots; z^f_{\tau_f}] is constructed from pure random Gaussian noise over ff warm-up steps.

    Input: Clip size ff, noise schedule {τi}i=0f\{\tau_i\}_{i=0}^f, prompt cc, denoiser Φ\Phi
    Output: Initial diagonal latents Q=[zτ11,…,zτff]Q = [z^1_{\tau_1}, \dots, z^f_{\tau_f}]
    Sample zτf1:f∼N(0,I)z^{1:f}_{\tau_f} \sim \mathcal{N}(0, I)
    Q←[zτf1,…,zτff]Q \leftarrow [z^1_{\tau_f}, \dots, z^f_{\tau_f}]
    τ←[τf,…,τf]\tau \leftarrow [\tau_f, \dots, \tau_f]
    for i=1i = 1 to ff do
        Q←Φ(Q,τ,c;ϵθ)Q \leftarrow \Phi(Q, \tau, c; \epsilon_\theta)
        Q.dequeue()Q.\text{dequeue}()
        Sample zτfi∼N(0,I)z^i_{\tau_f} \sim \mathcal{N}(0, I)
        Q.enqueue(zτfi)Q.\text{enqueue}(z^i_{\tau_f})
        τ←[τf−i,…,τf−i⏟f−i times,τf−i+1,…,τf⏟i times]\tau \leftarrow [\underbrace{\tau_{f-i}, \dots, \tau_{f-i}}_{f-i \text{ times}}, \underbrace{\tau_{f-i+1}, \dots, \tau_f}_{i \text{ times}}]
    end for
    return QQ

    Over ff iterations, the scheduled timesteps shift from uniform maximum noise [τf,…,τf][\tau_f, \dots, \tau_f] to the fully diagonal schedule [τ1,…,τf][\tau_1, \dots, \tau_f], producing the initial state needed for streaming video generation.

  7. Knowl 7 — Long Video Generation Benchmark on UCF-101

    data/table

    FIFO-Diffusion was quantitatively evaluated against long video generation baselines using the Latte diffusion transformer backbone on the UCF-101 dataset. Generation quality was measured by Fréchet Video Distance across 128 frames (FVD128\text{FVD}_{128}, lower is better) on 2,048 generated samples of length 128, and Inception Score (IS, higher is better) on randomly sampled 16-frame clips.

    Method FVD128\text{FVD}_{128} (↓\downarrow) IS (↑\uparrow)
    StyleGAN-V 1773.4 23.94±0.7323.94 \pm 0.73
    VIDM 1531.9 –
    PVDM-L (400-400s) 648.4 74.40±1.2574.40 \pm 1.25
    FIFO-Diffusion (ours) 596.64 74.44±1.17\textbf{74.44} \pm \textbf{1.17}

    FIFO-Diffusion (configured with n=4n=4 partitions and lookahead denoising over 64 DDPM inference steps) outperformed chunked autoregressive approaches such as PVDM-L (which required 400 diffusion steps conditioned on prior chunks), achieving a lower FVD128\text{FVD}_{128} score while operating completely training-free.

  8. Knowl 8 — Memory Scaling and Inference Speed across Video Lengths

    data/table

    Memory consumption (in megabytes) and inference latency per frame (in seconds per frame) were benchmarked using VideoCrafter2 as the baseline architecture with 64 inference steps on NVIDIA RTX A6000 GPUs.

    Method Memory usage [MB] (↓\downarrow) Inference time [s/frame] (↓\downarrow)
    128 frames 256 frames 512 frames
    FreeNoise 26,163 44,683 Out of Memory 6.09
    Gen-L-Video 10,913 10,937 10,965 22.07
    FIFO-Diffusion (1 GPU) 11,245 11,245 11,245 12.37
    FIFO-Diffusion (8 GPUs) 13,496 13,496 13,496 1.84

    FreeNoise exhibits linear memory growth with sequence length, leading to out-of-memory failure at 512 frames. FIFO-Diffusion maintains strict O(f)\mathcal{O}(f) constant memory usage regardless of video length (11,24511,245 MB on 1 GPU for 128 to 512+ frames). Running FIFO-Diffusion with n=4n=4 partitions across 8 GPUs reduces inference time from 12.37 s/frame to 1.84 s/frame through parallel block execution.

  9. Knowl 9 — Noise Prediction Error Reduction via Partitioning and Lookahead

    data/table

    The relative mean squared error (Relative MSE) measures the accuracy of noise prediction under diagonal denoising relative to standard uniform-noise video diffusion:

    Relative MSE=∥ϵθ(zdiag,τdiag)i−ϵ(zτivdm,τi⋅1)i∥2∥ϵθ(zτivdm,τi⋅1)i−ϵ(zτivdm,τi⋅1)i∥2\text{Relative MSE} = \frac{\|\epsilon_\theta(z^{\text{diag}}, \tau^{\text{diag}})^i - \epsilon(z^{\text{vdm}}_{\tau_i}, \tau_i \cdot \mathbf{1})^i\|^2}{\|\epsilon_\theta(z^{\text{vdm}}_{\tau_i}, \tau_i \cdot \mathbf{1})^i - \epsilon(z^{\text{vdm}}_{\tau_i}, \tau_i \cdot \mathbf{1})^i\|^2}

    Evaluations using VideoCrafter1 (f=16f=16, DDIM η=0.5\eta=0.5, 200 text prompts) demonstrate how latent partitioning (LP) and lookahead denoising (LD) progressively reduce prediction error:

    Configuration # of partitions (nn) without LD with LD
    without LP 1 1.09 1.01
    with LP 2 1.04 0.99
    with LP 4 1.02 0.98

    Increasing partition count nn systematically lowers relative MSE by reducing the input noise range. Combining n=4n=4 latent partitions with lookahead denoising achieves a relative MSE of 0.98 (below 1.0), indicating that conditioning noisier frames on cleaner preceding frames yields noise predictions superior to baseline uniform denoising.

  10. Knowl 10 — Multi-Prompt Video Generation with Continuous Sliding Context

    model/method

    FIFO-Diffusion supports long video generation conditioned on dynamic sequences of text prompts without abrupt visual scene cuts. Given a sequence of kk consecutive text prompts c1,c2,…,ckc_1, c_2, \dots, c_k and monotonic iteration thresholds 0=n0<n1<⋯<nk0 = n_0 < n_1 < \dots < n_k, the active conditioning prompt is switched at runtime such that prompt cjc_j is used during queue iterations i∈[nj−1+1,nj]i \in [n_{j-1} + 1, n_j].

    Because frames in the FIFO queue are updated continuously with a temporal stride of 1, the newly enqueued latents incorporate the new prompt condition while attending to preceding frames generated under the previous prompt. This produces smooth, semantically coherent scene and action transitions (e.g., an entity transitioning smoothly from walking to standing to resting) across a single continuous video stream.

  11. Knowl 11 — Distributional Input Mismatch in Training-Free Diagonal Denoising

    limitation

    A primary limitation of FIFO-Diffusion is the residual training-inference distribution shift. Because underlying video diffusion models are trained exclusively on latent tensors with spatially and temporally uniform noise levels (t1=t2=⋯=tft_1 = t_2 = \dots = t_f), passing input latents with varying diagonal noise levels induces an out-of-distribution condition.

    While latent partitioning narrows this noise gap and lookahead denoising provides cleaner reference frames, the fundamental mismatch in input distribution remains uneliminated in the training-free inference regime. Completely closing this gap requires fine-tuning or training base video diffusion architectures directly with diagonal noise schedules.

Coverage note — None was omitted; all key theoretical bounds, algorithms, mechanisms, quantitative experimental data tables, qualitative extensions, and limitations are fully represented.

References

  1. 1.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In CVPR, 2023.
  2. 2.Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, Chao Weng, and Ying Shan. VideoCrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023.
  3. 3.Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. VideoCrafter2: Overcoming data limitations for high-quality video diffusion models. arXiv preprint arXiv:2401.09047, 2024.
  4. 4.Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Seine: Short-to-long video diffusion model for generative transition and prediction. arXiv preprint arXiv:2310.20700, 2023.
  5. 5.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat GANs on image synthesis. In NeurIPS, 2021.
  6. 6.William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, and Frank Wood. Flexible diffusion modeling of long videos. In NeurIPS, 2022.
  7. 7.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221, 2022.
  8. 8.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020.
  9. 9.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models. In NeurIPS, 2022.
  10. 10.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. CogVideo: Large-scale pretraining for text-to-video generation via transformers. In ICLR, 2023.
  11. 11.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In NeurIPS, 2022.
  12. 12.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. In CVPR, 2023.
  13. 13.Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024.
  14. 14.Kangfu Mei and Vishal M. Patel. Vidm: Video implicit diffusion models. In AAAI, 2023.
  15. 15.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  16. 16.William Peebles and Saining Xie. Scalable diffusion models with transformers. In ICCV, 2023.
  17. 17.Haonan Qiu, Menghan Xia, Yong Zhang, Yingqing He, Xintao Wang, Ying Shan, and Ziwei Liu. FreeNoise: Tuning-free longer video diffusion via noise rescheduling. arXiv preprint arXiv:2310.15169, 2023.
  18. 18.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
  19. 19.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015.
  20. 20.David Ruhe, Jonathan Heek, Tim Salimans, and Emiel Hoogeboom. Rolling diffusion models. In ICML, 2024.
  21. 21.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In NeurIPS, 2016.
  22. 22.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-A-Video: Text-to-video generation without text-video data. In ICLR, 2022.
  23. 23.Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In CVPR, 2022.
  24. 24.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021.
  25. 25.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In ICLR, 2021.
  26. 26.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  27. 27.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.0171, 2018.
  28. 28.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description. In ICLR, 2023.
  29. 29.Vikram Voleti, Alexia Jolicoeur-Martineau, and Christopher Pal. MCVD: Masked conditional video diffusion for prediction, generation, and interpolation. In NeurIPS, 2022.
  30. 30.Fu-Yun Wang, Wenshuo Chen, Guanglu Song, Han-Jia Ye, Yu Liu, and Hongsheng Li. Gen-L-Video: Multi-text to long video generation via temporal co-denoising. arXiv preprint arXiv:2305.18264, 2023.
  31. 31.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. ModelScope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023.
  32. 32.Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, Yuwei Guo, Tianxing Wu, Chenyang Si, Yuming Jiang, Cunjian Chen, Chen Change Loy, Bo Dai, Dahua Lin, Yu Qiao, and Ziwei Liu. LaVie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023.
  33. 33.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. MSR-VTT: A large video description dataset for bridging video and language. In CVPR, 2016.
  34. 34.Shengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Gong Ming, Lijuan Wang, Zicheng Liu, Houqiang Li, and Nan Duan. NUWA-XL: Diffusion over diffusion for extremely long video generation. arXiv preprint arXiv:2303.12346, 2023.
  35. 35.Sihyun Yu, Kihyuk Sohn, Subin Kim, and Jinwoo Shin. Video probabilistic diffusion models in projected latent space. In CVPR, 2023.
  36. 36.Zihan Zhang, Richard Liu, Kfir Aberman, and Rana Hanocka. Tedi: Temporally-entangled diffusion for long-term motion synthesis. In SIGGRAPH, 2024.
  37. 37.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. MagicVideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022.

Citation

MLA
Kim, J., et al. “FIFO-Diffusion: Generating Infinite Videos from Text Without Training”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 89834–68, https://proceedings.neurips.cc/paper_files/paper/2024/file/a397986e0f34d4b1f0b640686ceaeff7-Paper-Conference.pdf.
APA
Kim, J., Kang, J., Choi, J., & Han, B. (2024). FIFO-Diffusion: Generating Infinite Videos from Text without Training. Advances in Neural Information Processing Systems, 37, 89834–89868. https://proceedings.neurips.cc/paper_files/paper/2024/file/a397986e0f34d4b1f0b640686ceaeff7-Paper-Conference.pdf
Chicago
Kim, J., J. Kang, J. Choi, and B. Han. 2024. “FIFO-Diffusion: Generating Infinite Videos from Text Without Training”. Advances in Neural Information Processing Systems 37: 89834–68. https://proceedings.neurips.cc/paper_files/paper/2024/file/a397986e0f34d4b1f0b640686ceaeff7-Paper-Conference.pdf.
Harvard
Kim, J. et al. (2024) “FIFO-Diffusion: Generating Infinite Videos from Text without Training”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 89834–89868. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/a397986e0f34d4b1f0b640686ceaeff7-Paper-Conference.pdf.
Vancouver
1. Kim J, Kang J, Choi J, Han B (2024) FIFO-Diffusion: Generating Infinite Videos from Text without Training. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 89834–89868

BibTeX

@inproceedings{kim2024fifo,
  title = {FIFO-Diffusion: Generating Infinite Videos from Text without Training},
  author = {Kim, Jihwan and Kang, Junoh and Choi, Jinyoung and Han, Bohyung},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {89834-89868},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/a397986e0f34d4b1f0b640686ceaeff7-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors