RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models

Ozgur KaraBariscan KurtkayaHidir YesiltepeJames M. RehgPinar Yanardag

article2024CVPR98 citations

Proposes a training-free video editing framework that uses randomized noise shuffling across diffusion steps to achieve temporally consistent, full-sequence edits with lower memory overhead and faster inference than prior methods.

Listen

Artificial intelligence tools have achieved significant success in generating and modifying still images from text descriptions, but translating these capabilities to video editing remains challenging. Existing video editing techniques frequently suffer from visual inconsistencies across frames, blurriness, severe memory limitations on longer sequences, and high computational costs from model retraining or extensive processing time. The article presents RAVE, a zero-shot, training-free framework designed to perform fast and temporally consistent video editing using off-the-shelf text-to-image diffusion models.

The framework organizes video frames into structured grids to process multiple frames as a unified image while incorporating spatial control tools for structural guidance. To overcome memory constraints that prevent processing entire videos in a single grid, RAVE randomly shuffles frame assignments across grids at each step of the iterative image generation process. This shuffling enforces global visual connections across all frames without demanding proportional GPU memory, effectively eliminating stylistic discrepancies between separate frame batches.

Empirical evaluations across 186 text-video pairs covering 8-, 36-, and 90-frame sequences demonstrate that RAVE outperforms existing state-of-the-art methods in overall edit quality and frame-to-frame consistency, with advantages becoming especially pronounced on longer videos. In a user study involving 130 participants, RAVE achieved the highest satisfaction, securing top ratings in 90.51% of evaluations for general editing, 82.82% for temporal consistency, and 86.67% for prompt alignment. Furthermore, RAVE executes edits approximately 25% faster than competing tools, processing a 90-frame video in 4 minutes and 28 seconds.

These findings indicate that high-quality, complex video modifications—including style transfers, local attribute modifications, and major shape changes—can be deployed economically without fine-tuning specialized models. Practitioners should consider adopting randomized noise shuffling architectures when low latency, limited hardware overhead, and long-sequence consistency are critical. However, minor visual artifacts and flickering can still arise during extreme shape edits due to approximation errors inherent in latent image reconstruction. Further development is recommended to extend the technique to related domains such as consistent avatar synthesis and three-dimensional texture editing.

arXiv: 2312.04524
Cover for RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models

Abstract

Recent advancements in diffusion-based models have demonstrated significant success in generating images from text. However, video editing models have not yet reached the same level of visual quality and user control. To address this, we introduce RAVE, a zero-shot video editing method that leverages pre-trained text-to-image diffusion models without additional training. RAVE takes an input video and a text prompt to produce high-quality videos while preserving the original motion and semantic structure. It employs a novel noise shuffling strategy, leveraging spatio-temporal interactions between frames, to produce temporally consistent videos faster than existing methods. It is also efficient in terms of memory requirements, allowing it to handle longer videos. RAVE is capable of a wide range of edits, from local attribute modifications to shape transformations. In order to demonstrate the versatility of RAVE, we create a comprehensive video evaluation dataset ranging from object-focused scenes to complex human activities like dancing and typing, and dynamic scenes featuring swimming fish and boats. Our qualitative and quantitative experiments highlight the effectiveness of RAVE in diverse video editing scenarios compared to existing methods. Our code, dataset

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. Preliminaries
  • 3.2. Our approach
  • 4. Experimentation
  • 4.1. Implementation details
  • 4.2. Evaluation
  • 5. Discussion and Conclusion
  • References

Knowls

  1. Knowl 1 — RAVE Framework for Zero-Shot Text-Guided Video Editing

    model/method

    RAVE is a training-free, zero-shot video editing framework that adapts pre-trained 2D text-to-image (T2I) latent diffusion models (such as Stable Diffusion) and spatial conditioning networks (such as ControlNet) to perform style, attribute, and shape editing on video sequences without requiring per-video optimization, text-video paired training, or user-defined masks.

    Given an input video VK={I1,I2,…,IK}\mathcal{V}_K = \{I_1, I_2, \dots, I_K\} of KK frames and a target text editing prompt P\mathcal{P}, the RAVE pipeline proceeds as follows:

    1. Preprocessing and Inversion: For each frame Ik∈RW×H×CI_k \in \mathbb{R}^{W \times H \times C} (where WW, HH, and CC denote width, height, and color channels), deterministic DDIM inversion maps the latent representation to inverted noise trajectories {ztk}t=1T\{z_t^k\}_{t=1}^T, where TT is the number of diffusion timesteps. Simultaneously, spatial condition maps ckc_k (e.g., depth maps, line art, or soft edge maps) are extracted from each frame using an off-the-shelf preprocessor.
    2. Grid Construction: Frame latents ztkz_t^k and conditions ckc_k are formatted into spatial multi-image grids of size N=n×mN = n \times m (nn rows and mm columns).
    3. Iterative Denoising with Noise Shuffling: At each reverse diffusion timestep tt from TT down to 11, the assignment of frames to grid cells across all L=K/NL = K / N grids is randomly shuffled. The 2D U-Net and ControlNet process these shuffled grids as single composite images.
    4. Reconstruction: After TT timesteps, the denoised latents are un-shuffled into their original sequential frame order 1,…,K1, \dots, K and decoded via the latent diffusion autoencoder decoder to yield the edited video VK∗={I1∗,…,IK∗}\mathcal{V}_K^* = \{I_1^*, \dots, I_K^*\}.
  2. Knowl 2 — Grid Partitioning and Randomized Frame Shuffling Formulation

    equation

    Let a video sequence be denoted as VK={I1,…,IK}\mathcal{V}_K = \{I_1, \dots, I_K\} with KK frames, where Ik∈RW×H×CI_k \in \mathbb{R}^{W \times H \times C} represents the kthk^{\text{th}} video frame with width WW, height HH, and CC color channels. Given a spatial grid dimension N=n×mN = n \times m (with nn rows and mm columns), frames or their latent representations are partitioned into an ordered set of L=K/NL = K / N composite grids (assuming KK is divisible by NN):

    video2grid(VK,N)={G1(1,…,N),…,GL(K−N+1,…,K)}=GL\text{video2grid}(\mathcal{V}_K, N) = \left\{ G_1^{(1, \dots, N)}, \dots, G_L^{(K - N + 1, \dots, K)} \right\} = \mathcal{G}_L

    where each composite grid Gl∈R(W⋅m)×(H⋅n)×CG_l \in \mathbb{R}^{(W \cdot m) \times (H \cdot n) \times C} concatenates NN individual frame images or latent maps into an n×mn \times m spatial layout.

    To prevent independent grids from following divergent sampling trajectories during reverse diffusion denoising, a randomized frame shuffling operation is applied at each timestep tt:

    shuffle(GL)={G1(i,…,j),…,GL(k,…,l)}\text{shuffle}(\mathcal{G}_L) = \left\{ G_1^{(i, \dots, j)}, \dots, G_L^{(k, \dots, l)} \right\}

    where (i,…,j),…,(k,…,l)(i, \dots, j), \dots, (k, \dots, l) are non-repeating index permutations uniformly sampled from the set {1,…,K}\{1, \dots, K\}. This operation randomly rearranges frame latents and their corresponding ControlNet conditioning maps into new grid assignments at each denoising step.

  3. Knowl 3 — Spatio-Temporal Consistency Mechanism of Grid Shuffling

    model/method

    When a 2D text-to-image (T2I) diffusion model processes a composite spatial grid G∈R(W⋅m)×(H⋅n)×CG \in \mathbb{R}^{(W \cdot m) \times (H \cdot n) \times C} containing N=n×mN = n \times m video frames, two internal structural modules enforce inter-frame consistency:

    1. Spatial Self-Attention Layers: Within each transformer block of the 2D U-Net, self-attention operates across all spatial positions in the composite grid. Consequently, patches from any frame in the grid attend to patches from all other frames within that grid, propagating global style, texture, and appearance attributes across frames.
    2. Convolutional Residual Blocks: 2D convolution filters operate across spatial boundaries between adjacent frames in the grid, smoothing transitions across latent regions and suppressing high-frequency pixel shifts that cause visual flickering.

    Processing a video by dividing it into fixed, independent grids causes each grid's sampling trajectory to diverge, leading to internal consistency within each grid but visible stylistic and temporal jumps across grid boundaries. By randomly re-permuting frame allocations into different grids at each reverse diffusion timestep t∈{T,…,1}t \in \{T, \dots, 1\}, each frame's latents interact with latents from throughout the entire video sequence over the course of TT denoising steps. This achieves global spatio-temporal coupling across the entire video without expanding the self-attention mechanism to full 3D attention or increasing memory requirements proportionally to video length.

  4. Knowl 4 — RAVE Video Editing Algorithm

    algorithm
    Input: Input video frames VK={I1,…,IK}\mathcal{V}_K = \{I_1, \dots, I_K\}, target prompt P\mathcal{P}, grid dimensions n×mn \times m yielding grid capacity N=n×mN = n \times m, total diffusion timesteps TT, pre-trained 2D U-Net ϵθ\epsilon_\theta, pre-trained autoencoder (E,D)(\mathcal{E}, \mathcal{D}), ControlNet C\mathcal{C}, classifier-free guidance scale ss.
    Output: Edited video frames VK∗={I1∗,…,IK∗}\mathcal{V}_K^* = \{I_1^*, \dots, I_K^*\}.
    // Step 1: Preprocessing and DDIM Inversion
    for k=1k = 1 to KK do
        z0k←E(Ik)z_0^k \leftarrow \mathcal{E}(I_k)
        ck←ExtractCondition(Ik)c_k \leftarrow \text{ExtractCondition}(I_k) // e.g. depth map, line art, or soft edge
        {ztk}t=1T←DDIM_Invert(z0k,Pnull,ϵθ)\{z_t^k\}_{t=1}^T \leftarrow \text{DDIM\_Invert}(z_0^k, \mathcal{P}_{\text{null}}, \epsilon_\theta)
    end for
    ZT←{zT1,…,zTK}\mathcal{Z}_T \leftarrow \{z_T^1, \dots, z_T^K\}
    Call←{c1,…,cK}\mathcal{C}_{\text{all}} \leftarrow \{c_1, \dots, c_K\}
    L←K/NL \leftarrow K / N
    // Step 2: Denoising with randomized noise shuffling
    for t=Tt = T down to 1 do
        π←RandomPermutation({1,…,K})\pi \leftarrow \text{RandomPermutation}(\{1, \dots, K\})
        GLt←video2grid(Zt[π],N)\mathcal{G}_L^t \leftarrow \text{video2grid}(\mathcal{Z}_t[\pi], N)
        CL←video2grid(Call[π],N)\mathcal{C}_L \leftarrow \text{video2grid}(\mathcal{C}_{\text{all}}[\pi], N)
        
        for l=1l = 1 to LL do
            G^lt−1←DenoiseStep(GLt[l],CL[l],P,t,ϵθ,C,s)\hat{G}_l^{t-1} \leftarrow \text{DenoiseStep}(\mathcal{G}_L^t[l], \mathcal{C}_L[l], \mathcal{P}, t, \epsilon_\theta, \mathcal{C}, s)
        end for
        
        Z~t−1←grid2video({G^1t−1,…,G^Lt−1})\tilde{\mathcal{Z}}_{t-1} \leftarrow \text{grid2video}(\{\hat{G}_1^{t-1}, \dots, \hat{G}_L^{t-1}\})
        Zt−1←Unshuffle(Z~t−1,π)\mathcal{Z}_{t-1} \leftarrow \text{Unshuffle}(\tilde{\mathcal{Z}}_{t-1}, \pi)
    end for
    // Step 3: Reconstruction
    for k=1k = 1 to KK do
        Ik∗←D(z0k)I_k^* \leftarrow \mathcal{D}(z_0^k)
    end for
    return VK∗={I1∗,…,IK∗}\mathcal{V}_K^* = \{I_1^*, \dots, I_K^*\}
  5. Knowl 5 — Evaluation Dataset and Experimental Benchmark Setup

    experimental setup

    To evaluate zero-shot text-guided video editing across diverse motions and scene geometries, a benchmark dataset of 31 video clips across three sequence lengths is used:

    • 10 clips of length 8 frames,
    • 15 clips of length 36 frames,
    • 6 clips of length 90 frames.

    The videos encompass object-centric scenes, complex human actions (such as dancing, stretching, and typing), and dynamic natural environments (such as swimming animals and moving vehicles). Each video is edited with 4 style prompts and 2 shape prompts, yielding 186 text-video evaluation pairs.

    Implementation Setup:

    • Base diffusion model: Stable Diffusion 1.5 (with Realistic Vision V5.1 for qualitative evaluations).
    • Inversion and Sampling: 50 DDIM steps with a fixed classifier-free guidance scale of 7.5; no positive or negative prompt filtering is used.
    • Spatial conditioning: Depth-conditioned ControlNet with grid size 2×22 \times 2 for 8-frame clips and 3×33 \times 3 for 36- and 90-frame clips.
    • Hardware: A single NVIDIA A40 GPU.

    Evaluation Metrics:

    • CLIP-F: Coarse temporal consistency measured as the average pairwise cosine similarity of CLIP image embeddings across consecutive frames.
    • WarpSSIM: Structural and temporal consistency calculated as the average SSIM between the edited video frame Ik∗I_k^* and the optical-flow-warped edited frame Warp(Ik−1∗)\text{Warp}(I_{k-1}^*), where optical flow is estimated on the source video using RAFT.
    • CLIP-T: Textual alignment measured as the average cosine similarity between the CLIP text embedding of prompt P\mathcal{P} and CLIP image embeddings of all edited frames.
    • QeditQ_{\text{edit}}: A composite editing quality metric defined as Qedit=WarpSSIM⋅CLIP-TQ_{\text{edit}} = \text{WarpSSIM} \cdot \text{CLIP-T}.
  6. Knowl 6 — Quantitative Evaluation and Runtime Comparison

    data/table

    Quantitative performance and runtime comparison across video sequence lengths of 8, 36, and 90 frames on a single NVIDIA A40 GPU:

    Method CLIP-F (×10−2\times 10^{-2}) ↑\uparrow WarpSSIM (×10−2\times 10^{-2}) ↑\uparrow CLIP-T (×10−2\times 10^{-2}) ↑\uparrow QeditQ_{\text{edit}} (×10−5\times 10^{-5}) ↑\uparrow
    8-fr 36-fr 90-fr 8-fr 36-fr 90-fr 8-fr 36-fr 90-fr 8-fr 36-fr 90-fr
    Text2Video-Zero 95.49 92.89 94.35 67.97 36.65 71.57 29.46 29.42 29.73 20.02 10.78 21.27
    Rerender 92.87 89.71 90.63 68.57 44.54 74.56 25.65 27.42 27.55 17.66 12.24 20.51
    TokenFlow 95.80 93.17 95.92 74.03 50.97 80.40 28.27 28.29 29.53 20.92 14.41 23.74
    Pix2Video 89.96 – – 24.78 – – 28.01 – – 5.61 – –
    RAVE (w/o shuffle) 93.98 89.90 92.49 71.78 47.26 76.58 28.78 29.49 29.71 20.66 13.94 22.76
    RAVE 95.95 93.18 95.99 71.44 48.81 80.51 29.51 29.93 29.76 21.08 14.60 23.95

    Pipeline Runtime for 90-Frame Videos (minutes:seconds):

    • Text2Video-Zero: 5:33
    • Rerender: 5:24
    • TokenFlow: 5:24 total (4:14 excluding preprocessing)
    • Pix2Video: Out-of-memory on GPU (fails on >45>45 frames; FateZero fails on >22>22 frames)
    • RAVE: 4:28 total (3:13 excluding preprocessing)

    RAVE achieves the top composite score QeditQ_{\text{edit}} across all frame lengths, achieves highest CLIP-F consistency across all lengths, surpasses all baselines in WarpSSIM on 90-frame videos (80.51×10−280.51 \times 10^{-2}), and executes ∼25%\sim 25\% faster than TokenFlow.

  7. Knowl 7 — User Preference Study on Video Editing Quality

    empirical result

    A user study conducted with 130 participants on the Prolific platform evaluated 23 randomly selected video-text pairs across three criteria, measuring the frequency with which each method was ranked in the top two choices:

    • General Editing (GE, Q1): Overall visual edit quality and prompt adherence.
      • RAVE: 90.51%
      • Text2Video-Zero: 47.95%
      • TokenFlow: 44.10%
      • Rerender: 17.44%
    • Temporal Consistency (TC, Q2): Smoothness and stability across frames without flickering or abrupt jumps.
      • RAVE: 82.82%
      • TokenFlow: 68.97%
      • Text2Video-Zero: 24.87%
      • Rerender: 23.33%
    • Textual Alignment (TA, Q3): Faithfulness of the edited video to the target text prompt.
      • RAVE: 86.67%
      • Text2Video-Zero: 52.56%
      • TokenFlow: 43.59%
      • Rerender: 17.18%

    RAVE achieved the highest preference rate across all three evaluation questions, with TokenFlow ranking second in temporal consistency.

  8. Knowl 8 — Ablation Analysis on Noise Shuffling, DDIM Inversion, and Conditioning Types

    empirical result

    Ablations demonstrate the necessity of each core component in RAVE:

    • Noise Shuffling: Removing shuffling (RAVE w/o shuffle) results in severe style drift between grids as the video length increases. On 36-frame videos, CLIP-F falls from 93.18×10−293.18 \times 10^{-2} to 89.90×10−289.90 \times 10^{-2} and WarpSSIM drops from 48.81×10−248.81 \times 10^{-2} to 47.26×10−247.26 \times 10^{-2}. On 90-frame videos, CLIP-F drops from 95.99×10−295.99 \times 10^{-2} to 92.49×10−292.49 \times 10^{-2} and WarpSSIM drops from 80.51×10−280.51 \times 10^{-2} to 76.58×10−276.58 \times 10^{-2}.
    • DDIM Inversion: Omitting DDIM inversion leads to a loss of the original video structure, motion trajectory, and spatial layout.
    • Conditioning Modality Flexibility: Replacing the default depth condition in ControlNet with line art or soft edge conditions preserves global and temporal consistency across frames, demonstrating that RAVE operates robustly across multiple structural guidance modalities.
  9. Knowl 9 — Limitations in Extreme Shape Editing and Inversion Error Accumulation

    limitation

    RAVE exhibits visual flickering and artifacts under two specific scenarios:

    1. Extreme Shape Transformations: Large topological or geometric modifications away from the source video structure can create conflicting signals between the spatial ControlNet condition and the text prompt.
    2. High-Frequency Local Attribute Edits: DDIM inversion relies on numerical approximations that accumulate error when combined with classifier-free guidance, and the autoencoder compression stage of Latent Diffusion Models introduces reconstruction distortions in high-frequency detail areas. These factors lead to localized flickering in regions with fine spatial patterns or rapid motion.

Coverage note — None was omitted; all key contributions including framework design, mathematical grid formulation, consistency mechanisms, full algorithm, benchmark dataset, quantitative data, user study, ablations, and stated limitations are fully represented.

References

  1. 1.Civitai. https://civitai.com/. Accessed: 2023-11-16. 6
  2. 2.Gridtrick. https : / / web . archive . org / web / 20231025170948 / https : / / semicolon . dev / midjourney / how - to - make - consistent - characters. Archived: 2023-10-25. 4
  3. 3.Pexels. https://www.pexels.com/. Accessed: 2023-11-16. 1, 2
  4. 4.Pixabay. https://pixabay.com/. Accessed: 2023-11-16. 1, 2
  5. 5.Prolific. https://www.prolific.com/. Accessed: 2023-11-18. 8
  6. 6.Omri Avrahami, Dani Lischinski, and Ohad Fried. Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18208–18218, 2022. 2, 3
  7. 7.Omri Avrahami, Ohad Fried, and Dani Lischinski. Blended latent diffusion. ACM Transactions on Graphics (TOG), 42 (4):1–11, 2023. 2, 3
  8. 8.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, pages 1728–1738, 2021. 2
  9. 9.Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Text2live: Text-driven layered image and video editing. In European conference on computer vision, pages 707–723. Springer, 2022. 2
  10. 10.Duygu Ceylan, Chun-Hao P Huang, and Niloy J Mitra. Pix2video: Video editing using image diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 23206–23217, 2023. 2, 3, 5, 7
  11. 11.Yuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen, Jiawei Ren, Yanping Xie, Juan-Manuel Perez-Rua, Bodo Rosenhahn, Tao Xiang, and Sen He. Flatten: optical flow-guided attention for consistent text-to-video editing. arXiv preprint arXiv:2310.05922, 2023. 7, 1, 2
  12. 12.Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based semantic image editing with mask guidance. In The Eleventh International Conference on Learning Representations, 2023. 2, 3
  13. 13.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. An image is worth one word: Personalizing text-to-image generation using textual inversion. In The Eleventh International Conference on Learning Representations, 2023. 3
  14. 14.Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing. arXiv preprint arXiv:2307.10373, 2023. 2, 3, 5, 7
  15. 15.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. Prompt-to-prompt image editing with cross-attention control. In The Eleventh International Conference on Learning Representations, 2023. 2, 3, 6
  16. 16.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 2, 4
  17. 17.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 2
  18. 18.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In Advances in Neural Information Processing Systems, pages 8633–8646. Curran Associates, Inc., 2022. 2, 3
  19. 19.Susung Hong, Gyuseong Lee, Wooseok Jang, and Seungryong Kim. Improving sample quality of diffusion models using self-attention guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7462–7471, 2023. 3
  20. 20.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In The Eleventh International Conference on Learning Representations, 2023. 2
  21. 21.Ondrej Jamriska. Ebsynth: Fast example-based image synthesis and style transfer, 2018. 4
  22. 22.Ondřej Jamriška, Šárka Sochorová, Ondřej Texler, Michal Lukáč, Jakub Fišer, Jingwan Lu, Eli Shechtman, and Daniel Sýkora. Stylizing video by example. ACM Transactions on Graphics (TOG), 38(4):1–11, 2019. 2
  23. 23.Yoni Kasten, Dolev Ofri, Oliver Wang, and Tali Dekel. Layered neural atlases for consistent video editing. ACM Transactions on Graphics (TOG), 40(6):1–12, 2021. 2
  24. 24.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6007–6017, 2023. 3
  25. 25.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15954–15964, 2023. 2, 3, 5
  26. 26.Yao-Chih Lee, Ji-Ze Genevieve Jang, Yi-Ting Chen, Elizabeth Qiu, and Jia-Bin Huang. Shape-aware text-driven layered video editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14317–14326, 2023. 2, 3, 1
  27. 27.Jun Hao Liew, Hanshu Yan, Jianfeng Zhang, Zhongcong Xu, and Jiashi Feng. Magicedit: High-fidelity and temporally coherent video editing. arXiv preprint arXiv:2308.14749, 2023. 2, 3
  28. 28.Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, and Jiaya Jia. Video-p2p: Video editing with cross-attention control. arXiv preprint arXiv:2303.04761, 2023. 3
  29. 29.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. Sdedit: Guided image synthesis and editing with stochastic differential equations. In ICLR, 2021. 2
  30. 30.Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6038–6047, 2023. 3
  31. 31.Eyal Molad, Eliahu Horwitz, Dani Valevski, Alex Rav Acha, Yossi Matias, Yael Pritch, Yaniv Leviathan, and Yedid Hoshen. Dreamix: Video diffusion models are general video editors. arXiv preprint arXiv:2302.01329, 2023. 3
  32. 32.Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu. Zero-shot image-to-image translation. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–11, 2023. 2
  33. 33.Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 724–732, 2016. 1, 2
  34. 34.Chenyang QI, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. Fatezero: Fusing attentions for zero-shot text-based video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15932–15942, 2023. 2, 3, 5, 1
  35. 35.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 7
  36. 36.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2
  37. 37.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 4
  38. 38.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015. 4
  39. 39.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510, 2023. 2, 3
  40. 40.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022. 2, 3
  41. 41.Chaehun Shin, Heeseung Kim, Che Hyun Lee, Sang-gil Lee, and Sungroh Yoon. Edit-a-video: Single video editing with object-aware consistency. arXiv preprint arXiv:2303.07945, 2023. 3
  42. 42.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. In The Eleventh International Conference on Learning Representations, 2023. 2, 3
  43. 43.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256–2265. PMLR, 2015. 2
  44. 44.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021. 2, 4
  45. 45.Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pages 402–419. Springer, 2020. 7
  46. 46.Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1921–1930, 2023. 2, 6
  47. 47.Dani Valevski, Matan Kalman, Eyal Molad, Eyal Segalis, Yossi Matias, and Yaniv Leviathan. Unitune: Text-driven image editing by fine tuning a diffusion model on a single image. ACM Transactions on Graphics (TOG), 42(4):1–10, 2023. 3
  48. 48.Wen Wang, Kangyang Xie, Zide Liu, Hao Chen, Yue Cao, Xinlong Wang, and Chunhua Shen. Zero-shot video editing using off-the-shelf image diffusion models. arXiv preprint arXiv:2303.17599, 2023. 3
  49. 49.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 7
  50. 50.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7623–7633, 2023. 2, 3
  51. 51.Shuai Yang, Yifan Zhou, Ziwei Liu, , and Chen Change Loy. Rerender a video: Zero-shot text-guided video-to-video translation. In ACM SIGGRAPH Asia Conference Proceedings, 2023. 2, 3, 5, 7
  52. 52.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 2, 3, 4

Citation

MLA
Kara, O., et al. “RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2312.04524v1.
APA
Kara, O., Kurtkaya, B., Yesiltepe, H., Rehg, J. M., & Yanardag, P. (2023). RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models. arXiv. http://arxiv.org/abs/2312.04524v1
Chicago
Kara, O., B. Kurtkaya, H. Yesiltepe, J. M. Rehg, and P. Yanardag. 2023. “RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models”. arXiv. http://arxiv.org/abs/2312.04524v1.
Harvard
Kara, O. et al. (2023) “RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.04524v1.
Vancouver
1. Kara O, Kurtkaya B, Yesiltepe H, Rehg JM, Yanardag P (2023) RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models. arXiv

BibTeX

@article{kara2023rave,
  title = {RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models},
  author = {Kara, Ozgur and Kurtkaya, Bariscan and Yesiltepe, Hidir and Rehg, James M. and Yanardag, Pinar},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.04524v1},
  eprint = {2312.04524}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE