VideoBooth: Diffusion-based Video Generation with Image Prompts

Yuming JiangTianxing WuShuai YangChenyang SiDahua LinYu QiaoChen Change LoyZiwei Liu

article2024CVPR155 citations

Proposes VideoBooth, a tuning-free feed-forward diffusion framework that generates customized, temporally consistent videos from image prompts by combining coarse semantic embeddings with multi-scale attention injection.

Listen

Text-driven video generation models have advanced rapidly, yet text descriptions alone often fail to capture specific visual appearances and fine details needed for customized content creation. Existing personalization methods typically rely on fine-tuning model parameters at inference time or requiring multiple reference images, both of which introduce high computational costs, deployment delays, and operational friction.

The article introduces VideoBooth, a framework designed to generate high-quality, temporally consistent videos conditioned on both a text prompt and a single reference image prompt without requiring any fine-tuning during inference.

To achieve this, the authors developed a coarse-to-fine visual embedding architecture built on latent video diffusion models. At the coarse level, an image encoder extracts high-level semantic features from the image prompt and maps them into text embedding space. At the fine level, multi-scale latent representations of the image prompt are injected into the model's cross-frame attention layers as additional keys and values. This attention injection directly refines spatial details in the initial frame and propagates them across subsequent frames to preserve temporal consistency. The framework is trained sequentially, optimizing the coarse encoder first to prevent feature leakage before training the fine attention module. The authors also established a dedicated dataset derived from WebVid, filtering down to 48,724 training pairs of video clips, text prompts, and segmented subject image prompts, along with a 650-pair benchmark for testing.

The experimental evaluation demonstrated key performance findings:

  • Superior subject fidelity: VideoBooth achieved state-of-the-art visual alignment scores, recording a 74.80 CLIP-Image score and a 65.10 DINO score, substantially outperforming existing customized generation baselines such as Textual Inversion, DreamBooth, and ELITE.
  • Preserved textual alignment: VideoBooth maintained strong prompt fidelity with a CLIP-Text score of 30.10, performing comparably to established alternatives while properly balancing text and image instructions.
  • Strong user preference: In a user study with 25 participants across multiple test scenarios, VideoBooth secured the highest user preference rates in image alignment, text alignment, and overall visual quality.
  • Critical ablation insights: Removing coarse embeddings caused temporal degradation and object distortion across later frames, omitting fine injection caused lost visual patterns, and training both components simultaneously degraded overall encoder performance.

These results show that high-fidelity subject customization in generative video can be accomplished through feed-forward inference alone. By eliminating test-time fine-tuning, the framework lowers inference latency and computational expenses, making customized video generation more viable for scalable commercial pipelines. Furthermore, the two-stage coarse-to-fine injection strategy solves the trade-off between retaining fine-grained spatial attributes and preserving fluid, consistent temporal motion.

Organizations evaluating this technology should adopt feed-forward, multi-scale visual embedding architectures to streamline customized video production workflows. Future research and development should focus on expanding the dataset with automated 3D image augmentation pipelines to enable diverse multi-angle viewpoints, and integrating deepfake detection safeguards to address synthetic media risks.

The framework's primary technical limitation is its reliance on training image prompts that closely align with the viewpoints of the target videos, meaning it cannot reliably synthesize extreme perspective shifts (such as generating a front-facing video from a rear-view prompt). Confidence in the reported image fidelity and temporal consistency gains is high, backed by comprehensive quantitative metrics, ablation studies, and qualitative user evaluations.

arXiv: 2312.00777
Cover for VideoBooth: Diffusion-based Video Generation with Image Prompts

Abstract

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized content creation. In this paper, we study the task of video generation with image prompts, which provide more accurate and direct content control beyond the text prompts. Specifically, we propose a feed-forward framework VideoBooth, with two dedicated designs: 1) We propose to embed image prompts in a coarse-to-fine manner. Coarse visual embeddings from image encoder provide high-level encodings of image prompts, while fine visual embeddings from the proposed attention injection module provide multi-scale and detailed encoding of image prompts. These two complementary embeddings can faithfully capture the desired appearance. 2) In the attention injection module at fine level, multi-scale image prompts are fed into different cross-frame attention layers as additional keys and values. This extra spatial in-formation refines the details in the first frame and then it is propagated to the remaining frames, which maintains temporal consistency. Extensive experiments demonstrate that VideoBooth achieves state-of-the-art performance in generating customized high-quality videos with subjects specified in image prompts. Notably, VideoBooth is a generalizable framework where a single model works for a wide range of image prompts with only feed-forward passes.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. VideoBooth
  • 3.1. Preliminary: Pretrained Text-to-Video Model
  • 3.2. Coarse Visual Embeddings via Image Encoder
  • 3.3. Fine Visual Embeddings via Attention Injection
  • 3.4. Coarse-to-Fine Training Strategy
  • 4. VideoBooth Dataset
  • 5. Experiments
  • 5.1. Comparison Methods
  • 5.2. Evaluation Metrics
  • 5.3. Quantitative Comparisons
  • 5.4. Qualitative Comparisons
  • 5.5. Ablation Study
  • 6. Discussion
  • References

Knowls

  1. Knowl 1 — VideoBooth Framework Architecture

    model/method

    VideoBooth is a feed-forward, inference-tuning-free text-to-video generation framework conditioned on a text prompt TT and a single image prompt II specifying subject appearance. The framework integrates visual prompt information through a coarse-to-fine dual-pathway architecture built on top of a latent video diffusion model:

    1. Coarse Visual Embedding Pathway: The image prompt II is processed by a frozen CLIP image encoder and a multi-layer perceptron (MLP) mapping network to produce a visual embedding vector fIf_I that replaces the subject token embeddings within the text conditioning vector cTc_T, which is subsequently fed into the cross-attention layers of the U-Net denoiser.
    2. Fine Visual Embedding Pathway: Multi-scale spatial latent representations of II, with timestep-matched noise added, are injected directly into the cross-frame attention layers of the U-Net. These features serve as additional keys and values (KI,VIK_I, V_I) via dedicated projection layers (Kimg,VimgK^{img}, V^{img}) to update the first frame, whose updated features are subsequently propagated across subsequent frames to preserve visual fidelity and temporal consistency across video frames.
  2. Knowl 2 — Coarse Visual Feature Embedding via Subject Token Replacement

    equation

    To condition the video diffusion model on the coarse semantics of the image prompt II, visual features are extracted and projected into the text embedding space to replace the word embeddings corresponding to the target subject in text prompt TT.

    First, visual feature vector fVf_V is extracted from image prompt II and mapped through an MLP network F(⋅)F(\cdot) to obtain visual embedding fIf_I: fV=CLIPI(I),fI=F(fV)f_V = \text{CLIP}_I(I), \quad f_I = F(f_V) where CLIPI\text{CLIP}_I is a pretrained CLIP vision transformer encoder and F:Rdv→RdtF: \mathbb{R}^{d_v} \to \mathbb{R}^{d_t} maps the visual embedding dimension dvd_v into the text token embedding dimension dtd_t.

    Given the text prompt TT, the text token embeddings fT=[fT0,fT1,…,fTk,… ]f_T = [f_T^0, f_T^1, \dots, f_T^k, \dots] are extracted by the CLIP text encoder. The final condition sequence cTc_T is constructed by substituting the nn consecutive token embeddings starting at index kk (which correspond to the target subject) with fIf_I: cT=[fT0,fT1,…,fTk−1,fI,fTk+n,… ]c_T = [f_T^0, f_T^1, \dots, f_T^{k-1}, f_I, f_T^{k+n}, \dots] This fused condition sequence cTc_T is passed to the cross-attention modules of the video diffusion U-Net.

  3. Knowl 3 — Fine-Grained Multi-Scale Attention Injection Mechanism

    model/method

    To preserve high-resolution spatial details without relying solely on the flattened coarse embedding fIf_I, VideoBooth injects noisy latent representations of the image prompt II into the cross-frame attention layers of the video diffusion model.

    Let x0Ix_0^I be the latent representation of II obtained via the Stable Diffusion VAE encoder. At diffusion timestep tt, forward diffusion noise is added to align the latent distribution: xtI=α‾tx0I+1−α‾tϵ,ϵ∼N(0,I)x_t^I = \sqrt{\overline{\alpha}_t} x_0^I + \sqrt{1 - \overline{\alpha}_t} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I) where α‾t=∏s=1t(1−βs)\overline{\alpha}_t = \prod_{s=1}^t (1 - \beta_s) denotes the cumulative noise schedule parameter.

    In each cross-frame attention layer, keys KIK_I and values VIV_I are derived from xtIx_t^I using dedicated projection matrices KimgK^{img} and VimgV^{img} (initialized from the original projection layers). The feature value update for the first frame (frame 00) is computed as: V0new=softmax(Q0KTd)V,K=[KI,K0],V=[VI,V0]V_0^{new} = \text{softmax}\left(\frac{Q_0 K^T}{\sqrt{d}}\right) V, \quad K = [K_I, K_0], \quad V = [V_I, V_0] where Q0,K0,V0Q_0, K_0, V_0 denote the query, key, and value matrices of frame 00, and [⋅,⋅][\cdot, \cdot] denotes concatenation along the sequence dimension.

    For each remaining frame i>0i > 0, the updated first-frame value V0newV_0^{new} is propagated: Vinew=softmax(QiKTd)V,K=[K0,Ki−1],V=[V0new,Vi−1]V_i^{new} = \text{softmax}\left(\frac{Q_i K^T}{\sqrt{d}}\right) V, \quad K = [K_0, K_{i-1}], \quad V = [V_0^{new}, V_{i-1}] This operation is executed across multiple cross-frame attention layers in different stages of the U-Net by providing xtIx_t^I at matching spatial resolutions.

  4. Knowl 4 — Two-Stage Coarse-to-Fine Training Strategy

    model/method

    VideoBooth uses a sequential two-stage training strategy to prevent representation collapse:

    • Stage 1 (Coarse Embedding Training): The CLIP image encoder is frozen. The MLP mapping layers F(⋅)F(\cdot) and the key/value linear projection weights in the cross-attention modules are trained to generate videos containing the subject category and general coarse attributes specified by fIf_I.
    • Stage 2 (Fine Attention Injection Training): With the coarse image encoder weights fixed, the attention injection modules—specifically the dedicated linear projection weights KimgK^{img} and VimgV^{img} for the image prompt in cross-frame attention layers—are trained.

    If both modules are trained jointly in a single stage (unified training), the direct connection of spatial tokens in attention injection creates a shortcut: the model relies exclusively on fine attention injection, causing the coarse image encoder to learn uninformative representations. During inference sampling, the absence of coarse semantic grounding leads to severe visual distortion in subsequent video frames.

  5. Knowl 5 — VideoBooth Dataset Construction Pipeline

    experimental setup

    To train and benchmark image-prompted text-to-video generation, the VideoBooth dataset was curated from WebVid:

    1. Subject Segmentation: Grounded-SAM receives noun chunks extracted from the video's original text caption via the spaCy NLP library and segments the corresponding subject from the first frame (t=0t=0) of each video to generate clean-background image prompts.
    2. Filtering by Object Scale: Objects occupying an overly small bounding ratio or an excessively large bounding ratio (approaching the full frame area) relative to the video frame are discarded.
    3. Filtering by Dynamic Classes: Videos are filtered to retain dynamic moving objects matching specific subject keywords: dog, cat, bear, car, panda, tiger, horse, elephant, and lion.

    From a 2.5M subset of WebVid, this filtering yields 48,724 paired training samples (text prompt, subject image prompt, target video). A dedicated evaluation benchmark of 650 non-overlapping test pairs was curated from WebVid-10M.

  6. Knowl 6 — Quantitative Evaluation of VideoBooth vs. Customization Baselines

    data/table

    VideoBooth was evaluated on 650 test pairs against adapted versions of text-to-image personalization methods (Textual Inversion, DreamBooth, and ELITE) using three metrics: CLIP-Text (cosine similarity between CLIP text embeddings of prompts and frame image embeddings), CLIP-Image (cosine similarity between CLIP image embeddings of image prompt and generated frames), and DINO (cosine similarity using ViT-S/16 self-supervised features to assess fine-grained subject identity).

    Method CLIP-Text ↑\uparrow CLIP-Image ↑\uparrow DINO ↑\uparrow
    Textual Inversion 29.9749 69.7995 45.3143
    DreamBooth 30.6877 71.2078 52.9661
    ELITE 30.0881 73.7518 58.9522
    VideoBooth (Ours) 30.0967 74.7971 65.0979

    VideoBooth outperforms baseline approaches on CLIP-Image (74.7971) and DINO (65.0979, an improvement of +6.1457 over ELITE), while maintaining competitive CLIP-Text alignment (30.0967). DreamBooth achieves a higher CLIP-Text score (30.6877) because its modifier token S∗S^* is appended rather than replacing word tokens, which can cause the model to ignore the visual reference in favor of generic text concepts.

  7. Knowl 7 — Ablation Study on Visual Embedding Modules and Training Strategy

    data/table

    An ablation study evaluated the contributions of the coarse visual embedding, fine visual embedding, and sequential training pipeline on a validation subset.

    Variants CLIP-Image ↑\uparrow DINO ↑\uparrow
    (a) Coarse Embeddings only 75.4366 64.9568
    (b) Fine Embeddings only 75.5553 66.0378
    (c) Unified Training 75.8254 67.4201
    (d) Full Model 76.1631 69.7374
    • Coarse Embeddings only: Captures high-level category and color but misses fine spatial patterns, yielding lower DINO (64.9568).
    • Fine Embeddings only: Matches the first frame accurately but lacks semantic guidance in subsequent frames, leading to distorted motion and lower DINO (66.0378).
    • Unified Training: Joint training degrades the coarse image encoder, causing over-reliance on the first frame and temporal degradation (67.4201 DINO).
    • Full Model: Sequential coarse-to-fine training with both coarse and fine visual embeddings achieves the highest identity preservation (76.1631 CLIP-Image, 69.7374 DINO).
  8. Knowl 8 — Viewpoint Diversity Limitation in VideoBooth

    limitation

    Because the training image prompts are extracted directly from the initial frames of target video sequences, the training distribution lacks large out-of-plane viewpoint variations between the prompt and the video subject. Consequently, if an input image prompt displays a subject exclusively from one viewpoint (such as a rear view), VideoBooth cannot synthesize videos depicting unseen viewpoints (such as a frontal view) with accurate identity preservation.

Coverage note — Omitted qualitative visual descriptions from the figures and demographic specifics of the 25-person user study, as these are qualitative summaries that reinforce the quantitative tables.

References

  1. 1.Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin. Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv preprint arXiv:2304.08477, 2023.
  2. 2.Moab Arar, Rinon Gal, Yuval Atzmon, Gal Chechik, Daniel Cohen-Or, Ariel Shamir, and Amit H Bermano. Domainagnostic tuning-encoder for fast personalization of text-toimage models. arXiv preprint arXiv:2307.06925, 2023.
  3. 3.Omri Avrahami, Kfir Aberman, Ohad Fried, Daniel CohenOr, and Dani Lischinski. Break-a-scene: Extracting multiple concepts from a single image. arXiv preprint arXiv:2305.16311, 2023.
  4. 4.Jinbin Bai, Zhen Dong, Aosong Feng, Xiao Zhang, Tian Ye, Kaicheng Zhou, and Mike Zheng Shou. Integrating view conditions for image synthesis. arXiv preprint arXiv:2310.16002, 2023.
  5. 5.Max Bain, Arsha Nagrani, Gõl Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In IEEE International Conference on Computer Vision, 2021.
  6. 6.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023.
  7. 7.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021.
  8. 8.Duygu Ceylan, Chun-Hao Huang, and Niloy J. Mitra. Pix2video: Video editing using image diffusion. arXiv:2303.12688, 2023.
  9. 9.Wenhao Chai, Xun Guo, Gaoang Wang, and Yan Lu. Stablevideo: Text-driven consistency-aware diffusion video editing. arXiv preprint arXiv:2308.09592, 2023.
  10. 10.Hong Chen, Xin Wang, Guanning Zeng, Yipeng Zhang, Yuwei Zhou, Feilin Han, and Wenwu Zhu. Videodreamer: Customized multi-subject text-to-video generation with disen-mix finetuning. arXiv preprint arXiv:2311.00990, 2023.
  11. 11.Hong Chen, Yipeng Zhang, Xin Wang, Xuguang Duan, Yuwei Zhou, and Wenwu Zhu. Disenbooth: Disentangled parameter-efficient tuning for subject-driven text-to-image generation. arXiv preprint arXiv:2305.03374, 2023.
  12. 12.Li Chen, Mengyi Zhao, Yiheng Liu, Mingxu Ding, Yangyang Song, Shizun Wang, Xu Wang, Hao Yang, Jing Liu, Kang Du, et al. Photoverse: Tuning-free image customization with text-to-image diffusion models. arXiv preprint arXiv:2309.05793, 2023.
  13. 13.Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, and Hengshuang Zhao. Anydoor: Zero-shot object-level image customization. arXiv preprint arXiv:2307.09481, 2023.
  14. 14.Zhuowei Chen, Shancheng Fang, Wei Liu, Qian He, Mengqi Huang, Yongdong Zhang, and Zhendong Mao. Dreamidentity: Improved editability for efficient face-identity preserved image generation. arXiv preprint arXiv:2307.00300, 2023.
  15. 15.Jooyoung Choi, Yunjey Choi, Yunji Kim, Junho Kim, and Sungroh Yoon. Custom-edit: Text-guided image editing with customized diffusion models. arXiv preprint arXiv:2305.15779, 2023.
  16. 16.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems, 34:19822–19835, 2021.
  17. 17.Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. Cogview2: Faster and better text-to-image generation via hierarchical transformers. arXiv preprint arXiv:2204.14217, 2022.
  18. 18.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. arXiv preprint arXiv:2302.03011, 2023.
  19. 19.Ruoyu Feng, Wenming Weng, Yanhui Wang, Yuhui Yuan, Jianmin Bao, Chong Luo, Zhibo Chen, and Baining Guo. Ccedit: Creative and controllable video editing via diffusion models. arXiv preprint arXiv:2309.16496, 2023.
  20. 20.Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scenebased text-to-image generation with human priors. arXiv preprint arXiv:2203.13131, 2022.
  21. 21.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H. Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion, 2022.
  22. 22.Rinon Gal, Moab Arar, Yuval Atzmon, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Encoder-based domain tuning for fast personalization of text-to-image models. ACM Transactions on Graphics (TOG), 42(4):1–13, 2023.
  23. 23.Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, MingYu Liu, and Yogesh Balaji. Preserve your own correlation: A noise prior for video diffusion models. In ICCV, 2023.
  24. 24.Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing. arXiv preprint arXiv:2307.10373, 2023.
  25. 25.Yuan Gong, Youxin Pang, Xiaodong Cun, Menghan Xia, Haoxin Chen, Longyue Wang, Yong Zhang, Xintao Wang, Ying Shan, and Yujiu Yang. Talecrafter: Interactive story visualization with multiple characters. arXiv preprint arXiv:2305.18247, 2023.
  26. 26.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022.
  27. 27.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023.
  28. 28.Ligong Han, Yinxiao Li, Han Zhang, Peyman Milanfar, Dimitris Metaxas, and Feng Yang. Svdiff: Compact parameter space for diffusion fine-tuning. arXiv preprint arXiv:2303.11305, 2023.
  29. 29.Xingzhe He, Zhiwen Cao, Nicholas Kolkin, Lantao Yu, Helge Rhodin, and Ratheesh Kalarot. A data perspective on enhanced identity preservation for diffusion personalization. arXiv preprint arXiv:2311.04315, 2023.
  30. 30.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity video generation with arbitrary lengths. arXiv preprint arXiv:2211.13221, 2022.
  31. 31.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022.
  32. 32.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  33. 33.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  34. 34.Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022.
  35. 35.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan AllenZhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022.
  36. 36.Zhihao Hu and Dong Xu. Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet. arXiv preprint arXiv:2307.14073, 2023.
  37. 37.Ziqi Huang, Kelvin CK Chan, Yuming Jiang, and Ziwei Liu. Collaborative diffusion for multi-modal face generation and editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6080–6090, 2023.
  38. 38.Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. arXiv preprint arXiv:2311.17982, 2023.
  39. 39.Ziqi Huang, Tianxing Wu, Yuming Jiang, Kelvin CK Chan, and Ziwei Liu. Reversion: Diffusion-based relation inversion from images. arXiv preprint arXiv:2303.13495, 2023.
  40. 40.Junha Hyung, Jaeyo Shin, and Jaegul Choo. Magicapture: High-resolution multi-concept portrait customization. arXiv preprint arXiv:2309.06895, 2023.
  41. 41.Hyeonho Jeong and Jong Chul Ye. Ground-a-video: Zeroshot grounded video editing using text-to-image diffusion models. arXiv preprint arXiv:2310.01107, 2023.
  42. 42.Xuhui Jia, Yang Zhao, Kelvin CK Chan, Yandong Li, Han Zhang, Boqing Gong, Tingbo Hou, Huisheng Wang, and Yu-Chuan Su. Taming encoder for zero fine-tuning image customization with text-to-image diffusion models. arXiv preprint arXiv:2304.02642, 2023.
  43. 43.Yuming Jiang, Ziqi Huang, Xingang Pan, Chen Change Loy, and Ziwei Liu. Talk-to-edit: Fine-grained facial editing via dialog. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13799–13808, 2021.
  44. 44.Yuming Jiang, Shuai Yang, Haonan Qiu, Wayne Wu, Chen Change Loy, and Ziwei Liu. Text2human: Text-driven controllable human image generation. ACM Transactions on Graphics (TOG), 41(4):1–11, 2022.
  45. 45.Yuming Jiang, Ziqi Huang, Tianxing Wu, Xingang Pan, Chen Change Loy, and Ziwei Liu. Talk-to-edit: Fine-grained 2d and 3d facial editing via dialog. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  46. 46.Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu, Chen Change Loy, and Ziwei Liu. Text2performer: Textdriven human video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.
  47. 47.Chen Jin, Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, and Philip Teare. An image is worth multiple words: Learning object level concepts using multi-concept prompt learning. arXiv preprint arXiv:2310.12274, 2023.
  48. 48.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-toimage diffusion models are zero-shot video generators. arXiv preprint arXiv:2303.13439, 2023.
  49. 49.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023.
  50. 50.Sumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu, Jimei Yang, Jingwan Lu, Alexei A Efros, and Krishna Kumar Singh. Putting people in their place: Affordance-aware human insertion into scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17089–17099, 2023.
  51. 51.Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. 2023.
  52. 52.Dongxu Li, Junnan Li, and Steven CH Hoi. Blipdiffusion: Pre-trained subject representation for controllable text-to-image generation and editing. arXiv preprint arXiv:2305.14720, 2023.
  53. 53.Yuheng Li, Haotian Liu, Yangming Wen, and Yong Jae Lee. Generate anything anywhere in any scene. arXiv preprint arXiv:2306.17154, 2023.
  54. 54.Jun Hao Liew, Hanshu Yan, Jianfeng Zhang, Zhongcong Xu, and Jiashi Feng. Magicedit: High-fidelity and temporally coherent video editing. In arXiv, 2023.
  55. 55.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023.
  56. 56.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9298–9309, 2023.
  57. 57.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  58. 58.Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin, and Jiaya Jia. Video-p2p: Video editing with cross-attention control. arXiv preprint arXiv:2303.04761, 2023.
  59. 59.Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng, Kai Zhu, Ruili Feng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. Cones 2: Customizable image synthesis with multiple subjects. arXiv preprint arXiv:2305.19327, 2023.
  60. 60.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10209–10218, 2023.
  61. 61.Jian Ma, Junhao Liang, Chen Chen, and Haonan Lu. Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning. arXiv preprint arXiv:2307.11410, 2023.
  62. 62.Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. arXiv preprint arXiv:2211.09794, 2022.
  63. 63.Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
  64. 64.Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Juntao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, and Yujun Shen. Codef: Content deformation fields for temporally consistent video processing. arXiv preprint arXiv:2308.07926, 2023.
  65. 65.Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. Fatezero: Fusing attentions for zero-shot text-based video editing. arXiv preprint arXiv:2303.09535, 2023.
  66. 66.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  67. 67.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021.
  68. 68.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  69. 69.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021.
  70. 70.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022.
  71. 71.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. 2022.
  72. 72.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, and Kfir Aberman. Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models. arXiv preprint arXiv:2307.06949, 2023.
  73. 73.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  74. 74.Chaehun Shin, Heeseung Kim, Che Hyun Lee, Sang-gil Lee, and Sungroh Yoon. Edit-a-video: Single video editing with object-aware consistency. arXiv preprint arXiv:2303.07945, 2023.
  75. 75.Chenyang Si, Ziqi Huang, Yuming Jiang, and Ziwei Liu. Freeu: Free lunch in diffusion u-net. arXiv preprint arXiv:2309.11497, 2023.
  76. 76.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  77. 77.Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian Price, Jianming Zhang, Soo Ye Kim, and Daniel Aliaga. Objectstitch: Object compositing with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18310–18319, 2023.
  78. 78.Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653, 2023.
  79. 79.Dani Valevski, Danny Wasserman, Yossi Matias, and Yaniv Leviathan. Face0: Instantaneously conditioning a text-toimage model on a face. arXiv preprint arXiv:2306.06638, 2023.
  80. 80.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description. arXiv preprint arXiv:2210.02399, 2022.
  81. 81.Wen Wang, kangyang Xie, Zide Liu, Hao Chen, Yue Cao, Xinlong Wang, and Chunhua Shen. Zero-shot video editing using off-the-shelf image diffusion models. arXiv preprint arXiv:2303.17599, 2023.
  82. 82.Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023.
  83. 83.Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. arXiv preprint arXiv:2302.13848, 2023.
  84. 84.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Weixian Lei, Yuchao Gu, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. arXiv preprint arXiv:2212.11565, 2022.
  85. 85.Tianxing Wu, Chenyang Si, Yuming Jiang, Ziqi Huang, and Ziwei Liu. Freeinit: Bridging initialization gap in video diffusion models. arXiv preprint arXiv:2312.07537, 2023.
  86. 86.Zijie Wu, Chaohui Yu, Zhen Zhu, Fan Wang, and Xiang Bai. Singleinsert: Inserting new concepts from a single image into text-to-image models for flexible editing. arXiv preprint arXiv:2310.08094, 2023.
  87. 87.Guangxuan Xiao, Tianwei Yin, William T Freeman, Frédo Durand, and Song Han. Fastcomposer: Tuning-free multisubject image generation with localized attention. arXiv preprint arXiv:2305.10431, 2023.
  88. 88.Xingqian Xu, Jiayi Guo, Zhangyang Wang, Gao Huang, Irfan Essa, and Humphrey Shi. Prompt-free diffusion: Taking ”text” out of text-to-image diffusion models. arXiv preprint arXiv:2305.16223, 2023.
  89. 89.Hanshu Yan, Jun Hao Liew, Long Mai, Shanchuan Lin, and Jiashi Feng. Magicprop: Diffusion-based video editing via motion-aware appearance propagation. arXiv preprint arXiv:2309.00908, 2023.
  90. 90.Shuai Yang, Yifan Zhou, Ziwei Liu, and Chen Change Loy. Rerender a video: Zero-shot text-guided video-to-video translation. arXiv preprint arXiv:2306.07954, 2023.
  91. 91.Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ipadapter: Text compatible image prompt adapter for text-toimage diffusion models. arXiv preprint arXiv:2308.06721, 2023.
  92. 92.Ge Yuan, Xiaodong Cun, Yong Zhang, Maomao Li, Chenyang Qi, Xintao Wang, Ying Shan, and Huicheng Zheng. Inserting anybody in diffusion models via celeb basis. arXiv preprint arXiv:2306.00926, 2023.
  93. 93.Ziyang Yuan, Mingdeng Cao, Xintao Wang, Zhongang Qi, Chun Yuan, and Ying Shan. Customnet: Zero-shot object customization with variable-viewpoints in text-to-image diffusion models. arXiv preprint arXiv:2310.19784, 2023.
  94. 94.Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. Controlvideo: Training-free controllable text-to-video generation. arXiv preprint arXiv:2305.13077, 2023.
  95. 95.Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jiawei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. Motiondirector: Motion customization of text-tovideo diffusion models. arXiv preprint arXiv:2310.08465, 2023.
  96. 96.Yuyang Zhao, Enze Xie, Lanqing Hong, Zhenguo Li, and Gim Hee Lee. Make-a-protagonist: Generic video editing with an ensemble of experts. arXiv preprint arXiv:2305.08850, 2023.
  97. 97.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022.
  98. 98.Yufan Zhou, Ruiyi Zhang, Tong Sun, and Jinhui Xu. Enhancing detail preservation for customized text-to-image generation: A regularization-free approach. arXiv preprint arXiv:2305.13579, 2023.

Citation

MLA
Jiang, Y., et al. “VideoBooth: Diffusion-based Video Generation with Image Prompts”. arXiv, 2023, http://arxiv.org/abs/2312.00777v1.
APA
Jiang, Y., Wu, T., Yang, S., Si, C., Lin, D., Qiao, Y., Loy, C. C., & Liu, Z. (2023). VideoBooth: Diffusion-based Video Generation with Image Prompts. arXiv. http://arxiv.org/abs/2312.00777v1
Chicago
Jiang, Y., T. Wu, S. Yang, et al. 2023. “VideoBooth: Diffusion-based Video Generation with Image Prompts”. arXiv. http://arxiv.org/abs/2312.00777v1.
Harvard
Jiang, Y. et al. (2023) “VideoBooth: Diffusion-based Video Generation with Image Prompts”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.00777v1.
Vancouver
1. Jiang Y, Wu T, Yang S, Si C, Lin D, Qiao Y, Loy CC, Liu Z (2023) VideoBooth: Diffusion-based Video Generation with Image Prompts. arXiv

BibTeX

@article{jiang2023videobooth,
  title = {VideoBooth: Diffusion-based Video Generation with Image Prompts},
  author = {Jiang, Yuming and Wu, Tianxing and Yang, Shuai and Si, Chenyang and Lin, Dahua and Qiao, Yu and Loy, Chen Change and Liu, Ziwei},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.00777v1},
  eprint = {2312.00777}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE