VideoBooth: Diffusion-based Video Generation with Image Prompts
Yuming JiangTianxing WuShuai YangChenyang SiDahua LinYu QiaoChen Change LoyZiwei Liu
Proposes VideoBooth, a tuning-free feed-forward diffusion framework that generates customized, temporally consistent videos from image prompts by combining coarse semantic embeddings with multi-scale attention injection.
Text-driven video generation models have advanced rapidly, yet text descriptions alone often fail to capture specific visual appearances and fine details needed for customized content creation. Existing personalization methods typically rely on fine-tuning model parameters at inference time or requiring multiple reference images, both of which introduce high computational costs, deployment delays, and operational friction.
The article introduces VideoBooth, a framework designed to generate high-quality, temporally consistent videos conditioned on both a text prompt and a single reference image prompt without requiring any fine-tuning during inference.
To achieve this, the authors developed a coarse-to-fine visual embedding architecture built on latent video diffusion models. At the coarse level, an image encoder extracts high-level semantic features from the image prompt and maps them into text embedding space. At the fine level, multi-scale latent representations of the image prompt are injected into the model's cross-frame attention layers as additional keys and values. This attention injection directly refines spatial details in the initial frame and propagates them across subsequent frames to preserve temporal consistency. The framework is trained sequentially, optimizing the coarse encoder first to prevent feature leakage before training the fine attention module. The authors also established a dedicated dataset derived from WebVid, filtering down to 48,724 training pairs of video clips, text prompts, and segmented subject image prompts, along with a 650-pair benchmark for testing.
The experimental evaluation demonstrated key performance findings:
- Superior subject fidelity: VideoBooth achieved state-of-the-art visual alignment scores, recording a 74.80 CLIP-Image score and a 65.10 DINO score, substantially outperforming existing customized generation baselines such as Textual Inversion, DreamBooth, and ELITE.
- Preserved textual alignment: VideoBooth maintained strong prompt fidelity with a CLIP-Text score of 30.10, performing comparably to established alternatives while properly balancing text and image instructions.
- Strong user preference: In a user study with 25 participants across multiple test scenarios, VideoBooth secured the highest user preference rates in image alignment, text alignment, and overall visual quality.
- Critical ablation insights: Removing coarse embeddings caused temporal degradation and object distortion across later frames, omitting fine injection caused lost visual patterns, and training both components simultaneously degraded overall encoder performance.
These results show that high-fidelity subject customization in generative video can be accomplished through feed-forward inference alone. By eliminating test-time fine-tuning, the framework lowers inference latency and computational expenses, making customized video generation more viable for scalable commercial pipelines. Furthermore, the two-stage coarse-to-fine injection strategy solves the trade-off between retaining fine-grained spatial attributes and preserving fluid, consistent temporal motion.
Organizations evaluating this technology should adopt feed-forward, multi-scale visual embedding architectures to streamline customized video production workflows. Future research and development should focus on expanding the dataset with automated 3D image augmentation pipelines to enable diverse multi-angle viewpoints, and integrating deepfake detection safeguards to address synthetic media risks.
The framework's primary technical limitation is its reliance on training image prompts that closely align with the viewpoints of the target videos, meaning it cannot reliably synthesize extreme perspective shifts (such as generating a front-facing video from a rear-view prompt). Confidence in the reported image fidelity and temporal consistency gains is high, backed by comprehensive quantitative metrics, ablation studies, and qualitative user evaluations.
- Paper: DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation, Nataniel Ruiz et al. (2023). It introduces the foundational subject-driven personalization framework that VideoBooth aims to replace with a training-free, feed-forward alternative for video generation.
- Paper: IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models, Hu Ye et al. (2023). It establishes the decoupled cross-attention mechanism for injecting visual image prompt features into diffusion models without full fine-tuning.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It establishes the core paradigm of turning 2D latent image diffusion models into video generators by inserting temporal attention layers.
- Paper: An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion, Rinon Gal et al. (2022). It introduces pseudo-word embedding optimization for visual concept personalization, serving as a primary baseline and conceptual precursor to VideoBooth's coarse visual mapping.
- Paper: AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. (2024). It demonstrates how to animate personalized text-to-image diffusion models using plug-and-play temporal motion modules without per-model fine-tuning.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It presents foundational architectures and curated training protocols for scaling latent video diffusion models conditioned on images.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It provides the foundational framework and 3D spatio-temporal architecture for extending diffusion models to video generation.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). It introduces the architectural blueprint of using zero convolutions and auxiliary paths to feed additional conditioning signals into diffusion models.
- Paper: Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning, Rohit Girdhar et al. (2024). It extends the concept of image-conditioned video generation by explicitly factorizing the text-to-video pipeline into sequential image synthesis and video animation stages.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). It advances space-time video diffusion modeling with unified architectures for coherent motion and versatile image-to-video editing tasks.
- Paper: Hierarchical Spatio-temporal Decoupling for Text-to- Video Generation, Zhiwu Qing et al. (2024). It expands spatio-temporal generation by hierarchically decoupling visual structure and content-guided motion dynamics.
- Paper: CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer, Zhuoyi Yang et al. (2025). It scales text-to-video diffusion through expert transformer backbones and 3D causal compression for extended, coherent video generation.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). It generalizes reference-based generation and subject customization into a unified multi-image editing framework grounded in video-derived dynamics.
