Breathing Life Into Sketches Using Text-to-Video Priors
Rinon GalYael VinkerYuval AlalufAmit BermanoDaniel Cohen-OrAriel ShamirGal Chechik
Presents an optimization framework that animates static vector sketches directly from text prompts by distilling motion priors from pretrained text-to-video diffusion models without requiring manual rigging or specialized training data.
Freehand sketches are a fundamental visual communication tool, but animating them traditionally requires specialized design skills and labor-intensive manual work. While modern generative artificial intelligence models can create realistic video from text prompts, existing tools struggle to animate abstract, sparse line drawings without introducing severe pixel distortions or altering the original style. The article evaluates an automated framework designed to generate smooth, short animations from a single static sketch and a written motion prompt, eliminating the need for manual joint rigging, skeletal keypoints, or reference videos.
The authors developed an optimization-based method that operates directly on vector graphics, representing strokes as parametric curves rather than pixels. To extract motion knowledge, the system uses a pretrained text-to-video diffusion model and distills its motion signals using a score-distillation sampling technique. The core architecture splits movement into two parallel pathways: an unconstrained local motion network for fine-grained deformations (such as bending a limb) and a constrained global motion network that applies uniform geometric transformations (such as scaling, rotation, and translation) across entire frames. This dual structure prevents excessive shape distortion while allowing natural, sweeping movements. The authors evaluated the approach across diverse categories including humans, animals, and objects, comparing performance against leading pixel-based video generation baselines.
The findings show that the proposed vector-based architecture consistently outperforms existing image-to-video models in both visual preservation and prompt alignment. Quantitatively, the method achieved a sketch-to-video consistency score of 0.965, substantially exceeding baselines such as ModelScope (0.779) and VideoCrafter (0.876). It also registered higher text alignment (0.142 versus 0.124 for VideoCrafter). In ablation studies and a 31-participant user study, removing either the neural network prior or the global-local motion split led to increased motion jitter, unrealistic wobbling, or failure to preserve the subject's original geometry. Crucially, the authors demonstrated that standard text-to-video backbones—even those that fail to produce clean sketches independently—contain robust semantic motion priors that can successfully drive abstract vector artwork.
These results demonstrate that organizations can automate early-stage animation, visual storytelling, and graphic design workflows with minimal human intervention and no costly model retraining. Because the output remains in vector format, assets retain infinite scalability and can be edited directly by downstream design teams. However, decision-makers should note several limitations: the current pipeline is designed for single-subject sketches and exhibits reduced performance on multi-object scenes; it requires 15 to 30 minutes of optimization per short video on high-end hardware; and it inherits occasional biases or motion inaccuracies present in the underlying generative video backbones. Future work should prioritize extending the architecture to multi-object scenes, incorporating mesh-based structural constraints for amateur drawings, and integrating faster, next-generation video backbones.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). Introduces Score Distillation Sampling (SDS), the foundational technique adapted by the source to distill generative priors into parametric vector sketch animation.
- Paper: ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation, Zhengyi Wang et al. (2023). Extends score distillation principles to variational formulations, providing vital mathematical context for how diffusion distillation operates during optimization.
- Paper: AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. (2024). Demonstrates how temporal motion modules inject dynamics into static generation models, offering key background for text-to-video motion priors used in the source.
- Paper: Shape-Aware Text-Driven Layered Video Editing, Yao-Chih Lee et al. (2023). Explores score distillation guidance for decomposing and modifying layered video shapes, establishing relevant methods for parametric deformation.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Establishes video diffusion modeling foundations that supply the pretrained spatio-temporal priors leveraged by the source.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). Introduces conditional structural controls such as sketches and edges into diffusion pipelines, serving as essential conceptual groundwork for sketch-driven generation.
- Paper: SVGDreamer: Text Guided SVG Generation with Diffusion Model, Ximing Xing et al. (2024). Extends score distillation sampling on parametric vector graphics to high-quality static SVG generation and distinct path editability.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). Explores training-free motion extraction directly from temporal attention layers to guide video generation without per-sample optimization.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). Builds on trajectory-guided motion generation by integrating explicit user-drawn motion paths into diffusion transformers.
- Paper: VideoBooth: Diffusion-based Video Generation with Image Prompts, Yuming Jiang et al. (2024). Applies cross-frame visual attention conditioning to ensure fine subject consistency across video frames without inference fine-tuning.
