Breathing Life Into Sketches Using Text-to-Video Priors

Rinon GalYael VinkerYuval AlalufAmit BermanoDaniel Cohen-OrAriel ShamirGal Chechik

article2024CVPR65 citations

Presents an optimization framework that animates static vector sketches directly from text prompts by distilling motion priors from pretrained text-to-video diffusion models without requiring manual rigging or specialized training data.

Listen

Freehand sketches are a fundamental visual communication tool, but animating them traditionally requires specialized design skills and labor-intensive manual work. While modern generative artificial intelligence models can create realistic video from text prompts, existing tools struggle to animate abstract, sparse line drawings without introducing severe pixel distortions or altering the original style. The article evaluates an automated framework designed to generate smooth, short animations from a single static sketch and a written motion prompt, eliminating the need for manual joint rigging, skeletal keypoints, or reference videos.

The authors developed an optimization-based method that operates directly on vector graphics, representing strokes as parametric curves rather than pixels. To extract motion knowledge, the system uses a pretrained text-to-video diffusion model and distills its motion signals using a score-distillation sampling technique. The core architecture splits movement into two parallel pathways: an unconstrained local motion network for fine-grained deformations (such as bending a limb) and a constrained global motion network that applies uniform geometric transformations (such as scaling, rotation, and translation) across entire frames. This dual structure prevents excessive shape distortion while allowing natural, sweeping movements. The authors evaluated the approach across diverse categories including humans, animals, and objects, comparing performance against leading pixel-based video generation baselines.

The findings show that the proposed vector-based architecture consistently outperforms existing image-to-video models in both visual preservation and prompt alignment. Quantitatively, the method achieved a sketch-to-video consistency score of 0.965, substantially exceeding baselines such as ModelScope (0.779) and VideoCrafter (0.876). It also registered higher text alignment (0.142 versus 0.124 for VideoCrafter). In ablation studies and a 31-participant user study, removing either the neural network prior or the global-local motion split led to increased motion jitter, unrealistic wobbling, or failure to preserve the subject's original geometry. Crucially, the authors demonstrated that standard text-to-video backbones—even those that fail to produce clean sketches independently—contain robust semantic motion priors that can successfully drive abstract vector artwork.

These results demonstrate that organizations can automate early-stage animation, visual storytelling, and graphic design workflows with minimal human intervention and no costly model retraining. Because the output remains in vector format, assets retain infinite scalability and can be edited directly by downstream design teams. However, decision-makers should note several limitations: the current pipeline is designed for single-subject sketches and exhibits reduced performance on multi-object scenes; it requires 15 to 30 minutes of optimization per short video on high-end hardware; and it inherits occasional biases or motion inaccuracies present in the underlying generative video backbones. Future work should prioritize extending the architecture to multi-object scenes, incorporating mesh-based structural constraints for amateur drawings, and integrating faster, next-generation video backbones.

arXiv: 2311.13608
Cover for Breathing Life Into Sketches Using Text-to-Video Priors

Abstract

A sketch is one of the most intuitive and versatile tools humans use to convey their ideas visually. An animated sketch opens another dimension to the expression of ideas and is widely used by designers for a variety of purposes. Animating sketches is a laborious process, requiring extensive experience and professional design skills. In this work, we present a method that automatically adds motion to a single-subject sketch (hence, “breathing life into it”), merely by providing a text prompt indicating the desired motion. The output is a short animation provided in vector representation, which can be easily edited. Our method does not require extensive training, but instead leverages the motion prior of a large pretrained text-to-video diffusion model using a score-distillation loss to guide the placement of strokes. To promote natural and smooth motion and to better preserve the sketch’s appearance, we model the learned motion through two components. The first governs small local deformations and the second controls global affine transformations. Surprisingly, we find that even models that struggle to generate sketch videos on their own can still serve as a useful backbone for animating abstract representations.

Table of Contents

  • 1. Introduction
  • 2. Previous Work
  • 3. Preliminaries
  • 4. Method
  • 4.1. Representation
  • 4.2. Text-Driven Optimization
  • 4.3. Neural Displacement Field
  • 4.4. Training Details
  • 5. Results
  • 5.1. Comparisons
  • 5.2. Ablation Study
  • 6. Limitations
  • 7. Conclusions
  • References

Knowls

  1. Knowl 1 — Vector Representation and Stroke Displacement Formulation for Sketch Animation

    model/method

    A static sketch is modeled as a set of two-dimensional cubic Bézier curves placed over a white canvas, where each stroke is parameterized by 4 control points. Let NN denote the total number of control points in the sketch. The geometry of the static sketch is defined by its control point coordinates:

    Pinit={p1init,p2init,…,pNinit}∈RN×2P^{\text{init}} = \{p_1^{\text{init}}, p_2^{\text{init}}, \dots, p_N^{\text{init}}\} \in \mathbb{R}^{N \times 2}

    where each point piinit=(xi,yi)∈R2p_i^{\text{init}} = (x_i, y_i) \in \mathbb{R}^2.

    To generate an animation of kk frames, the initial point set PinitP^{\text{init}} is replicated across kk temporal instances to form an initial static sequence Zinit={Pj}j=1k∈RN⋅k×2Z^{\text{init}} = \{P^j\}_{j=1}^k \in \mathbb{R}^{N \cdot k \times 2}. The animation task is formulated as learning a set of 2D displacements:

    ΔZ={Δpij}i=1,j=1N,k∈RN⋅k×2\Delta Z = \{\Delta p_i^j\}_{i=1, j=1}^{N, k} \in \mathbb{R}^{N \cdot k \times 2}

    where Δpij∈R2\Delta p_i^j \in \mathbb{R}^2 represents the spatial offset applied to point piinitp_i^{\text{init}} at frame j∈{1,…,k}j \in \{1, \dots, k\}. The number of strokes and control points NN is kept strictly fixed across all generated frames, preserving stroke topology and preventing rasterization artifacts like blurring or pixelation.

  2. Knowl 2 — Text-Driven Video Score Distillation Sampling for Vector Sketches

    equation

    To animate a vector sketch according to a textual description cc, stroke displacements ΔZ\Delta Z parameterized by a neural network with weights ϕ\phi are optimized using a pretrained text-to-video diffusion model ϵθ\epsilon_\theta via Score Distillation Sampling (SDS).

    At each iteration, the deformed control points for each frame j∈{1,…,k}j \in \{1, \dots, k\}, given by Pj=Pinit+ΔPjP^j = P^{\text{init}} + \Delta P^j, are rendered into raster frames Fj=R(Pj)F^j = \mathcal{R}(P^j) using a differentiable rasterizer R\mathcal{R}. The rendered frames are concatenated into a video tensor F={F1,…,Fk}∈Rh×w×kF = \{F^1, \dots, F^k\} \in \mathbb{R}^{h \times w \times k}, where hh and ww are the frame height and width.

    Gaussian noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) is added to FF at a randomly sampled diffusion timestep tt according to the noise schedule Ft=αtF+σtϵF_t = \alpha_t F + \sigma_t \epsilon. The diffusion model predicts the noise ϵθ(Ft,t,c)\epsilon_\theta(F_t, t, c) conditioned on prompt cc. The gradient updating the network parameters ϕ\phi is:

    ∇ϕLSDS=Et,ϵ[w(t)(ϵθ(Ft,t,c)−ϵ)∂F∂ϕ]\nabla_\phi \mathcal{L}_{\text{SDS}} = \mathbb{E}_{t, \epsilon} \left[ w(t) \left( \epsilon_\theta(F_t, t, c) - \epsilon \right) \frac{\partial F}{\partial \phi} \right]

    where w(t)w(t) is a time-dependent weighting scalar and ∂F∂ϕ=∂R(Z)∂Z∂Z∂ϕ\frac{\partial F}{\partial \phi} = \frac{\partial \mathcal{R}(Z)}{\partial Z} \frac{\partial Z}{\partial \phi} propagates gradients from the pixel space back into the vector parameters.

  3. Knowl 3 — Disentangled Local-Global Neural Displacement Field Architecture

    model/method

    Directly optimizing unconstrained displacements for all control points leads to jitter and severe shape distortion. The motion model M\mathcal{M} decouples motion into an unconstrained local deformation path and a constrained global affine transformation path: ΔZ=ΔZg+ΔZl\Delta Z = \Delta Z_g + \Delta Z_l.

    1. Shared Backbone: Each initial control point pijp_i^j is projected into a latent feature space via a shared linear transformation MsharedM_{\text{shared}} and summed with a positional encoding vector that depends on both the temporal frame index jj and the point index ii in the stroke sequence.

    2. Local Motion Path: An MLP Ml\mathcal{M}_l receives the shared latent embeddings and predicts per-point offsets ΔZl={Δpi,localj}i=1,j=1N,k\Delta Z_l = \{\Delta p_{i, \text{local}}^j\}_{i=1, j=1}^{N, k}, accommodating fine-grained, non-rigid semantic movements.

    3. Global Motion Path: A network Mg\mathcal{M}_g processes the shared features to predict a single affine transformation matrix Tj\mathcal{T}^j for each frame jj. This matrix is applied uniformly to all initial control points PinitP^{\text{init}} in frame jj, producing the global displacement:

    Δpi,globalj=Tj⊙piinit−piinit\Delta p_{i, \text{global}}^j = \mathcal{T}^j \odot p_i^{\text{init}} - p_i^{\text{init}}

  4. Knowl 4 — Parametric Global Transformation Matrix with Motion Attenuation Scaling

    model/method

    The per-frame global transformation matrix Tj\mathcal{T}^j for frame jj is parameterized by sequentially applying 2D scaling, shear, rotation, and translation:

    Tj=[sxjshxjsyjdxjshyjsxjsyjdyj001][cos⁡θj−sin⁡θj0sin⁡θjcos⁡θj0001]\mathcal{T}^j = \begin{bmatrix} s_x^j & sh_x^j s_y^j & d_x^j \\ sh_y^j s_x^j & s_y^j & d_y^j \\ 0 & 0 & 1 \end{bmatrix} \begin{bmatrix} \cos\theta^j & -\sin\theta^j & 0 \\ \sin\theta^j & \cos\theta^j & 0 \\ 0 & 0 & 1 \end{bmatrix}

    where (sxj,syj)(s_x^j, s_y^j) are scale factors, (shxj,shyj)(sh_x^j, sh_y^j) are shear parameters, θj\theta^j is the rotation angle, and (dxj,dyj)(d_x^j, d_y^j) are translation offsets.

    To give users explicit control over individual motion degrees of freedom, the network's raw predicted parameters (s~xj,s~yj,sh~xj,sh~yj,θ~j,d~xj,d~yj)(\tilde{s}_x^j, \tilde{s}_y^j, \tilde{sh}_x^j, \tilde{sh}_y^j, \tilde{\theta}^j, \tilde{d}_x^j, \tilde{d}_y^j) are scaled by user-defined attenuation coefficients (λs,λsh,λr,λt)(\lambda_s, \lambda_{sh}, \lambda_r, \lambda_t):

    (sxj,syj)=(1+λss~xj,1+λss~yj)(s_x^j, s_y^j) = (1 + \lambda_s \tilde{s}_x^j, 1 + \lambda_s \tilde{s}_y^j) (shxj,shyj)=(λshsh~xj,λshsh~yj)(sh_x^j, sh_y^j) = (\lambda_{sh} \tilde{sh}_x^j, \lambda_{sh} \tilde{sh}_y^j) θj=λrθ~j\theta^j = \lambda_r \tilde{\theta}^j (dxj,dyj)=(λtd~xj,λtd~yj)(d_x^j, d_y^j) = (\lambda_t \tilde{d}_x^j, \lambda_t \tilde{d}_y^j)

    Default attenuation parameters are set to λt=1.0\lambda_t = 1.0, λr=10−2\lambda_r = 10^{-2}, λs=5×10−2\lambda_s = 5 \times 10^{-2}, and λsh=10−1\lambda_{sh} = 10^{-1}. Setting λt=0\lambda_t = 0, for example, keeps the subject centered while allowing local articulation.

  5. Knowl 5 — Alternating Optimization Algorithm for Vector Sketch Animation

    algorithm

    The text-driven sketch animation optimizes the shared backbone MsharedM_{\text{shared}}, local MLP Ml\mathcal{M}_l, and global network Mg\mathcal{M}_g over T=1000T = 1000 iterations by alternating between local and global path updates.

    Input: Initial sketch control points Pinit∈RN×2P^{\text{init}} \in \mathbb{R}^{N \times 2}, motion text prompt cc, frame count kk, iterations T=1000T = 1000
    Output: Animated control points Z={Pj}j=1kZ = \{P^j\}_{j=1}^k
    Initialize parameters of MsharedM_{\text{shared}}, Ml\mathcal{M}_l, Mg\mathcal{M}_g
    Set hyperparameters: λt=1.0\lambda_t = 1.0, λr=0.01\lambda_r = 0.01, λs=0.05\lambda_s = 0.05, λsh=0.1\lambda_{sh} = 0.1
    Set Adam learning rates: ηglobal=10−4\eta_{\text{global}} = 10^{-4}, ηlocal=5×10−3\eta_{\text{local}} = 5 \times 10^{-3}
    Set SDS guidance scales: wglobal=40w_{\text{global}} = 40, wlocal=30w_{\text{local}} = 30
    for step =1= 1 to TT do
        Zinit←Z^{\text{init}} \leftarrow duplicate PinitP^{\text{init}} across kk frames
        Compute shared embeddings from ZinitZ^{\text{init}} using MsharedM_{\text{shared}} and positional encodings
        Predict global matrices {Tj}j=1k\{\mathcal{T}^j\}_{j=1}^k via Mg\mathcal{M}_g and apply scalings (λt,λr,λs,λsh)(\lambda_t, \lambda_r, \lambda_s, \lambda_{sh})
        Compute global displacements ΔZg\Delta Z_g
        Predict local offsets ΔZl\Delta Z_l via Ml\mathcal{M}_l
        Z←Zinit+ΔZg+ΔZlZ \leftarrow Z^{\text{init}} + \Delta Z_g + \Delta Z_l
        Render pixel video F={R(Pj)}j=1kF = \{\mathcal{R}(P^j)\}_{j=1}^k
        Apply spatial augmentations (random crops, perspective transforms) to FF
        Sample diffusion timestep tt and noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I)
        Compute ∇LSDS\nabla \mathcal{L}_{\text{SDS}} conditioned on prompt cc
        if step is odd then
            Update parameters of Ml\mathcal{M}_l and MsharedM_{\text{shared}} with learning rate ηlocal\eta_{\text{local}}
        else
            Update parameters of Mg\mathcal{M}_g and MsharedM_{\text{shared}} with learning rate ηglobal\eta_{\text{global}}
        end if
    end for
    return ZZ
  6. Knowl 6 — Quantitative Benchmark Against Pixel-Based Video Generation Baselines

    data/table

    The vector-based sketch animation method is evaluated across 30 sketches spanning humans, animals, and objects against open-source image-to-video baselines: ZeroScope, ModelScope, and VideoCrafter.

    Evaluation metrics include:

    1. Sketch-to-video consistency (↑\uparrow): Average cosine similarity in CLIP embedding space between individual generated video frames and the input static sketch.
    2. Text-to-video alignment (↑\uparrow): Cosine similarity between the generated video sequence and the input text prompt computed using X-CLIP.
    Method Sketch-to-video consistency (↑\uparrow) Text-to-Video alignment (↑\uparrow)
    ZeroScope 0.754±0.0090.754 \pm 0.009 –
    ModelScope 0.779±0.0090.779 \pm 0.009 –
    VideoCrafter 0.876±0.0070.876 \pm 0.007 0.124±0.0050.124 \pm 0.005
    Ours 0.965±0.003\mathbf{0.965 \pm 0.003} 0.142±0.005\mathbf{0.142 \pm 0.005}

    The vector representation constrained by score distillation achieves significantly higher visual consistency (0.9650.965) compared to direct raster generation models (0.7540.754--0.8760.876), which frequently introduce raster artifacts, deform sketch stroke structure, or hallucinate photorealistic textures outside the sketch domain.

  7. Knowl 7 — Ablation Study on Architecture Components and Global-Local Separation

    data/table

    The contributions of the neural displacement network and the separation into global and local pathways are evaluated quantitatively using CLIP sketch-to-video consistency and X-CLIP text-to-video alignment across four configurations:

    • Full: The complete proposed model with shared backbone, global affine branch, and local MLP branch.
    • No Net: Direct coordinate optimization of control point offsets without a neural network prior.
    • No Glob.: Model containing only the unconstrained local displacement branch.
    • No Local: Model containing only the global affine transformation branch.
    Setup Sketch-to-video consistency (↑\uparrow) Text-to-Video alignment (↑\uparrow)
    Full 0.965±0.0030.965 \pm 0.003 0.142±0.0050.142 \pm 0.005
    No Net 0.926±0.0070.926 \pm 0.007 0.142±0.0050.142 \pm 0.005
    No Glob. 0.936±0.0060.936 \pm 0.006 0.140±0.0050.140 \pm 0.005
    No Local 0.970±0.0020.970 \pm 0.002 0.140±0.0040.140 \pm 0.004

    Removing the neural network prior (No Net) or the global affine path (No Glob.) harms sketch consistency (0.9260.926 and 0.9360.936) and introduces high-frequency jitter. Removing the local path (No Local) produces the highest frame consistency (0.9700.970) because the sketch remains almost static, but it fails to generate expressive articulated motion. In a 2-alternative forced-choice user study (N=31N = 31 participants over 30 video pairs), responders preferred the Full model for both appearance preservation and motion plausibility over the ablated variants.

  8. Knowl 8 — Limitations and Failure Modes of Text-Driven Sketch Animation

    limitation

    The method is subject to four primary operational limitations:

    1. Vector Primitive Restrictions: The framework relies on cubic Bézier curves. It cannot handle vector curves with C1C^1 continuity constraints under the differentiable rasterizer, and applying the method to non-standard vector primitives can cause scale or geometric drift.
    2. Single-Subject Assumption: The optimization assumes a single isolated foreground subject. When applied to multi-object sketches or complete scenes, motion cannot be independently assigned to interacting objects (e.g., a basketball cannot detach from a dribbling basketball player).
    3. Fidelity-Motion Trade-Off: Text-to-video diffusion priors tend to "correct" abstract or amateur drawings toward typical natural depictions before synthesizing motion, causing unintended geometric drift in highly stylized sketches.
    4. Diffusion Prior Biases: The quality, variety, and physical accuracy of generated animations are bound by the semantic and temporal biases of the pretrained text-to-video diffusion backbone.

Coverage note — None was omitted; all primary methodological components, mathematical formulations, algorithmic training steps, quantitative/qualitative experimental benchmarks, ablation studies, and stated limitations are fully represented.

References

  1. 1.Aseem Agarwala, Aaron Hertzmann, David Salesin, and Steven Seitz. Keyframe-based tracking of rotoscoping and animation. ACM Trans. Graph., 23:584–591, 2004. 2
  2. 2.Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin. Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv preprint arXiv:2304.08477, 2023. 3
  3. 3.Maxime Aubert, Adam Brumm, Muhammad Ramli, Thomas Sutikna, E Wahyu Saptomo, Budianto Hakim, Michael J Morwood, Gerrit D van den Bergh, Leslie Kinsley, and Anthony Dosseto. Pleistocene cave art from sulawesi, indonesia. Nature, 514(7521):223–227, 2014. 1
  4. 4.Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H Campbell, and Sergey Levine. Stochastic variational video prediction. arXiv preprint arXiv:1710.11252, 2017. 3
  5. 5.Itamar Berger, Ariel Shamir, Moshe Mahler, Elizabeth Carter, and Jessica Hodgins. Style and abstraction in portrait sketching. ACM Trans. Graph., 32(4), 2013. 2
  6. 6.Ayan Kumar Bhunia, Ayan Das, Umar Riaz Muhammad, Yongxin Yang, Timothy M. Hospedales, Tao Xiang, Yulia Gryaditskaya, and Yi-Zhe Song. Pixelor: a competitive sketching ai agent. so you think you can sketch? ACM Trans. Graph., 39:166:1–166:15, 2020. 2
  7. 7.Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, and Michael Felsberg. Doodleformer: Creative sketch drawing with transformers. ECCV, 2022. 2
  8. 8.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  9. 9.Christoph Bregler, Lorie Loeb, Erika Chuang, and Hrishi Deshpande. Turning to the masters: Motion capturing cartoons. ACM Transactions on Graphics (TOG), 21(3):399–407, 2002. 1
  10. 10.Lluis Castrejon, Nicolas Ballas, and Aaron Courville. Improved conditional vrnns for video prediction. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7608–7617, 2019. 3
  11. 11.Caroline Chan, Fredo Durand, and Phillip Isola. Learning to generate line drawings that convey geometry and semantics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7915–7925, 2022. 2
  12. 12.Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter1: Open diffusion models for high-quality video generation, 2023. 3, 7
  13. 13.Yajing Chen, Shikui Tu, Yuqi Yi, and Lei Xu. Sketchpix2seq: a model to generate sketches of multiple categories. ArXiv, abs/1709.04121, 2017. 2
  14. 14.James Davis, Maneesh Agrawala, Erika Chuang, Zoran Popovic, and David Salesin. A sketching interface for articulated figure animation. In Acm siggraph 2006 courses, pages 15–es. 2006. 2
  15. 15.Richard C. Davis, Brien Colwell, and James A. Landay. Ksketch: A ’kinetic’ sketch pad for novice animators. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, page 413–422, New York, NY, USA, 2008. Association for Computing Machinery. 2
  16. 16.Emily Denton and Rob Fergus. Stochastic video generation with a learned prior. In International conference on machine learning, pages 1174–1183. PMLR, 2018. 3
  17. 17.Marek Dvorozˇnˇak, Wilmot Li, Vladimir G Kim, and Daniel Sýkora. Toonsynth: example-based synthesis of handcolored cartoon animations. ACM Transactions on Graphics (TOG), 37(4):1–11, 2018. 1, 2
  18. 18.Mathias Eitz, James Hays, and Marc Alexa. How do humans sketch objects? ACM Trans. Graph. (Proc. SIGGRAPH), 31(4):44:1–44:10, 2012. 7
  19. 19.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. arXiv preprint arXiv:2302.03011, 2023. 3, 7
  20. 20.Judy Fan, Wilma A. Bainbridge, Rebecca Chamberlain, and Jeffrey D. Wammes. Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2:556 – 568, 2023. 1, 2
  21. 21.Judith E. Fan, Daniel L. K. Yamins, and Nicholas B. Turk-Browne. Common object representations for visual production and recognition. Cognitive science, 42 8:2670–2698, 2018. 2
  22. 22.Kevin Frans, Lisa B. Soros, and Olaf Witkowski. Clipdraw: Exploring text-to-drawing synthesis through languageimage encoders. CoRR, abs/2106.14843, 2021. 2
  23. 23.Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing. arXiv preprint arxiv:2307.10373, 2023. 8
  24. 24.Michael Gleicher. Motion path editing. In Proceedings of the 2001 Symposium on Interactive 3D Graphics, page 195–202, New York, NY, USA, 2001. Association for Computing Machinery. 2
  25. 25.Ernst Hans Gombrich. The story of art. Phaidon London, 1995. 1
  26. 26.Martin Guay, Remi Ronfard, Michael Gleicher, and Marie-Paule Cani. Space-time sketching of character animation. ACM Transactions on Graphics (ToG), 34(4):1–10, 2015. 2
  27. 27.David Ha and Douglas Eck. A neural representation of sketch drawings. CoRR, abs/1704.03477, 2017. 2
  28. 28.Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for highfidelity long video generation. 2022. 3
  29. 29.Aaron Hertzmann. Why do line drawings work? a realism hypothesis. Perception, 49:439 – 451, 2020. 2
  30. 30.Tobias Hinz, Matthew Fisher, Oliver Wang, Eli Shechtman, and Stefan Wermter. Charactergan: Few-shot keypoint character animation and reposing. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1988–1997, 2022. 2
  31. 31.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 2, 3
  32. 32.Alexander Hornung, Ellen Dekkers, and Leif Kobbelt. Character animation from 2d pictures and 3d motion data. ACM Trans. Graph., 26(1):1–es, 2007. 2
  33. 33.Yaosi Hu, Chong Luo, and Zhenzhong Chen. Make it move: controllable image-to-video generation with text descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18219–18228, 2022. 3
  34. 34.Yaosi Hu, Zhenzhong Chen, and Chong Luo. Lamd: Latent motion diffusion for video generation, 2023. 3
  35. 35.Takeo Igarashi, Rieko Kadobayashi, Kenji Mase, and Hidehiko Tanaka. Path drawing for 3d walkthrough. In Proceedings of the 11th Annual ACM Symposium on User Interface Software and Technology, page 173–174, New York, NY, USA, 1998. Association for Computing Machinery. 2
  36. 36.Takeo Igarashi, Tomer Moscovich, and John F. Hughes. As-rigid-as-possible shape manipulation. ACM Trans. Graph., 24(3):1134–1141, 2005. 8
  37. 37.Shir Iluz, Yael Vinker, Amir Hertz, Daniel Berio, Daniel Cohen-Or, and Ariel Shamir. Word-as-image for semantic typography. ACM Trans. Graph., 42(4), 2023. 2, 5
  38. 38.Ajay Jain, Amber Xie, and Pieter Abbeel. Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models. arXiv, 2022. 2, 5
  39. 39.Moritz Kampelmuhler and Axel Pinz. Synthesizing human-like sketches from natural images using a conditional convolutional decoder. CoRR, abs/2003.07101, 2020. 2
  40. 40.Rubaiat Kazi, Fanny Chevalier, Tovi Grossman, Shengdong Zhao, and George Fitzmaurice. Draco: Bringing life to illustrations with kinetic textures. Conference on Human Factors in Computing Systems - Proceedings, 2014. 2
  41. 41.Levon Khachatryan. Tex-an mesh: Textured and animatable human body mesh reconstruction from a single image. https://github.com/lev1khachatryan/Tex-An_Mesh, 2020. 2
  42. 42.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv preprint arXiv:2303.13439, 2023. 2
  43. 43.Doyeon Kim, Donggyu Joo, and Junmo Kim. Tivgan: Text to image to video generation with step-by-step evolutionary generator. IEEE Access, 8:153113–153122, 2020. 3
  44. 44.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  45. 45.Zohar Levi and Craig Gotsman. ArtiSketch: A System for Articulated Sketch Modeling. Computer Graphics Forum, 2013. 2
  46. 46.Mengtian Li, Zhe Lin, Radomir Mech, Ersin Yumer, and Deva Ramanan. Photo-sketching: Inferring contour drawings from images, 2019. 2
  47. 47.Tzu-Mao Li, Michal Luka´c, Gharbi Micha¨el, and Jonathan Ragan-Kelley. Differentiable vector graphics rasterization for editing and learning. ACM Trans. Graph. (Proc. SIGGRAPH Asia), 39(6):193:1–193:15, 2020. 3, 4
  48. 48.Xin Li, Wenqing Chu, Ye Wu, Weihang Yuan, Fanglong Liu, Qi Zhang, Fu Li, Haocheng Feng, Errui Ding, and Jingdong Wang. Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation. arXiv preprint arXiv:2309.00398, 2023. 3
  49. 49.Yi Li, Yi-Zhe Song, Timothy M. Hospedales, and Shaogang Gong. Free-hand sketch synthesis with deformable stroke models. CoRR, abs/1510.02644, 2015. 2
  50. 50.Yitong Li, Martin Min, Dinghan Shen, David Carlson, and Lawrence Carin. Video generation from text. In Proceedings of the AAAI conference on artificial intelligence, 2018. 3
  51. 51.Hangyu Lin, Yanwei Fu, Yu-Gang Jiang, and X. Xue. Sketch-bert: Learning sketch bidirectional encoder representation from transformers by self-supervised learning of sketch gestalt. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6757–6766, 2020. 2
  52. 52.Difan Liu, Matthew Fisher, Aaron Hertzmann, and Evangelos Kalogerakis. Neural strokes: Stylized line drawing of 3d shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14204–14213, 2021. 2
  53. 53.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. Videofusion: Decomposed diffusion models for high-quality video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 3, 6, 7
  54. 54.Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. Latent-nerf for shape-guided generation of 3d shapes and textures. arXiv preprint arXiv:2211.07600, 2022. 2
  55. 55.Daniela Mihai and Jonathon Hare. Learning to draw: Emergent communication through sketching. Advances in Neural Information Processing Systems, 34:7153–7166, 2021. 2
  56. 56.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European Conference on Computer Vision, pages 405–421. Springer, 2020. 3
  57. 57.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. 2
  58. 58.Jianyuan Min, Yen-Lin Chen, and Jinxiang Chai. Interactive generation of human animation with deformable motion models. ACM Trans. Graph., 29(1), 2009. 2
  59. 59.Haoran Mo, Edgar Simo-Serra, Chengying Gao, Changqing Zou, and Ruomei Wang. General virtual sketching framework for vector line art. ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH 2021), 40(4):51:1–51:14, 2021. 2
  60. 60.Umar Riaz Muhammad, Yongxin Yang, Yi-Zhe Song, Tao Xiang, and Timothy M. Hospedales. Learning deep sketch abstraction. CoRR, abs/1804.04804, 2018. 2
  61. 61.Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jianlong Fu, Shiming Xiang, and Haibin Ling. Expanding language-image pretrained models for general video recognition. 2022. 7
  62. 62.Haomiao Ni, Changhao Shi, Kai Li, Sharon X. Huang, and Martin Renqiang Min. Conditional image-to-video generation with latent flow diffusion models, 2023. 2
  63. 63.A. Cengiz Öztireli, Ilya Baran, Tiberiu Popa, Boris Dalstein, Robert W. Sumner, and Markus Gross. Differential blending for expressive sketch-based posing. In Proceedings of the 2013 ACM SIGGRAPH/Eurographics Symposium on Computer Animation, New York, NY, USA, 2013. ACM. 2
  64. 64.Junjun Pan and Jian J. Zhang. Sketch-Based Skeleton-Driven 2D Animation and Motion Capture, pages 164–181. Springer Berlin Heidelberg, Berlin, Heidelberg, 2011. 2
  65. 65.Yingwei Pan, Zhaofan Qiu, Ting Yao, Houqiang Li, and Tao Mei. To create what you tell: Generating videos from captions. In Proceedings of the 25th ACM international conference on Multimedia, pages 1789–1798, 2017. 3
  66. 66.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representations, 2022. 2, 3
  67. 67.Omid Poursaeed, Vladimir Kim, Eli Shechtman, Jun Saito, and Serge Belongie. Neural puppet: Generative layered cartoon characters. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3346–3356, 2020. 2
  68. 68.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 7
  69. 69.Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, and Moacir Antonelli Ponti. Sketchformer: Transformer-based representation for sketched structure. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14141–14150, 2020. 2
  70. 70.Runway. Gen-2: Text driven video generation. https://research.runwayml.com/gen2, 2023. 3, 7
  71. 71.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 3
  72. 72.Uriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual, Iurii Makarov, Filippos Kokkinos, Naman Goyal, Andrea Vedaldi, Devi Parikh, Justin Johnson, et al. Text-to-4d dynamic scene generation. arXiv preprint arXiv:2301.11280, 2023. 2, 3
  73. 73.Harrison Jesse Smith, Qingyuan Zheng, Yifei Li, Somya Jain, and Jessica K Hodgins. A method for animating children’s drawings of the human figure. ACM Transactions on Graphics, 42(3):1–15, 2023. 1, 2, 7
  74. 74.Jifei Song, Kaiyue Pang, Yi-Zhe Song, Tao Xiang, and Timothy Hospedales. Learning to sketch with shortcut cycle consistency, 2018. 2
  75. 75.Qingkun Su, Xue Bai, Hongbo Fu, Chiew-Lan Tai, and Jue Wang. Live sketch: Video-driven dynamic deformation of static drawings. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, pages 1–12, 2018. 1, 2
  76. 76.Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal. Any-to-any generation via composable diffusion. 2023. 3
  77. 77.Matthew Thorne, David Burke, and Michiel Van De Panne. Motion doodles: an interface for sketching character motion. ACM Transactions on Graphics (ToG), 23(3):424–431, 2004. 2
  78. 78.Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N. Metaxas, and Sergey Tulyakov. A good image generator is what you need for high-resolution video synthesis. In International Conference on Learning Representations, 2021. 3
  79. 79.Yael Vinker, Yuval Alaluf, Daniel Cohen-Or, and Ariel Shamir. Clipascene: Scene sketching with different types and levels of abstraction. 2022. 2, 7
  80. 80.Yael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir. Clipasso: Semantically-aware object sketching. ACM Trans. Graph., 41(4), 2022. 2, 7, 8
  81. 81.Jue Wang, Yingqing Xu, Heung-Yeung Shum, and Michael F. Cohen. Video tooning. In ACM SIGGRAPH 2004 Papers, page 574–583, New York, NY, USA, 2004. Association for Computing Machinery. 2
  82. 82.Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. Modelscope text-to-video technical report. arXiv preprint arXiv:2308.06571, 2023. 3, 6, 7
  83. 83.Xiang* Wang, Hangjie* Yuan, Shiwei* Zhang, Dayou* Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. Videocomposer: Compositional video synthesis with motion controllability. arXiv preprint arXiv:2306.02018, 2023. 2, 4
  84. 84.Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023. 3
  85. 85.Dirk Weissenborn, Oscar Täckström, and Jakob Uszkoreit. Scaling autoregressive video models. arXiv preprint arXiv:1906.02634, 2019. 3
  86. 86.Chung-Yi Weng, Brian Curless, and Ira Kemelmacher-Shlizerman. Photo wake-up: 3d character animation from a single photo. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5901–5910, 2018. 2
  87. 87.Chung-Yi Weng, Brian Curless, and Ira Kemelmacher-Shlizerman. Photo wake-up: 3d character animation from a single photo. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5908–5917, 2019. 1
  88. 88.Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. Godiva: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021. 3
  89. 89.Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nüwa: Visual synthesis pre-training for neural visual world creation. In European conference on computer vision, pages 720–736. Springer, 2022. 3
  90. 90.Saining Xie and Zhuowen Tu. Holistically-nested edge detection. In Proceedings of the IEEE international conference on computer vision, pages 1395–1403, 2015. 2
  91. 91.Jun Xing, Li-Yi Wei, Takaaki Shiratori, and Koji Yatani. Autocomplete hand-drawn animations. ACM Trans. Graph., 34(6), 2015. 2
  92. 92.Jun Xing, Rubaiat Kazi, Tovi Grossman, Li-Yi Wei, Jos Stam, and George Fitzmaurice. Energy-brushes: Interactive tools for illustrating stylized elemental dynamics. pages 755–766, 2016. 2
  93. 93.Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Xintao Wang, Tien-Tsin Wong, and Ying Shan. Dynamicrafter: Animating open-domain images with video diffusion priors. arXiv preprint arXiv:2310.12190, 2023. 2
  94. 94.Peng Xu, Timothy M. Hospedales, Qiyue Yin, Yi-Zhe Song, Tao Xiang, and Liang Wang. Deep learning for free-hand sketch: A survey and a toolbox, 2020. 2
  95. 95.Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021. 3
  96. 96.Ran Yi, Yong-Jin Liu, Yu-Kun Lai, and Paul L Rosin. Unpaired portrait drawing generation via asymmetric cycle mapping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8217–8225, 2020. 2
  97. 97.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. 2022. 7
  98. 98.Sharon Zhang, Jiaju Ma, Jiajun Wu, Daniel Ritchie, and Maneesh Agrawala. Editing motion graphics video via motion vectorization and transformation. ACM Transactions on Graphics, 42(6):1–13, 2023. 2
  99. 99.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models, 2023. 3
  100. 100.Jingyuan Zhu, Huimin Ma, Jiansheng Chen, and Jian Yuan. Motionvideogan: A novel video generator based on the motion space learned from image pairs. IEEE Transactions on Multimedia, 2023. 3

Citation

MLA
Gal, R., et al. “Breathing Life Into Sketches Using Text-to-Video Priors”. arXiv, 2023, http://arxiv.org/abs/2311.13608v1.
APA
Gal, R., Vinker, Y., Alaluf, Y., Bermano, A. H., Cohen-Or, D., Shamir, A., & Chechik, G. (2023). Breathing Life Into Sketches Using Text-to-Video Priors. arXiv. http://arxiv.org/abs/2311.13608v1
Chicago
Gal, R., Y. Vinker, Y. Alaluf, et al. 2023. “Breathing Life Into Sketches Using Text-to-Video Priors”. arXiv. http://arxiv.org/abs/2311.13608v1.
Harvard
Gal, R. et al. (2023) “Breathing Life Into Sketches Using Text-to-Video Priors”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.13608v1.
Vancouver
1. Gal R, Vinker Y, Alaluf Y, Bermano AH, Cohen-Or D, Shamir A, Chechik G (2023) Breathing Life Into Sketches Using Text-to-Video Priors. arXiv

BibTeX

@article{gal2023breathing,
  title = {Breathing Life Into Sketches Using Text-to-Video Priors},
  author = {Gal, Rinon and Vinker, Yael and Alaluf, Yuval and Bermano, Amit H. and Cohen-Or, Daniel and Shamir, Ariel and Chechik, Gal},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.13608v1},
  eprint = {2311.13608}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE