Shape-Aware Text-Driven Layered Video Editing

Yao-Chih LeeJi-Ze Genevieve JangYi-Ting ChenElizabeth QiuJia-Bin Huang

article2023CVPR61 citations

Proposes a text-guided layered video editing framework that enables temporally consistent structural and shape modifications by propagating keyframe deformation fields across UV mapping spaces and completing occluded regions using pre-trained diffusion models.

Listen

While text-guided image manipulation has advanced rapidly, video editing remains challenging due to the need for temporal consistency across frames. Prior state-of-the-art video editing approaches decompose footage into unified 2D texture maps (atlases) with fixed pixel coordinate mapping fields to maintain consistency. However, because these coordinate mapping fields store the object's original geometry, existing frameworks are fundamentally restricted to changing surface appearance (such as color or texture) and fail when user prompts require structural or shape modifications.

The main objective of the article is to demonstrate a shape-aware, text-driven video editing method that simultaneously modifies both the geometric shape and appearance of a foreground object while preserving original motion dynamics and temporal consistency across sequential frames.

The proposed approach combines layered neural representations with image-level generative models. The process begins by decomposing an input video into foreground and background atlas layers. An off-the-shelf text-to-image diffusion model modifies a single representative keyframe to match a target text prompt. The method then extracts dense semantic correspondences between the original and edited keyframes to derive a pixel-level shape deformation field, projecting these shifts back onto the atlas representation to construct deformed per-frame coordinate maps. To address missing pixels from unseen viewpoints and correct noisy semantic alignments, the framework performs a test-time optimization across 3 to 5 viewpoint frames for 600 to 1,000 iterations (taking roughly 20 minutes on an A5000 GPU) using score distillation guidance from a pre-trained diffusion model alongside mask and smoothness constraints.

The evaluation yields several key findings: First, the method successfully executes complex shape transformations—such as converting a boat into a yacht or a dog into a cat—while matching target text prompts and preserving realistic video motion. Second, it resolves the limitations of existing layered editing frameworks, which preserve temporal stability but remain trapped in the original object's silhouette. Third, it avoids the severe temporal flickering seen in independent multi-frame baseline models and eliminates the visible propagation distortions found in single-frame flow-propagation methods. Fourth, interpolating the learned deformation maps enables smooth, gradual transitions between distinct object shapes across video sequences.

These findings demonstrate that generative video editing can achieve geometric modifications without requiring full 3D modeling or computationally massive video-diffusion training. This capability significantly expands automated video manipulation pipelines for creative production, lowering the time and manual skill barriers required for complex visual effects. Practitioners deploying such workflows must consider operational governance, including licensing and ethical controls, to mitigate the risks of synthetic video misuse.

Organizations evaluating this technology should integrate shape deformation modules into layered video workflows and leverage pre-trained 2D generative models for automated frame completion. When deploying into production pipelines, teams should incorporate optional user-guided correspondence tools to allow manual correction of alignment errors during initial keyframe matching.

The approach exhibits clear operational boundaries. Its success relies heavily on accurate initial layered video decomposition and reasonable initial semantic correspondence; severe mapping failures occur during highly complex motions (such as crossed limbs) or extreme shape discrepancies. While confidence in the demonstrated benchmark cases is high, users should apply caution when processing complex, non-rigid, multi-object interactions.

arXiv: 2301.13173
Cover for Shape-Aware Text-Driven Layered Video Editing

Abstract

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than object shape changes due to the limitation of using a fixed UV mapping field for texture atlas. We present a shape-aware, text-driven video editing method to tackle this challenge. To handle shape changes in video editing, we first propagate the deformation field between the input and edited keyframe to all frames. We then leverage a pre-trained text-conditioned diffusion model as guidance for refining shape distortion and completing unseen regions. The experimental results demonstrate that our method can achieve shape-aware consistent video editing and compare favorably with the state-of-the-art.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Keyframe editing
  • 3.2. Deformation formulation
  • 3.3. Atlas optimization
  • 4. Experimental Results
  • 4.1. Experimental Setup
  • 4.2. Visual Comparison
  • 4.3. Ablation Study
  • 4.4. Application
  • 4.5. Limitations
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Shape-aware layered video editing pipeline

    model/method

    The method edits a source video I1:NsI^s_{1:N} containing NN frames according to a text prompt and produces a target video I1:NtI^t_{1:N} with both appearance and object-shape changes while preserving the source motion. It assumes one moving foreground object and uses a pretrained Neural Layered Atlas (NLA) decomposition to obtain canonical foreground and background atlases plus per-frame UV mappings.

    The pipeline first selects one representative source keyframe IksI^s_k, edits it with a text-to-image diffusion model such as Stable Diffusion to obtain IktI^t_k, and estimates dense semantic correspondence between the source and edited keyframes. The correspondence initializes a deformation between the source and target shapes. The edited keyframe appearance and deformation are mapped into atlas space, propagated through the video, and used to deform the original per-frame UV maps and alpha masks. Finally, an atlas network and a semantic-correspondence network are optimized using text-conditioned diffusion guidance and reconstruction, smoothness, and mask-consistency losses. The method overview diagram on page 4 shows this separation between NLA preprocessing, keyframe editing, deformation initialization, and atlas optimization.

  2. Knowl 2 — NLA atlas representation and source-video rendering

    equation

    For a source frame IjsI^s_j at time j∈{1,…,N}j\in\{1,\ldots,N\}, the NLA representation contains a foreground atlas IAs,FGI_A^{s,\mathrm{FG}}, a background atlas IAs,BGI_A^{s,\mathrm{BG}}, foreground and background UV fields WA→js,gW^{s,g}_{A\to j}, and a foreground alpha mask αjs\alpha^s_j, where g∈{FG,BG}g\in\{\mathrm{FG},\mathrm{BG}\}. The rendered source frame is

    Ijs=Ijs,FG ∗ αjs+Ijs,BG ∗ (1−αjs),Ijs,g=WA→js,g⊗IAs,g.I^s_j=I^{s,\mathrm{FG}}_j\,*\,\alpha^s_j+I^{s,\mathrm{BG}}_j\,*\,(1-\alpha^s_j),\qquad I^{s,g}_j=W^{s,g}_{A\to j}\otimes I^{s,g}_A.

    Here ∗* denotes elementwise multiplication, ⊗\otimes denotes sampling or warping an atlas with a UV field, and AA denotes canonical atlas coordinates. The UV fields encode how atlas content is placed in each frame, so editing a fixed atlas can provide temporal consistency. However, fixed source UV fields preserve the original object geometry and therefore cannot by themselves represent a target object with a different shape.

  3. Knowl 3 — Keyframe editing and source-shaped atlas initialization

    model/method

    The selected source keyframe IksI^s_k is edited by a pretrained text-conditioned diffusion image editor, producing a target keyframe IktI^t_k. A pretrained dense semantic-correspondence model estimates a pixelwise displacement field Dkt→s∈RH×W×2D^{t\to s}_k\in\mathbb{R}^{H\times W\times 2}, where HH and WW are the keyframe height and width, that maps the target object into the coordinate system of the source object.

    The target keyframe is first warped into the source shape:

    Ikt→s=Dkt→s⊗Ikt.I^{t\to s}_k=D^{t\to s}_k\otimes I^t_k.

    The source-shaped edited keyframe is then back-projected into atlas coordinates using the original source UV map Wk→AsW^s_{k\to A}, yielding an edited atlas IAt→sI^{t\to s}_A. This intermediate atlas contains the target appearance while retaining the source atlas geometry, which makes it compatible with the original NLA mapping. The target shape is restored later by propagating the displacement field and deforming the per-frame UV maps.

  4. Knowl 4 — UV deformation propagation for target-shape rendering

    equation

    The method transports the keyframe deformation through atlas space while accounting for the fact that a warp changes the coordinate basis. Let WW be a smooth warp mapping a source pixel (x,y)(x,y) to (x′,y′)=W(x,y)(x',y')=W(x,y), and let D(x,y)∈R2D(x,y)\in\mathbb{R}^2 be a displacement vector in source coordinates. The corresponding displacement in warped coordinates is defined by

    D′(x′,y′)=MW(x,y)D(x,y),D'(x',y')=M_W(x,y)D(x,y),

    where the local vector-transformation matrix is estimated with small scalar offsets Δx\Delta x and Δy\Delta y as

    MW(x,y)=[(W(x+Δx,y)−W(x,y))⊤Δx(W(x,y+Δy)−W(x,y))⊤Δy].M_W(x,y)= \begin{bmatrix} \dfrac{\big(W(x+\Delta x,y)-W(x,y)\big)^\top}{\Delta x}\\[4pt] \dfrac{\big(W(x,y+\Delta y)-W(x,y)\big)^\top}{\Delta y} \end{bmatrix}.

    In practice, thin-plate splines provide a smooth approximation of the warp. With ⋆\star denoting pointwise matrix-vector multiplication and ⊗\otimes denoting warping, the keyframe displacement is propagated first to the atlas and then to frame jj:

    DAt→s=MWk→As⋆(Wk→As⊗Dkt→s),D^{t\to s}_A=M_{W^s_{k\to A}}\star\left(W^s_{k\to A}\otimes D^{t\to s}_k\right), Djt→s=MWA→js⋆(WA→js⊗DAt→s).D^{t\to s}_j=M_{W^s_{A\to j}}\star\left(W^s_{A\to j}\otimes D^{t\to s}_A\right).

    The corresponding forward displacement Djs→tD^{s\to t}_j deforms the source UV field and alpha mask:

    WA→jt=Djs→t⊗WA→js,αjt=Djs→t⊗αjs.W^t_{A\to j}=D^{s\to t}_j\otimes W^s_{A\to j},\qquad \alpha^t_j=D^{s\to t}_j\otimes\alpha^s_j.

    Using the deformed UV field, the initially edited target frame is rendered as

    Ijt=(WA→jt⊗IAt→s)∗αjt+IABG∗(1−αjt).I^t_j=\left(W^t_{A\to j}\otimes I^{t\to s}_A\right)*\alpha^t_j+I^{\mathrm{BG}}_A*(1-\alpha^t_j).

    Thus, appearance is propagated through the edited atlas while the UV sampling geometry and foreground mask are simultaneously changed to follow the target shape.

  5. Knowl 5 — Diffusion-guided atlas and correspondence optimization

    model/method

    The initial deformation can be inaccurate, and a single edited keyframe leaves unseen parts of the edited atlas incomplete. The method therefore introduces an atlas-refinement network FθAF_{\theta_A} and a semantic-correspondence refinement network FθSCF_{\theta_{SC}}. The atlas network takes the initial appearance and deformation atlases (IAt→s,DAt→s)(I^{t\to s}_A,D^{t\to s}_A) and outputs refined quantities (I~At→s,D~At→s)(\widetilde I^{t\to s}_A,\widetilde D^{t\to s}_A). The correspondence network refines the thin-plate-spline correspondence Dkt→sD^{t\to s}_k from its control points. The parameter set is θ={θA,θSC}\theta=\{\theta_A,\theta_{SC}\}.

    For a rendered edited frame ItI^t, Gaussian noise ϵ\epsilon is added at diffusion timestep ii. A frozen pretrained diffusion U-Net predicts noise ϵ^\hat\epsilon. The noise-residual guidance used to update the trainable parameters is

    ∇θLdiff(It)≜Ei,ϵ[w(i)(ϵ^−ϵ)∂It∂θ],\nabla_{\theta}\mathcal{L}_{\mathrm{diff}}(I^t) \triangleq \mathbb{E}_{i,\epsilon}\left[w(i)(\hat\epsilon-\epsilon)\frac{\partial I^t}{\partial\theta}\right],

    where w(i)w(i) is the timestep-dependent diffusion weight. Gradients are not backpropagated through the U-Net. Only a few frames with different viewpoints are rendered during optimization, while the shared atlas parameters enforce consistency across all frames.

    The auxiliary losses preserve the edited keyframe, smooth the edited atlas, and enforce semantic mask alignment. With I~kt\widetilde I^t_k the reconstructed edited keyframe, Ltv\mathcal{L}_{tv} the total-variation operator, and Mkt,MksM^t_k,M^s_k the target and source object masks, they are

    Lk=∥I~kt−Ikt∥1,\mathcal{L}_k=\left\|\widetilde I^t_k-I^t_k\right\|_1, LA=Ltv(I~At→s),\mathcal{L}_A=\mathcal{L}_{tv}\left(\widetilde I^{t\to s}_A\right), LSC=∥D~kt→s⊗Mkt−Mks∥1.\mathcal{L}_{SC}=\left\|\widetilde D^{t\to s}_k\otimes M^t_k-M^s_k\right\|_1.

    The total objective is

    L=Ldiff+λkLk+λALA+λSCLSC,\mathcal{L}=\mathcal{L}_{\mathrm{diff}}+\lambda_k\mathcal{L}_k+\lambda_A\mathcal{L}_A+\lambda_{SC}\mathcal{L}_{SC},

    with λk=106\lambda_k=10^6, λA=103\lambda_A=10^3, and λSC=103\lambda_{SC}=10^3. The optimized networks produce the final edited video, filling unseen atlas regions and correcting deformation artifacts while retaining temporal consistency.

  6. Knowl 6 — Implementation and evaluation protocol

    experimental setup

    The method is implemented in PyTorch at a video resolution of 768×432768\times432 pixels. Thin-plate splines are used to invert warp fields and avoid holes caused by forward warping. The atlas-refinement network follows the architecture used by Text2LIVE, and the semantic-correspondence refinement network follows TPS-STN. Optimization uses 3--5 selected frames, including the first frame, the edited keyframe, and the last frame, for 600--1000 iterations. A run takes approximately 20 minutes on a 24-GB NVIDIA A5000 GPU, and an off-the-shelf super-resolution model is applied to sharpen the final edited atlases.

    For evaluation, the authors select DAVIS videos containing one moving object over 50--70 frames and choose prompts describing objects with shapes different from the source objects. All compared methods use the same Stable Diffusion image editor. The multi-frame baseline edits multiple keyframes independently and uses FILM to interpolate nearby frames; the single-frame baseline edits one keyframe and propagates it with EbSynth; and Text2LIVE applies NLA-based text-driven atlas editing with a structure-preservation loss.

  7. Knowl 7 — Qualitative superiority in shape-aware video editing

    empirical result

    Across the qualitative comparisons, the proposed method changes both texture and object geometry while maintaining coherent motion over time. In the examples shown in the visual-comparison grid on page 7, the edits are black swan to duck, boat to yacht, and dog to cat. The multi-frame baseline produces inconsistent object edits across frames, while the single-frame baseline propagates source-shape motion incorrectly and introduces distortions. Text2LIVE gives temporally consistent target-like appearance but retains the original source shape. The proposed deformation-plus-atlas-optimization method instead produces plausible target appearance and target shape consistently, including for the non-rigid dog-to-cat sequence.

  8. Knowl 8 — Ablation of correspondence, UV deformation, and atlas optimization

    empirical result

    The ablation sequence on page 7 isolates the contributions of the proposed components. With fixed NLA UV maps, edited keyframe content is not correctly transformed across frames. Adding semantic correspondence while keeping the UV maps fixed maps the edited appearance into the atlas but leaves the rendered object in its original source shape. Adding the UV-deformation module changes the per-frame sampling fields and restores the target shape. This intermediate result can still contain artifacts because the correspondence is noisy and the edited atlas lacks pixels unseen in the single keyframe, notably around regions such as a car roof and back wheel. Adding atlas and correspondence optimization fills missing regions and reduces these distortions, producing the full result.

  9. Knowl 9 — Shape-aware interpolation through deformation maps

    model/method

    The learned representation also supports interpolation between the source and edited shapes without using an additional frame-interpolation model. Interpolating the atlas deformation maps produces a gradual transition from the source object geometry to the target object geometry over time; atlas textures can be interpolated similarly. The examples on page 8 illustrate smooth shape transitions for edited animals and vehicles.

    Background content can be edited directly in the background atlas because it behaves like a natural panorama. Foreground atlas textures are unwrapped object textures and are unsuitable as direct inputs to general-purpose image editors, so the method edits a video keyframe first and maps the result back into the foreground atlas. This allows users to employ arbitrary image-editing tools while retaining temporally consistent layered-video rendering.

  10. Knowl 10 — Dependence on atlas quality and semantic correspondence

    limitation

    The method relies on a many-to-one mapping from individual video frames to a unified NLA atlas. NLA can produce erroneous mappings for complex motions, causing visible artifacts in the corresponding regions; the failure example on page 8 shows distortion around crossing hind legs in a bear-to-lion edit. The method also depends on reliable semantic correspondence between different object categories. Severe false matches can prevent successful optimization even when the refinement networks are used. Manual user correction of the correspondence, illustrated on page 8, can improve the resulting video but introduces additional interaction and does not eliminate the underlying correspondence limitation.

Coverage note — No substantial contributed material was omitted; supplementary video results and the paper’s societal-impact statement were not expanded because they do not add separate load-bearing technical findings.

References

  1. 1.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022. 3
  2. 2.Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Text2live: Text-driven layered image and video editing. ECCV, 2022. 2, 3, 6, 7
  3. 3.Kiran S Bhat, Steven M Seitz, Jessica K Hodgins, and Pradeep K Khosla. Flow-based video synthesis and editing. In ACM SIGGRAPH. 2004. 3
  4. 4.Fred L. Bookstein. Principal warps: Thin-plate splines and the decomposition of deformations. IEEE TPAMI, 1989. 5
  5. 5.Tim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang, Timo Aila, Jaakko Lehtinen, Ming-Yu Liu, Alexei A Efros, and Tero Karras. Generating long videos of dynamic scenes. In NeurIPS, 2022. 3
  6. 6.Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. Vqgan-clip: Open domain image generation and editing with natural language guidance. ECCV, 2022. 2
  7. 7.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. NeurIPS, 2021. 3
  8. 8.Kevin Frans, Lisa B Soros, and Olaf Witkowski. Clipdraw: Exploring text-to-drawing synthesis through language-image encoders. arXiv preprint arXiv:2106.14843, 2021. 3
  9. 9.Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Stylegan-nada: Clip-guided domain adaptation of image generators. ACM TOG, 2022. 3
  10. 10.Chen Gao, Ayush Saraf, Jia-Bin Huang, and Johannes Kopf. Flow-edge guided video completion. In ECCV, 2020. 3
  11. 11.Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. In ECCV, 2022. 3
  12. 12.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. 3
  13. 13.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 3
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. NeurIPS, 2020. 3
  15. 15.Jia-Bin Huang, Sing Bing Kang, Narendra Ahuja, and Johannes Kopf. Temporally coherent completion of dynamic video. ACM Transactions on Graphics (TOG), 35(6):1–11, 2016. 3
  16. 16.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. NeurIPS, 2015. 6
  17. 17.Ondˇrej Jamriska, ˇ Sˇ arka Sochorov ´ a, Ond ´ ˇrej Texler, Michal Luka´c, Jakub Fi ˇ ser, Jingwan Lu, Eli Shechtman, and Daniel ˇ Sykora. Stylizing video by example. ` ACM TOG, 2019. 2, 3, 6
  18. 18.Yoni Kasten, Dolev Ofri, Oliver Wang, and Tali Dekel. Layered neural atlases for consistent video editing. ACM TOG, 2021. 1, 3, 4
  19. 19.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. arXiv preprint arXiv:2210.09276, 2022. 1
  20. 20.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In CVPR, 2022. 1
  21. 21.Gihyun Kwon and Jong Chul Ye. Clipstyler: Image style transfer with a single text condition. In CVPR, 2022. 3
  22. 22.Wei-Sheng Lai, Jia-Bin Huang, Oliver Wang, Eli Shechtman, Ersin Yumer, and Ming-Hsuan Yang. Learning blind video temporal consistency. In ECCV, 2018. 3
  23. 23.Chenyang Lei, Yazhou Xing, Hao Ouyang, and Qifeng Chen. Deep video prior for video consistency and propagation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022. 3
  24. 24.Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip HS Torr. Manigan: Text-guided image manipulation. In CVPR, 2020. 1, 2
  25. 25.Wenbo Li, Pengchuan Zhang, Lei Zhang, Qiuyuan Huang, Xiaodong He, Siwei Lyu, and Jianfeng Gao. Object-driven text-to-image synthesis via adversarial training. In CVPR, 2019. 2
  26. 26.Zhen Li, Cheng-Ze Lu, Jianhua Qin, Chun-Le Guo, and Ming-Ming Cheng. Towards an end-to-end framework for flow-guided video inpainting. In CVPR, 2022. 3
  27. 27.Wentong Liao, Kai Hu, Michael Ying Yang, and Bodo Rosenhahn. Text to image generation with semantic-spatial aware gan. In CVPR, 2022. 2
  28. 28.Sharon Lin, Matthew Fisher, Angela Dai, and Pat Hanrahan. Layerbuilder: Layer decomposition for interactive image and video color editing. arXiv preprint arXiv:1701.03754, 2017. 3
  29. 29.Feng-Lin Liu, Shu-Yu Chen, Yukun Lai, Chunpeng Li, Yue-Ren Jiang, Hongbo Fu, and Lin Gao. Deepfacevideoediting: Sketch-based deep editing of face videos. ACM TOG, 2022. 3
  30. 30.Xingchao Liu, Chengyue Gong, Lemeng Wu, Shujian Zhang, Hao Su, and Qiang Liu. Fusedream: Training-free text-to-image generation with improved clip+ gan space optimization. arXiv preprint arXiv:2112.01573, 2021. 2
  31. 31.Sebastian Loeschcke, Serge Belongie, and Sagie Benaim. Text-driven stylization of video objects. arXiv preprint arXiv:2206.12396, 2022. 3
  32. 32.Erika Lu, Forrester Cole, Tali Dekel, Weidi Xie, Andrew Zisserman, David Salesin, William T Freeman, and Michael Rubinstein. Layered neural rendering for retiming people in video. ACM TOG, 2020. 3
  33. 33.Erika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman, William T Freeman, and Michael Rubinstein. Omnimatte: Associating objects and their effects in video. In CVPR, 2021. 3
  34. 34.Seonghyeon Nam, Yunji Kim, and Seon Joo Kim. Text-adaptive generative adversarial networks: manipulating images with natural language. NeurIPS, 2018. 1
  35. 35.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In ICML, 2022. 3
  36. 36.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In ICCV, 2021. 3
  37. 37.Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbelaez, Alex Sorkine-Hornung, and Luc Van Gool. The 2017 davis challenge on video object segmentation. arXiv preprint arXiv:1704.00675, 2017. 6
  38. 38.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 2, 5
  39. 39.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 2
  40. 40.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML. PMLR, 2021. 1
  41. 41.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021. 2
  42. 42.Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, and Brian Curless. Film: Frame interpolation for large motion. ECCV, 2022. 2, 6
  43. 43.Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis. In ICML. PMLR, 2016. 2
  44. 44.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 1, 2, 3, 4, 5, 6
  45. 45.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242, 2022. 3
  46. 46.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022. 3
  47. 47.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022. 3
  48. 48.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 3
  49. 49.Ivan Skorokhodov, Sergey Tulyakov, and Mohamed Elhoseiny. Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. In CVPR, 2022. 3
  50. 50.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 3
  51. 51.Prune Truong, Martin Danelljan, Fisher Yu, and Luc Van Gool. Probabilistic warp consistency for weakly-supervised semantic correspondences. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8708–8718, 2022. 2, 4
  52. 52.Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description. arXiv preprint arXiv:2210.02399, 2022. 3
  53. 53.Xintao Wang, Liangbin Xie, Chao Dong, and Ying Shan. Real-esrgan: Training real-world blind super-resolution with pure synthetic data. In International Conference on Computer Vision Workshops (ICCVW). 6
  54. 54.Weihao Xia, Yujiu Yang, and Jing-Hao Xue. Gan inversion for consistent video interpolation and manipulation. arXiv preprint arXiv:2208.11197, 2022. 3
  55. 55.Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu. Tedigan: Text-guided diverse face image generation and manipulation. In CVPR, 2021. 2
  56. 56.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In CVPR, 2018. 2
  57. 57.Yiran Xu, Badour AlBahar, and Jia-Bin Huang. Temporally consistent semantic video editing. arXiv preprint arXiv:2206.10590, 2022. 3
  58. 58.Zipeng Xu, Tianwei Lin, Hao Tang, Fu Li, Dongliang He, Nicu Sebe, Radu Timofte, Luc Van Gool, and Errui Ding. Predict, prevent, and evaluate: Disentangled text-driven image manipulation empowered by pre-trained vision-language model. In CVPR, 2022. 3
  59. 59.Vickie Ye, Zhengqi Li, Richard Tucker, Angjoo Kanazawa, and Noah Snavely. Deformable sprites for unsupervised video decomposition. In CVPR, 2022. 3
  60. 60.Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. Generating videos with dynamics-aware implicit generative adversarial networks. In ICLR, 2022. 3
  61. 61.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In ICCV, 2017. 2

Citation

MLA
Lee, Y.-C., et al. “Shape-aware Text-driven Layered Video Editing”. arXiv, 2023, http://arxiv.org/abs/2301.13173v1.
APA
Lee, Y.-C., Jang, J.-Z. G., Chen, Y.-T., Qiu, E., & Huang, J.-B. (2023). Shape-aware Text-driven Layered Video Editing. arXiv. http://arxiv.org/abs/2301.13173v1
Chicago
Lee, Y.-C., J.-Z. G. Jang, Y.-T. Chen, E. Qiu, and J.-B. Huang. 2023. “Shape-aware Text-driven Layered Video Editing”. arXiv. http://arxiv.org/abs/2301.13173v1.
Harvard
Lee, Y.-C. et al. (2023) “Shape-aware Text-driven Layered Video Editing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.13173v1.
Vancouver
1. Lee Y-C, Jang J-ZG, Chen Y-T, Qiu E, Huang J-B (2023) Shape-aware Text-driven Layered Video Editing. arXiv

BibTeX

@article{lee2023shape,
  title = {Shape-aware Text-driven Layered Video Editing},
  author = {Lee, Yao-Chih and Jang, Ji-Ze Genevieve and Chen, Yi-Ting and Qiu, Elizabeth and Huang, Jia-Bin},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.13173v1},
  eprint = {2301.13173}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE