Shape-Aware Text-Driven Layered Video Editing
Yao-Chih LeeJi-Ze Genevieve JangYi-Ting ChenElizabeth QiuJia-Bin Huang
Proposes a text-guided layered video editing framework that enables temporally consistent structural and shape modifications by propagating keyframe deformation fields across UV mapping spaces and completing occluded regions using pre-trained diffusion models.
While text-guided image manipulation has advanced rapidly, video editing remains challenging due to the need for temporal consistency across frames. Prior state-of-the-art video editing approaches decompose footage into unified 2D texture maps (atlases) with fixed pixel coordinate mapping fields to maintain consistency. However, because these coordinate mapping fields store the object's original geometry, existing frameworks are fundamentally restricted to changing surface appearance (such as color or texture) and fail when user prompts require structural or shape modifications.
The main objective of the article is to demonstrate a shape-aware, text-driven video editing method that simultaneously modifies both the geometric shape and appearance of a foreground object while preserving original motion dynamics and temporal consistency across sequential frames.
The proposed approach combines layered neural representations with image-level generative models. The process begins by decomposing an input video into foreground and background atlas layers. An off-the-shelf text-to-image diffusion model modifies a single representative keyframe to match a target text prompt. The method then extracts dense semantic correspondences between the original and edited keyframes to derive a pixel-level shape deformation field, projecting these shifts back onto the atlas representation to construct deformed per-frame coordinate maps. To address missing pixels from unseen viewpoints and correct noisy semantic alignments, the framework performs a test-time optimization across 3 to 5 viewpoint frames for 600 to 1,000 iterations (taking roughly 20 minutes on an A5000 GPU) using score distillation guidance from a pre-trained diffusion model alongside mask and smoothness constraints.
The evaluation yields several key findings: First, the method successfully executes complex shape transformations—such as converting a boat into a yacht or a dog into a cat—while matching target text prompts and preserving realistic video motion. Second, it resolves the limitations of existing layered editing frameworks, which preserve temporal stability but remain trapped in the original object's silhouette. Third, it avoids the severe temporal flickering seen in independent multi-frame baseline models and eliminates the visible propagation distortions found in single-frame flow-propagation methods. Fourth, interpolating the learned deformation maps enables smooth, gradual transitions between distinct object shapes across video sequences.
These findings demonstrate that generative video editing can achieve geometric modifications without requiring full 3D modeling or computationally massive video-diffusion training. This capability significantly expands automated video manipulation pipelines for creative production, lowering the time and manual skill barriers required for complex visual effects. Practitioners deploying such workflows must consider operational governance, including licensing and ethical controls, to mitigate the risks of synthetic video misuse.
Organizations evaluating this technology should integrate shape deformation modules into layered video workflows and leverage pre-trained 2D generative models for automated frame completion. When deploying into production pipelines, teams should incorporate optional user-guided correspondence tools to allow manual correction of alignment errors during initial keyframe matching.
The approach exhibits clear operational boundaries. Its success relies heavily on accurate initial layered video decomposition and reasonable initial semantic correspondence; severe mapping failures occur during highly complex motions (such as crossed limbs) or extreme shape discrepancies. While confidence in the demonstrated benchmark cases is high, users should apply caution when processing complex, non-rigid, multi-object interactions.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). It introduces cross-attention control mechanisms in diffusion models that form the foundational paradigm for text-guided editing without structural collapse.
- Paper: Null-text Inversion for Editing Real Images using Guided Diffusion Models, Ron Mokady et al. (2022). It establishes pivotal null-text inversion for real image editing, providing the necessary prerequisite for faithfully inverting and guiding real keyframes with diffusion models.
- Paper: Imagic: Text-Based Real Image Editing with Diffusion Models, Bahjat Kawar et al. (2022). It presents methods for non-rigid, complex semantic edits on single images using diffusion models, motivating the keyframe editing strategies extended to video sequences.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). It introduces SDE-based stochastic guided denoising for stroke and distortion refinement, underpinning diffusion-guided completion of deformed regions.
- Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). It details blended diffusion for text-driven local manipulation and inpainting, which is foundational for completing and blending disoccluded or deformed video areas.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It establishes the core architectural principles of video diffusion models that enable temporal consistency in generative video modeling.
- Paper: RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models, Ozgur Kara et al. (2024). It advances text-driven video editing by introducing noise shuffling across multi-frame grids to achieve temporally consistent edits efficiently without heavy deformation pipelines.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). It expands diffusion-guided shape and region editing by introducing image-prompt encoding and deterministic rollback for precise local structural manipulations.
- Paper: VideoBooth: Diffusion-based Video Generation with Image Prompts, Yuming Jiang et al. (2024). It extends customized diffusion-based video manipulation by using image prompts and cross-frame attention injection to preserve consistent identity and fine geometry across frames.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). It builds upon space-time video synthesis concepts to provide a holistic space-time U-Net framework capable of direct text-to-video generation and consistent video editing.
- Paper: AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. (2024). It generalizes text-guided video generation by integrating plug-and-play motion modules into personalized diffusion models without per-case motion field fitting.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). It develops training-free motion cloning by isolating temporal attention maps to achieve versatile, motion-controlled video synthesis.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). It advances controllable video synthesis by incorporating trajectory-oriented diffusion transformers for precise motion and shape guidance across long sequences.
