Built independently by an author, for readers. Read the story and support ChapterPal

keyword

controllable video generation

Controllable video generation is an artificial intelligence process in which deep generative models synthesize coherent video sequences while adhering to precise, user-specified conditioning signals and structural constraints. Unlike unconstrained video synthesis that relies solely on high-level text descriptions or random sampling, controllable generation enables fine-grained governance over spatial, temporal, and physical dynamics within the scene. These steering mechanisms can encompass camera trajectories, localized object motions, human pose sequences, depth maps, or frame-by-frame interactive user actions. By incorporating multimodal guidance from reference images, motion vectors, spatial masks, or latent action inputs, these systems ensure that the generated video preserves visual fidelity, semantic alignment, and temporal consistency according to targeted visual behaviors.

2 items

Genie: Generative Interactive Environments

Genie: Generative Interactive Environments

Jake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Bechtle, Feryal M. P. Behbahani, Stephanie C. Y. Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott E. Reed, Jingwei Zhang, Konrad Zolna, Jeff Clune, Nando de Freitas, Satinder Singh, Tim Rocktäschel

OrganizationsGoogleUniversity of British Columbia

Why you should read this

Presents Genie, an 11-billion-parameter foundation world model that learns controllable interactive environments and latent actions directly from unlabeled internet video, enabling users to generate and play virtual worlds from prompts as simple as a single image or sketch.

We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a foundation world model. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis despite training without any ground-truth action labels or other domain-specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future.

Added

2026-09-28

MotionClone: Training-Free Motion Cloning for Controllable Video Generation

MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong, Yuhang Zang, Tong Wu, Huaian Chen, Jiaqi Wang, Yi Jin

OrganizationsShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong UniversityThe Chinese University of Hong KongUniversity of Science and Technology of China

Why you should read this

Proposes MotionClone, a training-free framework that transfers camera and object motions from reference videos to text-to-video and image-to-video generation by extracting motion guidance from sparse temporal attention maps in a single denoising step.

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain motion patterns, resulting in limited flexibility and generalization. In this work, we propose MotionClone, a training-free framework that enables motion cloning from reference videos to versatile motion-controlled video generation, including text-to-video and image-to-video. Based on the observation that the dominant components in temporal-attention maps drive motion synthesis, while the rest mainly capture noisy or very subtle motions, MotionClone utilizes sparse temporal attention weights as motion representations for motion guidance, facilitating diverse motion transfer across varying scenarios. Meanwhile, MotionClone allows for the direct extraction of motion representation through a single denoising step, bypassing the cumbersome inversion processes and thus promoting both efficiency and flexibility. Extensive experiments demonstrate that MotionClone exhibits proficiency in both global camera motion and local object motion, with notable superiority in terms of motion fidelity, textual alignment, and temporal consistency.

Added

2026-09-26