Boximator: Generating Rich and Controllable Motions for Video Synthesis
Jiawei WangYuchen ZhangJiaxin ZouYan ZengGuoqiang WeiLiping YuanHang Li
Proposes Boximator, a plug-in for video diffusion models that enables precise, visual motion control across frames using flexible bounding-box constraints and a novel self-tracking training technique.
Generating rich and precise object motion remains a central bottleneck in generative video synthesis. Current video diffusion models excel at producing realistic visuals from text or reference images, but they struggle to provide fine-grained, frame-by-frame control over where objects move, how their shapes change, or how multiple elements interact across time without requiring tedious textual descriptions for every entity.
The article demonstrates Boximator, an architecture designed to provide fine-grained motion control for video diffusion models using bounding-box constraints. It aims to establish a visually grounded, plug-and-play mechanism that enables precise object selection and trajectory definition while preserving the generative quality and pre-existing knowledge of underlying base models.
The authors implemented Boximator by introducing a lightweight control module integrated into the spatial attention layers of existing video models, keeping the original base model parameters completely frozen. The system uses two types of box constraints: hard boxes for exact boundaries and flexible soft boxes for approximate regions and motion paths. To resolve the optimization challenge of linking discrete coordinate signals to visual objects, the researchers developed an intermediate training technique called self-tracking, which trains the model to generate and track visible colored bounding boxes before disabling their visual appearance in the final stage. The system was trained on a curated dataset of 1.1 million dynamic video clips containing 2.4 million tracked objects and evaluated across standard benchmarks including MSR-VTT, ActivityNet, and UCF-101, alongside human blind comparisons.
The primary findings show significant gains in both motion precision and visual quality. Motion alignment precision improved drastically with box constraints, achieving a 1.9- to 3.7-fold increase in mean average precision on MSR-VTT and a 4.4- to 8.9-fold increase on highly dynamic ActivityNet sequences. Video quality scores improved notably over the base models (lowering Fréchet Video Distance from 237 to 174 on PixelDance and 239 to 216 on ModelScope), while adding less than 20% computational overhead in model parameters and inference latency. In blind human evaluations, raters preferred Boximator's motion control in 76.0% of cases compared to only 2.2% for the base model, and favored its overall video quality by a margin of 35.2% to 16.8%.
These results demonstrate that fine-grained motion control can be added to existing video generation platforms efficiently without expensive full-model retraining or degraded output quality. By enabling direct visual selection and trajectory shaping, Boximator lowers the operational barrier for complex video generation workflows, providing a predictable tool for applications in content creation, animation, and digital media production.
Organizations developing or deploying video generation tools should consider adopting visual box conditioning and self-tracking training paradigms rather than relying solely on text-prompt engineering. For future work, development efforts should focus on expanding training beyond the current WebVid-derived data to enhance domain generalization, integrating support for longer videos and widescreen aspect ratios, and pairing box conditioning with complementary controls such as text-driven rotations and skeletal pose guidance.
The study's primary limitations stem from its current evaluation scope: generated videos are restricted to 4-second clips at a 256x256 resolution with a 1:1 aspect ratio, and automated evaluation metrics rely on third-party object detectors that may introduce measurement noise. Despite these boundary conditions, the large improvements across automated benchmarks and human side-by-side reviews provide high confidence in Boximator's core control capabilities.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). ControlNet establishes how trainable control modules can steer a frozen diffusion model with spatial conditions, the plug-in design principle Boximator adapts to video.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). Stable Video Diffusion introduces the latent video-diffusion foundation and frozen-backbone training strategy that Boximator uses as its generative setting.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Video Diffusion Models lays out the diffusion-based video-generation foundations needed to understand the base models Boximator controls.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). Align Your Latents explains how image latent-diffusion models are adapted for video, providing context for Boximator’s plug-in control of video diffusion.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). MotionClone carries Boximator’s goal of controllable video motion into training-free transfer of movements from reference clips.
- Paper: MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling, Yifang Men et al. (2025). MIMO extends controllable video synthesis from box-guided motion toward spatially decomposed character motion, occlusion, and scene interaction.
- Paper: Tora: Trajectory-oriented Diffusion Transformer for Video Generation, Zhenghao Zhang et al. (2025). Tora advances trajectory-based motion control by bringing user-specified paths into scalable diffusion-transformer video generation.
