Boximator: Generating Rich and Controllable Motions for Video Synthesis

Jiawei WangYuchen ZhangJiaxin ZouYan ZengGuoqiang WeiLiping YuanHang Li

article2024ICML117 citations

Proposes Boximator, a plug-in for video diffusion models that enables precise, visual motion control across frames using flexible bounding-box constraints and a novel self-tracking training technique.

Listen

Generating rich and precise object motion remains a central bottleneck in generative video synthesis. Current video diffusion models excel at producing realistic visuals from text or reference images, but they struggle to provide fine-grained, frame-by-frame control over where objects move, how their shapes change, or how multiple elements interact across time without requiring tedious textual descriptions for every entity.

The article demonstrates Boximator, an architecture designed to provide fine-grained motion control for video diffusion models using bounding-box constraints. It aims to establish a visually grounded, plug-and-play mechanism that enables precise object selection and trajectory definition while preserving the generative quality and pre-existing knowledge of underlying base models.

The authors implemented Boximator by introducing a lightweight control module integrated into the spatial attention layers of existing video models, keeping the original base model parameters completely frozen. The system uses two types of box constraints: hard boxes for exact boundaries and flexible soft boxes for approximate regions and motion paths. To resolve the optimization challenge of linking discrete coordinate signals to visual objects, the researchers developed an intermediate training technique called self-tracking, which trains the model to generate and track visible colored bounding boxes before disabling their visual appearance in the final stage. The system was trained on a curated dataset of 1.1 million dynamic video clips containing 2.4 million tracked objects and evaluated across standard benchmarks including MSR-VTT, ActivityNet, and UCF-101, alongside human blind comparisons.

The primary findings show significant gains in both motion precision and visual quality. Motion alignment precision improved drastically with box constraints, achieving a 1.9- to 3.7-fold increase in mean average precision on MSR-VTT and a 4.4- to 8.9-fold increase on highly dynamic ActivityNet sequences. Video quality scores improved notably over the base models (lowering Fréchet Video Distance from 237 to 174 on PixelDance and 239 to 216 on ModelScope), while adding less than 20% computational overhead in model parameters and inference latency. In blind human evaluations, raters preferred Boximator's motion control in 76.0% of cases compared to only 2.2% for the base model, and favored its overall video quality by a margin of 35.2% to 16.8%.

These results demonstrate that fine-grained motion control can be added to existing video generation platforms efficiently without expensive full-model retraining or degraded output quality. By enabling direct visual selection and trajectory shaping, Boximator lowers the operational barrier for complex video generation workflows, providing a predictable tool for applications in content creation, animation, and digital media production.

Organizations developing or deploying video generation tools should consider adopting visual box conditioning and self-tracking training paradigms rather than relying solely on text-prompt engineering. For future work, development efforts should focus on expanding training beyond the current WebVid-derived data to enhance domain generalization, integrating support for longer videos and widescreen aspect ratios, and pairing box conditioning with complementary controls such as text-driven rotations and skeletal pose guidance.

The study's primary limitations stem from its current evaluation scope: generated videos are restricted to 4-second clips at a 256x256 resolution with a 1:1 aspect ratio, and automated evaluation metrics rely on third-party object detectors that may introduce measurement noise. Despite these boundary conditions, the large improvements across automated benchmarks and human side-by-side reviews provide high confidence in Boximator's core control capabilities.

Cover for Boximator: Generating Rich and Controllable Motions for Video Synthesis

Abstract

Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose Boximator, a new approach for fine-grained motion control. Boximator introduces two constraint types: hard box and soft box. Users select objects in the conditional frame using hard boxes and then use either type of boxes to roughly or rigorously define the object’s position, shape, or motion path in future frames. Boximator functions as a plug-in for existing video diffusion models. Its training process preserves the base model’s knowledge by freezing the original weights and training only the control module. To address training challenges, we introduce a novel self-tracking technique that greatly simplifies the learning of box-object correlations. Empirically, Boximator achieves state-of-the-art video quality (FVD) scores, improving on two base models, and further enhanced after incorporating box constraints. Its robust motion controllability is validated by drastic increases in the bounding box alignment metric. Human evaluation also shows that users favor Boximator generation results over the base model.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background: Video Diffusion Model
  • 4. Boximator: Box-guided Motion Control
  • 4.1. Model Architecture
  • 4.2. Data Pipeline
  • 4.3. Self-Tracking
  • 4.4. Multi-Stage Training Procedure
  • 4.5. Inference
  • 5. Experiments
  • 5.1. Experiment Settings
  • 5.2. Quantitative Evaluation
  • 5.3. Human Evaluation
  • 5.4. Ablation Study
  • 5.5. Case Study
  • 6. Conclusion
  • Limitations
  • Impact Statement
  • References
  • A. More Implementation Details
  • B. Results on UCF-101
  • C. Human Evaluation Details

Knowls

  1. Knowl 1 — Paper content unavailable for extraction

    limitation

    The source paper's content was not accessible: the attachment provided only the file name 22cac8ed-0675-4e43-85a3-73967543f721.pdf with no readable text, so no methods, models, theory, experiments, results, or stated limitations of the paper could be identified or extracted. Any knowls about this paper's actual contribution would require the document's text to be supplied.

Coverage note — The attached document could not be read: only the filename '22cac8ed-0675-4e43-85a3-73967543f721.pdf' was provided, with no extractable text or content from the paper itself. No knowls could therefore be extracted from the paper's contribution; no contributed material was deliberately omitted, since none was accessible.

Citation

MLA
Wang, J., et al. “Boximator: Generating Rich and Controllable Motions for Video Synthesis”. arXiv, 2024, http://arxiv.org/abs/2402.01566v1.
APA
Wang, J., Zhang, Y., Zou, J., Zeng, Y., Wei, G., Yuan, L., & Li, H. (2024). Boximator: Generating Rich and Controllable Motions for Video Synthesis. arXiv. http://arxiv.org/abs/2402.01566v1
Chicago
Wang, J., Y. Zhang, J. Zou, et al. 2024. “Boximator: Generating Rich and Controllable Motions for Video Synthesis”. arXiv. http://arxiv.org/abs/2402.01566v1.
Harvard
Wang, J. et al. (2024) “Boximator: Generating Rich and Controllable Motions for Video Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.01566v1.
Vancouver
1. Wang J, Zhang Y, Zou J, Zeng Y, Wei G, Yuan L, Li H (2024) Boximator: Generating Rich and Controllable Motions for Video Synthesis. arXiv

BibTeX

@article{wang2024boximator,
  title = {Boximator: Generating Rich and Controllable Motions for Video Synthesis},
  author = {Wang, Jiawei and Zhang, Yuchen and Zou, Jiaxin and Zeng, Yan and Wei, Guoqiang and Yuan, Liping and Li, Hang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.01566v1},
  eprint = {2402.01566}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/