Tora: Trajectory-oriented Diffusion Transformer for Video Generation
Zhenghao ZhangJunchao LiaoMenghao LiZuozhuo DaiBingxue QiuSiyu ZhuLong QinWeizhi Wang
Presents Tora, the first trajectory-controlled Diffusion Transformer framework that combines text, visual, and path guidance to generate videos across variable aspect ratios, resolutions, and durations up to 204 frames.
Recent advances in video generation have shifted toward transformer-based diffusion architectures, which excel at generating long, high-resolution videos across diverse aspect ratios. However, existing tools that allow users to control object motion rely largely on older network designs. These older models struggle to generate sequences longer than a few seconds, frequently causing visual distortions, unnatural drifting, and blurring during extended movements. Achieving precise, user-guided motion control within scalable transformer models has therefore become a critical goal for practical video synthesis.
The article demonstrates and evaluates Tora, the first trajectory-oriented framework built on diffusion transformers that concurrently integrates text, visual inputs, and arbitrary movement trajectories. The primary objective is to enable scalable, high-fidelity video generation while maintaining accurate motion control that respects real-world physical dynamics.
To achieve this, the authors designed two core components integrated into an open-source diffusion transformer backbone. The Trajectory Extractor converts user-drawn trajectories into compressed spacetime motion representations that match the video's underlying data format. The Motion-guidance Fuser then incorporates these motion representations into the transformer blocks using adaptive normalization. The system was trained on approximately 630,000 curated, high-quality video clips using an efficient two-stage process that transitioned from dense motion learning to flexible, sparse trajectory following. Crucially, training was restricted primarily to motion and temporal blocks, preserving the base model's core visual knowledge.
The evaluation yielded several key findings. First, on long 128-frame sequences, Tora achieved three to five times higher trajectory tracking accuracy and improved video quality by roughly 30% to 40% compared to leading motion-control methods. Second, Tora demonstrated robust scalability, maintaining stable motion control across varying aspect ratios, resolutions up to 720p, and lengths reaching 204 frames. Third, architectural tests showed that adaptive normalization combined with a three-dimensional autoencoder outperformed alternative fusion and compression methods in both motion fidelity and computational efficiency. Finally, testing on larger transformer architectures confirmed that trajectory accuracy continues to improve as model size and training data scale up.
These findings indicate that trajectory-guided control can be integrated into large-scale video models without sacrificing generative visual quality or incurring prohibitive retraining costs. This provides content creators and developers with a practical mechanism for fine-grained animation control and dynamic camera management, lowering the risk of unnatural motion artifacts in longer video productions.
For future development, the source supports adopting adapter-style training and adaptive normalization as standard baselines for motion control in transformer-based video models. When choosing trajectory integration methods, engineering teams should favor three-dimensional latent compression over simpler pooling or keyframe subsampling to prevent motion degradation. Continued exploration should examine expanding model scale and training dataset size, as the architecture exhibits clear performance gains when scaled.
While the findings demonstrate high confidence and clear advantages on curated benchmarks, users should note that performance depends on the quality of automated captions and the filtering of interfering camera motions during data preparation. Furthermore, error rates still show a slight, gradual rise as video length extends, indicating that extremely long generations may require careful trajectory specification.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). Introduces the Diffusion Transformer (DiT) architecture upon which Tora builds its foundational spatial-temporal video generation backbone.
- Paper: Photorealistic Video Generation with Diffusion Models, Agrim Gupta et al. (2024). Demonstrates training transformer-based diffusion models on compressed latent representations for photorealistic video generation, establishing key spatial-temporal design principles.
- Paper: AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning, Yuwei Guo et al. (2024). Provides foundational techniques for inserting dedicated motion modules into diffusion networks to achieve controllable temporal dynamics.
- Paper: T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models, Chong Mou et al. (2023). Establishes lightweight adapter mechanisms to inject external conditioning signals into pre-trained diffusion backbones, inspiring motion guidance fusion.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). Pioneers the fundamental formulation and joint image-video training paradigms for video diffusion models.
- Paper: MotionClone: Training-Free Motion Cloning for Controllable Video Generation, Pengyang Ling et al. (2025). Explores training-free motion transfer and cloning across video diffusion models, offering an alternative paradigm to explicit trajectory-guided conditioning.
- Paper: Accelerating Diffusion Transformers with Token-wise Feature Caching, Chang Zou et al. (2025). Presents token-wise caching methods to accelerate diffusion transformers during inference, directly addressing the computational overhead inherent to DiT architectures like Tora.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). Applies diffusion transformers learned from real-world video dynamics to unify complex multi-image generation and editing tasks.
