Tora: Trajectory-oriented Diffusion Transformer for Video Generation

Zhenghao ZhangJunchao LiaoMenghao LiZuozhuo DaiBingxue QiuSiyu ZhuLong QinWeizhi Wang

article2025CVPR184 citations

Presents Tora, the first trajectory-controlled Diffusion Transformer framework that combines text, visual, and path guidance to generate videos across variable aspect ratios, resolutions, and durations up to 204 frames.

Listen

Recent advances in video generation have shifted toward transformer-based diffusion architectures, which excel at generating long, high-resolution videos across diverse aspect ratios. However, existing tools that allow users to control object motion rely largely on older network designs. These older models struggle to generate sequences longer than a few seconds, frequently causing visual distortions, unnatural drifting, and blurring during extended movements. Achieving precise, user-guided motion control within scalable transformer models has therefore become a critical goal for practical video synthesis.

The article demonstrates and evaluates Tora, the first trajectory-oriented framework built on diffusion transformers that concurrently integrates text, visual inputs, and arbitrary movement trajectories. The primary objective is to enable scalable, high-fidelity video generation while maintaining accurate motion control that respects real-world physical dynamics.

To achieve this, the authors designed two core components integrated into an open-source diffusion transformer backbone. The Trajectory Extractor converts user-drawn trajectories into compressed spacetime motion representations that match the video's underlying data format. The Motion-guidance Fuser then incorporates these motion representations into the transformer blocks using adaptive normalization. The system was trained on approximately 630,000 curated, high-quality video clips using an efficient two-stage process that transitioned from dense motion learning to flexible, sparse trajectory following. Crucially, training was restricted primarily to motion and temporal blocks, preserving the base model's core visual knowledge.

The evaluation yielded several key findings. First, on long 128-frame sequences, Tora achieved three to five times higher trajectory tracking accuracy and improved video quality by roughly 30% to 40% compared to leading motion-control methods. Second, Tora demonstrated robust scalability, maintaining stable motion control across varying aspect ratios, resolutions up to 720p, and lengths reaching 204 frames. Third, architectural tests showed that adaptive normalization combined with a three-dimensional autoencoder outperformed alternative fusion and compression methods in both motion fidelity and computational efficiency. Finally, testing on larger transformer architectures confirmed that trajectory accuracy continues to improve as model size and training data scale up.

These findings indicate that trajectory-guided control can be integrated into large-scale video models without sacrificing generative visual quality or incurring prohibitive retraining costs. This provides content creators and developers with a practical mechanism for fine-grained animation control and dynamic camera management, lowering the risk of unnatural motion artifacts in longer video productions.

For future development, the source supports adopting adapter-style training and adaptive normalization as standard baselines for motion control in transformer-based video models. When choosing trajectory integration methods, engineering teams should favor three-dimensional latent compression over simpler pooling or keyframe subsampling to prevent motion degradation. Continued exploration should examine expanding model scale and training dataset size, as the architecture exhibits clear performance gains when scaled.

While the findings demonstrate high confidence and clear advantages on curated benchmarks, users should note that performance depends on the quality of automated captions and the filtering of interfering camera motions during data preparation. Furthermore, error rates still show a slight, gradual rise as video length extends, indicating that extremely long generations may require careful trajectory specification.

Cover for Tora: Trajectory-oriented Diffusion Transformer for Video Generation

Abstract

Recent advancements in Diffusion Transformer (DiT) have demonstrated remarkable proficiency in producing high-quality video content. Nonetheless, the potential of transformer-based diffusion models for effectively generating videos with controllable motion remains an area of limited exploration. This paper introduces Tora, the first trajectory-oriented DiT framework that concurrently integrates textual, visual, and trajectory conditions, thereby enabling scalable video generation with effective motion guidance. Specifically, Tora consists of a Trajectory Extractor (TE), a Spatial-Temporal DiT, and a Motion-guidance Fuser (MGF). The TE encodes arbitrary trajectories into hierarchical spacetime motion patches with a 3D motion compression network. The MGF integrates the motion patches into the DiT blocks to generate consistent videos that accurately follow designated trajectories. Our design aligns seamlessly with DiT's scalability, allowing precise control of video content's dynamics with diverse durations, aspect ratios, and resolutions. Extensive experiments demonstrate that Tora excels in achieving high motion fidelity compared to the foundational DiT model, while also accurately simulating the complex movements of the physical world. Code is made available at https://github.com/alibaba/Tora .

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Diffusion models for Video Generation
  • 2.2. Motion control in Video Generation
  • 3. Methodology
  • 3.1. Preliminary
  • 3.2. Tora
  • 3.3. Data Processing and Training Strategy
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Results
  • 4.3. Ablation study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Tora trajectory-oriented diffusion-transformer framework

    model/method

    Tora is a trajectory-controlled video-generation framework built on OpenSora’s Spatial-Temporal Diffusion Transformer (ST-DiT). It accepts text, optional input images, and one or more user-specified spatial trajectories, and produces videos with controllable object or camera motion across variable durations, aspect ratios, and resolutions. The framework adds two motion-specific components to the DiT pipeline: a Trajectory Extractor (TE), which converts trajectories into hierarchical spacetime motion patches, and a Motion-guidance Fuser (MGF), which injects these patches into the denoising transformer. Tora is designed to preserve the generative knowledge and scaling behavior of the foundation DiT while adding explicit motion control; the reported system generates videos up to 204 frames at 720p resolution. The architecture diagram on page 4 shows noisy video patches and text entering the ST-DiT, while trajectory inputs pass through the TE and are fused into successive transformer blocks by the MGF.

  2. Knowl 2 — Hierarchical spacetime trajectory extraction

    model/method

    Given a trajectory consisting of spatial positions (xi,yi)(x_i,y_i) at video frames i=0,,,,,L-1, Tora first represents the trajectory as a two-channel displacement map over a video of length LL, height HH, and width WW. The displacement components at a trajectory position are the horizontal and vertical frame-to-frame offsets, u(xi,yi)=xi+1−xiu(x_i,y_i)=x_{i+1}-x_i and v(xi,yi)=yi+1−yiv(x_i,y_i)=y_{i+1}-y_i. The sparse displacement map is Gaussian-filtered to reduce scattered motion signals, with an all-zero map for the first frame, and is then converted into a three-channel RGB flow visualization.

    A 3D variational autoencoder compresses the visualized trajectory by a factor of four temporally and eight spatially, producing a motion latent gm∈Rl×h×w×4g_m\in\mathbb{R}^{l\times h\times w\times4} with l=L/4l=L/4, h=H/8h=H/8, and w=W/8w=W/8. The VAE is trained with reconstruction loss and uses a simplified MAGVIT-v2-style design, with spatial compression initialized from the SDXL VAE. Applying the same patchification as the video latent and a stack of lightweight convolutional layers yields motion patches f∈Rl×s×d′f\in\mathbb{R}^{l\times s\times d'}, where s=hw/p2s=hw/p^2 is the number of spatial patches per latent frame, pp is the patch size, and d′d' is the motion-patch channel dimension. Each convolutional stage is skip-connected to the preceding stage, producing multiple motion-feature levels for injection into different ST-DiT blocks. This construction makes trajectory conditions inhabit the same compressed spacetime organization as the video tokens rather than using only independent frame-to-frame offsets.

  3. Knowl 3 — Adaptive-normalization motion-guidance fusion

    model/method

    The Motion-guidance Fuser injects the hierarchical motion features into the corresponding ST-DiT blocks. For transformer block ii, let hi−1∈Rl×s×dh_{i-1}\in\mathbb{R}^{l\times s\times d} be the incoming video-token representation and let fif_i be the motion feature assigned to that block. Tora’s default fusion mechanism transforms fif_i with two zero-initialized convolutional layers into a scale tensor γi\gamma_i and a shift tensor βi\beta_i, then modulates the video representation with a residual adaptive-normalization update:

    hi=γi⊙hi−1+βi+hi−1,h_i=\gamma_i\odot h_{i-1}+\beta_i+h_{i-1},

    where hih_i is the fused representation, ⊙\odot denotes elementwise multiplication, and γi\gamma_i and βi\beta_i have the same token shape as hi−1h_{i-1}. The zero initialization makes the added motion pathway initially close to an identity mapping, allowing the pretrained DiT behavior to be retained while the motion adapter learns.

    The authors also tested channel concatenation followed by an MLP and an additional cross-attention layer using motion patches as keys and values. Adaptive normalization was selected for the final MGF because it provided the best motion and visual metrics with lower computational cost. The MGF is placed in the temporal DiT blocks, where it has the strongest effect on trajectory consistency.

  4. Knowl 4 — Spatial-temporal DiT representation used by Tora

    model/method

    Tora uses OpenSora’s ST-DiT as its denoising backbone. The backbone alternates spatial DiT blocks and temporal DiT blocks. A spatial block applies spatial self-attention, text cross-attention, and a pointwise feed-forward transformation; a temporal block retains the same structure but replaces spatial self-attention with temporal self-attention. Residual skip connections preserve each block’s input after normalization and transformation, and the text condition is supplied by a T5 encoder.

    For an input video X∈RL×H×W×3X\in\mathbb{R}^{L\times H\times W\times3}, the video autoencoder produces a latent z0∈Rl×h×w×4z_0\in\mathbb{R}^{l\times h\times w\times4}, where l=L/4l=L/4, h=H/8h=H/8, and w=W/8w=W/8. Patchification converts this latent into tokens I∈Rl×s×dI\in\mathbb{R}^{l\times s\times d} with s=hw/p2s=hw/p^2, where pp is the spatial patch size and dd is the video-token dimension. Spatial and temporal attention therefore operate on a variable-length spacetime token sequence, enabling the same denoiser to process different video durations. The TE uses the matching latent and patch structure, which is the architectural basis for integrating trajectory information without the fixed-length limitations of earlier UNet-based controllers.

  5. Knowl 5 — Two-stage motion training and motion-focused data processing

    algorithm

    Tora’s training data pipeline is designed to retain videos with reliable object motion while suppressing samples dominated by camera motion. Raw videos are first divided into short clips using scene detection. Clips with encoding errors, zero duration, low resolution, or poor aesthetic and optical-flow scores are removed. Motion segmentation and camera-motion detection are then used to exclude clips whose motion is primarily caused by the camera. Videos with dramatic object motion that produces unusually large flow deviations are retained with probability 1−flow score/1001-\text{flow score}/100. Captions are generated with PLLaVA; prompts are refined with GPT-4o at inference time to match the training caption style. This process yields approximately 630,000 eligible training videos from Panda-70M’s high-quality 10M subset, Mixkit, Pexels, and internally annotated videos.

    Motion learning uses two stages. First, dense optical flow from each training video is used as the trajectory condition for two epochs, providing rich motion coverage. Second, the model is fine-tuned for one epoch using sparse, user-like trajectories: between one and NN object trajectories are randomly selected from motion-segmentation results and flow scores, with N=16N=16 as the maximum. A Gaussian filter refines the sparse trajectories before TE encoding. Only the temporal transformer blocks, TE, and MGF are trained in the adapter-style setup; the pretrained generative knowledge of the foundation DiT is otherwise retained. For image conditioning, randomly selected frames are unmasked during training and their corresponding video patches are kept noise-free.

  6. Knowl 6 — Motion-control comparison across video durations

    data/table

    The quantitative comparison reported on page 7 evaluates motion-guided video generation at 16, 64, and 128 frames. FVD measures video-distribution quality (lower is better), CLIPSIM measures text-video similarity (higher is better), and TrajError is the mean L1 distance between generated and prescribed trajectories (lower is better). Tora uses OpenSora as its DiT foundation and is compared with UNet-based controllers, an OpenSora baseline, and an OpenSora implementation of DragNUWA.

    Could not parse LaTeX table

    Tora’s 128-frame TrajError of 11.72 is lower than the best UNet-based value, 38.39, and its FVD of 494 is lower than the best UNet-based value, 731. Thus, the gap in both motion fidelity and visual quality grows with sequence length. Tora also improves on the OpenSora-based DragNUWA adaptation, whose motion features are less compatible with the DiT latent space and whose 128-frame FVD is 565. The page-8 qualitative comparison complements these numbers: Tora preserves realistic pedaling in a bicycle scene and avoids the lantern deformation or missing-object behavior shown by competing controllers.

  7. Knowl 7 — 3D motion VAE is the most effective trajectory-compression method

    empirical result

    Tora’s trajectory-compression ablation compares three ways of converting trajectories into DiT-compatible spacetime conditions. Sampling Frame selects a middle frame from each successive four-frame interval and uses Patch-Unshuffle for spatial compression. Average Pooling aggregates successive frames, while the proposed 3D VAE learns a temporally and spatially compressed motion representation with global context.

    Could not parse LaTeX table

    The 3D VAE is best on all three metrics. The paper attributes the weaker performance of frame sampling to flow-estimation errors during rapid motion or occlusion and to increased dissimilarity between adjacent compressed patches. Average pooling captures coarse movement but averages away trajectory direction and magnitude. The learned 3D VAE retains motion information across the successive frames that correspond to each compressed latent interval.

  8. Knowl 8 — Adaptive normalization is the best motion-fusion block

    empirical result

    The MGF ablation compares the three fusion mechanisms implemented in Tora: concatenating motion and video features along the channel dimension, adding motion-conditioned cross-attention, and using adaptive normalization. All variants are evaluated with the same motion-controlled video-generation setup.

    Could not parse LaTeX table

    Adaptive normalization achieves the lowest FVD and TrajError, the highest CLIPSIM, and the best computational efficiency. The authors explain that it dynamically modulates features without requiring strict token alignment, which can be difficult for motion cross-attention. Channel concatenation can congest the representation and weaken the motion signal. Identity initialization of the normalization pathway is important for performance, and placing MGF in the temporal rather than spatial transformer blocks reduces TrajError from 23.39 to 14.25.

  9. Knowl 9 — Hybrid dense-to-sparse trajectory training is necessary for flexible control

    empirical result

    The training-trajectory ablation evaluates dense optical flow alone, sparse object trajectories alone, and Tora’s hybrid two-stage schedule. Dense flow supplies detailed and abundant motion supervision but does not adequately represent the sparse trajectories users provide at inference. Sparse trajectories are user-friendly but provide too little information for the model to learn motion structure efficiently when used from the beginning.

    Could not parse LaTeX table

    The hybrid schedule first trains with dense optical flow and then fine-tunes with randomly selected sparse object trajectories. It substantially improves both visual quality and trajectory adherence over either supervision type used alone, supporting Tora’s ability to follow arbitrary sparse trajectories at inference.

  10. Knowl 10 — Motion control improves with DiT scale and training volume

    empirical result

    To test whether the motion modules preserve DiT scalability, the authors transfer Tora’s TE and MGF to CogVideoX backbones. The resulting Tora-CogVideoX2B has 2.5 billion parameters and Tora-CogVideoX5B has 6.3 billion parameters; both retain the base video VAE for motion compression, with MGF inserted before the Expert Transformer’s full-attention module.

    The scaling plot on page 8 shows TrajError decreasing as training volume increases for both models. The plotted values are 23.82, 15.65, and 13.28 for Tora-CogVideoX2B, and 15.37, 11.47, and 10.52 for Tora-CogVideoX5B, at successively larger training volumes. The larger model has lower error at each shown scale, and both models improve with additional data. These results support the paper’s claim that the trajectory representation and fusion design are compatible with the scaling behavior of diffusion transformers rather than being tied to one backbone size.

Coverage note — No substantial contributed material was omitted; implementation settings and the image-conditioning mask strategy were incorporated into the method and training knowls, while qualitative examples were summarized with the quantitative comparison.

References

  1. 1.Pexels. https://www.pexels.com/. 6
  2. 2.Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models. arXiv preprint arXiv:2405.04233, 2024. 3
  3. 3.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 1, 3
  4. 4.Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. https://openai.com/research/video-generation - models - as - world - simulators, 2024. 1, 3
  5. 5.Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual KITTI 2. arXiv preprint arXiv:2001.10773, 2020. 6
  6. 6.Haoxin Chen, Menghan Xia, Yin-Yin He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, Chao-Liang Weng, and Ying Shan. VideoCrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023. 3
  7. 7.Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, and Sergey Tulyakov. Panda-70M: Captioning 70M videos with multiple cross-modality teachers. In CVPR, pages 13320–13331. IEEE, 2024. 6
  8. 8.Zuozhuo Dai, Zhenghao Zhang, Yao Yao, Bingxue Qiu, Siyu Zhu, Long Qin, and Weizhi Wang. Fine-grained open domain image animation with motion guidance. arXiv preprint arXiv:2311.12886, 2023. 3, 7
  9. 9.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat GANs on image synthesis. In NeurIPS, pages 8780–8794, 2021. 1
  10. 10.Elements Envato. Mixkit: Free assets for your next video project. https://mixkit.co, 2024. 6
  11. 11.Daniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole, Serena Zhang, Tobias Pfaff, Tatiana Lopez-Guevara, Carl Doersch, Yusuf Aytar, Michael Rubinstein, Chen Sun, Oliver Wang, Andrew Owens, and Deqing Sun. Motion prompting: Controlling video generation with motion trajectories. arXiv preprint arXiv:2412.02700, 2024. 3
  12. 12.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 3
  13. 13.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In NeurIPS, 2022. 1, 3
  14. 14.Hyemi Jang, Junsung Park, Dahuin Jung, Jaihyun Lew, Ho Bae, and Sungroh Yoon. PUCA: Patch-unshuffle and channel attention for enhanced self-supervised image denoising. In NeurIPS, 2023. 7
  15. 15.Hyeonho Jeong, Geon Yeong Park, and Jong Chul Ye. VMC: Video motion customization using temporal attention adaption for text-to-video diffusion models. In CVPR, pages 9212–9221. IEEE, 2024. 3
  16. 16.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2Video-Zero: Text-to-image diffusion models are zero-shot video generators. In ICCV, pages 15908–15918. IEEE, 2023. 3
  17. 17.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 6
  18. 18.Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014. 3
  19. 19.Wan-Duo Kurt Ma, John P. Lewis, and W. Bastiaan Kleijn. Trailblazer: Trajectory control for diffusion-based video generation. In SIGGRAPH Asia, pages 97:1–97:11. ACM, 2024. 3, 7
  20. 20.Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In CVPR, pages 4040–4048. IEEE Computer Society, 2016. 6
  21. 21.Lukas Mehl, Jenny Schmalfuss, Azin Jahedi, Yaroslava Nalivayko, and Andres Bruhn. Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo. In CVPR, pages 4981–4991. IEEE, 2023. 6
  22. 22.Chong Mou, Mingdeng Cao, Xintao Wang, Zhaoyang Zhang, Ying Shan, and Jian Zhang. ReVideo: Remake a video with motion and content control. In NeurIPS, 2024. 3
  23. 23.Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In AAAI, pages 4296–4304. AAAI Press, 2024. 3
  24. 24.OpenAI. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 6
  25. 25.William Peebles and Saining Xie. Scalable diffusion models with transformers. In ICCV, pages 4195–4205, 2023. 1, 4
  26. 26.Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville. Film: Visual reasoning with a general conditioning layer. In AAAI, pages 3942–3951, 2018. 3
  27. 27.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. SDXL: Improving latent diffusion models for high-resolution image synthesis. In ICLR. OpenReview.net, 2024. 5
  28. 28.Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Jagadeesh, Kunpeng Li, Luxin Zhang, Mannat Singh, Mary Williamson, Matt Le, Matthew Yu, Mitesh Kumar Singh, Peizhao Zhang, Peter Vajda, Quentin Duval, Rohit Girdhar, Roshan Sumbaly, Sai Saketh Rambhatla, Sam S. Tsai, Samaneh Azadi, Samyak Datta, Sanyuan Chen, Sean Bell, Sharadh Ramaswamy, Shelly Sheynin, Siddharth Bhattacharya, Simran Motwani, Tao Xu, Tianhe Li, Tingbo Hou, Wei-Ning Hsu, Xi Yin, Xiaoliang Dai, Yaniv Taigman, Yaqiao Luo, Yen-Cheng Liu, Yi-Chiao Wu, Yue Zhao, Yuval Kirstain, Zecheng He, Zijian He, Albert Pumarola, Ali K. Thabet, Artsiom Sanakoyeu, Arun Mallya, Baishan Guo, Boris Araya, Breena Kerr, Carleigh Wood, Ce Liu, Cen Peng, Dmitry Vengertsev, Edgar Schonfeld, Elliot Blanchard, Felix Juefei-Xu, Fraylie Nord, Jeff Liang, John Hoffman, Jonas Kohler, Kaolin Fire, Karthik Sivakumar, Lawrence Chen, Licheng Yu, Luya Gao, Markos Georgopoulos, Rashel Moritz, Sara K. Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du. Movie Gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720, 2024. 3
  29. 29.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022. 1, 4
  30. 30.Anurag Ranjan, David T. Hoffmann, Dimitrios Tzionas, Siyu Tang, Javier Romero, and Michael J. Black. Learning multi-human optical flow. IJCV, 128(4):873–890, 2020. 6
  31. 31.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention, pages 234–241. Springer, 2015. 1
  32. 32.Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Motion-I2V: Consistent and controllable image-to-video generation with explicit motion modeling. In SIGGRAPH, page 111. ACM, 2024. 3
  33. 33.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022. 3
  34. 34.Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 6
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017. 4
  36. 36.Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. VideoComposer: Compositional video synthesis with motion controllability. In NeurIPS, 2023. 2, 3, 7
  37. 37.Xiang Wang, Shiwei Zhang, Hangjie Yuan, Zhiwu Qing, Biao Gong, Yingya Zhang, Yujun Shen, Changxin Gao, and Nong Sang. A recipe for scaling up text-to-video generation with text-free videos. In CVPR, 2024. 3
  38. 38.Zhouxia Wang, Ziyang Yuan, Xintao Wang, Tianshui Chen, Menghan Xia, Ping Luo, and Yin Shan. MotionCtrl: A unified and flexible motion controller for video generation. In SIGGRAPH, 2024. 2, 3, 7
  39. 39.Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. GODIVA: Generating open-domain videos from natural descriptions. arXiv preprint arXiv:2104.14806, 2021. 6
  40. 40.Weijia Wu, Zhuang Li, Yuchao Gu, Rui Zhao, Yefei He, David Junhao Zhang, Mike Zheng Shou, Yan Li, Tingting Gao, and Di Zhang. DragAnything: Motion control for anything using entity representation. In ECCV, pages 331–348. Springer, 2024. 3
  41. 41.Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying flow, stereo and depth estimation. IEEE TPAMI, 45(11):13941–13958, 2023. 3, 6
  42. 42.Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See-Kiong Ng, and Jiashi Feng. PLLaVA : Parameter-free LLaVA extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994, 2024. 6
  43. 43.Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. CogVideoX: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 3
  44. 44.Shengming Yin, Chenfei Wu, Jian Liang, Jie Shi, Houqiang Li, Gong Ming, and Nan Duan. DragNUWA: Fine-grained control in video generation by integrating text, image, and trajectory. arXiv preprint arXiv:2308.08089, 2023. 2, 3, 7
  45. 45.Lijun Yu, Yong Cheng, Kihyuk Sohn, Jose Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, et al. MAGVIT: Masked generative video transformer. In CVPR, 2023. 5
  46. 46.Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, and Lu Jiang. Language model beats diffusion - tokenizer is key to visual generation. In ICLR. OpenReview.net, 2024. 3
  47. 47.Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou. I2VGen-XL: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:2311.04145, 2023. 1
  48. 48.Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. ControlVideo: Training-free controllable text-to-video generation. In ICLR. OpenReview.net, 2024. 3
  49. 49.Rui Zhao, Yuchao Gu, Jay Zhangjie Wu, David Junhao Zhang, Jia-Wei Liu, Weijia Wu, Jussi Keppo, and Mike Zheng Shou. MotionDirector: Motion customization of text-to-video diffusion models. In ECCV, pages 273–290. Springer, 2024. 3
  50. 50.Wang Zhao, Shaohui Liu, Hengkai Guo, Wenping Wang, and Yong-Jin Liu. ParticleSfM: Exploiting dense point trajectories for localizing moving cameras in the wild. In ECCV, pages 523–542, 2022. 3, 6
  51. 51.Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-Sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404, 2024. 3, 7

Citation

MLA
Zhang, Z., et al. “Tora: Trajectory-oriented Diffusion Transformer for Video Generation”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 2063–73, https://doi.org/10.1109/CVPR52734.2025.00198.
APA
Zhang, Z., Liao, J., Li, M., Dai, Z., Qiu, B., Zhu, S., Qin, L., & Wang, W. (2025). Tora: Trajectory-oriented Diffusion Transformer for Video Generation. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2063–2073. https://doi.org/10.1109/CVPR52734.2025.00198
Chicago
Zhang, Z., J. Liao, M. Li, et al. 2025. “Tora: Trajectory-oriented Diffusion Transformer for Video Generation”. 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2063–73. https://doi.org/10.1109/CVPR52734.2025.00198.
Harvard
Zhang, Z. et al. (2025) “Tora: Trajectory-oriented Diffusion Transformer for Video Generation”, 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 2063–2073. Available at: https://doi.org/10.1109/CVPR52734.2025.00198.
Vancouver
1. Zhang Z, Liao J, Li M, Dai Z, Qiu B, Zhu S, Qin L, Wang W (2025) Tora: Trajectory-oriented Diffusion Transformer for Video Generation. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2063–2073

BibTeX

@inproceedings{Zhang_2025, title={Tora: Trajectory-oriented Diffusion Transformer for Video Generation}, url={http://dx.doi.org/10.1109/CVPR52734.2025.00198}, DOI={10.1109/cvpr52734.2025.00198}, booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhang, Zhenghao and Liao, Junchao and Li, Menghao and Dai, ZuoZhuo and Qiu, Bingxue and Zhu, Siyu and Qin, Long and Wang, Weizhi}, year={2025}, month=June, pages={2063–2073} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE