Built independently by an author, for readers. Read the story and support ChapterPal

keyword

motion modeling

Motion modeling is the computational and mathematical process of estimating, representing, and predicting the spatial and temporal movements of objects, surfaces, or camera viewpoints across sequences of visual data. In computer vision and video analysis, it establishes formal descriptions of dynamic behavior, ranging from parametric trajectory and kinematic state estimations to dense pixel flows and spatiotemporal feature representations. By capturing how entities shift, accelerate, or deform over time, motion modeling enables automated systems to maintain object associations through occlusions, analyze human actions, interpolate intermediate video frames, and synthesize coherent dynamic visual scenes.

6 items

Video Frame Interpolation Transformer

Video Frame Interpolation Transformer

Zhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen, Ming-Hsuan Yang

Why you should read this

Proposes a lightweight video frame interpolation transformer that adapts local self-attention into a space-time separable mechanism to capture long-range spatial-temporal dependencies with high memory efficiency.

Existing methods for video interpolation heavily rely on deep convolution neural networks, and thus suffer from their intrinsic limitations, such as content-agnostic kernel weights and restricted receptive field. To address these issues, we propose a Transformer-based video interpolation framework that allows content-aware aggregation weights and considers long-range dependencies with the self-attention operations. To avoid the high computational cost of global self-attention, we introduce the concept of local attention into video interpolation and extend it to the spatial-temporal domain. Furthermore, we propose a space-time separation strategy to save memory usage, which also improves performance. In addition, we develop a multi-scale frame synthesis scheme to fully realize the potential of Transformers. Extensive experiments demonstrate the proposed model performs favorably against the state-of-the-art methods both quantitatively and qualitatively on a variety of benchmark datasets. The code and models are released at https://github.com/zhshi0816/Video-Frame-Interpolation-Transformer.

Added

2026-10-05

How Far Is Video Generation from World Model: A Physical Law Perspective

How Far Is Video Generation from World Model: A Physical Law Perspective

Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, Jiashi Feng

OrganizationsByteDanceTechnion – Israel Institute of TechnologyTsinghua University

Why you should read this

Demonstrates through controlled 2D mechanics simulations that scaling video diffusion models improves in-distribution and combinatorial generalization but fails to discover fundamental physical laws for out-of-distribution extrapolation, relying instead on case-based mimicry prioritized by superficial visual attributes.

Scaling video generation models is believed to be promising in building world models that adhere to fundamental physical laws. However, whether these models can discover physical laws purely from vision can be questioned. A world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios. In this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization. We developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws. We focus on the scaling behavior of training diffusion-based video generation models to predict object movements based on initial frames. Our scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios. Further experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit “case-based” generalization behavior, i.e., mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color > size > velocity > shape. Our study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws.

Added

2026-10-01

DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion

DanceTrack: Multi-Object Tracking in Uniform Appearance and Diverse Motion

Peize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan, Song Bai, Kris Kitani, Ping Luo

OrganizationsByteDanceCarnegie Mellon UniversityUniversity of Hong Kong

Why you should read this

Introduces DanceTrack, a large-scale multi-human tracking benchmark of group dancing scenes featuring uniform appearance and complex motion patterns, exposing the vulnerabilities of conventional appearance-based tracking models and directing research toward motion-centric association.

A typical pipeline for multi-object tracking (MOT) is to use a detector for object localization, and following re-identification (re-ID) for object association. This pipeline is partially motivated by recent progress in both object detection and re-ID, and partially motivated by biases in existing tracking datasets, where most objects tend to have distinguishing appearance and re-ID models are sufficient for establishing associations. In response to such bias, we would like to re-emphasize that methods for multi-object tracking should also work when object appearance is not sufficiently discriminative. To this end, we propose a large-scale dataset for multi-human tracking, where humans have similar appearance, diverse motion and extreme articulation. As the dataset contains mostly group dancing videos, we name it “DanceTrack”. We expect DanceTrack to provide a better platform to develop more MOT algorithms that rely less on visual discrimination and depend more on motion analysis. We benchmark several state-of-the-art trackers on our dataset and observe a significant performance drop on DanceTrack when compared against existing benchmarks. The dataset, project code and competition is released at: https://github.com/DanceTrack.

Added

2026-09-26