How Far Is Video Generation from World Model: A Physical Law Perspective

Bingyi KangYang YueRui LuZhijie LinYang ZhaoKaixin WangGao HuangJiashi Feng

article2025ICML238 citations

Demonstrates through controlled 2D mechanics simulations that scaling video diffusion models improves in-distribution and combinatorial generalization but fails to discover fundamental physical laws for out-of-distribution extrapolation, relying instead on case-based mimicry prioritized by superficial visual attributes.

Listen

Recent advances in generative video modeling have raised expectations that scaling up data and computing power could yield effective "world models" capable of autonomously discovering physical principles from visual input alone. Such capabilities are considered vital for high-stakes domains such as robotics and autonomous vehicle simulation. The article evaluates whether scaling diffusion-based video generation architectures actually allows systems to extract fundamental physical laws or merely leads to superficial pattern memorization.

To test this, the article establishes a controlled evaluation framework using 2D deterministic physics engines (Box2D and PHYRE) to model classical mechanics scenarios such as uniform linear motion, elastic collisions, parabolic flight, and multi-object interactions. The evaluation isolates visual appearance by using simple geometric shapes and assesses models across three operational regimes: in-distribution scenarios (familiar parameters), out-of-distribution scenarios (novel velocities or masses outside training bounds), and combinatorial generalization (unseen combinations of previously observed objects and interactions). Model sizes ranged up to 456 million parameters and dataset sizes reached up to 6 million video examples.

The findings demonstrate a fundamental divide in model capabilities. For in-distribution tasks, models achieve near-perfect performance, reducing velocity errors to baseline system noise levels. However, in out-of-distribution tasks, models fail entirely: extrapolation error is roughly an order of magnitude higher than in-distribution error, and scaling data volume from 30,000 to 3 million samples or enlarging model capacity yields no measurable improvement. In contrast, combinatorial generalization scales effectively; expanding template variety from 6 to 60 combinations reduces human-evaluated physical abnormality rates from 67% down to 10%. Diagnostic tests reveal that models operate through "case-based" imitation—retrieving and adapting the nearest training sample rather than deducing underlying rules. When resolving conflicting cues during generation, models follow a strict visual priority hierarchy: color is prioritized over size, size over velocity, and velocity over shape, which explains frequent real-world generation flaws such as shape distortion and object inconsistency.

These results show that scaling training volume and parameters alone cannot produce true physical reasoning or reliable world models. Deploying vision-only video generation models into safety-critical applications—such as edge-case simulation for autonomous driving or robotic task planning—introduces significant operational risk, as these systems cannot extrapolate beyond their observed training envelopes and can hallucinate physically impossible behaviors when encountering unfamiliar inputs.

Organizations developing or applying physical world models should pivot investment strategies away from brute-force data volume scaling toward increasing combinatorial diversity across scenarios. Furthermore, developers should exercise caution regarding visual ambiguity and avoid assuming multimodal inputs (like text or numeric conditioning) naturally solve extrapolation; in the article's experiments, adding language annotations increased out-of-distribution error due to overfitting. Future research must explore architectures with stronger physical inductive biases or explicit neuro-symbolic reasoning rather than relying purely on standard visual diffusion models.

  • Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It introduces the video-diffusion approach that underlies the source’s models, making its prediction setup and scaling experiments easier to follow.

No sufficiently relevant recommendations were found.

Cover for How Far Is Video Generation from World Model: A Physical Law Perspective

Abstract

Scaling video generation models is believed to be promising in building world models that adhere to fundamental physical laws. However, whether these models can discover physical laws purely from vision can be questioned. A world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios. In this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization. We developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws. We focus on the scaling behavior of training diffusion-based video generation models to predict object movements based on initial frames. Our scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios. Further experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit “case-based” generalization behavior, i.e., mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color > size > velocity > shape. Our study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws.

Table of Contents

  • 1. Introduction
  • 2. Video Generation for Physical Law Discovery
  • 2.1. Problem Definition
  • 2.2. Video Generation Model
  • 2.3. On the Verification of Learned Laws
  • 3. In-Distribution and Out-of-Distribution Generalization
  • 3.1. Fundamental Physical Scenarios
  • 3.2. Perfect ID and Failed OOD Generalization
  • 4. Combinatorial Generalization
  • 4.1. Combinatorial Physical Scenarios
  • 4.2. Scaling Law Observed for Combinatorial Generalization
  • 5. Deeper Analysis
  • 5.1. Understanding Generalization from Interpolation and Extrapolation
  • 5.2. Memorization or Generalization
  • 5.3. How Does Diffusion Model Retrieve Data?
  • 5.4. Complex Combinatorial Generalization
  • 5.5. Is Video Sufficient for Complete Physics Modeling?
  • 6. Related Works
  • 7. Conclusion and Discussion
  • Impact Statement
  • References
  • A. More related works
  • B. Latent Video Diffusion Model
  • B.1. Diffusion preliminaries
  • B.2. VAE Architecture and Pretrain
  • B.3. VAE Reconstruction
  • B.4. DiT Implementation Details
  • C. Detailed Experimental setup
  • C.1. Fundamental Physical Scenarios Data
  • C.2. Combinatorial Experiments Evaluation Setup
  • D. Experiments on SOTA Video Generation Models
  • D.1. Is the Prioritization Relevant to VAE?
  • D.2. Can pretrained models learn physical laws?
  • E. More Experiments and Discussions
  • E.1. Can Language and Numerics Aid in Learning Physical Laws?
  • E.2. Continuous Experiment for Pairwise Comparison
  • E.3. Principle Behind Data Retrieval in the Diffusion Model
  • E.4. Failure Cases in Combinatorial Generalization
  • F. Comparison with ID/OOD Generalization Works
  • G. Visualization

Knowls

  1. Knowl 1 — Discrepancy Between In-Distribution and Out-of-Distribution Generalization Under Scaling

    empirical result

    When diffusion-based video generation models are trained to predict 2D kinematic motion from initial conditioning frames, scaling dataset size (from 30K30\text{K} to 3M3\text{M} examples) and model parameters (from DiT-S with 22.5M22.5\text{M} parameters to DiT-XL with 456M456\text{M} parameters) produces opposite outcomes for in-distribution (ID) versus out-of-distribution (OOD) settings:

    • In-Distribution (ID): Across uniform linear motion, two-ball elastic collisions, and parabolic motion under gravity, scaling consistently reduces velocity error. For example, in uniform linear motion, velocity error drops from 0.0220.022 (DiT-S on 30K30\text{K} data) to 0.0120.012 (DiT-L on 3M3\text{M} data), closely approaching the ground-truth video parsing lower bound (0.0100.010).

    • Out-of-Distribution (OOD): When evaluated on initial velocities or object radii outside the training ranges (such as velocities v∈[0,0.8]∪[4.5,6.0]v \in [0, 0.8] \cup [4.5, 6.0] and radii r∈[0.3,0.6]∪[1.5,2.0]r \in [0.3, 0.6] \cup [1.5, 2.0] versus training ranges v∈[1,4]v \in [1, 4] and r∈[0.7,1.5]r \in [0.7, 1.5]), prediction errors are an order of magnitude higher (e.g., 0.4270.427 for DiT-L on 3M3\text{M} data in uniform motion). Increasing dataset size or model parameters yields no systematic improvement, showing irregular fluctuations and failing to generalize physical laws.

  2. Knowl 2 — Attribute Prioritization Hierarchy in Video Diffusion Case Matching

    empirical result

    When video diffusion models encounter inputs with unseen combinations of attributes during uniform linear motion tasks, they resolve conflicting constraints by modifying low-priority attributes to match nearest training examples. Systematic pairwise comparisons across four attributes—color, size, velocity, and shape—reveal a strict prioritization hierarchy:

    Color>Size>Velocity>Shape\text{Color} > \text{Size} > \text{Velocity} > \text{Shape}

    Key empirical behaviors demonstrating this hierarchy include:

    • Color vs. Shape: When trained on red balls and blue squares and conditioned on a blue ball or red square, generated videos alter object shape (e.g., a blue ball transforms into a blue square immediately after conditioning frames) with zero exceptions across 1,400 test cases, preserving color.
    • Velocity vs. Shape: Low-speed balls and high-speed squares in training cause high-speed balls at test time to immediately morph into squares to match the speed-associated shape.
    • Color vs. Size: Red small balls and blue large balls in training cause small blue test balls to rapidly expand into large blue balls, keeping color intact while mutating size.
    • Size vs. Velocity: Small high-speed balls and large low-speed balls result in size dominating over velocity adjustments at extreme boundaries.

    This hierarchy corresponds to the magnitude of pixel/latent-space difference required for each modification: color changes alter pixel values across the entire object surface, size changes alter pixel counts, velocity alters spatial displacement over time, and shape differences involve localized boundary adjustments.

  3. Knowl 3 — Problem Formulation for Physical Law Discovery in Video Generation

    model/method

    Physical law discovery in generative video modeling is formalized by considering physical states governed by latent variables z=(z1,z2,…,zk)∈Z⊆Rkz = (z_1, z_2, \dots, z_k) \in \mathcal{Z} \subseteq \mathbb{R}^k (representing parameters such as position, velocity, and mass) evolving via differential dynamics z˙=F(z)\dot{z} = F(z), or discretely with frame interval δ\delta as:

    zt+1≈zt+δF(zt),t=1,…,L−1z_{t+1} \approx z_t + \delta F(z_t), \quad t = 1, \dots, L-1

    A rendering function R:Z→R3×H×WR: \mathcal{Z} \to \mathbb{R}^{3 \times H \times W} maps states to RGB video frames It=R(zt)I_t = R(z_t), producing a sequence V={I1,I2,…,IL}V = \{I_1, I_2, \dots, I_L\}.

    A parameterized video diffusion model pθp_\theta is conditioned on the first cc frames (typically c∈{1,3}c \in \{1, 3\}) to predict subsequent frames by sampling from pθ(Ic+1′,…,IL′∣I1,…,Ic)p_\theta(I'_{c+1}, \dots, I'_L \mid I_1, \dots, I_c). The model is trained using a velocity-prediction objective:

    Ldiff=EV∼p(x),t∼U(0,1),ϵ∼N(0,I)[∥y−pθ(Vt,c,t)∥2]\mathcal{L}_{\text{diff}} = \mathbb{E}_{V \sim p(x), t \sim \mathcal{U}(0, 1), \epsilon \sim \mathcal{N}(0, I)} \left[ \| y - p_\theta(V_t, c, t) \|^2 \right]

    where Vt=γtV+1−γtϵV_t = \sqrt{\gamma_t} V + \sqrt{1 - \gamma_t} \epsilon, γt\gamma_t is a monotonically decreasing noise schedule with γ1=1\gamma_1 = 1, and the training target is the velocity y=1−γtϵ−γtVy = \sqrt{1 - \gamma_t}\epsilon - \sqrt{\gamma_t}V. The input consists of concatenated noisy latent tokens, zero-padded conditioning frames, and a binary mask indicating which frames serve as conditioning.

  4. Knowl 4 — Combinatorial Generalization and Template Scaling in Multi-Object Physical Simulations

    data/table

    In multi-object 2D physical simulations (using the PHYRE benchmark with 8 object types forming (84)=70\binom{8}{4} = 70 unique four-object templates), combinatorial generalization was evaluated by training Diffusion Transformer models on subsets of templates (6, 30, and 60 templates) and testing on both held-in templates and 10 held-out combination templates. Evaluation metrics include Fréchet Video Distance (FVD), Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), Learned Perceptual Image Patch Similarity (LPIPS), and human-evaluated physically abnormal video rates.

    Model #Templates FVD (↓\downarrow) SSIM (↑\uparrow) PSNR (↑\uparrow) LPIPS (↓\downarrow) Abnormal (↓\downarrow)
    DiT-XL 6 18.2 / 22.1 0.973 / 0.943 32.8 / 25.5 0.028 / 0.082 3% / 67%
    DiT-XL 30 19.5 / 19.7 0.973 / 0.950 32.7 / 27.1 0.028 / 0.065 3% / 18%
    DiT-XL 60 17.6 / 18.7 0.972 / 0.951 32.4 / 27.3 0.030 / 0.062 2% / 10%
    DiT-B 60 18.4 / 21.4 0.967 / 0.949 30.9 / 27.0 0.035 / 0.066 3% / 24%

    The results are presented as {in-template result} / {out-of-template result}. Scaling combination diversity from 6 to 60 templates on DiT-XL substantially reduces the out-of-template abnormal rate from 67% to 10% and improves out-of-template visual metrics. Scaling model size from DiT-B (89.5M parameters) to DiT-XL (456M parameters) on 60 templates reduces the abnormal rate from 24% to 10%, indicating that combinatorial generalization depends critically on both combination coverage in training data and model capacity.

  5. Knowl 5 — Case-Based Imitation and Directional Reversal Artifacts in Video Diffusion

    empirical result

    Rather than learning abstract physical invariants (such as the law of inertia), video diffusion models generalize by retrieving and mimicking nearest-neighbor training samples when presented with out-of-distribution inputs.

    In uniform linear motion experiments where models are evaluated on unseen low velocities v∈[1.0,2.5]v \in [1.0, 2.5] given initial frames:

    • When trained exclusively on left-to-right trajectories (v∈[2.5,4.0]v \in [2.5, 4.0]), the model generates forward motion biased toward the high-speed training range.
    • When trained on the same data augmented with horizontal flipping (introducing right-to-left motions with negative velocities), the model frequently generates videos where a ball initially moving to the right abruptly reverses direction and moves backward after the conditioning frames.

    This occurs because the model identifies reversed videos in the training set as the closest match to the low-speed conditioning input, demonstrating reliance on memorized case retrieval rather than learned physical principles.

  6. Knowl 6 — Convex Hull Boundaries of Generalization in Kinematic Latent Spaces

    empirical result

    The generalization capability of video diffusion models in physical prediction is bounded by the convex hull of the training distribution in latent parameter space:

    • Continuous Parameter Gaps (Uniform Motion): When a middle velocity range is omitted from training (e.g., training on [1.0,1.25][1.0, 1.25] and [3.75,4.0][3.75, 4.0] with [1.25,3.75][1.25, 3.75] absent), testing on the missing range causes predicted velocities to drift toward the nearest training clusters, violating inertia. As the gap narrows or intermediate anchor values are introduced, the model transitions from failed extrapolation to continuous interpolation.
    • Multi-Variable Dynamics (Elastic Collision): In two-ball collisions parameterized by initial velocities (v1,v2)(v_1, v_2), when rectangular regions inside the velocity space are omitted from training, the model successfully interpolates post-collision velocities for test combinations situated inside the convex hull of the training set. However, prediction errors rise sharply for test combinations located outside the convex hull.
  7. Knowl 7 — Kinematic Simulation Benchmark and Velocity Parsing Evaluation Protocol

    experimental setup

    To quantitatively evaluate physical law discovery without visual confounders, a 2D Box2D kinematic simulation benchmark generates 32-frame videos (128×128128 \times 128 resolution, 0.1 s timestep) across three deterministic classical mechanics scenarios:

    1. Uniform Linear Motion: A single ball moving at constant velocity (v∈[1,4]v \in [1, 4], radius r∈[0.7,1.5]r \in [0.7, 1.5] for in-distribution; v∈[0,0.8]∪[4.5,6.0]v \in [0, 0.8] \cup [4.5, 6.0], r∈[0.3,0.6]∪[1.5,2.0]r \in [0.3, 0.6] \cup [1.5, 2.0] for out-of-distribution).
    2. Elastic Collision: Two balls colliding horizontally with conservation of momentum and energy (44 degrees of freedom: initial velocities v1,v2v_1, v_2 and radii r1,r2r_1, r_2).
    3. Parabolic Motion: A ball with initial horizontal velocity subject to constant vertical gravitational acceleration.

    Kinematic accuracy is evaluated by tracking ball centroids xtix_t^i via color-thresholded pixel averaging over valid frames TT (frames where balls are entirely within view, and post-collision frames for collision scenarios), computing differentiated velocities vtiv_t^i, and calculating mean absolute error against simulator ground-truth velocities v^ti\hat{v}_t^i:

    e=1N∣T∣∑i=1N∑t∈T∣vti−v^ti∣e = \frac{1}{N |T|} \sum_{i=1}^N \sum_{t \in T} \left| v_t^i - \hat{v}_t^i \right|

    where NN is the number of balls and ∣T∣|T| is the number of evaluated frames.

  8. Knowl 8 — Effect of Numeric and Natural Language Conditioning on Physical Extrapolation

    empirical result

    Augmenting video diffusion models (DiT-B) with multimodal physical state information does not resolve out-of-distribution (OOD) extrapolation failures in elastic collision modeling:

    • In-Distribution (ID): Adding state vectors (mapped to layer-wise feature embeddings added to video tokens) or T5-encoded natural language descriptions (integrated via cross-attention) achieves velocity errors comparable to vision-only conditioning (approximately 0.0120.012 to 0.0150.015). Visual frames already provide sufficient information for ID prediction.
    • Out-of-Distribution (OOD): Numeric conditioning slightly increases velocity error over vision-only conditioning, while natural language conditioning causes a substantial increase in OOD error across dataset scales (30K30\text{K}, 300K300\text{K}, 3M3\text{M}).

    The heightened failure under text conditioning is attributed to the discrete nature and higher token variability of language embeddings, which induce stronger overfitting to training distribution combinations.

  9. Knowl 9 — Physical Law Generalization in Pretrained Video Foundation Models

    empirical result

    To assess whether large-scale pretraining on internet videos alleviates out-of-distribution physical reasoning failures, Stable Video Diffusion (SVD) was fine-tuned for 300K steps on 2D uniform linear motion and compared against a DiT-B model trained from scratch.

    Model ID Error OOD Error
    Ground Truth Video Parsing Error 0.0099 0.0104
    DiT-B (trained from scratch) 0.0138 0.3583
    SVD-VAE Reconstruction 0.0103 0.0107
    SVD-Finetuned 0.0505 0.9081

    While SVD's pretrained VAE accurately preserves kinematics with reconstruction error (0.01030.0103 ID, 0.01070.0107 OOD) close to ground-truth video parsing, the fine-tuned SVD diffusion model exhibits an OOD velocity error of 0.90810.9081—an order of magnitude higher than its ID error (0.05050.0505) and higher than DiT-B trained from scratch. Pretraining on diverse natural video distributions does not instill universal physical law extrapolation.

  10. Knowl 10 — Visual Ambiguity as an Intrinsic Limitation in Video-Based Physics Modeling

    limitation

    Video generation models that operate purely on visual frame representations encounter fundamental limits in fine-grained physical prediction due to visual ambiguity. When geometric parameters (such as the clearance between a falling ball and a narrow gap, or the exact horizontal alignment relative to a corner) differ at the single-pixel or sub-pixel level, the visual conditioning input contains insufficient precision to uniquely determine the true physical trajectory. Consequently, the model generates visually plausible yet physically incorrect outcomes (such as a ball passing through a gap it should collide with, or vice versa), demonstrating that raw pixel observation alone is inadequate for deterministic physical simulation in sensitive dynamical regimes.

Coverage note — Omitted qualitative visual walkthroughs of specific PHYRE collision failure cases (e.g., bars splitting, sticks floating) and minor DiT implementation hyperparameter details (e.g., patch size, AdamW weight decay values), as the core scientific findings and metrics are fully captured in the knowls.

References

  1. 1.1x world model. 2024. URL https://www.1x.tech/discover/1x-world-model. 9
  2. 2.Allen, K. R., Smith, K. A., and Tenenbaum, J. B. Rapid trial-and-error learning with simulation supports flexible tool use and physical reasoning. Proceedings of the National Academy of Sciences, 117(47):29302–29310, 2020. 15
  3. 3.Ates, T., Atesoglu, M. S., Yigit, C., Kesen, I., Kobas, M., Erdem, E., Erdem, A., Goksun, T., and Yuret, D. Craft: A benchmark for causal reasoning about forces and interactions. arXiv preprint arXiv:2012.04293, 2020. 15
  4. 4.Bakhtin, A., van der Maaten, L., Johnson, J., Gustafson, L., and Girshick, R. Phyre: A new benchmark for physical reasoning. Advances in Neural Information Processing Systems, 32, 2019. 5, 15
  5. 5.Balestriero, R., Pesenti, J., and LeCun, Y. Learning in high dimension always amounts to extrapolation. arXiv preprint arXiv:2110.09485, 2021. 6
  6. 6.Bansal, H., Lin, Z., Xie, T., Zong, Z., Yarom, M., Bitton, Y., Jiang, C., Sun, Y., Chang, K.-W., and Grover, A. Videophy: Evaluating physical commonsense for video generation. arXiv preprint arXiv:2406.03520, 2024. 9
  7. 7.Beaumont, R. and Schuhmann, C. Laion-aesthetics v1. https://github.com/LAION-AI/laion-datasets/blob/main/laion-aesthetic.md, 2022. 16
  8. 8.Bell, S., Upchurch, P., Snavely, N., and Bala, K. Material recognition in the wild with the materials in context database. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3479–3487, 2015. 15
  9. 9.Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S. Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639, 2023. 9
  10. 10.Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 17
  11. 11.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021. 1
  12. 12.Bouman, K. L., Xiao, B., Battaglia, P., and Freeman, W. T. Estimating the material properties of fabric from video. In Proceedings of the IEEE international conference on computer vision, pp. 1984–1991, 2013. 15
  13. 13.Brooks, T., Peebles, B., Holmes, C., DePue, W., Guo, Y., Jing, L., Schnurr, D., Taylor, J., Luhman, T., Luhman, E., Ng, C., Wang, R., and Ramesh, A. Video generation models as world simulators. 2024. URL https://openai.com/research. 1, 3, 9, 16, 22
  14. 14.Brown, T. B. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 1
  15. 15.Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al. Genie: Generative interactive environments. In Forty-first International Conference on Machine Learning, 2024. 9
  16. 16.Byeon, M., Park, B., Kim, H., Lee, S., Baek, W., and Kim, S. Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset, 2022. 16
  17. 17.Cao, Q., Wang, D., Li, X., Chen, Y., Ma, C., and Yang, X. Teaching video diffusion model with latent physical phenomenon knowledge. arXiv preprint arXiv:2411.11343, 2024. 15
  18. 18.Carreira, J. and Zisserman, A. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299–6308, 2017. 6
  19. 19.de Silva, B. M., Higdon, D. M., Brunton, S. L., and Kutz, J. N. Discovery of physics from data: Universal laws and discrepancies. Frontiers in artificial intelligence, 3:25, 2020. 15
  20. 20.Diederik, P. K. and Ba, J. Adam: A method for stochastic optimization. In ICLR, 2015. 16
  21. 21.Du, Y. and Kaelbling, L. Compositional generative modeling: A single model is not all you need. arXiv preprint arXiv:2402.01103, 2024. 2
  22. 22.Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., Orgad, E., Entezari, R., Daras, G., Pratt, S., Ramanujan, V., Bitton, Y., Marathe, K., Mussmann, S., Vencu, R., Cherti, M., Krishna, R., Koh, P. W., Saukh, O., Ratner, A. J., Song, S., Hajishirzi, H., Farhadi, A., Beaumont, R., Oh, S., Dimakis, A. G., Jitsev, J., Carmon, Y., Shankar, V., and Schmidt, L. Datacomp: In search of the next generation of multimodal datasets. ArXiv, abs/2304.14108, 2023. URL https://api.semanticscholar.org/CorpusID:258352812. 16
  23. 23.Gao, S., Yang, J., Chen, L., Chitta, K., Qiu, Y., Geiger, A., Zhang, J., and Li, H. Vista: A generalizable driving world model with high fidelity and versatile controllability. arXiv preprint arXiv:2405.17398, 2024. 9
  24. 24.Girdhar, R., Gustafson, L., Adcock, A., and van der Maaten, L. Forward prediction for physical reasoning. arXiv preprint arXiv:2006.10734, 2020. 15
  25. 25.Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I. Emu video: Factorizing text-to-video generation by explicit image conditioning. arXiv preprint arXiv:2311.10709, 2023. 9
  26. 26.Groth, O., Fuchs, F. B., Posner, I., and Vedaldi, A. Shapestacks: Learning vision-based physical intuition for generalised object stacking. In Proceedings of the european conference on computer vision (eccv), pp. 702–717, 2018. 15
  27. 27.Ha, D. and Schmidhuber, J. Recurrent world models facilitate policy evolution. Advances in neural information processing systems, 31, 2018. 9
  28. 28.Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603, 2019. 9
  29. 29.Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. Mastering atari with discrete world models. arXiv preprint arXiv:2010.02193, 2020. 9
  30. 30.Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023. 19
  31. 31.He, C., Shen, Y., Fang, C., Xiao, F., Tang, L., Zhang, Y., Zuo, W., Guo, Z., and Li, X. Diffusion models in low-level vision: A survey. arXiv preprint arXiv:2406.11138, 2024. 9
  32. 32.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 15
  33. 33.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022a. 9
  34. 34.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. Advances in Neural Information Processing Systems, 35:8633–8646, 2022b. 9
  35. 35.Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 17
  36. 36.Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G. Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023. 1, 9
  37. 37.Hu, Y., Tang, X., Yang, H., and Zhang, M. Case-based or rule-based: How do transformers do the math? ICML, 2024. 2, 7, 22
  38. 38.Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 21807–21818, 2024. 9
  39. 39.Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A. Image-to-image translation with conditional adversarial networks. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5967–5976, 2016. URL https://api.semanticscholar.org/CorpusID:6200260. 16
  40. 40.Jia, F., Mao, W., Liu, Y., Zhao, Y., Wen, Y., Zhang, C., Zhang, X., and Wang, T. Adriver-i: A general world model for autonomous driving. arXiv preprint arXiv:2311.13549, 2023. 9
  41. 41.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. 1
  42. 42.Khachatryan, L., Movsisyan, A., Tadevosyan, V., Henschel, R., Wang, Z., Navasardyan, S., and Shi, H. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 15954–15964, 2023. 9
  43. 43.Kingma, D. P. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 9
  44. 44.Kondratyuk, D., Yu, L., Gu, X., Lezama, J., Huang, J., Hornung, R., Adam, H., Akbari, H., Alon, Y., Birodkar, V., et al. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125, 2023. 9
  45. 45.Liao, M., Ye, Q., Zuo, W., Wan, F., Wang, T., Zhao, Y., Wang, J., Zhang, X., et al. Evaluation of text-to-video generation models: A dynamics perspective. Advances in Neural Information Processing Systems, 37:109790–109816, 2024. 9
  46. 46.Lin, S., Liu, B., Li, J., and Yang, X. Common diffusion noise schedules and sample steps are flawed. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 5404–5411, 2024. 16
  47. 47.Liu, Y., Cun, X., Liu, X., Wang, X., Zhang, Y., Chen, H., Liu, Y., Zeng, T., Chan, R., and Shan, Y. Evalcrafter: Benchmarking and evaluating large video generation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22139–22149, 2024. 9
  48. 48.Loshchilov, I. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 16
  49. 49.Melnik, A., Schiewer, R., Lange, M., Muresanu, A., Saeidi, M., Garg, A., and Ritter, H. Benchmarks for physical reasoning ai. arXiv preprint arXiv:2312.10728, 2023. 15
  50. 50.Peebles, W. and Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4195–4205, 2023. 9, 16
  51. 51.Riveland, R. and Pouget, A. Natural language instructions induce compositional generalization in networks of neurons. Nature Neuroscience, 27(5):988–999, 2024. 3, 19
  52. 52.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022. 9
  53. 53.Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, pp. 234–241. Springer, 2015. 9
  54. 54.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, 2022. 16
  55. 55.Schott, L., Von Kugelgen, J., Träuble, F., Gehler, P., Russell, C., Bethge, M., Scholkopf, B., Locatello, F., and Brendel, W. Visual representation learning does not generalize strongly within the same domain. arXiv preprint arXiv:2107.08221, 2021. 22
  56. 56.Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. Mastering atari, go, chess and shogi by planning with a learned model. Nature, 588(7839):604–609, 2020. 9
  57. 57.Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. Mastering the game of go with deep neural networks and tree search. nature, 529(7587):484–489, 2016. 9
  58. 58.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020. 15
  59. 59.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. 3
  60. 60.Sun, K., Huang, K., Liu, X., Wu, Y., Xu, Z., Li, Z., and Liu, X. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation. arXiv preprint arXiv:2407.14505, 2024. 9
  61. 61.Sutton, R. S. Reinforcement learning: An introduction. A Bradford Book, 2018. 9
  62. 62.Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 6
  63. 63.Van Den Oord, A., Vinyals, O., et al. Neural discrete representation learning. Advances in neural information processing systems, 30, 2017. 9
  64. 64.Wang, W., Yang, H., Tuo, Z., He, H., Zhu, J., Fu, J., and Liu, J. Swap attention in spatiotemporal diffusions for text-to-video generation. 2023. URL https://api.semanticscholar.org/CorpusID:258762479. 16
  65. 65.Wang, X., Zhang, X., Zhu, Y., Guo, Y., Yuan, X., Xiang, L., Wang, Z., Ding, G., Brady, D., Dai, Q., and Fang, L. Panda: A gigapixel-level human-centric video dataset. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3265–3275, 2020. doi: 10.1109/CVPR42600.2020.00333. 16
  66. 66.Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 6
  67. 67.Weitnauer, E., Goldstone, R. L., and Ritter, H. Perception and simulation during concept learning. Psychological Review, 2023. 15
  68. 68.Wu, J., Lim, J. J., Zhang, H., Tenenbaum, J. B., and Freeman, W. T. Physics 101: Learning physical object properties from unlabeled videos. In BMVC, volume 2, pp. 7, 2016. 15
  69. 69.Wu, J. Z., Ge, Y., Wang, X., Lei, S. W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., and Shou, M. Z. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7623–7633, 2023. 9
  70. 70.Xu, K., Zhang, M., Li, J., Du, S. S., Kawarabayashi, K.-i., and Jegelka, S. How neural networks extrapolate: From feedforward to graph neural networks. arXiv preprint arXiv:2009.11848, 2020. 6
  71. 71.Xue, C., Pinto, V., Gamage, C., Nikonova, E., Zhang, P., and Renz, J. Phy-q as a measure for physical reasoning intelligence. Nature Machine Intelligence, 5(1):83–93, 2023. 15
  72. 72.Xue, T., Chen, B., Wu, J., Wei, D., and Freeman, W. T. Video enhancement with task-oriented flow. International Journal of Computer Vision, pp. 1–20, 2017. URL https://api.semanticscholar.org/CorpusID:40412298. 16
  73. 73.Yang, J., Gao, S., Qiu, Y., Chen, L., Li, T., Dai, B., Chitta, K., Wu, P., Zeng, J., Luo, P., et al. Generalized predictive model for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14662–14672, 2024. 1
  74. 74.Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 2023. 1, 9
  75. 75.Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., and Tenenbaum, J. B. Clevrer: Collision events for video representation and reasoning. arXiv preprint arXiv:1910.01442, 2019. 15
  76. 76.Yu, L., Cheng, Y., Sohn, K., Lezama, J., Zhang, H., Chang, H., Hauptmann, A. G., Yang, M.-H., Hao, Y., Essa, I., et al. Magvit: Masked generative video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10459–10469, 2023a. 9
  77. 77.Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., et al. Language model beats diffusion–tokenizer is key to visual generation. arXiv preprint arXiv:2310.05737, 2023b. 3, 9, 16
  78. 78.Yue, Y., Kang, B., Xu, Z., Huang, G., and Yan, S. Value-consistent representation learning for data-efficient reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2023. 9
  79. 79.Yue, Y., Wang, Y., Kang, B., Han, Y., Wang, S., Song, S., Feng, J., and Huang, G. Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution. Advances in Neural Information Processing Systems, 2024. 1
  80. 80.Zeng, Y., Wei, G., Zheng, J., Zou, J., Wei, Y., Zhang, Y., and Li, H. Make pixels dance: High-dynamic video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8850–8860, 2024. 9
  81. 81.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595, 2018. 6
  82. 82.Zhang, S., Wang, J., Zhang, Y., Zhao, K., Yuan, H., Qin, Z., Wang, X., Zhao, D., and Zhou, J. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:2311.04145, 2023a. 9
  83. 83.Zhang, Y., Wei, Y., Jiang, D., Zhang, X., Zuo, W., and Tian, Q. Controlvideo: Training-free controllable text-to-video generation. arXiv preprint arXiv:2305.13077, 2023b. 9
  84. 84.Zheng, W., Song, R., Guo, X., and Chen, L. Genad: Generative end-to-end autonomous driving. arXiv preprint arXiv:2402.11502, 2024. 9

Citation

MLA
Kang, B., et al. “How Far Is Video Generation from World Model: A Physical Law Perspective”. arXiv, 2024, http://arxiv.org/abs/2411.02385v2.
APA
Kang, B., Yue, Y., Lu, R., Lin, Z., Zhao, Y., Wang, K., Huang, G., & Feng, J. (2024). How Far is Video Generation from World Model: A Physical Law Perspective. arXiv. http://arxiv.org/abs/2411.02385v2
Chicago
Kang, B., Y. Yue, R. Lu, et al. 2024. “How Far Is Video Generation from World Model: A Physical Law Perspective”. arXiv. http://arxiv.org/abs/2411.02385v2.
Harvard
Kang, B. et al. (2024) “How Far is Video Generation from World Model: A Physical Law Perspective”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2411.02385v2.
Vancouver
1. Kang B, Yue Y, Lu R, Lin Z, Zhao Y, Wang K, Huang G, Feng J (2024) How Far is Video Generation from World Model: A Physical Law Perspective. arXiv

BibTeX

@article{kang2024how,
  title = {How Far is Video Generation from World Model: A Physical Law Perspective},
  author = {Kang, Bingyi and Yue, Yang and Lu, Rui and Lin, Zhijie and Zhao, Yang and Wang, Kaixin and Huang, Gao and Feng, Jiashi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2411.02385v2},
  eprint = {2411.02385}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/