Model as a Game: On Numerical and Spatial Consistency for Generative Games

Jingye ChenYuzhong ZhaoYupan HuangLei CuiLi DongTengchao LvQifeng ChenFuru Wei

article2025ICCV7 citationsBest Paper Runnerup

Develops specialized numerical and spatial consistency modules for Diffusion Transformers, solving persistent score-tracking and environment-continuity failures in generative game systems with minimal latency overhead.

Listen

Recent advances in generative artificial intelligence have enabled real-time game video generation directly from player inputs, presenting a potential alternative to labor-intensive traditional game engines. However, current models treat interactive game simulation merely as a next-frame pixel prediction task. This simplification causes critical failures in gameplay logic: numerical scores fluctuate erratically regardless of player actions, and previously visited environments morph or disappear when revisited, breaking player immersion.

The article evaluates and demonstrates a new framework designed to enforce numerical and spatial consistency in generative gameplay. The objective is to establish an architecture where state changes and persistent environmental layouts remain stable across indefinite play sessions.

The authors conducted an empirical study using three 2D games of varying complexity: Traveler, Pong, and Pac-Man. They enhanced a Diffusion Transformer baseline by integrating two explicit modules. First, a lightweight neural network called LogicNet predicts gameplay event triggers, combining with an external numerical record that supplies explicit digit tokens to condition frame generation. Second, an external spatial module maintains a persistent map of explored areas, retrieving local map tokens to guide rendering and linking newly generated frames back to the map using a sliding-window algorithm. These modules add less than 2% to the baseline's total parameter count.

The findings show substantial improvements in gameplay fidelity across all test environments. In the Traveler game, the proposed modules increased the numerical consistency score from 0.3245 to 0.9141 and spatial consistency signal quality from 16.15 to 33.64. Text rendering guided by digit tokens aligned with target scores in 99% of cases. Furthermore, visual fidelity remained stable even when scaling generation from 64 to 256 consecutive frames, while incurring negligible computational overhead—LogicNet required only 0.0004 seconds and spatial matching 0.015 seconds per inference step.

These results demonstrate that pure end-to-end pixel generation is insufficient for interactive game simulations; separating explicit game logic and persistent spatial memory from visual synthesis resolves the primary usability bottlenecks of generative engines. The approach enables practical features such as map customization and precise player tracking without increasing deployment costs or hardware requirements.

Organizations developing generative interactive media should adopt hybrid architectures that decouple explicit logical state management from generative rendering pipelines. Moving forward, engineering teams should conduct pilot implementations to adapt this framework to complex 3D environments, evaluate dynamic multiplayer states, and increase inference throughput to achieve standard 30 frames-per-second interactive rates.

Confidence in these findings is high for controlled 2D environments, supported by quantitative benchmarks and user studies. However, key limitations remain: the sliding-window matching method fails in visually repetitive or monochrome backgrounds, and the current models occasionally generate physics anomalies without explicit physical constraints. Additional testing is required before applying this technique to complex, high-resolution 3D worlds.

arXiv: 2503.21172
Cover for Model as a Game: On Numerical and Spatial Consistency for Generative Games

Abstract

Recent advances in generative models have significantly impacted game generation. However, despite producing high-quality graphics and adequately receiving player input, existing models often fail to maintain fundamental game properties such as numerical and spatial consistency. Numerical consistency ensures gameplay mechanics correctly reflect score changes and other quantitative elements, while spatial consistency prevents jarring scene transitions, providing seamless player experiences. In this paper, we revisit the paradigm of generative games to explore what truly constitutes a Model as a Game (MaaG) with a well-developed mechanism. We begin with an empirical study on ``Traveler'', a 2D game created by an LLM featuring minimalist rules yet challenging generative models in maintaining consistency. Based on the DiT architecture, we design two specialized modules: (1) a numerical module that integrates a LogicNet to determine event triggers, with calculations processed externally as conditions for image generation; and (2) a spatial module that maintains a map of explored areas, retrieving location-specific information during generation and linking new observations to ensure continuity. Experiments across three games demonstrate that our integrated modules significantly enhance performance on consistency metrics compared to baselines, while incurring minimal time overhead during inference.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Game Creation
  • 2.2 Controllable Generative Models
  • 3 Methodology
  • 3.1 Preliminary
  • 3.2 Numerical Module
  • 3.3 Spatial Module
  • 3.4 Module Generalization to Diverse Games
  • 4 Experiments
  • 4.1 Implementation Details
  • 4.2 Evaluation Metrics
  • 4.3 Quantitative Results
  • 4.4 Qualitative Results
  • 5 Discussion
  • 6 Conclusion
  • A Detailed Architecture of LogicNet
  • B Details of CNN in the Spatial Module
  • C Examine the Effectiveness of Map Construction for Traveler
  • D Constructing 2D Map for Pac-Man
  • E Details of Using LLMs for Creating Games
  • F Details of Models Provided by PGG
  • G Detailed Architecture of Valid Action Model and Valid Numerical Model
  • H Visualizations of Failure Cases
  • References

Knowls

  1. Knowl 1 — Consistency-Aware Generative Game Framework

    model/method

    The Model as a Game (MaaG) architecture extends recurrent diffusion transformers to maintain both numerical and spatial consistency during interactive game generation. Rather than treating game synthesis purely as unconstrained pixel prediction, the framework augments a backbone Diffusion Transformer (DiT) with two specialized components: a trainable numerical module (LogicNet) and an explicit spatial module.

    At each time step nn, a frame In∈RH×W×3I_n \in \mathbb{R}^{H \times W \times 3} is encoded by a pre-trained Variational Autoencoder (VAE) into a latent representation zn∈RH4×W4×4z_n \in \mathbb{R}^{\frac{H}{4} \times \frac{W}{4} \times 4}. The DiT takes as input the noisy latent znz_n, the previous hidden state hn−1∈RH4×W4×32h_{n-1} \in \mathbb{R}^{\frac{H}{4} \times \frac{W}{4} \times 32} concatenated along the channel dimension, the player action an−1a_{n-1} injected via cross-attention, digit tokens representing explicit game scores, and local map tokens extracted from an explicit global map. The DiT outputs an updated hidden state hnh_n, which is decoded by a convolutional network and VAE decoder to reconstruct the generated frame InI_n.

  2. Knowl 2 — Numerical Consistency Module via LogicNet and Digit Token Conditioning

    model/method

    To prevent arbitrary score fluctuations and maintain game logic, the numerical module decouples logic event prediction from visual rendering.

    1. Event Prediction (LogicNet): LogicNet is a lightweight neural network (0.6M parameters) that accepts the prior hidden state hn−1∈RH4×W4×32h_{n-1} \in \mathbb{R}^{\frac{H}{4} \times \frac{W}{4} \times 32} and player action an−1a_{n-1}. It processes hn−1h_{n-1} through two convolutional and downsampling layers, concatenates the resulting flattened features with a learnable action embedding, and feeds the fused representation into an MLP to predict whether a discrete game event (e.g., score increment) occurs in frame nn.

    2. External Calculator & Digit Token Conditioning: When LogicNet triggers an event, an external deterministic logic calculator computes the updated numerical value. The resulting number is decomposed into individual digit characters (e.g., hundreds, tens, and units). Each digit is converted into a learnable embedding (digit tokens) and fed as conditional tokens to the Diffusion Transformer. This offloads arithmetic operations from the generative diffusion backbone, restricting its role to visual rendering of the provided score digits.

  3. Knowl 3 — Spatial Consistency Module via Global Map Maintenance and Sliding Window Linking

    model/method

    To ensure that revisited game areas retain visual coherence, the spatial consistency module maintains an explicit persistent canvas/map MM throughout gameplay instead of relying solely on recurrent hidden states.

    1. Auxiliary Context Retrieval: Given the stored map Mn−1M_{n-1} and the player's central position xn−1x_{n-1}, the module extracts an extended local map slice mn−1m_{n-1} covering a region (xn−1−δ1,xn−1+δ1)(x_{n-1} - \delta_1, x_{n-1} + \delta_1), where the window width 2δ1>W2\delta_1 > W extends beyond the camera observation width WW. Unexplored areas within this slice are filled with black pixels. This local map is encoded via a lightweight CNN (0.1M parameters) into map tokens that condition the Diffusion Transformer during frame generation.

    2. Observation Registration & Linking: Once the new frame InI_n is generated, it is registered against the existing global map Mn−1M_{n-1} via sliding window matching over a local search interval (xn−1−δ2,xn−1+δ2)(x_{n-1} - \delta_2, x_{n-1} + \delta_2), where δ2\delta_2 reflects the maximum movement per step. For each candidate offset, Peak Signal-to-Noise Ratio (PSNR) is evaluated across the overlapping regions. The offset maximizing PSNR determines the updated player coordinate xnx_n, and newly observed pixels are written into the map canvas to form MnM_n.

  4. Knowl 4 — Joint Optimization Objective for Consistency-Aware Generative Games

    equation

    The generative model is trained end-to-end on sequence samples using a composite loss that combines diffusion denoising with logic event classification:

    L=Ldenoise+λLlogic\mathcal{L} = \mathcal{L}_{\text{denoise}} + \lambda \mathcal{L}_{\text{logic}}

    where λ\lambda is a balancing hyperparameter set to 10−410^{-4}.

    The denoising loss Ldenoise\mathcal{L}_{\text{denoise}} follows the Denoising Diffusion Probabilistic Model (DDPM) formulation:

    Ldenoise=Eϵ,n[∥ϵ−ϵθ(zn,an−1,hn−1)∥22]\mathcal{L}_{\text{denoise}} = \mathbb{E}_{\epsilon, n} \left[ \left\| \epsilon - \epsilon_\theta(z_n, a_{n-1}, h_{n-1}) \right\|_2^2 \right]

    where ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) is Gaussian noise, ϵθ\epsilon_\theta denotes the DiT parameterized by θ\theta, hn−1h_{n-1} is the recurrent hidden state, an−1a_{n-1} is the input action, and the noisy latent znz_n at diffusion step tt is given by:

    zn=αˉtVAE(In)+1−αˉtϵz_n = \sqrt{\bar{\alpha}_t} \text{VAE}(I_n) + \sqrt{1 - \bar{\alpha}_t} \epsilon

    with αˉt\bar{\alpha}_t representing the cumulative noise schedule parameter. Llogic\mathcal{L}_{\text{logic}} is the cross-entropy loss between LogicNet's event trigger prediction and the ground truth game logic state.

  5. Knowl 5 — Spatial Consistency Evaluation Metric (SpaCon)

    definition

    Spatial Consistency (SpaCon) is an evaluation metric designed to assess whether generative game models preserve the visual structure of previously visited areas when players revisit them. Because newly explored regions outside the initial view are generated stochastically, standard full-frame ground-truth comparisons (such as baseline PSNR or LPIPS) cannot evaluate consistency across dynamic rollouts.

    SpaCon is formally defined as:

    SpaCon:=E[PSNR(In′,mn−1′)]\text{SpaCon} := \mathbb{E} \left[ \text{PSNR}\left(I'_n, m'_{n-1}\right) \right]

    where In′I'_n represents the newly visible image area at time step nn, and mn−1′m'_{n-1} denotes the corresponding spatial region recorded in the persistent map up to step n−1n-1. A higher SpaCon indicates that scenes reproduced upon revisiting match the historical canvas rendered during earlier visits.

  6. Knowl 6 — Numerical Consistency Evaluation Metric (NumCon)

    definition

    Numerical Consistency (NumCon) measures whether numeric gameplay indicators (e.g., scores, counters, health) update accurately in response to triggering events rather than fluctuating arbitrarily or failing to update.

    NumCon is computed as the F-measure of predicted numerical transitions between consecutive frames:

    1. For models with explicit score trackers, event triggers are verified against game mechanics.
    2. For baseline models without internal score representations, an optical character recognition model (TrOCR) reads score text directly from generated video frames.
    3. A separate pre-trained Valid Numerical Model processes frame-action pairs to predict the expected ground-truth event transition, and the F-measure between actual observed score transitions and predicted transitions determines the NumCon score.
  7. Knowl 7 — Impact of Consistency Modules Across 2D Generative Games

    data/table

    Integrating numerical and spatial consistency modules into the baseline recurrent DiT architecture provides substantial gains in numerical consistency (NumCon), spatial consistency (SpaCon), visual quality (FID, FVD), and action accuracy (ActAcc) across multiple 2D games (Traveler, Pong, Pac-Man), evaluated across 100 generated episodes with 256-frame rollouts:

    Game Steps Modules Metrics
    FID↓\downarrow FVD↓\downarrow ActAcc↑\uparrow NumCon↑\uparrow SpaCon↑\uparrow
    Traveler 8 56.85 84.93 0.9657 0.3245 16.15
    Traveler 8 ✓ 43.76 51.58 0.9909 0.9141 33.64
    Traveler 16 51.62 83.95 0.9658 0.3252 16.01
    Traveler 16 ✓ 43.75 52.12 0.9916 0.9219 31.39
    Pong 8 29.30 73.62 0.6534 0.0889 -
    Pong 8 ✓ 24.67 78.50 0.7871 0.6847 -
    Pong 16 26.40 78.22 0.6161 0.0465 -
    Pong 16 ✓ 25.01 82.18 0.8717 0.5911 -
    Pac-Man 8 21.70 1472.87 0.7382 0.2667 13.54
    Pac-Man 8 ✓ 15.43 793.96 0.6862 0.6087 18.56
    Pac-Man 16 19.61 1194.74 0.7260 0.3871 13.70
    Pac-Man 16 ✓ 14.24 751.17 0.7869 0.7917 17.88

    In Traveler, NumCon improves by +0.5896 (from 0.3245 to 0.9141 at 8 steps) and SpaCon improves by +17.49 dB (from 16.15 to 33.64). In Pong and Pac-Man, NumCon increases substantially while FVD drops significantly in Pac-Man (from 1472.87 to 793.96 at 8 steps).

  8. Knowl 8 — Robustness Across Denoising Steps and Prediction Sequence Lengths

    empirical result

    Evaluation of the consistency-enhanced model across varying sequence lengths (64, 128, and 256 frames) and denoising steps (8 vs 16 steps) demonstrates that the framework scales to long autoregressive rollouts without quality degradation:

    1. Autoregressive Rollout Length: In Traveler with 8 denoising steps, extending sequence length from 64 to 256 frames improves FID from 47.43 to 43.76, decreases FVD from 79.06 to 51.58, and maintains stable numerical consistency (0.9315 to 0.9141) and spatial consistency (33.01 to 33.64 dB). This demonstrates that recurrent hidden states paired with explicit map conditioning avoid progressive drift.

    2. Denoising Steps Effect: Increasing diffusion sampling steps from 8 to 16 yields minor visual gains in Traveler and Pac-Man (e.g., Pac-Man FID improves from 15.43 to 14.24 at length 256), but can slightly degrade temporal FVD metrics in Pong (worsening from 78.50 to 82.18), indicating that 8 sampling steps achieves an optimal balance between quality and computational cost.

  9. Knowl 9 — Computational Overhead and Real-Time Inference Efficiency

    empirical result

    The consistency framework introduces negligible parameter and runtime overhead relative to the baseline generation model:

    1. Model Capacity: The baseline DiT has 33M parameters. LogicNet adds 0.6M parameters and the spatial CNN adds 0.1M parameters, representing an overall increase of under 2% in total trainable parameters.
    2. Inference Latency: On a single NVIDIA A100 40GB GPU, LogicNet requires 0.0004 seconds per step, and spatial map retrieval and sliding window linking requires 0.015 seconds per step. Total memory consumption is 1.2 GB.
    3. Throughput: Interactive inference achieves 10 FPS with 8 diffusion steps and 6 FPS with 16 diffusion steps, matching baseline generation speed while providing numerical and spatial stability.
  10. Knowl 10 — Failure Modes of Sliding Window Map Registration and Physics Simulation

    limitation

    The framework exhibits two principal failure modes:

    1. Repetitive/Uniform Textures: The sliding window registration algorithm relies on PSNR variations over overlapping visual features to detect player translation. When entering scenes with homogeneous or uniformly colored backgrounds (e.g., a continuous single-colored structure), the PSNR remains constant across candidate window positions. As a result, the registration fails to detect movement, and the global map does not expand.

    2. Unsupervised Physics Violations: In games with complex projectile or particle dynamics (such as the ball in Pong), the diffusion model occasionally produces unphysical trajectories (e.g., a ball suddenly reversing vertical direction mid-flight) due to the lack of explicit physical dynamical laws in the supervision loss.

Coverage note — Omitted minor implementation details such as the prompt templates used for LLM-assisted Pygame generation, baseline VAE pretraining hyperparameters, and the 8-participant subjective user study summary.

References

  1. 1.Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. In Forty-first International Conference on Machine Learning, 2024. 3
  2. 2.Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin, and Hao Chen. Gamegen-x: Interactive open-world game video generation. arXiv preprint arXiv:2411.00769, 2024. 2, 3
  3. 3.Boyuan Chen, Diego Martı Monso, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token  prediction meets full-sequence diffusion. Advances in Neural Information Processing Systems, 37:24081–24125, 2025. 3
  4. 4.Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei. Textdiffuser: Diffusion models as text painters. Advances in Neural Information Processing Systems, 36:9353–9387, 2023. 3
  5. 5.Jingye Chen, Yupan Huang, Tengchao Lv, Lei Cui, Qifeng Chen, and Furu Wei. Textdiffuser-2: Unleashing the power of language models for text rendering. In European Conference on Computer Vision, pages 386–402. Springer, 2024. 3, 5
  6. 6.Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu, Pan Xie, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. Control-a-video: Controllable text-to-video generation with diffusion models. arXiv e-prints, pages arXiv–2305, 2023. 3
  7. 7.claude. Link: https://www.anthropic.com/news/claude-3-5-sonnet, 2024. 5
  8. 8.Epic Games. Unreal engine. 3
  9. 9.Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer,  Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, 2024. 3
  10. 10.Ruili Feng, Han Zhang, Zhantao Yang, Jie Xiao, Zhilei Shu, Zhiheng Liu, Andy Zheng, Yukun Huang, Yu Liu, and Hongyang Zhang. The matrix: Infinite-horizon world generation with real-time moving control. arXiv preprint arXiv:2412.03568, 2024. 3
  11. 11.flux. Link: https://github.com/black-forest-labs/flux, 2024. 3
  12. 12.genie 2. Link: https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/, 2024. 3
  13. 13.John K Haas. A history of the unity game engine. 2014. 3
  14. 14.Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for text-to-video generation. arXiv preprint arXiv:2404.02101, 2024. 3
  15. 15.Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian, Hongyu Liu, Xiaowei Chi, Runtao Liu, Ruibin Yuan, Yazhou Xing, Wenhai Wang, et al. Llms meet multimodal generation and editing: A survey. arXiv preprint arXiv:2405.19334, 2024. 3
  16. 16.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017. 6
  17. 17.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 4
  18. 18.Zhihao Hu and Dong Xu. Videocontrolnet: A motion-guided video-to-video translation framework by using diffusion model with controlnet. arXiv preprint arXiv:2307.14073, 2023. 3
  19. 19.Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou. Composer: Creative and controllable image synthesis with composable conditions. arXiv preprint arXiv:2302.09778, 2023. 3
  20. 20.Anssi Kanervisto, Dave Bignell, Linda Yilin Wen, Martin Grayson, Raluca Georgescu, Sergio Valcarcel Macua, Shan Zheng Tan, Tabish Rashid, Tim Pearce, Yuhan Cao, et al. World and human action models towards gameplay ideation. Nature, 638(8051): 656–663, 2025. 3
  21. 21.Seung Wook Kim, Yuhao Zhou, Jonah Philion, Antonio Torralba, and Sanja Fidler. Learning to simulate dynamic environments with gamegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1231–1240, 2020. 2, 3, 6, 7
  22. 22.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 4
  23. 23.Dan Kondratyuk, Lijun Yu, Xiuye Gu, Jose Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy  Yan, Ming-Chang Chiu, et al. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125, 2023. 3
  24. 24.Jialu Li, Yuanzhen Li, Neal Wadhwa, Yael Pritch, David E Jacobs, Michael Rubinstein, Mohit Bansal, and Nataniel Ruiz. Unbounded: A generative infinite game of character life simulation. arXiv preprint arXiv:2410.18975, 2024. 3
  25. 25.Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei. Trocr: Transformer-based optical character recognition with pre-trained models. In Proceedings of the AAAI conference on artificial intelligence, pages 13094–13102, 2023. 6, 7, 14
  26. 26.Ming Li, Taojiannan Yang, Huafeng Kuang, Jie Wu, Zhaoning Wang, Xuefeng Xiao, and Chen Chen. Controlnet++: Improving conditional controls with efficient consistency feedback: Project page: liming-ai. github. io/controlnet_plus_plus. In European Conference on Computer Vision, pages 129–147. Springer, 2024. 3
  27. 27.Willi Menapace, Stephane Lathuiliere, Sergey Tulyakov, Aliaksandr Siarohin, and Elisa Ricci. Playable video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10061–10070, 2021. 3
  28. 28.Willi Menapace, Aliaksandr Siarohin, Stephane Lathuili ere, Panos Achlioptas, Vladislav Golyanik, Sergey Tulyakov, and Elisa Ricci.  Promptable game models: Text-guided game simulation via masked diffusion models. ACM Transactions on Graphics, 43(2):1–16, 2024. 3
  29. 29.Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proceedings of the AAAI conference on artificial intelligence, pages 4296–4304, 2024. 3
  30. 30.oasis. Link: https://github.com/etched-ai/open-oasis, 2024. 2, 3
  31. 31.Hao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai, Juntao Zhang, Kecheng Zheng, Xiaowei Zhou, Qifeng Chen, and Yujun Shen. Codef: Content deformation fields for temporally consistent video processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8089–8099, 2024. 3
  32. 32.William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023. 2, 3, 4
  33. 33.Bohao Peng, Jian Wang, Yuechen Zhang, Wenbo Li, Ming-Chang Yang, and Jiaya Jia. Controlnext: Powerful and efficient control for image and video generation. arXiv preprint arXiv:2408.06070, 2024. 3
  34. 34.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent  diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 3
  35. 35.Shyam Sudhakaran, Miguel Gonzalez-Duque, Matthias Freiberger, Claire Glanois, Elias Najarro, and Sebastian Risi. Mariogpt:  Open-ended text2level generation through large language models. Advances in Neural Information Processing Systems, 36:54213–54227, 2023. 3
  36. 36.Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 6
  37. 37.Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter. Diffusion models are real-time game engines. arXiv preprint arXiv:2408.14837, 2024. 2, 3
  38. 38.Mingyu Yang, Junyou Li, Zhongbin Fang, Sheng Chen, Yangbin Yu, Qiang Fu, Wei Yang, and Deheng Ye. Playable game generation. arXiv preprint arXiv:2412.00887, 2024. 2, 3, 6, 7, 12
  39. 39.Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao. Direct-a-video: Customized video generation with user-directed camera movement and object motion. In ACM SIGGRAPH 2024 Conference Papers, pages 1–12, 2024. 3
  40. 40.ying. Link: https://giantailab.github.io/yinggame/, 2024. 3
  41. 41.Jiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan, Di Zhang, and Xihui Liu. Gamefactory: Creating new games with generative interactive videos. arXiv preprint arXiv:2501.08325, 2025. 2, 3
  42. 42.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3836–3847, 2023. 3
  43. 43.Qihang Zhang, Shuangfei Zhai, Miguel Angel Bautista, Kevin Miao, Alexander Toshev, Joshua Susskind, and Jiatao Gu. World-consistent video diffusion with explicit 3d modeling. arXiv preprint arXiv:2412.01821, 2024. 3
  44. 44.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018. 6
  45. 45.Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, and Qi Tian. Controlvideo: Training-free controllable text-to-video generation. arXiv preprint arXiv:2305.13077, 2023. 3
  46. 46.Yiyuan Zhang, Yuhao Kang, Zhixin Zhang, Xiaohan Ding, Sanyuan Zhao, and Xiangyu Yue. Interactivevideo: User-centric controllable video generation with synergistic multimodal instructions. arXiv preprint arXiv:2402.03040, 2024. 3
  47. 47.Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models. Advances in Neural Information Processing Systems, 36:11127–11150, 2023. 3
  48. 48.Haitao Zhou, Chuang Wang, Rui Nie, Jinxiao Lin, Dongdong Yu, Qian Yu, and Changhu Wang. Trackgo: A flexible and efficient method for controllable video generation. arXiv preprint arXiv:2408.11475, 2024. 3

Citation

MLA
Chen, J., et al. “Model as a Game: On Numerical and Spatial Consistency for Generative Games”. arXiv, 2025, http://arxiv.org/abs/2503.21172v1.
APA
Chen, J., Zhao, Y., Huang, Y., Cui, L., Dong, L., Lv, T., Chen, Q., & Wei, F. (2025). Model as a Game: On Numerical and Spatial Consistency for Generative Games. arXiv. http://arxiv.org/abs/2503.21172v1
Chicago
Chen, J., Y. Zhao, Y. Huang, et al. 2025. “Model as a Game: On Numerical and Spatial Consistency for Generative Games”. arXiv. http://arxiv.org/abs/2503.21172v1.
Harvard
Chen, J. et al. (2025) “Model as a Game: On Numerical and Spatial Consistency for Generative Games”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2503.21172v1.
Vancouver
1. Chen J, Zhao Y, Huang Y, Cui L, Dong L, Lv T, Chen Q, Wei F (2025) Model as a Game: On Numerical and Spatial Consistency for Generative Games. arXiv

BibTeX

@article{chen2025model,
  title = {Model as a Game: On Numerical and Spatial Consistency for Generative Games},
  author = {Chen, Jingye and Zhao, Yuzhong and Huang, Yupan and Cui, Lei and Dong, Li and Lv, Tengchao and Chen, Qifeng and Wei, Furu},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2503.21172v1},
  eprint = {2503.21172}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/