Model as a Game: On Numerical and Spatial Consistency for Generative Games
Jingye ChenYuzhong ZhaoYupan HuangLei CuiLi DongTengchao LvQifeng ChenFuru Wei
Develops specialized numerical and spatial consistency modules for Diffusion Transformers, solving persistent score-tracking and environment-continuity failures in generative game systems with minimal latency overhead.
Recent advances in generative artificial intelligence have enabled real-time game video generation directly from player inputs, presenting a potential alternative to labor-intensive traditional game engines. However, current models treat interactive game simulation merely as a next-frame pixel prediction task. This simplification causes critical failures in gameplay logic: numerical scores fluctuate erratically regardless of player actions, and previously visited environments morph or disappear when revisited, breaking player immersion.
The article evaluates and demonstrates a new framework designed to enforce numerical and spatial consistency in generative gameplay. The objective is to establish an architecture where state changes and persistent environmental layouts remain stable across indefinite play sessions.
The authors conducted an empirical study using three 2D games of varying complexity: Traveler, Pong, and Pac-Man. They enhanced a Diffusion Transformer baseline by integrating two explicit modules. First, a lightweight neural network called LogicNet predicts gameplay event triggers, combining with an external numerical record that supplies explicit digit tokens to condition frame generation. Second, an external spatial module maintains a persistent map of explored areas, retrieving local map tokens to guide rendering and linking newly generated frames back to the map using a sliding-window algorithm. These modules add less than 2% to the baseline's total parameter count.
The findings show substantial improvements in gameplay fidelity across all test environments. In the Traveler game, the proposed modules increased the numerical consistency score from 0.3245 to 0.9141 and spatial consistency signal quality from 16.15 to 33.64. Text rendering guided by digit tokens aligned with target scores in 99% of cases. Furthermore, visual fidelity remained stable even when scaling generation from 64 to 256 consecutive frames, while incurring negligible computational overhead—LogicNet required only 0.0004 seconds and spatial matching 0.015 seconds per inference step.
These results demonstrate that pure end-to-end pixel generation is insufficient for interactive game simulations; separating explicit game logic and persistent spatial memory from visual synthesis resolves the primary usability bottlenecks of generative engines. The approach enables practical features such as map customization and precise player tracking without increasing deployment costs or hardware requirements.
Organizations developing generative interactive media should adopt hybrid architectures that decouple explicit logical state management from generative rendering pipelines. Moving forward, engineering teams should conduct pilot implementations to adapt this framework to complex 3D environments, evaluate dynamic multiplayer states, and increase inference throughput to achieve standard 30 frames-per-second interactive rates.
Confidence in these findings is high for controlled 2D environments, supported by quantitative benchmarks and user studies. However, key limitations remain: the sliding-window matching method fails in visually repetitive or monochrome backgrounds, and the current models occasionally generate physics anomalies without explicit physical constraints. Additional testing is required before applying this technique to complex, high-resolution 3D worlds.
- Paper: Mastering Atari with Discrete World Models, Danijar Hafner et al. (2021). Provides the foundational framework for learning discrete world models from gameplay observations to support consistent state representation and planning.
- Paper: Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs, Ling Yang et al. (2024). Introduces spatial planning and regional conditioning mechanisms for diffusion models, which informs the spatial mapping and continuity techniques used in generative games.
- Paper: Generative Agents: Interactive Simulacra of Human Behavior, Joon Sung Park et al. (2023). Establishes memory and reflection architectures for interactive agents in simulated game environments, laying the groundwork for maintaining long-term state coherence.
- Paper: Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models, Chang Liu et al. (2024). Demonstrates methods for maintaining visual and narrative context across sequential image generations in diffusion models, a core challenge in generative game scenes.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). Presents attention control mechanisms to enforce spatial layout and numerical instance constraints in diffusion generation.
- Paper: Code World Models for General Game Playing, Wolfgang Lehrach et al. (2026). Extends generative game modeling by translating natural language rules into executable Python code world models to enforce strict state transitions and legal move consistency.
- Paper: T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation, Kaiyue Sun 0001 et al. (2025). Provides a comprehensive benchmark for evaluating dynamic spatio-temporal consistency and generative numeracy in video-like dynamic environments.
