Evaluating the World Model Implicit in a Generative Model
Keyon VafaJustin Y. ChenAshesh RambachanJon M. KleinbergSendhil Mullainathan
Proposes formal, Myhill-Nerode-inspired metrics to assess whether generative models truly learn underlying deterministic finite automata, revealing that high next-token accuracy often conceals fragile, incoherent internal world models across navigation, logic, and game-playing domains.
Large language models and generative sequence models demonstrate remarkable performance across diverse applications, from strategic games to route planning and scientific discovery. These achievements have led researchers to hypothesize that these models implicitly learn coherent internal representations of the underlying systems they reflect, commonly known as world models. However, existing methods for verifying this claim typically rely on superficial diagnostics, such as checking whether the model predicts a valid immediate next step or using linear probes to decode state information from internal network activations. These standard diagnostics can give a misleading impression of competence while masking critical structural deficiencies.
The article develops and demonstrates a rigorous, theoretically grounded evaluation framework to determine whether generative models actually learn accurate, coherent world models. The analysis focuses on environments governed by deterministic finite automata—systems with well-defined discrete states and transition rules, which encompass domains such as geographic navigation, board games, and formal logic. Drawing on the classic Myhill-Nerode theorem from formal language theory, the article introduces two model-agnostic evaluation criteria: sequence compression, which assesses whether two distinct paths leading to the identical state generate the same set of allowable continuations, and sequence distinction, which tests whether paths leading to different states correctly yield distinct continuations.
To evaluate this framework empirically, the authors trained transformer models on millions of New York City taxi ride trajectories representing shortest paths, traffic-perturbed paths, and random walks across Manhattan's street network. They also evaluated models trained on game transcripts of Othello and several commercial and open-source large language models—including GPT-4, Llama, and Qwen—on seating arrangement logic puzzles. These experiments assessed next-step validity, internal probe accuracy, and the proposed sequence compression and distinction metrics, accompanied by visual graph reconstructions and stress tests involving forced detours.
The findings reveal a sharp disconnect between standard task performance and underlying structural coherence. On the navigation benchmark, models trained on shortest paths achieved nearly 100% valid next-step accuracy and high probing recovery of current locations (over 90%), yet achieved a sequence compression precision of only 10% and distinction recall of only 20%. When implicit street maps were reconstructed from model outputs, they revealed severe physical incoherencies, including impossible street angles and non-existent flyovers. In detour stress tests, the performance of shortest-path models collapsed from 99% to 8% valid traversals when exposed to a 10% detour rate, dropping to 0% at higher rates. Models trained on random walks proved more robust, attaining higher distinction metrics and retaining 97% validity under moderate detours. Similarly, leading language models solved fully specified logic puzzles with up to 100% accuracy, yet none achieved compression precision above 40%, routinely declaring valid continuations impossible when presented with equivalent premise sequences.
These results demonstrate that strong generative performance on familiar sequences does not imply the presence of a robust world model. In practice, models rely on statistical heuristics that create severe operational fragility when faced with out-of-distribution conditions, minor disruptions, or altered intermediate steps. Relying on these systems for autonomous planning, critical decision-making, or scientific hypothesis generation introduces substantial safety and operational risks if structural coherence is assumed without rigorous boundary testing. The findings also highlight that training data diversity fundamentally influences structure recovery: models trained on diverse exploratory paths (such as random walks) acquire far more coherent representations than those exposed only to optimized routes.
Organizations developing or deploying generative models for planning, navigation, or reasoning tasks should immediately incorporate sequence-level boundary evaluations rather than relying solely on next-step accuracy or probe diagnostics. Where feasible, training pipelines should prioritize diverse and exploratory sequence data over exclusively optimized trajectories to improve underlying robustness. Further research should focus on extending these Myhill-Nerode evaluation frameworks beyond deterministic finite automata to non-deterministic, continuous, and unknown latent environments. Because current metrics rely on Monte Carlo approximations over bounded suffix lengths, decision-makers should treat reported compression scores as optimistic upper bounds and exercise caution when deploying generative systems in mission-critical environments requiring continuous state coherence.
- Paper: Recurrent World Models Facilitate Policy Evolution, David Ha et al. (2018). This foundational work establishes the concept of learning implicit, predictive world models for planning and control in generative neural architectures.
- Paper: Future Lens: Anticipating Subsequent Tokens from a Single Hidden State, Koyena Pal et al. (2023). This paper examines how single internal hidden states encode multi-step future information, providing foundational context for probing representations and testing forward state transitions in transformers.
- Paper: On the Origins of Linear Representations in Large Language Models, Yibo Jiang et al. (2024). This study analyzes why linear representations and internal state geometry emerge in language models, directly informing the evaluation and probing of implicit state spaces.
- Paper: Exploring Length Generalization in Large Language Models, Cem Anil et al. (2022). This work demonstrates how autoregressive transformers rely on shortcuts and fail at out-of-distribution sequential algorithmic generalization, motivating the need for formal structure-preservation evaluations.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). This paper shows that language models often achieve high local validity while failing global sequence planning on formal logical tasks, underscoring the structural deficits targeted by the source.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). This paper investigates the capacity of autoregressive language models to generate executable multi-step plans in structured environments, serving as a key predecessor to formal world-model verification.
- Paper: Next-Latent Prediction Transformers Learn Compact World Models, Jayden Teoh et al. (2025). This work directly builds on the source's findings and Manhattan navigation benchmark to introduce next-latent prediction objectives that induce compact, coherent world models.
- Paper: WorldSimBench: Towards Video Generation Models as World Simulators, Yiran Qin et al. (2025). This benchmark extends the evaluation of generative models as implicit world simulators from discrete automata to embodied physical and visual simulation environments.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). This paper develops foundation models that learn explicit predictive state transitions and planning from video trajectories, directly addressing the representational limitations identified in generative planners.
- Paper: LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels, Lucas Maes et al. (2026). This study proposes joint-embedding predictive architectures trained to learn non-collapsing, stable latent world models for physical navigation and planning.
- Paper: The Topological Trouble With Transformers, Michael C. Mozer et al. (2026). This work provides a theoretical and architectural analysis of why standard transformer architectures suffer topological bottlenecks in tracking evolving internal states across sequential tasks.
- Paper: Hierarchical Reasoning Model, Guan Wang et al. (2025). This architecture replaces standard next-token sequence models with hierarchical recurrent modules to prevent the structural planning collapse seen in long-horizon navigation and maze tasks.
