Evaluating the World Model Implicit in a Generative Model

Keyon VafaJustin Y. ChenAshesh RambachanJon M. KleinbergSendhil Mullainathan

article2024NeurIPS98 citations

Proposes formal, Myhill-Nerode-inspired metrics to assess whether generative models truly learn underlying deterministic finite automata, revealing that high next-token accuracy often conceals fragile, incoherent internal world models across navigation, logic, and game-playing domains.

Listen

Large language models and generative sequence models demonstrate remarkable performance across diverse applications, from strategic games to route planning and scientific discovery. These achievements have led researchers to hypothesize that these models implicitly learn coherent internal representations of the underlying systems they reflect, commonly known as world models. However, existing methods for verifying this claim typically rely on superficial diagnostics, such as checking whether the model predicts a valid immediate next step or using linear probes to decode state information from internal network activations. These standard diagnostics can give a misleading impression of competence while masking critical structural deficiencies.

The article develops and demonstrates a rigorous, theoretically grounded evaluation framework to determine whether generative models actually learn accurate, coherent world models. The analysis focuses on environments governed by deterministic finite automata—systems with well-defined discrete states and transition rules, which encompass domains such as geographic navigation, board games, and formal logic. Drawing on the classic Myhill-Nerode theorem from formal language theory, the article introduces two model-agnostic evaluation criteria: sequence compression, which assesses whether two distinct paths leading to the identical state generate the same set of allowable continuations, and sequence distinction, which tests whether paths leading to different states correctly yield distinct continuations.

To evaluate this framework empirically, the authors trained transformer models on millions of New York City taxi ride trajectories representing shortest paths, traffic-perturbed paths, and random walks across Manhattan's street network. They also evaluated models trained on game transcripts of Othello and several commercial and open-source large language models—including GPT-4, Llama, and Qwen—on seating arrangement logic puzzles. These experiments assessed next-step validity, internal probe accuracy, and the proposed sequence compression and distinction metrics, accompanied by visual graph reconstructions and stress tests involving forced detours.

The findings reveal a sharp disconnect between standard task performance and underlying structural coherence. On the navigation benchmark, models trained on shortest paths achieved nearly 100% valid next-step accuracy and high probing recovery of current locations (over 90%), yet achieved a sequence compression precision of only 10% and distinction recall of only 20%. When implicit street maps were reconstructed from model outputs, they revealed severe physical incoherencies, including impossible street angles and non-existent flyovers. In detour stress tests, the performance of shortest-path models collapsed from 99% to 8% valid traversals when exposed to a 10% detour rate, dropping to 0% at higher rates. Models trained on random walks proved more robust, attaining higher distinction metrics and retaining 97% validity under moderate detours. Similarly, leading language models solved fully specified logic puzzles with up to 100% accuracy, yet none achieved compression precision above 40%, routinely declaring valid continuations impossible when presented with equivalent premise sequences.

These results demonstrate that strong generative performance on familiar sequences does not imply the presence of a robust world model. In practice, models rely on statistical heuristics that create severe operational fragility when faced with out-of-distribution conditions, minor disruptions, or altered intermediate steps. Relying on these systems for autonomous planning, critical decision-making, or scientific hypothesis generation introduces substantial safety and operational risks if structural coherence is assumed without rigorous boundary testing. The findings also highlight that training data diversity fundamentally influences structure recovery: models trained on diverse exploratory paths (such as random walks) acquire far more coherent representations than those exposed only to optimized routes.

Organizations developing or deploying generative models for planning, navigation, or reasoning tasks should immediately incorporate sequence-level boundary evaluations rather than relying solely on next-step accuracy or probe diagnostics. Where feasible, training pipelines should prioritize diverse and exploratory sequence data over exclusively optimized trajectories to improve underlying robustness. Further research should focus on extending these Myhill-Nerode evaluation frameworks beyond deterministic finite automata to non-deterministic, continuous, and unknown latent environments. Because current metrics rely on Monte Carlo approximations over bounded suffix lengths, decision-makers should treat reported compression scores as optimistic upper bounds and exercise caution when deploying generative systems in mission-critical environments requiring continuous state coherence.

arXiv: 2406.03689
Cover for Evaluating the World Model Implicit in a Generative Model

Abstract

Recent work suggests that large language models may implicitly learn world models. How should we assess this possibility? We formalize this question for the case where the underlying reality is governed by a deterministic finite automaton. This includes problems as diverse as simple logical reasoning, geographic navigation, game-playing, and chemistry. We propose new evaluation metrics for world model recovery inspired by the classic Myhill-Nerode theorem from language theory. We illustrate their utility in three domains: game playing, logic puzzles, and navigation. In all domains, the generative models we consider do well on existing diagnostics for assessing world models, but our evaluation metrics reveal their world models to be far less coherent than they appear. Such incoherence creates fragility: using a generative model to solve related but subtly different tasks can lead to failures. Building generative models that meaningfully capture the underlying logic of the domains they model would be immensely valuable; our results suggest new ways to assess how close a given model is to that goal.

Table of Contents

  • 1 Introduction
  • 2 Framework
  • 2.1 Recovering world models
  • 2.2 Next-token prediction is a fragile metric for recovering structure
  • 2.3 The Myhill-Nerode interior and boundary
  • 2.4 Compression and distinction metrics for evaluating world models
  • 3 Illustration: Do Transformers Recover the Street Map of New York City?
  • 3.1 Data and models
  • 3.2 Evaluating world models
  • 3.3 Reconstructing implicit maps
  • 3.4 Implication of failing to recover the world model: detour fragility
  • 4 Other Applications: Othello and Logic Puzzles
  • 5 Conclusion
  • Acknowledgements
  • References
  • A Proof
  • B Reconstructed Maps
  • B.1 Algorithm
  • B.2 Maps
  • C Deterministic Finite Automata and Myhill-Nerode
  • D Additional results
  • E Evaluation metric details and ablations
  • E.1 Implementation details
  • E.2 Ablations
  • F Rides data construction and training
  • G Additional maps

Knowls

  1. Knowl 1 — DFA-based definition of world-model recovery

    theoretical result

    The paper formalizes a world as a deterministic finite automaton (DFA) W=(Q,Σ,δ,q0,F)W=(Q,\Sigma,\delta,q_0,F), where QQ is a finite state set, Σ\Sigma is a finite token alphabet, δ:Q×Σ→Q\delta:Q\times\Sigma\to Q is the transition function, q0q_0 is the initial state, and FF is the set of valid states. A special rejecting state qrejectq_{\mathrm{reject}} has no outgoing transitions, and the simplifying assumption is F=Q∖{qreject}F=Q\setminus\{q_{\mathrm{reject}}\}. The extended transition function δ^\hat\delta applies δ\delta token by token. For a state qq, LW(q)L^W(q) is the set of nonempty token sequences that can be followed from qq without reaching qrejectq_{\mathrm{reject}}; S(q)S(q) is the set of prefixes that take the automaton from q0q_0 to qq.

    A generative sequence model is a conditional distribution m(⋅∣s)m(\cdot\mid s) over the next token given a prefix s∈Σ∗s\in\Sigma^*. Its accepted language from prefix ss is the set of nonempty sequences whose successive tokens receive positive conditional probability:

    Lm(s)={a1⋯ak:k≥1, m(aj+1∣sa1⋯aj)>0 for every j<k}.L^m(s)=\{a_1\cdots a_k: k\ge 1,\ m(a_{j+1}\mid sa_1\cdots a_j)>0\text{ for every }j<k\}.

    The model recovers the DFA exactly if every prefix that reaches the same valid state induces exactly the language of that state:

    ∀q∈F, ∀s∈S(q):Lm(s)=LW(q).\forall q\in F,\ \forall s\in S(q):\qquad L^m(s)=L^W(q).

    The paper proves that this language-level recovery condition is equivalent to exact next-token support matching: for every valid state qq, every prefix s∈S(q)s\in S(q), and every token a∈Σa\in\Sigma,

    m(a∣s)>0⟺δ(q,a)≠qreject.m(a\mid s)>0\quad\Longleftrightarrow\quad \delta(q,a)\ne q_{\mathrm{reject}}.

    Thus, perfect next-token support prediction is sufficient for exact recovery, but approximate next-token accuracy need not imply approximate recovery.

  2. Knowl 2 — Myhill–Nerode interiors expose the fragility of next-token tests

    definition

    For two valid DFA states q1,q2∈Fq_1,q_2\in F, the Myhill–Nerode interior is the set of suffixes accepted from both states:

    MNI⁡W(q1,q2)=LW(q1)∩LW(q2).\operatorname{MNI}^W(q_1,q_2)=L^W(q_1)\cap L^W(q_2).

    The Myhill–Nerode boundary consists of minimal suffixes accepted from q1q_1 but not q2q_2. A suffix x=a1⋯akx=a_1\cdots a_k belongs to the boundary when x∈LW(q1)∖LW(q2)x\in L^W(q_1)\setminus L^W(q_2) and every proper prefix a1⋯aja_1\cdots a_j, j<kj<k, lies in the interior:

    MNB⁡W(q1,q2)={a1⋯ak: a1⋯ak∈LW(q1)∖LW(q2), a1⋯aj∈MNI⁡W(q1,q2) ∀j<k}.\operatorname{MNB}^W(q_1,q_2)=\{a_1\cdots a_k:\ a_1\cdots a_k\in L^W(q_1)\setminus L^W(q_2),\ a_1\cdots a_j\in\operatorname{MNI}^W(q_1,q_2)\ \forall j<k\}.

    The Myhill–Nerode theorem guarantees a distinguishing suffix for every pair of distinct states in a minimal DFA, but that suffix can be much longer than one token. Consequently, a model can appear accurate under a next-token legality test while failing to represent the underlying state.

    The paper illustrates this with cumulative Connect-4 on a board with nn rows and seven columns. A state records the number of disks in each column, and a column is legal until it contains nn disks. A model that assigns probability 1/71/7 to every column regardless of the prefix encodes no board information. Nevertheless, when n=1000n=1000, it predicts a legal next move for more than 99%99\% of states because most columns remain unfilled. Distinct boards can therefore share the same legal next-token set for many steps, forming a large Myhill–Nerode interior; their differences only become visible at longer boundary suffixes.

  3. Knowl 3 — Boundary precision, recall, compression, and distinction metrics

    model/method

    The evaluation compares the true Myhill–Nerode boundary of a DFA with the boundary implied by a generative model. For prefixes s1,s2∈Σ∗s_1,s_2\in\Sigma^*, the model boundary is the set of minimal suffixes accepted after s1s_1 but not after s2s_2:

    MNB⁡m(s1,s2)={x=x1⋯xk: x∈Lm(s1)∖Lm(s2), x1⋯xj∈Lm(s1)∩Lm(s2) ∀j<k}.\operatorname{MNB}^m(s_1,s_2)=\{x=x_1\cdots x_k:\ x\in L^m(s_1)\setminus L^m(s_2),\ x_1\cdots x_j\in L^m(s_1)\cap L^m(s_2)\ \forall j<k\}.

    For prefixes s1∈S(q1)s_1\in S(q_1) and s2∈S(q2)s_2\in S(q_2), boundary recall measures how much of the true boundary the model distinguishes, while boundary precision measures how much of the model boundary is genuinely valid under the DFA:

    Recall⁡=∣MNB⁡W(q1,q2)∩(Lm(s1)∖Lm(s2))∣∣MNB⁡W(q1,q2)∣,\operatorname{Recall}=\frac{|\operatorname{MNB}^W(q_1,q_2)\cap (L^m(s_1)\setminus L^m(s_2))|}{|\operatorname{MNB}^W(q_1,q_2)|}, Precision⁡=∣MNB⁡m(s1,s2)∩(LW(q1)∖LW(q2))∣∣MNB⁡m(s1,s2)∣.\operatorname{Precision}=\frac{|\operatorname{MNB}^m(s_1,s_2)\cap (L^W(q_1)\setminus L^W(q_2))|}{|\operatorname{MNB}^m(s_1,s_2)|}.

    The sequence-compression metric samples two distinct prefixes s1,s2s_1,s_2 that reach the same state, q1=q2q_1=q_2, and reports the model-boundary precision. Because the true boundary is empty for equal states, a model receives compression precision 11 when it correctly gives both prefixes identical continuation languages. The sequence-distinction metric samples distinct states, q1≠q2q_1\ne q_2, and reports both boundary precision and boundary recall.

    In experiments, a token is treated as accepted only when its model probability exceeds an acceptance threshold ϵ\epsilon. Scores are averaged first over prefix pairs and then over sampled states or state pairs. Exact model boundaries are generally intractable, so the paper uses Monte Carlo continuations: M=30M=30 samples for map and Othello experiments, maximum true-boundary suffix length k=5k=5, and ϵ=0.01\epsilon=0.01 in the main experiments. These metrics explicitly test whether a model both compresses different histories of the same state and distinguishes histories of different states.

  4. Knowl 4 — New York navigation benchmark and transformer training setup

    experimental setup

    The navigation experiments represent Manhattan as a weighted directed graph G=(V,E,W)G=(V,E,W), where vertices are intersections, edges are streets, and W:E→R+W:E\to\mathbb{R}_+ gives street distances. Each edge receives one of eight direction labels—N, S, E, W, NE, NW, SE, or SW—and each intersection has at most one outgoing edge in each direction. The graph contains 4,580 vertices and 9,846 edges.

    The data originate from 2014 NYC taxi rides. After restricting rides to Manhattan, removing duplicates, mapping endpoints to nearby intersections, and retaining traversals of at most 100 directions, the processed data contain 3,358,737 sequences. Every sequence contains an origin, a destination, direction tokens, and an end token. Origin–destination pairs are disjoint between training and test sets.

    Three traversal distributions are used. Shortest-path data contain approximately 2.9 million sequences and 120 million tokens. Noisy-shortest-path data contain approximately 31 million sequences and 1.7 billion tokens; their edge weights are perturbed as W~(i,j)=W(i,j)+ϵij\widetilde W(i,j)=W(i,j)+\epsilon_{ij} with ϵij∼Gamma⁡(W(i,j),1)\epsilon_{ij}\sim\operatorname{Gamma}(W(i,j),1), and 50 perturbed graphs are used. Random-walk data contain approximately 91 million sequences and 4.7 billion tokens, generated by choosing a starting vertex uniformly, choosing a length uniformly from 3 through 100, and sampling outgoing edges uniformly.

    GPT-2-style transformers are trained from scratch with next-token prediction. The smaller model has 89.3 million parameters, 12 layers, hidden dimension 768, and 12 attention heads; the larger model has 1.5 billion parameters, 48 layers, hidden dimension 1,600, and 25 attention heads. The best 89.3-million-parameter model is used for shortest paths, while the best 1.5-billion-parameter models are used for noisy shortest paths and random walks.

  5. Knowl 5 — Navigation models pass conventional diagnostics but fail structural tests

    empirical result

    On unseen origin–destination pairs, the navigation transformers generate valid traversals 96–99% of the time. The shortest-path model generates the true shortest route for 97% of prompts, and the noisy-shortest-path model finds a shortest route in one of the perturbed training graphs for 94% of prompts. Conventional diagnostics also look strong: the next-token test checks whether the top prediction is a legal turn, while a linear probe from the final transformer layer predicts the current intersection.

    The proposed metrics reveal substantially weaker world-model recovery. The following are means with standard errors in parentheses; the columns are next-token legality, current-intersection probe accuracy, compression precision, distinction precision, and distinction recall:

    • Untrained transformer: 0.03 (0.00)0.03\ (0.00), 0.10 (0.00)0.10\ (0.00), 0.00 (0.00)0.00\ (0.00), 0.00 (0.00)0.00\ (0.00), 0.00 (0.00)0.00\ (0.00).
    • Shortest-path transformer: 1.00 (0.00)1.00\ (0.00), 0.91 (0.00)0.91\ (0.00), 0.10 (0.01)0.10\ (0.01), 0.35 (0.02)0.35\ (0.02), 0.20 (0.01)0.20\ (0.01).
    • Noisy-shortest-path transformer: 1.00 (0.00)1.00\ (0.00), 0.92 (0.00)0.92\ (0.00), 0.05 (0.01)0.05\ (0.01), 0.37 (0.02)0.37\ (0.02), 0.24 (0.01)0.24\ (0.01).
    • Random-walk transformer: 1.00 (0.00)1.00\ (0.00), 0.99 (0.00)0.99\ (0.00), 0.50 (0.02)0.50\ (0.02), 0.99 (0.00)0.99\ (0.00), 1.00 (0.00)1.00\ (0.00).
    • True world model: 1.001.00, not applicable, 1.001.00, 1.001.00, 1.001.00.

    Thus, models can be nearly perfect on legal next directions and current-state probes while failing to give identical continuations to prefixes that reach the same state. The random-walk model recovers distinctions well but still compresses only half of the tested same-state histories, showing that compression and distinction measure different structural properties.

  6. Knowl 6 — Graph reconstruction reveals incoherent implicit street maps

    model/method

    To visualize the world model implied by a navigation transformer, the paper reconstructs a graph whose vertices are fixed to the 4,580 real Manhattan intersections and whose coordinates are their actual latitude–longitude positions. The reconstruction enforces one outgoing edge per direction, a maximum vertex degree, and a maximum Euclidean edge length.

    For each generated sequence, reconstruction starts at its source and follows existing edges matching each direction token. When no matching edge exists, it first adds the corresponding edge from the true graph if one is available. If the model requests a direction absent from the true graph, the algorithm considers nearby candidate vertices within the maximum distance and adds the candidate that permits the longest subsequent continuation before violating the degree or reconstruction constraints. A sequence is marked unreconstructable if the degree limit prevents an edge from being added or if the traversal does not finish at its stated destination.

    Using 6,400 sampled origin–destination pairs, the true-world sequences reconstruct the Manhattan graph, whereas transformer-generated sequences produce many false edges. The map visualization on page 8 shows that the random-walk transformer's reconstructed graph contains streets with physically implausible orientations, including edges labeled northwest that point east, and apparent flyovers crossing other streets. Artificially relabeling true-world sequences at the same error rate as the transformer produces a reconstruction much closer to the real map, indicating that the transformer's errors reflect an incoherent graph rather than a small number of independent transcription mistakes. Similar incoherence appears for shortest-path and noisy-shortest-path transformers and across reconstruction constraints.

  7. Knowl 7 — Incoherent maps cause severe detour fragility

    empirical result

    The paper tests whether structural incoherence matters for downstream routing by introducing detours during greedy decoding. For each token, with probability pp, a random detour replaces the model's proposed token with a uniformly chosen legal token, while an adversarial detour replaces it with the model's lowest-ranked legal token. After every replacement, the experiment retains only cases for which a valid route to the destination shorter than 100 steps still exists.

    The fraction of valid completed traversals is reported below as pp increases. Values are means with standard errors in parentheses.

    For random detours, at p=0,0.01,0.10,0.50,0.75p=0,0.01,0.10,0.50,0.75 respectively:

    • Shortest paths: 0.99 (0.01)0.99\ (0.01), 0.69 (0.05)0.69\ (0.05), 0.08 (0.03)0.08\ (0.03), 0.00 (0.00)0.00\ (0.00), 0.00 (0.00)0.00\ (0.00).
    • Noisy shortest paths: 0.96 (0.02)0.96\ (0.02), 0.52 (0.05)0.52\ (0.05), 0.03 (0.02)0.03\ (0.02), 0.00 (0.00)0.00\ (0.00), 0.00 (0.00)0.00\ (0.00).
    • Random walks: 0.99 (0.01)0.99\ (0.01), 0.99 (0.01)0.99\ (0.01), 1.00 (0.00)1.00\ (0.00), 0.97 (0.02)0.97\ (0.02), 0.74 (0.04)0.74\ (0.04).
    • True world model: 1.001.00 at every probability.

    For adversarial detours, at p=0,0.01,0.10,0.50,0.75p=0,0.01,0.10,0.50,0.75 respectively:

    • Shortest paths: 0.99 (0.01)0.99\ (0.01), 0.66 (0.05)0.66\ (0.05), 0.06 (0.02)0.06\ (0.02), 0.00 (0.00)0.00\ (0.00), 0.00 (0.00)0.00\ (0.00).
    • Noisy shortest paths: 0.96 (0.02)0.96\ (0.02), 0.64 (0.05)0.64\ (0.05), 0.04 (0.02)0.04\ (0.02), 0.00 (0.00)0.00\ (0.00), 0.00 (0.00)0.00\ (0.00).
    • Random walks: 0.99 (0.01)0.99\ (0.01), 1.00 (0.00)1.00\ (0.00), 1.00 (0.00)1.00\ (0.00), 0.93 (0.03)0.93\ (0.03), 0.51 (0.05)0.51\ (0.05).
    • True world model: 1.001.00 at every probability.

    The random-walk model is markedly more robust, consistent with its stronger distinction scores, while models trained on shortest-path data fail rapidly when their generated route is perturbed.

  8. Knowl 8 — Othello evaluation separates synthetic-game and tournament-game models

    empirical result

    The same metrics are applied to transformers trained to predict moves in 8×8 Othello. The compression test asks whether different openings reaching the same board position induce the same continuation language; distinction tests ask whether different board positions receive different continuation languages.

    For an untrained transformer, a model trained on championship tournament games, a model trained on synthetic games, and the true Othello rules respectively, the reported compression precision, distinction precision, and distinction recall are:

    • Untrained transformer: 0.00 (0.00)0.00\ (0.00), 0.02 (0.00)0.02\ (0.00), 0.14 (0.01)0.14\ (0.01).
    • Championship-game model: 0.00 (0.00)0.00\ (0.00), 0.65 (0.01)0.65\ (0.01), 0.27 (0.01)0.27\ (0.01).
    • Synthetic-game model: 0.98 (0.00)0.98\ (0.00), 0.99 (0.00)0.99\ (0.00), 1.00 (0.00)1.00\ (0.00).
    • True world model: 1.001.00, 1.001.00, 1.001.00.

    The championship model therefore fails to compress most openings that lead to the same board and has weak distinction recall, whereas the synthetic-game model nearly recovers the true automaton. A detour test replaces a proposed move with either a random legal move or the model's lowest-ranked legal move. With random detour probabilities p=0,0.01,0.10,0.25,0.50p=0,0.01,0.10,0.25,0.50, championship-model validity is 1.00 (0.00),0.66 (0.05),0.05 (0.02),0.01 (0.01),0.01 (0.01)1.00\ (0.00),0.66\ (0.05),0.05\ (0.02),0.01\ (0.01),0.01\ (0.01), while synthetic-model validity is 1.00 (0.00),0.99 (0.01),0.97 (0.02),0.97 (0.02),0.99 (0.01)1.00\ (0.00),0.99\ (0.01),0.97\ (0.02),0.97\ (0.02),0.99\ (0.01). Under adversarial detours, the corresponding values are 1.00 (0.00),0.70 (0.05),0.01 (0.01),0.01 (0.01),0.00 (0.00)1.00\ (0.00),0.70\ (0.05),0.01\ (0.01),0.01\ (0.01),0.00\ (0.00) and 1.00 (0.00),0.98 (0.01),0.99 (0.01),0.96 (0.02),0.97 (0.02)1.00\ (0.00),0.98\ (0.01),0.99\ (0.01),0.96\ (0.02),0.97\ (0.02). The detour results validate the metrics' distinction between coherent and incoherent game models.

  9. Knowl 9 — Logic-puzzle LLMs solve specified instances without coherent state tracking

    empirical result

    The paper evaluates LLMs on a seating-arrangement puzzle with n=3n=3 individuals and three seats. A state is the set of seating arrangements consistent with the statements seen so far, and a statement is valid when it does not eliminate every arrangement in the current state. The task asks whether a proposed seating relation is possible. The models are prompted to use chain-of-thought reasoning and are evaluated with greedy decoding over 100 samples.

    The reported columns are fully specified task accuracy, compression precision, and distinction recall. Distinction precision is omitted because approximating each LLM's model boundary is too expensive.

    • Llama-2 70B: 0.77 (0.03)0.77\ (0.03), 0.08 (0.03)0.08\ (0.03), 0.42 (0.04)0.42\ (0.04).
    • Llama-3 8B: 0.85 (0.02)0.85\ (0.02), 0.18 (0.04)0.18\ (0.04), 0.23 (0.03)0.23\ (0.03).
    • Llama-3 70B: 0.98 (0.00)0.98\ (0.00), 0.25 (0.04)0.25\ (0.04), 0.57 (0.04)0.57\ (0.04).
    • Mixtral-8×22B: 0.88 (0.01)0.88\ (0.01), 0.35 (0.05)0.35\ (0.05), 0.57 (0.05)0.57\ (0.05).
    • Qwen 1.5 72B: 0.88 (0.02)0.88\ (0.02), 0.21 (0.04)0.21\ (0.04), 0.56 (0.03)0.56\ (0.03).
    • Qwen 1.5 110B: 0.98 (0.00)0.98\ (0.00), 0.53 (0.05)0.53\ (0.05), 0.53 (0.04)0.53\ (0.04).
    • GPT-3.5 turbo: 0.83 (0.02)0.83\ (0.02), 0.33 (0.05)0.33\ (0.05), 0.18 (0.03)0.18\ (0.03).
    • GPT-4: 1.00 (0.00)1.00\ (0.00), 0.21 (0.04)0.21\ (0.04), 0.56 (0.03)0.56\ (0.03).
    • True world model: 1.001.00, 1.001.00, 1.001.00.

    The results show a separation between task success and coherent state representation: even models with near-perfect puzzle accuracy have low compression and distinction scores. The example prompt illustrated on page 10 shows GPT-4 judging the same underlying seating state differently depending on which logically equivalent set of statements supplied the context.

  10. Knowl 10 — Scope, approximation dependence, and stated limitation

    limitation

    The framework assumes that the true world is known and can be represented by a finite deterministic automaton, including a rejecting sink state. This covers the paper's navigation, Othello, and seating-puzzle settings, but not unknown world models, nondeterministic processes, or domains whose state is not finite. The authors identify extending sequence compression and distinction to richer or unknown world models as the primary limitation and leave it for future work.

    The empirical metrics also require an operational definition of model acceptance. Because neural models often assign nonzero probability to nearly every sequence, the experiments use a probability threshold ϵ\epsilon, top-kk acceptance, or top-pp acceptance, and approximate boundaries with finite suffix lengths and Monte Carlo samples. These choices create a precision–recall tradeoff. For example, in the map experiments the random-walk model has compression precision 0.160.16, distinction precision 1.001.00, and distinction recall 1.001.00 at ϵ=0.10\epsilon=0.10, but compression precision 0.990.99, distinction precision 0.990.99, and distinction recall 0.110.11 at ϵ=10−6\epsilon=10^{-6}.

    The analyses nevertheless show that the qualitative conclusions are not tied to one acceptance rule: top-kk and top-pp variants also find that no map model fully recovers the true world model, with random-walk training performing best overall. Boundary length is important as well; the random-walk model achieves 1.001.00 compression precision when only length-one boundaries are considered but approximately 0.500.50 when the full boundary is considered. This demonstrates both the usefulness of the metrics and the limitation that finite approximations can obscure longer-range incoherence.

Coverage note — Extended appendix map galleries, secondary reconstruction-parameter grids, proof derivations, and detailed Monte Carlo-sample ablations were omitted because they support the main findings without adding separate load-bearing contributions.

References

  1. 1.Abdou, M., Kulmizev, A., Hershcovich, D., Frank, S., Pavlick, E., and Søgaard, A. Can language models encode perceptual structure without grounding? A case study in color. arXiv preprint arXiv:2109.06129, 2021.
  2. 2.Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  3. 3.Benegas, G., Batra, S. S., and Song, Y. S. DNA language models are powerful predictors of genome-wide variant effects. Proceedings of the National Academy of Sciences, 120(44):e2311219120, 2023.
  4. 4.Bhattamishra, S., Ahuja, K., and Goyal, N. On the ability and limitations of transformers to recognize formal languages. arXiv preprint arXiv:2009.11264, 2020.
  5. 5.Boeing, G. Modeling and analyzing urban networks and amenities with OSMnx. 2024.
  6. 6.Boiko, D. A., MacKnight, R., Kline, B., and Gomes, G. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023.
  7. 7.Chowdhury, R., Bouatta, N., Biswas, S., Floristean, C., Kharkar, A., Roy, K., Rochereau, C., Ahdritz, G., Zhang, J., Church, G. M., Sorger, P. K., and AlQuraishi, M. Single-sequence protein structure prediction using a language model and deep learning. Nature Biotechnology, 40(11):1617–1623, 2022.
  8. 8.Fan, A., Lewis, M., and Dauphin, Y. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833, 2018.
  9. 9.Guan, L., Valmeekam, K., Sreedharan, S., and Kambhampati, S. Leveraging pre-trained large language models to construct and utilize world models for model-based task planning. Advances in Neural Information Processing Systems, 36:79081–79094, 2023.
  10. 10.Hazineh, D. S., Zhang, Z., and Chiu, J. Linear latent world models in simple transformers: A case study on Othello-GPT. arXiv preprint arXiv:2310.07582, 2023.
  11. 11.Hewitt, J. and Liang, P. Designing and interpreting probes with control tasks. arXiv preprint arXiv:1909.03368, 2019.
  12. 12.Hewitt, J., Manning, C. D., and Liang, P. Truncation sampling as language model desmoothing. arXiv preprint arXiv:2210.15191, 2022.
  13. 13.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019.
  14. 14.Jablonka, K. M., Schwaller, P., Ortega-Guerrero, A., and Smit, B. Leveraging large language models for predictive chemistry. Nature Machine Intelligence, pp. 1–9, 2024.
  15. 15.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  16. 16.Jin, C. and Rinard, M. Evidence of meaning in language models trained on programs. arXiv preprint arXiv:2305.11169, 2023.
  17. 17.Kıcıman, E., Ness, R., Sharma, A., and Tan, C. Causal reasoning and large language models: Opening a new frontier for causality. arXiv preprint arXiv:2305.00050, 2023.
  18. 18.Kuo, M.-T., Hsueh, C.-C., and Tsai, R. T.-H. Large language models on the chessboard: A study on ChatGPT’s formal language comprehension and complex reasoning skills. arXiv preprint arXiv:2308.15118, 2023.
  19. 19.Li, B. Z., Nye, M., and Andreas, J. Implicit representations of meaning in neural language models. arXiv preprint arXiv:2106.00737, 2021.
  20. 20.Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M. Emergent world representations: Exploring a sequence model trained on a synthetic task. In International Conference on Learning Representations, 2023.
  21. 21.Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S., and Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637):1123–1130, 2023.
  22. 22.Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C. Transformers learn shortcuts to automata. arXiv preprint arXiv:2210.10749, 2022.
  23. 23.Merrill, W. and Sabharwal, A. The parallelism tradeoff: Limitations of log-precision transformers. Transactions of the Association for Computational Linguistics, 11:531–545, 2023.
  24. 24.Merrill, W., Petty, J., and Sabharwal, A. The illusion of state in state-space models. arXiv preprint arXiv:2404.08819, 2024.
  25. 25.Murray, K. W. 2014 New York City taxi trips. https://www.kaggle.com/datasets/kentonnlp/2014-new-york-city-taxi-trips, 2017. Accessed: 2024-10-24.
  26. 26.Myhill, J. Finite automata and the representation of events. WADD Technical Report, 57:112–137, 1957.
  27. 27.Nerode, A. Linear automaton transformations. Proceedings of the American Mathematical Society, 9(4):541–544, 1958.
  28. 28.Patel, R. and Pavlick, E. Mapping language models to grounded conceptual spaces. In International Conference on Learning Representations, 2021.
  29. 29.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  30. 30.Schumann, R. and Riezler, S. Generating landmark navigation instructions from maps as a graph-to-text problem. Association for Computational Linguistics, 2021.
  31. 31.Schumann, R. and Riezler, S. Analyzing generalization of vision and language navigation to unseen outdoor areas. Association for Computational Linguistics, 2022.
  32. 32.Schumann, R., Zhu, W., Feng, W., Fu, T.-J., Riezler, S., and Wang, W. Y. VELMA: Verbalization embodiment of LLM agents for vision and language navigation in street view. In AAAI Conference on Artificial Intelligence, 2024.
  33. 33.Sipser, M. Introduction to the Theory of Computation, Third Edition. Cengage Learning, 2013.
  34. 34.Suzgun, M., Belinkov, Y., and Shieber, S. M. On evaluating the generalization of LSTM models in formal languages. arXiv preprint arXiv:1811.01001, 2018.
  35. 35.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. Challenging BIG-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  36. 36.Toshniwal, S., Wiseman, S., Livescu, K., and Gimpel, K. Chess as a testbed for language model state tracking. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 11385–11393, 2022.
  37. 37.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  38. 38.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Neural Information Processing Systems, 2017.
  39. 39.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682, 2022.
  40. 40.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Neural Information Processing Systems, 35:24824–24837, 2022.

Citation

MLA
Vafa, K., et al. “Evaluating the World Model Implicit in a Generative Model”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 26941–75, https://proceedings.neurips.cc/paper_files/paper/2024/file/2f6a6317bada76b26a4f61bb70a7db59-Paper-Conference.pdf.
APA
Vafa, K., Chen, J. Y., Rambachan, A., Kleinberg, J., & Mullainathan, S. (2024). Evaluating the World Model Implicit in a Generative Model. Advances in Neural Information Processing Systems, 37, 26941–26975. https://proceedings.neurips.cc/paper_files/paper/2024/file/2f6a6317bada76b26a4f61bb70a7db59-Paper-Conference.pdf
Chicago
Vafa, K., J. Y. Chen, A. Rambachan, J. Kleinberg, and S. Mullainathan. 2024. “Evaluating the World Model Implicit in a Generative Model”. Advances in Neural Information Processing Systems 37: 26941–75. https://proceedings.neurips.cc/paper_files/paper/2024/file/2f6a6317bada76b26a4f61bb70a7db59-Paper-Conference.pdf.
Harvard
Vafa, K. et al. (2024) “Evaluating the World Model Implicit in a Generative Model”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 26941–26975. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/2f6a6317bada76b26a4f61bb70a7db59-Paper-Conference.pdf.
Vancouver
1. Vafa K, Chen JY, Rambachan A, Kleinberg J, Mullainathan S (2024) Evaluating the World Model Implicit in a Generative Model. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 26941–26975

BibTeX

@inproceedings{vafa2024evaluating,
  title = {Evaluating the World Model Implicit in a Generative Model},
  author = {Vafa, Keyon and Chen, Justin Y. and Rambachan, Ashesh and Kleinberg, Jon and Mullainathan, Sendhil},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {26941-26975},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/2f6a6317bada76b26a4f61bb70a7db59-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors