ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Xinye LiLingshuai LinLei WangLiuzhou ZhangJialin CuiQingshang LiGuanchu WangQingbin LiuXi ChenJiang Bian

article2026arXiv1 citations

Presents a progressive causal distillation framework that transforms bidirectional video generators into efficient one-to-four-step action-conditioned world models with precise keyboard and mouse control for low-latency interactive gameplay.

Listen

Generative video models can simulate dynamic interactive environments without traditional graphics engines, but deploying them interactively requires low latency, strict frame-by-frame causality, and precise responsiveness to user inputs. Standard video generation techniques rely on processing entire video sequences simultaneously and using many iterative refinement steps, which causes intolerable delays for real-time applications. Conversely, aggressively reducing the number of computation steps often causes visual quality to degrade rapidly, destabilizes long-term generation, and leads models to ignore complex user inputs such as continuous mouse movements and discrete keystrokes.

The article introduces and evaluates ForgeWM, a progressive four-stage training framework designed to convert full-sequence video generators into highly efficient, interactive causal world models. The primary objective is to demonstrate that a unified distillation and training pipeline can produce specialized models capable of running in one to four processing steps while maintaining high visual fidelity, temporal stability, and strict alignment with native game controls.

The approach uses a structured training progression: initial domain adaptation on full video sequences, teacher-forced causal training, online causal consistency distillation to enable few-step generation, and on-policy distribution matching to bridge the gap between clean training data and imperfect self-generated histories. To test this approach, the authors trained models on 40,000 Minecraft video clips and evaluated them across 1,000 paired control trajectories against leading baseline interactive world models. The framework was assessed across visual quality benchmarks, perceptual difference metrics, directional control accuracy tests, and a blind 41-participant human preference study. The training recipe was also tested on a cross-game dataset spanning seven gamepad-controlled first-person shooter titles.

The evaluation produced four key findings. First, ForgeWM models outperformed baselines across major quality and control benchmarks on Minecraft, with the one-step student achieving 72.1 frames per second (a latency of 168.2 milliseconds per chunk) compared to 32.4 frames per second for Matrix-Game 2.0 and 7.5 frames per second for HY-WorldPlay. Second, human evaluators strongly preferred ForgeWM-4 over baselines, awarding it 60.7% of total preference votes across visual quality, action accuracy, and spatiotemporal consistency. Third, a dual-path replay protocol allowed a saved one-step interactive draft to be refined offline with four processing steps using the exact same model weights; this matched the reference quality of direct four-step generation while staying roughly three times closer to the user's experienced trajectory. Fourth, the four-stage recipe transferred successfully to a four-channel gamepad-driven shooter domain across seven distinct game titles without structural backbone modifications.

These findings demonstrate that organizations developing interactive digital environments, training simulations, or physical AI systems can achieve real-time streaming performance without needing dedicated offline rendering clusters or compromising control responsiveness. Furthermore, the dual-path replay mechanism shows that low-latency live interaction and high-fidelity archival video can be supported by a single, lightweight model checkpoint, reducing infrastructure and model maintenance costs.

For engineering and product teams building real-time simulators, the article supports adopting ForgeWM's four-stage distillation recipe and modular action conditioning. Teams should deploy the one-step or two-step student for live user control and apply the post-interaction replay-refinement path when high-fidelity recordings are required. Before deploying in open-ended or safety-critical settings, practitioners should conduct pilot testing to address three identified limitations: visual drift and color artifacts during extended rollouts beyond 20 seconds, a tendency to produce stronger camera motion than requested in shooter environments (averaging 45% higher motion magnitude than reference footage), and domain generalization constraints outside the evaluated game environments.

  • Paper: Progressive Distillation for Fast Sampling of Diffusion Models, Tim Salimans et al. (2022). Progressive distillation introduces the teacher-to-student step-halving strategy that ForgeWM adapts to compress video generation into few-step sampling.
  • Paper: One-Step Diffusion Distillation through Score Implicit Matching, Weijian Luo et al. (2024). Its one-step score-matching distillation provides useful grounding for ForgeWM’s goal of preserving video quality at extremely low denoising budgets.
  • Paper: Genie: Generative Interactive Environments, Jake Bruce et al. (2024). Genie establishes the action-conditioned generative-environment setting that ForgeWM targets with more temporally compressed, low-latency control.
  • Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). Flexible Diffusion Model’s long-video and Minecraft generation gives context for the temporal-coherence challenges underlying ForgeWM’s causal rollout training.
  • Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). Diffusion Forcing’s framewise noise levels and history conditioning clarify the causal video-diffusion techniques relevant to ForgeWM’s training and rollout.
  • Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This foundational video-diffusion work explains the iterative generation framework that ForgeWM adapts for action-conditioned, few-step video synthesis.

No sufficiently relevant recommendations were found.

Cover for ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Abstract

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.

Citation

MLA
Li, X., et al. “ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models”. arXiv, 2026, http://arxiv.org/abs/2608.14022v1.
APA
Li, X., Lin, L., Wang, L., Zhang, L., Cui, J., Li, Q., Wang, G., Liu, Q., Chen, X., Bian, J., & Lam, W. (2026). ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models. arXiv. http://arxiv.org/abs/2608.14022v1
Chicago
Li, X., L. Lin, L. Wang, et al. 2026. “ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models”. arXiv. http://arxiv.org/abs/2608.14022v1.
Harvard
Li, X. et al. (2026) “ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2608.14022v1.
Vancouver
1. Li X, Lin L, Wang L, et al (2026) ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models. arXiv

BibTeX

@article{li2026forgewm,
  title = {ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models},
  author = {Li, Xinye and Lin, Lingshuai and Wang, Lei and Zhang, Liuzhou and Cui, Jialin and Li, Qingshan and Wang, Guanchu and Liu, Qingbin and Chen, Xi and Bian, Jiang and Lam, Wai},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2608.14022v1},
  eprint = {2608.14022}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/