ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Xinye LiLingshuai LinLei WangLiuzhou ZhangJialin CuiQingshang LiGuanchu WangQingbin LiuXi ChenJiang Bian
Presents a progressive causal distillation framework that transforms bidirectional video generators into efficient one-to-four-step action-conditioned world models with precise keyboard and mouse control for low-latency interactive gameplay.
Generative video models can simulate dynamic interactive environments without traditional graphics engines, but deploying them interactively requires low latency, strict frame-by-frame causality, and precise responsiveness to user inputs. Standard video generation techniques rely on processing entire video sequences simultaneously and using many iterative refinement steps, which causes intolerable delays for real-time applications. Conversely, aggressively reducing the number of computation steps often causes visual quality to degrade rapidly, destabilizes long-term generation, and leads models to ignore complex user inputs such as continuous mouse movements and discrete keystrokes.
The article introduces and evaluates ForgeWM, a progressive four-stage training framework designed to convert full-sequence video generators into highly efficient, interactive causal world models. The primary objective is to demonstrate that a unified distillation and training pipeline can produce specialized models capable of running in one to four processing steps while maintaining high visual fidelity, temporal stability, and strict alignment with native game controls.
The approach uses a structured training progression: initial domain adaptation on full video sequences, teacher-forced causal training, online causal consistency distillation to enable few-step generation, and on-policy distribution matching to bridge the gap between clean training data and imperfect self-generated histories. To test this approach, the authors trained models on 40,000 Minecraft video clips and evaluated them across 1,000 paired control trajectories against leading baseline interactive world models. The framework was assessed across visual quality benchmarks, perceptual difference metrics, directional control accuracy tests, and a blind 41-participant human preference study. The training recipe was also tested on a cross-game dataset spanning seven gamepad-controlled first-person shooter titles.
The evaluation produced four key findings. First, ForgeWM models outperformed baselines across major quality and control benchmarks on Minecraft, with the one-step student achieving 72.1 frames per second (a latency of 168.2 milliseconds per chunk) compared to 32.4 frames per second for Matrix-Game 2.0 and 7.5 frames per second for HY-WorldPlay. Second, human evaluators strongly preferred ForgeWM-4 over baselines, awarding it 60.7% of total preference votes across visual quality, action accuracy, and spatiotemporal consistency. Third, a dual-path replay protocol allowed a saved one-step interactive draft to be refined offline with four processing steps using the exact same model weights; this matched the reference quality of direct four-step generation while staying roughly three times closer to the user's experienced trajectory. Fourth, the four-stage recipe transferred successfully to a four-channel gamepad-driven shooter domain across seven distinct game titles without structural backbone modifications.
These findings demonstrate that organizations developing interactive digital environments, training simulations, or physical AI systems can achieve real-time streaming performance without needing dedicated offline rendering clusters or compromising control responsiveness. Furthermore, the dual-path replay mechanism shows that low-latency live interaction and high-fidelity archival video can be supported by a single, lightweight model checkpoint, reducing infrastructure and model maintenance costs.
For engineering and product teams building real-time simulators, the article supports adopting ForgeWM's four-stage distillation recipe and modular action conditioning. Teams should deploy the one-step or two-step student for live user control and apply the post-interaction replay-refinement path when high-fidelity recordings are required. Before deploying in open-ended or safety-critical settings, practitioners should conduct pilot testing to address three identified limitations: visual drift and color artifacts during extended rollouts beyond 20 seconds, a tendency to produce stronger camera motion than requested in shooter environments (averaging 45% higher motion magnitude than reference footage), and domain generalization constraints outside the evaluated game environments.
- Paper: Progressive Distillation for Fast Sampling of Diffusion Models, Tim Salimans et al. (2022). Progressive distillation introduces the teacher-to-student step-halving strategy that ForgeWM adapts to compress video generation into few-step sampling.
- Paper: One-Step Diffusion Distillation through Score Implicit Matching, Weijian Luo et al. (2024). Its one-step score-matching distillation provides useful grounding for ForgeWM’s goal of preserving video quality at extremely low denoising budgets.
- Paper: Genie: Generative Interactive Environments, Jake Bruce et al. (2024). Genie establishes the action-conditioned generative-environment setting that ForgeWM targets with more temporally compressed, low-latency control.
- Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). Flexible Diffusion Model’s long-video and Minecraft generation gives context for the temporal-coherence challenges underlying ForgeWM’s causal rollout training.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). Diffusion Forcing’s framewise noise levels and history conditioning clarify the causal video-diffusion techniques relevant to ForgeWM’s training and rollout.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). This foundational video-diffusion work explains the iterative generation framework that ForgeWM adapts for action-conditioned, few-step video synthesis.
No sufficiently relevant recommendations were found.
