acceptodds
Under review as a conference paper at ICLR 2027

SHUGYO: On-Policy Reinforcement Learning of Game Agents in a Generative World Model

Abstract

Game agents trained by supervised imitation on video–action data drift into states that demonstrations rarely cover, where small errors compound into task failure. Reinforcement learning (RL) can provide the missing experience of acting, failing, and recovering, but closed-source commercial games lack the fast dynamics, programmatic resets, and reward signals that RL requires. We present SHUGYO, a framework that adapts a pretrained video diffusion model into a learned environment for on-policy RL in the AAA action game Black Myth: Wukong. The simulator predicts future frames from rendered observations and native keyboard-and-mouse controls, resets rollouts from specified starting frames, and generates each chunk in four denoising steps. A rubric-guided vision–language judge scores the resulting rollouts, providing rewards without access to internal game state. Starting from a supervised vision–language agent, we perform RL post-training with GRPO and a two-stage curriculum entirely within the simulator, using a load-balanced client–server architecture for scalable rollout generation. Experiments show that the simulator follows native controls, remains visually coherent over extended rollouts, and adapts to a second game with only three hours of data. Per-task success rates in the simulator correlate with those in the real game, and RL post-training in the simulator raises real-game success from 58.6% to 68.6%. These results suggest that action-conditioned video world models can serve as practical training environments for game agents when the real game is not RL-ready.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.