Learning Agentic Vision-Language Game Agents from Trajectory-Derived Supervision
Abstract
Vision-language game agents are typically trained to predict the next action from gameplay demonstrations in a supervised fine-tuning (SFT) manner. While effective for learning short-horizon behavior, this training view underuses the process structure already present in trajectories, including how plans persist across multiple steps, how actions are carried out under a plan, and how execution outcomes signal when a new plan is needed. Gameplay trajectories contain richer supervision than action labels alone. When organized into plan-conditioned segments, they naturally specify how an agent should form a plan, act under that plan over time, and track whether execution is still aligned with it. In this paper, we propose a trajectory-to-process supervision framework for learning plan-act-verify vision-language game agents from gameplay trajectories. Unlike conventional CoT-before-action training, where plans mainly serve as intermediate prefixes for action prediction, our formulation explicitly trains the model to carry out a plan across multiple steps and to generate the next plan when the current one succeeds, fails, or is no longer valid. The resulting agent moves beyond direct acting toward closed-loop gameplay, using trajectory-derived process annotations as supervision for agentic behavior. Across Tetris, Sokoban, and MiniGrid, our method consistently improves over direct action prediction, step-wise plan-as-CoT, and open-loop plan-chunk baselines, and transfers across Qwen3-VL-8B, InternVL3-8B, and Gemma3-12B backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.