Self-Taught Policy Improvement for Vision-Language Game Agents
Abstract
Vision-language models (VLMs) have achieved strong performance on visual question answering (QA), yet they are ineffective at playing games directly from raw pixels. This gap stems from a fundamental mismatch: VLM pre-training is based on image-grounded QA, whereas game playing requires goal-directed action generation, which is far outside the VLM pre-training distribution. Compared with raw action generation, in-game reasoning about the visual scene is structurally closer to image QA. Based on this insight, we propose a self-taught policy improvement framework that uses reasoning tokens, rather than action tokens, as the primary alignment target for VLM post-training in games, and requires no human annotation. Concretely, we first generate high-quality, state-grounded traces by rolling out the initial policy with privileged environment state information injected into the prompt, then distill these traces back into a pixel-only policy; iterating this rollout-and-distillation loop progressively improves reasoning and decision-making. Our approach raises Qwen3-VL-8B-Instruct's win rate from 0.0% to 31.3% on ProcGen-bigfish and from 13.8% to 58.1% on MiniHack, while more than doubling average Crafter achievements from 3.3 to 7.5. It also outperforms Gemini-3.5-Flash, SFT, GRPO, and GiGPO in task reward across all three games. Our approach outperforms all trained baselines in task reward with substantially fewer training samples.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.