VLNAPO: Gradient-Guided Harness Engineering with Key-Token Supervision for Full-Game VLNA Policy Optimization in OpenRA
Abstract
Full-game control in real-time strategy (RTS) games requires agents to transform partial, multimodal observations into long-horizon strategic decisions while satisfying rapidly changing game-state constraints. The central challenge is that strategic actions are combinatorial: an action that is desirable in isolation may be infeasible because of resource limits, technology prerequisites, unit availability or production-queue conflicts. Thus, an agent must jointly determine what should be done from visual, linguistic and numerical context and generate command sequences that are both strategically effective and executable. We propose VLNA, a structured multimodal policy for full-game control in OpenRA. VLNA fuses tactical renderings, language context, and numerical game states to autoregressively generate structured macro-actions, which a deterministic compiler translates into native commands. A cascaded conditional action space uses prefix-conditioned masks to restrict each action while dynamically updating resource and production-queue budgets. Through training gradients, Harness Engineering combines rule-generated local hard labels with the PPO objective, embedding infrastructure and production rules without replacing policy actions during interaction. Across five seeds with 100 evaluation games each on the singles map, VLNA achieved a win rate of 77%±2.00%, rising to 90%±1.55% with persistent Harness supervision. The cascaded action space achieved a joint validity rate of 99.96%. These findings show that structured multimodal reasoning, prefix-conditioned action constraints, and gradient-guided Harness Engineering can jointly improve executability and strategic performance in full-game RTS agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.