acceptodds
Under review as a conference paper at ICLR 2027

RTSGameBench: Evaluating VLMs For Strategic Reasoning in Game Environments

Abstract

Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., deciding under uncertainty while coordinating with and adapting to other agents. Real-time strategy (RTS) games can be a natural testbed for this capability, demanding long-horizon planning, coordination with allies, and adaptation to opponents under partial observability. But existing RTS benchmarks cover limited agent configurations, lack controlled evaluation across distinct strategic demands, and remain confined to fixed scenarios. To address these gaps, we present RTSGAMEBENCH, built on Beyond All Reason, an RTS game whose larger scale broadens the strategic space beyond prior testbeds. In RTSGameBench, we offer holistic evaluation across diverse full-game match-ups, diagnostic mini-games each emphasizing a focal strategic demand, and extensible coverage via a selfevolving framework that turns free-form queries into new mini-games, improving with accumulated experience. We further introduce RTSGameAgen, a baseline agent harness that combines scalable group-level trajectory planning with short- and long-term memory and outperforms existing agents. Using it, we find that state-of-the-art VLMs struggle where humans succeed, particularly when they must coordinate with allies or contend with multiple opponents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.