SC2Arena: Evaluating Large Language Models in Full Games of StarCraft II
Abstract
Evaluating large language models (LLMs) as agents in competitive environments requires examining how they pursue objectives while responding to opponents whose actions can disrupt their progress. StarCraft II provides a challenging setting where agents must coordinate resource management, production, and unit control against an opponent over a complete match. We introduce SC2Arena, a benchmark for evaluating LLMs in full games of StarCraft II through a standardized text interface. SC2Arena supports all three playable races, fine-grained actions, and matches against both built-in AI and other agents. It combines game outcomes with process metrics to assess action validity, resource management, and token consumption, and incorporates unit ordering and aggregation to structure textual observations. We also develop a hierarchical baseline that separates command generation from action execution, incorporates iterative verification, and supports supervised fine-tuning on selected gameplay data. Experiments reveal substantial performance differences across evaluated models. Within the tested settings, the hierarchical baseline with verification improves win rates and action validity, while fine-tuning yields further gains. These results highlight the importance of considering agent design alongside model choice when evaluating LLMs in competitive environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.