SC2Forge: A Calibrated Benchmark and a Program-Writing Agent for Full-Game StarCraft II
Abstract
A language model can play a real-time strategy game from inside the control loop or write the program that plays it. The first needs a model call per decision and, on our benchmark, at least 5.5 wall-clock seconds per game second; the second delivers a compiled package that plays a typical game faster than real time, with no model call. We contribute SC2FullGame-Bench, which scores a StarCraft II bot package on 93 full-game cells, 90 against the built-in opponents (three maps, ten levels and three races) and 3 against a community bot. It also rates each package with four probe games (economy, tech, combat and defense). Before any submission is scored, the benchmark is calibrated: two ladders of known order and a package's agreement with itself on a repeated seed show that it separates agents of different strength. We also contribute SC2Forge, a framework that turns a minimal demonstration package, the runtime's documentation and a playbook of named openings into a compiled package through a write-test-revise loop with a verifier ladder, a tool-using rewrite arm and a capability-keyed archive. SC2Forge's package scores 0.943, its best of three runs, at the level of VibeCraft's expert-system Protoss package (0.932), the bot its playbook was written from. The package is separated above an iterated prompt (0.767) and an off-the-shelf coding agent (0.803); a coding agent given the playbook under the framework's own model budget reaches 0.900, within 0.05 per cell of the framework's package on sealed suite A's six seeds and separated below it on the released suite's six seeds. The framework also generates Terran and Zerg packages, each scoring above VibeCraft's package for that race; on connect-four the loop, its arms adapted to the game, delivers a policy scoring 0.977 against a scripted 0.611.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.