ALERT: Agents for Language-Guided Execution of Revocable Tasks in Real-Time Strategy Games
Abstract
An AI that can win a game while ignoring human instructions is a frightening AI. Real-time strategy benchmarks measure how well agents play, but winning does not show whether an agent followed a human instruction. We introduce ALERT, a benchmark for executing human instructions that can be revised or revoked during full *Command & Conquer: Red Alert 2* matches. The engine continues to advance during model inference, so instructions must remain in force while the agent reacts to a changing game. A reference harness connects language understanding and planning (System 2) to fast execution (System 1), recording active instructions, controller ownership, and actions. We evaluate requested world-state effects separately from match outcomes and use these records to distinguish interpretation, timing, and control failures, rather than relying only on wins and losses. We also use intent and value models to monitor game replays. Testing our early real-time agent, ALERT-Executor, reveals a failure that action logs alone would miss. In seven stored matches, all seven retreat instructions produced engine-accepted movement orders, but controller ownership ended before confirmed group arrival in six. Issuing an accepted retreat order, therefore, did not mean the agent carried out the retreat. ALERT makes such failures visible and provides a way to study instruction following as System 1 and System 2 agents improve.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.