Red Alert Bench: A Diagnostic Benchmark for Multimodal Planning and Execution
Abstract
Multimodal agents in robotics and other embodied settings need to use observations to plan actions, monitor their effects, and achieve a goal. Final success rates alone do not reveal where this process breaks down. We introduce Red Alert Bench, a diagnostic benchmark and reproducible simulation environment for multimodal planning and execution, based on the real-time strategy game Red Alert. Agents gather resources, build, scout, and fight on a map where actions take effect over time, with configurable observation modality (text, image, or both) and map visibility (limited or full). Researchers can evaluate their own agents on 217 human-authored tasks, each with a verifiable goal, and in one-on-one matches where two agents control opposing sides. Evaluating eight LLMs, we find that text-only input outperforms image-primary input by 7–11 percentage points in each of the four models tested for this comparison. We also find that models repeat blocked commands, such as orders they cannot afford, without adjusting their plans. Agent-to-agent matches reveal a perspective-dependent spatial bias: on a symmetric map, the agent starting on the east side wins 68% of all games and 75% of self-play games. Red Alert Bench thus offers a dynamic world in which agents compete against one another, with execution records that link outcomes to the observations and actions behind them. We release the environment, tasks, evaluation tools, and interaction records.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.