acceptodds
Under review as a conference paper at ICLR 2027

Red Alert Bench: A Diagnostic Benchmark for Multimodal Planning and Execution

Abstract

Multimodal agents in robotics and other embodied settings need to use observations to plan actions, monitor their effects, and achieve a goal. Final success rates alone do not reveal where this process breaks down. We introduce Red Alert Bench, a diagnostic benchmark and reproducible simulation environment for multimodal planning and execution, based on the real-time strategy game Red Alert. Agents gather resources, build, scout, and fight on a map where actions take effect over time, with configurable observation modality (text, image, or both) and map visibility (limited or full). Researchers can evaluate their own agents on 217 human-authored tasks, each with a verifiable goal, and in one-on-one matches where two agents control opposing sides. Evaluating eight LLMs, we find that text-only input outperforms image-primary input by 7–11 percentage points in each of the four models tested for this comparison. We also find that models repeat blocked commands, such as orders they cannot afford, without adjusting their plans. Agent-to-agent matches reveal a perspective-dependent spatial bias: on a symmetric map, the agent starting on the east side wins 68% of all games and 75% of self-play games. Red Alert Bench thus offers a dynamic world in which agents compete against one another, with execution records that link outcomes to the observations and actions behind them. We release the environment, tasks, evaluation tools, and interaction records.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.