acceptodds
Under review as a conference paper at ICLR 2027

ALAMEINBENCH: EVALUATING LONG-HORIZON STRATEGIC PLANNING OF LLM AGENTS IN A HISTORICAL WARGAME

Abstract

Evaluating LLM agents in strategic environments requires tasks that combine long-horizon planning, coordinated resource allocation, and adaptation to an opponent under explicit rules. We introduce AlameinBench, a benchmark built on a turn-based hex-grid wargame for evaluating these capabilities over complete campaigns. It comprises three scenarios centered on supplied advance, mine clearance, and timed withdrawal, each evaluated from two playable roles. The resulting six scenario–role conditions couple multi-unit coordination, logistics, phase-dependent actions, and delayed scoring within a shared game environment. AlameinBench combines a common observation and action interface, rule-validated execution, and an evaluation protocol covering terminal outcomes, scenario-specific progress, execution reliability, and computational cost. We also introduce Strategy–Allocation–Execution (SAE), a modular neuro-symbolic reference agent that connects strategic objectives to unit assignments and executable actions. Using this framework, we evaluate six LLM profiles alongside four LLM-free heuristic policies and a rule-based self-play reference. In the reported fixed-seed evaluation, performance varies across scenarios and roles, and a deterministic domain-informed heuristic matches or exceeds the best SAE profile in four of six conditions. Campaign traces further reveal differences in objective maintenance, task execution, and score sources. AlameinBench provides a structured testbed and reference baselines for studying long-horizon planning and strategic decision-making in LLM-based agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.