acceptodds
Under review as a conference paper at ICLR 2027

CivWorldBench: Evaluating Language World Models in Structured Multi-Agent Sandbox Worlds

Abstract

The growing use of language models for forecasting motivates a broader evaluation of their world-model capabilities. Such evaluation must go beyond predictive accuracy to assess hidden-state inference and intervention prediction. Real-world forecasting benchmarks offer limited access to hidden-state ground truth and controlled experiments for evaluating intervention predictions. We introduce CivWorldBench, a benchmark built on a multi-agent strategy game engine with recorded internal states, controlled observations, and paired intervention rollouts. Its 1,900 decision points from 475 worlds support three task levels: association forecasting, belief inference, and intervention prediction. Controlled observations and reference predictors support attributable evaluation of capability gaps, with an expert harness and empirical ceiling estimates contextualizing scores. Stored trajectories make intermediate predictions verifiable, while agent-play experiments test whether world-model capabilities are usable for decision-making. The main panel comprises eleven frontier language models spanning different capability levels. Composite benchmark scores correlate with real-world forecasting leaderboards on matched model panels, more closely than with a general-capability index, supporting its use as a proxy for real-world forecasting capability. Strong association forecasts do not consistently extend to belief inference or intervention prediction, and effect magnitudes remain poorly calibrated despite informative direction predictions. Process verification reveals errors missed by final-answer evaluation, and structured reasoning improves intervention prediction for the strongest model in the process study. Agent-play experiments identify useful assisted plans and further motivate evaluating decision-relevant capabilities beyond association forecasting. These findings identify hidden-state inference, intermediate-state prediction, and effect calibration as priorities for world-model training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.