acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Open-Ended Decision-Making with LLM-as-Environment

Abstract

We study how to evaluate open-ended decision-making. Because outcomes may take years to become clear and alternative choices cannot be observed, evaluating these decisions requires simulating how different choices shape future outcomes. We introduce , a framework for constructing environments for this purpose. Our key idea is to use an LLM as part of the environment to simulate changes that are difficult to specify with fixed rules, while code handles calculations, constraints, and state updates. We also develop design principles and skill files that guide coding agents in constructing, calibrating, and stress-testing these environments. Using this framework, we build , with environments for undergraduate career planning, startup strategy, and NFL head coaching. We validate the environments using human judgments, historical data, and a controlled comparison with an existing rule-based simulator. Replacing rule-based transitions of a programmatic startup simulator, YC-Bench, with LLM-predicted transitions preserves model rankings across eight models (Spearman ), compared with between independent runs. We then evaluate 11 frontier language models across all three environments. Even the strongest models close only – of the utility gap between a scripted baseline and a privileged search reference, leaving substantial room for improvement. Advances in LLM capabilities enable a new approach to building decision-making environments: LLMs can help construct them and also serve as part of the environment itself. We release the environments, evaluation harness, and builder so that researchers can refine these environments and construct new ones for other decision problems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.