TerraWatt: Discovering Hidden World Dynamics Through Irreversible Capital Allocation in a Physically Grounded Voxel Environment
Abstract
Long-horizon agent benchmarks divide into two families. One tests business coherence over simulated months or years, but hands the agent a world whose rules are legible from its prompt. The other withholds the rules, but tests discovery in abstract, deterministic puzzles with no capital, no risk and no delayed payoff. Neither asks whether an agent can infer how an unfamiliar world works and then commit irreversible resources to what it learned. We introduce TerraWatt, in which an agent runs a newly founded renewable-energy developer with $40M of equity on a 2.56 km voxel map. It sites solar and wind plant on land that binds, exports through a capped grid connection into an hourly market that its own output depresses, pays for permits and overruns as they fall due, and can go bankrupt. It is given the goal, its tools and a handbook of priors, but not the site-specific quantities that decide success, such as the wind resource, each vendor's true panel degradation, price elasticity and its own operating costs. These are drawn per seed from documented real-world distributions and recoverable only through costly, delayed measurement. Every exogenous process is action-invariant, so runs are bit-reproducible and agents can be compared on identical worlds, including against a hindsight planner. Running each agent both blind and briefed with the hidden parameters yields a discovery gap that separates failing to find out from failing to plan. In a pilot with three frontier models the gap takes both signs, from −0.57 to +0.11; with the parameters revealed, one model outplans a 112-candidate hindsight search; stated beliefs expose confident errors that performance hides; and the agents found nine defects in the environment that several hundred tests had missed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.