acceptodds
Under review as a conference paper at ICLR 2027

Thaw: Benchmarking LLM Controllers Under Inference Delay

Abstract

An agent’s command can arrive too late when the environment continues to change during inference. Thaw is a benchmark for studying this timing problem in a deterministic outpost task: a large language model (LLM) directs simulated scientists to deliver supplies before shortages cause failure. We measure survival as the share of a fixed-duration episode completed before failure. To show that the task can be completed, we replay commands found by a planning program or fixed-rule controller. Every replay keeps the outpost operating for the entire fixed duration without any LLM calls. We compare a clock that pauses during inference with one that advances proportionally to each model response’s measured duration. In a task variant with brief gathering, processing and travel times, each evaluated LLM has lower survival when response duration advances the simulation. As the simulation advances more per second of inference, one LLM goes from completing the episode with the clock stopped to surviving less than half of it on average. In a separate timing test, each response advances the simulation by a delay sampled from that model’s earlier response times, independently of the current response’s duration. As we increase how far each sampled delay advances the simulation, survival falls for every model. Benchmarks that pause the environment during inference can miss this source of control failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.