acceptodds
Under review as a conference paper at ICLR 2027

Prompt-Attributable Feasibility Violations of Medium-Weight LLM-Based Control of a Cascaded Five-Reservoir System

Abstract

When an LLM is tasked to control a complex real-world problem, a physically infeasible control action is usually projected onto the feasible operating range by the environment, allowing the simulation to continue. Yet, this can mask repeated constraint violations behind strong aggregate performance. Here, we investigate prompt-attributable feasibility failures in a twenty-year, non-resettable, multi-reservoir hydrological testbed, using two LLM families – DeepSeek-R1-Distill and Gemma 4 – as controllers making sequential operational decisions under physical constraints. Empirically, despite achieving near-optimal aggregate performance, the LLM agents violate physical constraints in 10 to 28% of decisions. We show that these violations are systematic rather than stochastic and arise from two reproducible, prompt-attributable reading failures. The first, constraint anchoring, occurs when a field specifies a maximum output limit while explicitly excluding an incoming quantity. DeepSeek-R1-Distill interprets this limit as total availability and systematically under-commits, whereas Gemma 4 does not. The second, forecast presence bias, occurs when a predictive field is introduced into an observation whose other fields describe currently available quantities, causing agents to spend the predicted quantity as if it was immediately available. The effect is driven by forecast presence, rather than magnitude: scaling the forecast to an unambiguous underestimate still triggers violation rates similar to the unscaled forecast case, while removing the forecast drops up to 87% of the violations. We investigate both mechanisms with a five-step attribution protocol – open-loop replay, single-element perturbation with the remainder byte-identical, failure-rate comparison, domain-stripped simplified replay, and token-preserved identity substitution. The results isolate structural misreading from domain reasoning. Constraint anchoring survives complete removal of domain vocabulary, and is eliminated by a single arithmetic clarification. Conversely, forecast presence bias disappears in domain-stripped simplified replay, but returns intact under identity substitution, so neither violation is semantic. Our findings therefore show that the observed feasibility violations arise not from hydrology semantics or the specific benchmark, but from a specific model family's anchoring tendency and a general disposition to remain optimistic even with an underestimated forecast.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.