acceptodds
Under review as a conference paper at ICLR 2027

The Tool-Agent Observability Boundary: Predicting Recoverable Tool-Call Failures Before Any Recovery Run

Abstract

Adding more validation to a tool-using agent, whether schemas, entity checks, ranges, or self-consistency, feels like it should always help. It does not, and which failures it can reach is decidable in advance. Tool-output reliability is governed by an observability boundary: some corruptions leave evidence in the returned observation, and some cannot. A wrong-entity, partial, or malformed response violates a formal invariant and is recoverable. A stale-valid observation, a real past state of the same entity, is schema-valid, id-consistent, in-enum, and in-range. We prove a no-go theorem: no gate in the non-temporal invariant family (schema, id-echo, enum, range) separates it from a fresh record, so more reasoning or a larger model provably cannot cross the boundary. A freshness sweep confirms this: recovery falls to 0% when the verifier returns only same-cache evidence. Our contribution is to make the boundary predictable rather than to assert it. LOCO assigns per-observation causal criticality labels by leave-one-corruption-out replay, and shows that the obvious structural proxies fail unpredictably: read→write recall ranges from 0.0 to 1.0 across domains, so the boundary cannot be eyeballed. StructAudit compiles admission predicates from tool schemas and clean examples alone, then emits a frozen coverage map that places each corruption class on the correct side a priori and label-free, at 87.9% per-break agreement across three causal domains and predicate-level precision 1.000 with recall 0.700 on retail. FreshVERIFY supplies the fresh read where the map predicts a freshness limit. Measured against the strongest cheap gate rather than the weakest, one invariant does most of the work: on gpt-oss-120B, id-echo alone matches or exceeds the full predicate family on recovery, and it is the most robust cheap gate net of lost clean rescues; where the fuller family recovers more, it does so at higher rescue-loss, and declared-schema validation trails both. This corroborates the boundary rather than undercutting it, since the recoverable corruptions are exactly the entity-observable ones. The resulting operating rule for tool-output admission: validate what is observable, verify what is not, and do not expect in-trajectory reasoning to manufacture missing freshness. Whether the same informational constraint governs verification signals more broadly remains open; we settle it where it can be measured.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.