acceptodds
Under review as a conference paper at ICLR 2027

When State Information Pays: Cost Separation in Sequential Multi-Model Routing

Abstract

Errors in agent trajectories can change the value of later model calls. We isolate this coupling in a two-state transmission model and compare state-conditioned policies with both deterministic and randomized open-loop schedules. Under capability ordering, the best schedule with a fixed number of strong calls has a strong suffix. This yields an exact deterministic cost in an explicit target window and reduces the randomized optimum to the lower convex envelope of K+1 suffix points. A feasible state policy then gives a matched-class cost-separation certificate. In a controlled instance, randomization reduces the open-loop cost from 12 to 5.32. The exact randomized state-history optimum costs 3.56, leaving a matched information gap of 1.76. A noisy-signal certificate vanishes under an uninformative signal, separating information from exogenous randomization. We evaluate a gold-free detector cascade on a real-execution data workflow. After calibrating a state-independent random trigger on costs alone, we freeze it before testing on 200 disjoint tasks with three fresh repetitions per policy (1,200 episodes). StopRoute reaches 79.50% success versus 74.50%, a 5.00-point paired gain on average, while the nominal-cost gap is 2.67%, inside the prespecified 5% point-estimate tolerance. Together, the exact model analysis and controlled online comparison isolate the value of where additional model calls are allocated, beyond randomization or heterogeneous model assignment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.