Paired Hybrid Policy Improvement (PHPI): The Geometry of Deployment Evidence
Abstract
Paired Hybrid Policy Improvement (PHPI) asks when observations justify replacing a baseline with a frozen learned controller: which feasible worlds still require opposite decisions? Under common real-chart, coverage and attainability conditions, we characterize their separation exponent by valuation ratios and provide exact primal–dual certificates for a restricted positive-monomial class. Source-linked recurrent-control constructions show that timing supports produce exact ambiguity or first-, third- and fifth-order separation, while complementary calibrations can remove collisions at prohibitive finite evidence cost. A locked 48-world study of a 16-state AC/DC model with a frozen GRU tests numerical predictions, revealing processor-dependent rankings, errors, abstentions and matched-method ties. An OPAL-RT RT-SIL study reproduces four complete replacement directions and all 36 sign/status outcomes of a precomputed digital stress map. These exact-source, numerical and RT-SIL results support different levels of evidence; continuous physical deployment remains uncertified.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.