acceptodds
Under review as a conference paper at ICLR 2027

Paired Hybrid Policy Improvement (PHPI): The Geometry of Deployment Evidence

Abstract

Paired Hybrid Policy Improvement (PHPI) asks when observations justify replacing a baseline with a frozen learned controller: which feasible worlds still require opposite decisions? Under common real-chart, coverage and attainability conditions, we characterize their separation exponent by valuation ratios and provide exact primal–dual certificates for a restricted positive-monomial class. Source-linked recurrent-control constructions show that timing supports produce exact ambiguity or first-, third- and fifth-order separation, while complementary calibrations can remove collisions at prohibitive finite evidence cost. A locked 48-world study of a 16-state AC/DC model with a frozen GRU tests numerical predictions, revealing processor-dependent rankings, errors, abstentions and matched-method ties. An OPAL-RT RT-SIL study reproduces four complete replacement directions and all 36 sign/status outcomes of a precomputed digital stress map. These exact-source, numerical and RT-SIL results support different levels of evidence; continuous physical deployment remains uncertified.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.