When Learned Options Become Infrastructure: Promotion-Induced Maintenance in Control Lifecycles
Abstract
When a learned option becomes an action for another policy, a change to its behavior can also invalidate that policy. Which callers should be suspended, and when is collecting more evidence cheaper than revalidating them? We study these questions as promotion-induced maintenance. Our rule uses the option-to-policy dependency graph twice: to block potentially stale calls and to choose which option to probe next. Unresolved dependencies remain blocked. On identical action–consequence data from 96 fresh CARL PointMass worlds, removing the graph eliminates the allocation response, supplying the wrong graph reverses it, and restoring the graph restores the original decisions. Correct topology lowers a source-calibrated maintenance-cost objective by 0.9% on average. One option-reliability failure keeps the preregistered composite endpoint unconfirmed. On one released SIL-C checkpoint, dependency-local revalidation preserves guarded execution of two policies outside the affected closure during validation and saves 4.17% of executed validation-plus-service actions, with identical paired trajectories for all 24 policies. These results show how downstream dependencies can guide evidence collection as well as maintenance. They do not establish benefits across natural task streams.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.