acceptodds
Under review as a conference paper at ICLR 2027

Monitorability Is Not Portability: When Can Runtime Monitors Be Reused?

Abstract

Runtime monitors are typically validated for a particular agent configuration, yet deployment can change the model, tools, permissions, or task environment. A monitor may continue to produce confident scores after such changes even when the ranking that made those scores useful no longer transfers. The central question is therefore not whether the monitor still runs, but whether its evidence remains valid on the new target. Across GPT 4.1 mini, GPT 4.1, and GPT 5.5, HarnessAudit remains strongly monitorable, while ToolSandbox to HarnessAudit portability is 0.520, 0.822, and 0.421. Reversing the GPT 5.5 transfer gives 0.815, and a separate transfer falls significantly below chance. These results reveal a separation between monitorability and portability. The target can retain a clear harm signal even when a source trained monitor no longer preserves the same ordering between harmful and clean episodes. This separation changes what is required for reuse. Capability, marginal distribution shift, and the tested zero label adaptations are less informative than direct target probing on the transfer grid, while a small stratified probe gives the most accurate portability estimates among the tested alternatives. SELECT TRUST converts sequential target comparisons into an anytime valid reuse decision with false trust control under i.i.d. target episodes. Finally, an independently fitted and calibrated HarnessAudit breaker raises alarms before the audited harm locus on 91.7% of harmful evaluation episodes at 10.9% false cuts. Taken together, the results show that source validation is only the starting point for reuse. The target must remain monitorable, the source ranking must remain portable in the intended direction, and the resulting score must become actionable in time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.