acceptodds
Under review as a conference paper at ICLR 2027

DETECTOR-COUPLED SURROGATE REGRET FOR TEST-TIME ADAPTATION IN PIECEWISE-STATIONARY CMDPS

Abstract

Constrained reinforcement-learning policies deployed in provisioning and finance must determine not only how to adapt under environmental shifts, but also when adaptation is warranted. Keeping a policy frozen can reduce reward and move its cost behavior away from the levels it was trained to satisfy, whereas indiscriminate test-time adaptation can overfit transient noise. We study this timing problem in piecewise-stationary constrained Markov decision processes (CMDPs) with imperfect change-point scores. Existing detection-and-restart analyses price detection delay and false alarms but omit constraints, while online-CMDP analyses price constraint costs but assume no external detector. Neither targets the timing decision. We propose AdaptiveCMDP, which closes both gaps: constraint exposure is fixed before deployment, and detector error is priced inside the regret bound. Specifically, it leaves the trained policy frozen and adapts only a bounded adapter whose radius, set from the Slater margin, fixes that exposure. Updates are reward-free, so adaptation can start before returns are informative; each detector firing restarts the step-size clock. The surrogate dynamic-regret bound then prices the detector directly—a detector tax that is square-root in horizon, regime count and expected false alarms, plus linear total detection delay. Under explicit value sufficiency and realizability, this transfers to policy value on an adapter-independent target stream. A frozen exploitability–delay rule scores every registered window: it correctly keeps the policy frozen in 15 of 16 financial panels where adaptation does not pay, and fires on all three provisioning domains—public traffic, bike-sharing and electricity-load data. There a capacity budget replaces the financial drawdown limit; strict regret wins exclude zero, at the cost of more stale-capacity breaches in two. A matching lower bound shows these terms are unavoidable: detector quality, not adapter capacity alone, decides when adaptation pays.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.