acceptodds
Under review as a conference paper at ICLR 2027

CCPL: Stress-Testing Delayed-Cost Attribution under Safety Constraints

Abstract

Constrained reinforcement learning often pairs an observed safety cost with the most recent action. When consequences arrive after stochastic delays, this temporal pairing can send the learning signal to the wrong decision and leave the responsible decision uncorrected. We introduce Causal Consequence-Penalized Learning (CCPL), an auditable delayed-cost attribution layer for constrained reinforcement learning. Its central mechanism is an explicit posterior-to-penalty interface: a label-free operational posterior over source lags combines a learned delay prior with a consequence-model likelihood, assigns the identifiable share of a cost to earlier actions, and retains unresolved responsibility as a state-level risk signal. For jointly caused events, a Shapley split distributes the explainable share across candidate actions. Thus CCPL separates uncertainty about which action caused a cost from uncertainty about its magnitude. Across five seeds in a bimodal delayed-cost benchmark, the hindsight assignment reduced mean consequence by with a paired 95% interval of and increased constraint satisfaction from to , while the reward difference was not statistically distinguishable from zero. With source-lag labels withheld from the operational assignment path, attribution overlap was – across three delay profiles; a uniform split reached higher raw overlap because responsibility was diffuse, but did not translate into better policy performance. A simulator intervention diagnostic with found that responsibility shares tracked measured intervention effects with Pearson , compared with for path-Shapley and for visible-step attribution. The learned delay distribution recovered a known constant delay only after extended training, reaching a predicted mean of for a true delay of . The gains are conditional: the mechanism improves safety in the binding delayed regime studied here, while other held-out conditions do not separate the methods. The action-level consequence contrast uses simulator structural-equation supervision for training and intervention validation; label-free refers specifically to the operational assignment path, which receives no source-lag or SCM label at assignment time. We therefore present CCPL as a stress-tested attribution mechanism and evaluation framework for delayed safety costs, not as a general causal-identification procedure or a guarantee for full neural training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.