acceptodds
Under review as a conference paper at ICLR 2027

COUNTERFACTUAL PLASTICITY: RISK-AWARE SELECTION OF MODEL UPDATES BEFORE EXECUTION

Abstract

An update can teach a language model new information while harming behavior that was previously reliable. When several update procedures are available, the question is not only whether to update, but which update to apply. We introduce Counterfactual Plasticity (CP), a framework that predicts the likely benefit, collateral damage, and damage uncertainty of each candidate before it is run. CP calibrates the resulting selection rule with held-out split conformal risk control, then chooses the highest-benefit admissible candidate or abstains. Controlled experiments show that benefit is easier to forecast than damage, and that simple forecast-based selection can fail when the candidate menu has real benefit and damage trade-offs. In frozen factual-editing evaluations, CP substantially reduces harmful selections in two higher-risk regimes. In an independent Qwen2.5-7B evaluation where the baseline is already benign, calibration selects λ = 0 and leaves the baseline decision unchanged. Together, these results show how a model can compare plausible updates before committing to one, while adding conservatism only when the evidence supports it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.