acceptodds
Under review as a conference paper at ICLR 2027

Interval-Width Drift: A Controlled Study of Localization, Detection, and Error Ranking

Abstract

Can interval widths from a frozen imputer support variable localization, independent-window detection, and individual error ranking interchangeably? We study the evidence needed for each use. With draws from one coronary-trained flow imputer, raw-width localization AUROC on 40 paired PhysioNet 2012 encounters is 0.987 with shared imputation randomness, but 0.386 with independent randomness; disjoint groups of 40 give 0.447. The latter two estimates are compatible with chance. Monte Carlo-only controls locate a dependence on numerical noise cancellation before patient pairing is removed; paired localization alone is therefore insufficient monitoring evidence. Across window sizes 20–100, width-statistic responses remain small while a direct timing diagnostic increasingly detects the intervention. Source-validation pseudo-masking reveals worse reconstruction than simple linear interpolation and small, measurable widening after context removal. A separate encounter–feature-matched design changes the hidden positions; harder selections increase error with weak average response in both local width and the deployed representation, including a 50-draw sensitivity check. In a separate supervised error-ranking analysis, restoring probability information to the baseline nearly eliminates the apparent benefit of widths. We provide an evaluation protocol that tests these uses separately, varying patient pairing, random-stream coupling and baseline information. The findings delimit this imputer's behavior rather than establishing a universal failure of width monitoring.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.