Measuring the Temporal and Identity Readability of Operational Logs Across Six Domains
Abstract
Retrospective event-prediction evaluations can reward recognition of temporal position or unit identity, making apparent predictive performance difficult to interpret. We introduce a reusable audit that measures these signals in operational logs before an event-prediction model is evaluated. Using windows from units without recorded adverse events, the audit asks two questions: which of two windows from the same unit came later, and whether two windows matched on time and offset came from the same unit. These tasks measure temporal readability and identity exposure, respectively; under their stated exchangeability nulls, expected AUC is exactly 0.5. We release both audits as one harness, driven by a small adapter and a declaration file, and apply it across six public domains. Removing amortization clocks reduces temporal AUC in mortgage records from 0.834 to 0.549, while identity AUC stays at 0.994. In ICU records, measurement counts alone reveal temporal position (AUC 0.634) and distinguish same-patient from different-patient pairs (AUC 0.879). A controlled case-control experiment connects identity exposure to evaluation inflation: carry-forward imputation increases the AUC gap between overlapping-unit and unit-disjoint splits by 0.118. This experiment measures susceptibility to unit overlap rather than prospective prediction skill. Together, the audits characterize temporal and identity signals that can confound evaluation and motivate controls appropriate to the intended prediction task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.