TRACE: Temporal Relevance Attribution for Clinical Explanations
Abstract
Irregularly sampled multivariate clinical time series are central to intensive care unit (ICU) risk prediction, yet deployable models must be both highly accurate and faithfully interpretable. Practitioners currently face a sharp trade-off: inherently interpretable architectures often rely on heuristic attention mechanisms without strict faithfulness guarantees, while rigorous post-hoc methods such as Integrated Gradients incur prohibitive computational costs. We introduce Temporal Relevance Attribution for Clinical Explanations (TRACE), a framework that resolves this tension through architecture-attribution co-design, achieving the strongest predictive performance together with single-pass, mathematically exact attribution. TRACE structurally reduces to an affine map for any fixed input, enabling a single backward pass to exactly decompose the prediction logit into signed per-variable, per-timestep contributions with a completeness gap at floating-point round-off. On PhysioNet-2012, PhysioNet-2019, and MIMIC-III, TRACE achieves the highest AUROC and AUPRC of all baselines. Under a graded remove-and-retrain evaluation (gROAR) that sweeps retention budgets with trend and equivalence testing, TRACE consistently outperforms existing inherently interpretable models. Its top- degradation is comparable to that of Integrated Gradients, yet it runs two orders of magnitude faster and preserves machine-precision faithfulness, with a relative completeness gap on the order of . Together, these results show that predictive accuracy, attribution faithfulness, and computational efficiency need not be traded off in clinical time-series modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.