SHIR: Evidence-Guided Explanations for Time-Series Predictions
Abstract
Explaining a time-series prediction requires identifying which parts of the input affect it. Perturbation responses recorded on reference inputs can guide query selection, but do not establish the effects on a new input. We introduce SHIR for frozen time-series predictors, separating unit scoring from query selection through evidence-guided experimental design. A language model proposes unit-effect comparisons from recorded evidence and an explicit account of which endpoint effects remain unmeasured. Under a fixed perturbation rule, SHIR selects the query batch whose endpoints would complete the most candidate comparisons within the budget. An evidence graph enables many-to-many evidence reuse: one execution record serves multiple applicable registered comparisons while retaining its input and configuration identity. Reference construction explores inputs in separate episodes and freezes the merged graph and reference scores. For each new input, SHIR preselects one batch of at most three perturbation queries and evaluates the separately perturbed inputs in parallel. Each perturbation starts from the same original input. Valid current effects replace the corresponding reference scores, while unmeasured units retain their defaults. On forecasting tasks built from WeatherBench2 and VitalDB, SHIR reduces mean sufficiency error by approximately 11.2% and 18.7%, respectively, relative to each task's strongest baseline for this metric among seven external methods. Our code is available anonymously.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.