acceptodds
Under review as a conference paper at ICLR 2027

Architectural Leakage: Per-Timestep Sequence Models That Quietly See the Future

Abstract

Leakage in time-series evaluation is usually a data problem: a split that lets training see the test period, or a feature computed after the label. We document a form those checks cannot see, architectural leakage, in which the split and the features are clean and the model itself has a path from future inputs to past outputs. It arises whenever a per-timestep model conditions on a whole-sequence quantity—a bidirectional encoder, unmasked attention, symmetric padding, a pooled context vector—and it is invisible to split- and feature-level checks because the model is trained and scored under the same protocol. We give a model-agnostic detector, the future-perturbation score (FPS): perturb timesteps after , measure the change at timesteps ; a causal model scores exactly zero. It needs only forward(x, t): a replay on the truncated prefix detects the leak, and four value perturbations of the future narrow down what is read; a shuffle-only test would pass the form in the audited model, a mean-pooled context, because a mean is permutation-invariant. We then measure what such a leak is worth: on a monotone early-warning label, an oracle -bit channel carrying the onset admits a label-only score computed from the labels alone; on PhysioNet2019 a single such bit reaches (at least under a validation-selected partition), above every causal model we can train, and a bit predicted from whole-record summaries of the inputs already reaches . An audit of matched causal/non-causal pairs from four families (LSTM/BiLSTM, causal/symmetric TCN, masked/unmasked Transformer, running/whole-sequence pooling) finds every causal member of the four at – and every non-causal member beating its causal twin on all twenty splits, up to ; the gap survives per-patient scoring and the challenge utility. A routed additive model reported at AUROC over its causal baseline serves as the case study: a fair comparison gives , and reading its measurement process causally recovers AUROC on 19 of 20 splits ( normalised utility)—most of it without any router. The detector (under a minute per checkpoint) and the label-only scores (minutes on a CPU) are released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.