acceptodds
Under review as a conference paper at ICLR 2027

Information Without Geometry: The Encoding–Reliability Dissociation In Time Series Foundation Models

Abstract

Time series foundation models (TSFMs) are increasingly used as embedding backends for retrieval, anomaly detection, and clustering, all of which assume that embedding proximity reflects predictive similarity. This assumption is rarely tested directly: standard forecasting metrics say nothing about the internal geometry of a representation, and the geometric analysis long applied to vision and language encoders has not been extended to time series. We introduce a conceptual decomposition of “representation quality,” usually treated as a single notion, into three properties: Information (is task-relevant structure encoded?), Geometry (does embedding proximity predict forecast difficulty?), and Utility (does that geometry actually improve a downstream operation?). We show these properties can come apart: a representation may contain useful temporal information, and its local geometry may be meaningful, without making nearest-neighbour retrieval useful. We call this pattern the Encoding–Reliability Dissociation (ERD). A supporting theorem shows the dissociation is possible for any injective encoder, so it is not just an artifact of the specific models we happen to test. Across four architectures spanning three pretraining paradigms, four benchmarks, and matched random-weight controls, we find exactly this pattern: temporal structure is often encoded, local geometry is often real, and yet embedding-based retrieval never outperforms a trivial, model-free raw-window baseline, even at the one point where the geometric signal is strongest. A metric-learning control confirms that geometry can be engineered on demand, but utility still does not reliably follow; utility has to be earned on its own, not assumed from information or geometry. This matters for anyone building retrieval, anomaly-detection, or clustering systems on top of these representations: a diagnostic can be statistically sound and still fail to tell you whether the downstream operation will work, so it should guide further investigation rather than stand in for testing that operation directly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.