acceptodds
Under review as a conference paper at ICLR 2027

Similarity Is Not Functional Alignment: What Time-Series–Language Alignment Must Establish

Abstract

Time-series–language alignment is commonly evaluated through retrieval and representational similarity. These scores measure correspondence between paired series and descriptions; they do not by themselves certify whether a statement is supported by the observed history, adds new conditional-mean information about the future, or improves a fitted predictor. We formalize these distinctions as Functional Alignment (FA), which separates representation correspondence, historical support, predictive innovation, and fitted-model assistance. Our central result is an identification limit: with the history–future process held fixed, two ways of selecting true historical statements can induce the same history–text distribution and the same support labels, yet different predictive innovation. No correspondence measure of fixed encoders can therefore tell the two apart, and a history-derived feature with zero innovation can still assist a restricted predictor. Across 20 encoder configurations, retrieval reaches 51.8 to 156.4 times chance, yet the strongest retriever has temporal factual-reversal accuracy of 0.4984. Executable Claim Alignment (ECA), which checks typed interpretations against observations, reaches domain-macro accuracy 0.9874 in controlled verification across 11 held-out CaTS domains. On Time-MMD, task-specific projection training nearly doubles Linear CKA, yet its forecasting increments over matched untrained projections range from −3.53% to 5.67% across six backbones. Functional claims about alignment require evidence matched to the target being claimed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.