acceptodds
Under review as a conference paper at ICLR 2027

Phase Dependence Is a Task Property: Measuring Classification Accuracy Beyond the Fourier Magnitude in Time-Series Classification

Abstract

Complex-valued and other phase-aware architectures are regularly proposed for time-series tasks, yet their reported benefits are inconsistent. We argue the question "are phase-aware models better?" is underdetermined: phase sensitivity is first a measurable property of tasks. Because x -> |Fx| is not injective (its fibers contain all circular shifts, reflections and negations), any amplitude-spectrum classifier is blind to label structure varying within fibers, so the Bayes accuracy reachable from the global amplitude spectrum is a true ceiling for that class. We measure not that ceiling but how much of it classifiers reach: the phase-accessibility gap Delta_phi, the test-accuracy difference between the best model given the full signal and given only the amplitude spectrum, with model and representation selected by cross-validation on training data alone. It screens a model-agnostic target with a calibrated model family and is stable across disjoint estimator families. Against exact synthetic ground truth it gives no false positives in 30 amplitude-only tasks or 36 phase-uninformative nulls, but 11 of 12 on redundancy nulls (the label a function of |Fx| yet the phase encoding it), so what it identifies is a finite-family bound, not a Bayes gap. It decomposes, relative to an explicit shift-invariant repertoire and up to a measured estimator term, into a position component a fixed-window spectrogram largely recovers (81-86%) and a relative-phase component it recovers much less of (54-66%). Across 114 UCR datasets the estimated gap spans -0.20 to 0.53. Estimated from training data alone it predicts held-out gain across five model families with raw-signal access (1D CNN, MiniRocket, InceptionTime, a complex-valued CNN, 1-NN DTW; partial correlations 0.52-0.68 against untouched test performance, family-clustered intervals excluding zero) and is near zero where the gap is; an intervention randomizing phase while preserving amplitude drives the gap and an independent family's gain down together (r=+0.91). On the UEA multivariate archive, which took no part in its development, it again predicts held-out gain (r=+0.70, +0.62) and sorts it into the same semantic categories. Per dataset it is a screen: under train-selection uncertainty 33 of 114 datasets are confidently positive, 24 show no detected gap and 53 are inconclusive. As a case study, a real CNN given the analytic-signal input [x, Hx] matches or beats its complex counterpart on the trunk and budget we test, so what helped here was a phase-bearing input, not arithmetic. The literature's inconsistency is what a task property, interacting with input and readout, predicts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.