acceptodds
Under review as a conference paper at ICLR 2027

Reconstruction Wastes Rate: Selective Prediction over Code Targets for Biosignal Foundation Models

Abstract

Discrete tokenizers for physiological signals are typically trained to reconstruct waveforms or spectral features. Reconstruction rewards encoding any component that lowers squared error, including noise that the surrounding context cannot predict, whereas a context model benefits from codes that the context can predict. We introduce (Selective Prediction Over Targets), which trains a context encoder to identify the quantized teacher code of a masked signal cell among candidates, without reconstructing the waveform or spectrum. On clinical EEG with injected noise, a probe recovers more of the noise from representations learned with reconstruction than from those learned with SPOT. We use SPOT to train CORTEX, a channel-agnostic encoder for EEG, EOG, EMG, ECG, and respiratory channels, and compare it with eleven baselines on eight clinical tasks spanning sleep and epilepsy. We show that is robust to increasing input noise and largely retains its clean-data performance under muscle-band contamination. A controlled comparison on an existing backbone, varying only the pretraining objective, attributes this robustness to the objective rather than the architecture. These results suggest that discriminative, prediction-based objectives yield representations that are more robust to noise than reconstructive ones, with direct implications for physiological signal encoders deployed in real-world, low-SNR clinical settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.