acceptodds
Under review as a conference paper at ICLR 2027

A Physiology-Aware Hierarchical Foundation Model for Long-Context Waveform Representation Learning

Abstract

Physiological waveforms contain information at multiple timescales, from morphology within a cardiac cycle to changes across beats, minutes, and hours. Most existing waveform foundation models, however, divide signals into fixed-length patches and encode short windows independently. Consequently, their representation units may cut across meaningful cardiac structures, and their encoders cannot carry information from one window to the next. We introduce Physio-HNet, a causal hierarchical foundation model that addresses both limitations through physiology-aware dynamic chunking and continuous long-context modeling. Waveform landmarks and cardiac-cycle structure guide variable-length sub-beat and beat-level chunking, replacing fixed-length patches with physiologically informed units. Hierarchical compression and recurrent state preserve information across successive windows, allowing context to accumulate beyond isolated short segments. Following self-supervised waveform pretraining, routinely measured vital signs, including heart rate, respiratory rate, oxygen saturation, and blood pressure, supervise higher-level representations at successive measurement times along each recording. We evaluate Physio-HNet through frozen probes on two external clinical cohorts, context-length experiments, and controlled ablations. With 1.2 million parameters, the self-supervised encoder achieves performance comparable to existing waveform foundation models. Clinically anchored continued pretraining further improves blood-pressure estimation, yielding the highest and the lowest MAE in eleven of twelve blood-pressure comparisons across two cohorts and two probing protocols, including all six operating-room comparisons. Increasing available history to two hours improves estimation for most targets. Physiology-aware dynamic chunking improves several clinical estimates over fixed-length patching, while the segmentation achieving the best waveform reconstruction does not yield the best downstream representation. These findings support combining physiologically informed segmentation, continuous temporal context, and vital-sign supervision to learn transferable waveform representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.