Large Language Models for Multimodal Irregular Time Series with Informative Sampling
Abstract
Multimodal irregular time series (MITS) consist of asynchronous observations from heterogeneous numerical and textual channels, such as lab measurements and clinical notes in electronic health records (EHR). LLMs are natural candidates for processing such heterogeneous data, yet existing MITS models rely on specialized fusion mechanisms rather than leveraging LLMs for unified MITS processing, and LLM-based approaches for irregular time series address purely numerical observations. Moreover, no prior work has examined whether LLMs can learn to use informative sampling patterns. We introduce MILM (Multimodal Irregular time series Language Model), which serializes MITS into a time-ordered sequence of time-channel-value triplets and fine-tunes an LLM through a two-stage strategy for MITS classification. The first stage trains on value-redacted MITS to predict from sampling patterns alone, and the second stage trains on full MITS to jointly model sampling patterns and observed values. MILM achieves the best average rank of 1.1 against 15 baselines on 4 EHR datasets, improving average precision by up to 3.1 points over the strongest baseline. The two-stage strategy also outperforms directly fine-tuning on full MITS in a single stage (MILM-Direct), even when the single-stage baseline receives the same total number of training epochs. Value redaction evaluations confirm that sampling patterns carry predictive signal and that MILM learns to exploit them. Under the value pending evaluation we introduce, where some values are unavailable at prediction time, MILM remains ahead of MILM-Direct when the affected observations are dropped, and its lead widens when their time and channel are shown. Our code is available at https://anonymous.4open.science/r/milm-7750.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.