EpiTAR-Bench: Benchmarking Transferability, Adaptation, and Reliability of Epilepsy Models
Abstract
Automated seizure analysis increasingly relies on deep learning and EEG foundation models, yet their transferability, adaptation, and reliability under heterogeneous recording conditions remain poorly characterized. In this paper, we introduce EpiTAR-Bench, a unified benchmark for the Transferability, Adaptation, and Reliability of epilepsy models across seizure detection and prediction. EpiTAR-Bench spans five datasets, seven transfer directions covering cross-dataset, cross-modality, and cross-time shifts, sixteen models from four families, two window lengths, and five levels of target supervision. We further present RIVER, a transfer-specific model designed for partially observed and heterogeneous electrode configurations. Across zero-shot cross-dataset transfer, EEG foundation models form the strongest baseline family, while longer window lengths provide no consistent advantage. Adaptation gains are concentrated at lower target supervision levels for both cross-dataset tasks and cross-modality detection (p ≤ 0.0004 after Holm correction), whereas cross-modality prediction shows a positive but nonsignificant contrast (p = 0.1605) and cross-time transfer shows no comparable pattern. For seizure detection, scalp EEG transfer followed by target iEEG supervision narrows the gap to direct iEEG training, while prediction retains a larger residual gap. Finally, discrimination, calibration, and seizure onset zone (SOZ) evidence are not consistently aligned. These results establish EpiTAR-Bench as a systematic framework for evaluating how epilepsy models transfer, adapt, and remain reliable across heterogeneous EEG settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.