acceptodds
Under review as a conference paper at ICLR 2027

HINT-AD: Hierarchical and Interpretable Transformer for Speech-Based Alzheimer's Disease Staging and Marker Discovery

Abstract

Accurate staging within the mild cognitive impairment (MCI) continuum is challenging because progression from early MCI (eMCI) to late MCI (lMCI) manifests through subtle, interacting changes in speech, language, and cognition. Can eMCI and lMCI stages be distinguished from connected speech by explicitly modeling the hierarchy of cognitive-linguistic disruption while retaining interpretable evidence for each prediction? We introduce HINT-AD, a hierarchically interpretable Transformer that integrates phonological, syntactic, semantic, and cognitive-linguistic markers through marker-guided Transformer stages for eMCI–lMCI staging. Rather than treating heterogeneous markers as independent features, HINT-AD progressively models their interactions while retaining links from marker summaries to supporting acoustic representations and transcript spans. Across 420 English and Chinese picture-description recordings from 218 participants, using MCI-stage labels transferred from the Alzheimer's Disease Neuroimaging Initiative (ADNI) reference cohort, with an additional 89 Spanish passage readings used for auxiliary training only, HINT-AD achieves 80% and 79% ROC-AUC and 69% and 74% average precision (AP) in English and Chinese, respectively. Compared with the strongest baselines, these results correspond to relative improvements of 8.1% and 5.3% in ROC-AUC and 7.8% and 10.4% in AP, respectively. Joint English–Chinese training improves ROC-AUC by 1.4% and 8.3% relative to language-specific training, while joint training with Spanish auxiliary data yields relative gains of 14.3% and 9.7% over the same baselines, respectively. Marker-group ablations further show that masking any linguistic level reduces performance in both languages, with the largest ROC-AUC reductions arising from semantics in English and cognition in Chinese. SHAP analysis identifies semantic markers as the strongest positive contributors in both languages, including role identification and lexical use, with additional language-specific contributions from scene integration and contrast use in English and thematic coherence and verb use in Chinese. Together, these findings provide evidence that hierarchical modeling of connected-speech markers improves fine-grained staging within the MCI continuum while exposing interpretable cognitive-linguistic patterns associated with stage. HINT-AD reframes speech-based within-MCI staging from flat classification over heterogeneous biomarkers to evidence-linked hierarchical representation learning, offering a clinically motivated and interpretable framework for computational modeling of subtle neurocognitive change.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.