HealthTTS: Symptom-Aware Text-to-Speech with Clinical Acoustic Cues
Abstract
Speech in healthcare interactions conveys not only verbal content but also symptom-related acoustic cues, such as coughs, throat clearing, and sighs. However, our evaluation of representative text-to-speech (TTS) models shows that existing models often fail to reproduce the requested cue type, count, and order, while even correct cue sequences may not integrate naturally into speech. To address these challenges, we propose **HealthTTS**, a symptom-aware TTS model that directly generates complete utterances from text, a speaker reference, and an optional cue configuration. HealthTTS follows a three-stage training procedure that combines cue-aware language-model adaptation, localized acoustic adaptation, and global preference refinement to improve both cue generation and natural cue–speech integration. We construct SymCue, a paired speech resource for training, and **SymCueBench**, an evaluation benchmark covering seven symptom-related acoustic cue types. Under the shared-cue setting, HealthTTS achieves 87.3% Cue Sequence Accuracy, outperforming Fish Speech S2 Pro by 25.1 percentage points while maintaining high Speech Naturalness and Cue Realism. Beyond benchmark evaluation, a virtual standardized patient study with 30 medical students across 14 scenarios shows that HealthTTS achieves higher Patient Realism and Perceived Usefulness than Cue Insertion, a post-hoc insertion reference. To support future research and the open-source community, we will make our code publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.