TIDE: Topological Identity Dynamics and Evidence Learning for Short-Utterance Speaker Verification
Abstract
Short-utterance speaker verification is challenging because only limited speaker information is observed. Short speech has poor phonetic coverage and misses many identity-related transitions. Its representation is therefore sparse, fragmented, and sensitive to nuisance factors. Most existing systems still compress the available frames into a fixed-dimensional embedding, without explicitly modeling missing identity structure or evidence reliability. We propose TIDE, which views a short utterance as an incomplete observation of a speaker trajectory. Frame-level features and their local temporal differences form position velocity state clouds. Multi-scale persistent homology captures stable identity structures across temporal scales. During training, a momentum long-view branch provides a richer topological target. A duration-gated residual module estimates the missing structural information from the short view without reconstructing speech. The completed topology then modulates the vector field of a neural controlled differential equation, enabling continuous identity evolution along the observed acoustic path. A thinning-consistency loss improves robustness to missing observations. For verification, each dynamic state produces directional evidence with a nonnegative weight. These contributions are accumulated over time to construct a von Mises Fisher posterior, whose mean direction represents speaker identity and whose concentration reflects the coherence of the observed evidence. Experiments on VoxCeleb under short-duration conditions demonstrate consistent improvements over strong speaker verification baselines, especially when speech duration is limited. These results show the value of jointly modeling identity structure, dynamics, and evidence for short-utterance speaker verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.