acceptodds
Under review as a conference paper at ICLR 2027

A Study on the Geometry of Accent in Speech Foundation Models

Abstract

Speech foundation models are known to encode phonetic, speaker, and higher-level linguistic information, as measured by decoding probes. Within the representation space, how these factors are geometrically organized is comparatively underexplored. We study this organization through accent, an attribute that self-supervised learning (SSL) models are not trained to represent. We ask how accent is arranged in representation space, how that arrangement changes across network depth, and whether it is shared across independently trained models, using three corpora (L2-ARCTIC, VCTK, GLOBE) and three encoders (WavLM, HuBERT, Whisper). We report three properties. First, accent variation forms a continuous, low-dimensional structure rather than isolated clusters, with softer cluster boundaries than speaker identity, populated interpolation paths, a connected accent graph, and a low-rank set of accent displacement directions (effective rank 5 on the parallel corpora). Removing the top five accent directions reduces a speaker-disjoint accent probe to near chance while leaving speaker identity largely intact, and ablating random subspaces of equal rank does neither. Second, because accent is deterministic given speaker, a probe whose training and test folds share speakers can classify accent by memorizing speaker identity. Under speaker-disjoint evaluation, where no speaker appears in both training and test folds, accent decodability is weakest in early layers and strengthens to a mid-to-late peak, whereas speaker identity peaks in the earliest layers. The apparent early accent structure under speaker-shared evaluation is largely speaker leakage, which shrinks with depth. Third, accent geometry agrees across the two SSL encoders (–, permutation ). On GLOBE this agreement exceeds both the untrained encoder and acoustic (MFCC) baselines and reaches the split-half noise ceiling, while on the smaller corpora the baselines approach the SSL–SSL correlation, so our learned-convergence claim is based on GLOBE. One initial expectation fails, where accent centroids cannot be expressed as convex combinations of other accents, so continuity does not imply compositionality. The speaker-first depth ordering we observe matches prior layer-wise analyses of these encoders. We release and report all code and results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.