acceptodds
Under review as a conference paper at ICLR 2027

AERIN: Acoustic Evidence Routing and Integration for Compact Speech Representation

Abstract

Different speech tasks depend on different properties of the same signal, yet frozen self-supervised encoders leave the organization of those properties largely to pretraining. Downstream systems can select or combine internal representations, but only after those representations have already been formed. We ask whether a reusable speech representation can instead be structured around explicit acoustic evidence while still learning how that evidence should be used. We introduce AERIN (Acoustic Evidence Routing and Integration Network), a raw-waveform encoder that organizes speech into Boundary, Envelope, Phonation, and Texture evidence streams and learns their contribution through fine-grained routing and contextual integration. Thus, the sources of evidence are specified by design, while their encoding and combination remain learned from data. AERIN is distilled from WavLM-base without downstream task labels, frozen, and evaluated across 11 clinical, phonetic, and affective tasks. It remains competitive with substantially larger pretrained encoders and with its WavLM-base teacher despite using 25.66 M parameters. Routing and stream-ablation analyzes further show that the learned representation is not governed by one globally dominant stream or a fixed mixture: routing is selective at individual frame-feature positions, and downstream sensitivity to the streams varies across tasks. These results support evidence-organized representation learning as a broadly transferable alternative to relying only on emergent layer representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.