ProbeCRC: Reliable Probe-Guided Speech Attribute Steering with Conformal Risk Control
Abstract
Modern text-to-speech models generate highly natural speech, yet controlling specific speaker attributes typically requires dedicated model adaptation and training. Existing inference-time alternatives can substantially degrade speech quality and provide no guarantees on the resulting attribute change. We introduce ProbeCRC, a probe-guided activation-steering framework with risk-controlled strength selection for pretrained autoregressive text-to-speech models. Layer-wise linear probes identify attribute-related directions, which we use to construct regularized hidden-state interventions. A shared steering parameter, defined in probe-logit space, translates the desired steering strength into layer-specific perturbations. We calibrate this parameter on held-out data using Conformal Risk Control (CRC), enabling attribute manipulation under a user-specified bound on failure risk. Experiments demonstrate effective bidirectional shifts in gender, emotion, and age. ProbeCRC outperforms baseline steering approaches and satisfies the prescribed risk guarantees while minimally affecting speech quality. Demo samples are available online.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.