acceptodds
Under review as a conference paper at ICLR 2027

Provable Speech Attribute Control via Latent Independence

Abstract

Conditional generation and disentangled representation learning are central to controlled generation across audio, vision, and multimodal domains. However, despite strong empirical progress, particularly in speech style transfer, most existing approaches rely on heuristic objectives and architectural choices, offering limited theoretical understanding of when and why reliable attribute control is achievable. In this work, we develop a formal framework for speech attribute conversion and provide a theoretical analysis of sufficient conditions for exact and consistent transfer. Our analysis focuses on a deterministic autoencoder setting augmented with an independence constraint between the learned latent representation and the controllable attribute. Under explicit population-level assumptions about the data-generating process, we establish guarantees linking reconstruction, independence, and the feasibility of attribute manipulation while preserving task-relevant content. We further show how the theoretical framework translates into practice by proposing a practical voice conversion method that directly implements its core principles. Experimental evaluations on voice and pitch conversion tasks demonstrate the applicability of the theoretical analysis to real-world speech conversion settings and show that the resulting method achieves competitive performance against existing approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.