VoxTact: Preserving Vocal Nuance for Acoustically Appropriate Responses
Abstract
The same words can call for different responses when spoken with different affect, yet responding distinctively can also introduce unsupported assumptions. We present VoxTact, a spoken-dialogue framework that learns when vocal nuance should change a reply while preserving factual and role consistency. ToneTrace retains temporally ordered acoustic summaries alongside the existing speech representation. Appropriateness-Relative Policy Optimization (ARPO) evaluates each candidate reply across paired contexts and rewards supported specificity under quality and relation constraints. ToneFoil provides 100,000 training recordings with reviewed response relations; ToneFoil-Bench evaluates strictly matched wording and distinguishes justified adaptation from appropriate shared responses. ToneTrace preserves the cues; ARPO learns their relevance to response suitability; ToneFoil tests the distinction. In matched response-learning comparisons, VoxTact achieves 63.34% bilateral response success, compared with 49.94% for sample-level policy optimization, while attaining 94.10% quality acceptance and 2.90% unsupported assertions. Representation and objective controls distinguish the contribution of acoustic input from additional capacity and expose the cost of rewarding specificity without quality constraints. The results support selective acoustic conditioning rather than indiscriminate emotional differentiation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.