acceptodds
Under review as a conference paper at ICLR 2027

A Fixed-Text Diagnostic for Selector-Aligned Source-Acoustic Prompting in Speech-to-Speech Translation

Abstract

Speech-to-speech translation can preserve target words while changing delivery. We examine a source-acoustic policy that measures source speech to select a target-language prompt. Its selector-aligned score uses the same four acoustic summaries. We introduce a counterfactual protocol that fixes the text-to-speech (TTS) input, renderer, candidate bank, decoding, and utterance identifier while changing only prompt policy. We test 160 read-speech utterances from the Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) validation set across English-to-German, English-to-French, English-to-Dutch, and English-to-Turkish. We compare source-acoustic prompting with function-label-only prompting, one recorded shuffled-source allocation, and three recorded random allocations. On the selector-aligned score, source-acoustic prompting exceeds the recorded shuffled-source allocation but is inconclusive against the function-label and random controls. Against function labels, the score comparing generated speech with held-out target references has a lower point estimate in all four language pairs. Lower values mean less acoustic similarity. The English-to-Dutch interval remains inconclusive. The duration score is higher for English-to-German but lower in the other three pairs, so it does not show a stable timing benefit. A stricter automatic audit does not use the selector's four stored summary values when it scores the outputs. In that audit, only the pooled pitch contrast against shuffled source is positive, and we do not treat that isolated contrast as standalone evidence. Energy, duration, and pause remain inconclusive, and some Dutch contrasts still have no paired support. On the 80-item support-screened WavLM follow-up, frozen embeddings are lower against function labels and inconclusive against shuffled source; 29 of its IDs overlap the pilot, so it is not an independent-cohort replication. The predeclared positive rule requires both contrasts to be above zero and is not met. Together, these automatic checks show movement in the selector's preferred signal, not reliable improvement in held-out target speech.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.