acceptodds
Under review as a conference paper at ICLR 2027

Structured Prosody Modeling for Long-Form Text-to-Speech via Semantic-Conditioned Cross-Scale Regularization

Abstract

Despite recent advances in neural text-to-speech synthesis, long-form generation still suffers from prosodic degradation, including inappropriate pauses, pitch drift, unstable energy patterns, and rhythmic discontinuities. We argue that these failures arise because prosody is often modeled as a flat acoustic sequence, a global style variable, or a reference-conditioned latent factor, without explicitly enforcing structural consistency across phrase-, sentence-, and discourse-level contexts. To address this limitation, we formulate long-form prosody modeling as a cross-scale structural-consistency problem and propose SPM-TTS, a structured prosody modeling framework for neural speech synthesis. SPM-TTS couples three components: 1. **Segmental Information-Guided Generation**: introduces linguistically motivated boundaries, segmental position encodings, and boundary-conditioned acoustic prediction to improve local prosodic organization. 2. **Semantic-conditioned Prosody Consistency Regularization**: uses global prosodic anchors, semantic offsets, and adaptive consistency weights to suppress harmful long-range drift while preserving semantically justified local prosodic variation. This regularization is applied only during training and introduces no additional decoding cost at inference time. 3. **Neural codec tokenizer**: provides expressive acoustic representations that support stable reconstruction and prosody-aware generation. Experiments on AISHELL-3 and long-form Chinese inputs show improvements in speech quality, intelligibility, speaker similarity, and prosodic stability over representative baselines, with more stable pitch and energy trajectories in long-form generation. Ablation studies further indicate that segment-aware local modeling, semantic-conditioned global regularization, and codec-level acoustic representation provide complementary gains. These results suggest that explicit cross-scale structure provides an effective inductive bias for robust long-form TTS.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.