acceptodds
Under review as a conference paper at ICLR 2027

TrinityASR: An End-to-End Model for Long-Form Speech Recognition and Speaker Diarization

Abstract

In autoregressive transcription of long multi-speaker conversations, incorrect speaker labels can influence subsequent predictions and lead to persistent attribution errors. We present TrinityASR, an end-to-end model that jointly generates transcript text, speaker identities, and segment timestamps, supporting up to 90 minutes of audio in a single decoding run. To improve robustness to erroneous speaker-label histories, we propose Speaker Identity-Preserving Optimization (SIPO), which combines corrective supervised fine-tuning with reinforcement learning. Its supervised stage, Speaker Identity Decoupling (SID), perturbs historical speaker tags and trains the model to make corrective attribution decisions while accounting for the equivalence of speaker-ID renamings. Reinforcement learning then optimizes complete transcripts using a reward based on concatenated minimum-permutation character error rate (cpCER) and diarization error rate (DER). Across AISHELL-4, AliMeeting, Podcast, and Movies, TrinityASR achieves the lowest cpCER among the evaluated systems with available results. On AISHELL-4, RL initialized from SID training achieves 15.83% cpCER, compared with 24.30% for RL initialized without SID. Analysis of 32 internal two-speaker recordings further shows less persistent attribution errors and fewer long error runs for the SID-trained model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.