acceptodds
Under review as a conference paper at ICLR 2027

Depth-Conditioned looped Transformers for Adaptive ASR Inference

Abstract

End-to-end automatic speech recognition (ASR) systems use fixed-depth encoders, so each accuracy–compute budget typically requires its own trained model. We ask whether a single compact encoder can instead make its inference budget a deployment-time choice. A natural approach is to reuse a shared Transformer block recurrently, but we find that naive looping does not fully exploit the additional recurrent compute. We introduce LARM, a depth-conditioned looped Transformer that combines sparse connectionist temporal classification (CTC) checkpoints, supervision-clock embeddings, FiLM depth conditioning, and soft-posterior feedback. These components structure the loop into recognition checkpoints separated by latent refinement phases and let shared weights specialize across recurrent steps. A single trained model can thus stop at any checkpoint within its trained loop budget . On LibriSpeech, LARM improves WER as more loops are executed within this budget, and in a common pipeline it achieves lower WER than closely related self-conditioning and recurrent CTC methods. The design also transfers to a Conformer backbone. Compared with deeper unshared encoders in the same pipeline, LARM reaches equal or better WER with far fewer parameters and less peak memory, making it well suited to memory-constrained deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.