acceptodds
Under review as a conference paper at ICLR 2027

Streaming Audio-to-Score Transcription

Abstract

We present the first streaming music transcription system that converts live audio directly into human-readable sheet music. Our system is comprised of a cache-aware Conformer–RNN-T that emits a playback-ordered Intervals-and-Moments (InterMo) notation sequence in real-time, with an incremental sheet music rendering system for low latency. Our model emits InterMo tokens as solo piano, solo violin, and violin–piano duet performances unfold, directly producing time-aligned scores spanning up to three staves, without intermediate stages, updated at each musical moment. We benchmark streaming transcription on different latency configurations between 0.16 and 2.24 seconds, and show that our streaming model performs competitively with conventional offline audio-to-MIDI-to-score cascades. We additionally introduce a musically motivated *conductor* buffer, and show that a startup delay of only 5 seconds substantially increases streaming quality, especially in piano music. By bringing metrically structured, polyphonic notation into the streaming setting, our approach moves real-time music transcription toward human-facing applications in which live performances can be immediately captured, read, and communicated as conventional musical scores.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.