PianoPoet: Unifying Piano Transcription and Rendering via Autoregressive Modeling
Abstract
Piano transcription and rendering connect the acoustic and symbolic domains of music in opposite directions; however, unified models that can perform both tasks lag behind task-specific methods. One approach to unification is to formulate both tasks as inverse sequence-to-sequence problems which share a token representation space; however, audio input and output formats appropriate for the tasks differ significantly. In this paper, we investigate this challenge and present PianoPoet, a single-checkpoint autoregressive framework that unifies both tasks with a shared transformer backbone. PianoPoet shares a common MIDI tokenizer for both directions while decoupling audio representations: continuous semantic audio frames for fine-grained transcription, and a piano-specific discrete audio tokenizer for realistic piano rendering. PianoPoet outperforms prior unified models, achieves the best results among the evaluated generative transcription methods, and outperforms specialized rendering methods in most evaluated metrics. It supports full-band stereo rendering, enabling stereo-position control through acoustic prompts. Ablations show that the proposed PianoPoet framework allows both directions to share a common backbone without quality degradation, maintaining performance comparable to its single-task variants.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.