acceptodds
Under review as a conference paper at ICLR 2027

AJAMT: REAL-TIME MULTI-AGENT MUSIC GENERA- TION WITH PREDICTIVE LOOKAHEAD AND FALLBACK SCHEDULING

Abstract

Real-time music accompaniment requires generative models to produce musically coherent responses under strict timing constraints, yet the variable inference latency of large language models (LLMs) makes their use in live interaction challenging. We present AJAMT, a real-time multi-agent framework for multi-instrument symbolic music accompaniment. AJAMT combines predictive generation with buffered scheduling and fallback mechanisms to decouple individual generation deadline misses from playback interruptions, while a lightweight symbolic evaluator reduces the overhead of iterative agentic refinement. Across currently available inference backends, the fastest evaluated configuration achieves a mean end-to-end latency of 1.21 seconds and maintains 100% playback continuity at 120 BPM. Under a multi-tempo stress test, AJAMT further maintains 100% playback continuity at 180 BPM despite 24.4% of generation requests missing their deadlines. Under a controlled shared-melody evaluation, AJAMT also remains competitive with the offline ComposerX baseline across the evaluated audio-aesthetic and symbolic-structure metrics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.