acceptodds
Under review as a conference paper at ICLR 2027

SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving

Abstract

In multi-model LLM serving, decode execution remains inefficient due to model-specific resource partitioning: since cross-model batching is not possible, memory-bound decoding often suffers from severe GPU underutilization, especially under skewed workloads. We propose Shared Use of Next-token Prediction (SUN), the first approach that enables cross-model sharing of decode execution in disaggregated multi-LLM serving. SUN decomposes the execution of a decoder-only Transformer into a prefill module and a decode module, and fine-tunes only the task-specific prefill module, enabling a frozen decode module to be shared across models. This design enables a model-agnostic decode routing policy that balances decode requests across shared workers to maximize utilization. Across diverse tasks and model families, SUN achieves accuracy comparable to full fine-tuning while maintaining system throughput with fewer decode workers. In particular, SUN improves throughput per decode GPU by up to 2.0 over conventional disaggregation while keeping time-per-output-token (TPOT) degradation within 10%. In eight-model multi-node serving, SUN retains 98.7% of baseline throughput with 25% fewer total GPUs at moderate load and increases throughput by 29.1% on the same hardware at higher load. SUN also outperforms popularity-aware allocation and GPU colocation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.