acceptodds
Under review as a conference paper at ICLR 2027

FastEPD: In-Network Scheduling for Disaggregated LMM Inference

Abstract

Encode–Prefill–Decode (EPD) disaggregation has been adopted to improve the performance of Large Multimodal Model (LMM) services. It separates the encoding (E), prefill (P), and decoding (D) stages across dedicated GPU servers while generating multi-stage data traffic among different servers. These data flows are commonly transmitted by switches in data centers, where queuing management and scheduling mechanisms significantly affect the overall performance. However, existing EPD inference systems lack in-network optimization which is necessary for efficient flow identification and queue scheduling, given the limited network resources and complex traffic characteristics. In this work, we propose FastEPD, an in-network identification and scheduling framework to accelerate LMM EPD inference by reducing the service completion time (SCT) for network transmissions. It consists of two main components: 1) lightweight multidimensional flow identification, and 2) service-oriented multi-queue scheduling. Experiments on a DPDK prototype driven by real model workloads show that FastEPD outperforms canonical network schemes, improving the overall SCT by 57.67 - 79.09% over FIFO and 41.02 - 68.10% over RSS, respectively. Experiments on a two-tier fat-tree also demonstrate that FastEPD remains beneficial as the topology scales from 32 to 128 endpoints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.