Protect Before Evicting: Lagrangian-Guided Encoder Caching for Disaggregated Multimodal LLM Serving
Abstract
For multimodal models with separate encoders, repeated inputs allow encoder outputs to be reused across requests, even when the questions differ. A bounded encoder cache must decide which outputs to retain because their sizes, recomputation costs, and reuse frequencies differ. We present Protect Before Evicting, a Lagrangian-guided cache policy for disaggregated multimodal LLM serving. Offline calibration solves a scalar problem using an object profile and a cache budget. Online decisions use its solution to grant temporary protection, while allocations may override that protection when space is needed. The implementation preserves active-reference restrictions and hard capacity checks, without running an optimization solver on each request. We integrate the policy into vLLM and evaluate controlled image workloads with fixed profiles using Qwen3.5-9B on four H800 GPUs. At 16 requests/s, mean and p99 time to first token decrease by 10.03% and 13.04%, respectively, relative to vLLM's native encoder-cache replacement in the same serving stack. These results show a practical benefit from managing encoder-output retention separately from language generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.