VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
Abstract
For years, we have built LLM serving systems like any other critical infrastructure, with a single general-purpose stack hand-tuned over many engineer-years to support every model and workload. In this paper, we take the opposite bet and let coding agents automatically synthesize bespoke serving systems for different usage scenarios. We propose VibeServe, the first agentic loop that generates entire LLM serving stacks end-to-end. VibeServe uses an outer loop to plan and track the search over system designs, and an inner loop to implement candidates, check correctness, and measure performance on the target benchmark. In the standard deployment setting, where existing stacks are highly optimized, VibeServe remains competitive with vLLM, showing that generation-time specialization need not come at the cost of performance. More interestingly, VibeServe outperforms existing systems in six non-standard scenarios by exploiting opportunities that generic systems miss. These scenarios involve non-standard model architectures, workload knowledge, and hardware-specific optimizations. Together, these results argue that infrastructure software should favor generation-time specialization over runtime generality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.