acceptodds
Under review as a conference paper at ICLR 2027

SEIS: Self-Evolving Inference Systems

Abstract

The concept of "intelligence per watts" indicates the significance of efficient large language model inference systems, which has been an important research area. However, prior works focus mostly on optimizing certain parts, such as kernels or memory, within the large system. In this work, we take a holistic approach where we apply an agentic self-evolution approach to optimize the whole system end-to-end. Our SEIS (Self-Evolving Inference Systems) autonomously optimizes the entire mini-SGLang engine without any human interventions through iterative sessions with inherited experiences and code changes. Serving Qwen3-0.6B on H100, the resulting engine reaches 3.27x throughput than the original mini-SGLang implementation and beats SoTA engines like vLLM, TensorRT, and SGLang in the single-request workload. The correctness of the optimized inference engine by SEIS is tested in terms of numerical difference and downstream accuracy on math and long-context retrieval tasks. To further understand the mechanisms behind the speedup, we conduct detailed trajectory and final optimized code artifact analysis, which provides significant insight on the evolution dynamics and concrete system optimization techniques. The success of SEIS implies that agentic self-evolution, when done right, can achieve impressive results given a complex objective, and the core challenge might shift to the meta-evolution of the concrete objective given an abstract task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.