MIA: Runtime-Reconfigurable Access to Model Internals in LLM Inference Engines
Abstract
As Large Language Models (LLMs) and agents see broader deployment, runtime-time monitoring and intervention become increasingly important. This requires an inference engine to accept diverse requests asking for different layers, tokens, and vectors within one batched step, while meeting the latency targets. However, existing inference engines, which are primarily optimized for high-throughput text generation, are not designed to handle such heterogeneous request efficiently. Recent systems narrow this gap, but each pays for it by trading either throughput, write access, model support, or internal types. Therefore, we present MIA, a plugin for native vLLM server that captures and steers model internals at near-native efficiency. MIA consists of two primary parts, a Scheduler and a Courier. Before each step, the Scheduler decides what each request should read or write based on the batch and turns that decision into a routing entry. The GPU then replays the CUDA graph that reads and writes following those entries with minimal overhead. The Courier moves what the graph produced off the GPU in the background, returning it together with the model's own output. Under open-loop arrivals with requests asking for different layers, MIA generates each output token 2.9-14.1 faster than existing systems that support per-request hidden state capture during serving from 1 to 16 requests per second.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.