acceptodds
Under review as a conference paper at ICLR 2027

FlexBI: Regaining Execution Flexibility for Batch-Invariant LLM Inference

Abstract

Batch-invariant (BI) inference preserves bitwise-identical results for each request regardless of its surrounding batch. Existing BI implementations couple this numerical guarantee to fixed physical schedules, limiting their ability to exploit changing concurrency and prefix sharing. FlexBI regains execution flexibility by separating a fixed request-local numerical contract from adaptive physical scheduling. Flextile regains dynamic GEMM output tiling by keeping each output's reduction trace invariant across tile changes. Flexload uses partition-invariant sharing to share identical physical key-value (KV) page loads while preserving each query's logical reduction. In online serving across two models and two datasets, FlexBI delivers 25–33% higher output throughput and 16–32% lower P99 time to first token than same-engine BI in vLLM and SGLang.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.