Critical Points for LLM Admission
Abstract
Shared large language model (LLM) services must use limited capacity efficiently without repeatedly postponing certain workloads. Admission controllers can balance request value against service debt, the service owed but not yet delivered, by ranking requests with a weighted score and allocating resources in that order. However, fixed weight grids can repeat allocations while missing useful request orders in narrow weight intervals. We present TERRA-CP (Temporal Entitlement Ranking with Critical Points), an admission controller that searches queue-specific score crossings, where request rankings change. It combines explicit crossing enumeration for small queues with implicit search for large queues, constructs resource-feasible candidates, and selects the one with the smallest predicted maximum service debt. A welfare reserve constrains cumulative estimated utility relative to greedy allocations on each current queue. We prove coverage of the ranking-and-packing family under uncapped exact-arithmetic search and characterize fixed-grid omissions and finite-budget loss. Extensive experiments demonstrate that TERRA-CP reduces maximum cumulative deficit by 14.4% relative to a runtime-matched grid and by 29.0% relative to independently tuned, reserve-adapted drift-plus-penalty scheduling. In six-tenant vLLM experiments, it reduces completion deficit by 11.6% at moderate load and mean response time by 4.85% at high load relative to welfare-constrained token-counter scheduling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.