An Empirical Ceiling on Per-Request Compute Allocation in Diffusion Serving
Abstract
Interactive generative products run under a hard latency constraint: a tap must return a rendered clip within a service-level objective or the interaction breaks. This makes compute an allocation problem—spend more on requests whose quality still climbs, less on those that have saturated—and a growing body of harness and serving work assumes such allocation is worth doing. Whether it is worth doing is an empirical question about the backend. We measure it. We formalise budget-aware generative orchestration as a deadline-filtered multiple-choice knapsack over per-request quality–compute curves and build a complete reference allocator—online curve estimation, marginal-return allocation, admission control, mid-flight revision, speculative prefetch—not as a system to advocate but as the instrument the measurement requires. Against 7,200 profiled SDXL renders and numeric gates frozen before the first render, we report a ceiling: an oracle given the true per-request curves, an upper bound no online estimator can pass, beats the best static configuration by +0.0118 [+0.0090, +0.0146] delivered quality while spending 3.5× its compute. Per-request allocation is not merely hard to realise here; there is almost nothing for it to win. The realised system behaves as that bound predicts: it is beaten by a single static configuration in all 45 cells of the sweep, by the same configuration in every cell, and in 30 cells a static setting strictly dominates it—more quality and less compute. Cost, by contrast, is separable in the knobs to 0.6% held-out error, so one profiling pass prices any configuration. A pre-registered human study (259 pairs) shows the quality proxy is calibrated, making these facts about the backend rather than the score. The mechanism is distillation: a few-step sampler collapses the useful cost axis to a handful of non-dominated rungs. We give the conditions under which the ceiling should rise, and release the measured tables so the bound can be recomputed on any backend before an allocator is built.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.