Arrival-Conditioned Voting Power in Heterogeneous Language Model Ensembles
Abstract
Large language model (LLM) ensembles can reduce latency by voting before every response finishes. With heterogeneous models, completion order also determines whose responses count. We derive how this selection changes the response distribution and final decision, separating changes in model participation from selection within each model. Across two disjoint five-model panels spanning seven benchmarks, voting on the first three responses lowers selected-response accuracy by 5.09 and 11.84 percentage points. Changes in model participation explain most of both losses, and exhaustive subset controls distinguish arrival selection from the effect of using fewer votes. Changing delivery times alone realizes all 128 predicted decision changes with fixed responses. We then derive a completion-invariant rule: commit when every possible completion gives the same full-panel decision. This rule matches all 672 full-panel decisions while reducing mean latency by 40.83% and 22.49% on the two panels. Completion order thus changes effective voting weights, but some latency savings remain available without changing the full-panel decision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.