The Right Kind of Disagreement: Dynamic Collaborator Pruning for Speculative Token-Level Collaboration
Abstract
Speculative token-level collaboration reduces collaboration overhead by having a main model generate candidate tokens while auxiliary models verify them in parallel. However, existing methods typically preselect collaborator models and keep them active throughout the generation process, so verification costs accumulate across decoding rounds. We propose CoStay, a disagreement-driven selective-exit framework for token-level collaboration systems. Guided by the principle that collaboration requires a balance between consensus and informative disagreement, CoStay uses disagreements between the main and collaborator models during token-level decisions as an online signal. When multiple collaborators participate, aggregate disagreement cannot be attributed to any single collaborator. CoStay therefore applies leave-one-out analysis to quantify each collaborator’s marginal impact on the aggregated decision and uses a two-sided confidence-bound rule to determine when a collaborator should stop participating in subsequent rounds. Across six benchmarks and eight model configurations, CoStay achieves 99.12% of the mean accuracy of SAFE, CoSD, and UniTE, with a gap of only 0.64 percentage points, while reducing the average collaborator token cost by 44.43%. In end-to-end runtime experiments, this reduced collaboration overhead translates into a 40.8%–57.3% throughput improvement over full-participation collaboration methods. The anonymized source code is available at https://anonymous.4open.science/r/CoStay-0732/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.