Quality-Aware Cross-Model Computation Reuse
Abstract
An intermediate result computed by one model can be reused by other models to perform their tasks. Existing work mainly focuses on practical execution, leaving a theoretical gap in optimizing reuse decisions. This optimization faces two challenges: quality uncertainty, because the effect of reuse on task quality is uncertain across models, and coupled scheduling, because tasks need to share the cost of preparing reusable results. These challenges compound each other: quality must be learned online, but the shared preparation structure makes scheduling NP-hard even with known quality, breaking the key assumption in existing methods. We formulate cross-model computation reuse as an online decision problem and develop the Quality-Aware Reuse Scheduling (QARS) algorithm to address it. For quality uncertainty, QARS learns task-dependent reuse quality from selected, possibly delayed feedback and uses optimistic estimates to guide decisions. For coupled scheduling, it jointly chooses which results to prepare and which tasks should use them, adapting scheduling accuracy to the remaining quality uncertainty. For the considered problem, our analysis separates learning and optimization error in regret and quantifies the tradeoff between scheduling accuracy and computation. Completing the quality-aware stopping rule yields regret while preserving feasibility. Experiments demonstrate the effectiveness of QARS in optimizing cross-model reuse, reducing the combined cost of computation and quality loss by up to 18.0%, and mean regret by 63.9% over the strongest scheduling baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.