acceptodds
Under review as a conference paper at ICLR 2027

PipeSearch: Buying Composition Labels Only Where They Change the Pipeline

Abstract

Serving a large language model on a single accelerator asks quantization, pruning, low-rank factorization and cache compression to share one memory and throughput budget, so what is deployed is an ordered composition of methods rather than a single choice. Choosing that composition needs the pairwise interactions of the methods, yet they have to be bought: a pairwise label costs a full compression pass, reversing two methods yields a different model, and some pairings refuse rather than degrade. We propose PipeSearch, which buys a pairwise label only while it can still change the returned pipeline, that is, while the pipelines of its optimistic and pessimistic completions disagree. Decision-relevant acquisition observes the cheapest unobserved pair either completion applies, order-aware prefix search carries which method acted last and keeps a sequence apart from its reverse, and feasibility maintenance excludes a refused pair at expansion time. The key insight is that the observations which reveal the interaction structure are part of the decision rather than a preliminary to it. Across catalogues of four to twelve atoms, PipeSearch returns the pipeline of complete offline labelling while building 0.201 to 0.500 of the ordered pairs, at 75.1 accelerator-hours against 367.4 at . Paying for every candidate pair instead leaves perplexity, zero-shot average, peak memory and decode rate unchanged at both memory budgets, at 15.0 and 3.6 times the search cost. Under a cap of the fp16 footprint, where no single method fits, PipeSearch fits Llama-3.1-8B at a predicted 8.79 GB and 9.59 perplexity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.