acceptodds
Under review as a conference paper at ICLR 2027

When Is Local Validation Cheaper for a Text-Update Composition?

Abstract

Can evaluations of individual text updates validate their composition more cheaply than direct execution? We study both whether local results determine the full gain and how combining their estimates affects sampling cost, for paired binary outcomes and a fixed composition of updates. Under a degree- mean-interaction model, we derive explicit weights that recover the full gain from average gains over local subsets of each size (shell means). With local prices and full-query price , the minimax expected validation cost is . The lower bound covers adaptive mixtures, task reuse and random stopping; a cost-aware local sampler or direct execution attains the rate. We distinguish mean additivity (gains add in expectation) from taskwise additivity (gains add on each task): their singleton-query factors are and , even with identical mean responses. Without a verified structural identity, we correct cached predictions using sampled full outcomes to obtain an exact finite-pool gain interval. On 13 complete public GEPA quartets, additive proxies overstate every full gain, while a fixed retrospective audit reduces new full-composition queries from 3,400 to 2,780 (18.24%) with either branch or clipped predictions. Across 1,000 paired replays, mean query savings remain 15.52–16.06%. These results identify how execution prices, paired structure and cached predictions govern validation reuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.