Semantic Coverage for Test-Time Reasoning: The Value and Cost of Early Information
Abstract
An additional reasoning trace can be individually valuable while adding little to the work already done. We study test-time allocation through semantic coverage, an objective that values the useful support a set of outputs provides jointly. Predicting this support before completion may avoid redundant work, but consumes computation that could otherwise be spent completing another trace. In a structured overlap model, we show that optimal coordination need not require a full model of possible outputs: learning which branches share useful contributions can suffice. We characterize how many observations are needed to approach optimal coverage in the worst case. In a screening model, we determine exactly when checking whether a candidate will add something new reduces the average computation needed to obtain a useful contribution. The check need not be perfect, and we derive how much it is worth paying for its ability to distinguish new contributions from duplicates. We then show how repeated checking in a discovery model can obtain every distinct useful contribution with high probability using only one full completion for each. Duplicates still have to be encountered, but need not be fully generated; we characterize the minimum total cost of checks and completions up to constant factors. In multi-answer question answering, output format changes the benefit of coordinated generation: naming answers before explaining them increases its coverage gain. In a controlled study using the same preview-based predictions, selecting unfinished traces for their predicted joint contribution recovers more distinct correct answers on average than individual ranking. In a separate comparison of complete procedures with all online work charged, agreement-filtered answer lists achieve higher mean F1 at lower mean token cost than iterative preview selection on retrieved evidence. These results distinguish improvements in selection from improvements in the overall quality–cost tradeoff.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.