acceptodds
Under review as a conference paper at ICLR 2027

Allocating Additional Labels Between Cascade Predictors and Selectors

Abstract

A useful learned selector does not establish that the next labels should train the selector rather than the predictors. Controlled image-cascade experiments separate these questions. In an adapted CIFAR-100 long-tailed condition, class-conditioned routing improves on a strong scalar rule by 2.206 percentage points (pp), but expanding gate support contributes only 0.200pp. Allocating the identical additional labels to the predictors instead improves the source-selected cascade by 0.881pp; two additional paired training repetitions give 1.612pp. Tiny ImageNet preserves the investment advantage, including after SOURCE-selected answer correction. A complete endpoint-support × sampling-prior factorial identifies a 5.901pp gain from distinct examples at the expanded prior, rising to 7.088pp under a separately fixed fourfold training horizon. Prior-only and routing-interaction contrasts remain unresolved. Component controls connect the support effect to an executable alternative: source-selected upgraded cheap-only prediction exceeds the old Tiny cascade by 2.965pp. Corrected Tiny and stronger CIFAR predictors retain fixed-budget investment gains without crossing the corresponding old-pair hindsight ceilings. When corrected budgets are also SOURCE-selected, cheap-only prediction has lower resident request time, but its same-label accuracy advantage remains unresolved. Thus useful routing, the return to further gate labels, and the answers made available by predictor training are distinct quantities. Reported intervals condition on fitted models and exposed images; the comparisons do not establish a universal allocation rule or a lifecycle speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.