acceptodds
Under review as a conference paper at ICLR 2027

When Is Test-Time Allocation Actionable? Exchangeability and Breadth Dominance in Hierarchical Generation

Abstract

Retrospective oracle gains and accurate value prediction do not by themselves establish that adaptive test-time allocation can improve generation. We audit when information that distinguishes actions becomes available. Under joint conditional exchangeability of future child observations and costs, unseen child labels reduce to one expand action per parent. With hidden fresh randomness and no candidate observation before spawn, policies cannot select a particular unseen parent realization. These conditions restrict action identity but do not establish breadth optimality. In a frozen two-stage TRELLIS generator, value models achieve pooled Spearman correlations of 0.935–0.970, while only 0.108–0.534% of retrospective target variation lies within decision sets. For the primary input-silhouette proxy, the empirical within-parent sum-of-squares share is 4.038% and the ANOVA one-child ICC is 0.959. Three development attempts fail to improve on the selected heuristic comparator. A one-shot 80-unit test of a common spawn-first controller with learned branch selection yields a normalized-utility gain of 0.004281 (95% CI [0.000903, 0.008504]) over uniform branching. Raw silhouette-Chamfer improves by 0.409%, with an interval including zero. A zero-parameter rule attains 92.2% of the learned gain's point estimate; the incremental benefit of the 76-feature Q model remains unresolved. These are finite-table replay results under an estimated cost contract, not end-to-end latency or perceptual quality improvements. The study provides an actionability audit and evidence of small tested-policy gains in one breadth-dominant hierarchy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.