acceptodds
Under review as a conference paper at ICLR 2027

Measure, Don't Average: Small-Cohort Measurement of VLA Policies

Abstract

Vision-language-action (VLA) studies often compare only four to six policies, yet report benchmark means without testing whether items, ranks, and uncertainty are identified. We audit a 3,155-episode matrix over six variants and a 14,412-episode LIBERO-Plus spatial bank. Standard 2PL estimation converges while item discrimination is prior-sensitive at . A hierarchical Bernoulli model grounded in observed rollout structure predicts unseen initializations (log-loss .1865 versus .1978) and has the best unseen-task point loss (.2039 versus a predeclared .2596 baseline, but not separated from a .2084 suite mean). Under a common observed-compatible data-generating process, nominal-90% measurement-replication coverage is .912 for the hierarchy, .898 for 1PL, and .931 for 2PL; the hierarchy has the best interval score but is not uniformly better calibrated. An equal-12,800-observation control reverses our earlier policy-count explanation: discrimination recovery is .764 with repeats and .731 with . For acquisition, we implement outcome-enumerated joint weak-rank information gain and compare matched atomic task queries on independent future matrices. It beats uniform, but does not robustly beat task EIG or Fisher: 25% rank-difference intervals include zero on both pre-specified data-generating processes, no method has a median sustained- budget within 50%, and the joint method triggers no model-based stops. Architecture-DIF tests also remain negative and partly unidentified. The contribution is validated small-cohort measurement with transparent negative boundaries, not generic adaptive or certified robot evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.