What Does a Skill Vocabulary Cost? A Closed-Loop Decomposition over Demonstrated Motions
Abstract
Discrete skill policies execute motions drawn from a finite vocabulary. In most current methods this vocabulary is a latent codebook, learned jointly with a selector and a decoder and validated by offline reconstruction. When such a policy fails in closed loop, the failure cannot be traced to the vocabulary, to the choice of skill, or to its execution. We present Movement Prototype Quantization (MPQ), which separates these costs by placing the vocabulary after the policy instead of inside it. The library stores demonstration segments as B-spline control points. At execution time, MPQ replaces each predicted action chunk with its nearest library element and corrects that element toward the prediction, within a bound given by the spread of the demonstrations themselves. MPQ trains no parameters. Across six architectures in simulation, including a 7B vision-language-action model, the wrapped policy is statistically equivalent to the original. Its success changes by +0.6 points on average, less than the spread between repeated evaluations of the same checkpoint. A reference policy executed through the library therefore sets a ceiling for any selector over the same library. Replacing the reference's components with those of a trained selector, one at a time, then attributes the selector's shortfall to its parts. On LIBERO-90 with 256 prototypes, quantization costs nothing and the selector's choice costs 2 points, whereas its learned heads lose 15, nearly the entire gap. Offline reconstruction predicts none of these differences, and it misjudges library size, library design, and transfer across benchmark suites.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.