acceptodds
Under review as a conference paper at ICLR 2027

From 201 Bins to 1 Bit: What a Binary Advantage Interface Retains in Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models are promising for general robot control, yet advantage-conditioned post-training can use a fine-grained categorical value critic even when the actor receives only a binary condition. We study what information this interface retains. Equality of the induced observation-action-label law is sufficient for equality of the supervised actor objective. For a mean-based residual, equal critic means suffice but are not necessary; we also bound objective changes under value and empirical-quantile perturbations. A covariance identity characterizes when ideal positive conditioning improves a chosen action utility. Direct binary prediction has a Bernoulli posterior as its Bayes target. In an idealized finite-alphabet model, mean-estimation sample complexity is independent of bin count, unlike full categorical recovery. Guided by this analysis, we train a scalar bootstrap score, cache recorded residuals in a Persistent Advantage Buffer, and fine-tune a prompt-conditioned flow-matching actor with classifier-free guidance. Physical evaluations on four bimanual manipulation tasks illustrate the pipeline against a shared behavioral-cloning initialization. The results separate sufficiency at the actor interface from the challenge of learning useful labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.