Select to Succeed: Pareto-Consistent Action Selection for VLAs via Augmented Tchebycheff Scalarization
Abstract
Scaling test-time computation in Vision-Language-Action (VLA) models necessitates robust arbitration over stochastically generated candidate trajectories. We formalize test-time candidate selection as a finite-set multi-objective optimization problem under a receding-horizon control framework, introducing S2S-VLA (Select to Succeed) for frozen VLA policies. Specifically, we map candidate action prefixes into a predefined cost space comprising the marginal probabilities of task failure and local anomalies, alongside an analytic metric of command roughness. To systematically resolve objective trade-offs, we construct the selection rule via an augmented Tchebycheff scalarization and prove that its exact minimizers are nondominated within the evaluated candidate set in the predicted objective space. To align a lightweight neural evaluator with this selection criterion, we optimize a theoretically compatible dual-objective formulation: a pointwise binary cross-entropy loss anchoring marginal event probabilities to empirical offline branch rollouts, and a pairwise logistic loss enforcing scalarization-induced soft pairwise preferences. Following temperature scaling for post-hoc probability calibration, S2S-VLA deterministically minimizes the scalarized score, enabling mathematically consistent, rollout-free, and computationally efficient test-time scaling for embodied agents. Extensive experiments in simulated and real-world environments demonstrate substantial improvements in task success over baseline methods. Our project is available at S2S-VLA.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.