Where Refinement Pays: Measuring and Allocating Sampling Compute in Flow-Matching Manipulation Policies
Abstract
Flow-matching manipulation policies typically use the same number of function evaluations (NFE) at every call, assigning the same sampling budget to free-space motion and precision insertion. Does each decision benefit from the same amount of computation? Whole-episode comparisons cannot isolate the value of one budget choice because changing the budget also changes the states visited afterward. We introduce counterfactual state branching, a testbed that restores exact simulator states and policy histories, matches sampler noise, and varies the budget at selected calls. All branches then use the same continuation budget. We compare terminal success and cumulative NFE from the intervention to termination. Across four flow-matching policies on nine VLA-Arena suites, budget responses vary strongly with the policy and decision state. Using more computation throughout an episode often fails to improve success and can reduce it. We use these paired success and cost outcomes to train FlowBudgetAllocator (FBA), which operates across suites while keeping the action policy frozen. In simulation, it achieves higher observed success with fewer NFE than the default high budget. On basic configurations, it improves task success over the best fixed budget with only a modest increase in NFE. FBA exhibits distinct budget-selection patterns across tasks and execution phases. On the real robot, it uses substantially lower sampling budgets than the default high-budget policy without reducing observed task success. By measuring the consequences of individual budget choices, we turn sampling computation into a resource to allocate during control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.