When can test-time search be skipped? Successor-Channel Criticality for Action-Chunked Policies
Abstract
Test-time search can improve action-chunked robot policies, but running it at every decision is expensive and often does not change the outcome. We ask when search can be skipped. Our answer uses the policy's own candidate actions: if every candidate leads to a state with a similar capacity for future control, search has little to choose between. We measure successor-channel criticality as the variance across candidate actions of a future-controllability score. The score summarizes the spread of object positions reached by subsequent policy rollouts at a fixed resolution, and is computed in a resettable simulator without task reward or search labels. We prove that low criticality bounds the value of search, provided that value differences the score does not capture are also small. We evaluate gates with paired per-state interventions on fourteen datasets from three benchmark suites and two policy classes; criticality's ranking lies above the matched random band on all fourteen. Skipping search where estimated criticality is zero avoids about half of decisions while giving up 19.9% of the net benefit, less than the mean of the same score or a cheap rule based on episode time and sampled gripper disagreement. Distilled into a small network and added to that rule, a criticality head raises success on 600 held-out episodes by 4.2 points with 20.5% fewer search calls on a LIBERO image policy, and by 7.0 points at similar call counts on Stack. A network trained on the mean score also improves on that rule; on the image policy, averaged over three head seeds, the criticality head reaches 2.2 points higher success at similar calls when search is scarce and similar success with fewer calls at .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.