When to Launch Inference for Action Chunking under Latency?
Abstract
Policies parameterized by expressive generative models promise a powerful interface for real-time decision making. To boost performance and promote execution smoothness, these models predict a chunk of future actions in a single inference call. Yet, the non-negligible inference latency introduces a delicate trade-off. Early inference may prevent the model from reacting to critical information revealed shortly afterward; late inference suffers from staleness before the next action chunk prediction takes effect. Existing work mainly improves the chunk prediction policy itself, but sidesteps the equally essential question: when to launch inference under latency? This paper bridges this gap by formulating inference timing as a meta control policy in a semi-MDP induced by the original decision-making problem. We show that value can vary non-monotonically with inference timing, illustrate its consequences in a hidden-danger example, and adapt a fitted -iteration algorithm to learn the meta control policy. Theoretically, we prove a performance bound that disentangles sources of sub-optimality, notably including action chunk generation quality and fitted estimation error. Empirically, experiments on Kinetix and DOM span different environments and action-policy scales. Adding learned launch timing improves task success while reducing action-model calls compared with the same action generators under their default launch schedules, with larger gains on harder tasks and abrupt dynamic adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.