acceptodds
Under review as a conference paper at ICLR 2027

When to Launch Inference for Action Chunking under Latency?

Abstract

Policies parameterized by expressive generative models promise a powerful interface for real-time decision making. To boost performance and promote execution smoothness, these models predict a chunk of future actions in a single inference call. Yet, the non-negligible inference latency introduces a delicate trade-off. Early inference may prevent the model from reacting to critical information revealed shortly afterward; late inference suffers from staleness before the next action chunk prediction takes effect. Existing work mainly improves the chunk prediction policy itself, but sidesteps the equally essential question: when to launch inference under latency? This paper bridges this gap by formulating inference timing as a meta control policy in a semi-MDP induced by the original decision-making problem. We show that value can vary non-monotonically with inference timing, illustrate its consequences in a hidden-danger example, and adapt a fitted -iteration algorithm to learn the meta control policy. Theoretically, we prove a performance bound that disentangles sources of sub-optimality, notably including action chunk generation quality and fitted estimation error. Empirically, experiments on Kinetix and DOM span different environments and action-policy scales. Adding learned launch timing improves task success while reducing action-model calls compared with the same action generators under their default launch schedules, with larger gains on harder tasks and abrupt dynamic adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.