acceptodds
Under review as a conference paper at ICLR 2027

Stage-Wise Deterministic Inference for Vision-Language-Action Models

Abstract

Batched inference improves the efficiency of vision-language-action (VLA) models in shared serving, parallel evaluation, and policy-learning rollouts. However, changes in batch context can affect tensor shapes and kernel selection, altering the order of floating-point accumulation. Action outputs can therefore differ even when request inputs and sampling noise remain fixed, leading to divergent robot trajectories and changes in task outcomes. To address this reproducibility issue, we propose stage-wise deterministic inference (Stage). Stage groups and pads requests at independently selected fixed widths for vision encoding, language encoding, and action generation, allowing each stage to balance batching efficiency and padding overhead. This organization supports autoregressive decoding, parallel regression, and iterative generation, and extends to dynamic batching with fixed stage widths and per-request noise. Experiments with four VLA models show that Stage maintains byte-identical action outputs across the tested batch contexts. Stage achieves to LLM-42 throughput at batch size and to the throughput of the standard batched inference implementation (Native) with stage widths selected for the target batch size, while maintaining comparable task success rates in closed-loop evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.