acceptodds
Under review as a conference paper at ICLR 2027

Training-Free Adaptive Time Grids for Flow-Based Vision-Language-Action Models

Abstract

Flow-based vision-language-action (VLA) policies generate action sequences by numerically integrating a learned probability-flow ordinary differential equation. Under a limited inference budget, the choice of time discretization can substantially affect control quality, especially after reinforcement-learning (RL) post-training, which may change the velocity field away from the geometry induced by conditional flow matching. We study this problem through a training-free, inference-time perspective. We first separate three quantities that are often conflated: marginal entropy rate, local velocity-field variation, and numerical integration error. Based on this analysis, we propose Entropy- and Geometry-aware Time Grids (EATG), which estimates trajectory-dependent complexity from a calibration rollout and constructs a constrained non-uniform time grid under an explicit inference-cost budget. The method combines entropy-rate proxies with along-trajectory velocity variation, supports fallback to uniform discretization when the signals are unreliable, and requires no parameter updates or auxiliary networks. We evaluate EATG on flow-based VLA policies before and after RL post-training across simulated and real-world manipulation tasks. Our experiments compare entropy-only, geometry-only, fixed-grid, and equal-budget numerical baselines, and analyze robustness across initial noise realizations. The results show when adaptive discretization improves closed-loop success and when reducing ODE integration error alone is insufficient, highlighting the distinction between generative numerical accuracy and task-level control quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.