Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models
Abstract
A generative policy acting in a closed loop must decide what action to produce, how much inference to spend producing it, and when to observe again. Flow-based vision-language-action (VLA) policies typically fix the latter two choices across an entire task, even as the value of computation and feedback changes across states. We formulate each policy query as a variable-duration decision over latent input, denoising steps, and execution length, and introduce Elastic Queries Reinforcement Learning (EQRL) to learn these choices while keeping the VLA frozen. A lightweight actor jointly learns latent steering and two schedule residuals around a difficulty-guided reference, using duration-aware query-level RL with a number-of-function-evaluations (NFE) objective. A local compute–feedback analysis distinguishes schedules with equal amortized inference cost. On the long-horizon LIBERO-10 task of putting both moka pots on the stove, EQRL raises five-seed success from 0.411 to 0.861 while reducing requested NFE from 0.500 to 0.441. On four LIBERO-90 tasks and ALOHA Cube Transfer, EQRL improves aggregate success and learning AUC while reducing requested NFE by 39.4% and 14.0% relative to DSRL. Scheduling ablations support the contributions of joint control and difficulty guidance, and physical-robot evaluations match or exceed DSRL's success while reducing inference cost and execution time across five tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.