APAS-VLA: Adaptive Physical Action Scaling for Faster and Reliable Robot Execution
Abstract
Physical execution of standard Vision-language-action (VLA) is limited by the robot demonstration speed, because there is a one-to-one mapping between action token and future demonstration step. To break this implicit coupling, we introduce APAS, an adaptive action-generation framework that allows each action token to represent a different amount of progress along the demonstration trajectory. APAS combines a lightweight Token-Speed Predictor, which predicts an observation-dependent token-speed profile, with a Profile-Conditioned Action Expert, which directly generates actions under the predicted profile. To adapt actions with different precision requirements, we develop an offline teacher to assign larger token speeds for stable motion and smaller token speeds for precise interactions, and then distill this knowledge into a lightweight predictor for fast deployment. On LIBERO, APAS retains near-base success (96.25% versus 96.85%) while reducing successful-episode execution from 148 to 86 control steps. On CALVIN, it improves average completed sequence length from 3.887 to 4.096 while reducing successful-subtask execution from 88.96 to 55.89 steps. On the real robot, APAS reduces successful-trial completion time by 32.3–49.0% while matching or exceeding Vanilla's success rate across three evaluated conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.