AHR: Amplitude and Horizon Refinement for Vision-Language-Action Manipulation
Abstract
Vision-language-action (VLA) policies provide strong semantic grounding and coherent motion proposals for robotic manipulation. However, small execution errors can prevent these capabilities from translating into reliable task completion. We introduce Amplitude–Horizon Refinement (AHR), which improves the execution of frozen VLA policies by jointly learning how to correct their action proposals and how long to execute each correction before re-observation. The key insight is that the intended correction horizon can change the appropriate correction from the first command. AHR therefore represents each refinement decision as a correction unit: a horizon-conditioned residual sequence paired with its execution length. Given RGB observations, proprioception, and a base proposal, a lightweight actor generates bounded residual candidates and selects a complete unit, allowing different corrections to overlapping commands across horizons. An off-policy actor–critic objective jointly optimizes residual generation and horizon selection using shared unit-level value estimates, duration-aware semi-Markov targets, and mixed demonstration and online replay. With a frozen GR00T N1.6 base policy, AHR achieves average success rates of 96.1% on RoboMimic-Pixel and 96.8% on DexMimicGen, outperforming the compared methods on both benchmarks. Across two real-world tasks, AHR achieves a 95.6% average success rate,compared with 65.0% for the frozen base policy. AHR thus provides an effective framework for refining frozen VLA policies by jointly learning action corrections and correction horizons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.