HERA: Hierarchical Response-Aware Supervision Allocation for Efficient VLA Fine-Tuning
Abstract
Vision-language-action (VLA) models provide a foundation for robotic manipulation. Fine-tuning on demonstrations collected in target environments enables adaptation, but can require lengthy training. One source of this inefficiency is repeatedly training on demonstrations with little additional benefit at the current learning state. Existing data-selection methods concentrate training on useful supervision through subsets selected by behavioral structure or estimated utility. Yet fixed subsets cannot track changing utility: retained samples may become less useful, while excluded samples remain unavailable when their value grows. Our key insight is that training utility evolves at two coupled granularities: learning from one action chunk can change the value of physically related chunks, while the remaining benefits within a chunk can evolve differently, or oppositely, across its action subspaces. These effects depend on trajectory structure and the current learning state. We introduce HERA, a HiErarchical Response-Aware Supervision Allocation framework for flow-based VLA fine-tuning. A shared response estimator combines physical correspondences with the current learning state to predict changes to a fixed training-side reference objective, sharing sparse probe responses across related chunks. Dynamic chunk sampling adapts training frequencies while keeping all chunks eligible. Dynamic action-subspace weighting assigns normalized, chunk-specific weights over native action subspaces. Both allocations are jointly optimized and periodically refreshed. On LIBERO, HERA reaches 80% success with 53.6% fewer optimizer updates than LoRA, yielding a 1.93× wall-clock fine-tuning speedup including response-estimation and allocation overhead. At 20K updates, HERA improves average success rate over LoRA by 9.38 and 10.75 percentage points on LIBERO and PiPER-X, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.