EDGE: RESIDUAL-CONDITIONED CORRECTION ACQUISITION FOR VLA POST-TRAINING
Abstract
A failed robot trajectory is not automatically a useful training example. Under a fixed correction budget, the relevant question is which logged states deserve new expert actions. We study this offline acquisition problem for vision-language-action policies and introduce EDGE (Experience Discovery via Gap-aware Evaluation). A nominally trained, frozen latent world-action ensemble supplies three ranking signals: action-consequence surprise, the local mass of progress-improving action alternatives, and a penalty for uncertain residual transport. The observed transition residual anchors local counterfactuals without fitting the verifier to the target pool. With rollout pool, query count, replay mixture, and policy-update schedule held fixed, EDGE reaches 58.9±0.7% on LIBERO-PRO and 56.2±0.7% on PhysicalShift. A development-frozen G-bandpass center remains about 3pp behind on ten new paired validation seeds. More importantly, a same-snapshot PhysicalShift rollback study holds the learned WAM, verifier, 16 candidate actions, G, and budget fixed while changing only residual transport. Constant anchoring reduces soft-T error and raises downstream success from 52.0 to 55.5%; kernel decay further improves ranking and reaches 56.2% overall, with gains concentrated in friction, mass, damping, and gain shifts and no gain under delay. This factor-wise pattern connects the local transport assumption to where it succeeds and fails. Five-seed mechanism inference remains limited, but the empirical chain now spans counterfactual prediction, acquisition ranking, and equal-budget post-training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.