DGA-VLA: A Lightweight Depth-Guided Grasp Enhancement for VLA Models without Retraining
Abstract
Vision-language-action (VLA) models provide a unified policy foundation for a wide range of robotic manipulation tasks. However, their grasping performance remains constrained by imprecise gripper-object alignment at closure, leading to empty or unstable grasps and task failures. To address this limitation, we propose DGA-VLA, a lightweight depth-guided grasp refinement framework that augments pretrained VLA policies without VLA architectural changes or policy retraining. DGA-VLA couples the VLA's global motion generation with a wrist-depth feedback controller, enabling local correction without changing the underlying policy. When the policy indicates a closure intent, the controller uses wrist-depth feedback to estimate local gripper-object offsets and apply bounded corrective motions before closure. Experiments demonstrate that DGA-VLA improves average success rates on grasping tasks by 12.0 percentage points on RoboTwin, 8.4 points on SimplerEnv, and 14.6 points on real-robot pick-and-place tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.