R2G-VLA: Risk- and Role-Guided Grounding Correction for Robust VLA Policies
Abstract
Vision-language-action (VLA) policies have advanced language-conditioned manipulation, yet their closed-loop robustness remains limited by visual grounding shifts. In real scenes, task semantics may stay unchanged while target positions, receiver locations, object relations, or viewpoints change, inducing visual-token-level distribution shifts in frozen policies. Such shifts can mislead action prediction through distractors or perturbed relational cues, making inference-time grounding correction important for robust VLA execution during deployment. We propose R2G-VLA, a risk- and role-guided grounding correction framework for robust VLA policies under physical scene perturbations. R2G-VLA uses action-token uncertainty as a self-diagnostic signal and propagates it to visual tokens to estimate token-level grounding risk. Instead of intervening on scattered high-response tokens, it aggregates risky tokens into spatially coherent regions that better match object-level perturbations in manipulation scenes. To avoid role-agnostic suppression, R2G-VLA further performs semantic role-aware token correction over target, receiver, support, background, and distractor regions, preserving task-relevant evidence while reducing misleading visual cues. The corrected visual representation is re-injected into the same policy for a second forward pass, and a safety gate accepts the correction only when uncertainty and action-shift constraints are satisfied. Experiments on LIBERO-Spatial show that R2G-VLA improves robustness across five physical scene perturbations, highlighting the effectiveness of risk- and role-guided internal grounding correction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.