RN-VLA: Reference-Point Normalized Vision-Language-Action Models for Sub-Centimeter Fine-Grained Manipulation
Abstract
Fine-grained robotic manipulation—such as pressing a specific key on a dense keyboard or performing micro-adjustments on household dials—remains challenging, as moderate per-step errors compound exponentially over long-horizon sequences. Monolithic Vision-Language-Action (VLA) models struggle here due to instance-level semantic ambiguity, geometric attention dispersion, and coordinate scale mismatch between reaching and fine manipulation. We propose RN-VLA, which grounds task-specified targets via Semantic Query Grounding (SQG), suppresses background clutter through FiLM-based Cross-Modal Gated Modulation (CMGM), and re-centers the action space onto the target part via Reference Point Normalization (RPN), enabling a diffusion policy to achieve sub-centimeter precision. Experiments in Isaac Sim and on a physical xArm6 robot show RN-VLA achieves over 80% atomic success across seven household tasks, outperforming state-of-the-art VLA and 3D diffusion baselines, with strong robustness under out-of-distribution placements. By eliminating atomic-level errors, reliable single-step policies compose zero-shot into long-horizon tasks—without LLM planning or online verification—achieving up to a 68% relative improvement over the strongest baselines (up to 1.68× higher success rate).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.