acceptodds
Under review as a conference paper at ICLR 2027

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Abstract

Pretrained robot manipulator policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features only describe local shape without stating where it lies with respect to the robot and the most effective method to deliver this information is unsettled. We propose **Spatial Grafting**, a versatile, lightweight spatial module, which binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting supplies the flow-matching-based action model with metric-grounded spatial tokens, and grafts the resulting spatial tokens by cross-attention into the flow-matching action expert, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs, gaining 11.3% and 15.6% on clean and randomized scenes. It reaches 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks, by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.