Dual-Role Geometry, Contextual RGB: Representing Point Clouds for 3D Robot Policies
Abstract
Colored point clouds provide robot policies with spatial geometry and appearance cues, whose relative importance varies across manipulation tasks. Jointly encoding XYZ and RGB at each point can reduce performance on tasks requiring precise spatial execution, while separate pointwise encoding still lacks explicit interactions among neighboring appearance features. We propose GeoRGB, a 3D robot policy that preserves pointwise spatial references while organizing appearance features through geometry-guided interactions. Its asymmetric dual-tower encoding (ADTE) pairs a dedicated pointwise XYZ pathway with a multiscale RGB hierarchy, separating spatial reference encoding from appearance contextualization before fusion. Within the RGB hierarchy, geometry-guided RGB contextualization (GRC) aggregates neighboring appearance features within and across scales, using XYZ-defined neighborhoods and directional weights derived from relative positions. Across 50 RoboTwin 2.0 tasks, GeoRGB achieves 77.5% average success, improving over the strongest evaluated baseline by 14.3 percentage points. It further achieves 93.8% success across four real-world tasks. On MetaWorld, a deeper variant reaches 75.0% success, matching the best-performing baseline with fewer than half the parameters and less than one-quarter the training time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.