RouteWAM: Hierarchical Routing of Heterogeneous Future Representations for Robot Manipulation
Abstract
Predicting future representations can provide vision-language-action policies with useful cues for robot manipulation, but the relevance of visual, geometric, and motion information varies across tasks and execution stages. We introduce RouteWAM, a hierarchical mixture-of-experts policy that adaptively selects and combines heterogeneous future representations. RouteWAM organizes six future-prediction experts into visual, geometric, and motion groups, separating the selection of alternative representations within a group from the combination of complementary information across groups. For each conditioning token, within-group routing produces one candidate per group, and a cross-group router selects and fuses two candidates to guide action generation. Both routing levels score candidates using the current observation and instruction together with each candidate’s own predicted future content. The model jointly optimizes future prediction and action imitation, while a routing curriculum gradually transitions from dense fusion to sparse selection. At inference, RouteWAM requires only current observations, language instructions, and robot proprioception. Experiments on LIBERO and RoboCasa show that RouteWAM outperforms the compared methods across all four LIBERO suites and both RoboCasa task categories. It achieves 92.5% success on LIBERO-Long, improving over DreamVLA by 3.0 percentage points, and reaches 75.2% and 85.6% on RoboCasa pick-and-place and articulated-object tasks, exceeding DIAL by 6.3 and 11.3 percentage points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.