Complementary Goal Grounding via Prototype Convolution for OmniGoal Navigation
Abstract
OmniGoal Navigation requires a single policy to locate goals in different forms, such as categories, images, or instance descriptions. Existing OmniGoal approaches primarily match heterogeneous goals with current observations, while navigation experience associated with different goal forms remains largely unmodeled. However, experience across goal forms can provide complementary cues for locating the same or related concepts, such as appearance from image goals and detailed textual descriptions from instance goals. To incorporate such complementary experience into a unified navigation policy, this work introduces Prototype Convolution(ProtoConv). A prototype graph organizes prior navigation experience into multi-form concept prototypes and their spatial relations. Given a goal, ProtoConv retrieves goal-conditioned prototype kernels and convolves them with an online semantic graph constructed from current observations. The convolution combines cross-form cues that transfer appearance and descriptive attributes for the same concept with cross-concept spatial correlations that support target localization through related observations. By integrating these cues, ProtoConv guides navigation with complementary experience across goal forms and concepts. Evaluations on HM3D, MP3D, and GOAT-Bench demonstrate consistent effectiveness across ObjectGoal, ImageGoal, and Instance Navigation without form-specific policy design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.