acceptodds
Under review as a conference paper at ICLR 2027

Physically Consistent Multimodal Representations for Persistent Target Tracking and Capture with Legged Manipulators

Abstract

Persistent target tracking and capture in unstructured environments require physically grounded multimodal representations that preserve object identity and support interaction under uncertain observations. We present a unified framework that integrates physical consistency into multimodal representation, persistent target reasoning, and coordinated capture for legged manipulators. First, synchronized visual, thermal, and radar observations are organized into a pixel-aligned physical representation for constructing metric-semantic maps and estimating target states. Subsequently, an acquire–maintain–recover mechanism combines semantic compatibility with geometric, motion, and thermal consistency to preserve target identity under occlusion and visual degradation. The resulting representation and target confidence guide coordinated locomotion and manipulation during approach, alignment, and capture. Extensive real-robot experiments cover indoor and outdoor environments, day and night conditions, and four target categories. The framework achieves 91.3% capture success across 480 episodes per method and 84.8% post-occlusion recovery, compared with 76.7% and 61.9% for the strongest baseline. Component ablations further demonstrate complementary contributions from physical representation, semantic mapping, persistent reasoning, and coordinated execution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.