acceptodds
Under review as a conference paper at ICLR 2027

Preserving How Across Objects: Execution-Conditioned Bimanual Contact Grounding

Abstract

Most cross-object interaction models learn where an action can be performed on a new object. Describing an interaction by its action makes this knowledge transferable, but can obscure how two hands carry it out: different executions of the same action assign different roles to the hands and may require different contacts. We represent an execution by its action and the functional roles of both hands, independently of the target object. We learn this representation from the visual features of regions contacted under the same execution across training objects. Given a new target image, we use the learned representation directly as the grounding query to predict where the queried hand should make contact. We evaluate on unseen instances of GigaHands training categories, in external transfer to HOI4D, and through execution contrasts that change hand roles while fixing the object, action, and queried hand. Our method achieves the highest mean localization scores and better distinguishes executions that require different contacts, with a score of 26.4% versus 18.8% for a non-learned prototype baseline and 12.0% for the strongest learned baseline. It also reaches 21.8% versus 10.1% for that baseline on role combinations never seen in training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.