CAP-7: Binding-Aware and Commitment-Complete Verification for Task-Level Affordance Reasoning
Abstract
Task-level affordance reasoning requires vision-language models to turn images and manipulation instructions into plans that connect objects, actions, states, and motions. Yet evaluating these elements separately leaves two failures unresolved: plausible elements may be connected incorrectly, which we term , and required steps may be omitted without being penalized, which we term . To expose incorrect bindings, we introduce CAP-7, a typed program representation linking objects and roles to actions, states, dependencies, and motions. We construct CAP-7-PD for supervision and a extension for community research. Controlled changes to roles, action order, and motion paths test whether the resulting plans preserve the relations required by the task. To detect omissions, we develop commitment-complete verification that retains missing requirements in the score and penalizes unsupported additions. Combined with a check of whether the submitted action order executes legally, this verification provides a training reward. Rescoring unchanged model outputs reveals misleading gains under the original verifier. Under matched training conditions, replacing only the reward raises the mean share of sampled orders judged legal by an independent symbolic engine from 51.3% to 94.4%, compared with 85.2% for the shared supervised initialization; unmatched step pairs also decrease. These results show that complete task verification can improve the structure and symbolic executability of learned manipulation plans.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.