A2-CAP: A Causal Attribute–Affordance Graph Harness for Frozen VLA Manipulation
Abstract
Long-horizon robotic manipulation requires reasoning about which object properties enable skills and whether their intended effects occur. We introduce **A2-CAP**, an inference-time harness around a frozen vision-language-action (VLA) policy. Its registry of 12 mechanism-level strategies, plus an LLM fallback, defines a causal attribute–affordance graph whose directed edges encode preconditions and effects. **PiCausal** grounds video evidence in this graph; **CAP-Plan** selects skills, verifies their outcomes, and handles failures without updating the policy. We introduce **A2-CausalHOI**, 1,305 hand–object interaction videos spanning 85 object categories and 256 tasks, with registry-derived target graphs. On a 100-clip graph-grounding sample, PiCausal raises edge F1 from to over a matched VLM-only baseline. On 50 scored EPIC-KITCHENS-100 segments, gains are to across three backbones; on 14 segments annotated before viewing model drafts, they are to . With pre-generated graphs on two dual-arm tasks, the frozen policy under the harness scores task SR over eight T1 rounds versus over ten rounds for each of two baselines, under a joint execution-and-planning rubric with partial credit. On T2, which reverses the arm roles, it scores task SR over 11 rounds without a baseline comparison.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.