acceptodds
Under review as a conference paper at ICLR 2027

What Should a VLM Tell an Action Expert? Comparing Conditioning Interfaces for Long-Horizon Furniture Assembly

Abstract

Translating the semantic understanding, task decomposition, and visual grounding capabilities of vision–language models (VLMs) into continuous robot actions typically requires a conditioning interface that conveys upstream predictions to an action expert. However, systematic evidence remains limited on what information this interface should contain and whether downstream execution can tolerate errors in actual VLM predictions. We use multitask, long-horizon furniture assembly as a diagnostic setting. This domain requires switching across tasks and manipulation stages, localizing parts and their targets, and executing contact-rich skills such as insertion and screwing, thereby amplifying and exposing the downstream effects of interface design and upstream error at the complete-task level. We compare conditioning interfaces constructed from different combinations of three information types: spatial targets, gripper-action cues, and end-effector rotation. Based on this comparison, we recommend the Target–Action Guidance Point (TAGPoint) as an intermediate representation: it encodes a coarse gripper-action cue in the color of a target point, integrating spatial and action information within a unified, image-aligned interface. We further train the action expert with noisy guidance to mitigate performance degradation under inaccurate upstream predictions. Our evaluation covers clean FurnitureBench assembly, guidance-noise training and evaluation, end-to-end execution driven by actual VLM predictions, joint FurnitureBench–AutoMate training, and physical-robot execution. Together, these results provide an empirical basis for selecting measurable and intervenable intermediate representations in VLM–action expert systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.