acceptodds
Under review as a conference paper at ICLR 2027

SelectionVLM: Enabling VLMs to Interact Programmatically with CAD

Abstract

Modifying CAD geometry programmatically requires first identifying the entities to manipulate and then passing references to them to the appropriate functions in the CAD system's API. Vision-language models (VLMs), given only a canvas screenshot, can perceive these entities but cannot directly reference them, requiring the correspondence to be inferred from geometric or textual clues. We introduce SelectionVLM, a VLM augmented with a new modality that exposes information from the selection buffer: a per-pixel integer-valued image containing the pick IDs which 3D graphics software uses to allow interactive picking. Unlike RGB-based strategies such as object coloring, 2D bounding boxes, or numbered image markers, the selection buffer provides a direct correspondence between image pixels and CAD entities. We show that this representation enables SelectionVLM to substantially outperform RGB-only approaches on a range of entity identification tasks. Because pick IDs can also index structured per-entity metadata, supplied as part of the prompt, SelectionVLM can directly connect visual entities with their attributes, yielding further improvements on relational and spatial reasoning tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.