CALIPER: Direct Point-Text Alignment for Dimension-Bearing CAD Retrieval
Abstract
Retrieving CAD models from a geometry-only gallery with natural-language text differs from open-vocabulary 3D retrieval in what the query contains. Engineers describe a part by its dimensions, so retrieval must match stated numbers against geometry. The prevailing design, which anchors a point encoder to a frozen CLIP text tower, discards this information on both sides, truncating expert captions before their dimensions and normalizing away the scale those dimensions refer to, and it freezes the tower that would have to learn them. This paper presents , a point–text bi-encoder for geometry-only CAD galleries that keeps the whole caption and each shape's corpus-frame scale and adapts both towers to paired CAD data. Ablations show that each preserved quantity and each adapted tower is load-bearing. Because the standard Text2CAD split leaks a majority of its evaluation shapes into training through near-duplicates, results are reported on a component-disjoint split, where reaches 10 0.955 under cluster credit and 0.848 by exact uid. Counterfactual probes establish aggregate-level grounding of the numbers in a query, and on human-written CADPrompt queries over held-out geometry retrieves dimension-bearing prompts seven times as often as the strongest zero-shot baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.