acceptodds
Under review as a conference paper at ICLR 2027

CAD-Ladder: A Source-Grounded Benchmark for Text-to-CAD Models and Agents

Abstract

Generating CAD models by synthesizing executable code from natural-language descriptions is promising, but evaluation depends on whether each description and reference program denote the same geometry. Misaligned pairs can penalize faithful systems and obscure differences among models and agents. We introduce CAD-Ladder, a source-grounded annotation method and benchmark that derives five increasingly abstract descriptions through sequential code-to-language rewriting from execution-verified CadQuery programs. Semantic auditing checks whether the reference solid satisfies the description’s geometric constraints and estimates essential source-constraint retention. Repeated blind reconstruction measures geometric fidelity and inter-run agreement through deterministic geometric scoring. CAD-Ladder achieves a semantic consistency rate of 77.9%, compared with 42.5% for CADPrompt. Audit-guided revision of 38 descriptions that fail the semantic criterion raises their joint reconstruction pass rate from 55.3% to 73.7%. Applying CAD-Ladder's annotation principles to the same 200 CADPrompt targets increases mean reconstruction IoU from 0.7748 to 0.9295. The resulting benchmark separates standalone LLM and agentic configurations on T1 at L5, with delivery quality ranging from 49.7% to 95.1%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.