acceptodds
Under review as a conference paper at ICLR 2027

MCAD-Bench: Benchmarking CAD Editing on Real Supplier Parts with a Controlled Answer Channel

Abstract

Benchmarks that ask a language model to produce CAD almost always require the answer as a program in a geometry-kernel API and then score the solid that program builds. We show that this choice is part of the measurement, not an implementation detail. CAD systems share features but not syntax, so no canonical API settles the matter, and a representation that moves the result should be controlled rather than ranked. MCAD-Bench is built from 144 real supplier products stored as STEP, with three task families (reconstruction, instruction-guided editing by defect injection, and parameter-specified editing) on single-part and assembly tracks. Alongside the code channel we provide a structured feature-op answer channel, in which the model emits an operation and its parameters and we apply verified primitives. Instruction-guided editing has two tiers, one that gives the defect's location and one in which the model must find it. Over five models, editing success in code spans up to 39 points, mostly through executability, and the structured channel narrows this to at most 8 in either tier. A model whose code runs on only 7% of instances when it must find the defect reaches frontier-level accuracy in the channel. The narrow band does not mean the models have converged. Many instances lie beyond what the channel can express and many are passed by every model, so fewer than 50 of the 301 single-part instances separate any two models in either tier, against at least 155 in code. Where the channel can express the repair, the two frontier models are indistinguishable when told the location and differ significantly when they must find it. Assemblies remain limited by the models, with wide spreads and no model near the channel's ceiling. We also report achievability oracles, a two-arm protocol run before any near-zero result is published. On our own pipeline they caught results that looked like model findings but were properties of the benchmark, including thirteen instances on which the metric rejects its own reference and 28 on which it accepts the unedited input.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.