DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?
Abstract
Generative CAD evaluation has increasingly moved beyond static geometry toward executability, parameter reasoning, and editability, yet how behavioral evidence should be allocated under a finite evaluation budget remains largely unresolved. We show that this omission creates a measurement paradox: auditing each generated program more deeply can reduce the reliability of model-level conclusions by sacrificing task and generation coverage. We formalize this problem as a three-level evidence-allocation framework over task templates, independent generations, and counterfactual edit states. DepthBenchCAD defines an audit-depth-invariant average failure-risk estimand, derives finite-population variance under nested sampling, and combines variance with measured execution costs to predict when additional auditing improves or harms model-level estimation. Across two CAD environments and five LLM-based generation systems, deeper auditing improves individual-program measurement. We find a clear answer to when deeper auditing is most valuable: when a program’s behavior varies substantially across edit states and generating new programs is expensive; when uncertainty is driven mainly by task differences or generation randomness, broader coverage gives more reliable model-level estimates. Calibration estimates predict the direction of most such audit-depth changes on held-out tasks. These results advance behavioral CAD evaluation from checking whether generated programs survive edits to determining how much and where to audit for reliable model-level conclusions under finite budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.