2D-FET-BENCH: FROM SPATIAL REASONING TO FET DESIGN ON FLAKES
Abstract
Field-effect transistor (FET) layouts on exfoliated two-dimensional flakes are typ- ically drawn by hand for each flake, placing contacts and gates to match its posi- tion and outline in optical micrographs. To our knowledge, no executable bench- mark tests whether language-model agents can perform this flake-specific con- struction reliably. We introduce 2D-FET-Bench V2, a benchmark of 128 lay- out tasks built from microscopy-derived flake contours, including hole-containing flakes and multi-flake tasks. Each task supplies a textual device specification and contour coordinates. An agent generates typed polygon and path operations ren- dered to GDSII. A deterministic verifier checks geometric and structural require- ments, and a separate integrity check verifies that the supplied contours remain unchanged. Scripted reference layouts pass all 128 tasks, showing that every task is solvable. We evaluate six models and seven workflow and scaffold variants of GPT5.6-Luna, with five attempts per task. The best-performing configuration in the six-model panel, GPT5.6-Luna with ReAct-3, passes 62.3% of attempts and solves 80.5% of tasks at least once (coverage) and 43.8% in all five attempts (con- sistency). ReAct-3 exceeds the one-pass Plan-and-Execute by 27.0 pass@1 points at 2.46× the tokens. An expert audit of one sampled verifier-passing layout per covered task, across five ReAct-3 configurations, accepts 56.4%–63.5% of them. The benchmark evaluates geometric and structural FET layout construction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.