PICAgentBench: Benchmarking Multimodal Agents for Photonic Integrated Circuit Design
Abstract
Agents for photonic integrated circuit (PIC) design must interpret visual and structural evidence, reason about constraints, and produce valid design artifacts. Final-artifact success alone provides limited insight into which engineering tasks agents perform reliably or where errors arise. We introduce PICAgentBench, a multimodal benchmark with a selected bank of 982 cases across six separately scored PIC engineering tasks. These tasks comprise requirement-to-design, device recognition, circuit recognition, layout-to-netlist extraction, fault localization, and fault repair. They place overlapping demands on perception, reasoning, and action through task-specific combinations of specifications, visual layouts, and structured circuit representations. Fault repair evaluates constrained geometric editing given a local crop and known fault type. Expert-developed checking scripts and manual review support reference construction, with explicit observation and verification boundaries defining what each score measures. A shared agent environment supports model comparisons under common interaction conditions, while task-specific acceptance criteria define six-task performance profiles. Separate native metrics distinguish device classification from localization and circuit identity from potential function. Together with task-level breakdowns and resource reporting, these profiles provide a structured basis for identifying model strengths and limitations in PIC design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.