CPSE-Bench: Benchmarking Coding Agents on Cyber-Physical Systems Engineering
Abstract
Coding agents perform strongly on generic coding benchmarks, but it remains unknown whether they can reliably engineer cyber-physical systems (CPS), such as driving simulators, robotics platforms, and autonomy stacks whose correctness requires closed-loop validation rather than unit testing alone. We introduce CPSE-Bench, a repository-level, execution-grounded benchmark of real cyber-physical engineering tasks across thirteen production repositories spanning driving, robotics, and aerial autonomy. Each task pairs a historical issue with a containerized behavioral reproducer verified against the buggy version and maintainer repair. We organize the released tasks into eight literature-grounded cyber-physical capability suites and evaluate nine models with mini-SWE-agent and find that even the strongest resolves only 73.5% of tasks. Our failure investigation reveals errors in enforcing physical constraints and propagating state across components, motivating models and scaffolds that reason about and validate the behavioral effects of code changes. We will publicly release the benchmark's public split and toolbox, comprising the collection pipeline, evaluation suite, and diagnostic toolbox, to facilitate further research in this field.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.