From Dead Code and Static Page to Working Engines: Software Revival with Coding Agents
Abstract
Coding agents bring two interesting questions into software engineering: can they bring scientific and commercial software that no longer runs back to a working state without changing its method, and can they build the core engine of commercial or industrial software from an open specification alone? We introduce ReviveBench, a benchmark that tests whether coding agents can restore nonfunctional scientific and commercial software while preserving its underlying methods, and reconstruct industrial software engines from open specifications alone. ReviveBench combines two complementary task families with hidden verifiers calibrated against native execution environments or industrial-grade reference systems. The revival family comprises ten tasks spanning dependency incompatibilities with Keras 3 and NumPy 2, as well as a GPU-based foundation model; every starting workspace fails verification. The strongest evaluated model restores all ten. In contamination-control experiments, line similarity to the original implementations drops from 0.51–0.96 to 0.03–0.44 without reducing any evaluated model’s pass rate. On repositories created after the cutoff, the strongest model succeeds in eight of nine runs, while weaker models achieve fewer successes. The reconstruction family comprises thirteen clean-room tasks covering statistics, SPICE circuit simulation, programmable logic controllers, finite elements, 2D and 3D CAD, logic synthesis, chip place-and-route, computational fluid dynamics, and enterprise and manufacturing systems. These engines are evaluated using numerical tolerances, formal proofs, or zero-tolerance transactional invariants; two frontier models satisfy the verification criteria for all thirteen. Crucially, constructing and validating the benchmark uncovers 28 verifier defects, including 24 false negatives and two false positives—more failures attributable to verification than to the agents themselves. Our contribution is both a benchmark for software revival and industrial engine reconstruction, and empirical evidence that verifier reliability can become the dominant evaluation bottleneck as agent capabilities improve. These findings demonstrate substantial capabilities under the tested specifications while establishing verifier validation as a central requirement for trustworthy agent evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.