acceptodds
Under review as a conference paper at ICLR 2027

CarryBench: A Multi-Role, Multi-Maintainer Benchmark for the Maintainability of Agent-Generated Code

Abstract

Agent-generated code is often the starting point for later work. Functional correctness tests show whether the current task is done. They do not show how difficult the next task will be. Existing maintainability benchmarks usually measure one downstream activity, and most let an agent work on its own output, so the score mixes the artifact with the worker. CarryBench separates the two. Nine producer models write repository-scale programs for 13 cases in Python, Go, Rust, and C++, and frozen outputs are evaluated on three maintenance demands. Each demand is a targeted downstream task under three fixed maintainer profiles. An observer predicts behavior from source for Analysability, a consumer builder constructs a new use for Reusability, and a tester writes a black-box test suite for Testability. Static Modularity is measured on the same source as a traditional proxy. Rankings agree most across maintainers for Reusability, and less for Analysability and Testability. Different downstream outcomes produce different producer orderings. Systematic within-case analysis shows that higher state cohesion need not coincide with easier downstream work. Ablations of prespecified source-feature blocks find conditional associations with lexical organization, observable entry-to-result paths, and behavior localization. Those associations differ across maintenance demands and maintainer profiles. CarryBench treats maintainability as the difficulty a fixed downstream worker experiences on a specified next task over a frozen artifact. It reports that difficulty for each maintenance demand and each maintainer profile, from the executable tasks and the source-level associations, without combining those results into a universal score.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.