Generalization Beyond Coding: An Empirical Comparison of Coding Agents and Purpose-Built Agents
Abstract
Recent advances in large language models (LLMs) have enabled both flexible coding agents and specialized, purpose-built agents to tackle complex tasks. General-purpose coding agents such as Codex show strong adaptability, while purpose-built agents rely on domain-specific structures to solve specialized problems. This prompts a key question: Are coding agents now advanced enough to make purpose-built engineering obsolete? This paper presents a comprehensive empirical comparison between Codex and comparable state-of-the-art purpose-built LLM agents across 20 non-coding benchmarks spanning finance, clinical reasoning, data analysis, information retrieval, and scientific research. For each benchmark, an AI-assisted procedure finds the comparable state-of-the-art purpose-built agent with released code, verifies its reported score, and re-runs it with the same language model as Codex; both agents are scored on the same instances by the same grader. Using the same underlying model and evaluation methods, Codex significantly outperforms on 5 benchmarks, shows no significant difference on 13, and underperforms on only 2. Many of the advantages previously attributed to purpose-built agents diminish when controlling for the backbone model and evaluation protocols. In certain cases, the additional complexity introduced by staged data processing even results in poorer performance compared to Codex’s direct, end-to-end approach. Where purpose-built agents do win, much of their advantage is portable: their domain tools, handed to Codex, carry much of the benefit, and their decisive workflow steps encode procedural knowledge that Codex can execute but rarely finds on its own; once given them, it matches the purpose-built agents on most of the selected instances it had lost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.