From Scores to Evidence: A Framework for Epistemically Valid Evaluation of LLM Coding Age
Abstract
Large language models (LLMs) are increasingly evaluated as autonomous coding agents, yet existing benchmarks typically reduce performance to task-levelcsuccess metrics that obscure why agents succeed, fail, or change behaviour across interaction stages. We introduce EDASE-AIO, an evidence-based framework for analysing coding agents through claim-level trajectories rather than binary task outcomes. The framework models agent interactions as structured evidence chains spanning initial attempts, feedback-driven revisions, independent validation, and final outcomes. We apply the framework to a controlled study of coding agents across 180 claim trajectories, 900 claim-level evaluation cells, and 66 replicated tasks spanning multiple programming languages and task categories. Through blinded dual-pass annotation, deterministic evidence freezing, and independent adjudication, we identify behavioural patterns including persistent degradation, stable partial align- ment, soft improvement, and recovery trajectories. Our framework treats evaluation evidence as a first-class object and enables more reliable conclusions about agent capabilities, limitations, and behavioural change under realistic software engineering workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.