Agentic Auto-Rubrics: Verifiable Rubrics for Judging Multi-Turn Coding Agents
Abstract
Judging whether a coding agent fulfills a user's request is challenging in multi-turn sessions. In multi-turn sessions, user requirements constantly change and replace earlier requests, while user reactions can unintentionally reveal task outcome and bias evaluation. We introduce , a framework that compiles a multi-turn interaction into a verifiable specification of what the agent is responsible for. Given reaction-masked user turns, AAR constructs the active requirements and compiles them into fine-grained atomic rubrics, which specify explicit evidence types and codebase locations. To increase the judge's efficiency, AAR compresses similar atomic rubrics into a concise set of high-level parent rubrics. An agentic verifier then inspects the agent's workspace through read-only tools, evaluating these rubrics from concrete file-based evidence and references. Across multi- and single-turn coding benchmarks, AAR consistently outperforms direct scoring judges and existing rubric-based verifiers. While baseline verifiers degrade as sessions lengthen, AAR maintains stable accuracy across session lengths and five judge models, showing that compiling evolving interactions into verifiable specifications enables robust evaluation of coding agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.