Little Engines of Evidence: Codification of Task-Specific Natural Language Rubrics
Abstract
Language-model-based evaluators can grade outputs of complex tasks using natural-language rubrics as specifications. High-quality evaluators often use agentic workflows that make multiple LLM and tool calls. While effective, such evaluators are expensive, noisy, and difficult to run at high cadence. We present an approach for rubric codification: converting each natural-language rubric into a small executable program that servers as the evaluator. We focus on spreadsheet tasks where an input workbook is transformed into an output workbook based on a user query, and codify the NL rubrics associated with such tasks. A codification agent receives a rubric, training examples consisting of initial and final workbooks plus verdicts from an (expensive) agentic evaluator, and it emits Python code. We evaluate the effectiveness of the codification approach on a finance-spreadsheet benchmark set with 36 queries and 194 rubric criteria.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.