RAF: RUBRICS AUTO-GENERATION FRAMEWORK FOR AUDITABLE EVALUATION OF LONG-HORIZON LLM AGENTS
Abstract
Evaluating LLM agents with rubrics is a popular approach, which specifies a set of textual criteria with per-criterion weights, together with a scoring rule that ag- gregates criterion-level judgments into a single score. However, it is challenging to design effective rubrics due to rubric bias and score instability, which can lead to unreliable evaluation results that deviate substantially from human judgments. In this paper, we address these challenges by proposing a novel rubric generation method. First, we introduce an AND–OR tree as the underlying rubric structure, which organizes criteria into two types of nodes: leaf nodes, each representing an atomic and verifiable criterion, and condition nodes, which hierarchically ag- gregate the judgments of leaf nodes through AND/OR conditions. Secondly, we introduce Rubric Auto-generation Framework (RAF) that automatically and itera- tively optimizes rubrics within AND-OR tree structure to improve their alignment with human judgments. Furthermore, we construct a benchmark comprising 121 long-horizon tasks from personal-assistant scenarios, annotated by human experts. On this benchmark, RAF outperforms baselines by a clear margin: it improves the Pearson correlation with human by 0.136 over the strongest baseline, reaching 0.8102, and reduces the average scoring variance from 10.18 to 3.36.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.