acceptodds
Under review as a conference paper at ICLR 2027

Preserving Rule Uncertainty for Auditable Coding-Agent Evaluation

Abstract

Evaluating coding agents requires interpreting both execution outcomes and trajectory-level behavior, yet deterministic rules can be misleading when their preconditions fail, their implementations change, or supporting evidence is absent. We present RAVEL, a framework that represents each rule as a versioned sensor contract specifying applicability, provenance, evidence requirements, and an admissible reliability set. Two model proxies inform these sets, but their agreement is not treated as semantic correctness and their disagreement is retained as uncertainty. A compact generative judge is supervised to emit criterion scores and event citations, then trained with matched reward variants, including a worst-case rule-agreement objective over admissible reliability weights. An audit of 402K production trajectories shows that one parser correction changes 36.9% of grades and that 42.0% of sampled pairwise orders remain uncertified across two detector versions. For diagnostic evaluation, we construct 670 instances from automatically extracted execution signals and 2,017 metamorphic relations. Across 2,680 evaluation views, supervised initialization raises strict structured-output coverage from 0% to 88.1%. The four reward variants, however, have overlapping uncertainty intervals and change rank between 128 and 256 training steps; the proposed set objective does not show a stable advantage. These results support the measurement framework—auditing applicability, provenance, and version sensitivity—while delimiting what the present execution-anchored study can establish about process-level judgment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.