Auditing LLM API Quality Beyond Task Success
Abstract
Black-box access makes it difficult to determine whether the quality delivered by an LLM API has deteriorated. For agentic workloads, task pass rates summarize outcomes but do not reveal whether decisions were grounded in observations, instructions were followed, or errors were handled appropriately. We study whether process-quality measurements provide actionable auditing information beyond final outcomes. We propose a trajectory-based auditing framework that uses generative judges to produce dimension-specific assessments with supporting evidence and compares them with reference measurements under a controlled workload. To accommodate long trajectories, the framework combines compressed views with selective access to omitted evidence. We develop an evaluation protocol using natural trajectories, targeted quality interventions, and null edits, with independent checks of the realized effects. The protocol compares outcome-only, process-only, and combined audit signals, and evaluates detection, false alarms, evidence localization, and judging cost. Comparisons across judge models and execution harnesses assess the dependence of the measurements on the evaluator. The study examines when observable process quality complements pass rates as a practical signal for black-box API auditing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.