Certificate Semantics Determines Query Geometry in LLM Evaluation
Abstract
Large language model (LLM) evaluations commonly aggregate responses across prompt formats, option orders, and other representation variants. Such aggregation can obscure whether behavior is consistent within every registered representation group. We formulate representation-conditioned certification as a group-level ternary decision: for a fixed item, determine whether all group-conditioned response probabilities lie above a threshold, all lie below it, or groups occur on both sides. Certificate semantics induces distinct first-order information geometry. Robustness is universal, requiring dense resolution of potential blockers; heterogeneity is existential, requiring only a positive and a negative witness. By specializing established fixed-confidence pure-exploration geometry, we derive closed-form characteristic times and introduce ActiveOrbit, a sequential allocator that tracks the weakest blocker or balances opposite witnesses. We establish item-level anytime correctness and selector-specific first-order behavior. A mechanism intervention matches the predicted regime asymmetry: witness-selective allocation yields exactly query reduction in all six robust scenarios and positive reductions in all six heterogeneous scenarios (mean ). Across six LLM model–task cells spanning BoolQ, PIQA, and RTE and three open-weight model families, ActiveOrbit reduces query cost by – relative to round robin, with every paired 95% confidence interval excluding zero. These results establish a practical principle: query allocation should be designed around the logical structure of the claim being certified.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.