acceptodds
Under review as a conference paper at ICLR 2027

How Much Evaluation Is Enough? Risk-Aware Adaptive Evaluation (RAAE) of LLMs in High-Stakes Financial Applications

Abstract

Large language models are increasingly used in financial applications where evaluation errors can have highly asymmetric consequences: overlooking a minor formatting issue is fundamentally different from missing an unsafe financial recommendation, a privacy violation, a fraud-enabling response, or sensitive-data disclosure. Yet current LLM evaluation pipelines commonly rely on fixed-size test sets and aggregate metrics that treat evaluation examples uniformly. We propose Risk-Aware Adaptive Evaluation (RAAE), a sequential evaluation strategy that allocates LLM-as-a-Judge effort by both statistical uncertainty and domain-specific risk, stopping only once an anytime-valid confidence sequence – constructed correctly for sampling without replacement from each finite risk stratum – is tight enough and every critical risk stratum has received sufficient coverage. On a 267-item, 8-stratum financial evaluation suite with 3 candidate LLMs and 2 independent judges, adaptive sampling shows a real, statistically significant advantage over naive sampling in the low-to-mid evaluation-budget range; that advantage narrows as budget approaches the adaptive methods’ own (necessarily late, since the underlying guarantee is strict) natural stopping point, where all methods converge toward near-exhaustive recall – reported plainly rather than only at the budget most flattering to our own method. RAAE’s specific contribution over plain uncertainty-driven sampling is not a general recall improvement – at a moderate coverage requirement it is statistically indistinguishable, and can even slightly underperform – but a structural capability: it can enforce a pre-declared minimum coverage requirement for critical-risk strata that plain uncertainty-driven sampling has no mechanism to offer at any setting. Achieving complete coverage costs additional evaluation budget, and matched-budget comparisons show no consistent recall advantage for RAAE, confirming that its benefit is the enforceable coverage guarantee itself, not superior sampling efficiency. We further find that judge choice alone changes the detected critical-failure count by more than 2× on this dataset, underscoring that evaluation reliability depends on judge selection as much as on sampling strategy. We provide a formal anytime-valid stopping guarantee grounded in confidencesequence theory, with explicitly stated assumptions and limitations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.