acceptodds
Under review as a conference paper at ICLR 2027

TriEval: A Three-Facet Framework for Modeling Capability, Task Difficulty, and Prompt Sensitivity in LLM Evaluation

Abstract

Average benchmark accuracy remains the dominant metric in LLM evaluation. However, it conflates model capability, task difficulty, and prompt-dependent variation across models and tasks. In addition, characterizing prompt sensitivity can require costly evaluation across thousands of configurations. To address these challenges, we introduce TriEval, a three-facet framework that jointly estimates baseline model capability under a reference prompt, task difficulty, and prompt effects on models and tasks. A regularized Gaussian variational expectation-maximization (RGVEM) algorithm with closed-form updates is developed for scalable estimation. We also introduce prompt sampling to reduce the number of configurations in the evaluation while preserving estimation quality. We evaluate TriEval through simulations and an empirical analysis of five open-weight LLMs across prompt configurations and six benchmarks. Simulations show that TriEval more accurately recovers reference-prompt model rankings than conventional methods. Empirical results show that (i) a small set of tasks exhibits anomalous difficulty shifts across prompts; (ii) model rankings vary across prompts; (iii) prompt effects vary systematically across models; and (iv) estimates from sampled configurations closely match those from the full configuration set.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.