TREEval: Allocating Compute for Efficient LLM Safety Evaluation
Abstract
Safety evaluation of large language models often relies on repeatedly generating and judging complete responses, making thorough evaluation computationally expensive. We introduce TREEval, a framework that formulates safety evaluation as a compute-allocation problem over response probability mass, reusing alternative continuations exposed during generation instead of repeatedly generating isolated full rollouts. A simple instantiation, TREEval-DFS, achieves over 98% correlation with repeated-sampling safety measurements with an average estimated compute gain of 69x. We find that this efficiency is enabled by broader response-space coverage and by noisy but informative partial-response judgments that preserve much of the model-level safety signal. We further characterize the tradeoff between allocating compute within versus across prompts: greater inter-prompt exploration exposes more harmful behavior, and these differences show similar trends to standard measures of response diversity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.