acceptodds
Under review as a conference paper at ICLR 2027

Spend Where It Counts: Dynamic Budget Allocation for Multi-Turn LLM Evaluation

Abstract

Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events—e.g., jailbreaks or successful task completion by an agent—often emerge only after repeated interactions. These events might be rare and, under any feasible computational budget, remain unobserved. Recent conformal survival frameworks construct reliable lower predictive bounds (LPBs) on the number of iterations to trigger the event of interest, but rely on static budget allocation that prohibits adaptivity in multi-turn setups. To address this, we introduce Dynamic Allocation via PRojected Optimization (DAPRO), a theoretically valid dynamic budget allocation framework for bounding the time-to-event in multi-turn LLM interactions. We prove that DAPRO satisfies the expected budget constraint and provides distribution-free, finite-sample coverage guarantees without requiring the conditional independence between censoring and event times assumed by prior conformal survival approaches. A key theoretical contribution is a novel coverage bound that scales with the square root of the mean censoring weight over only a subset of the samples rather than the worst-case weight, yielding tighter guarantees than prior work. Furthermore, DAPRO can be employed to obtain unbiased estimates of population-level evaluation metrics, such as the jailbreak rate, under limited computing resources. Comprehensive experiments across agentic task success, LLM jailbreaks, toxic content generation, and RAG hallucinations demonstrate that DAPRO achieves coverage closer to the nominal level with lower variance than static baselines, while not exceeding the budget constraint in expectation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.