Amplify, Locate, Estimate (ALE) on Rare-Failure Rates of Fine-Tuned LLMs
Abstract
Safety fine-tuning can make undesirable model behaviors sufficiently rare that measuring their frequency by direct sampling becomes prohibitively expensive. We introduce Amplify, Locate, Estimate (ALE), a method for estimating such rare failure probabilities efficiently. ALE uses the change in token log-probabilities induced by the safety fine-tuning to define an amplification direction, along which we increase the failure rate for cheap measurement. We develop a score-function estimator that uses the selected amplified batch to recover the local slope of the log failure rate from the token probabilities of the sampled failures, enabling extrapolation back to rare target rates. We also use the amplified batch to obtain an importance-sampling estimate. Across 41 evaluation targets, ALE's estimate-to-reference ratios have a median close to one, with 37 ALE estimates within a factor of three of reference values obtained from independent direct measurements. Ablation tests show that with a budget of up to 1000 labeled rollouts, in addition to a cheap unlabeled scan, ALE estimates failure probabilities for Qwen3-8B fine-tunes down to about , where direct measurement can require hundreds of thousands of rollouts, as the required sampling budget scales inversely with failure probability at fixed relative precision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.