Estimating Rare Events in Language Models via Activation-Steered Committor Proposals
Abstract
At deployment scale, stochastic language generation can expose rare unintended or harmful behaviors that standard evaluation may never observe. We present SIREN, a white-box framework for rare-event simulation and low-probability estimation in large language models (LLMs). SIREN uses activation-steered lookahead rollouts to estimate candidate-wise future event probabilities and converts these estimates into a support-preserving sequential importance-sampling proposal over the model's high-probability token nucleus. To amortize the cost of nested lookahead, we further learn a lightweight committor-tilt surrogate: an MLP maps each candidate token's hidden state to a scalar score and is trained by cross-entropy to match the rollout-derived tilted distribution over the nucleus. At inference, this replaces candidate-specific sub-rollouts with a single cached candidate evaluation plus the MLP, while exact likelihood-ratio correction preserves unbiased probability estimation regardless of surrogate error. Across rare reading-level and jailbreak events, including probabilities on the order of , SIREN reaches event-bearing trajectories with far fewer outer samples than direct Monte Carlo, while the learned surrogate provides a practical accuracy–computation tradeoff.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.