acceptodds
Under review as a conference paper at ICLR 2027

Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism?

Abstract

Activation probes are increasingly used to monitor LLMs in deployment. A probe is typically trained once per model and dataset under one inference configuration, then used under whatever batch size and numerical precision the serving stack happens to use. Because common GPU kernels are not batch-invariant and floating-point formats round differently, the activations seen at deployment are not the ones the probe was trained on. We measure what that costs, for Llama-3.1-8B, Qwen3-8B and Gemma-3-4B, across batch sizes 4, 8 and 16 and float32, bfloat16 and float16, at four depths and four token positions. We train 768 probes on one configuration, evaluate each on every other, and compare the same probe's verdicts between two runs example by example. Our central finding is that probes are stable under these perturbations, but that aggregate accuracy is the wrong instrument for showing it: it understates how many individual verdicts change by a factor of two to nine, and that churn, once measured, proves symmetric and leaves the probe's ranking intact. At the prompt, probe accuracy never moves by more than 0.47 percentage points across transfers, and only 0.076% of individual verdicts change; under float32 with only the batch size varied, not one verdict changes in evaluations. During decoding the flip rate rises to 2.8%, but it separates cleanly: rows whose realised tokens matched flip in 0.12–0.15% of cases at every position, while rows whose tokens diverged flip in 12.9%. Reconstructing each run's greedy trajectory from stored logits shows why, and shows that the cause is the text rather than the arithmetic: a bfloat16 batch-size change flips the first generated token for 2.1% of rows and leaves the two runs on different tokens in 25% of rows by token 20, while float32 diverges nowhere. Flips are symmetric between classes, Cohen's stays above 0.94, and AUROC moves by at most 0.05 points, so what the perturbation disturbs is the handful of examples within rounding distance of the boundary. Underneath, the activations move about as much as the format's rounding: in bfloat16 a batch-size change perturbs them by a median relative of , roughly the same change in float16, matching float16's three extra mantissa bits, and for Gemma, where both runs share a GPU, more than the entire float32-to-bfloat16 conversion. Probes absorb this; the model's own next-token argmax does not. Robustness evaluations of activation monitors should therefore report per-example agreement rather than aggregate accuracy, separate representational noise from input change, and state the serving configuration alongside the result. We will release our extraction and analysis pipeline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.