acceptodds
Under review as a conference paper at ICLR 2027

WHEN ROUTING PROBES IMPROVE WITHOUT NEW INFORMATION: AUDITING UNCERTAINTY BEYOND MODEL OUTPUTS

Abstract

Improved error prediction from routing traces does not by itself establish information beyond a model's outputs. We examine this inference with real output–routing pairs from vision transformers and correctness labels redrawn from a fixed output-only function fitted on disjoint data, so that routing adds no information by construction. A width-matched MLP comparison nevertheless reports a routing gain in 308 of 600 confidence-only evaluations, whereas a linear comparison reports none. On a six-model panel, holding each training trajectory fixed and choosing the checkpoint by validation log loss instead of validation accuracy reduces detections from 50/120 to 0/120, and an independently implemented MLP shows the same contrast (83/120 to 0/120). Log-loss selection removes all 528 raw detections among the 1,920 MLP null evaluations but does not make the raw comparison sensitive: in two matched synthetic settings it detects a signal of about 0.005 nats in 0/20 replicates each, while a conditional permutation test based on an estimated routing law detects it in 11/20 and 10/20 and has low measured rejection rates on the null benchmark. On real correctness labels, the primary conditional analysis finds model-relative evidence of conditional-assignment dependence in five DeiT attention-residual families; in four of them the evidence persists under two specified variants of the conditional model, and no family passes an additional noise criterion. Improving a probe's fit and testing for incremental information are separate problems that need separate validation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.