Beyond Confidence: Log-Alignment Profiles for Monitoring Internal Computation
Abstract
Distribution shift can alter a network's internal computation while leaving its predictions confident. We construct log-alignment profiles that record layerwise amplification and activation scale for individual inputs. Calibrated against reference data, these compact profiles detect atypical processing and reveal its distribution across network depth. We call this reference-relative typicality structural familiarity. Across image classifiers, language models and visuomotor policies, the profiles detect shifts that output scores detect poorly. On CIFAR-10 and SVHN inputs matched by maximum-softmax-probability bins, an alignment-only profile achieves detection AUROC. In a supervised clean-plus-shift evaluation, adding an alignment-only profile to entropy raises ResNet-50 error-prediction AUROC from to . In language models from M to B parameters, the profile detects Python code relative to a natural-text reference at – AUROC. On LIBERO behavior-cloning policies, profile alarms detect corruption onset – steps before action entropy; on ManiSkill, profiles add episode-failure information beyond ensemble disagreement. We also test alignment interventions during learning. In grokking transformers, late-layer interventions reduce mean post-fork grokking delay by in a three-seed comparison and accelerate grokking in all eight runs of a separate paired experiment. Their effects depend on the targeted layer and subspace.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.