FSR: Assessing Reliability in Language Model Interpretability
Abstract
Affine lenses translate the hidden states of large language models into intermediate predictions. But when should a researcher trust those predictions to draw a conclusion about the full model? We introduce Fisher–Schur Reliability (FSR), which uses readout errors on reference prompts to assess reliability for the specific quantity being measured. For token distributions, FSR measures how readout errors align with distinctions between likely tokens. This information reduces error-ranking loss (EAURC) by 4.1–27.3% beyond confidence, input atypicality, and token frequency across five models, helps richer predictors, and remains useful on a disjoint cohort with the score frozen. For intervention effects, FSR compares each estimate with its residual error scale to select directions to report. With the readout fixed, it improves lens and attribution-patching reporting. A small model of each reading's error scale further improves selection on a fresh 512-prompt cohort, and both forms of reliability scoring replicate on Gemma beyond the two Qwen models used for development. FSR thus complements effect estimation: it assesses which readings support an interpretation and which warrant exact checking.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.