A Reference Alone Is Not a Bias Estimator: Identifiability in Logit Correction
Abstract
Blank images and content-free text are often queried to expose a model's class preference, and their logits are then subtracted from ordinary predictions. We ask when such a reference identifies a target-optimal correction. For affine readouts, reference subtraction is exactly geometric recentering, but the induced zero point need not match the optimum selected by a target risk. We prove two complementary non-identifiability results: off-support responses are unconstrained by in-domain behavior, and fixed responses cannot track optima that move with the target distribution. We also give a sufficient condition for identification. Frozen-checkpoint audits in long-tailed semi-supervised learning, including a three-seed replication, show why references can nevertheless help: heterogeneous responses share a training-induced prior-like direction, while their scale and orthogonal content are unreliable. Controlled experiments with two language-model families reproduce the target-risk gap under prior and domain interventions. References are useful inductive probes, but target-optimal estimation requires structure or target information.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.