acceptodds
Under review as a conference paper at ICLR 2027

Per-Input Prior Adequacy: Sources of Information and Conditions of Failure

Abstract

Tabular foundation models (TFMs) acquire their entire competence from a synthetic prior sampled at meta-training time, and their prediction errors are far from uniform across inputs. Prior evaluation scores a whole dataset with one number, and such a score cannot say where a model is misaligned. Applying those scores per input does not help either. A coverage-style score is uninformative, and a paired improvement ratio is blind to uniform prior gaps because it is a rate statistic. To address this, we propose the reference-averaged log-error ratio (, RA-ratio), a per-input, training-free diagnostic composed of two components, the excess loss of the model over a Bayesian reference and an average of that excess over the reference's prior strength. First, for each input we take the model's log error in excess of that of a closed-form Bayesian reference, the Bayes gap of in-context learning theory. Second, we average that excess over a grid of reference strengths, which removes the reference's strength as a tuning knob rather than selecting it. Third, we validate the diagnostic causally, injecting meta-training tasks from the region it flags and verifying that the error falls there and not in an off-target region. Notably, the RA-ratio attains a mean rank correlation of with the per-point error on regression datasets with seeds, ahead of every reference-free baseline we tried including three conformal variants, and replicates on a second architecture. On a held-out classification family, prior-based per-input diagnostics help where the model's own uncertainty has stopped resolving difficulty, and hurt where it already succeeds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.