acceptodds
Under review as a conference paper at ICLR 2027

Bayes-Equivalent Prompts Reveal Hidden Nuisance Dependence in In-Context Learning

Abstract

In-context learners can closely match Bayesian predictions on synthetic tasks. But do they also learn which aspects of the input a Bayesian predictor knows to ignore? We introduce Bayes-equivalent prompts: inputs that differ in raw form but induce exactly the same posterior and posterior predictive distribution, a direct test of whether learned predictors acquire the Bayesian solution's invariances. In multilayer perceptrons (MLPs) and Transformers, sensitivity to Bayes-irrelevant variation falls during training and is attenuated more in later layers, yet the identity of the nuisance perturbation remains linearly decodable from final representations. In a separate twelve-seed Transformer study trained to 800k updates, fit keeps improving and responses to task-relevant changes stay close to the Bayesian reference, while sensitivity to the same physical change when it is Bayes-irrelevant declines but stays above a pre-specified, study-specific benchmark. A consistency penalty (one weight, short continuation) sharply reduces disagreement on the targeted perturbation but raises held-out Kullback–Leibler (KL) divergence; in the sampled-target branch, this gain does not clearly carry over to a second Bayes-equivalent transformation. A public pretrained regression Transformer also shows small but measurable sensitivity to an exactly posterior-preserving recoding. Predictive fit therefore does not determine invariance, and invariance gained on one transformation need not transfer: evaluating whether a model has captured a Bayesian procedure requires testing its invariances, not only its predictions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.