Is In-Context Reinforcement Learning A Posterior Sampler? An Empirical Study
Abstract
In-context reinforcement learning (ICRL) exploits the impressive in-context learning ability of large language models to zero-shot generalize to new, unseen tasks at inference time without any parameter updates. Among many approaches, the Decision-Pretrained Transformer (DPT) is particularly appealing for its theoretical grounding: under idealized assumptions such as a sufficiently expressive model and infinite pretraining data, DPT can perform Bayesian posterior sampling. Whether this characterization survives in practice remains open. Therefore, we ask: is ICRL, specifically DPT and our proposed DPT variants, a Bayesian posterior sampler in practice? To approach this question, we establish two necessary conditions that a Bayesian posterior sampler must satisfy: sufficient-statistic invariance and the martingale property. We then construct statistical hypothesis tests to statistically assess if practically-trained ICRL models satisfy these conditions. Applying these tests in Gaussian, Bernoulli, and linear bandit settings, we find that both conditions are consistently violated, suggesting that ICRL may not behave as Bayesian posterior samplers in practice. Beyond this negative result, our tests provide a principled toolkit for testing the Bayesian behavior of ICRL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.