Amortizing Sequential Experimental Design in Reasoning LLMs
Abstract
Despite their strong reasoning capabilities and extensive pretraining knowledge, large language models (LLMs) struggle to adaptively gather information in experimental design tasks. We study how to amortize sequential Bayesian experimental design (BED) into a reasoning LLM, using information gain as the training signal. Our approach decouples the experimental generative model from the policy, providing a fixed training distribution and avoiding reliance on the LLM's own potentially biased beliefs about the experiment. This construction also allows us to marginalize over hypotheses and outcomes to train on expected rewards with less variance, and to extend the objective across future interactions for non-myopic training. On Twenty Questions, the resulting policies improve information acquisition and task performance, outperform inference- and training-time BED baselines at substantially lower computational cost, and transfer to unseen domains and to Medi-Q, a real-world clinical dialogue dataset. These results suggest that information gathering can be learned as a transferable reasoning skill in LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.