Bayesian Intelligence from the Outside
Abstract
Inferring mechanisms underlying intelligence from observable behavior is a foundational challenge in artificial intelligence. We develop a theory of Bayesian intelligence for agents such as language models, which tests whether observed behavior could be generated by a rational agent responding to information revealed by inference-time reasoning. Each prompt induces a possibly imperfect internal experiment; the agent updates a full-support prior by Bayes' rule and faithfully reports its posterior over the possible answers to the question. Repetitions draw fresh, independent outcomes from the same unobserved experiment at one fixed state. We show that the agent's behavior admits this explanation if and only if its reports are not fully contradictory, i.e., some state remains possible under every report across all prompts. Report frequencies and the sizes of positive probabilities impose no further restrictions. We further propose and characterize the behavioral implications of one agent having access to a more informative experiment: there should exist a coupling of report distributions such that the more informative agent's report excludes every answer excluded by its counterpart. Finally, we show the difficulty of aggregating coarse reports: unless the agent reports a belief about the complete state of the world, the optimal aggregation can assign arbitrary weights to states that have not been excluded. These results provide a basis for understanding when agents' behavior can be modeled as revealing information through reasoning and highlight the difficulty of rejecting Bayesian rationality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.