A Causal Hierarchy for Agents
Abstract
Agents are defined by the preferences they induce over the different actions available to them. Understanding an agent involves predicting their preferences in any scenario they might face at deployment, as well as in counterfactual circumstances 'had things been different'. We might ask whether, even in principle, we could determine an agent's preferences (i.e., ranking) over candidate actions empirically, from observations of behaviour? In this paper, we show that this is not possible in general. We show a preference-level counterpart to the Causal Hierarchy Theorem (Bareinboim et al., 2022). Specifically, while the CHT states that probabilities at a higher layer (e.g., counterfactual, interventional regimes) are under-determined from lower-layer probabilities (e.g., interventional, observational regimes, respectively), we show that also the ranking of actions (a strictly coarser object than a distribution) at a higher layer is under-determined with increasing probability as the number of actions grows. For example, the question of whether one action produces more harm than another is a statement that is not typically deducible from behavioural data. We conclude with a discussion of the implications of this result for the evaluation of agents, AI Safety, and related disciplines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.