acceptodds
Under review as a conference paper at ICLR 2027

A Causal Hierarchy for Agents

Abstract

Agents are defined by the preferences they induce over the different actions available to them. Understanding an agent involves predicting their preferences in any scenario they might face at deployment, as well as in counterfactual circumstances 'had things been different'. We might ask whether, even in principle, we could determine an agent's preferences (i.e., ranking) over candidate actions empirically, from observations of behaviour? In this paper, we show that this is not possible in general. We show a preference-level counterpart to the Causal Hierarchy Theorem (Bareinboim et al., 2022). Specifically, while the CHT states that probabilities at a higher layer (e.g., counterfactual, interventional regimes) are under-determined from lower-layer probabilities (e.g., interventional, observational regimes, respectively), we show that also the ranking of actions (a strictly coarser object than a distribution) at a higher layer is under-determined with increasing probability as the number of actions grows. For example, the question of whether one action produces more harm than another is a statement that is not typically deducible from behavioural data. We conclude with a discussion of the implications of this result for the evaluation of agents, AI Safety, and related disciplines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.