Readable Coordinates, Distributed Concepts: Concept Supervision in DreamerV3
Abstract
Concept supervision gives a world model a vocabulary: named coordinates report facts such as whether an agent has wood or is running low on energy. Understanding what these reports reveal requires connecting the readable interface to the representation behind it. We study this connection in DreamerV3 on Craftax-Classic through a paired comparison of concept supervision, localization regularization, and sparse-representation objectives. Across five starting checkpoints, supervised scalar readouts consistently outperform their raw-coordinate controls, while the same concepts remain highly predictable from the surrounding state. The most revealing differences emerge beneath the aggregate scores: objectives with similar overall readability expose different inventory, spatial, and physiological variables, producing distinct views of the agent’s state alongside different training-return profiles. Two analytical results sharpen this interpretation. Identical decoding measurements can accompany different causal roles, and shared temporal structure can preserve predictability even after labels are reassigned between episodes. Together, these findings characterize concept supervision as a way to construct readable interfaces within distributed representations. They show how the value of a named coordinate depends on which concept it exposes, how reliably it reports that concept, and what the available evidence establishes about its relationship to the wider model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.