SUPERVISION BEFORE STRUCTURE: WHAT A WORLD MODEL NEEDS TO KNOW WHAT IT CANNOT ANSWER
Abstract
A world model asked a counterfactual question should first decide whether its evidence determines the answer at all. We study what such a model needs for that decision on a benchmark where the answer is known exactly: ten environments over seven systems (six physical, one structural causal model) whose hidden mechanism is drawn from a finite bank, so compatible laws, answer dispersion and the value of every probe are enumerated rather than estimated. We test the causal world-model design the field is converging on — a mechanism posterior with per-mechanism counterfactual heads — against a supervision-matched control with none of that machinery. Three findings follow. With an answerability label on every training query, learned structure is redundant: eight of nine ablations are null, the posterior's own dispersion scores answerability far below a learned head, and the one margin the architecture showed was an artifact of the baseline's training loss. With scarce labels — as few as forty — a factorised posterior improves label efficiency by up to AUROC, needs no mechanism labels, and transfers to a non-physics causal model; auxiliary supervision, generic factorised capacity, independent heads and width do not account for it, nor, on most environments, do the posterior's own auxiliary losses. Acting on the query is at oracle level for both models; stopping and answering after self-gathered evidence are weak for both. We release the benchmark, every control, and the failed hypotheses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.