The Experiment Defines What Is Measurable
Abstract
Must causal representation learning reconstruct an entire latent world when the scientific task asks for one target-wise measurement? Two stochastic hard interventions on a known target define such a measurement before any representation model is chosen: their density ratio cancels every non-target mechanism, and the Bayes environment logit reads out the remaining scalar quantity. Observation determines what this latent estimand becomes. Under an arbitrary invariant observation channel, the observable ratio is a posterior average of the latent intervention ratio, yielding a coordinate, a quotient, or a posterior observable. Alignment across entities or acquisition conditions imposes a separate requirement. The intervention laws also control the semantics and conditioning of the measurement. Exponential tilts program a chosen target statistic; among monotone tilts, affine statistics uniquely optimize worst-case conditioning under a weak symmetric-KL budget. Controlled experiments recover the programmed statistic across linear, monotone, and non-injective choices and predict a renderer-induced quotient before training. Matched comparisons with ILCM, CITRIS, BISCUIT, iVAE, and GSCALE-I produce the corresponding crossed profile: target-wise classification favors the programmed quotient, whereas interfaces designed to recover state or intervention structure retain more signed-state and intervention information. Natural-image and chemical-sensor studies further separate within-context discrimination from alignment under unseen entities and acquisition drift. The first identification question is therefore which measurement the experiment induces, before attempting full latent reconstruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.