An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules
Abstract
Masked prediction learns to infer missing variables from visible context selected by a mask schedule. When does small masked-prediction excess risk guarantee recovery of the true joint distribution? We study this question on finite product spaces under the joint masked-block log loss, with conditionals induced by a single joint distribution. To quantify recovery, we introduce an -identifiability modulus that measures the worst-case error among joint distributions with population excess risk at most . For data with separated modes pinned down by sufficient context, mask schedules supported only on large contexts can permit nonvanishing mode-weight errors at an exponentially small excess risk. An information decomposition explains why: the loss captures mode-weight mismatch only where context leaves the mode uncertain. We prove two-sided bounds showing that, over a fixed range of mode weights, masked-prediction sensitivity to mode reweighting is governed by schedule-averaged residual mode uncertainty. Assigning schedule mass to low-visibility masks that retain this uncertainty yields recovery bounds within the reweighting family. Moreover, positive full-mask probability characterizes uniform control of joint KL divergence by excess risk. Exact calculations and controlled gradient optimization validate these predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.