Mask Schedules Are a Programmable Supervision Graph
Abstract
Masked pretraining treats the mask distribution as a hyperparameter. But a mask draw also assigns regions to the context and to the target, and so decides which pairs of regions the objective makes predict one another. Averaged over draws this gives a co-exposure matrix over pairs of regions, a supervision graph that the mask distribution, or schedule, determines completely and that can be rewritten with the data, the network and the objective held fixed. Removing one edge takes predictive influence away from that pair alone (AUC 0.996) and drives its relational readout to chance, while each region's own content remains readable. It does so even when both regions are predicted more often than before and still appear together in every input. The effect reverses when two schedules are exchanged and is graded in co-exposure. It holds on real images: with the regions of CelebA faces tiled onto noise, raising the co-exposure between one eye and the mouth lifts that eye's readout of mouth opening from 65.4% to 71.1%, the level of the mouth region itself (70.9%). The graphs a fixed geometry admits form a polytope with exact linear-programming certificates, and a specified graph can be compiled by constrained maximum-entropy projection, which gets the direction of the effect right but, at the same co-exposure of the edited pair, achieves less of it than placing target blocks by hand. The effect shrinks with the share of the objective the relation carries and vanishes when the image's own coherence already delivers the information, as on whole faces.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.