Label Sharing through Sufficient Graph Explanations
Abstract
Predictive substructures such as functional groups recur across graphs with different backgrounds, so graphs sharing an explanation could share labels. Existing guarantees require such graphs to carry nearly the same label. We show that label sharing needs only sufficiency, meaning that the rest of the graph adds no label information once the explanation is known. Our learner, pooled ERM, gives each training graph the label frequency of its explanation class, fits the resulting weighted majority targets, and averages fitted predictions within observed classes. Since pooling reuses the fitting sample, the proof separates within-class representatives from class counts. Let be the VC dimension of a binary class across the classes of an explanation map . If , the learner returns with probability a deterministic classifier within of the best risk in from labels, even when the ordinary VC dimension is infinite. At label noise within classes is unrestricted, and the dependence cannot be improved in general. Risk-dependent rates refine this bound, and a rate of order holds under exact sufficiency, Bayes realizability, and a positive margin. With supplied explanations, GNNs trained by pooled ERM improve significantly in every controlled setting, with large gains under label noise. With mined keys on real graphs, pooling only within gated keys raises accuracy significantly on TOX21 and AIDS under a strict label budget and never lowers it significantly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.