Observation Limits of Node Attribution under Graph-Level Supervision
Abstract
Node attribution ranks nodes by their relevance to a graph-level outcome, supporting applications such as localizing faulty program statements from test outcomes. Existing evaluations emphasize faithfulness to a trained model, leaving unresolved whether the observations distinguish the responsible node from other candidates. We study the limits that available observations impose on node attribution in real software faults. Our analysis characterizes attainable localization accuracy, relates indistinguishability under shared message passing to equal node scores, and gives upper confidence bounds from observation counts collected separately for each candidate target. To assess the prevalence of this ambiguity, we enrich test coverage with local program descriptors and identify repaired and non-repaired statements that remain indistinguishable under these observations. Such pairs occur in 34.8% of eligible Defects4J faults and 33.1% of eligible Codeflaws faults. The three evaluated architectures and three coverage-based scoring formulas tie on every matched pair, yielding pair AUC 0.50. We then separate the effects of target information and training supervision by independently varying node labels and a node feature encoding oracle target information. On structurally interchangeable EPFL circuit nodes, this target anchor alone gives mean AUC 0.980, compared with 0.511 for node supervision alone, while combining both gives mean pair AUC 1.0. These results connect ambiguity in real workloads to the information required for localization and motivate evaluating node attribution separately from graph-level prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.