Finding Is Not Certifying: Evidence Reports for Redundant Mechanisms
Abstract
Finding a model's minimal harmful ablation groups, ruling out additional groups under an assumption, and predicting untested interventions are different tasks. We develop evidence reports that keep these claims separate. A conditional rank frontier states which orders of additional monotone interactions have been excluded by the observations, without claiming to infer the model's true interaction order. Sharp bounds for restricted response classes show that a small record can identify an explanation under monotonicity while remaining fragile even to responses with one minimal harmful set and one recovery per intervention chain; a bound on the total number of descents changes the scaling. Fragility warnings describe compatible alternatives, not actual prediction error. We identify observed-only warnings with elementary proof completions and treat neighborhood optimization as optional rather than a default validation method. A 36-seed tiny-transformer study evaluates the actual finite-candidate routine against cheap matched controls, while overlapping-reference tests separate probability-counting, query, and optimization costs. A pretrained-model reanalysis distinguishes prompt uncertainty from mask-pair uncertainty. The resulting recommendation is to report conditional coverage, retain concrete contradiction witnesses, and use independent behavioral validation for acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.