Black-Box Attribution under Domain Shift: When Are Generative Mechanisms Identifiable?
Abstract
We study when a black-box behavioral response that separates generators on a source domain remains valid attribution evidence after domain shift. We distinguish source recognition, population identifiability over a declared nuisance family, and finite-development certification. Response-family overlap gives a global testing obstruction. In a structured common-location model, finitely observed domains cannot certify unseen-domain attribution without a coverage assumption; under declared linear coverage, only the nuisance-orthogonal mechanism contrast survives, with an exact binary Gaussian benchmark and a finite-development recovery bound. In a controlled training-parameterization stress test, models are 9/9 separable within MNIST and Fashion-MNIST but 3/9 under frozen transfer. A nuisance projection reaches 17/18 on the development domains yet falls to 4/9 on held-out KMNIST. Deployed diffusion, Flow Matching, and consistency pipelines remain direction-dependent even with a frozen S-CLIP representation (8/9 versus 5/9). Synthetic checks recover the predicted Gaussian risk and coverage-rank behavior. Our claim is methodological: attribution validity is relative to a declared nuisance family, and finite source domains do not certify unseen
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.