Auditing Model Fusion under Asymmetric Loss
Abstract
Systems that fuse several models under an asymmetric operational loss usually summarise each model by a scalar: a hard vote, or a reliability weight in place of a likelihood. We ask three questions about that practice, and answer each with a finite-sample object. Is it the scalar that hurts? A matched control fuses hard labels through one code path and varies only how each model is described: its fitted confusion matrix, against the symmetric channel carrying that matrix's own accuracy. This separates error structure from the change of representation that confounded earlier comparisons. The scalar description loses on all seeds of the primary domain, by a paired risk-weighted loss; the secondary domains show smaller or more sample-sensitive effects. Can the evaluation tell? An evaluation can be quantised more coarsely than the effect it measures. For a separate, smaller contrast against weighted averaging, we distinguish observed ranking sensitivity from uncertainty using an exact leave-one-out diagnostic— tighter than the bound it replaces—and a retrospective empirical-Bernstein calculation. What can a deployed rule certify? We give a one-sided finite-sample certificate that controls, at level , the probability of falsely declaring a deployed rule worse than a fixed comparator, built on where in the loss matrix the rule sends the rare class, and costing only the calibration split rather than the rare rows alone. It fires on all seeds of the decisive domain under the weaker of two concentration bounds. Where it cannot fire it abstains, and we report the abstention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.