Earned Redundancy: Auditing LLM Skill-Promotion Quorums
Abstract
Promotion gates combine language-model reports to decide whether reusable agent skills can be admitted. Different model names, prompts, or endpoints do not establish that agreement reduces malicious admission while retaining benign skills. We introduce an intervention-conditioned audit that separates observer identity from acquired evidence, measures joint failure, and tests decision benefit against a declared reference under an explicit missing-report policy. On an identity-disjoint panel of 120 malicious packages, the sole preregistered dependence contrast yields (95% interval , one-sided ). Yet the least dependent pair jointly clears 14/120 packages, versus 13/120 for the more dependent prompt pair, because their error margins differ. In calibrated simulation, target clones reverse the sign of proxy-adjusted dependence. A paired evidence-view audit changes joint clearance from 16/28 to 0/28 on an outcome-selected cohort. A benign-inclusive Claude–Qwen extension finds lower raw cross-family correlation but no resolved balanced-accuracy gain for the calibration-selected rule, with the dependence advantage sensitive to marginal-rate normalization. Redundancy credit therefore requires a declared intervention, evidence view, and measured risk-utility improvement against the reference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.