Auditing Behavioral Coverage in Model Ecosystems
Abstract
A target model can differ from each of its peers yet be closely approximated by their fixed combination. A similarly small reconstruction error can also depend almost entirely on one peer. We study these distinct sources of behavioral coverage through matched model responses. The Peer-Inexpressible Residual (\PIER) measures what remains after a fixed convex peer fit, and \DISCO estimates the fit and its residual on separate samples. Comparing the group with a selected single peer, and refitting after peer removal, identifies who supplies the coverage. In an eight-model MMLU-Pro audit, our primary response is the reference-answer probability averaged over cyclic option rotations. Six targets benefit from collective approximation under two stress families and matched squared- and absolute-error fitting. Qwen2.5 instead obtains essentially all available approximation quality from its reasoning-tuned sibling; the reverse direction gains a further – from other peers. Mistral's group fit reduces error by –. A separate complete-option-distribution analysis retains the Qwen directionality and Mistral's group advantage. Matched stress increases residuals for some targets and decreases them for others. In vision, context changes both the supporting peer combination and the concentration of residuals on individual images. These results characterize how accurately a collection reproduces a target, which peers supply that approximation, and which inputs account for the remaining gaps.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.