acceptodds
Under review as a conference paper at ICLR 2027

CapAtlas: Resolution-Specific Validity of Sparse Behavioral Model Maps

Abstract

Thousands of language-model checkpoints are poorly documented, yet comprehensively benchmarking every model is impractical. Sparse model maps offer a cheaper alternative, but current evaluations collapse map quality into one score and leave their downstream validity unclear. We introduce CapAtlas, a leakage-controlled protocol that learns a capability map from documented references, freezes it, places an unseen model from 16 response distributions, and only then unseals its evaluation profile. We evaluate a 46-model atlas on 70 targets from 43 unseen lineages, using 84,000 held-out response records and 109,480 sealed oracle outcomes for evaluation. Reference supervision improves graded local retrieval from 0.8181 to 0.8609 lineage-macro NDCG@5 (+0.0428; 95% CI [0.0229, 0.0640]; random floor 0.6048), with a positive effect from checkpoint- to architecture-level aggregation. Six calibration summaries reach 0.8611, showing that calibration explains the aggregate score while itemwise geometry contributes selectively by domain. Under a disjoint capability construct, global ranking improves by 0.1691 Kendall's tau; the registered local-retrieval gate identifies a separate neighborhood-calibration target. On 50 prospective targets, a frozen reliability score stratifies placement risk (AUROC 0.731, 95% CI [0.582, 0.866]), and its 80.95%-precision operating point quantifies a 9.05-point gap to the registered 90% standard. CapAtlas turns sparse model mapping into an auditable instrument whose claims are matched to their decision resolution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.