SE-Eval: Verifiable Benchmark Evolution for Adjacent Model Releases
Abstract
Static benchmarks can lose resolution precisely where rapid model development needs it most: between adjacent releases near ceiling. We introduce SE-Eval, a programmatically grounded loop that generates candidate items, verifies unique solutions, measures an ordered model-version sequence, and selects a bounded item pool for the next round. On three pre-specified Qwen seeds, SE-Eval increases adjacent-release disagreement over a frozen pool by +0.0451 [+0.0251, +0.0660], a repeatable but sub-minimum-detectable-effect gain. A fresh-seed replication analyzed under a protocol fixed before accepted outcomes reveals a family-dependent boundary: +0.0396 on Qwen, -0.0893 on DeepSeek, and +0.1673 on Kimi with a family-disjoint generator. In a saturation diagnostic, the three newest Qwen releases score 100% on frozen items but only 83-93% on generated items, showing retained frontier headroom. A scoring-fault diagnostic shows why resolution alone is insufficient: placeholder labels reduce true-gold boundary disagreement from 0.56 to 0.28 while still producing smooth trajectories. Together, these results establish self-evolving benchmarks as longitudinal measurement instruments whose value depends on transition structure, verifiable grounding, and validation across model families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.