SepaRank: Measuring Intelligence Beyond Human Scale
Abstract
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. We develop a novel, fully endogenous protocol that naturally scales with agent capabilities, requires no human-authored questions, labels or judgements, and is also judge-free. We instantiate this framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.