Kepler: Auditable World Models for ARC-AGI-3
Abstract
Final scores can conceal invalid behavior in interactive-agent benchmarks. We present Kepler, an ARC-AGI-3 harness in which agents express hypotheses as executable world models. Its protocol combines complete-history retrodiction, conditional next-observation checks, an append-only transition ledger, and mechanically replayed scoring. Checked mismatches interrupt plans; missing predictions do not block execution. Across two frontier coding models, 48 of 50 game-model cells reached the score cap. One frozen Claude Opus 5 configuration obtained a server-verified 100.00 over all 25 public games. On 181 of 183 completed levels, its final attempt used no more actions than the median-human baseline. This measures final-attempt efficiency, not learning cost or human-like cognition. Trajectory audits exposed failures hidden by final scores: an invalid perfect run after source access, harness reconstruction by all six agents in an intended control, instruction rewriting in 26 of 26 workspaces, and silent replacement of a broken planner. A single-game visual continuation resolved a mechanic missed during earlier text-only sessions. A code regression suite detects 11 of 13 hand-built threat fixtures and flags none of five benign controls; it misses unlogged behavior and a semantically wrong but runnable tool. These findings motivate reporting trajectory integrity, selection rules, resource accounting, and executable belief checks alongside outcome scores. They do not establish held-out generalization or a causal harness advantage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.