acceptodds
Under review as a conference paper at ICLR 2027

Adversarial Self-Play for Scientific-Computing Agents: An Arena Where a Model Writes Finite-Element Problems It Cannot Solve

Abstract

Static benchmarks for tool-using scientific agents saturate quickly and are vulnerable to contamination, while human-authored problems are expensive and rarely calibrated to the frontier of what an agent can do. We introduce an adversarial, zero-sum arena in which a single language-model agent plays both sides of a game about numerical finite-element computation: a Generator designs a multiple-choice question, together with a ground-truth value, three distractors, and four verification scripts that independently reproduce every option with FEniCSx on a fixed, tagged mesh; a Solver, the same model, must answer the question, first without tools and then with a sandboxed Python/FEniCSx environment. The arena executes the generator's scripts, forfeits any round whose question is not fully verified, presents each question under four cyclic option orderings, and requires the solver to be correct on all four. Every invocation runs under fail-closed filesystem, CPU, and memory isolation, and every call is metered for time, tokens, and cost. We report a match of rounds played by a single frontier coding model. Of the verified questions, the solver answered and failed , despite a per-attempt tool-solver accuracy of . The generator's strategy visibly evolves: median question length grows from k to k characters and questions progress from single PDEs to chains of six to eight coupled sub-problems; the no-tool solver, meanwhile, exploits structure in the distractors to answer of verified questions without computation. We analyze these dynamics, the forfeit rate caused by the generator's own verification scripts, and the cost profile of the match ($1.4k, 5.9B tokens). The arena yields a self-refreshing, verifiably correct, and difficulty-calibrated question bank, and a companion evaluation mode replays that bank against agents that must implement the finite-element method themselves. Code, prompts, and the full match log are released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.