acceptodds
Under review as a conference paper at ICLR 2027

A Verifiable Environment for Agents That Design Neural Architectures

Abstract

Reinforcement learning and agent evaluation both need a verifier: a function that scores a candidate cheaply, with no human in the loop. While software development relies on compilers and mathematics on automated answer checkers, neural architecture design has lacked a deterministic, public verifier. We present an environment in which an agent designs a neural network by editing a typed graph, and a deterministic verifier grades the result in under a millisecond against structural, budget, and serving-cost constraints. No human and no LLM takes part in scoring. The environment includes a procedural task generator (ten families, with satisfiability and non-vacuity checks and red-teamed anti-gaming rules), a difficulty-calibration harness, and a grounding study. In that study, across architectures, the verifier blocks every corrupted graph (), no blocked graph survives a real PyTorch forward pass (), and every clean graph runs (). The to health score is a validity margin, partly a size proxy, and not a quality ranking. We report four findings. As a repair signal, verifier feedback lifts claude-sonnet-4-6 from to , grok-4 from to , and deepseek-chat from to ; the lift concentrates on arithmetic defects, while cross-branch topology is largely unfixed. As a correction to how design agents are scored, a rubric that does not propagate shapes rates the same models and points higher than they are. As a training-data generator, fine-tuning a 1.5B model on verified pairs lifts pass@1 from to on held-out tasks from the same ten families, and GRPO steps on top reach (). Under a stricter rubric the same recipe runs , and a 7B policy starts lower than the 1.5B (, because it is more verbose) and ends at . Transfer is not general: four families held out of training collapse to . A contamination audit we ran on ourselves shows the naive protocol would have reported . As ground truth, eleven LLM reward models have a false-positive rate on blatantly broken designs (ten of eleven when re-measured), yet approve to of near-miss designs that are one edit from correct. The verifier is on both tiers. The same function serves as a leaderboard, an RL gym, and a verified-data generator. We release all of it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.