acceptodds
Under review as a conference paper at ICLR 2027

KS Gov: Long-Running AI Discovery and Adversarial Testing from a System Prompt

Abstract

Users ask coding agents to make a system faster or more robust and let them run for hours. Such a run outlives the context window, and one untested change can undo a day’s work. KS Gov is a coding agent built for such runs: five Python modules and 4,449 significant lines under a 3,851-word system prompt. The prompt states engineering rules and two procedures: an AI discovery and optimization loop that logs tried ideas to files, and adversarial testing by alternating break and fix sub-tasks. One task-prompt sentence adds a read-only reviewer from a second vendor. In two long runs the agent built a key-value store that sustains 5.50 Mops/s on YCSB-A over 17.2 agent hours, and made a TPC-H engine 34.5× faster than its single-threaded artifact. The first ran under an earlier one-paragraph discovery instruction, with the adversarial steps written in its task prompt; only the second ran under the seven-step loop. In both, the reviewer found defects the benchmark hid. Each case study is one run without a baseline agent. Terminal-Bench 2.0 tests the loop and tools without the two procedures: under a shorter coding prompt on the HarnessTax study’s 30 tasks and seven models, the agent solves 75.6% of attempts (74.1% counting solves past the study’s 100-turn cap as failures) against 70.0% for Pi, the best published harness, a success gap within sampling error. Against Pi rerun by us on one model, the paired success gap is +11.1 points, at a cost we cannot tell apart.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.