acceptodds
Under review as a conference paper at ICLR 2027

Code-to-Harness: Learning Language Policies through Executable Practice

Abstract

Learning reusable policies for frozen large language models (LLMs) raises a representation question: must a policy be discovered in the same form in which it is deployed? We introduce Code-to-Harness, which uses executable programs as the discovery representation and compact instructions as the deployment representation. A practice agent proposes and revises optimizer programs from executable feedback; a development competence screen determines which completed runs are eligible for promotion; and a one-shot in-context synthesis compresses the accumulated executable practice into an approximately 200-word frozen policy. Only policies from promoted runs enter the primary transfer evaluation, and no source program is executed at deployment. Across ten completed discovery runs, four satisfy the recorded promotion criterion. On 30 genuinely fresh objective instances per task, promoted policies reduce mean regret relative to the unaided executor by 44.4–60.7% on quadratics and 32.5–50.0% on Rastrigin. One fixed policy attains lower mean regret than all five numerical baselines on eight of ten BBOB families, and all four promoted texts improve a second vendor's executor on both core tasks. On three LCBench/YAHPO-Gym HPO tasks, frozen policies attain the lowest reported mean on two datasets. A five-batch audit further shows that the recorded threshold behaves as a conservative competence screen rather than a calibrated predictor of downstream text quality. Together, the results show that useful search-strategy information can cross an executable-to-language representation boundary without requiring the discovery program at deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.