acceptodds
Under review as a conference paper at ICLR 2027

When Is Program Structure Required? Certificate-Grounded Calibration of Program-Induction Difficulty

Abstract

Program-induction benchmarks require a distinction between the structure of a generator and the structure demanded by its examples. We introduce certificate-grounded calibration for programming-by-example: complete one-rule enumeration and a two-rule witness certify a change in the shortest legal solution from one rule to two. Paired specifications share a reference program, contain six examples each, and exactly match five edit features of the changed example. On 120 new validation families, Qwen3-8B and Olmo-3-7B-Think exhibit displayed-task success decreases of 45.83 and 30.83 percentage points alongside gains of 7.50 and 9.17 points on a common nine-example audit. A supplemental analysis of a shared, wholly unshown witness pair finds gains of 9.17 and 11.67 points. An independent bounded symbolic solver solves all controls with one-rule programs, whose unseen-audit errors follow from the construction; on separation, it solves 95/120 tasks and passes the unshown audit on 87/120. These results distinguish a certified requirement, the cost of finding and delivering a solution, and agreement with a specified unseen target. The witness gains do not establish improved accuracy on broader rule-active inputs; separate readouts locate failures in final-answer delivery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.