acceptodds
Under review as a conference paper at ICLR 2027

Matched Twins Show What an Evidence Contract Cannot Carry to a Verifier of ML-Experiment Claims

Abstract

An ML-engineering agent reports that adding one engineered feature improved its validation score. A verifier bounded by the trace evidence — the claim, the reporting discipline, the outcome of plain re-execution, and the agent's one-line description of the change — must decide whether the gain is real. We construct matched twins: pairs of executed runs on the same task whose sentence, model types and budget block are identical, whose five-seed mean gains agree within , and which differ only in whether the added feature is a label leak. Scope-controlled re-execution separates the twins; the contract's fields carry almost no signal (the best single visible field reaches AUC 0.67, and it is the seed variance, which we did not manage to match: larger on the honest twin in 12 of 14 pairs). Two frontier reader models, 189 blind readers each on identical prompts, the primary reading fixed before dispatch: neither separates the pair. One flags leak and twin at 11/14 and 11/14 (exact ; bootstrap 95% interval on the difference to ), the other at 5/14 and 3/14 (; to ); with 14 pairs the design can exclude only a gap above . A second, secondary result is an intervention on one field for one reader: the agent's sentence, which carries no truth value because the twins share it, moves that reader's verdicts on corrupt and honest traces alike — "add one engineered feature" raises rejection of the leak (1/14 to 11/14) and of its matched twin (2/14 to 11/14); the tree-count sentence, worded truthfully for each class, lowers rejection of a starved baseline (18/18 to 4/18) and of an honest budget change (9/22 to 3/22). For the second reader the sentence does nothing; removing the budget block instead drops its rejection of the honest budget change (22/22 to 6/22) and of the starved baseline (17/18 to 9/18) and removes every unfair-baseline call, an ablation run after the replication and registered before dispatch. The replication did not hold as registered: two of its four readings replicated (twin non-separability and the legible classes) and two did not (the sentence effect and the honest false-positive bound). A test-retest of the primary reader two days later bounds what an unpinned reader can support: it agrees with itself on 25 of 46 prompts (kappa 0.15, below the registered bar of 0.6); the tree-sentence effect reproduces exactly (18/18 to 4/18 on both days) while its flag rate on the matched pairs falls from 11/14 to 3/14 with the sentence and to 0/14 without, so the feature-sentence effect keeps its direction and loses its size, and the same-day contrast reads indeterminate. A registered prediction that surfacing the labeler's two scope-controlled re-execution deltas inside the contract would make the pair legible failed: with a leak-free delta of printed on every leak, the primary reader flagged the leak 0/14 against 7/14 in a same-day control (), reading the field as confirmation of the gain. Re-encoded as a comparison, the plain gain beside the scoped gain and their gap, the same re-executions made the pair fully legible to the same reader (leak 14/14, matched twin 0/14, ) and the starved baseline too (18/18), while turning the honest budget change into a false-positive class (3/22 to 17/22, fourteen called unfair on gaps of at most 0.01 or none at all). Replicated the same day with a fresh control and under the second reader, the separation held for both (14/14 to 0/14; 14/14 to 2/14) and the magnitude was read in opposite directions: the primary reader called the honest budget change unfair again (18/22), the second read the gap's size and stopped calling it unfair (20/22 to 6/22). The pair rates took a new value on each of four days while the budget classes reproduced task for task on all four. Under every encoding that omits the scoped re-execution the matched pair is non-separable on every day; under the one that carries it as a comparison it is separable on the day we ran it. We reached this measurement by catching ourselves three times, each time with a pre-registered blind pilot on the same corpus: a 0.909 false-positive rate that was a contract encoding defect, a leakage recall that swung from 0/21 to 15/21 on a corpus artifact, and a 0.90-against-0.59 twin separation that was claim magnitude. The corpus (242 executed traces, 22 tasks, 11 scripted proposers) is auto-labeled by counterfactual re-execution. A deterministic contract-bounded stand-in, not an LLM, reaches F1 0.82 at precision 1.0 on it while never firing on leakage, a class its rule set does not encode; in candidate selection the stand-in filter with the standard claimed-gain ranking falls to of the oracle gap while re-executed credit closes 0.88. No funded API call was made; the readers are unpinned session models, and one reader model violated the no-tools instruction in two dispatches in a way we audit and disclose.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.