acceptodds
Under review as a conference paper at ICLR 2027

Auditing Harness Evolution: Selected Gains, Persistent Edits, and Runtime Behavior

Abstract

An accepted harness edit can earn selected credit without retaining that gain, and can persist without activating its intended behavior. We introduce the RSI Information Ledger to connect disclosure accounting, frozen-state comparisons and execution traces. Across six post-selection 4B coding lineages, the evolved harness gains 1.55 percentage points on the MBPP reserve, with no detectable advantage over an audit-selected single edit at matched candidate-evaluation budget and no detectable positive terminal interaction. An exploratory search under Qwen3.5-4B and a new serving configuration also retains gains; its lineage-minus-single-edit difference is −1.899 points. This comparison matches scheduled candidate slots, with unequal realized scoring and token costs, and establishes neither superiority nor equivalence. Registered eight-seed follow-ups detect neither an acceptance-margin contrast nor a fixed-gate disclosure effect on the audit–reserve gap. In a complementary 27B telecom case, an accepted guard never activates in 3,960 recorded checks. Local wrapper controls yield differing next actions in selected contexts, while synthetic triggers establish conditional blocking and feedback delivery. A subsequent four-condition panel comprises 592 full-task executions on 74 reused tasks. It records zero guard hits and separates content-clearing behavior from the guard; all task-score contrast intervals cross zero, leaving benefit, harm and practical equivalence unresolved. These studies distinguish retained improvement from successive-edit value, and rule persistence from runtime behavior; they do not establish a general causal account of recursive improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.