acceptodds
Under review as a conference paper at ICLR 2027

A HARNESS CAN REPLICATE AND STILL REGRESS: TASK-CONDITIONAL CAUSAL ADMISSION FOR SELF- IMPROVING AGENTS

Abstract

Self-improving agents often deploy harness changes globally after observing gains on a small development set. This can overfit twice: selection overfitting can mistake noise for improvement; scope overfitting can turn a real local improvement into harm on other task types. We introduce Task-Conditional Causal Admission (TCCA). It evaluates each candidate harness component change together with a frozen deployment rule. TCCA uses controlled comparisons with equal compute and tool budgets to estimate task-conditional causal effects. Same-arm replicates calibrate noise; a holdout set tests transfer; and harm limits protect prespecified task types. TCCA admits global deployment only when supported by the evidence. In controlled synthetic studies, the same context-ordering change increased named-record lookup accuracy by 70.8 percentage points but reduced population-mean accuracy by 66.7 points. When pre-treatment task features predicted effect direction, a frozen admission rule recovered 98.8% of oracle policy value; without such predictability, it provided no held-out benefit. These results show why replication alone cannot justify global deployment and how task-conditional admission can preserve supported local gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.