acceptodds
Under review as a conference paper at ICLR 2027

Learning What Not to Learn: Target-Conditioned Update Eligibility for Self-Improving Agents

Abstract

Self-improving large language model (LLM) agents distill interaction feedback into persistent memory to enhance future capabilities, but local task success alone does not warrant persistent reuse. We formulate target-conditioned update eligibility to assess whether observable source evidence supports admitting a feedback span under specified persistence and reuse contracts. Controlled interventions demonstrate contract dependence under construction-defined labels and show admission changes learner states, while downstream utility varies across reuse settings. To examine current evaluation paradigms, we develop a Span-PRM that passes construction-defined development tests but fails confirmation against fresh human judgments. It admits 0 of 300 generated reflections at the frozen threshold under a cross-task rule contract and achieves 37.5% ranking agreement with human majorities on a separate challenge set. A post hoc diagnostic reveals a fixed contract ordering satisfies all 120 human-majority comparisons without processing feedback or evidence, demonstrating success on this ranking target alone cannot establish evidence-sensitive judgment. The frozen model fails all three prespecified tests in a prospectively frozen counterfactual audit against construction-defined references. Our findings highlight vulnerabilities in self-improving agent evaluations, showing local metric optimization alone does not establish evidence-based reasoning and that construction fit, reference-aligned input sensitivity, human agreement, and downstream utility must be validated separately.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.