acceptodds
Under review as a conference paper at ICLR 2027

PDSA: Separate Policy and Memory State for Resource-Bounded Adaptation

Abstract

An adaptive agent can improve by retaining facts, changing its policy, or trying more answers. Predictive Dual-State Adaptation (PDSA) separates bounded memory from an independently restorable numerical policy and proposes action-conditioned future utility for admission. We study controlled arithmetic and list-program repair with a frozen shared semantic base, public execution feedback, and hidden audit cases excluded from updates and selection. Five studies examine shared semantic qualification, candidate/update budget, calibrated update strength, policy-carrier design, and feedback resolution. Larger candidate budgets with evolving public feedback improve selected answers equally in updating and frozen regimes; neither calibrated strength nor an internal carrier establishes its prespecified feedback-specific gain. In the final 32-problem comparison, dense feedback distinguishes every active candidate batch, drives 83 optimizer steps across problems (at most three per problem), and changes attained-prefix distributions and most candidate streams. Yet dense-learning audit loss is 0.703125, versus 0.7015625 for frozen selection, 0.653125 for exact-reward updates, and 0.73125 for dense cyclic mismatch; none of the registered dense-learning contrasts meets the materiality and uncertainty rule. Configured search reaches 0.028125. Subsequent post-hoc CPU scoring finds no missed audit-perfect solutions in the saved pools: all four methods retain only two fully solved problems under pool-best selection. Pool-best mean loss remains 0.609375–0.684375 on the same audit cases, locating substantial residual error in observed candidate coverage. Separately, admission improves over no-write but not a strong rule. These results distinguish feedback variability, policy reuse, candidate coverage and selection from incremental task benefit; learned predictive admission and full PDSA advantage remain unestablished.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.