From Past Attempts to Better Decisions: Generalizing Verified Experience through Parametric Memory
Abstract
A language model's first attempt on a difficult reasoning problem is often wrong. Correct answers may emerge only after repeated sampling or revision. Yet the failed and successful attempts are then discarded, so the next problem starts from scratch and the same mistakes can recur. We ask which lessons from this trial-and-error history are useful, how to store them in model parameters, and whether they generalize to unseen problems. We introduce BVEd, a parametric experience memory for frozen solvers. Starting from repeated training attempts, BVEd removes final predictions and copied answer text, then uses a teacher to distill successes and failures into candidate lessons about useful reasoning steps, missed constraints, and recurring mistakes. Source-side screens measure whether lessons increase support for the known answer, limit answer steering without the problem, and detect copied answer cues; at most one training target is selected from each admitted source. The selected lessons train a helper to transform a new, label-free problem into fresh problem-specific guidance, which the unchanged solver uses to produce its answer. Across four reasoning benchmarks, BVEd yields absolute accuracy gains of –. Together, these results suggest that trial-and-error histories can supply useful supervision: BVEd extracts experience from past attempts, consolidates it into model parameters, and generates guidance for previously unseen problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.