S³Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Abstract
Large language models (LLMs) accumulate interaction experience, but this experience does not necessarily improve their subsequent behavior. We introduce S3Gym, a benchmark connecting Self-Testing, Self-Judging, and Self-Improvement across seven text-based games. The protocol separates permissive exploration from strict held-out evaluation and uses executable environment verifiers to audit self-judgments and performance. We compare History ICL and score-conditioned Summary Memory across seven proprietary models, and separately study parameter training with Qwen3-8B. The results reveal model- and task-dependent benefits rather than a universally effective experience representation. Self-Judging is only partially reliable: agreement about positive rewards can coexist with large normalized magnitude discrepancies, and local judgment accuracy does not reliably predict subsequent gains. Summaries do not consistently outperform direct history. Selected cases suggest that explicit, state-dependent action rules grounded in observed outcomes provide more useful iteration directions than generic reminders. Parameter training improves some tasks but also produces transient gains and persistent degradation. Late-stage declines may reflect overfitting to exploration patterns and reduced generalization to held-out conditions, although the experiment does not establish this mechanism. These findings motivate evaluating judgment quality, actionable experience reuse, and retained performance separately. SGym supports this diagnosis by distinguishing accumulated experience from measurable improvement in held-out behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.