acceptodds
Under review as a conference paper at ICLR 2027

GET: Direct Peer Evidence for Agentic Self-Distillation

Abstract

Outcome-supervised reinforcement learning provides language agents with limited guidance for intermediate decisions. On-policy self-distillation (OPSD) supplies token-level supervision from privileged context, making the selection and organization of decision-relevant experience a central design question. We introduce Group Evidence Teacher (GET), which organizes observation-matched peer experience into a decision-local reference alongside immediate feedback on the student's own action. For each decision, GET selects transitions from other rollouts of the same task and deterministically compiles their actions and source episodes' outcomes into an Evidence Card, without generating skills or corrective guidance. During training, the same policy re-scores the student-sampled response using a decision-local context built from available action feedback and peer evidence; the resulting token-level log-probability differences determine the weights of an auxiliary update alongside group-relative reinforcement learning. On ALFWorld, GET achieves mean success across three seeds. Adding peer evidence to action feedback yields a -percentage-point gain in task success in a seed-0 component comparison; environment-verified replay of these policies shows a -point gain in required object-state changes across cleaning, heating, and cooling tasks. On Search-QA, GET outperforms Skill in both micro- and macro-averaged exact match across both backbones and weighting rules. With Qwen3-1.7B and Raw OPSD, GET achieves higher test accuracy with fewer search calls per question than Skill.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.