acceptodds
Under review as a conference paper at ICLR 2027

Group-Relative Skill Optimization for Frozen Language Models

Abstract

Optimizing external natural-language skills has emerged as an effective approach to adapting frozen large language models. Existing methods have improved task performance and stabilized skill updates through held-out validation and incremental editing. However, we identify two limitations of batch-level skill optimization: (1) common-denominator bias, where reflection over trajectories from different questions favors generic guidance over question-specific knowledge; and (2) all-or-nothing acceptance, where edits from different questions are accepted or rejected as a single update, making it difficult to retain useful edits while rejecting harmful ones. We introduce Group-Relative Skill Optimization (GRSO), which aligns reflection and verification at the question level. For each question, GRSO samples a group of trajectories and uses within-question reflection to propose itemized edits, contrasting successful and failed attempts. Each edit bundle is verified on its source question before accepted bundles are merged. Dynamic sampling and retry concentrate additional computation on difficult questions. This design preserves question-specific learning signals without held-out validation during optimization. Across SearchQA, DocVQA, and LiveMathBench, GRSO achieves the best or tied-best accuracy in all six target–benchmark settings. Compared with the unadapted models, it improves average accuracy across the three benchmarks by 15.2 and 22.0 percentage points for Qwen3.5-4B and Seed1.8, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.