acceptodds
Under review as a conference paper at ICLR 2027

Task Coverage or Repeated Exposure? A Study of Learning from Execution Feedback

Abstract

Task-set scaling in execution-feedback post-training changes both task coverage and repetition under a fixed budget, so nominal task count does not specify the training allocation. We study this problem with Qwen3.5-4B and group-normalized policy gradient (PG) on nested Structured Query Language (SQL) and arithmetic task sets of \(N=64\), \(256\), and \(1,024\). We compare rollout-count and generated-response-token endpoints, measure realized task coverage, generated tokens, updates, wall-clock time, and acquisition cost, vary task composition and reference complexity, and evaluate six adaptation methods on reserved final tasks under a registered 7,200-second pipeline budget. \(N=256\) gives the highest development mean, while \(N=1,024\) does not continue the gain and the original quantity endpoints are close on the reserved test. Rollout-count and token endpoints produce different exposure and cost, task composition and reference complexity change outcomes at fixed \(N\), and no tested adaptation method exceeds the unadapted checkpoint under the common time budget. These results establish that task-set scaling is a training-allocation claim: interpretation requires separating new task coverage from repeated feedback and accounting for endpoint and resource costs. This allocation diagnosis gives researchers a basis for deciding whether limited budget should purchase new tasks, repeated feedback, or further adaptation, and for judging whether an observed gain transfers when the budget or endpoint changes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.