Task Coverage or Repeated Exposure? A Study of Learning from Execution Feedback
Abstract
Task-set scaling in execution-feedback post-training changes both task coverage and repetition under a fixed budget, so nominal task count does not specify the training allocation. We study this problem with Qwen3.5-4B and group-normalized policy gradient (PG) on nested Structured Query Language (SQL) and arithmetic task sets of \(N=64\), \(256\), and \(1,024\). We compare rollout-count and generated-response-token endpoints, measure realized task coverage, generated tokens, updates, wall-clock time, and acquisition cost, vary task composition and reference complexity, and evaluate six adaptation methods on reserved final tasks under a registered 7,200-second pipeline budget. \(N=256\) gives the highest development mean, while \(N=1,024\) does not continue the gain and the original quantity endpoints are close on the reserved test. Rollout-count and token endpoints produce different exposure and cost, task composition and reference complexity change outcomes at fixed \(N\), and no tested adaptation method exceeds the unadapted checkpoint under the common time budget. These results establish that task-set scaling is a training-allocation claim: interpretation requires separating new task coverage from repeated feedback and accounting for endpoint and resource costs. This allocation diagnosis gives researchers a basis for deciding whether limited budget should purchase new tasks, repeated feedback, or further adaptation, and for judging whether an observed gain transfers when the budget or endpoint changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.