The Protocol Is the Agent: Decomposing RL, SFT, and Scaffold Effects in Tool-Use Agents from 0.8B to 27B
Abstract
On a single frozen instrument (the same 100 financial tool-agent tasks, execution harness, and analyzer, with three evaluation seeds per arm and pooled paired exact McNemar tests), we decompose the measured task-completion capability of one Qwen3.x family (0.8B, 2B, 9B, 27B, plus a 4B case study) into protocol, supervised-fine-tuning (SFT), and reinforcement-learning (RL) contributions. Three findings. (1) A consistent RL null: the same GRPO recipe with verifiable structural rewards changes nothing at any scale (+0.7 at 0.8B, p=.75; +2.0 at 2B, p=.27; 0.0 at 9B, p=1.0; -0.7 at 27B, p=.64), including the scale with the largest SFT headroom. We further document how a shared-but-wrong evaluation template alone manufactured p=.024 in our own 2B cell before correction. (2) The protocol dominates: switching the surrounding stack from a text-protocol scaffold to native tool-calling is worth +59 points on the same 2B SFT weights, versus +4.3 for SFT itself and +2.0 for RL. (3) SFT's contribution is non-monotone and can be catastrophic: the identical recipe adds +18 points at 0.8B and +11 at 9B but destroys 45 points at 4B by inducing temperature-dependent format fragility in an otherwise format-clean base, after which RL repairs the format while capturing a rewarded shortcut. Corollaries: measure the protocol before the weights, the headroom before the RL budget, and never compare training arms without tokenizer-exact evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.