Learning to Ask for Help: Budget-Constrained Decision Routing via Dual-Advantage GRPO
Abstract
Small language models (SLMs) on devices answer queries locally and request cloud assistance when needed. Each cloud request consumes resources, motivating a prescribed allowance for collaboration. Existing collaborative post-training methods use a fixed assistance reward, but the resulting cloud usage can drift as local competence changes. DA-GRPO treats expected cloud usage as a constraint: task reward and cloud cost yield two group-relative advantages, combined by a multiplier updated from measured usage relative to the target. The method retains the clipped GRPO surrogate and adds no measurable training overhead in our configuration. On Qwen2.5-1.5B-Instruct and Llama-3.2-3B-Instruct across mathematical reasoning, code generation, and knowledge benchmarks, DA-GRPO achieves the highest mean joint accuracy in all eight settings under common task-specific training budget targets, improving on the strongest fixed-reward baseline by 1.8 to 7.0 percentage points. In a separate budget-control experiment, mean usage is 0.313 at the end of training and 0.314 at in-distribution test time against a target of 0.30. Cloud usage increases from 0.10 on the easiest MATH level to 0.49 on the hardest. The same framework supports changing targets during continued training and cloud-token costs, while test-time usage under distribution shift can exceed the training target.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.