CAGE: Blockwise Policy Gradients for Explicitly Routed Language-Model Workflows
Abstract
Learning from terminal rewards in routed language-model workflows requires two distinct comparisons: between routes to learn which workflow to select, and within routes to learn how to execute it. Small execution budgets make both comparisons difficult: independent sampling can omit routes, while global centering mixes route-level reward differences into execution credit. We introduce Credit Assignment with Graph-aware Estimation (CAGE), which allocates executions across routes before observing rewards. It sums the Router gradient over the finite route set and uses a route-local leave-one-out baseline for the Specialist. Live-policy weights correct for the allocated sample counts. Under independent conditional sampling and sampling-consistent scores, both raw estimators are conditionally unbiased for the original current-policy objective. We establish deterministic route coverage, invariance of Specialist credit to route-specific reward offsets, and removal of the additional categorical sampling variance of a same-batch sampled-Router estimator. An exact covariance decomposition quantifies the noise added by estimating the Specialist's local baselines from finite samples. A fixed-policy diagnostic confirms the coverage mechanism: IID sampling misses routes on 29 of 32 questions, whereas CAGE covers every route. A one-question full-LoRA-gradient check on a frozen 4B model also verifies samplewise route-shift invariance. These diagnostics support the coverage and local-credit mechanisms; an end-to-end reward advantage remains unestablished. CAGE estimates both route-selection and route-execution gradients from the same fixed complete-workflow budget while preserving the current-policy objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.