BitRouter: Learning Interpretable Model Routing Policies for Agentic Workflows from Deployment Feedback
Abstract
Agentic workflows involve state-dependent model choices whose combined effects are often evaluated through delayed, session-level outcomes. Building on bandit-based routing and cost-subsidy formulations, we introduce BitRouter, a framework for learning interpretable workflow routing policies under cost and quality requirements. Each complete policy is a bandit arm retained throughout a session, while explicit rules adapt model choices to workflow state. We reuse session-level evidence across policies with shared rules, accounting for uncertainty in the rules’ combined effects relative to quality-feasibility margins when deciding whether additional policy-specific evaluation is needed. Across repeated runs on 80 common-valid Terminal-Bench 2.1 tasks, an initial routing configuration achieved 81.25% task success, versus 81.56% for a fixed strong-model baseline and 76.50% for same-pool random routing. Mean nominal API cost per trial, counting accepted valid-path requests, was 40.93% lower than that of the strong-model baseline. End-to-end learning experiments show lower combined execution and verification costs than independent policy evaluation, while keeping the observed rate of selecting policies that fail to meet the quality requirement within the prescribed limit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.