TACO: A Non-Intrusive RL Framework for Online, Black-Box Agent Services
Abstract
A growing number of deployed agents operate at the service level: they run online as black boxes, managing their own loops, tools, and model providers, with fixed configurations after deployment. Existing reinforcement learning (RL) frameworks, however, generally assume the agent loop can be relocated into the RL runtime, making it difficult to improve deployed agents without rebuilding them. We present TACO, a non-intrusive RL framework that trains such agents directly in their existing execution environments. TACO requires no changes to the agent source code; instead, before the service starts, it uses only a single configurable field the agent passes to the model unchanged (e.g., the model endpoint or API key) to carry a worker identifier and track rollout identity. To improve training efficiency, TACO packs rollout turns using bounded token-level edit distance to avoid repeated context computation, and overlaps rollout with training while draining admitted tasks before weight synchronization. We evaluate TACO on three diverse agents: a ReTool-style math agent, a TritonForge-style kernel agent, and ZeroClaw, a long-running Rust service. Integrating them with TACO requires only 122–334 lines of adapter code. On the multi-turn math agent, packing reduces processed tokens by 3.0×, and combining packing with asynchrony gives a 2.51× end-to-end speedup. TACO is now available on a commercial cloud AI platform as a selectable built-in training image for service-level agentic RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.