acceptodds
Under review as a conference paper at ICLR 2027

Do Modules Stay in Their Lane? Role Drift in Compound AI Systems

Abstract

When a compound AI system is trained end-to-end with reinforcement learning, the terminal reward only shows how good the final output is, but cannot tell whether each module performed its assigned job. A module can deviate from its assigned role while still improving the terminal reward. We call this failure role drift, and demonstrate it in four compound systems. For example, in a multihop QA pipeline, a decomposer is supposed to pose subquestions and a solver is supposed to answer them concisely. After training, the decomposer begins writing the gold answer into its subquestions, with leakage rising from 14% to 73.5%. Similar drift appears in a retrieval-augmented QA pipeline, a code generation pipeline, and an LLM routing system. Each module is capable of performing its assigned role, but the terminal reward pays more for leaving the role than for keeping it. The drift is invisible in the reward curve and becomes harmful when deployment conditions change: when the leaked answers in the decomposer's subquestions are masked, the system's performance drops by 8.9 points. Monitoring the terminal reward alone is not sufficient to ensure that a compound system's modules behave as intended.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.