When Do Behavior Traces Validate MoE Interventions? Auditing Labels, Selection, and Decision Impact
Abstract
Behavior traces can validate an MoE cache policy when the policy preserves the ordered execution and replay matches its initial state and accounting. Interventions that change expert outputs require evidence about the execution that follows. We study these requirements in Qwen3-MoE, DeepSeek-V2-Lite, and GPT-OSS-20B, separating the accuracy of a candidate's value from the validity of selecting it. Matching execution order and initial state repairs the tested fixed-route labels. Replay and execution still select different windows in 9/40 Qwen records with FP32 routing: modified expert outputs alter the inputs to later routers. Cache history adds a delayed effect, allowing costs to differ even after routes agree. Yet label errors often leave decisions intact. In a separate two-family audit, route-local correction leaves 74/96 labels inexact, yet only 3/32 choices have positive regret, all within 1% of native loads. We construct two runtimes with the same checkpoint and observed behavior but opposite intervention rankings when action-path execution semantics are unobserved. The Trace-to-Action Audit Protocol (TAP) checks the selected action's value, its alternatives, and the loss from choosing it. These checks distinguish a wrong cost estimate from a suboptimal choice and from a loss that exceeds the evaluator's budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.