Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
Abstract
One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth to over a blind one on four open-weight backbones served with thinking disabled, and on the two Qwen backbones the blind inspector changes no item at all. The inspector helps only when it can see the problem. That head-to-head is exploratory. We manipulate it directly. One deterministic fault enters the first stage, and we re-expose the original problem to of the downstream stages with agents, items, fault and topology held fixed, on 120 gsm_hard items per arm at temperature zero. Accuracy under fault rises on four of four backbones, from to , the largest Holm-adjusted being . A registered kill test rules out tokens. Blanking every word holds the word slots fixed, and retention tracks the visible fraction on four of four, climbing from 0.221 to 0.692 on the primary. The prompt-token ratio between the ends runs 0.90 to 1.03, above one only on Phi-4, whose masked arm still scores worse. Those two families are confirmatory and everything else here is exploratory. The interaction excludes zero on two of four backbones under the registered pipeline, four of four under a three-stage pipeline, and three of four under full message history, which passes prior stage outputs, the primary at . On Llama-3.1-8B the fault carries no detectable cost at any dose, so the other three carry every claim about what a fault costs. Re-derivability also sets what the architecture costs, and no decomposition we measured reliably beats one direct call. With no fault injected the registered pipeline loses to that call by , and , and on Phi-4 reads at , which the test fails to separate from zero. The repair that works is cheap and front-loaded: the first re-grounded stage buys of matched retention for tokens per item on Qwen3-14B, and the stages after it buy nothing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.