Risk Before the Action: What Pre-Action Hidden States Predict in LLM Agents
Abstract
An LLM agent may issue a tool call that returns an error. It may also fail its task without any tool error. We study what hidden states predict before the agent generates a response. Our main target is tool rejection, meaning an error reply from the environment. We fit linear probes on frozen hidden states from two Qwen3 model sizes in the retail and airline domains of \bench. Task-grouped cross-validation shows gains over the tested text baselines in rejection ranking across all four settings. On unseen tasks, frozen B probes reach AUROC in retail and in airline. These scores compare calls within the same task and call position. Repeated sampling from fixed conversations and environment states shows little variation in first-call rejection, yet probe probabilities remain imperfect. We next fit separate B probes for rejection and final task failure using the same pre-action layer and prefixes. In cross-validation on training tasks, rejection ranking is stronger on shared task and call-position groups. Retail AUROC is for rejection and for final failure. We also show when a fixed predictor using only binary calling and rejection records cannot give accurate task-success probabilities across populations. Finally, the tested verification policy shows no detectable advantage over random, observable, or timing-matched triggers. Local risk prediction, task-success estimation, and intervention value therefore require separate evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.