Predictive Trajectory World Models for Process Risk in LLM Agents
Abstract
An LLM agent can reach a correct outcome while following an unreliable execution process. Outcome-level evaluation can therefore miss process failures, while fixed anomaly taxonomies provide task- and dataset-dependent definitions of abnormal behavior. Instead of defining abnormal agent behavior with a fixed taxonomy, we model normal execution and measure how a candidate action is predicted to depart from it. Given a completed trajectory prefix and a candidate next action, we study prospective process risk by predicting before execution whether that action will move the process outside its normal support. Our analysis finds cross-model consistency for normal executions of the same exact task but limited evidence that a whole task category forms one compact process cluster. This motivates task-conditioned local normal support. This also separates executor-specific variation from task information, which remains necessary to define what constitutes normal execution. We align same-task executions across executor models and use retrieved normal context in a trajectory world model that predicts the next latent process state under the candidate action. The world model learns normal process dynamics rather than directly predicting anomaly labels, leaving anomaly supervision to the final onset-aware risk mapping. Predicted future-state and transition deviations then inform an onset-aware risk score focused on the first abnormal step. Experiments support each part of this chain. Geometry analysis shows clear same-task cross-model structure, but does not support a single compact center for a whole task category. Retrieval-conditioned prediction achieves the lowest mean next-state error across cross-model, held-exact-task, and leave-one-subset-out settings. At the 10% development normal-trajectory FAR budget, the final predictive-risk model reaches 56.8% first-alarm exact-onset hit, improving over the matched direct-only baseline by 17.3 percentage points at similar observed FAR. These results show that process risk can be grounded in task-conditioned normal structure and predicted future deviation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.