acceptodds
Under review as a conference paper at ICLR 2027

Probing Task-Level Agent Reliability Through Rollout and Replay

Abstract

An agent's execution record can be available before an external evaluator returns a task outcome. We ask whether the acting model's own unlabeled reasoning-and-action traces reveal how reliably it handles a fixed task. Our rollout–replay protocol collects repeated natural executions, replays their recorded histories through the original model, and reads intrinsic statistics at action and post-observation events. Discovery batch A fixes the measurement rules before evaluation on new execution seeds in batch B. Across benchmarks, event-aligned readouts rank task reliability more accurately than matched whole-trajectory statistics. Level separates groups whose sampled executions consistently succeed or fail, while Dispersion distinguishes higher- from lower-success groups with mixed outcomes. Scores from A also rank outcomes from independent B executions; comparator gains persist before explicit outcome disclosure and on held-out MCPMark. Cross-model replay shows that reader capability alone does not replace compatibility with the acting model. These results support task-level reliability measurement from unlabeled interaction traces. They do not establish calibrated success probabilities or a benefit from using the scores to change agent behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.