acceptodds
Under review as a conference paper at ICLR 2027

ROBUST AI-AGENT REPUTATION VIA DELAYED- VALIDATION ONLINE LEARNING

Abstract

An audit can miss a deployment failure; a user report can reveal it, but can also be manipulated. We study this tension as online decision-making for nonstationary entities under strategic multi-source evidence and delayed, correctable outcomes.AI-agent reputation is one instance: the learner estimates an instance or version’s current task reliability before outcomes mature. We propose FARR-SC, which separates bounded use of a report from credit earned by its source. Current reports inform the estimate; matured validation and exact correction update persistent source credit, specialist losses, and calibration. For its fixed-anchor capped-evidence problem, we prove matching strategic-mass bounds with Θ(Σt εt) dependence, alongside implementation-aligned fixed-specialist and realized-path calibration guarantees. With one frozen head, causal replays favor FARR-SC over a descriptive envelope of predeclared non-FARR-SC baselines on four of six traces in each of two blindspot regimes under adaptive attack; these envelope comparisons are descriptive, and conservative baselines win all six clean cells. On two public trust networks, the default has lower decision loss than every non-FARR-SC baseline, although its no-source ablation is lower still. Source scores learned without fraud-label supervision achieve fraud-source AUC 0.91/0.85. Together, these results identify where validated feedback aids recovery and where conservative operation remains preferable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.