acceptodds
Under review as a conference paper at ICLR 2027

Authorization and Endorsement Forgery in Multi-Agent Evaluation: Propagation and Defense

Abstract

Recent LLM-based evaluation systems increasingly rely on multi-agent collaboration, where judges exchange assessments and revise decisions before final aggregation to improve evaluation quality. However, such collaboration introduces an evidence-authorization vulnerability: unsupported authority claims can acquire apparent authority when relayed by internal agents, leading downstream judges to treat them as credible peer evidence. We characterize this vulnerability as Authorization and Endorsement Forgery (AEF), an IPI-derived attack pattern that manipulates how evaluation evidence is interpreted and trusted. Unlike conventional instruction-oriented IPI payloads, AEF injects task-relevant but fabricated or unsupported claims of expert endorsement, prior adjudication, authorization, or consensus rather than directly steering the evaluator through injected instructions. Across three LLM evaluation benchmarks, AEF propagates through multi-agent evaluation pipelines and increases downstream target-selection rates by up to 28.03 percentage points, demonstrating its ability to manipulate evidence trust beyond direct exposure. To mitigate this vulnerability, we propose TRUSTGUARD, an orchestration-layer defense that separates evidence registration from authorization. TRUSTGUARD prevents unsupported authority claims from gaining decision influence by tracking source and derived evidence and permitting downstream use only when provenance and policy-defined authorization requirements are satisfied. Evaluation results show that TRUSTGUARD reduces macro-average ASR-any from 61.03% to 34.41% while maintaining clean accuracy comparable to the undefended system.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.