acceptodds
Under review as a conference paper at ICLR 2027

Neighbors as Top Suspects: Diagnosing Ranking Shortfalls in LLM Adversary Detection

Abstract

LLM agents increasingly make decisions from other agents' claims. Even when evaluators observe the received messages, a scalar outcome does not reveal where the LLM's final ranking differs from a reference derived from that stream. We record the messages delivered to one LLM seat in a five-round Mafia game with fixed-policy surroundings and derive an offline accusation-minus-defense shadow ranking. We separate non-neighbor placement (the full-ranking slots occupied by players not heard directly) from ordering within those slots. Across four models and eleven configurations, model mean reciprocal rank (MRR) is below the shadow in 34 of 44 cells. On complete raw rankings with a non-neighbor adversary, placement's signed contribution accounts for 76–100% of the mean shadow-minus-model reciprocal-rank gap and remains the larger contribution in connected configurations. A separate delivered-accusation analysis shows that the adversary can rank highly among non-neighbors yet below neighbors overall. On the same 224 out-of-view GPT-5.6 games across seven relay configurations, with complete raw rankings in every arm, premise correction accompanies higher model MRR and improved mean within-group adversary rank. The ordering contribution decreases without a detectable placement-contribution change; structure disclosure shows no detectable change in either contribution. Corrected-premise profiles vary across models, with unequal complete-output coverage. These post-hoc, reference-relative diagnostics describe one scripted environment without identifying mechanisms, demonstrating repair, or establishing transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.