acceptodds
Under review as a conference paper at ICLR 2027

PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency

Abstract

Large language model (LLM)-based persona agents are increasingly used as scalable proxies for human participants across diverse domains. A prerequisite for their reliable use is that the factual and biographical content they assert remains consistent throughout an interaction. Existing evaluations, however, often rely on independent or loosely connected questions and do not jointly examine complementary forms of consistency within a single interaction. We propose **PICon**, a black-box evaluation framework that probes persona agents through logically chained multi-turn questioning and evaluates the consistency of their asserted content along three complementary dimensions: **internal consistency**, which captures contradictions across accumulated responses; **external consistency**, which assesses compatibility with externally verifiable real-world constraints; and **retest consistency**, which measures stability under repeated questioning. We evaluate eight groups of persona agents alongside 63 human participants under the same protocol. Our results show that no evaluated persona-agent group is as balanced as the human reference across all three dimensions, with distinct failure modes including contradictions across accumulated responses, limited or externally refuted factual claims, and instability under repeated questioning. These findings show that persona consistency depends not only on how responses are scored, but also on how commitments are elicited, motivating sustained, interaction-level evaluation before persona agents are used as human proxies. Code is available at https://anonymous.4open.science/r/picon-8745

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.