acceptodds
Under review as a conference paper at ICLR 2027

From Automation to Role-Play: A Population-Level Evaluation of Jev

Abstract

Can a model designed for fast automation support the role-conditioned decisions needed in a simulated population? We investigate this question in Jev, a newly released System One model that returns structured decisions rather than generated text. Our evaluation proceeds from individual capabilities to population-level behavior, using controlled testbeds to make black-box responses interpretable. PersonaEval and synthetic policy tasks first test role identification and action selection against source-derived or program-computed answers. We then adapt FlockVote into a fixed-population experiment that crosses role representation with shared information. This design distinguishes a change in an initial answer from a change in how the same roles respond to new information. Jev achieves approximately 97% modal-action accuracy on matched policy tasks. Yet, in one historical voting window, converting direct role queries into member-addressed requests changes the aggregate information response by 11.1 percentage points on matched complete records. The corresponding vector interaction remains above a specified tolerance under conservative missing-response bounds. Explicit role instructions and native object inputs leave residual representation effects. Member-level analysis further shows how aligned changes survive aggregation, while cancellation can conceal individual sensitivity. These findings identify a gap between individual decision competence and population-level role-play. They establish concrete validation requirements for using Jev as a shared decision backend before introducing memory, communication, and feedback; the reported probabilities are model responses, not validated human behavioral propensities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.