RealEstateBench: Measuring Fiduciary Performance of AI Agents
Abstract
Task-success benchmarks ask whether an agent can complete a task; safety benchmarks ask whether it will refuse harmful requests. Neither alone establishes whether an agent honors professional obligations while pursuing its objective. To address this gap, we define fiduciary performance: how well an agent achieves the outcome it was retained for, subject to the duties of the role. We operationalize this in RealEstateBench, a long-running simulated real estate market where seller agents represent homeowner clients, market properties, and negotiate sales through structured actions and open-ended dialogue. To make conduct measurable, we derive rubrics from statute and case law, assess behavior against simulator ground truth, and validate the examiner on 300 seller–buyer trajectories restructured from court cases and verified against source facts and holdings. Across eight models and 40 sessions, compliance and commission are nearly uncorrelated (), yet five model means are dominated, showing that much of the observed noncompliance yields no economic benefit. The three nondominated models span the observed frontier from higher compliance to higher commission, revealing a remaining tradeoff between economic performance and professional conduct; the highest earner makes 55% more commission while scoring 17 percentage points lower on compliance. Models with similar earnings also differ in aggregate compliance and conduct profiles, showing that similar economic outcomes can arise from qualitatively different behavior. Together, these results show that fiduciary performance separates avoidable noncompliance from the tradeoffs that remain on the observed frontier, providing a concrete target for improving both economic performance and professional conduct.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.