acceptodds
Under review as a conference paper at ICLR 2027

You Got the Right Answer for the Wrong Reason: Process Fairness in Agentic AI via Structured Deliberation

Abstract

Tool-using language model agents can treat matched applicants differently in their decisions, explanations, and evidence-gathering trajectories. We formalize these channels and introduce PRISM (Pre-Registered Inference with Stereotype Masked Scaffolding), a black-box protocol that fixes a case-independent rubric, threshold, fallback policy, and required tools before demographic exposure, then audits case-level adherence. On the evaluated DiscrimEval-derived and Legal/GRC audits, PRISM reduces paired outcome divergence by 52–62% and 82–84%, respectively, replicates across four backbones including a 30B agent-first model, and lowers reasoning-trace divergence on both open-world profiles. Matched controls suggest that committing to criteria before identity exposure, rather than who writes them, carries the effect. With executed tools, PRISM also raises required-tool coverage from 76% to 99.8% and removes tool-set mismatch between matched applicants, and an independent re-run reproduces 99% of its decisions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.