The Instrument Is Not the Model: Measuring How Much of an LLM Hiring Disparity Comes from Unreported Design Choices
Abstract
Audits of large language models for hiring discrimination carry legal weight: New York City's Local Law 144 makes it unlawful to use an automated employment decision tool that has not been audited for bias. Such an audit reports a demographic effect, meaning how differently a model treats two resumes differing only in the name, and whatever the two share should cancel. It does not, by enough to change what an audit concludes. Holding the model fixed and varying only choices published audits do not report, over 31,468 matched pairs from 64,397 calls on 6 open-weight and 5 frontier checkpoints, the instruction wording moves the effect by 25% to 64% of itself across the 6 cells where the effect is separable from zero, including under edits that change no word; where the sign flips, the resume template outweighs the wording. Only 25% to 33% of name pairs from a standard validated list are token-matched. Choices made after the data exists matter as much. Resampling rows rather than name pairs narrows intervals 3.0 to 4.8x, and a fixed operating point misstates the percentage-point conversion by 2.0x to 484x. The measurement itself reproduces bitwise only once request batching and cache residency are controlled. Of 13 LLM hiring audits read in full text, none reports enough to reconstruct its number. We give the minimum an audit must report, a screening rule with a calibrated false-positive rate, and the full pipeline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.