Evidence Response Is Not Decision Retention: An Audit of an Olfactory LLM Interface
Abstract
We present a decision-level audit of predictor-backed language interfaces, separating behavioral response, assignment advantage, retention, and added value. We apply it to a fixed olfactory predictor, using released odor annotations and, separately, human applicability ratings. On 101 floral annotation pairs, two complete Qwen3-4B interfaces both make 158 correct choices across 202 presentations, yet retain 67/78 and 78/78 rule-correct pairs in both orders. A fixed vote across four label–position encodings recovers all 78 for the first interface without correcting any rule error relative to the annotations. Voting is not uniformly beneficial: on odorless pairs, it makes 18 errors among 63 accepted pairs, compared with 10 for numerical gap selection at the same coverage. In the separate human-rating task, Phi-4-mini and Qwen3-4B reproduce the score rule on every clear floral presentation under revised wording. One 32B configuration correctly reads all 160 supplied reference-rating presentations, but shows no observed matched-coverage risk advantage over numerical selectors on clear pairs when given predictions. The audit distinguishes recovery of correct decisions from improvement beyond the visible-score rule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.