Measure Before They Debate: Auditing LLM Response Composition
Abstract
Response probes measure how an LLM reacts to individual inputs, but using those measurements to predict responses to joint inputs requires a separate composition assumption. We study this assumption through a finite-input audit that holds the interface fixed, excludes combinations from calibration, and separates prediction error from response noise. In a joint study of responses, summed single-sender secants fall below common-input secants by and ; both approximate simultaneous intervals exclude zero. One configuration also answers all nonzero-target arithmetic queries correctly, although the prespecified global screen fails. A prospective test on eight new, separately calibrated propositions replaces numerical cards with prose comments. Across responses, of composition contrasts have approximate simultaneous intervals below zero; five remain unresolved. Addition retains the lowest observed overall held-out error, while averaging measured increments improves mixed-input prediction. Further tests show how calibration support redistributes error and how remeasuring archived replies changes a gain ordering. The audit distinguishes failed composition identities from useful predictions and makes sensitivity to input families and readouts explicit. The evidence covers numerical cards and controlled prose substitutions; it does not establish internal mechanisms or performance in unrestricted discussion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.