ProtocolBench-Collab: Renaming a Role Changes What a Collaborative Language Model Answers
Abstract
Collaborative systems built from language models, such as actor–critic pipelines, are evaluated through the role interface their designers wrote, and that interface is rarely varied. We show that their answers depend on it. PROTOCOLBENCH-COLLAB holds each question fixed and changes only the role interface, across nine renderings that range from a neutral change of role to a collaborator who prefers a named wrong option. All eight untrained models we evaluate, from four families, change their answers with the interface, and the strongest remains inconsistent on a fifth of items. We then train a shared actor–critic adapter on preference pairs each presented under four interfaces. This narrows the accuracy spread across interfaces from ten points to one and a half and raises the share of items answered identically under all nine from 65% to 73%. Presenting each pair four times also quadruples the training steps, and a control with the same number of steps and a single interface does not improve agreement (+0.6, t=1.23), which identifies interface variety as the source of the gain in agreement. Enlarging the adapter improves robustness further, and the two effects combine.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.