Learning to Explain Where Models Disagree
Abstract
Aggregate benchmark performance often does not reveal *where* two classifiers disagree. We study how to turn a small set of labeled agreement and disagreement examples into a concise natural-language description that transfers to new inputs. We introduce **SEAM**, an explanation policy trained with reinforcement learning to maximize pairwise simulatability: the accuracy with which an observer can use an explanation to predict whether the classifiers agree on an unseen input. The reward uses automatically obtained agreement between models and requires no human-written explanations or preference labels. Across 15 structured and text classification tasks held out from training, **SEAM** achieves 0.75 simulatability, compared with 0.69 for GPT-4o and 0.51 for its untrained base model. Its advantage persists for observer models excluded from reward computation. On 4 held-out multimodal tasks, a jointly trained vision-language variant achieves 0.68 simulatability. The advantage measured by the LLM observer also appears for human participants: in a controlled study with 60 participants, **SEAM** explanations support 74.0% agreement-prediction accuracy, compared with 55.4% for GPT-4o explanations. These results show that direct optimization can produce transferable natural-language summaries of pairwise model behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.