acceptodds
Under review as a conference paper at ICLR 2027

A Benchmark with Known Causes: Can Language Models Explain ML and RL Decisions?

Abstract

The quality of an explanation for an ML model or RL policy in a given situation is typically judged by readability or by alignment with another explainer. A more insightful test is whether the explanation identifies the cause of the action that the policy proposed. For such a test, we need decisions for which the cause is known. We build a benchmark of fifty such decisions: fifteen from tabular classifiers, where each feature’s contribution is exact, and thirty-five from three reinforcement learning agents (value iteration on a grid world, tabular Q-learning on cart-pole and an upper-confidence-bound contextual bandit), where we keep a case only if removing one candidate cause and re-running the agent changes the decision. Each case carries four probe situations answered by re-running the system, so the answer key is exact. A program scores faithfulness against the known cause with no language model involved, and a weaker reader from a different model family scores comprehension by predicting what the system does in situations it has not seen. The audience named in the prompt decides whether the cause survives: a plain-language request scores 0.327 on faithfulness against 0.668 for a technical one, one added instruction to name the deciding factor recovers half that gap for five points of reading ease, and across ten open-weight models from six families (1 − 42 gigabytes) the plain request never exceeds 0.42. In the model’s own next-word distribution, naming the audience alone lowers the probability of the true cause by 0.137 (p = 0.0005). A model judges how plainly a passage reads, agreeing with the formula on 97% of texts, however, it cannot count its words. Counting constraints therefore belong in code, while judgements of style can stay in the prompt. On the benchmark, a four-stage pipeline that fixes the cause before simplifying the language matches a single well-specified prompt on mean faithfulness (0.696 against 0.700); what it adds is that it checks the rules in code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.