acceptodds
Under review as a conference paper at ICLR 2027

Beyond Flat Attribution: Recovering Boolean Evidence Formulas for Black-Box Language Models

Abstract

Post-hoc context attribution for closed LLMs asks which spans of a prompt support a claim the target asserts. Most black-box methods return flat per-span scores, which cannot express that either of two spans may satisfy the same evidence requirement. We formulate attribution as the recovery of a Boolean evidence formula , a monotone AND-of-OR representation for which flat span sets have a clause-averaged F1 ceiling of . LEAF builds a fixed set of candidate formulas from LLM-derived span roles and uses counterfactual queries to select among them. Under a registry-only answering instruction, we evaluate on synthetic Easy/Hard, two long-context extensions, and a curated 500-example fact-verification benchmark from HoVer, FEVER, and SciFact. Under the primary binary-feedback protocol, LEAF achieves Formula F1 of 0.83–0.99. Its margin over the strongest reported baseline ranges from +0.10 on Natural to +0.49 on the 100-span setting. The Hard result extends across four closed targets and Llama-3.3-70B. Held-out behaviour and weak-cartographer tests probe the roles of the metric and LLM prior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.