When Stronger Baselines Mislead: Admissibility Auditing for Representation-Level Claims
Abstract
**A baseline is evidence against an alternative explanation only if it is *invariant* to the claimed property**; so a stronger baseline need not be a stronger control. On AI Feynman, a zero-parameter bag of variable tokens scores against a Tree-LSTM's , so that encoder's score is no evidence about invariance to variable identity. **A baseline comparison estimates *relative performance*; *admissibility auditing* estimates the *ceiling* of a *declared* family : the best score *any* control strong on the task yet blind to the claimed property can reach, *trained* readouts included, at three cue levels S1–S3.** That direction of comparison is forced — ; *which* family we declare is the choice. A pass rules out the declared alternatives, and *no* admissible family can turn it into a certificate: **the audit falsifies by design**. The levels name *cues*, not algebra; the audit ports to language and to code, where *our own* inversion fails to replicate. Against a shape-matched twin: same tokens and operators, different *arrangement*; **every arrangement-invariant representation, declared or not, is pinned at exactly by proof**, and a Tree-LSTM reaches ; on unseen classes it scores against for the strongest order-blind control an adversarial search finds. **Four of the audited results are other people's**; ten in all, eight of them already published: a released 80M-parameter integration encoder scores where a zero-parameter operator/arity bag scores ; (an S2 task correctly identified, not a defect), while a 14-corpus leaderboard is *upheld*. So "strongest baseline" is not an evidential principle: the ceiling is *non-monotone* in resolution, and widening the family can only cost a pass. **What survives every control is narrower than "compositional generalisation".** A Tree-LSTM trained on *single* rewrites identifies held-out *compositions* at depth 8 on `poly8` () and clears a composition split designed by others, so it *passes the audit*; as *composition-of-known-transformations* generalization, scoped to here. That is not compositional reasoning: four primitives are not a library. **The audit is the deliverable; the findings are its test.**
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.