acceptodds
Under review as a conference paper at ICLR 2027

The Conjunctive Fraction of Embodied Benchmarks: Composition Cannot Be Measured Where the Write Path Has Already Destroyed It

Abstract

The Capability Convergence Hypothesis (CCH) holds that capability is purchased by access structure: the joint possession of a compressive bounded-state channel and a scalable verbatim-index channel. It proves composition strictly super-additive on a witness family that obstructs each channel separately, and names as its largest open risk a measurement it reports "has not been run by anyone, including us": how often natural workloads instantiate that joint load. We run it for embodied multimodal streams. An arithmetic criterion, computable before any run, first decides whether a benchmark can exhibit the effect at all: a bounded score cannot host an interaction whose components already sum past the bound. A census of public video question-answering corpora then puts the conjunctive fraction in the low single digits, so solving that subset perfectly would move the richest benchmark by less than the run-to-run floor at which these benchmarks are compared. We supply the witness the hypothesis implies (735 judge-free items over head-mounted kitchen streams), and in the registered grid every prediction fails and the channel dichotomy reverses. A diagnostic locates the failure upstream of every arm, in the shared single-frame write path, which decides the locator verb only 11-23 the ledger from annotations instead, composition by address beats a same-cost different-occurrence control by 14-26$ points on every backbone and is null on both single-axis families (the corpus is a valid witness and the loss is on the write), yet even then it answers less than half.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.