An Estimand-Based Evaluation Framework for Compositional Reasoning Pipelines
Abstract
Stage-level scores for compositional reasoning depend on which examples are scored, what upstream information is available, and which chance reference is used for comparison. These are distinct choices: scoring a stage only on examples whose upstream stages were answered correctly selects a model-specific population and changes the reference against which its score should be read, whereas supplying ground-truth upstream answers changes what the model receives while retaining all applicable examples. We present an estimand-based framework that makes these choices explicit. It distinguishes natural, prefix-conditional, and oracle-conditioned accuracy, pairing each with a chance reference matched to both its population and its upstream information. It additionally measures a signed selection gap by scoring the same oracle predictions on different populations. In a controlled shape-grammar study with five multimodal models, final-stage conditional accuracy reaches 98.45–99.67% but covers only 6.2–15.1% of applicable examples and ranks models differently from full-population accuracy. Holding oracle predictions fixed for two models, restricting evaluation to natural survivors raises accuracy by 7.15 and 12.97 percentage points at the rule stage and by 7.07 and 11.30 points at the final stage. Final-stage survivors reach 93.37–96.49% accuracy under fixed oracle predictions, still below their natural prefix-conditional scores. Furthermore, survivor-matched chance references deviate from full-cohort baselines. In a reanalysis of published results, accounting for coverage changes the preferred retrieval context size in five of eight comparisons. These results suggest the necessity of reporting population, input regime, coverage, and a matched chance reference alongside every stage score.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.