Reference-Dependent Scores, Margin-Stable Decisions: Calibrating MRI Reconstruction Benchmarks
Abstract
Reference-based MRI benchmarks conflate reconstruction quality with the statistical compatibility between an input and its evaluation target. We introduce a calibration framework for this dependence in repeated-acquisition MRI. When a reference contains the same measured repetition used to form an accelerated input, realization-specific reconstruction error can correlate with reference noise and shift method scores even when repeat count and averaging are matched. We separate this method-level score distortion from the downstream decision instability it may induce. A margin-normalized certificate, ρ, makes the distinction explicit: the disjoint-reference winner is guaranteed to remain the winner when ρ < 1, whereas near-tied candidates are vulnerable to much smaller differential score shifts. On M4Raw 0.3T data, an all-slice T1/T2 study over 540 subject–contrast–slice units yields positive independent-validator loss for shared-reference selection (0.00916 NRMSE on V1 and 0.01389 on V2), and the direction transfers to 270 FLAIR units. Adding the official PI-VarNet nearly removes selection harm because the learned model creates a large candidate margin, even though its shared–disjoint score shift is the largest in the pool. A second competitive learned candidate restores 70/540 per-unit selector changes while independent-validator loss remains small. Within the classical pool, a runner-up margin approximation explains 205 of 220 observed selector changes with no false positives. These results motivate a practical reporting protocol for repeated-acquisition benchmarks: pair reference provenance and score sensitivity with candidate-margin-based decision stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.