Hidden Misalignment in Low-Field MRI Benchmark References: The Leaderboard Depends on How the Reference Is Built
Abstract
Deep-learning enhancement is the most promising route to clinical-quality low-field MRI. Its progress is measured on M4Raw, the only public low-field (0.3 T) raw k-space benchmark. There, every method is scored against the average of repeated acquisitions. We show that this reference is systematically misaligned, and that the leaderboard built on it depends on how the reference is constructed. An audit of all 240 subject-contrast groups finds sub-pixel displacements between repetitions. Registering them lowers the reference's own noise level by 6-14% (validation) and 15-25% (test). The displacement lies along the readout axis, has the same sign in every validation group and 99% of test groups, and grows linearly with acquisition order (about 2.3 ppm per repetition). This is the signature of an uncompensated frequency drift of the permanent magnet, with head motion on top in the labelled motion subset. We build a corrected reference by exact k-space phase-ramp registration and phase-aligned complex averaging, with a per-volume error budget. It is validated by synthetic-truth recovery and leave-one-repetition-out prediction. We also measure the reference's own resolution. Two disjoint halves of the published reference agree only to 28.8-31.7 dB, with a per-subject spread of 1.0-1.7 dB. Registration raises the agreement by 1.1-1.8 dB and shrinks the spread two- to three-fold. We re-score the publicly released baselines under a decision rule frozen before unblinding. For every network-vs-BM3D pair the reference x method interaction is Holm-significant and same-signed under every reference construction we test. Registering the reference under the published protocol, and changing nothing else, removes about half of NAFNet's published lead over BM3D (+0.94/+0.87/+0.37 -> +0.50/+0.49/+0.20 dB). Whether the ordering then inverts depends on the reference construction (one, four or five pairs under the three constructions we test) and on the metric and region. The published reference cannot order these methods. We release the corrected reference (CC BY), the error budgets and the code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.