acceptodds
Under review as a conference paper at ICLR 2027

Downstream Task Scores in Super-Resolution Measure the Pipeline, Not the Pixels Alone

Abstract

Super-resolution (SR) is increasingly evaluated by downstream task scores, but these scores depend on the reconstructed images, the downstream model, and its decision rule. We audit how decision rules affect the performance and rankings attributed to SR. In remote-sensing segmentation, we hold reconstruction pixels, model weights, and continuous score maps fixed while comparing native decisions with a global threshold selected on validation data. Several SR pipelines score zero under the native rule even though their score maps retain discrimination above chance, but below that on matched high-resolution (HR) inputs. With HR-trained models, threshold selection nearly matches the best performance among the tested global thresholds, yet leaves a substantial gap to HR under matched conditions. Threshold gains persist when downstream models are fine-tuned on SR images. Large utility gains need not change method rankings, while changing the decision rule can alter which method ranks first. Controlled photography and microscopy analyses show that decision policies and target definitions can reverse comparisons, and audits of published downstream benchmarks show the same decision sensitivity, while the gap to matched references remains larger than the threshold gain from the default operating point. These findings motivate four reporting requirements: the decision budget, the matched reference, the target and cohort definition, and the claim scope. Together, they make clear what a downstream score measures and which comparisons it supports.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.