acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Multimodal Inverse Problems: A Protocol for Per-Event and Population-Level Fidelity

Abstract

Evaluation in inverse problems is dominated by pointwise metrics - RMSE, MAE, per-event resolution - under the implicit assumption that lower error means better reconstruction. That the conditional mean poorly represents multimodal structure is well known but existing evaluation protocols do not score its consequences for the reported marginal. When one inverse problem is instantiated over many independent events and the scientific deliverable is the marginal over the dataset, as in unfolding, a second failure mode appears that pointwise metrics cannot capture. By the law of total variance, point estimators trained to minimize MSE produce a marginal spectrum strictly narrower than the truth whenever the posterior has nonzero width. The resulting bias is independent of architecture, training, and dataset size, and it compresses precisely the spectral features like shapes, tails and modes that scientific measurements rely on. We propose a three-part evaluation protocol where each step targets a failure mode the others miss: per-event distributional accuracy via CRPS, population-level marginal accuracy via a spectrum-fidelity diagnostic, and uncertainty trustworthiness via coverage-based calibration. We instantiate the protocol on a synthetic benchmark with an analytic posterior to verify the mechanism, and on a realistic many-to-one inverse problem from particle physics we show that model rankings reverse between pointwise and distributional metrics, and calibration further separates architectures indistinguishable under CRPS. The choice of evaluation protocol can therefore determine the scientific conclusion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.