acceptodds
Under review as a conference paper at ICLR 2027

Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative

Abstract

Diffusion models can predict a follow-up medical scan from a baseline, and a clinician reading such a prediction needs a per-voxel map of where it should be trusted. Many of these maps approximate the diagonal of the Tweedie posterior covariance, and they are evaluated by comparing one approximation against another, so it remains unclear whether the estimator or the target is the limiting factor. We therefore compute the exact diagonal on six checkpoints across fourteen model–corpus conditions. A Hutchinson estimator at M=200 tracks it at rank agreement of at least 0.92 in every condition. However, in four of the fourteen the exact diagonal is anti-correlated with the denoising error, reaching −0.13, so a faithful estimator reproduces that reversal. All four are real-image conditions, and in these checkpoints the reversal does not appear on the models' own samples, so evaluation on model-generated samples gives a favourable reading of this family. From this one may conclude that the limiting factor is the target rather than the estimator. We then introduce Tweedie Probe-Tangent (T-PT), a gradient-free residual probe that corrupts one model-supported prediction repeatedly and measures the voxel-wise variance of the denoiser's response. T-PT reads a different functional of the same Jacobian, and its exact second-order form ranks identically with the diagonal wherever the diagonal reverses. At the working budget of thirty probes, however, T-PT returns a map too unstable to reproduce that ranking, whilst a Hutchinson estimator at M=5 is already stable enough to reproduce it; T-PT at that budget is therefore too unstable to be read as evidence against the reversal. Where the diagonal is not reversed, T-PT tracks the diagonal closely in most conditions, at a small fraction of the cost of the exact computation. We therefore offer T-PT as an instrument rather than as a better approximation, and we evaluate it on clinical follow-up prediction from a single reverse chain, on brain magnetic resonance imaging (MRI) and on lung computed tomography (CT). On brain MRI at full resolution, where every Jacobian-based estimator we test exhausts memory, T-PT leads the twenty-chain Monte-Carlo ensemble on five of the eight endpoints measured inside tissue and trails it on none, at 16× fewer network evaluations. Over the whole volume the ensemble leads instead, on a comparison carried by the background that fills three quarters of it, and a fifty-chain ensemble closes the tissue gap on a thirty-volume subset. On lung CT the ensemble is ahead on every endpoint, whilst a Hutchinson estimate of that diagonal ranks the error well below T-PT. Both estimators lose most of their discrimination in the regions of longitudinal change, which remains an open problem.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.