acceptodds
Under review as a conference paper at ICLR 2027

Calibrate Before You Measure: Conditioning Interfaces Decide What a Generative Probe Reports

Abstract

Frozen diffusion models serve neural decoding as renderers and, through conditional likelihood, as instruments that measure how much a brain recording reveals about a stimulus. Both uses route the signal through a conditioning interface that is treated as a neutral pipe. We show that it is not. We construct a chain of stimulus channels ordered by the data-processing inequality, so that any valid instrument must report a non-increasing curve, and use this ground truth to audit four conditioning interfaces of one diffusion backbone. The single-token image-embedding interface adopted by most EEG and fMRI pipelines scores the true stimulus below a deliberately mismatched condition on spatial layout, whereas a 257-token readout of the same encoder pass reads layout correctly. The gap persists across a second backbone and two disjoint stimulus sets, and a far more bottlenecked caption interface reads layout correctly, so the failure cannot be anticipated from an interface's description and must be measured. A dose–response identifies conditioning strength as the mechanism: raising it improves the axes an interface can express while pushing the axes it cannot express below chance. Deployed decoders exhibit the same over-pushing: their decoded vector is a worse estimate of the true embedding than a constant in all ten subjects and is optimal at one third of its deployed strength, and two independent decoders are overconfident by a conformally guaranteed margin that rendering at the optimal strength narrows but does not close. We release the phantom as a pre-flight check for any conditioned generative probe.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.