Readout Fidelity Bottlenecks Self-Verification in Unified Multimodal Models
Abstract
Unified multimodal models combine image generation and visual understanding, allowing them to evaluate and select among their own generations without an external verifier. Yet self-verified Best-of-N recovers only a small fraction of the quality available in candidate pools. We identify readout fidelity as an independent bottleneck: useful candidate preferences can already exist in verifier evidence but are compressed by discrete Yes/No decoding. Across seven models, discrete readout produces extensive score collisions and can even underperform random selection despite substantial oracle headroom. We model discrete verification as finite observations of an underlying Yes/No preference, with a formulation agnostic to how the verifier evidence is produced. In the direct-answer setting, repeated stochastic observations converge to the pre-decoding logit margin, while full stochastic decoding closely tracks the predicted finite-observation behavior. We use Continuous Best-of-N (C-BoN) as a controlled intervention that directly exposes this limiting preference while keeping the remaining selection pipeline fixed. Across seven models and four benchmarks, C-BoN improves all 28 settings without additional models, training, generation, or verifier calls, and enables stronger Best-of-N scaling. Larger continuous margins also yield more reliable ordering within discrete ties, while recovered ordering tracks downstream gains. These results establish readout fidelity as an information bottleneck in self-verification and show that verifier evidence can contain richer preferences than discrete verdicts reveal.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.