acceptodds
Under review as a conference paper at ICLR 2027

CrossRefEval: Paired Evaluation of Genomic Models Across Reference Assemblies

Abstract

Genomic models are usually evaluated on a single reference genome, leaving open how the same predictor responds when corresponding regions are represented by another assembly. This comparison requires explicit correspondence: for position-resolved prediction, corresponding bases can occur at different output indices, so the rule used to pair predictions can itself affect the measured comparison. We introduce CrossRefEval, a paired framework that holds the fitted predictor and decision rule fixed across reference representations and evaluates task performance, cross-reference consistency, and paired correctness at corresponding outputs. We instantiate CrossRefEval on the maternal haplotype assembly of T2T-YAO v2.0 and GRCh38, constructing 254,366 paired windows across five genomic prediction tasks. We evaluate five frozen DNA encoders with readouts trained separately on YAO and GRCh38. Holding prediction arrays fixed, changing only the output correspondence alters the comparison between readout variants: across ten encoder–training-reference combinations, the absolute correspondence-induced shift exceeds the magnitude of the readout difference measured under mapped correspondence in nine cases and reverses the direction of the three-seed mean comparison in six. Better task performance can also accompany greater cross-reference disagreement, while increased agreement can include shared errors. These results show that per-reference task scores alone do not characterize how predictions change across reference representations. CrossRefEval makes reference representation and output correspondence explicit parts of genomic model evaluation rather than fixed background assumptions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.