Annotate or Filter? The Redundancy–Rescue Mechanism and Phase Boundaries in LLM Fusion Pipelines
Abstract
Multi-component pipelines such as Archon place a verifier in front of a fuser, assuming that discarding incorrect candidates can only help synthesis. We show that this selection-before-synthesis prior is mechanistically inverted. On MATH-500, five 8B process reward models all degrade fusion: verifier-filtered pipelines reach 0.818–0.866 accuracy versus 0.888 for unfiltered critique-and-fuse, a loss of 2.2–7.0 points. The cause is not verifier incompetence but a redundancy–rescue mechanism. Generator outputs are strongly polarized—most problems are either entirely correct or entirely wrong—so a verifier can act only in the thin mixed tier, where false positives pass errors through while false negatives are absorbed by redundancy (80.2% of problems have at least two correct candidates). The fuser, conversely, is an error corrector: it rescues 41.0% of majority-vote failures and corrupts 1.4%, synthesizing correct answers even from all-wrong sets and thereby breaking the generator ceiling (0.834 0.888). Filtering therefore starves fusion of the incorrect candidates that constitute its corrective signal. We derive a three-tier accuracy decomposition that predicts both the FP/FN asymmetry and a filtering/fusion phase boundary from two observable statistics—candidate redundancy and fuser rescue rate. Strong fusors should receive annotated but unfiltered candidates; weak fusors still benefit from gating. The crossover lies near FP 0.25, well below the operating regime of current 8B PRMs (0.41–0.64). Dual-path cross-validation against a 60-point FP/FN grid and a weak-fuser replication confirm both. The implication is architectural: verifiers should annotate, not filter.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.